Downloads · 30 days
0
a12354/AnomSeer
AnomSeer is a machine learning model from a12354. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<h2 align="center" AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly Detection <br<br <img src="./assets/logo.jpg" width="300" alt="AnomSeer Logo" <br <b@ <a href="https://icml.cc/"ICML 2026</a</…
Downloads · 30 days
0
Access
Public
Updated Jun 19, 2026
Repo size
610 MB
Likes
0
Public
Click a slice to open those files.
.parquet569 MB · 92%
From the Hugging Face model README
AnomSeer is a reinforcement-learning post-training framework for time-series anomaly detection (TSAD) with multimodal LLMs. It unifies anomaly classification, localization, and explanation by grounding MLLM reasoning in fine-grained, classical TSAD evidence through two components:
With Qwen2.5-VL-3B/7B-Instruct as backbone, AnomSeer outperforms GPT-4o and Gemini-2.5-Pro on anomaly classification and localization across AnomLLM, VisualTimeAnomaly, and TSB-UAD benchmarks.
conda create -n anomseer python=3.11 -y
conda activate anomseer
pip3 install -r requirements.txt
pip3 install flash-attn==2.7.4.post1 --no-build-isolation
pip3 install -e .
pip3 install -r requirements_anomseer.txt
pip3 install git+https://github.com/ahstat/affiliation-metrics-py.git
Requires CUDA 12.4 (for
torch 2.6.0andflash-attn).flash-attnmust be installed separately with--no-build-isolation.
AnomSeer is trained on AnomLLM (synthetic) and evaluated on AnomLLM, VisualTimeAnomaly, and TSB-UAD.
Clone AnomLLM, set up its environment, and generate the synthetic dataset:
git clone https://github.com/Rose-STL-Lab/AnomLLM.git
cd AnomLLM
conda env create --file environment.yml && conda activate anomllm && poetry install --no-root
bash synthesize.sh
Download from: https://github.com/mllm-ts/VisualTimeAnomaly
Download from: https://github.com/decisionintelligence/TAB
python multimodal_data_processing/anom.py \
--root_dir /path/to/AnomLLM/data/synthetic \
--out_dir ./data/anol_processed_mllm_data
This produces train_full.parquet (3,200 samples) and test_full.parquet under ./data/anol_processed_mllm_data/.
Edit MODEL_PATH in example/timerpo_trainer/run_anomseer.sh to select the backbone:
MODEL_PATH=Qwen/Qwen2.5-VL-3B-Instruct # 3B variant
# MODEL_PATH=Qwen/Qwen2.5-VL-7B-Instruct # 7B variant
Then launch training (requires 4 GPUs by default):
bash example/timerpo_trainer/run_anomseer.sh
Checkpoints are saved to ./checkpoints/.
To override the data paths without editing the script:
TRAIN_FILE=./data/anol_processed_mllm_data/train_full.parquet \
VAL_FILE=./data/anol_processed_mllm_data/test_full.parquet \
bash example/timerpo_trainer/run_anomseer.sh
Set EVAL=True inside example/timerpo_trainer/run_anomseer.sh and point VAL_FILE to the desired benchmark:
EVAL=True
VAL_FILE=./data/anol_processed_mllm_data/test_full.parquet
Then run:
bash example/timerpo_trainer/run_anomseer.sh
Metrics reported in the logs:
To evaluate on VisualTimeAnomaly or TSB-UAD, preprocess those datasets with multimodal_data_processing/anom.py into the same parquet format and set VAL_FILE accordingly.
@inproceedings{zhang2026anomseer,
title = {AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly Detection},
author = {Zhang, Junru and Feng, Lang and Shi, Haoran and Guo, Xu and Yu, Han and Dong, Yabo and Xu, Duanqing},
booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
year = {2026}
}
We thank the veRL project for foundational RL infrastructure and AnomLLM for the synthetic TSAD benchmark and data generation code.