Downloads · 30 days
0
fengxr93/ts-align-benchmarks
ts-align-benchmarks is a machine learning model from fengxr93. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
TS-Align aligns multivariate time series with a language model for numeric metric question-answering and captioning. A frozen MOMENT time-series encoder feeds a fully fine-tuned Qwen3-4B-Instruct through a per-channel
Downloads · 30 days
0
Access
Public
Updated Jul 23, 2026
Repo size
18 GB
Likes
0
Public
Click a slice to open those files.
.pt16.4 GB · 91%
From the Hugging Face model README
TS-Align aligns multivariate time series with a language model for numeric metric question-answering and captioning. A frozen MOMENT time-series encoder feeds a fully fine-tuned Qwen3-4B-Instruct through a per-channel ("2-tunnel") projection with attention pooling, trained end-to-end with GRPO reinforcement learning against verifiable metric and caption rewards.
This repository is a self-contained benchmark bundle: the trained TS-Align checkpoints, the canonical evaluation data, and a reproduction harness for every baseline we compare against.
| Path | What |
|---|---|
Models/4level_grpo_min2/ | TS-Align checkpoint — metric-QA + captioning |
Models/joint_grpo_min2/ | TS-Align checkpoint — metric-QA + captioning + TSQA (jointly trained) |
Models/<peer>/ | third-party baselines — placeholder only (weights not redistributed) |
Benchmark_eval/ | canonical evaluation inputs + scorers (TS_Caption_test, TSQA, ChatTS_test) |
Benchmarks/ | source datasets and their provenance — see Benchmarks/README.md |
results/ | one reproduction bundle per evaluated model (runner + run.sh + README) |
Both checkpoints share the same architecture — MOMENT-1-base (frozen encoder) +
Qwen3-4B-Instruct-2507 (fully fine-tuned), joined by a 2-tunnel per-channel projection
(1-layer, 8-head channel transformer) with attention pooling gated at
channel_pool_attn_min = 2. Trained with GRPO (2000 steps, PATRA-balanced caption reward,
split-wise Gaussian tolerance σ = {univar 0.5, bivar 0.5, multivar 1.0}).
| checkpoint | training tasks |
|---|---|
4level_grpo_min2 | metric-QA + captioning |
joint_grpo_min2 | metric-QA + captioning + TSQA (joint) |
Each best-checkpoint/ holds:
model-state-0000{1..6}.pt — the trainable weights (Qwen + projection + special-token
embeddings). The MOMENT encoder is frozen and not shipped; load it from its base repo.tokenizer/ — the Qwen tokenizer with 4 added special tokens: <ts> </ts> <scale> </scale>.training.pt — metadata only (≈1.6 KB: shard manifest, special-token ids, trainer
scalars). The optimizer state is intentionally not included.You need the two base models plus the TS-Align inference code:
infer/infer_local.py in the companion ts-align-rl code repositoryload_checkpoint reads training.pt for the shard manifest and special-token ids, then loads
the model-state-*.pt shards onto a freshly built MOMENT + Qwen3-4B model:
python infer/infer_local.py \
--moment-path <path>/MOMENT-1-base \
--llm-path <path>/Qwen3-4B-Instruct-2507 \
--tokenizer-path Models/joint_grpo_min2/best-checkpoint/tokenizer \
--checkpoint-path Models/joint_grpo_min2/best-checkpoint/training.pt \
--channel-pool-attn-min-channels 2 \
--series "10,12,11,13,12,14,13,15" \
--prompt "<ts></ts> What is the mean of this series?"
Models/<peer>/ — ChatTS-14B, GPT-OSS-20B, Qwen2.5-7B-Instruct, Qwen3-8B, Qwen3-14B,
TimeOmni-1-7B, ITFormer-7B, MOMENT-1-large, Mistral-7B-v0.3, Time-MQA-Mistral-7B — each contain
only a PLACE_MODEL_FILE_HERE.txt. These are third-party models we do not redistribute:
download each from its official source and drop the weights in place to reproduce.
Every evaluated model has a bundle under results/<model>/ with a run.sh, an _infer/
harness, and its own README. Inputs are read from Benchmark_eval/; scoring uses a multi-kernel
tolerance metric for numeric QA, precision/recall for captions, and RougeL/accuracy for TSQA.
Predictions and scores are regenerated by the runners and are not checked in.
Caption precision/recall uses an NLP-extraction step: GLM-5.1 (temperature=0, the same un-tuned prompt for all models) reads each caption and extracts the numeric value it predicts for each mentioned metric. The prompt provides GLM with the metric key and an explanation of its statistical property, so it can align free-text mentions to the correct ground-truth metrics; extracted values are then scored against the ground truth under the same tolerance kernel (precision = correct / extracted, recall = ground-truth metrics correctly covered). Extractor: Benchmark_eval/TS_Caption_test/extract_metrics_withdef.py with the key→explanation map metric_defs.json.
Provenance for the third-party benchmarks is documented in
Benchmarks/README.md; TS_Caption_test is our own out-of-sample set.
training.pt.full_bak optimizer states are kept locally only and are not published.This repository's own content — the fine-tuned TS-Align checkpoints
(Models/4level_grpo_min2/, Models/joint_grpo_min2/) and the TS_Caption_test evaluation set
— is not yet released under a license (all rights reserved by the authors for now).
All third-party components keep their original licenses:
| Model | Original license |
|---|---|
| Qwen3-4B-Instruct-2507, Qwen2.5-7B-Instruct, Qwen3-8B, Qwen3-14B | Apache-2.0 |
| MOMENT-1-base, MOMENT-1-large | MIT |
| GPT-OSS-20B | Apache-2.0 |
| ChatTS-14B | Apache-2.0 |
| TimeOmni-1-7B | Apache-2.0 |
| Mistral-7B-v0.3 (and its bnb-4bit quantization) | Apache-2.0 |
| Time-MQA-Mistral-7B (LoRA adapter) | see the Time-MQA project |
| ITFormer-7B | see the model's original repository |
| Dataset | Original license |
|---|---|
| TSQA (Time-MQA) | Apache-2.0 |
| ChatTS_test (ChatTS) | composite, per source component (e.g. NAB: Apache-2.0, Weather: CC-BY-4.0); see the ChatTS release (Zenodo 14349206) |
| MTBench (Yale GGLab) | see the dataset's Hugging Face page |
| TimeSeriesExam (CMU Auton Lab) | see the dataset's original source |
Sources and links for every dataset are in Benchmarks/README.md.