Downloads · 30 days
12
100% of all-time downloads
TimeDuet/TimeDuet-Qwen
TimeDuet-Qwen is a machine learning model from TimeDuet. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft.
Final SFT + GRPO adapter for TimeDuet-Qwen, the model reported in TimeDuet: A Benchmark for Audio-Visual Temporal Grounding on Misaligned Streams.
Downloads · 30 days
12
100% of all-time downloads
All-time downloads
12
Public
Repo size
80.8 MB
Likes
1
Trending 1
Click a slice to open those files.
.safetensors80.8 MB · 100%
From the Hugging Face model README
Final SFT + GRPO adapter for TimeDuet-Qwen, the model reported in TimeDuet: A Benchmark for Audio-Visual Temporal Grounding on Misaligned Streams.
This repository contains the final LoRA adapter, not standalone merged base-model weights. Load it on Qwen/Qwen2.5-Omni-7B. The adapter was initialized from the SFT adapter and then optimized with GRPO; do not apply it on top of an SFT-merged model or stack a separate SFT adapter. This matches the recorded evaluation configuration.
This is the final model used for the paper results below. Adapter rank is 8, alpha 32, with language-model projection targets. Detailed training provenance is recorded in release_provenance.json for reproducibility.
import torch
from peft import PeftModel
from transformers import Qwen2_5OmniForConditionalGeneration, Qwen2_5OmniProcessor
base_id = "Qwen/Qwen2.5-Omni-7B"
adapter_id = "TimeDuet/TimeDuet-Qwen"
model = Qwen2_5OmniForConditionalGeneration.from_pretrained(
base_id,
torch_dtype=torch.bfloat16,
device_map="auto",
attn_implementation="sdpa",
)
model.disable_talker()
model = PeftModel.from_pretrained(model, adapter_id)
model.eval()
processor = Qwen2_5OmniProcessor.from_pretrained(base_id)
Use compatible torch, transformers, peft, accelerate, and qwen-omni-utils versions; the adapter's saved configuration was generated with PEFT 0.19.1. This is a loading example, not the full benchmark evaluation script. Refer to the project code for audiovisual preprocessing and scoring.
For grounding, provide both the full video and its audio. Recorded evaluation uses 1 FPS video sampling, the full audio track, deterministic generation, max_new_tokens=1024, use_audio_in_video=True, and return_audio=False. Do not silently truncate frames or omit the audio. evaluation_protocol.json contains the canonical prompt and protocol; schema_contract.json describes the output format.
The model returns one interval set inside <answer>{"intervals":[[start,end], ...]}</answer> for times when the event is audible, visible, or both. Modality-specific accuracy is evaluated from that same prediction against separate ground truth.
The model underwent supervised fine-tuning on TimeDuet followed by GRPO on 12,536 shifted training queries. The reward is:
0.5 * T-IoU + 0.5 * min(A-IoU, V-IoU) - 0.3 * abs(A-IoU - V-IoU)
Aligned queries are excluded from this RL stage. Optimizer state, intermediate checkpoints and training-resume files are not part of this inference release.
Recorded TimeDuet test performance over 2,202 queries:
| Metric | Score |
|---|---|
| mIoU | 0.641 |
| [email protected] | 0.911 |
| [email protected] | 0.833 |
| [email protected] | 0.445 |
| A-IoU | 0.596 |
| V-IoU | 0.663 |
| AV overlap | 0.609 |
| Audio-only | 0.374 |
| Visual-only | 0.524 |
evaluation_metrics.json contains the original recorded scores. These are the paper experiment results, not a newly executed inference benchmark of this upload.
The model inherits the base model's limitations and can produce incorrect or malformed intervals. Constructed temporal shifts do not cover every natural alignment condition. It is not validated for safety-critical decisions. Base model use is subject to Qwen2.5-Omni-7B's license; an additional license for this trained adapter has not yet been specified by the authors.
Sihyeong Kim, Jaeyeong Choi, Joohyun Oh, Taeyeong Jeong, Bomin Kang, Daehee Park, and Jisoo Mok. TimeDuet: A Benchmark for Audio-Visual Temporal Grounding on Misaligned Streams.