Downloads · 30 days
7
16% of all-time downloads
SunsetCollector/Enhanced-TGRPO-Phase1
Enhanced-TGRPO-Phase1 is a machine learning model from SunsetCollector. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Project: CS41 Enhanced T-GRPO for Video Temporal Reasoning (USYD Capstone)
Downloads · 30 days
7
16% of all-time downloads
All-time downloads
45
Public
Parameters
8.3B
16.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
Project: CS41 Enhanced T-GRPO for Video Temporal Reasoning (USYD Capstone)
| Parameter | Value |
|---|---|
| num_generations | 4 |
| max_pixels | 200704 |
| max_prompt_length | 8192 |
| learning_rate | 1e-6 |
| beta (KL coeff) | 0.04 |
| corruption_strength | 1.0 (with curriculum) |
| margin_scale | 0.5 |
| reward_threshold | 0.1 |
| gradient_accumulation | 2 |
| DeepSpeed | ZeRO-3 + CPU offload |
WandB: https://wandb.ai/wuguangbo464-the-university-of-sydney/huggingface/runs/lep2va20
| Benchmark | Score |
|---|---|
| MMVU (mc) | 59.0% |
| Model | MMVU (mc) |
|---|---|
| Qwen2.5-VL-7B-SFT (starting point) | 61.3% |
| Baseline T-GRPO (P0 fixes only) | 61.1% |
| Enhanced T-GRPO (this model) | 59.0% |
| Video-R1-7B (paper, full data) | 64.2% |
Phase 1 small-scale validation. MMVU score drop is expected — trained on 0.2% of full data, and MMVU tests domain knowledge, not temporal reasoning. Training curves confirm the algorithm works (reward trending up, temporal signal active). Phase 2 will use full 260k dataset and evaluate on VSI-Bench/TempCompass (temporal reasoning benchmarks).