Downloads · 30 days
21
2% of all-time downloads
CodeGoat24/UnifiedReward-2.0-qwen3vl-4b
UnifiedReward-2.0-qwen3vl-4b is a machine learning model from CodeGoat24. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
UnifiedReward-2.0-qwen3vl-4b is the first unified reward model based on Qwen/Qwen3-VL-4B-Instruct for multimodal understanding and generation assessment, enabling both pairwise ranking and pointwise scoring, which can…
Downloads · 30 days
21
2% of all-time downloads
All-time downloads
893
Public
Parameters
4.4B
8.9 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors8.9 GB · 100%
From the Hugging Face model README
UnifiedReward-2.0-qwen3vl-4b is the first unified reward model based on Qwen/Qwen3-VL-4B-Instruct for multimodal understanding and generation assessment, enabling both pairwise ranking and pointwise scoring, which can be employed for vision model preference alignment.
For further details, please refer to the following resources:
| Reward Model | Method | Image Generation | Image Understanding | Video Generation | Video Understanding |
|---|---|---|---|---|---|
| PickScore | Point | √ | |||
| HPS | Point | √ | |||
| ImageReward | Point | √ | |||
| LLaVA-Critic | Pair/Point | √ | |||
| IXC-2.5-Reward | Pair/Point | √ | √ | ||
| VideoScore | Point | √ | |||
| LiFT | Point | √ | |||
| VisionReward | Point | √ | √ | ||
| VideoReward | Point | √ | |||
| UnifiedReward (Ours) | Pair/Point | √ | √ | √ | √ |
@article{unifiedreward,
title={Unified reward model for multimodal understanding and generation},
author={Wang, Yibin and Zang, Yuhang and Li, Hao and Jin, Cheng and Wang, Jiaqi},
journal={arXiv preprint arXiv:2503.05236},
year={2025}
}