Downloads · 30 days
30
24% of all-time downloads
MYJOKERML/SDiaReward-7B
SDiaReward-7B is a audio classification model from MYJOKERML. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
SDiaReward-7B is a reward model for spoken dialogue, built on Qwen2.5-Omni-7B with a pooling layer and a linear scoring head (QwenOmniThinkerReward). Given a multi-turn conversation containing interleaved speech and t…
Downloads · 30 days
30
24% of all-time downloads
All-time downloads
127
Public
Parameters
8.9B
17.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors17.9 GB · 100%
From the Hugging Face model README
SDiaReward-7B is a reward model for spoken dialogue, built on
Qwen2.5-Omni-7B with a pooling layer and a linear scoring head
(QwenOmniThinkerReward). Given a multi-turn conversation containing interleaved
speech and text, it produces a scalar reward that reflects two qualities:
Modality-awareness — prosody, emotion, and acoustic naturalness (real human speech vs. synthetic TTS).
Colloquialness — spontaneous spoken style vs. scripted written style.
📄 Paper: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness (ACL 2026 Main Conference) — arXiv:2603.14889
📚 Training data (gated): SDiaReward dataset.
🧩 Smaller variant: SDiaReward-3B.
On the ESDR-Bench validation set:
| Model | Accuracy | Eval loss | Margin |
|---|---|---|---|
| SDiaReward-7B | 0.971 | 0.358 | 1.167 |
| SDiaReward-3B | 0.917 | 0.419 | 0.869 |
This checkpoint uses a custom reward architecture (QwenOmniThinkerReward: a
Qwen2.5-Omni backbone with a pooling layer and a scalar reward head). The custom
modeling code is not bundled in this repository, so the weights cannot be
loaded with a plain AutoModel/pipeline call. To load the model and score
conversations, follow the loading and inference instructions in the official code
repository:
👉 https://github.com/MM-Speech/SDiaReward (checkpoint id: MYJOKERML/SDiaReward-7B)
Trained with TRL's reward trainer on the SDiaReward preference dataset
(11,630 episode-level preference pairs, ~200 hours of paired speech). Backbone:
Qwen2.5-Omni-7B; pooling variant: mean_center; 2 epochs.
Released under Apache-2.0 for research use. The reward signal is intended for evaluating and improving spoken-dialogue systems; it is not a safety classifier.
@article{lu2026modeling,
title={Modeling and benchmarking spoken dialogue rewards with modality and colloquialness},
author={Lu, Jingyu and Wang, Yuhan and Zhuo, Fan and Cheng, Xize and Pan, Changhao and Pu, Xueyi and Chen, Yifu and Wen, Chenyuhao and Liang, Tianle and Zhao, Zhou},
journal={arXiv preprint arXiv:2603.14889},
year={2026}
}