Downloads · 30 days
116
83% of all-time downloads
brgroup/BR-Voice-Reasoner
BR-Voice-Reasoner is a audio-text-to-text model from brgroup. Use it for the audio-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
BR-Voice-Reasoner is a multimodal reasoning model for spoken interaction. It understands spoken requests and responds with knowledge, reasoning, and instruction-aware text. It operates directly on audio, allowing appl…
Downloads · 30 days
116
83% of all-time downloads
All-time downloads
139
Public
Parameters
31.7B
63.4 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors63.4 GB · 100%
From the Hugging Face model README
BR-Voice-Reasoner is a multimodal reasoning model for spoken interaction. It understands spoken requests and responds with knowledge, reasoning, and instruction-aware text. It operates directly on audio, allowing applications to reason over what was said without inserting a separate transcription model into the pipeline.
The model is designed for spoken question answering, knowledge-intensive voice queries, multi-step reasoning, instruction following, open-ended interaction, and spoken interaction involving safety and refusal behavior.
| Item | Value |
|---|---|
| Input | Audio, text, image, and video |
| Output | Text |
| Architecture | Multimodal Mixture-of-Experts |
| Thinker LM | 30B total / 3B activated MoE |
| Primary evaluation language | English |
| Weights | BF16 Safetensors |
Post-training primarily targets speech-conditioned interaction; image and video capabilities are inherited from the base model and were not comprehensively re-evaluated.
BR-Voice-Reasoner is trained from Qwen3-Omni-30B-A3B-Thinking using cross-modal on-policy distillation. During training, the student processes spoken requests, while a frozen teacher uses their aligned text forms to provide token-level learning signals along student-generated trajectories.
Optimization is restricted to two audio-language projection modules. The language model, vision encoder, and remaining audio encoder parameters stay frozen, focusing post-training on the audio-language interface used to access the model's existing knowledge, reasoning, and instruction-following capabilities.
Training data cover spoken knowledge, reasoning, instruction following, open-ended interaction, and safety-oriented tasks using both real-world and synthesized speech.
BR-Voice-Reasoner is evaluated on nine VoiceBench subsets covering knowledge, reasoning, instruction following, safety, and open-ended spoken interaction. All scores are reported on a 0–100 scale; higher is better.
<table align="center"> <thead> <tr> <th align="left">VoiceBench subset</th> <th><div align="center">Qwen3-Omni<br>30B-A3B-Thinking</div></th> <th><div align="center">Nemotron 3 Nano Omni</div></th> <th><div align="center">BR-Voice-Reasoner</div></th> </tr> </thead> <tbody> <tr><td align="left">IFEval</td><td><div align="center">80.6</div></td><td><div align="center"><strong>88.7</strong></div></td><td><div align="center">83.2</div></td></tr> <tr><td align="left">BBH</td><td><div align="center">88.9</div></td><td><div align="center"><strong>91.1</strong></div></td><td><div align="center">90.4</div></td></tr> <tr><td align="left">AdvBench</td><td><div align="center">97.2</div></td><td><div align="center"><strong>100.0</strong></div></td><td><div align="center">99.8</div></td></tr> <tr><td align="left">AlpacaEval</td><td><div align="center">96.4</div></td><td><div align="center">95.0</div></td><td><div align="center"><strong>97.3</strong></div></td></tr> <tr><td align="left">CommonEval</td><td><div align="center">90.5</div></td><td><div align="center">91.3</div></td><td><div align="center"><strong>94.2</strong></div></td></tr> <tr><td align="left">WildVoice</td><td><div align="center">90.5</div></td><td><div align="center">91.7</div></td><td><div align="center"><strong>93.2</strong></div></td></tr> <tr><td align="left">OpenBookQA</td><td><div align="center">94.3</div></td><td><div align="center">93.0</div></td><td><div align="center"><strong>96.0</strong></div></td></tr> <tr><td align="left">MMSU</td><td><div align="center">83.0</div></td><td><div align="center">82.3</div></td><td><div align="center"><strong>85.7</strong></div></td></tr> <tr><td align="left">SD-QA</td><td><div align="center"><strong>78.1</strong></div></td><td><div align="center">71.4</div></td><td><div align="center">74.3</div></td></tr> <tr><td align="left"><strong>VoiceBench Avg</strong></td><td><div align="center">88.8</div></td><td><div align="center">89.4</div></td><td><div align="center"><strong>90.5</strong></div></td></tr> </tbody> </table>Qwen3-Omni results are taken from the official Qwen3-Omni model card, while Nemotron 3 Nano Omni results are taken from its official model card. BR-Voice-Reasoner results were obtained using the evaluation protocol described below. BR-Voice-Reasoner values are rounded to the same one-decimal format. The external columns and BR-Voice-Reasoner were not produced by a single shared evaluation run; cross-column differences are therefore reported as references rather than as a strictly controlled comparison.
| Item | Setting |
|---|---|
| Reasoning | Enabled |
| Generation | temperature 0.6, top-p 0.95, top-k 20 |
| Model judge | GPT-4o-mini, three judgments |
| Overall | Mean of nine normalized subset scores |
AlpacaEval, CommonEval, and WildVoice ratings are normalized to a 0–100 scale before aggregation. The VoiceBench average is the arithmetic mean of the nine normalized subset scores.
Full generation, seeding, evaluator, and retry settings are provided in
evaluation_protocol.json.
pip install "transformers==5.12.1" "accelerate" "qwen-omni-utils==0.0.9"
ffmpeg must also be available on the system for media loading. FlashAttention
2 is optional; install it separately with
pip install flash-attn --no-build-isolation on compatible hardware.
from transformers import (
Qwen3OmniMoeForConditionalGeneration,
Qwen3OmniMoeProcessor,
)
from qwen_omni_utils import process_mm_info
MODEL_ID = "brgroup/BR-Voice-Reasoner"
model = Qwen3OmniMoeForConditionalGeneration.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto",
# attn_implementation="flash_attention_2", # optional
)
processor = Qwen3OmniMoeProcessor.from_pretrained(MODEL_ID)
conversation = [
{
"role": "user",
"content": [
{"type": "audio", "audio": "example.wav"},
{"type": "text", "text": "Answer the question in the audio."},
],
}
]
prompt = processor.apply_chat_template(
conversation,
add_generation_prompt=True,
tokenize=False,
)
audios, images, videos = process_mm_info(conversation)
inputs = processor(
text=prompt,
audio=audios,
images=images,
videos=videos,
return_tensors="pt",
padding=True,
)
inputs = inputs.to(model.device).to(model.dtype)
text_ids, _ = model.generate(
**inputs,
return_audio=False,
thinker_return_dict_in_generate=True,
max_new_tokens=2048,
)
response = processor.batch_decode(
text_ids.sequences[:, inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)
print(response[0])
BR-Voice-Reasoner is released under the Apache License 2.0 and is derived from
Qwen3-Omni-30B-A3B-Thinking,
whose original model copyright is Copyright 2025 Alibaba Cloud.
Modifications are Copyright 2026 Bairong Inc. See LICENSE and
NOTICE for license terms and attribution.