Downloads · 30 days
62
11% of all-time downloads
inclusionAI/AudioMCQ-Weak-To-Strong
AudioMCQ-Weak-To-Strong is a machine learning model from inclusionAI. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
62
11% of all-time downloads
All-time downloads
571
Public
Parameters
10.7B
21.5 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors21.5 GB · 100%
From the Hugging Face model README
Based on community feedback, we identified a flaw in our evaluation script that artificially inflated the MMSU scores of our released models by ignoring sequence order. We sincerely apologize for this oversight. Crucially, please note that our AudioMCQ training data, the paper's conclusions regarding audio-contribution, and the MMAR/MMAU metrics remain completely unaffected. When comparing against our work, we recommend reporting the MMAR/MMAU results or re-evaluating our published checkpoints using your own exact-match algorithm. We deeply apologize for any inconvenience this may have caused to the research community.
This repository contains the Weak-to-Strong model checkpoint from our paper "Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models". This model demonstrates state-of-the-art performance on audio question-answering benchmarks through our novel audio-contribution-aware post-training approach.
The Weak-to-Strong training paradigm follows a two-stage approach:
Stage 1: SFT on weak audio-contribution data
Stage 2: GRPO (RL) on strong audio-contribution data
This paradigm begins with supervised fine-tuning on samples with weak audio contribution (where visual or textual cues provide substantial information), then applies reinforcement learning on challenging strong audio-contribution samples to enhance audio-specific understanding capabilities.
Our model loading and usage methods are identical to those of Qwen2.5-Omni. Please refer to the official documentation.
The evaluation input prompt structure is:
[Question] Please choose the answer from the following options: ['Option1', 'Option2', 'Option3', 'Option4']. Output the final answer in <answer> </answer>.
# Load model following Qwen2.5-Omni documentation
# Apply system prompt: "You are an audio understanding model that answers multiple choice questions based on audio content."
# Format your question with the input structure above
The Weak-to-Strong model achieves competitive performance across multiple benchmarks:
For detailed performance metrics and comparisons, please refer to our paper.
If you find this model useful in your research, please cite:
@inproceedings{he2025audiomcq,
title={Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models},
author={He, Haolin and others},
booktitle={Proceedings of the International Conference on Learning Representations (ICLR)},
year={2026}
}
We thank the organizers of DCASE 2025 and the research community for their valuable feedback and support.