Downloads · 30 days
14
15% of all-time downloads
zhoujiaming777/DIFFA-2
DIFFA-2 is a machine learning model from zhoujiaming777. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
[](https://arxiv.org/abs/2601.23161v1) [](https://huggingface.co/zhoujiaming777/DIFFA-2) [](https://github.com/NKU-HLT/DIFFA)
Downloads · 30 days
14
15% of all-time downloads
All-time downloads
92
Public
Repo size
734 MB
Likes
2
Public
Click a slice to open those files.
.bin734 MB · 100%
From the Hugging Face model README
In this paper, We introduce DIFFA-2, a practical diffusion-based LALM for general audio understanding. DIFFA-2 upgrades the speech encoder, employs dual semantic and acoustic adapters, and is trained with a four-stage curriculum that combines semantic and acoustic alignment, large-scale supervised fine-tuning, and variance-reduced preference optimization, using only fully open-source corpora. Experiments on MMSU, MMAU, and MMAR show that DIFFA-2 consistently improves over DIFFA and is competitive to strong AR LALMs under practical training budgets, supporting diffusion-based modeling is a viable backbone for large-scale audio understanding.
We have open-sourced the checkpoints for stage 1 and stage 4. The files in the root directory of the repository are for stage4, and stage1 is located in the stage1 folder.