Downloads · 30 days
0
YuCeong-May/MLC-SLM
MLC-SLM is a automatic speech recognition model from YuCeong-May. Use it when you need speech turned into text. The card lists the license as apache-2.0.
This repository contains the models and code presented in the paper Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR.
Downloads · 30 days
0
Access
Public
Updated Jan 15, 2026
Repo size
26.7 GB
Likes
1
Public
Click a slice to open those files.
.pt26.7 GB · 100%
From the Hugging Face model README
This repository contains the models and code presented in the paper Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR.
The project was developed for the INTERSPEECH 2025 Challenge on Multilingual Conversational Speech Language Models (MLC-SLM).
The proposed Speech-LLM is an enhanced framework that integrates fine-tuned Whisper and mHuBERT encoders with a Large Language Model (Qwen2.5-7B) to enrich speech representations for multilingual conversational ASR. It utilizes cross-attention-based fusion mechanisms to exploit complementary information between generative (Whisper) and discriminative (mHuBERT) speech features.
Performance (CER/WER) on the MLC-SLM Challenge datasets:
| System | Dev | Eval | CV-Test |
|---|---|---|---|
| Whisper (LoRA-fine-tuned) | 11.40 | 10.71 | 11.47 |
| Whisper (Full-fine-tuned) | 10.99 | 10.07 | 13.11 |
| Proposed Speech-LLM | 11.74 | 10.69 | 15.26 |
The models were trained on the official ~1500h training set from the MLC-SLM Challenge, covering 11 languages and 15 categories (including various English accents).
@article{mlcslm2025bridging,
title={Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR},
author={Yuxiang Mei, Dongxing Xu, Jiaen Liang and Yanhua Long},
journal={arXiv preprint arXiv:2601.01461},
year={2025}
}