Downloads · 30 days
22
17% of all-time downloads
BUT-FIT/Dixtral_QA
Dixtral_QA is a automatic speech recognition model from BUT-FIT. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
This repository hosts DixtralQA, developed by BUT Speech@FIT. Dixtral couples the Voxtral-Mini-3B spoken-language model with the DiCoW diarization-conditioned encoder, giving the LLM target-speaker awareness in multi-…
Downloads · 30 days
22
17% of all-time downloads
All-time downloads
128
Public
Parameters
4.7B
18.7 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.4 GB · 100%
From the Hugging Face model README
This repository hosts Dixtral_QA, developed by BUT Speech@FIT. Dixtral couples the Voxtral-Mini-3B spoken-language model with the DiCoW diarization-conditioned encoder, giving the LLM target-speaker awareness in multi-talker audio.
This checkpoint is tuned for spoken question answering over conversational/meeting audio. For pure target-speaker transcription, use Dixtral_TS-ASR instead.
from transformers import AutoModel, AutoProcessor
MODEL_NAME = "BUT-FIT/Dixtral_QA"
model = AutoModel.from_pretrained(MODEL_NAME, trust_remote_code=True)
processor = AutoProcessor.from_pretrained(MODEL_NAME)
➡️ For full inference pipelines (diarization → FDDT masks → generation), see the Dixtral GitHub repository.
📧 Email: ipoloka@fit.vut.cz 🏢 Affiliation: BUT Speech@FIT, Brno University of Technology 🔗 GitHub: BUTSpeechFIT