Downloads · 30 days
0
mistral-hackaton-2026/voxtral_model
voxtral_model is a automatic speech recognition model from mistral-hackaton-2026. Use it when you need speech turned into text. It is set up for custom. The card lists the license as apache-2.0.
Streaming speech-to-text model with ~4 billion parameters. Weights in BF16 safetensors format, extracted from mistralai/Voxtral-Mini-4B-Realtime-2602.
Downloads · 30 days
0
Access
Public
Updated Mar 1, 2026
Repo size
—
Likes
1
Public
Click a slice to open those files.
.md3.7 KB · 71%
From the Hugging Face model README
Streaming speech-to-text model with ~4 billion parameters. Weights in BF16 safetensors format, extracted from mistralai/Voxtral-Mini-4B-Realtime-2602.
Pipeline:
WAV → 16kHz → Mel Spectrogram → Conv Stem → Encoder → Downsample 4x → Adapter → Decoder → Tokens
| Parameter | Value |
|---|---|
| Sample rate | 16,000 Hz |
| Frame rate | 12.5 Hz |
| Mel bins | 128 |
| Hop length | 160 samples (10ms) |
| Window size | 400 samples (25ms) |
| 1 text token | 80ms of audio |
| Parameter | Value |
|---|---|
| dim | 1280 |
| layers | 32 |
| heads | 32 (MHA) |
| head_dim | 64 |
| hidden_dim | 5120 |
| FFN | SwiGLU |
| Norm | RMSNorm (eps=1e-5) |
| Position | RoPE (theta=1e6, interleaved) |
| Attention | causal, sliding window=750 |
Conv stem: conv1d(128→1280, k=3, s=1) → GELU → conv1d(1280→1280, k=3, s=2) → GELU
[seq/4, 5120] → Linear(5120→3072) → GELU → Linear(3072→3072) → [seq/4, 3072]
| Parameter | Value |
|---|---|
| dim | 3072 |
| layers | 26 |
| heads | 32 |
| KV heads | 8 (GQA 4:1) |
| head_dim | 128 |
| hidden_dim | 9216 |
| Norm | RMSNorm (eps=1e-5) |
| Position | RoPE (theta=1e6) |
| Attention | causal, sliding window=8192 |
| Vocab size | 131,072 |
| Tied embeddings | yes |
The decoder uses adaptive RMS normalization conditioned on transcription delay (6 delay tokens = 480ms).
consolidated.safetensors (8.3 GB) — 711 tensors, all BF16params.json — model configtekken.json (14.9 MB) — Tekken tokenizer| Token | ID |
|---|---|
| BOS | 1 |
| EOS | 2 |
| STREAMING_PAD | 32 |
Token IDs 0–999 are special tokens. IDs 1000+ index into the vocabulary (base64-encoded byte sequences in tekken.json).
| Parameter | Value |
|---|---|
| sampling_rate | 16,000 |
| frame_rate | 12.5 (80ms per token) |
| transcription_delay_ms | 480 (6 delay tokens) |
| left_pad_tokens | 32 |
| right_pad_tokens (offline) | 17 |
[BOS] + [STREAMING_PAD] × 38 (1 + 32 left-pad + 6 delay)audio_embed[i] + tok_embed(prompt[i]) for positions 0..L-2audio_embed[pos] + tok_embed(prev_token), greedy argmaxA pure C implementation of this model is available at voxtral.c — runs on Apple Silicon (Metal) and CPU (BLAS), with streaming microphone input.
Original model by Mistral AI: mistralai/Voxtral-Mini-4B-Realtime-2602
Built for the Mistral Hackathon 2026.