Downloads · 30 days
230
63% of all-time downloads
FermionResearch/Phonon-1-Micro
Phonon-1-Micro is a automatic speech recognition model from FermionResearch. Use it when you need speech turned into text. It is set up for mlx. The card lists the license as apache-2.0.
Phonon-1 Micro is an open speech recognition model for English, the smallest build of the Phonon-1 family. It downloads in 285 MB, runs on a laptop or a datacenter GPU, and beats Moonshine base on all eight benchmarks…
Downloads · 30 days
230
63% of all-time downloads
All-time downloads
363
Public
Repo size
285 MB
Likes
3
Public
Click a slice to open those files.
.zst285 MB · 100%
From the Hugging Face model README
Phonon-1 Micro is an open speech recognition model for English, the smallest build of the Phonon-1 family. It downloads in 285 MB, runs on a laptop or a datacenter GPU, and beats Moonshine base on all eight benchmarks below. It was trained at 2.4 bits per weight from the start.
| Benchmark | Phonon-1 (415 MB) | Phonon-1 Micro (285 MB) | Parakeet-0.6B 4-bit (637 MB) | Moonshine base (248 MB) | Whisper large-v3-turbo (1,619 MB) | Whisper small (967 MB) | wav2vec2-large (1,262 MB) | Qwen3-ASR teacher (1,569 MB) |
|---|---|---|---|---|---|---|---|---|
| LibriSpeech test-clean | 2.640 | 3.002 | 2.186 | 3.417 | 2.10 | 3.4† | 2.8† | 2.235 |
| LibriSpeech test-other | 5.699 | 6.511 | 3.937 | 8.262 | 4.07 | 7.6† | 6.3† | 4.618 |
| TED-LIUM | 3.421 | 3.878 | 2.829 | 5.272 | — | — | — | 2.889 |
| SPGISpeech | 4.163 | 4.858 | 4.104 | 5.731 | 2.79† | — | 13.31† | 3.074 |
| VoxPopuli | 8.394 | 9.177 | 6.345 | 10.470 | 11.22† | — | — | 7.151 |
| GigaSpeech | 11.396 | 11.882 | 9.614 | 12.114 | 8.52† | — | — | 9.321 |
| Earnings-22 | 12.571 | 14.771 | 11.190 | 17.872 | 11.07† | — | 36.28† | 11.188 |
| AMI | 13.084 | 14.094 | 12.723 | 17.790 | 15.16† | — | — | 12.560 |
| Macro (eight benchmarks) | 7.67 | 8.52 | 6.62 | 10.1 | — | — | — | 6.63 |
Word error rate, lower is better. Unmarked cells: measured by us — full test sets, Whisper English text normalizer, greedy decoding. † = published figure (model card, paper, or the Open ASR Leaderboard). Dash = no comparable measurement.
pip install fermion-research
fermion transcribe phonon-1-micro recording.wav
Or serve an OpenAI-compatible endpoint:
fermion serve phonon-1-micro
curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
-F "[email protected]" \
-F "model=FermionResearch/Phonon-1-Micro"
The same weights run on a Mac (via MLX), on an NVIDIA GPU, or on a plain CPU; the runtimes and Docker images are in the GitHub repo.
Apache License 2.0 for the weights and the command line. Base model: Qwen/Qwen3-ASR-0.6B, Apache-2.0.