Downloads · 30 days
52
13% of all-time downloads
nile2000/Fun-ASR-Nano-GGUF
Fun-ASR-Nano-GGUF is a automatic speech recognition model from nile2000. Use it when you need speech turned into text. It is set up for gguf. The card lists the license as apache-2.0.
GGUF build of Fun-ASR-Nano (SenseVoice SAN-M encoder + adaptor + Qwen3-0.6B LLM decoder) for the zero-Python, CPU/edge FunASR llama.cpp runtime — the accuracy leader (LLM decoder), single C++ binary.
Downloads · 30 days
52
13% of all-time downloads
All-time downloads
389
Public
Repo size
2.3 GB
Likes
0
Public
Click a slice to open those files.
.gguf2.3 GB · 100%
From the Hugging Face model README
GGUF build of Fun-ASR-Nano (SenseVoice SAN-M encoder + adaptor + Qwen3-0.6B LLM decoder) for the zero-Python, CPU/edge FunASR llama.cpp runtime — the accuracy leader (LLM decoder), single C++ binary.
The Fun-ASR-Nano LLM (Qwen3-0.6B) ships in three tiers — all within 0.1% CER (184-file micro-CER). Pair any with funasr-encoder-f16.gguf (470 MB).
| LLM file | size | CER ↓ | speed |
|---|---|---|---|
qwen3-0.6b-q4km.gguf | 484 MB | 8.35% | 6.1× |
qwen3-0.6b-q5km.gguf | 551 MB | 8.25% | 5.7× |
qwen3-0.6b-q8_0.gguf | 805 MB | 8.30% | 6.0× |
Recommended: q4_K_M (smallest) or q5_K_M (best).
These are GGUF weights for the FunASR llama.cpp runtime — a whisper.cpp-style, single self-contained binary for CPU / edge. Grab a prebuilt binary, then fetch this model and run:
runtime-llamacpp-v*)bash download-funasr-model.sh nano ./gguf
llama-funasr-cli --enc ./gguf/funasr-encoder-f16.gguf -m ./gguf/qwen3-0.6b-q8_0.gguf --vad ./gguf/fsmn-vad.gguf -a audio.wav
| file | size | notes |
|---|---|---|
funasr-encoder-f16.gguf | 470 MB | audio encoder + adaptor (f16) |
qwen3-0.6b-q8_0.gguf | 805 MB | LLM decoder, recommended (Q8_0) |
qwen3-0.6b-q4km.gguf | 484 MB | LLM decoder, smaller (Q4_K_M) |
llama-funasr-cli --enc funasr-encoder-f16.gguf -m qwen3-0.6b-q8_0.gguf -a audio.wav --vad fsmn-vad.gguf
On CPU: 8.30 % CER on the 184-clip Mandarin benchmark (vs whisper.cpp 22–31 %).