Downloads · 30 days
2.1K
26% of all-time downloads
FunAudioLLM/fsmn-vad-GGUF
fsmn-vad-GGUF is a voice activity detection model from FunAudioLLM. Use it for the voice activity detection task on the model card, and read the license before you ship it in a product. It is set up for gguf. The card lists the license as apache-2.0.
GGUF build of FunASR's FSMN-VAD for the zero-Python, CPU/edge FunASR llama.cpp runtime. Native ggml voice-activity detection: segment long audio entirely in C++, no Python at runtime.
Downloads · 30 days
2.1K
26% of all-time downloads
All-time downloads
8.2K
Public
Repo size
1.7 MB
Likes
3
Public
Click a slice to open those files.
.gguf1.7 MB · 100%
From the Hugging Face model README
GGUF build of FunASR's FSMN-VAD for the zero-Python, CPU/edge FunASR llama.cpp runtime. Native ggml voice-activity detection: segment long audio entirely in C++, no Python at runtime.
These are GGUF weights for the FunASR llama.cpp runtime — a whisper.cpp-style, single self-contained binary for CPU / edge. Grab a prebuilt binary, then fetch this model and run:
runtime-llamacpp-v*)bash download-funasr-model.sh fsmn-vad ./gguf
# fsmn-vad is the VAD used by the ASR runtimes via --vad (see the SenseVoice / Paraformer / Fun-ASR-Nano GGUF repos)
| file | size | notes |
|---|---|---|
fsmn-vad.gguf | 1.7 MB | FSMN encoder + CMVN |
Pass --vad to any FunASR llama.cpp tool to segment long audio internally:
llama-funasr-sensevoice -m sensevoice-small.gguf -a long.wav --vad fsmn-vad.gguf
llama-funasr-cli --enc funasr-encoder-f16.gguf -m qwen3-0.6b-q8_0.gguf -a long.wav --vad fsmn-vad.gguf
Segment boundaries match the PyTorch fsmn-vad front end within ~10 ms.