Downloads · 30 days
3.1K
15% of all-time downloads
funasr/fsmn-vad
fsmn-vad is a voice activity detection model from funasr. Use it for the voice activity detection task on the model card, and read the license before you ship it in a product. It is set up for funasr. The card lists the license as apache-2.0.
Downloads · 30 days
3.1K
15% of all-time downloads
All-time downloads
21.4K
Public
Repo size
4 MB
Likes
26
Public
Click a slice to open those files.
.wav2.3 MB · 56%
From the Hugging Face model README
This model is part of the FunASR ecosystem — one industrial-grade open-source toolkit for ASR · VAD · punctuation · speaker diarization · emotion / event · LLM-ASR. A Star really helps the project (and keeps you updated):
🌟 FunASR · 🌟 SenseVoice · 🌟 Fun-ASR · 🌟 FunClip
</div>Voice Activity Detection — accurately detect speech segments in audio, essential for long-audio processing pipelines.
FSMN-VAD uses a Feedforward Sequential Memory Network to detect speech/non-speech boundaries with high precision and low latency. It supports both streaming and offline modes.
from funasr import AutoModel
# Standalone VAD
model = AutoModel(model="funasr/fsmn-vad", hub="hf", device="cuda")
result = model.generate(input="long_audio.wav")
# Returns speech segments: [[start_ms, end_ms], [start_ms, end_ms], ...]
print(result[0]["value"])
from funasr import AutoModel
# VAD automatically segments long audio before ASR
model = AutoModel(
model="funasr/paraformer-zh",
hub="hf",
vad_model="funasr/fsmn-vad",
device="cuda",
)
result = model.generate(input="meeting_2hours.wav")
print(result[0]["text"])
max_single_segment_time)| Property | Value |
|---|---|
| Architecture | FSMN (Feedforward Sequential Memory Network) |
| Sample Rate | 16kHz |
| Modes | Streaming + Offline |