Downloads · 30 days
0
TransWithAI/Whisper-Vad-EncDec-ASMR-onnx
Whisper-Vad-EncDec-ASMR-onnx is a audio classification model from TransWithAI. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
This is a refined Whisper-based Voice Activity Detection (VAD) model that leverages the pre-trained Whisper encoder with a lightweight non-autoregressive decoder for high-precision speech activity detection. While fin…
Downloads · 30 days
0
Access
Public
Updated Nov 17, 2025
Repo size
119 MB
Likes
16
Public
Click a slice to open those files.
.onnx119 MB · 100%
From the Hugging Face model README
This is a refined Whisper-based Voice Activity Detection (VAD) model that leverages the pre-trained Whisper encoder with a lightweight non-autoregressive decoder for high-precision speech activity detection. While fine-tuned on Japanese ASMR content for optimal performance on soft speech and whispers, the model retains Whisper's robust multilingual foundation, enabling effective speech detection across diverse languages and acoustic conditions. It has been optimized and exported to ONNX format for efficient inference across different platforms. For the full training code, configs, and ONNX export utilities, see the GitHub repository: TransWithAI/whisper-vad.
This work builds upon recent research demonstrating the positive transfer of Whisper's speech representations to VAD tasks, as shown in WhisperSeg and related work.
import numpy as np
import onnxruntime as ort
from transformers import WhisperFeatureExtractor
import librosa
# Load model
session = ort.InferenceSession("model.onnx")
feature_extractor = WhisperFeatureExtractor.from_pretrained("openai/whisper-base")
# Load and preprocess audio
audio, sr = librosa.load("audio.wav", sr=16000)
audio_chunk = audio[:480000] # 30 seconds
# Extract features
inputs = feature_extractor(
audio_chunk,
sampling_rate=16000,
return_tensors="np"
)
# Run inference
outputs = session.run(None, {session.get_inputs()[0].name: inputs.input_features})
predictions = outputs[0] # Shape: [1, 1500] - 1500 frames of 20ms each
# Apply threshold
speech_frames = predictions[0] > 0.5
The model repository includes a comprehensive inference.py script with advanced features:
from inference import WhisperVADInference
# Initialize model
vad = WhisperVADInference(
model_path="model.onnx",
threshold=0.5, # Speech detection threshold
min_speech_duration=0.25, # Minimum speech segment duration
min_silence_duration=0.1 # Minimum silence between segments
)
# Process audio file
segments = vad.process_audio("audio.wav")
# Segments format: List of (start_time, end_time) tuples
for start, end in segments:
print(f"Speech detected: {start:.2f}s - {end:.2f}s")
# Process audio stream in chunks
vad = WhisperVADInference("model.onnx", streaming=True)
for audio_chunk in audio_stream:
speech_active = vad.process_chunk(audio_chunk)
if speech_active:
# Handle speech detection
pass
[1, 80, 3000] (batch size fixed to 1 - see note below)[1, 1500] (batch size fixed to 1)Note on Batch Processing: Currently, the ONNX model only supports batch size of 1 due to export limitations between PyTorch transformers and ONNX. However, single-sample inference is highly optimized and runs extremely fast (~100x real-time on CPU), making sequential processing still very efficient for most use cases.
model.onnx: ONNX model filemodel_metadata.json: Model configuration and parametersinference.py: Ready-to-use inference script with post-processingrequirements.txt: Python dependenciespip install onnxruntime # or onnxruntime-gpu for GPU support
pip install librosa transformers numpy
If you use this model, please cite:
@misc{whisper-vad,
title={Whisper-VAD: Whisper-based Voice Activity Detection},
author={Grider},
year={2025},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/TransWithAI/Whisper-Vad-EncDec-ASMR-onnx}}
}
MIT License
This model builds upon OpenAI's Whisper model and implements architectural refinements for efficient voice activity detection.