Downloads ยท 30 days
0
Nitinbudania/tiny-turn-detector
tiny-turn-detector is a audio classification model from Nitinbudania. Use it for the audio classification task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.
A lightweight real-time audio turn detection model that predicts whether a speaker is DONE speaking or PAUSING/CONTINUING in conversational audio.
Downloads ยท 30 days
0
Access
Public
Updated Aug 24, 2026
Repo size
217 MB
Likes
1
Public
Click a slice to open those files.
.pt217 MB ยท 100%
From the Hugging Face model README
A lightweight real-time audio turn detection model that predicts whether a speaker is DONE speaking or PAUSING/CONTINUING in conversational audio.
This model addresses a critical challenge in building responsive voice assistants and conversation systems: determining when a speaker has actually finished their turn versus just pausing mid-sentence.
Key Features:
Audio (8 sec, 16kHz)
โ
Whisper Tiny Encoder (frozen/fine-tuned)
โ
Mean Pooling
โ
MLP Head (384 โ 64 โ 1)
โ
Sigmoid โ P(end_turn)
โ
Binary Decision: END or CONTINUE
Components:
The model was trained on the pipecat-ai/smart-turn-data-v3.2-train dataset.
| Split | Loss | Accuracy | Precision | Recall | F1 Score |
|---|---|---|---|---|---|
| Train | 0.1391 | 95.50% | 95.90% | 95.34% | 95.62% |
| Val | 4.9075 | 73.00% | 70.00% | 89.09% | 78.40% |
Training Configuration:
Note: The validation loss is higher due to the model being optimized for F1 score rather than loss. The high recall (89%) indicates the model is conservative about marking turn endings, which is desirable for real-time applications to avoid premature interruptions.
from huggingface_hub import hf_hub_download
import torch
# Download model
model_path = hf_hub_download(
repo_id="YOUR_USERNAME/tiny-turn-detector",
filename="best_model.pt"
)
# Load model
model = torch.load(model_path, map_location='cpu')
model.eval()
import torch
import librosa
from transformers import WhisperProcessor
# Load processor
processor = WhisperProcessor.from_pretrained("openai/whisper-tiny")
# Load audio (8 seconds at 16kHz)
audio, sr = librosa.load("your_audio.wav", sr=16000, duration=8.0)
# Process audio
inputs = processor(audio, sampling_rate=16000, return_tensors="pt")
# Predict
with torch.no_grad():
outputs = model(inputs.input_features)
probability = torch.sigmoid(outputs).item()
# Decision
threshold = 0.5
decision = "END" if probability > threshold else "CONTINUE"
print(f"Probability: {probability:.4f}")
print(f"Decision: {decision}")
For complete inference code with audio loading, preprocessing, and visualization, see the GitHub repository.
pip install torch torchaudio transformers librosa huggingface_hub
Dataset: pipecat-ai/smart-turn-data-v3.2-train
The dataset contains conversational audio clips labeled with turn-taking information:
endpoint_bool: Binary label (0=continue, 1=end)Validation Gap: The model shows some overfitting (95.5% train vs 73% val accuracy). This could be improved with:
8-Second Window: Requires exactly 8 seconds of audio context
English Focus: Primarily trained on English conversations
VAD Dependency: Works best when combined with Voice Activity Detection (VAD) for silence removal
Full training code, evaluation scripts, and inference examples:
๐ https://github.com/Nitin1613/Turn_detector/tree/main
The repository includes:
If you use this model in your research or application, please cite:
@misc{tiny-turn-detector-2026,
title={Tiny Turn Detector: Real-time Audio Turn Detection with Whisper},
author=Nitinbudania,
year={2026},
publisher={Hugging Face},
howpublished={\url{https://huggingface.co/Nitinbudania/tiny-turn-detector}}
}
MIT License - See LICENSE file for details
Model Card Authors: YOUR_NAME
Contact: YOUR_EMAIL or GitHub
Last Updated: August 2026