Downloads · 30 days
0
Baiji-Team/TurnSense
TurnSense is a machine learning model from Baiji-Team. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Apr 23, 2026
Repo size
255 MB
Likes
1
Public
Click a slice to open those files.
.onnx253 MB · 99%
From the Hugging Face model README
Language: English | 中文
<br/><br/>⭐ If TurnSense is useful to you, please give us a Star! It helps us keep improving the model and documentation.
| Dimension | TurnSense Performance |
|---|---|
| 🎯 Accuracy | F1 96.35% (easyturn_real_test_ZH) — best in class |
| ⚡ Inference Latency | CPU p50 ≈ 54.65ms — real-time interaction ready |
| 📦 Model Size | Only 47M parameters, INT8 version only ~50MB |
| 🧠 Classification | First open-source model natively supporting complete / incomplete / invalid three-class detection |
| 🚫 Invalid Filtering | Invalid utterance F1 reaches 94.34%, effectively suppressing noise-triggered responses |
| 🤗 Open-Source Friendly | FP32 / INT8 ONNX provided, ready to use out of the box |
TurnSense is a three-class semantic detection model designed for human-machine voice interaction, focused on solving a critical problem in dialogue systems:
During a user's speech, should the system respond immediately, or continue waiting?
Traditional approaches typically rely on a simple binary classification — "finished or not." TurnSense goes further by simultaneously modeling semantic completeness and invalid input detection, enabling more natural turn-taking in complex real-world scenarios and significantly reducing false interruptions, premature responses, and noise-triggered activations.
<div align="center"> <img src="./TenSense/image/TurnSense.svg" alt="TurnSense Three-Class Illustration" width="820"/> </div> <br/>TurnSense classifies user input into three semantic states:
| State | Description | Example |
|---|---|---|
| ✅ Complete | The user has expressed a complete intent; the system can respond | "Check tomorrow's weather in Shanghai for me." |
| ⏳ Incomplete | The user's expression is unfinished — truncated, paused, or trailing off | "I'd like to ask about that order from yesterday..." |
| 🔇 Invalid | The input does not constitute meaningful speech and should not trigger a response | "...(continuous noise / non-verbal vocalization)" |
These three labels enable the system to determine not only "should I respond?" but also "is it worth responding to?" — significantly improving interaction naturalness and system stability in voice assistants, real-time calls, intelligent customer service, and more.
<br/>Simultaneously models complete / incomplete / invalid states — closer to real conversational behavior than traditional binary classification, and currently the only open-source solution with native invalid utterance detection.
Only 47M parameters (INT8 version ~50MB). CPU inference latency: p50 ≈ 54.65ms, p90 ≈ 58.00ms — meets the strict requirements of real-time interaction without a GPU.
Achieves F1 96.35% (complete) and F1 96.32% (incomplete) on easyturn_real_test_ZH (300 samples), and F1 92.30% (complete) and F1 91.62% (incomplete) on semantic_test_ZH (2000 samples) — best or runner-up among all comparable models.
On the NonverbalVocalization test set, invalid utterance precision reaches 100% with recall of 90.37% (F1 = 94.34%), effectively suppressing false triggers from non-verbal sounds and noise.
Balances precision and recall in semantically ambiguous, pause-heavy, or colloquial scenarios, reducing both premature responses and missed responses.
Ships with a complete evaluation pipeline and scripts, supporting unified metric comparison and performance regression analysis for full reproducibility.
Standardized repository structure with FP32 / INT8 ONNX models — from installation to inference in just a few minutes.
<br/>| Model | Parameters | Three-Class | Link |
|---|---|---|---|
| TEN-Turn | 7B | ❌ | TEN-framework/TEN_Turn_Detection |
| Easy-Turn | 850M | ❌ | ASLP-lab/Easy-Turn |
| NAMO-Turn-Detector (ZH) | 66M | ❌ | videosdk-live/Namo-Turn-Detector-v1-Multilingual |
| ⭐ TurnSense | 47M | ✅ | Baiji-Team/TurnSense |
| Smart-Turn-v3 | 8M | ❌ | pipecat-ai/smart-turn-v3 |
| FireRedChat-turn-detector | -- | ❌ | FireRedTeam/FireRedChat-turn-detector |
<br/>💡 With only 47M parameters, TurnSense achieves three-class capability — the best balance between accuracy and model size.
<br/>All results below are based on open-source Chinese evaluation sets. Latency marked with
(GPU)indicates GPU environment; otherwise, latency was measured on CPU.
Data source: Real data samples from Easy-Turn-Testset
| Model | P (complete) | R (complete) | F1 (complete) | P (incomplete) | R (incomplete) | F1 (incomplete) | p50 Latency | p90 Latency |
|---|---|---|---|---|---|---|---|---|
| Easy-Turn | 97.26% | 94.67% | 95.95% | 94.81% | 97.33% | 96.05% | 183.87 (GPU) | 300.37 (GPU) |
| Smart-Turn-v3 | 64.97% | 76.67% | 70.34% | 71.54% | 58.67% | 64.47% | 36.84 | 39.10 |
| TEN-Turn | 99.25% | 88.00% | 93.29% | 89.22% | 99.33% | 94.01% | 17.66 (GPU) | 19.41 (GPU) |
| FireRedChat | 70.65% | 94.67% | 80.91% | 91.92% | 60.67% | 73.09% | 98.30 | 99.42 |
| NAMO-Turn | 81.53% | 85.33% | 83.39% | 84.62% | 80.67% | 82.59% | 3.60 | 83.44 |
| ⭐ TurnSense | 96.03% | 96.67% | 🏆 96.35% | 96.64% | 96.00% | 🏆 96.32% | 54.65 | 58.00 |
<br/>🔍 Key Finding: TurnSense achieves the highest F1 on both complete and incomplete classes, and is the only model with CPU p50 < 60ms while maintaining F1 > 96%.
Data source: Chinese test split from KE-Team/SemanticVAD-Dataset
| Model | P (complete) | R (complete) | F1 (complete) | P (incomplete) | R (incomplete) | F1 (incomplete) | p50 Latency | p90 Latency |
|---|---|---|---|---|---|---|---|---|
| Easy-Turn | 78.14% | 98.30% | 87.07% | 97.64% | 70.30% | 81.74% | 183.87 (GPU) | 300.37 (GPU) |
| Smart-Turn-v3 | 59.25% | 88.10% | 70.85% | 76.80% | 39.40% | 52.08% | 36.84 | 39.10 |
| TEN-Turn | 85.25% | 99.60% | 91.87% | 99.52% | 82.70% | 90.33% | 17.66 (GPU) | 19.41 (GPU) |
| FireRedChat | 66.76% | 99.40% | 79.87% | 98.83% | 50.50% | 66.84% | 98.30 | 99.42 |
| NAMO-Turn | 71.48% | 86.70% | 78.36% | 83.10% | 65.40% | 73.20% | 3.60 | 83.44 |
| ⭐ TurnSense | 88.96% | 95.90% | 🏆 92.30% | 95.55% | 88.00% | 🏆 91.62% | 54.65 | 58.00 |
<br/>🔍 Key Finding: On the larger 2000-sample test set, TurnSense still maintains the best F1, demonstrating strong generalization capability.
Data source: OpenSLR Deeply Nonverbal Vocalization Dataset (SLR99)
| Model | P (invalid) | R (invalid) | F1 (invalid) |
|---|---|---|---|
| ⭐ TurnSense | 100.00% | 90.37% | 🏆 94.34% |
<br/>🔍 Key Finding: TurnSense is currently the only model that supports invalid utterance detection. A precision of 100% means zero false positives — effectively preventing noise from triggering system responses.
git clone https://github.com/Baiji-Team/TurnSense.git
cd TurnSense
pip install -U numpy onnxruntime torch librosa soundfile pandas scikit-learn huggingface_hub
TurnSense model weights are available on Hugging Face: Baiji-Team/TurnSense
| Version | Size | Use Case |
|---|---|---|
| FP32 | ~191 MB | Accuracy-first |
| INT8 | ~50 MB | Deployment-first (recommended) |
Download Options:
Option 1: Auto-download (Recommended) The inference script includes built-in Hugging Face download logic. The model will be automatically fetched and cached on first run.
Option 2: Git LFS
git lfs install
git clone https://huggingface.co/Baiji-Team/TurnSense
Option 3: Hugging Face Hub
from huggingface_hub import snapshot_download
snapshot_download(repo_id="Baiji-Team/TurnSense")
python infer.py
Example output:
Loading model from Baiji-Team/TurnSense...
Running inference on: "我想问一下那个订单就是昨天..."
Results:
Input: "我想问一下那个订单就是昨天..."
TurnSense Detection Result: "incomplete"
<br/>
.jsonl test dataset (line-by-line JSONL)warmup_iters=20)Output files include:
| File | Description |
|---|---|
report.md | Summary evaluation report |
results.json | Structured evaluation results |
config.json | Evaluation configuration |
per_sample__*.jsonl | Per-sample prediction details |
Each line is a JSON object containing at least the following fields:
| Field | Description |
|---|---|
audio_path | Path to the audio file |
text | Text content |
label | Label (complete / incomplete / invalid) |
Example:
{"audio_path":"/001.wav","text":"帮我查一下明天上海天气","label":"complete"}
{"audio_path":"/002.wav","text":"我想问一下那个订单就是昨天...","label":"incomplete"}
{"audio_path":"/003.wav","text":"啊…嗯…(持续噪声)","label":"invalid"}
python TurnSense/Turn_benchmark/benchmark.py
<br/>
If you use TurnSense in your research or product, please cite:
@misc{turnsense2026,
author = {Baiji Team},
title = {TurnSense: A Three-Class Semantic Detection Model for Complete, Incomplete, and Invalid Utterances},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Baiji-Team/TurnSense}},
}
<br/>
If you have questions or suggestions, feel free to reach out:
| Channel | Contact |
|---|---|
| [email protected] · [email protected] · [email protected] | |
| h2538406363 | |
| 👥 WeChat Group | Scan the QR code to join the group<br><img src="TenSense/image/wechat.jpg" alt="WeChat group QR code" width="220" /> |
| 🐛 Issues | GitHub Issues |
| 🔀 PR | Pull Requests |
This project is released under the Apache License 2.0 with certain additional conditions. See LICENSE for details.
<br/>Built with ❤️ by Baiji Team
</div>