Downloads · 30 days
0
NagaYu/rubato-timing
rubato-timing is a voice activity detection model from NagaYu. Use it for the voice activity detection task on the model card, and read the license before you ship it in a product. It is set up for rubato. The card lists the license as apache-2.0.
Predict the silence as a distribution, then decide when to speak.
Downloads · 30 days
0
Access
Public
Updated Aug 15, 2026
Repo size
347 KB
Likes
0
Public
Click a slice to open those files.
.json818 KB · 59%
From the Hugging Face model README
Predict the silence as a distribution, then decide when to speak.
A 3,137-parameter model that answers one question at 50 Hz: given everything heard so far, how much longer will this silence last? A dynamic program turns that distribution into a start/wait decision under an explicit, tunable asymmetry between talking over someone and answering late.
It decides when to talk. It does not decide what to say, and it contains no speech recogniser, no language model and no synthesiser. It is a middle layer you drop into an existing pipeline.
from rubato import load_pretrained, CostWeights
taker = load_pretrained(hf_repo="NagaYu/rubato-timing", weights=CostWeights.from_seconds_per_collision(2.0))
should_speak = taker.push_audio(chunk).should_speak # one 20 ms frame of float audio
seconds_per_collision is the whole configuration surface: how many seconds of
extra latency is one talk-over worth to you? Small values give an eager agent,
large values a patient one, and the sweep between them is the Pareto front below.
Framework adapters: rubato.integrations.pipecat_processor.RubatoTurnGate and
rubato.integrations.livekit_plugin.RubatoTurnDetector.
Source, benchmark protocol, ablations and sensitivity analyses: https://github.com/NagaYu/rubato

Both axes are minimised. A fixed threshold can only ever trace the outer curve -- one threshold buys one point. The inset is the band real systems ship in.
Held out on 5988 silences from speakers never seen in training (maptask, CC-BY-4.0).
| covers the fixed-threshold frontier | 86% of its points |
| covers the semantic-completeness frontier | 100% of its points |
| latency saved at matched talk-over | 84 ms vs fixed, 96 ms vs semantic |
| talk-over removed at matched latency | 2.8 pp vs fixed, 2.1 pp vs semantic |
| hazard calibration (ECE) | 0.0007 |
| CRPS skill over a covariate-free hazard | +0.173 |
| layer cost per 20 ms frame | 0.26 ms median, 0.49 ms p99 |
At the operating point matched to a 1000 ms threshold: latency 1000 → 772 ms (95 % CI 707–831), talk-over 6.3% → 6.5% (95 % CI 5.6%–7.6%). Intervals are a conversation-level cluster bootstrap.
h_k = P(partner resumes in frame k+1 | silent through k, evidence up to k), from a one-hidden-layer network over a causal
feature vector: a radial-basis expansion of log elapsed silence, turn-so-far
duration and pause count, transcript completeness (gated by ASR lag, because a
real recogniser has not delivered the last word yet), terminal prosody where
audio exists, acoustic precursor cues such as in-breaths, and a running
per-partner posterior.alpha * P(collision) + beta * latency. Being pre-empted --
the human carries on while the agent is still silent -- costs nothing, which
is what produces the human-like behaviour: when a resumption looks likely,
waiting is nearly free; when the floor is clearly open, the agent can start at
zero gap.| you provide | per frame |
|---|---|
| 20 ms of mono audio (any rate; 16 kHz assumed) | required |
| the ASR's partial transcript | optional, improves accuracy |
| a partner id | optional, enables entrainment |
| you get back | |
|---|---|
should_speak | the decision |
decision.planned_onset_s | when it currently intends to start |
decision.p_overlap_now | collision probability if it started this instant |
prediction.future_hazards | the full predicted silence distribution |
seconds_per_collision.Better timing makes an assistant less irritating. It also makes a synthetic voice harder to distinguish from a person, and that is a use this model is not for.
maptask (CC-BY-4.0). Anderson et al. (1991), The HCRC Map Task Corpus. Language and Speech 34(4). Annotations (c) 2007 HCRC, Univ. of Edinburgh & Univ. of Glasgow. CC BY 4.0. https://groups.inf.ed.ac.uk/maptask/
No audio was redistributed in building this model.
@software{rubato,
title = {Rubato: predicting silence distributions for spoken-dialogue turn-taking},
year = {2026},
url = {https://huggingface.co/NagaYu/rubato-timing}
}