Downloads · 30 days
172
29% of all-time downloads
dockhardman/gemma-4-E2B-duplex
gemma-4-E2B-duplex is a machine learning model from dockhardman. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
A frame-synchronous full-duplex speech LoRA for google/gemma-4-E2B-it. It listens and "thinks" on the same clock as the incoming audio, so turn-taking, barge-in (interrupting the model mid-reply), and multi-turn conve…
Downloads · 30 days
172
29% of all-time downloads
All-time downloads
585
Public
Repo size
96.9 MB
Likes
3
Public
Click a slice to open those files.
.safetensors96.7 MB · 100%
From the Hugging Face model README
A frame-synchronous full-duplex speech LoRA for google/gemma-4-E2B-it. It listens and
"thinks" on the same clock as the incoming audio, so turn-taking, barge-in (interrupting the
model mid-reply), and multi-turn conversation are handled natively by the model — not by an
external VAD or scripting layer.
The model emits text frame-by-frame; speech is produced by any downstream TTS. This keeps the adapter tiny (24M trainable params, ~96MB) while the base model does the heavy lifting.

Portfolio / technical-sharing release. Runnable demo web app: companion repo https://github.com/allen2c/gemma-4-E2B-duplex.
All in one adapter, bilingual (Chinese / English).
Base Gemma is turn-based: you send a whole message, it generates a whole reply, done. There is no notion of time — nothing happens "while you are still talking."
This adapter makes the same model frame-synchronous. Time is sliced into 80ms frames, and on every single frame the model does two things at once: it reads the latest audio and decides what to emit. Each frame is 2 audio tokens + 1 text token, and that one text token is a live decision:
time → frame 1 frame 2 frame 3 … frame k frame k+1 …
audio [you ...... speaking .............. pause] (silence)
model <wait> <wait> <wait> … <wait> "Sure," "here's"…
└─ stays quiet while you talk ─┘ └─ opens up on its own ─┘
<wait> = "keep listening." Emitting a word = "I'm speaking now."<wait> to words?<wait> within a few frames.Because every decision lives on a shared clock with the audio, these behaviors are the model's own, learned end-to-end — not bolted on by a separate turn-detector.
The LoRA sits only on the language model's attention/MLP projections; the audio encoder and embeddings stay frozen — we teach behavior, not perception.
The key finding behind this release: opening up, shutting up (on barge-in), and staying in a
multi-turn conversation are the same problem — the "act now" signal is a tiny fraction of the
supervised positions and gets drowned out by all the <wait> frames. Amplifying the training loss
on those few decisive positions cures all three at once, with plain supervised fine-tuning.
Bare model, all runtime heuristics off (no VAD, no nudge) — this is the model on its own:
| Capability | Metric | Result |
|---|---|---|
| Opening (standard) | replies after you finish | 11/12 |
| Opening (rambling questions) | " | 12/12 |
| Barge-in | chars still spoken after interrupt (median) | 7 |
| Multi-turn | 5-turn survival (turn 1 / overall) | 8/8 · 5/8 stable |
| Text channel | replies to typed input | 24/24 |
| Tool calling | correct call + spoken summary | 11/12 |
With a light runtime layer (energy-gated barge, silence-freeze), the live experience is smoother still — see the demo repo.
This is a PEFT LoRA adapter for google/gemma-4-E2B-it.
from peft import PeftModel
from transformers import Gemma4ForConditionalGeneration, AutoProcessor
base = Gemma4ForConditionalGeneration.from_pretrained("google/gemma-4-E2B-it")
model = PeftModel.from_pretrained(base, "dockhardman/gemma-4-E2B-duplex")
processor = AutoProcessor.from_pretrained("google/gemma-4-E2B-it")
The adapter runs frame-synchronously (80ms frames), so it needs a streaming driver to feed audio and drain text/tool events. A companion demo repo with a ready-to-run web app (microphone UI
gemma-4-E2B-it base).google/gemma-4-E2B-it@misc{gemma4e2bduplex2026,
author = {allen2c},
title = {Gemma-4-E2B-Duplex: a frame-synchronous full-duplex speech LoRA},
year = {2026},
howpublished = {\url{https://huggingface.co/dockhardman/gemma-4-E2B-duplex}}
}