Downloads · 30 days
0
interfaze-ai/diffusion-gemma-asr-small
diffusion-gemma-asr-small is a automatic speech recognition model from interfaze-ai. Use it when you need speech turned into text. It is set up for transformers. The card lists the license as apache-2.0.
Downloads · 30 days
0
Access
Public
Updated Jul 22, 2026
Repo size
191 MB
Likes
13
Trending 1
Click a slice to open those files.
.pt169 MB · 89%
From the Hugging Face model README
📝 Links: Blog · Demo Space · Code
Audio-native, multilingual speech recognition that transcribes through DiffusionGemma's own discrete-diffusion decoder — not autoregressive, not an external ASR decoder. Audio is projected directly into the Gemma embedding space, and the transcript is produced by parallel diffusion denoising (~8–16 steps), giving real-time-plus throughput where cost is set by the number of denoising steps, not the length of the transcript.
This repo ships the trained adapter only (projector + LoRA, ~42M params — 0.16% of the model). The frozen 26B DiffusionGemma backbone and the frozen whisper-small encoder load from their own repos.
raw audio ─► whisper-small encoder (frozen) ─► projector (trained, ~19M)
─► scatter into <audio> token slots of DiffusionGemma's encoder
─► DiffusionGemma decoder denoises a 192-token canvas (bidirectional, cross-attends audio)
─► transcript
google/diffusiongemma-26B-A4B-it — frozen, small LoRA adapters on encoder/decoder attention.openai/whisper-small encoder — frozen feature extractor (NOT a decoder).lm_head (the key unlock that makes the
audio embeddings transcript-predictive).Install
pip install torch peft soundfile librosa huggingface_hub \
"transformers @ git+https://github.com/huggingface/transformers.git" # DiffusionGemma support
Transcribe in Python
import sys, soundfile as sf
from huggingface_hub import snapshot_download
repo = snapshot_download("interfaze-ai/diffusion-gemma-asr-small") # this adapter (~170 MB)
sys.path.insert(0, repo)
from inference import load, transcribe # bundled in this repo
# Loads frozen DiffusionGemma-26B + whisper-small + this adapter (downloads bases on first run).
model, tok, fe = load(f"{repo}/diffusion_asr_small.pt", device="cuda")
wav, sr = sf.read("audio.wav") # 16 kHz mono float32 (inference.py resamples if needed)
print(transcribe(wav, model, tok, fe, max_steps=16))
Or from the command line
python inference.py audio.wav # run inside the downloaded repo dir
Long audio is split at silence (the encoder has a 30 s window, like Whisper). max_steps trades
speed for accuracy — 8 is near-best and fastest, 16 is the default.
Trained on FLEURS (6 languages) + LibriSpeech (en) + VoxPopuli (en/de/fr/es). WER/CER are Whisper-normalized (Open-ASR / Artificial-Analysis convention), 16 diffusion steps:
| benchmark | metric | score |
|---|---|---|
| LibriSpeech test-clean (en) | WER | 6.6% |
| FLEURS English | WER | 15.7% |
| VoxPopuli English | WER | 18.5% |
| FLEURS Hindi | CER | 15.8% |
| FLEURS Mandarin | CER | 29.6% |
Among diffusion / non-autoregressive ASR it leads (6.6% on LibriSpeech vs Whisfusion's 8.3%, with a smaller encoder). It trails autoregressive Whisper — a training-data gap (~219 h seen), not architecture.
diffusion_asr_small.pt — trained adapter ({"projector": ..., "lora": ...})model.py, audio.py — model definition (self-contained)inference.py — runnable example (load + segment + transcribe)requirements.txttransformers from main (DiffusionGemma support) + torch, peft.google/diffusiongemma-26B-A4B-it (Gemma terms) and openai/whisper-small (MIT).@misc{khurdula2026audionativespeechrecognitionfrozen,
title={Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model},
author={Harsha Vardhan Khurdula and Abhinav Kumar Singh and Yoeven D Khemlani and Vineet Agarwal},
year={2026},
eprint={2607.13013},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2607.13013},
}