Downloads · 30 days
0
Howest-AI-Lab/babbeldoos-flemish-tts
babbeldoos-flemish-tts is a text-to-speech model from Howest-AI-Lab. Use it when you need text read aloud. It is set up for chatterbox. The card lists the license as cc-by-nc-4.0.
Type Flemish, pick a voice or record your own, hear it back. The demo also plays pre-rendered examples that cost no GPU time, so you can hear what the adapters sound like before spending any.
Downloads · 30 days
0
Access
Public
Updated Sep 10, 2026
Repo size
181 MB
Likes
4
Public
Click a slice to open those files.
.safetensors181 MB · 100%
From the Hugging Face model README
Type Flemish, pick a voice or record your own, hear it back. The demo also plays pre-rendered examples that cost no GPU time, so you can hear what the adapters sound like before spending any.
Two LoRA adapters that shift Chatterbox Multilingual from Netherlandic Dutch towards Flemish (Belgian Dutch). Voice identity stays an inference-time argument — you supply a reference clip. These adapters change accent and register, not who is speaking.
| adapter | register | trained on | size |
|---|---|---|---|
Conversational | spontaneous Flemish | 18.6 h Flemish broadcast speech | 22.6 M params |
Narration | polite Belgian Dutch, read aloud | 11.4 h Flemish broadcast speech | 22.6 M params |
Base model weights are not included or modified — the adapters are a 90 MB
delta on t3.tfmr and Chatterbox is downloaded separately at load time.
Conversational reads the register off the text. Give it spontaneous Flemish and
it drops word-final -n and -t the way many Flemish speakers do; give it a
formal script and it tightens up on its own. Measured word-final consonant
retention: 0.818 on conversational text against 0.948 on narration text —
the same weights, a 0.13 swing.
Narration keeps those consonants in both registers (0.915 / 0.979) and is the
more stable of the two across reference voices. Use it for anything read aloud.
from babbeldoos import Babbeldoos
tts = Babbeldoos.from_pretrained("Howest-AI-Lab/babbeldoos-flemish-tts")
wav = tts.generate("Goeiemiddag, waarmee kan ik u vandaag helpen?",
reference="your_voice.wav", model="Narration")
tts.save(wav, "out.wav")
Adapter names are case-insensitive — "Narration", "narration" and
"NARRATION" all resolve to the same adapter. The repo directories are
lowercase (conversational/, narration/); nothing depends on how you spell
the name at the call site.
babbeldoos.py is vendored in this repo — Chatterbox's T3 has no adapter API, so
the adapters attach by wrapping nn.Linear.forward on the modules they were
trained for. Copy the file next to your code, or hf download the repo and add
it to sys.path.
You must supply one, and none ships with this repo — see Reference voices below. 5 s of clean speech is enough. Chatterbox reads only the first 6 s for the T3 prompt and 10 s for the vocoder, so anything past ~10 s is ignored: a clean 8-second recording beats a noisy 30-second one.
Chatterbox stops at about 1000 speech tokens (~40 s) and these adapters saw
nothing over 15 s in training. read() splits at sentence boundaries, generates
each piece against the same reference and joins them:
wav = tts.read(open("article.txt").read(), reference="your_voice.wav",
model="Narration")
tts.save(wav, "article.wav")
A LoRA is an additive low-rank delta, so scaling it walks the straight line between base and adapter — nothing retrained, nothing approximated:
tts.generate(text, reference="v.wav", model="Conversational", strength=0.6)
strength=0.0 reproduces the base model bit for bit; 1.0 is the adapter as
trained. Lower values trade Flemish accent for crisper final consonants
(retention rises monotonically: 0.874 at 1.0 → 0.946 at 0.5 → 0.966 at base).
We ship 1.0 because that is what listeners preferred. At 0.4 and 0.7 the Netherlandic accent starts coming back — the thing these adapters exist to remove — and the silences get less stable. The metric prefers low strength; the ear does not.
The deltas also compose, since they add:
tts.attach("Conversational", 0.7)
tts.add("Narration", 0.5) # W + 0.7·d1 + 0.5·d2
Fitted independently, so a blend is worth listening to rather than assuming.
No reference clips are included in this repo, deliberately. The voices used during development were either synthesised by a commercial service, whose terms we will not extend to redistribution, or real broadcast speakers who never consented to having their voice cloned by strangers. Record five seconds of your own voice, or of someone who agreed to it.
The demo Space offers a handful of voices so you can try the adapters without recording anything first; it also takes an upload or a microphone recording.
Reference choice matters more than you would expect: across ten references on the same adapter, measured articulation spanned 0.084–0.094 — a bigger spread than between the two adapters on the same voice. It is worth trying several, which is quickest in the demo.
LoRA r=32, α=64, dropout 0.05, on t3.tfmr q/k/v/o/gate/up/down projections
(210 layers, 22.6 M parameters). 400 optimizer steps, batch 1 × grad-accum 8,
lr 1e-5, fp32 master weights. Audio is Flemish broadcast speech, segmented at
sentence boundaries, transcribed with
NeLF — a Flemish ASR, because
general-purpose models normalise the Flemish away (Whisper rewrites
goeiemiddag → goedemiddag and da's → dat is, which teaches the TTS model
to do the same). Loudness-normalised, gender-balanced, language-filtered.
The corpus is not published.
Four things bit us getting the demo Space to run. All four produce error messages that point somewhere other than the cause, so they are written out here rather than left as folklore. A working configuration is in the demo Space.
Build the model at startup, not inside the GPU function. Constructing
Chatterbox pulls ~2 GB of base weights, and the wall clock inside a
@spaces.GPU(duration=...) call is budgeted — a first request that also has to
download the model gets killed. Load on CPU while the app boots, move in per
request:
tts = Babbeldoos.from_pretrained(REPO, device="cpu") # at import
@spaces.GPU(duration=120)
def speak(...):
tts.to("cuda") # ~0.5 s; the adapters move with it
return tts.generate(...)
to() re-attaches the adapter branches afterwards so their buffers follow the
model. Moving back to CPU works too.
Install chatterbox-tts in a separate pip pass, and let a newer torch win.
It pins torch==2.6.0, whose wheels carry kernels only up to sm_90. ZeroGPU now
runs Blackwell (sm_120), so a generation dies with "no kernel image is
available for execution on the device". sm_120 needs CUDA 12.8, first in
torch 2.8's default wheel. Put chatterbox-tts alone in
pre-requirements.txt and torch==2.8.0 / torchaudio==2.8.0 in
requirements.txt: chatterbox is not part of the second resolve, so its pin
becomes a warning rather than a constraint. pip's "requires torch==2.6.0, but
you have 2.8.0" line is expected — these adapters are developed against
torch 2.13.
Keep sdk_version equal to your gradio pin. Spaces installs
gradio[oauth,mcp]==<sdk_version> from the Space README, and every
chatterbox-tts release pins gradio with ==. If the two disagree, pip
backtracks through older chatterbox releases and lands on 0.1.4, the last one
depending on the unmaintained pkuseg — which ships no wheels and fails to
build with "ModuleNotFoundError: No module named 'numpy'". The numpy error is
a symptom of the gradio mismatch.
Pin setuptools<81. resemble-perth, which Chatterbox uses for its
watermark, imports pkg_resources at module load. Newer setuptools removed it;
perth catches the ImportError and sets its watermarker to None, and
Chatterbox then calls None() — "TypeError: 'NoneType' object is not
callable" while building the model.
TTS comes out as "bijt", T T S as
"thees". Expand them in the text before synthesis (TTS → text-to-speech).read() does the splitting and joining.cc-by-nc-4.0. The base model (Chatterbox, Resemble AI) is MIT, but the
transcription model used to build the training corpus is CC-BY-NC, so
non-commercial is the consistent choice downstream.
Built for the PWO Physical AI project at Howest (Kortrijk, Belgium). Version 1 — a first step, not an endpoint.
@misc{babbeldoos2026,
title = {Babbeldoos: Flemish LoRA adapters for Chatterbox},
author = {Howest AI Lab},
year = {2026},
note = {https://huggingface.co/Howest-AI-Lab/babbeldoos-flemish-tts}
}