Downloads · 30 days
0
gabar-tech/chatterbox-amharic
chatterbox-amharic is a text-to-speech model from gabar-tech. Use it when you need text read aloud. It is set up for peft. The card lists the license as cc-by-sa-4.0.
A LoRA adapter and an extended Fidel tokenizer that teach Chatterbox Multilingual v3 (Resemble AI, MIT) to speak Amharic, with voice cloning from about ten seconds of reference audio. Trained only on speech we own or…
Downloads · 30 days
0
Access
Public
Updated Aug 23, 2026
Repo size
214 MB
Likes
0
Public
Click a slice to open those files.
.safetensors203 MB · 95%
From the Hugging Face model README
A LoRA adapter and an extended Fidel tokenizer that teach Chatterbox Multilingual v3 (Resemble AI, MIT) to speak Amharic, with voice cloning from about ten seconds of reference audio. Trained only on speech we own or that is licensed for it.
Stock Chatterbox cannot read Amharic at all: its tokenizer maps every Ge'ez
character to [UNK]. So the "before" clips below aren't a weaker version of
the same thing; they're the model guessing at unknown tokens. We add 244
tokens for the script and teach the model what they sound like.
The repo also has amharic_text.py, the text normalizer
the model was trained through. No dependencies, works on its own
(below).
Each pair uses the same sentence, reference audio, settings and seed. No language tag on either: the adapter was trained without one, and the stock tokenizer has no Ge'ez characters.
| # | Test case | Stock Chatterbox v3 | + Gabar adapter |
|---|---|---|---|
| 1 | Ordinary prose | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/01_stock.wav"></audio> | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/01_finetuned.wav"></audio> |
| 2 | Prose, ፥ punctuation | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/02_stock.wav"></audio> | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/02_finetuned.wav"></audio> |
| 3 | Prose (ejectives ቡ/ጅ) | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/03_stock.wav"></audio> | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/03_finetuned.wav"></audio> |
| 4 | Numbers + ዓ.ም. date abbreviation | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/04_stock.wav"></audio> | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/04_finetuned.wav"></audio> |
| 5 | ዶ/ር title abbreviation | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/05_stock.wav"></audio> | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/05_finetuned.wav"></audio> |
| 6 | Question intonation | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/06_stock.wav"></audio> | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/06_finetuned.wav"></audio> |
| 7 | Technical prose | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/07_stock.wav"></audio> | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/07_finetuned.wav"></audio> |
| 8 | Mixed punctuation + question | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/08_stock.wav"></audio> | <audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/08_finetuned.wav"></audio> |
Reading this on GitHub? The players only render on Hugging Face. Click a
clip in demo/ to play it, or watch
demo/before_after.mp4 (all eight pairs, 1:47).
The weights (new_lang_adapter/, 194 MB) are only on
Hugging Face;
everything else is mirrored here.
Texts: demo/sentences.txt. Reference voice:
demo/reference.wav, one of us
(<audio controls preload="none" src="https://huggingface.co/gabar-tech/chatterbox-amharic/resolve/main/demo/reference.wav"></audio>). Both models got the text after
amharic_text.normalize, temperature=0.6, cfg_weight=0.5, seed 1234.
One take per sentence per model, no picking. Known weaknesses are under
Limitations.
An adapter: LoRA weights on the T3 text-to-speech-token transformer, full-rank embeddings for the new tokens, and the extended tokenizer. You apply it on top of Chatterbox Multilingual v3, which you download from Resemble. We ship our delta, not a copy of their model.
Not merged weights, not a standalone model. Amharic only: Tigrinya and Ge'ez use the same script, but the adapter was not trained on them.
amharic_text.py built the training labels, and the
loader runs it on every input, so training and inference see the same text.
One file, standard library only, same licence as the adapter, usable
without the model. (One fix since training: a dotted abbreviation's trailing
dot mid-sentence, as in … ዓ.ም. የአገሪቱ …, is no longer read as a full stop.
Labels were built with the version whose SHA-256 is in
training_config.json; the only effect on the model is one fewer spurious
pause.)
It converts Ge'ez numerals, digits, decimals, percentages and clock times to
words (1500 → አንድ ሺህ አምስት መቶ, 75% → ሰባ አምስት በመቶ, 3:30 → ሶስት ሰዓት
ተኩል); expands about a hundred common abbreviations, keeping the inflected
suffix (ዶ/ር → ዶክተር, ዓ.ም. → ዓመተ ምሕረት, መ/ቤቱ → መሥሪያ ቤቱ); collapses the
four consonant families that are spelled several ways but pronounced the same
(ሐ ኀ ኅ → ሀ, ሠ → ሰ, ዐ → አ, ፀ → ጸ); reduces punctuation to the five marks that
change how you say something (። ፣ ፤ ? !); strips URLs, emoji and
control characters. It's a subset of what we run in production.
from amharic_text import normalize, split_sentences
normalize("ዶ/ር አበበ በ2018 ዓ.ም በተደረገው ምርጫ 75% ድምፅ አገኙ።")
# 'ዶክተር አበበ በሁለት ሺህ አስራ ስምንት ኣመተ ምህረት በተደረገው ምርጫ ሰባ አምስት በመቶ ድምጽ አገኙ።'
Three sources. Every clip's filename starts with its corpus prefix; the
assembled training directory was audited before training and the output is
committed as is (audit/corpus_audit.txt). The
adapter was trained from scratch on exactly that directory, starting from
Resemble's stock v3 T3.
| prefix | source | licence | clips | hours |
|---|---|---|---|---|
ih_ | Our own studio recordings | ours | 573 | 1.25 |
wxl_ | WaxalNLP Amharic (Digital Umuganda / Google) | CC-BY-SA-4.0 | 40921 | 190.93 |
cv_ | Common Voice Amharic | CC0-1.0 | 1055 | 1.45 |
| total | 42549 | 193.63 |
WaxalNLP comes as 48 kHz and Common Voice as 32/48 kHz MP3; both were resampled to 24 kHz. All Waxal Amharic clips that fit the trainer's 3 to 25 second window were used (190.93 h of roughly 191).
WaxalNLP is licensed under CC BY-SA 4.0, so the adapter is released under CC BY-SA 4.0. Credit to Digital Umuganda, the WaxalNLP contributors, and the Common Voice contributors. The base model is MIT; this licence covers what we add.
Base: Chatterbox Multilingual v3, ResembleAI/chatterbox at revision
5bb1f6ee58e50c3b8d408bc82a6d3740c2db6e18, T3 file t3_mtl23ls_v3.safetensors. Pinned on purpose:
the adapter only makes sense on that exact T3. v3 has Resemble's
hallucination and speaker-similarity fixes over v2; S3Gen, the voice encoder
and the tokenizer are the same as v2 and untouched.
| component | treatment |
|---|---|
| T3 (text→speech-token transformer, 0.5 B) | LoRA r=64, α=128, dropout 0.05 on q_proj k_proj v_proj o_proj gate_proj up_proj down_proj + spkr_enc; base weights frozen |
text_emb / text_head | trained full-rank and shipped whole (PEFT modules_to_save). New vocabulary rows can't be learned through a low-rank delta. |
| Tokenizer | base multilingual tokenizer + 244 added tokens: every Ge'ez character seen in the corpus, plus ። ፣ ፤ ? ! and U+135F. "ሰላም" in the stock tokenizer is [UNK] [UNK] [UNK]. |
| S3Gen (speech tokens → waveform, includes the PerTh watermark) | frozen, not shipped |
| Voice encoder | frozen, not shipped |
Notes for anyone building on this:
No language token. The base tokenizer has [fr], [de] and so on; there
is no [am] and we didn't add one. The adapter was trained on plain
normalized text, and the loader tokenizes without a language prefix, without
lower-casing or NFKD. language_id="am" on stock
ChatterboxMultilingualTTS.generate raises ValueError; use the loader.
Alignment guard is off. Upstream enables its attention-alignment
hallucination guard only when text_tokens_dict_size == 2454. With the
extended vocabulary it's off, in training and at inference. The loader chunks
by sentence instead, which handles the common failure (T3 stopping at the
first sentence-final mark).
Train/inference parity. The shipped amharic_text.py differs from the file
that built the training labels by later punctuation fixes. Both hashes and the
reason are recorded in training_config.json as
label_frontend_sha256, label_frontend_sha256_shipped and
label_frontend_changed_since_training. Property checks over every training
transcript are in audit/frontend_check.txt.
Held-out set: 100 clips from the same corpus, split before training
by a seeded speaker-disjoint rule, so whole speakers are held out and none of
their sentences appear in training. Checked independently of the trainer's
own assertion: HELD OUT: eval ∩ train = ∅ at clip, speaker and sentence level; all eval stems are ih_/wxl_/cv_.
(audit/holdout_verify.txt). Both models ran on
the same clips with the same per-clip reference audio (the held-out
speaker's own recording), through the same code: the released loader for the
adapter, stock v3 with the same normalized text and no language tag.
| metric | stock Chatterbox v3 | + Gabar adapter |
|---|---|---|
| Amharic CER ↓ (Meta omniASR-CTC-3B) | 0.932 | 0.095 |
| UTMOS ↑ (naturalness MOS predictor) | 2.359 | 2.711 |
| ECAPA cosine ↑ (speaker similarity to reference) | 0.610 | 0.860 |
| generation failures (empty / <0.5 s / error) | 0.0% | 1.0% |
Per-clip numbers: audit/eval/.
How CER is measured. We transcribe the generated audio with Meta's stock
omniASR-CTC-3B, which we
didn't train, and compare to the reference text after normalize_for_metric
(collapses the homophone families so ሀ/ሐ/ኀ spellings don't count as errors,
strips punctuation). Same ASR, same normalization, both models. The stock
column is a floor: the base model can't read Fidel, so most of what it
produces isn't Amharic. UTMOS and ECAPA involve no ASR. UTMOS was trained on
English MOS ratings and both outputs are Amharic, so its near-tie says more
about the metric than the models. Means are over clips that produced audio;
failures are on their own row so they can't hide in an average.
Our own listening verdict: intelligible Amharic, not yet fully natural (read-aloud cadence, occasional flat prosody); the stock output is garbled and not Amharic.
pip install "chatterbox-tts @ git+https://github.com/resemble-ai/chatterbox@5de7a54aa4e5e2baadb0182dde554908b48b85c2" peft safetensors huggingface_hub torchaudio
from huggingface_hub import hf_hub_download
import importlib.util, torchaudio
# the loader + text front-end ship in this repo
spec = importlib.util.spec_from_file_location(
"amharic_tts", hf_hub_download("gabar-tech/chatterbox-amharic", "amharic_tts.py"))
amharic_tts = importlib.util.module_from_spec(spec); spec.loader.exec_module(amharic_tts)
tts = amharic_tts.load_amharic_tts(device="cuda") # downloads base v3 (pinned) + adapter
wav = tts.generate(
"ሰላም! ይህ ከጽሑፍ በቀጥታ የተፈጠረ የአማርኛ ድምፅ ነው። ዛሬ ነሐሴ 11 ቀን 2018 ዓ.ም. ነው።",
audio_prompt_path="reference.wav", # ~10 s of the voice to clone, with consent
temperature=0.6, cfg_weight=0.5)
torchaudio.save("out.wav", wav, tts.sr) # 24 kHz, PerTh-watermarked
print(tts.normalize("ዛሬ ነሐሴ 11 ቀን 2018 ዓ.ም. ነው።")) # what the model actually read
Or from a checkout: python amharic_tts.py "ሰላም ዓለም።" --ref reference.wav --out out.wav.
generate() normalizes the text, splits at sentence-final marks (T3 tends to
stop at the first ። / ? / !), synthesizes each sentence against the
reference and joins them. normalize=False / split_sentences=False turn
those off. We ran the snippet above as written in a fresh virtualenv with
only the packages listed, on NVIDIA RTX A6000, Ubuntu 22.04.5 LTS, python 3.11.10, torch 2.6.0+cu124, chatterbox-tts@5de7a54, peft 0.20.0, before publishing.
Chatterbox puts Resemble's PerTh watermark in every waveform it generates.
We left that alone; nothing in this adapter touches S3Gen or the vocoder,
where it happens. We ran the public resemble-perth detector over every demo
clip from both models and the held-out eval outputs, with the natural
reference recording as a negative control: present on all 115 checked files (detector confidence ≥ 0.5 on every generated file; the natural reference recording scores 0.0, so the detector is discriminating)
(audit/watermark_verify.txt). If you build on
this, leave it in. It's about the only provenance signal that exists for
synthetic Amharic right now.
amharic_text.py. Skip it, or feed it things it doesn't handle (currency
symbols, odd date formats, Latin words), and the model gets them raw.This model can clone a voice from roughly ten seconds of reference audio, so it can be misused for impersonation or fraud. Use it only with informed consent and human review. Outputs are watermarked by the shipped inference path. Every voice in the training data was recorded under a licence that permits this use.
@misc{gabar2026chatterboxamharic,
title = {Chatterbox Amharic: a Fidel extension of Chatterbox Multilingual},
author = {{Gabar Technologies}},
year = {2026},
url = {https://huggingface.co/gabar-tech/chatterbox-amharic}
}