Downloads · 30 days
4.6K
66% of all-time downloads
thepatch/t5gemma-b-b-ul2-GGUF
t5gemma-b-b-ul2-GGUF is a text-to-audio model from thepatch. Use it for the text-to-audio task on the model card, and read the license before you ship it in a product. The card lists the license as gemma.
The shared text encoder + tokenizer for sa3.cpp. GGUF conversion of the frozen google/t5gemma-b-b-ul2 encoder (encoder-only at inference) that Stable Audio 3 uses to embed text prompts.
Downloads · 30 days
4.6K
66% of all-time downloads
All-time downloads
7K
Public
Repo size
2 GB
Likes
1
Public
Click a slice to open those files.
.gguf2 GB · 100%
From the Hugging Face model README
The shared text encoder + tokenizer for sa3.cpp. GGUF conversion of the frozen google/t5gemma-b-b-ul2 encoder (encoder-only at inference) that Stable Audio 3 uses to embed text prompts.
This component is identical across all three SA3 variants (medium, small-music, small-sfx), so it lives in its own repo and is fetched once — the per-variant conditioner ships separately in each model repo. Validated against the PyTorch reference at cosine similarity ~1.0.
| component | file | size | notes |
|---|---|---|---|
| text encoder | t5gemma-b-b-ul2-encoder-0.3B-v1.0-F16.gguf | 537 MiB | default — equivalent to F32 |
| text encoder | t5gemma-b-b-ul2-encoder-0.3B-v1.0-F32.gguf | 1074 MiB | reference precision |
| text encoder | t5gemma-b-b-ul2-encoder-0.3B-v1.0-Q8_0.gguf | 285 MiB | smallest; a small, real tradeoff |
| tokenizer | t5gemma-b-b-ul2-v1.0-vocab.gguf | 14 MiB | Gemma byte-fallback BPE |
The encoder carries no conditioner — that ships per-variant in each model repo.
Measured against an F32 control, swapping only the encoder:
| encoder | conditioning cosine | generated-audio cosine |
|---|---|---|
| F16 | 0.999995 | 0.999848 |
| Q8_0 | 0.999547 | 0.999108 |
F16 is what sa3.cpp downloads and resolves by default, and it is a size win rather than a
fidelity substitution. Q8_0 is offered for tight-memory installs; it is a small but real
tradeoff, so it is never auto-selected — ask for it by name.
Q4_K_M is not published for this encoder, on evidence rather than caution: it is slower
than F32 here (the T5 forward is too small to amortize dequant, unlike the DiT) while dropping
generated-audio cosine to ~0.80, which is a different piece of music rather than a degraded one.
Q8_0 dominates it on every axis.
Select a precision with --t5-encoding f16|f32|q8_0, independently of the --encoding that picks
the DiT/SAME tier — the combination worth having on a small device is a quantized DiT with an F16
encoder. This works the same way at inference (sa3-generate) and at training (sa3-train).
Pair these with any SA3 variant repo's DiT + SAME + conditioner:
medium ·
small-music ·
small-sfx.
tools/download_models.py fetches this repo automatically alongside whichever variant you pick.
This is a format conversion of google/t5gemma-b-b-ul2, released under the Gemma Terms of Use (including the use restrictions in Section 3.2). Those terms carry over to this converted encoder + tokenizer.
A format conversion (weights → GGUF) for inference in sa3.cpp — no retraining. See sa3.cpp/docs/DISTRIBUTION.md.