Downloads · 30 days
0
bschooled/caeleste-speech
caeleste-speech is a machine learning model from bschooled. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
4-bit NF4 (bitsandbytes) quantizations of the open models used by a self-hosted speech stack (speech-to-text, text-to-speech, and a small text LLM).
Downloads · 30 days
0
Access
Public
Updated Aug 19, 2026
Repo size
3.3 GB
Likes
0
Public
Click a slice to open those files.
.safetensors3.2 GB · 99%
From the Hugging Face model README
4-bit NF4 (bitsandbytes) quantizations of the open models used by a self-hosted speech stack (speech-to-text, text-to-speech, and a small text LLM).
Each model lives in its own repository (listed below) and is a standalone,
standard transformers checkpoint. Nothing here is a new model: these are
quantized copies of the upstream repositories linked below, at the exact pinned
revisions shown.
Quantizing at load time means every deployment downloads full-precision weights and repeats the quantization pass on every model load. Publishing the quantized weights removes both costs. Measured on an RTX 5080:
| Model | Upstream size | Here | Load (runtime quant) | Load (prequantized) |
|---|---|---|---|---|
MOSS-TTS-v1.5 (8B) | 15.83 GiB | 6.23 GiB | does not fit 16 GB | fits |
higgs-audio-v3-stt | 5.00 GiB | 2.61 GiB | 9.2 s | 0.8 s |
Qwen2.5-0.5B-Instruct | 0.93 GiB | 0.44 GiB | 1.2 s | 0.2 s |
MOSS-TTS is the clearest case: quantizing it at load time needs more VRAM than the quantized model itself occupies, so a 16 GB GPU cannot produce it locally but can run it once quantized.
caeleste-speech-higgs-audio-v3-stt-nf4 — Speech-to-text. Quantization of bosonai/higgs-audio-v3-stt at db4966839bef. Whisper-large-v3 encoder retained in bf16; Qwen3 decoder quantized to NF4.caeleste-speech-moss-tts-v1.5-nf4 — Quality text-to-speech. Quantization of OpenMOSS-Team/MOSS-TTS-v1.5 at cdd3b911b158. Fully NF4. Quantized on an AMD gfx1030; the artifact is hardware-independent.caeleste-speech-qwen2.5-0.5b-instruct-nf4 — Text LLM. Quantization of Qwen/Qwen2.5-0.5B-Instruct at 7ae557604adf. Fully NF4.Fetch the subfolder you need, then load it as a local checkpoint:
from huggingface_hub import snapshot_download
from transformers import AutoModel
path = snapshot_download("bschooled/caeleste-speech", allow_patterns="higgs-audio-v3-stt-nf4/*")
model = AutoModel.from_pretrained(
f"{path}/higgs-audio-v3-stt-nf4",
device_map={"": "cuda:0"},
trust_remote_code=True,
)
allow_patterns keeps the download to only the model you want.
Note: for models that use
trust_remote_code, load from a downloaded local directory as shown above rather than passingsubfolder=tofrom_pretrained.transformersresolves remote-code modules relative to the repository root, sosubfolder=cannot find them.
Models without remote code load directly:
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
"bschooled/caeleste-speech", subfolder="qwen2.5-0.5b-instruct-nf4", device_map={"": "cuda:0"},
)
Each subfolder also contains its own README.md with per-model provenance.
Loading requires bitsandbytes and a supported accelerator:
| Vendor | Support |
|---|---|
| NVIDIA | compute capability sm_75+ (Turing / RTX 2000 and newer) |
| AMD ROCm | RDNA3 (gfx1100–gfx1102), RDNA3.5 (gfx1150–gfx1152), CDNA |
NF4 was chosen over FP8 and NVFP4 specifically for this reach: FP8 requires
sm_89+ and NVFP4 requires sm_120+, which would exclude most consumer GPUs.
Weights were quantized with bitsandbytes NF4 (double quantization, bf16 compute
dtype) and saved with save_pretrained. Tokenizer, processor, and any
trust_remote_code modules are copied unmodified from the upstream revision.
No architecture, vocabulary, or generation defaults were changed.
Some architectures leave a submodule in full precision — this is recorded in the
per-model card. It is required for correctness: transformers matches a
parameter's quantization state by substring rather than by prefix, so when one
quantized module's path is a suffix of another's (for example layers.0.… and
audio_tower.layers.0.…), the saved checkpoint cannot be reloaded. Leaving one
side unquantized keeps the artifact loadable by stock transformers, and has the
side effect of preserving full precision in the audio encoder.
Every artifact in this repository was reloaded after saving and checked to be still 4-bit before upload.
These are derivative works. Each is redistributed under its upstream licence, with all upstream copyright, patent, trademark and attribution notices retained. The upstream repository is the authoritative source for licence text; consult it before use.
| Subfolder | Upstream | Upstream licence |
|---|---|---|
caeleste-speech-higgs-audio-v3-stt-nf4 | bosonai/higgs-audio-v3-stt | apache-2.0 |
caeleste-speech-moss-tts-v1.5-nf4 | OpenMOSS-Team/MOSS-TTS-v1.5 | apache-2.0 |
caeleste-speech-qwen2.5-0.5b-instruct-nf4 | Qwen/Qwen2.5-0.5B-Instruct | apache-2.0 |
Changes made relative to upstream: weight tensors replaced with NF4-quantized
equivalents, and a quantization_config block added to config.json.
Some models used by the stack are not redistributed here because their licences impose obligations that are better handled directly between the user and the upstream publisher. These are downloaded from upstream and quantized at load time instead:
canopylabs/orpheus-tts-0.1-finetune-prod / unsloth/orpheus-3b-0.1-ft — Derived from meta-llama/Llama-3.2-3B-Instruct. The Llama 3.2 Community License attaches naming, attribution and usage-scale obligations to derivative models, so it is not redistributed here.LiquidAI/LFM2.5-Audio-1.5B — Distributed under the LFM Open License v1.0, which permits redistribution of derivative works but attaches a commercial-use threshold. Not redistributed here so that the licence relationship stays directly between the user and the upstream publisher.