Downloads · 30 days
3
8% of all-time downloads
aiseosae/2026-TTS
2026-TTS is a text-to-speech model from aiseosae. Use it when you need text read aloud. It is set up for transformers. The card lists the license as cc-by-nc-sa-4.0.
A naturalness-first prompt-driven TTS, built on top of magma90909/vocenceminerv7. Two things distinguish this checkpoint:
Downloads · 30 days
3
8% of all-time downloads
All-time downloads
37
Public
Parameters
1.9B
4.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.5 GB · 100%
From the Hugging Face model README
A naturalness-first prompt-driven TTS, built on top of magma90909/vocence_miner_v7. Two things distinguish this checkpoint:
24 kHz mono WAV output, single forward call, no reference audio, no PEFT runtime. Everything ships in this repo.
pip install qwen-tts transformers torch soundfile
from qwen_tts import Qwen3TTSModel
import soundfile as sf
m = Qwen3TTSModel.from_pretrained("magma90909/vocence_miner_v8")
wavs, sr = m.generate_voice_design(
text="The train to Edinburgh departs from platform four.",
instruct="A man with a British English accent, calm and natural.",
language="english",
)
sf.write("out.wav", wavs[0], sr)
demo.py walks through three preset prompts.
instructThe model responds best to subtle, conversational language — not intensifiers like "intensely sad" or "nearly shouting". Stack these elements freely:
| Layer | Phrasings |
|---|---|
| Accent / region | British English, Scottish, Welsh, Northern Irish, Irish, unspecified |
| Gender | a man, a woman, a British woman |
| Mood | speaking warmly, softly sad, quietly pleased, with a touch of anger |
| Persona | bedtime storyteller, soft and warm; news anchor, professional and neutral; meditation guide, soft and serene |
| Pace | unhurried, brisk steady, naturally measured |
Some example prompts that work well:
A British man speaks calmly and naturally.
A woman with a Scottish accent, in an everyday speaking tone.
A man, softly sad, calm and unhurried.
A British news anchor, professional and neutral, at a brisk steady pace.
A clear, neutral voice reading the sentence.
Best at:
Less suited for:
CC BY-NC-SA 4.0 — research and non-commercial use only.
model.safetensors # merged Talker weights (3.6 GB)
speech_tokenizer/ # Qwen3 12 Hz audio codec (~650 MB)
tokenizer.json + ... # text tokenizer
config.json + ... # model configs
miner.py # Vocence engine
chute_config.yml # Chutes build (TEE / pro_6000)
vocence_config.yaml # runtime knobs
demo.py # quick smoke test
The Vocence files make this repo deployable on Bittensor SN78 (Vocence) via the canonical Vocence/Chutes wrapper without modification.