Downloads · 30 days
8
20% of all-time downloads
vaghawan/tts-last
tts-last is a machine learning model from vaghawan. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
- Speaker Name: hausaspeaker - Base Model: Qwen/Qwen3-TTS-12Hz-1.7B-Base - Number of Layers: 28 - Hidden Size: 2048 - Training Epoch: 0 - Training Step: 30 - Best Validation Loss: 0.0001
Downloads · 30 days
8
20% of all-time downloads
All-time downloads
41
Public
Parameters
1.9B
47 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.4 GB · 100%
From the Hugging Face model README
config.json - Model configurationgeneration_config.json - Generation parametersmodel.safetensors - Model weights (includes speaker encoder weights)tokenizer_config.json - Tokenizer configurationvocab.json - Vocabulary filemerges.txt - BPE merges filepreprocessor_config.json - Preprocessor configurationspeech_tokenizer/ - Speech tokenizer model and configspeaker_encoder/speaker_config.json - Speaker encoder configurationtraining_state.json - Training state and configurationfrom qwen_tts.inference.qwen3_tts_model import Qwen3TTSModel
# Load the fine-tuned model
model = Qwen3TTSModel.from_pretrained(
"./output/checkpoint",
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2"
)
# Generate speech with the new speaker
text = "Your text here"
ref_audio = "path/to/reference_audio.wav"
wavs, sr = model.generate_voice_clone(
text=text,
language="Auto",
ref_audio=ref_audio,
ref_text="Reference text for ICL mode",
x_vector_only_mode=False
)
The model has been fine-tuned for speaker: hausa_speaker Speaker embedding is stored at index 3000 in the codec embedding layer. Speaker encoder weights are included in the checkpoint and have been fine-tuned.