Downloads · 30 days
550
41% of all-time downloads
Trendyol/Trendyol-TTS
Trendyol-TTS is a text-to-speech model from Trendyol. Use it when you need text read aloud. The card lists the license as mit.
Trendyol-TTS is a Turkish text-to-speech research model built on top of openbmb/VoxCPM2. It is based on the Trendyol Turkish TTS LoRA training run and uses the manually selected step0002000 checkpoint as the released…
Downloads · 30 days
550
41% of all-time downloads
All-time downloads
1.4K
Public
Parameters
2.3B
10.2 GB on disk
Likes
38
Trending 2
Click a slice to open those files.
.safetensors4.7 GB · 93%
From the Hugging Face model README
Trendyol-TTS is a Turkish text-to-speech research model built on top of openbmb/VoxCPM2. It is based on the Trendyol Turkish TTS LoRA training run and uses the manually selected step_0002000 checkpoint as the released default. The repository contains a merged standalone model artifact: the Turkish LoRA adapter was applied to the VoxCPM2 base weights, while the original LoRA adapter files are kept for provenance and auditability. The model is optimized for Turkish speech synthesis experiments, evaluation, and controlled prototyping rather than general-purpose production serving.
openbmb/VoxCPM2 and released as a merged model artifact for simpler loading.step_0002000 Checkpoint - Chosen as the default after manual listening and checkpoint comparison.cfg_value=2.0 with inference_timesteps=16.Install the VoxCPM2 runtime requirements on a GPU-capable environment, then load the merged model directly from the Hub.
import soundfile as sf
from voxcpm.core import VoxCPM
model_id = "Trendyol/Trendyol-TTS"
model = VoxCPM.from_pretrained(
hf_model_id=model_id,
load_denoiser=False,
optimize=True,
)
audio = model.generate(
text="Merhaba, Trendyol TTS modelinden bir Türkçe ses örneği dinliyorsunuz.",
cfg_value=2.0,
inference_timesteps=16,
max_len=4096,
normalize=True,
denoise=False,
)
sf.write("trendyol_tts_sample.wav", audio, model.tts_model.sample_rate)
The currently recommended clean default is:
cfg_value = 2.0
inference_timesteps = 16
A more expressive setting worth testing is:
cfg_value = 1.5
inference_timesteps = 16
Avoid using cfg_value=2.5 as a general production default. Internal proxy checks showed some generated samples peaking too close to 0 dBFS, even when the measured clipping fraction was zero.
openbmb/VoxCPM2step_0002000+20 Hours Turkish Speaktr)merge_manifest.json, and preserved lora_adapter/ filesThe step_0002000 checkpoint was selected as the default based on manual listening preference, response quality, and available audio-quality checks. Later continuation checkpoints such as step_0002250 and step_0002500 were healthy training runs, but they did not replace step_0002000 as the recommended release default.
Evaluation artifacts used during development include checkpoint sweeps, inference-parameter sweeps, and blind-listening package tooling. This model card should not be read as a claim of formal MOS, large-scale production stress testing, or complete downstream safety validation.
The model card declares MIT metadata following the upstream VoxCPM2 model metadata where applicable. Users are responsible for checking the licenses and usage terms of openbmb/VoxCPM2, the private training dataset, and any downstream deployment environment.