Downloads · 30 days
0
ZDisket/echolancer-stage3-zs
echolancer-stage3-zs is a text-to-speech model from ZDisket. Use it when you need text read aloud. It is set up for pytorch. The card lists the license as mit.
A decoder-only transformer text-to-speech model capable of zero-shot voice cloning from a single reference audio sample.
Downloads · 30 days
0
Access
Public
Updated Feb 9, 2026
Repo size
21.6 GB
Likes
0
Public
Click a slice to open those files.
.pt21.6 GB · 100%
From the Hugging Face model README
A decoder-only transformer text-to-speech model capable of zero-shot voice cloning from a single reference audio sample.
Echolancer is a neural codec language model for text-to-speech synthesis. This Stage 3 checkpoint enables zero-shot voice cloning - the ability to synthesize speech in any voice given just a few seconds of reference audio.
| Component | Details |
|---|---|
| Architecture | Decoder-only Transformer |
| Positional Encoding | ALiBi (Attention with Linear Biases) |
| Audio Tokenizer | NeuCodec (65,536 vocab) |
| Speaker Encoder | ECAPA-TDNN (192-dim embeddings) |
| Text Tokenizer | Character-level |
This model is designed for:
git clone https://github.com/ZDisket/Echolancer
cd Echolancer
pip install -r requirements.txt
import torch
import torchaudio
from speechbrain.inference.speaker import EncoderClassifier
from echolancerfe import EcholancerFE
from neucodecfe import NeuCodecFE
# Load models
echolancer = EcholancerFE(model_config_path="config/model_stage3_zs.yaml")
echolancer.load_checkpoint("path/to/checkpoint.pt")
neu_codec = NeuCodecFE(is_cuda=True, offset=echolancer.get_vocab_offset())
speaker_encoder = EncoderClassifier.from_hparams(source="speechbrain/spkrec-ecapa-voxceleb")
# Extract speaker embedding from reference audio
def get_speaker_embedding(audio_path):
signal, fs = torchaudio.load(audio_path)
signal = signal.mean(dim=0, keepdim=True) # mono
if fs != 16000:
signal = torchaudio.transforms.Resample(fs, 16000)(signal)
return speaker_encoder.encode_batch(signal).squeeze(0)
# Generate speech
speaker_emb = get_speaker_embedding("reference.wav")
with torch.amp.autocast(device_type='cuda', dtype=torch.bfloat16):
codes = echolancer.infer(
text="Hello, this is a test.",
speaker_id=speaker_emb,
temperature=0.8,
top_p=0.92,
max_length=1024
)
# Decode to audio
waveform = neu_codec.decode_codes(codes.unsqueeze(1))
torchaudio.save("output.wav", waveform[0].cpu(), 24000)
This is a Stage 3 zero-shot checkpoint, trained to:
This technology can generate speech that sounds like real people. Users should: