Downloads · 30 days
0
jspaulsen/unmute-encoder
unmute-encoder is a machine learning model from jspaulsen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
A speaker embedding encoder trained to replicate Kyutai's unreleased "unmute encoder". This model extracts speaker embeddings from audio for use with Kyutai's Moshi TTS system.
Downloads · 30 days
0
Access
Public
Updated Feb 5, 2026
Parameters
115M
920 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors460 MB · 100%
From the Hugging Face model README
A speaker embedding encoder trained to replicate Kyutai's unreleased "unmute encoder". This model extracts speaker embeddings from audio for use with Kyutai's Moshi TTS system.
The encoder is built on top of Kyutai's Mimi neural audio codec:
[512, 125] (512 channels, 125 time steps for 10s audio)Audio (24kHz, 10s) -> Mimi Encoder -> Latent [512, T] -> MLP Projector -> Embedding [512, 125]
from src.models.mimi import MimiEncoder
# Load the encoder
encoder = MimiEncoder.from_pretrained(
model_name="jspaulsen/unmute-encoder",
device="cuda",
num_codebooks=32,
)
# Create embedding from audio tensor [1, 1, T] at 24kHz
output = encoder(audio_tensor)
embedding = output.embedding # [1, 512, 125]
Trained using supervised learning with a hybrid loss (L1 + cosine similarity) against speaker embeddings from kyutai/tts-voices.