Downloads · 30 days
0
sarulab-speech/DialogueSidon
DialogueSidon is a machine learning model from sarulab-speech. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-4.0.
Two-speaker dialogue separation model based on diffusion with a VAE-32 latent space.
Downloads · 30 days
0
Access
Public
Updated Jul 27, 2026
Repo size
3.7 GB
Likes
17
Public
Click a slice to open those files.
.pt21.9 GB · 51%
From the Hugging Face model README
Two-speaker dialogue separation model based on diffusion with a VAE-32 latent space.
| File | Description |
|---|---|
ssl_encoder.pt2 | w2v-BERT 2.0 backbone + latent projection heads |
diffusion_head.pt2 | DiffusionTransformerHead — single denoising step |
vae_decoder.pt2 | DAC VAE decoder: latents → 24 kHz audio |
metadata.json | Latent normalisation stats, scheduler config, model dims |