Downloads · 30 days
0
RedbeardNZ/AV_MossFormer2_TSE_16K
AV_MossFormer2_TSE_16K is a machine learning model from RedbeardNZ. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
The AVMossFormer2TSE16K model weights for 16 kHz audio-visual target speaker extraction in ClearerVoice-Studio repo.
Downloads · 30 days
0
Access
Public
Updated May 16, 2025
Repo size
2.2 GB
Likes
0
Public
Click a slice to open those files.
.pt1.5 GB · 100%
From the Hugging Face model README
The AV_MossFormer2_TSE_16K model weights for 16 kHz audio-visual target speaker extraction in ClearerVoice-Studio repo.
This model is trained on large scale open-sourced datasets.
It extracts each speaker's voice from a multi-speaker video using facial recognition.