Downloads · 30 days
0
vumichien/AV-HuBERT
AV-HuBERT is a automatic speech recognition model from vumichien. Use it when you need speech turned into text. The card lists the license as apache-2.0.
These are model weights originally provided by the authors of the paper Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.
Downloads · 30 days
0
Access
Public
Updated Jan 17, 2023
Repo size
1.9 GB
Likes
13
Public
Click a slice to open those files.
.pt1.9 GB · 100%
From the Hugging Face model README
These are model weights originally provided by the authors of the paper Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction.
<figure> <img src="https://huggingface.co/vumichien/AV-HuBERT/resolve/main/HuBert.png" alt="Audio-visual HuBERT"> <figcaption>Audio-visual HuBERT </figcaption> </figure>Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker’s lip movements and the produced sound.
Audio-Visual Hidden Unit BERT (AV-HuBERT), a self-supervised representation learning framework for audio-visual speech, which masks multi-stream video input and predicts automatically discovered and iteratively refined multimodal hidden units. AV-HuBERT learns powerful audio-visual speech representation benefiting both lip-reading and automatic speech recognition.
The official code of this paper in here
The authors trained the model on LRS3 with 433 hours of transcribed English videos and English portion of VoxCeleb2, which amounts to 1,326 hours