Downloads · 30 days
0
tiiuae/visper
visper is a machine learning model from tiiuae. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-nc-2.0.
ViSPer is a model for audio visual speech recognition (VSR/AVSR). Trained on 5500 hours of labelled video data.
Downloads · 30 days
0
Access
Public
Updated Jun 5, 2024
Repo size
4.2 GB
Likes
10
Public
Click a slice to open those files.
.pth3.1 GB · 100%
From the Hugging Face model README
ViSPer is a model for audio visual speech recognition (VSR/AVSR). Trained on 5500 hours of labelled video data.
We use our proposed dataset to train a encoder-decoder model in a fully-supervised manner under a multi-lingual setting. While the encoder size is 12 layers, the decoder size is 6 layers. The hidden size, MLP and number of heads are set to 768, 3072 and 12, respectively. The unigram tokenizers are learned for all languages combined and have a vocabulary size of 21k. The models are trained for 150 epochs on 64 Nvidia A100 GPUs (40GB) using AdamW optimizer with max LR of 1e-3 and a weight decay of 0.1. A cosine scheduler with a warm-up of 5 epochs is used for training. The maximum batch size per GPU is set to 1800 video frames.
We provide the results of the model on our proposed benchmarks in this table:
| Language | VSR (WER/CER) | AVSR (WER/CER) |
|---|---|---|
| French | 29.8 | 5.7 |
| Spanish | 39.4 | 4.4 |
| Arabic | 47.8 | 8.4 |
| Chinese | 51.3 (CER) | 15.4 (CER) |
| English | 49.1 | 8.1 |
In essence, while we hope that ViSPer will open the doors for new research questions and opportunities, and should only be used for this purpose. There are also potential dual use concerns that come with releasing ViSPer (dataset and models), trained on a substantial corpus of multilingual video data. While the technology behind ViSPer offers significant advances in multimodal speech recognition, it should only be used for research purposes.
@inproceedings{djilali2023lip2vec,
title={Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation Mapping},
author={Djilali, Yasser Abdelaziz Dahou and Narayan, Sanath and Boussaid, Haithem and Almazrouei, Ebtessam and Debbah, Merouane},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
pages={13790--13801},
year={2023}
}
@inproceedings{djilali2024vsr,
title={Do VSR Models Generalize Beyond LRS3?},
author={Djilali, Yasser Abdelaziz Dahou and Narayan, Sanath and LeBihan, Eustache and Boussaid, Haithem and Almazrouei, Ebtesam and Debbah, Merouane},
booktitle={Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision},
pages={6635--6644},
year={2024}
}