Downloads · 30 days
0
wissemkarous/LIPREAD
LIPREAD is a machine learning model from wissemkarous. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Lipreading is an advanced neural network model designed for accurate lip reading by incorporating lip landmark coordinates as a supplementary input to the traditional image sequence input. This enhancement to the orig…
Downloads · 30 days
0
Access
Public
Updated Mar 24, 2024
Repo size
93.6 MB
Likes
8
Public
Click a slice to open those files.
.dat66.4 MB · 70%
From the Hugging Face model README
Lipreading is an advanced neural network model designed for accurate lip reading by incorporating lip landmark coordinates as a supplementary input to the traditional image sequence input. This enhancement to the original LipReading architecture aims to improve the precision of sentence predictions by providing additional geometric context to the model.
| Scenario | Image Size (W x H) | CER | WER |
|---|---|---|---|
| Unseen speakers (Original) | 100 x 50 | 6.7% | 13.6% |
| Overlapped speakers (Original) | 100 x 50 | 2.0% | 5.6% |
| Unseen speakers (LipReading-final-year-project) | 128 x 64 | 6.7% | 13.3% |
| Overlapped speakers ( LipReading-final-year-project) | 128 x 64 | 1.9% | 4.6% |
| Overlapped speakers (LipReading) | 128 x 64 | 0.6% | 1.7% |
requirements.txt.Clone the repository:
git clone https://huggingface.co/wissemkarous/LIPREAD
Navigate to the project directory:
cd LIPREAD
Install the required dependencies:
pip install -r requirements.txt
To train the LipReading model with your dataset, first update the options.py file with the appropriate paths to your dataset and pretrained weights (comment out the weights if you want to start from scratch). Then, run the following command:
python train.py
To perform sentence prediction using the pre-trained model:
python inference.py --input_video <path_to_video>
note: ffmpeg is required to convert video to image sequence and run the inference script.

This model is built on top of the LipReading project on GitHub. The training process if similar to the original LipReading model, with the addition of landmark coordinates as a supplementary input. We used the pretrained weights from the original LipReading model as a starting point for training our model, froze the weights for the original LipNet layers, and trained the new layers for the landmark coordinates.
The dataset used to train this model is the Lipreading dataset. The dataset is not included in this repository, but can be downloaded from the link above.
Total training time: 12 h Total epochs: 51 Training hardware: NVIDIA GeForce RTX 3050 6GB

For an interactive view of the training curves, please refer to the tensorboard logs in the runs directory.
Use this command to view the logs:
tensorboard --logdir runs
We achieved a lowest WER of 1.7%, CER of 0.6% and a loss of 0.0256 on the validation dataset.
This project is licensed under the MIT License.
This model, LipReading, has been developed for academic purposes as a final year project. Special thanks to everyone who provided assistance and all references
Project Link: https://github.com/wissemkarous/End-of-year-Project