Downloads · 30 days
0
0% of all-time downloads
hslay13/vit_lstm_jester
vit_lstm_jester is a machine learning model from hslay13. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as apache-2.0.
This repository contains a Bidirectional LSTM model for dynamic hand gesture recognition, trained on features extracted from the Jester dataset. The model is designed to classify temporal sequences of frame-level embe…
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
24
Public
Repo size
6.2 MB
Likes
0
Public
Click a slice to open those files.
.keras6.2 MB · 100%
From the Hugging Face model README
This repository contains a Bidirectional LSTM model for dynamic hand gesture recognition, trained on features extracted from the Jester dataset. The model is designed to classify temporal sequences of frame-level embeddings.
The overall architecture is a two-stage process:
Model Version: clean-baseline-v2.0-hf Architecture: Bidirectional LSTM with ViT-B/16 feature extractor Parameters (LSTM): ~514k Framework: TensorFlow + Keras (for LSTM), PyTorch + timm (for ViT feature extraction)
| Metric | Value |
|---|---|
| Validation Accuracy | 0.8100 (81.00%) |
| Validation Loss | 0.7010 |
| Weighted Average F1 Score | 0.81 |
Name: 20bn-Jester Dataset Source: local Train samples: 50,420 Validation samples: 7,047
Batch Size: 8 Optimizer: Adam Loss: Categorical Cross-Entropy Learning Rate: 0.001 (reduced by ReduceLROnPlateau) Epochs Trained: Up to 30 (best weights restored from epoch 29 based on validation accuracy) Regularization: Dropout (0.4, 0.3), L2 (1e-4) on Dense layer Callbacks: ModelCheckpoint, ReduceLROnPlateau, EarlyStopping (patience=7)
Preprocessing: Single deterministic pipeline (no augmentation) Resize: 224×224 Normalization: ImageNet statistics (mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]) Temporal Sampling: Fixed sequence length of 37 frames Padding: Zero-vector with masking Augmentation: none (clean baseline principle) Test-Time Augmentation: False (disabled)
Type: standard ViT-B/16 (no modifications, frozen during LSTM training) ViT Output: 768-D feature embeddings per frame Classification Head (ViT): Removed for feature extraction LSTM Model: - Input: Sequences of ViT embeddings (37 frames, 768 dimensions) - Layers: Masking, Bidirectional LSTM (256 units, return_sequences=True), Dropout (0.4), Bidirectional LSTM (128 units), Dropout (0.4), Dense (128 units, ReLU, L2 regularization), Dropout (0.3), Dense (27 classes, Softmax) Total Trainable Parameters (LSTM): ~514k Dropout: 0.4, 0.3
'Doing other things', 'Drumming Fingers', 'No gesture', 'Pulling Hand In', 'Pulling Two Fingers In', 'Pushing Hand Away', 'Pushing Two Fingers Away', 'Rolling Hand Backward', 'Rolling Hand Forward', 'Shaking Hand', 'Sliding Two Fingers Down', 'Sliding Two Fingers Left', 'Sliding Two Fingers Right', 'Sliding Two Fingers Up', 'Stop Sign', 'Swiping Down', 'Swiping Left', 'Swiping Right', 'Swiping Up', 'Thumb Down', 'Thumb Up', 'Turning Hand Clockwise', 'Turning Hand Counterclockwise', 'Zooming In With Full Hand', 'Zooming In With Two Fingers', 'Zooming Out With Full Hand', 'Zooming Out With Two Fingers'
Refer to how_to_use.py for a complete end-to-end example.
• Requires consistent frame rate and sampling • Sensitive to heavy occlusion and motion blur • Assumes a single dominant gesture per clip • Performance depends on ViT embedding quality
MIT License
This model was trained on the Jester Dataset, a large-scale video dataset for hand gesture recognition. We would like to thank Twenty Billion Neurons (TwentyBN) for creating and sharing this dataset.
This model consists of two parts: a Vision Transformer (ViT) feature extractor and an LSTM-based temporal classifier.
First, make sure you have the required libraries installed:
pip install tensorflow torch torchvision timm transformers
The ViT backbone can be loaded from the Hugging Face Hub here, while the trained LSTM model can be loaded from the best_lstm_model.keras file in this repository.