Downloads · 30 days
0
NanG01/Action_detection
Action_detection is a video classification model from NanG01. Use it for the video classification task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
This model performs human action classification on videos using a CNN-GRU architecture built on top of MobileNetV2 (1.0, 224) features and trained on the UCF101 dataset. It is well-suited for recognizing actions from…
Downloads · 30 days
0
Access
Public
Updated Aug 27, 2025
Repo size
11.3 MB
Likes
4
Public
Click a slice to open those files.
.pt11.3 MB · 99%
From the Hugging Face model README
This model performs human action classification on videos using a CNN-GRU architecture built on top of MobileNetV2 (1.0, 224) features and trained on the UCF101 dataset.
It is well-suited for recognizing actions from short trimmed video clips.
Base model: google/mobilenet_v2_1.0_224
Architecture: CNN-GRU

Dataset: UCF101 - Action Recognition Dataset (https://www.kaggle.com/datasets/abdallahwagih/ucf101-videos)
Task: Video Classification (Action Recognition)
Metrics: Accuracy
License: MIT
pip install torch torchvision opencv-python
from action_model import load_action_model, preprocess_frames, predict_action
import cv2
# Load model
model = load_action_model(model_path="best_model.pt", device="cpu", num_classes=5)
# Read frames from video
cap = cv2.VideoCapture("path_to_video.mp4")
frames = []
while True:
ret, frame = cap.read()
if not ret:
break
frames.append(frame)
cap.release()
# Preprocess frames for model input
clip_tensor = preprocess_frames(frames[:16], seq_len=16, resize=(112,112))
# Predict action
result = predict_action(model, clip_tensor, device="cpu")
print(result)
Intended for:
Limitations:
action · cnn-gru · video-classification · ucf101 · mobilenetv2 · deep-learning · torch