Downloads · 30 days
37
67% of all-time downloads
detker/temporal-vit-85M
temporal-vit-85M is a machine learning model from detker. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This repository contains the trained weights for the Temporal Vision Transformer (ViT) model fine-tuned with LoRA (Low-Rank Adaptation). The model is designed for temporal video classification tasks.
Downloads · 30 days
37
67% of all-time downloads
All-time downloads
55
Public
Parameters
85.9M
344 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors344 MB · 100%
From the Hugging Face model README
This repository contains the trained weights for the Temporal Vision Transformer (ViT) model fine-tuned with LoRA (Low-Rank Adaptation). The model is designed for temporal video classification tasks.
.safetensorsconfig.jsonThe available weights were obtained by training on a single RTX 5090 GPU for approximately 150 epochs and 144 effective batch size per GPU.
You can load the model using Hugging Face's AutoModel and AutoConfig classes:
from transformers import AutoModel, AutoConfig
from hf_pretrained_model import TemporalViTConfig, TemporalViTHF
# Register model
AutoConfig.register('temporal-vit', TemporalViTConfig)
AutoModel.register(TemporalViTConfig, TemporalViTHF)
# Load the model
model = AutoModel.from_pretrained('detker/temporal-vit-85M',
trust_remote_code=True)
# Example usage
inputs = ... # Prepare your input tensor
outputs = model(inputs)
model.safetensors: Trained model weights.config.json: Model configuration file.