Downloads · 30 days
10
36% of all-time downloads
eagle0504/llava-video-text-model
llava-video-text-model is a machine learning model from eagle0504. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Fine-tuned LLaVA model on video-text data using DeepSpeed.
Downloads · 30 days
10
36% of all-time downloads
All-time downloads
28
Public
Parameters
8.1B
32.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.3 GB · 100%
From the Hugging Face model README
Fine-tuned LLaVA model on video-text data using DeepSpeed.
import requests
from PIL import Image
import torch
from transformers import AutoProcessor, LlavaForConditionalGeneration
# Load model and processor
processor = AutoProcessor.from_pretrained("eagle0504/llava-video-text-model")
model = LlavaForConditionalGeneration.from_pretrained(
"eagle0504/llava-video-text-model",
torch_dtype=torch.float16,
low_cpu_mem_usage=True,
).to(0)
# Define conversation with multiple images for video
conversation = [
{
"role": "user",
"content": [
{"type": "text", "text": "What is in this video?"},
{"type": "image"},
{"type": "image"},
{"type": "image"},
{"type": "image"},
{"type": "image"},
],
},
]
prompt = processor.apply_chat_template(conversation, add_generation_prompt=True)
# Process video frames (you need to extract frames from your video)
video_frames = [...] # List of PIL Images from video
inputs = processor(images=video_frames, text=prompt, return_tensors='pt').to(0, torch.float16)
# Generate response
output = model.generate(**inputs, max_new_tokens=200, do_sample=False)
response = processor.decode(output[0], skip_special_tokens=True)
print(response)
This model expects 5 frames extracted from each video. For best results: