Downloads · 30 days
0
gurumurthy3/gpt2vl-stackformer-v1
gpt2vl-stackformer-v1 is a image-to-text model from gurumurthy3. Use it when you need a caption or text from an image. The card lists the license as mit.
A multimodal model combining Vision Transformer (ViT-B/16) and GPT-2 for image captioning, trained on Flickr8K dataset.
Downloads · 30 days
0
Access
Public
Updated Oct 13, 2025
Repo size
892 MB
Likes
0
Public
Click a slice to open those files.
.pth1.8 GB · 99%
From the Hugging Face model README
A multimodal model combining Vision Transformer (ViT-B/16) and GPT-2 for image captioning, trained on Flickr8K dataset.
This model generates natural language captions for images by:
pip install torch torchvision transformers pillow huggingface_hub
import torch
from transformers import GPT2Tokenizer
from PIL import Image
from torchvision import transforms
# Load checkpoint
checkpoint = torch.load("model_fp32/model_checkpoint.pth", map_location="cpu")
# Load tokenizer
tokenizer = GPT2Tokenizer.from_pretrained("model_fp32/tokenizer")
# Load your model architecture (you need to define this)
# model = YourVisionGPTModel(config)
# model.load_state_dict(checkpoint['model_state_dict'])
# model.eval()
print("Model loaded successfully!")
# Load FP16 checkpoint
checkpoint = torch.load("model_fp16/model_checkpoint.pth", map_location="cpu")
# Load model and convert to FP16
# model = YourVisionGPTModel(config)
# model.load_state_dict(checkpoint['model_state_dict'])
# model.half() # Convert to FP16
# model.eval()
# For inference with FP16, also convert input images to FP16
image_transform = transforms.Compose([
transforms.Resize((224, 224)),
transforms.Lambda(lambda x: x.convert('RGB')),
transforms.ToTensor(),
transforms.Normalize(
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
),
])
# Load and preprocess image
image = Image.open("your_image.jpg")
image_tensor = image_transform(image).unsqueeze(0) # Add batch dimension
# Generate caption
with torch.no_grad():
# Forward pass
generated_ids = model.generate(
image_tensor,
max_length=50,
num_beams=5,
temperature=0.7
)
# Decode caption
caption = tokenizer.decode(generated_ids[0], skip_special_tokens=True)
print(f"Generated caption: {caption}")
┌─────────────────┐
│ Input Image │
│ (224x224) │
└────────┬────────┘
│
▼
┌─────────────────┐
│ ViT-B/16 │
│ (frozen) │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Projection │
│ (trainable) │
└────────┬────────┘
│
▼
┌─────────────────┐
│ GPT-2 │
│ (frozen) │
└────────┬────────┘
│
▼
┌─────────────────┐
│ Caption Output │
└─────────────────┘
If you use this model, please cite:
@misc{vision-gpt-flickr8k,
author = {gurumurthy3},
title = {Vision-GPT: Image Captioning with ViT and GPT-2},
year = {2025},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/gurumurthy3/vision-gpt-flickr8k}}
}
MIT License