Skip to content

gurumurthy3

gpt2vl-stackformer-v1

gurumurthy3/gpt2vl-stackformer-v1

gpt2vl-stackformer-v1 is a image-to-text model from gurumurthy3. Use it when you need a caption or text from an image. The card lists the license as mit.

A multimodal model combining Vision Transformer (ViT-B/16) and GPT-2 for image captioning, trained on Flickr8K dataset.

Downloads · 30 days

0

Access

Public

Updated Oct 13, 2025

Repo size

892 MB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.pth1.8 GB · 99%

At a glance

Task
Image-to-Text
License
mit
Access
Public
Created
Oct 13, 2025
Updated
Oct 13, 2025
SHA
cef7c920
Task
Image-to-Text
License
mit
Languages
en
Created
Oct 13, 2025
Updated
Oct 13, 2025