Skip to content

atasoglu

vit-gpt2-flickr8k

atasoglu/vit-gpt2-flickr8k

vit-gpt2-flickr8k is a image-to-text model from atasoglu. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as apache-2.0.

Vision Encoder Decoder (ViT + GPT2) model that fine-tuned on flickr8k-dataset for image-to-text task.

Downloads · 30 days

11

1% of all-time downloads

All-time downloads

2.1K

Public

Parameters

264M

2 GB on disk

Likes

1

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.bin982 MB · 50%

Parameter types

How the weights are stored.

F32239M · 90%

Task
Image-to-Text
Library
transformers
Type
vision-encoder-decoder
License
apache-2.0
Languages
en
Created
May 29, 2023
Updated
Aug 2, 2023