Skip to content

sachin

vit2distilgpt2

sachin/vit2distilgpt2

vit2distilgpt2 is a image-to-text model from sachin. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as mit.

This model takes in an image and outputs a caption. It was trained using the Coco dataset and the full training script can be found in this kaggle kernel

Downloads · 30 days

39

0% of all-time downloads

All-time downloads

11.2K

Public

Parameters

195M

1.5 GB on disk

Likes

8

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.bin743 MB · 50%

Parameter types

How the weights are stored.

F32182M · 94%

Datasets

Task
Image-to-Text
Library
transformers
Type
vision-encoder-decoder
License
mit
Languages
en
Created
Mar 2, 2022
Updated
Aug 17, 2023