Skip to content

AnyModal

Image-Captioning-Llama-3.2-1B

AnyModal/Image-Captioning-Llama-3.2-1B

Image-Captioning-Llama-3.2-1B is a image-to-text model from AnyModal. Use it when you need a caption or text from an image. It is set up for AnyModal. The card lists the license as mit.

AnyModal/Image-Captioning-Llama-3.2-1B is an image captioning model built within the AnyModal framework. It integrates a Vision Transformer (ViT) encoder with the Llama 3.2-1B language model and has been trained on th…

Downloads · 30 days

0

Access

Public

Updated Dec 5, 2024

Repo size

172 MB

Likes

1

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.pt39.9 MB · 64%

At a glance

Task
Image-to-Text
Library
AnyModal
License
mit
Access
Public
Created
Dec 1, 2024
Updated
Dec 5, 2024
SHA
76977040

Base models

Task
Image-to-Text
Library
AnyModal
License
mit
Languages
en
Created
Dec 1, 2024
Updated
Dec 5, 2024