Downloads · 30 days
0
poonai/imagenet-caption
imagenet-caption is a machine learning model from poonai. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This project provides an image captioning model trained on the visual-layer/imagenet-1k-vl-enriched dataset. The model architecture combines a ViT backbone timm/vitmediumdpatch16reg4gap384.sbb2e200in12kftin1k for imag…
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2025
Repo size
1.8 GB
Likes
0
Public
Click a slice to open those files.
.ckpt1.8 GB · 100%
From the Hugging Face model README
This project provides an image captioning model trained on the visual-layer/imagenet-1k-vl-enriched dataset. The model architecture combines a ViT backbone timm/vit_mediumd_patch16_reg4_gap_384.sbb2_e200_in12k_ft_in1k for image feature extraction and a GPT-2 language model openai-community/gpt2 for text generation.
A custom projection layer was implemented to map the image features from the vision backbone to the input space of the language model, enabling seamless integration between the two modalities.
To run this app, follow these steps:
This project uses uv for fast dependency management. To install all dependencies, run:
uv sync
Run inference To test the model and generate captions, run:
uv run inference.py
This will process your input images and output captions using the trained model.

a boy holding a fish in the woods