Downloads · 30 days
0
JP106978/flickr8k-image-captioner
flickr8k-image-captioner is a image-to-text model from JP106978. Use it when you need a caption or text from an image. The card lists the license as mit.
This model generates natural language descriptions from input images by combining a Pretrained ResNet50 CNN visual feature extractor with an LSTM sequence decoder.
Downloads · 30 days
0
Access
Public
Updated Aug 28, 2026
Repo size
8.8 MB
Likes
0
Public
Click a slice to open those files.
.pt8.8 MB · 100%
From the Hugging Face model README
This model generates natural language descriptions from input images by combining a Pretrained ResNet50 CNN visual feature extractor with an LSTM sequence decoder.
import torch
from predict import CaptionPredictor
# Load directly from checkpoint
predictor = CaptionPredictor("image_caption_model.pt")
caption = predictor.predict("sample.jpg", method="beam", beam_width=3)
print(caption)