Downloads · 30 days
23
26% of all-time downloads
Pokzy/flickr8k-finetuned
flickr8k-finetuned is a image-to-text model from Pokzy. Use it when you need a caption or text from an image. The card lists the license as apache-2.0.
This repository contains a Fine-Tuned BLIP (Bootstrapping Language-Image Pre-training) model for Image Captioning, trained on the Flickr8k dataset. - Base Model: Salesforce/blip-image-captioning-base - Developer: Pokz…
Downloads · 30 days
23
26% of all-time downloads
All-time downloads
87
Public
Parameters
224M
896 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors896 MB · 100%
From the Hugging Face model README
This repository contains a Fine-Tuned BLIP (Bootstrapping Language-Image Pre-training) model for Image Captioning, trained on the Flickr8k dataset.
Salesforce/blip-image-captioning-basefacebook/nllb-200-distilled-600M).You can easily load and use this fine-tuned model using Hugging Face's transformers library:
import torch
from PIL import Image
from transformers import BlipProcessor, BlipForConditionalGeneration
# 1. Load Preprocessor & Fine-Tuned Model
processor = BlipProcessor.from_pretrained("Salesforce/blip-image-captioning-base")
model = BlipForConditionalGeneration.from_pretrained("Pokzy/flickr8k-finetuned")
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)
model.eval()
# 2. Load & Preprocess Image
image_path = "example.jpg"
image = Image.open(image_path).convert("RGB")
inputs = processor(images=image, return_tensors="pt").to(device)
# 3. Generate Caption
with torch.no_grad():
output_ids = model.generate(
**inputs,
do_sample=True,
temperature=0.7,
top_p=0.9,
max_length=50
)
caption = processor.decode(output_ids[0], skip_special_tokens=True)
print("Generated Caption:", caption)