Downloads · 30 days
0
ciphermosaic/blip-image-captioning
blip-image-captioning is a machine learning model from ciphermosaic. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A LoRA fine tuned version of Salesforce/blip-image-captioning-base for image caption generation using the Flickr8k dataset. This project demonstrates parameter efficient fine tuning with PEFT while keeping the base mo…
Downloads · 30 days
0
Access
Public
Updated Jul 10, 2026
Repo size
1.8 MB
Likes
0
Public
Click a slice to open those files.
.safetensors1.8 MB · 71%
From the Hugging Face model README
A LoRA fine tuned version of Salesforce/blip-image-captioning-base for image caption generation using the Flickr8k dataset. This project demonstrates parameter efficient fine tuning with PEFT while keeping the base model frozen.
ciphermosaic/blip-image-captioningSalesforce/blip-image-captioning-baseThe model was fine tuned on the jxie/flickr8k dataset.
Dataset split:
| Split | Samples |
|---|---|
| Train | 6,000 |
| Validation | 1,000 |
| Test | 1,000 |
5e-5LoraConfig(
r=8,
lora_alpha=16,
lora_dropout=0.1,
target_modules=[
"qkv",
"projection"
],
bias="none"
)
| Metric | Value |
|---|---|
| Validation Loss | 8.0478 |
This checkpoint is intended as an educational fine tuning project and baseline implementation for BLIP image captioning with LoRA.
from transformers import BlipProcessor, BlipForConditionalGeneration
from peft import PeftModel
base_model = BlipForConditionalGeneration.from_pretrained(
"Salesforce/blip-image-captioning-base"
)
model = PeftModel.from_pretrained(
base_model,
"ciphermosaic/blip-image-captioning"
)
processor = BlipProcessor.from_pretrained(
"Salesforce/blip-image-captioning-base"
)
from PIL import Image
import requests
image = Image.open(requests.get(image_url, stream=True).raw).convert("RGB")
inputs = processor(images=image, return_tensors="pt")
outputs = model.generate(**inputs)
caption = processor.decode(outputs[0], skip_special_tokens=True)
print(caption)
The model can also be loaded using standard Hugging Face Transformers workflows together with the PEFT adapter.
This work builds upon:
Special thanks to the open source community for making these resources available.
CipherMosaic
GitHub: https://github.com/ciphermosaic