Downloads · 30 days
12
1% of all-time downloads
Mouwiya/BLIP_image_captioning
BLIP_image_captioning is a image-to-text model from Mouwiya. Use it when you need a caption or text from an image. It is set up for transformers.
BLIPimagecaptioning is a model based on the BLIP (Bootstrapping Language-Image Pre-training) architecture, specifically designed for image captioning tasks. The model has been fine-tuned on the "image-in-words400" dat…
Downloads · 30 days
12
1% of all-time downloads
All-time downloads
816
Public
Parameters
247M
990 MB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors990 MB · 100%
From the Hugging Face model README
BLIP_image_captioning is a model based on the BLIP (Bootstrapping Language-Image Pre-training) architecture, specifically designed for image captioning tasks. The model has been fine-tuned on the "image-in-words400" dataset, which consists of images and their corresponding descriptive captions. This model leverages both visual and textual data to generate accurate and contextually relevant captions for images.
The model was fine-tuned on a shuffled and subsetted version of the "image-in-words400" dataset. A total of 400 examples were used during the fine-tuning process to allow for faster iteration and development.
To use this model for image captioning, you can load it using the Hugging Face transformers library and perform inference as shown below:
from transformers import BlipProcessor, BlipForConditionalGeneration
from PIL import Image
import requests
from io import BytesIO
# Load the processor and model
model_name = "Mouwiya/BLIP_image_captioning"
processor = BlipProcessor.from_pretrained(model_name)
model = BlipForConditionalGeneration.from_pretrained(model_name)
# Example usage
image_url = "URL_OF_THE_IMAGE"
response = requests.get(image_url)
image = Image.open(BytesIO(response.content)).convert("RGB")
inputs = processor(images=image, return_tensors="pt")
outputs = model.generate(**inputs)
caption = processor.decode(outputs[0], skip_special_tokens=True)
print(caption)
The model was evaluated on a subset of the "image-in-words400" dataset using the BLEU score. The evaluation results are as follows:
Mouwiya S. A. Al-Qaisieh [email protected]