Downloads · 30 days
33
1% of all-time downloads
gospacedev/blip-image-captioning-base-bf16
blip-image-captioning-base-bf16 is a image-to-text model from gospacedev. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as mit.
This model is a quantized version of the Salesforce/blip-image-captioning-base, an image-to-text model. From a memory footprint of 989 MBs - 494 MBs by quantizing the percision of float32 to bfloat 16, reducing the mo…
Downloads · 30 days
33
1% of all-time downloads
All-time downloads
2.7K
Public
Parameters
247M
501 MB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors495 MB · 98%
From the Hugging Face model README
This model is a quantized version of the Salesforce/blip-image-captioning-base, an image-to-text model. From a memory footprint of 989 MBs -> 494 MBs by quantizing the percision of float32 to bfloat 16, reducing the model's memory size by 50 percent.
| <img src="https://huggingface.co/gospacedev/blip-image-captioning-base-bf16/resolve/main/cat%20in%20currents.png" width="316" height="316"> |
|---|
| a cat sitting on top of a purple and red striped carpet |
Use the code below to get started with the model.
from transformers import BlipForConditionalGeneration, BlipProcessor
import requests
from PIL import Image
model = BlipForConditionalGeneration.from_pretrained("gospacedev/blip-image-captioning-base-bf16")
processor = BlipProcessor.from_pretrained("gospacedev/blip-image-captioning-base-bf16")
# Load sample image
image = Image.open(requests.get(img_url, stream=True).raw).convert('RGB')
# Generate output
inputs = processor(image, return_tensors="pt")
output = model.generate(**inputs)
result = processor.decode(out[0], skip_special_tokens=True)
print(results)