Downloads · 30 days
19
3% of all-time downloads
Fer14/paligemma_coffee_machine_caption
paligemma_coffee_machine_caption is a visual question answering model from Fer14. Use it for the visual question answering task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
Google's Paligemma VLM (Vision Language Model) finetuned to provide captions to coffe machine images
Downloads · 30 days
19
3% of all-time downloads
All-time downloads
554
Public
Repo size
47 GB
Likes
2
Public
Click a slice to open those files.
.bin11.7 GB · 99%
From the Hugging Face model README
Google's Paligemma VLM (Vision Language Model) finetuned to provide captions to coffe machine images
from transformers import PaliGemmaForConditionalGeneration, PaliGemmaProcessor
from PIL import Image
model_id = "Fer14/paligemma_coffee_machine_caption"
model = PaliGemmaForConditionalGeneration.from_pretrained(model_id)
processor = PaliGemmaProcessor.from_pretrained(model_id)
image = Image.open("path to your image").convert("RGB")
prompt = (
f"Generate a caption for the following coffee maker image. The caption has to be of the following structure:\n"
"\"A <color> <type>, <accessories>, <shape> shaped, with <screen> and <number> <b_color> butons\"\n\n"
"in which:\n"
"- color: red, black, blue...\n"
"- type: coffee machine, coffee maker, espresso coffee machine...\n"
"- accessories: a list of accessories like the ones described above\n"
"- shape: cubed, round...\n"
"- screen: screen, no screen.\n"
"- number: amount of buttons to add\n"
"- b_color: color of the buttons"
)
inputs = processor(
text=prompt,
images=image,
return_tensors="pt",
padding="longest",
)
output = model.generate(**inputs, max_length=1000)
decoded_output = processor.decode(output[0], skip_special_tokens=True)[len(prompt) :]