Downloads · 30 days
15
3% of all-time downloads
aloobun/rmfg
rmfg is a image-to-text model from aloobun. Use it when you need a caption or text from an image. It is set up for transformers. The card lists the license as apache-2.0.
<img src="https://i.pinimg.com/736x/7e/46/a6/7e46a6881623dfd3e1a2a5a2ae692374.jpg" width="300"
Downloads · 30 days
15
3% of all-time downloads
All-time downloads
561
Public
Parameters
1.9B
3.7 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors3.7 GB · 100%
From the Hugging Face model README
Image <img src="https://media-cldnry.s-nbcnews.com/image/upload/t_fit-760w,f_auto,q_auto:best/rockcms/2023-12/231202-elon-musk-mjf-1715-fc0be2.jpg" width="300"> Output
A man in a black cowboy hat and sunglasses stands in front of a white car, holding a microphone and speaking into it.
from transformers import AutoModelForCausalLM, AutoTokenizer
from PIL import Image
model_id = "aloobun/rmfg"
model = AutoModelForCausalLM.from_pretrained(
model_id, trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
image = Image.open('692374.jpg')
enc_image = model.encode_image(image)
print(model.answer_question(enc_image, "Describe this image.", tokenizer))