Downloads · 30 days
2
5% of all-time downloads
BINOMDA/OMDA-PROMPTER
OMDA-PROMPTER is a machine learning model from BINOMDA. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This model is a core member of the OMDA Family by BINOMDA. It bridge the gap between visual perception and detailed linguistic description.
Downloads · 30 days
2
5% of all-time downloads
All-time downloads
37
Public
Parameters
250M
4.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors5 GB · 77%
From the Hugging Face model README
This model is a core member of the OMDA Family by BINOMDA. It bridge the gap between visual perception and detailed linguistic description.
from transformers import AutoModelForCausalLM, AutoTokenizer, AutoProcessor
from PIL import Image
import torch
# Load the specialized OMDA architecture
model = AutoModelForCausalLM.from_pretrained("BINOMDA/OMDA-PROMPTER", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("BINOMDA/OMDA-PROMPTER")
processor = AutoProcessor.from_pretrained("google/siglip-base-patch16-224") # The vision processor is for SigLIP
# Generate description
image = Image.open("your-image.jpg").convert("RGB")
pixel_values = processor(images=image, return_tensors="pt").pixel_values.to(model.device)
generated_ids = model.generate(pixel_values, max_new_tokens=800, pad_token_id=tokenizer.pad_token_id, eos_token_id=tokenizer.eos_token_id)
description = tokenizer.decode(generated_ids[0], skip_special_tokens=True)
print(description)
Training Data: Curated dataset of images with detailed descriptions
Max Sequence Length: 2048 tokens
Training Epochs: 5
Learning Rate: 2e-5