Downloads · 30 days
34
5% of all-time downloads
fal/moondream2-docci-instruct
moondream2-docci-instruct is a image-text-to-text model from fal. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
Fine tuned version of moondream2 model using gokaygokay/randominstructdocci dataset. Which gives extremely detailed captions of the images.
Downloads · 30 days
34
5% of all-time downloads
All-time downloads
691
Public
Parameters
1.9B
3.7 GB on disk
Likes
9
Public
Click a slice to open those files.
.safetensors3.7 GB · 100%
From the Hugging Face model README
Fine tuned version of moondream2 model using gokaygokay/random_instruct_docci dataset. Which gives extremely detailed captions of the images.
pip install transformers timm einops bitsandbytes accelerate flash-attn
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from PIL import Image
DEVICE = "cuda"
DTYPE = (
torch.float32 if DEVICE == "cpu" else torch.float16
) # CPU doesn't support float16
revision = "3ec40c7b6b5d87bc0c51edee45e21f5f29b449d8"
tokenizer = AutoTokenizer.from_pretrained(
"fal-ai/moondream2-docci-instruct",
trust_remote_code=True,
revision=revision
)
moondream = AutoModelForCausalLM.from_pretrained(
"fal-ai/moondream2-docci-instruct",
trust_remote_code=True,
torch_dtype=DTYPE,
device_map={"": DEVICE},
attn_implementation="flash_attention_2",
revision=revision
)
moondream.eval()
image_path = "<your_image_path>"
image = Image.open(image_path).convert("RGB")
md_answer = moondream.answer_question(
moondream.encode_image(image),
"what is this picture about",
tokenizer=tokenizer,
)
print(md_answer)