Downloads · 30 days
53
31% of all-time downloads
PlusMinus1/omniparser-icon-caption-mlx
omniparser-icon-caption-mlx is a image-text-to-text model from PlusMinus1. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as mit.
MLX (bfloat16) conversion of microsoft/OmniParser-v2.0's iconcaption — a Florence-2 fine-tuned on UI elements, for captioning interactive icons in screenshots. Runs on Apple Silicon via mlxvlm with no PyTorch.
Downloads · 30 days
53
31% of all-time downloads
All-time downloads
172
Public
Parameters
271M
542 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors542 MB · 99%
From the Hugging Face model README
MLX (bfloat16) conversion of microsoft/OmniParser-v2.0's icon_caption — a Florence-2 fine-tuned on UI elements, for captioning interactive icons in screenshots. Runs on Apple Silicon via mlx_vlm with no PyTorch.
License: MIT (© Microsoft Corporation) — see LICENSE. This repo redistributes the original MIT-licensed weights converted to MLX format.
Needs a small no-torch patch on transformers 5.x (register florence2_language, route the image processor to CLIPImageProcessorPil). See the conversion recipe at the bottom.
# apply the florence2 no-torch patch first (see recipe), then:
from mlx_vlm import load, generate
model, processor = load("PlusMinus1/omniparser-icon-caption-mlx")
out = generate(model, processor, "<CAPTION>", image=["icon_crop.png"], max_tokens=20)
icon_caption (MIT)mlx_vlm.convert (bfloat16) + a transformers-5.x no-torch compatibility patch.