Downloads · 30 days
596
28% of all-time downloads
LiquidAI/LFM2.5-VL-3B-MLX-8bit
LFM2.5-VL-3B-MLX-8bit is a image-text-to-text model from LiquidAI. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as other.
<div align="center" <img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%;"
Downloads · 30 days
596
28% of all-time downloads
All-time downloads
2.1K
Public
Parameters
3.1B
3.7 GB on disk
Likes
17
Public
Click a slice to open those files.
.safetensors3.7 GB · 100%
How the weights are stored.
U322.7B · 86%
From the Hugging Face model README
MLX export of LFM2.5-VL-3B for Apple Silicon inference.
LFM2.5-VL-3B is a vision-language model built on the LFM2.5-2.6B backbone with a SigLIP2 NaFlex vision encoder (400M). It supports OCR, document comprehension, multilingual vision understanding, bounding box prediction, and function calling.
uv run --with mlx-vlm mlx_vlm.generate --model LiquidAI/LFM2.5-VL-3B-MLX-8bit --max-tokens 100 --temperature 0.2 --image https://placecats.com/neo/300/200 --prompt "how many animals are in the picture?"
from mlx_vlm import apply_chat_template, generate, load
from mlx_vlm.utils import load_image
model, processor = load("LiquidAI/LFM2.5-VL-3B-MLX-8bit")
image = load_image("https://placecats.com/neo/300/200")
messages = [
{
"role": "user",
"content": [
{"type": "image"},
{"type": "text", "text": "What do you see in this image?"},
],
}
]
prompt = apply_chat_template(
processor,
model.config,
messages,
add_generation_prompt=True,
num_images=1,
)
result = generate(
model,
processor,
prompt,
[image],
temp=0.2,
top_k=50,
repetition_penalty=1.0,
verbose=True,
)
print(result.text)