Downloads · 30 days
32
35% of all-time downloads
KaedeTai/VisionPsy-Nano-460M-MLX
VisionPsy-Nano-460M-MLX is a image-text-to-text model from KaedeTai. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as apache-2.0.
MLX port of qvac/VisionPsy-Nano-460M, a compact 460M-parameter vision-language model from Tether AI Research, converted to run natively on Apple Silicon.
Downloads · 30 days
32
35% of all-time downloads
All-time downloads
91
Public
Parameters
507M
1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1 GB · 100%
From the Hugging Face model README
MLX port of qvac/VisionPsy-Nano-460M, a compact 460M-parameter vision-language model from Tether AI Research, converted to run natively on Apple Silicon.
Measured across 7 images x 5 prompts, 64 max new tokens, greedy decode:
| Metric | Standard | Flash |
|---|---|---|
| Avg decode tok/s | 99 | 152 |
| Median decode tok/s | 90 | 157 |
| Avg peak GPU memory | 2.64 GB | 2.64 GB |
| Load time | ~0.4 s | ~0.7 s |
Per-prompt-type medians (Standard):
| Prompt type | Median tok/s | Example |
|---|---|---|
| Describe (EN, 1 sentence) | 158.7 | "A smiling man in a white lab coat gestures with his right hand..." |
| What text appears? | 90.3 | "OICOMELVANG" |
| Count objects/people | 40.0 | "There are 3 people in the image." |
| Main subject | 59.2 | "The main subject is a man wearing a white lab coat." |
| Describe (ZH) | 130.4 | "他說:"OICOMELVANG, 25158"" |
Full 70-run matrix (Standard + Flash) is at github.com/KaedeTai/mlx-video/tree/visionpsy-mlx-port.
This repo uses the mlx-vlm-style layout (text_config + vision_config top-level, language_model.* / vision_tower.* / multi_modal_projector.* tensor prefixes) so it slots cleanly into mlx-vlm once a visionpsy_nano handler lands there. Until then, load it via the MLX port bundled in mlx-video (branch visionpsy-mlx-port):
pip install mlx safetensors transformers pillow
git clone -b visionpsy-mlx-port https://github.com/KaedeTai/mlx-video.git
cd mlx-video
from huggingface_hub import snapshot_download
from mlx_video.models.visionpsy_nano import load_visionpsy_nano
from mlx_video.models.visionpsy_nano.processor import load_processor
from PIL import Image
# Snapshot from HF (or point at your local folder)
path = snapshot_download("KaedeTai/VisionPsy-Nano-460M-MLX")
model, cfg = load_visionpsy_nano(path)
proc = load_processor(path, cfg=cfg)
img = Image.open("photo.jpg").convert("RGB")
batch = proc("Describe this image in one sentence.", image=img)
tokens = list(model.generate(
batch["input_ids"],
pixel_values=batch["pixel_values"],
image_token_id=batch["image_token_id"],
max_new_tokens=64,
eos_token_id=proc.tokenizer.eos_token_id,
))
print(proc.decode(tokens, skip_special_tokens=True))
Note: the port's load_visionpsy_nano reads the repacked config via a compat shim; the original _original_config block is retained inside config.json for round-tripping.
decoder.* / vision_encoder.* / MP.* to language_model.* / vision_tower.* / multi_modal_projector.* to match mlx-vlm conventions.lm_* / vit_* keys refactored into nested text_config / vision_config blocks with standard field names (hidden_size, num_hidden_layers, etc.).decoder.rotary_embd.* buffers removed — MLX's nn.RoPE computes frequencies on the fly.If you use these weights, please also cite the original QVAC release and the base model authors.