Downloads · 30 days
0
coreai-community/MiniCPM-V-4.6-CoreAI
MiniCPM-V-4.6-CoreAI is a image-text-to-text model from coreai-community. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or t…
Downloads · 30 days
0
Access
Public
Updated Sep 5, 2026
Repo size
4.1 GB
Likes
0
Public
Click a slice to open those files.
.mlirb4 GB · 99%
From the Hugging Face model README
Core AI is Apple's on-device ML runtime in iOS 27 / macOS 27 and the successor to Core ML: PyTorch models are exported with Apple's coreai-torch (LLMs: coreai.llm.export) into .aimodel bundles that run on the GPU or the Neural Engine, e.g. Qwen3-8B 4-bit decodes at 94 tok/s on an M4 Max GPU, MLX 90 under the same protocol (apple-silicon-llm-bench, macOS 27 beta, 2026-06).
Mirror of
mlboydaisuke/MiniCPM-V-4.6-CoreAI— the canonical repo (CoreAI Model Zoo). Updates land there first.
On-device vision-language model for iPhone / Apple Silicon. A Core AI port of
openbmb/MiniCPM-V-4.6 — the strongest
sub-2B open VLM — running fully local on the GPU via the Core AI pipelined engine:
pick a photo, ask about it, stream the answer.
Verified on iPhone 17 Pro: image → grounded answer at ~51.5 tok/s decode, all local.
<p align="center"> <img src="https://github.com/user-attachments/assets/c4baa524-5217-4bb3-a23f-b0acd6249bd4" width="300" alt="MiniCPM-V 4.6 on iPhone — a fridge photo becomes recipe ideas, fully on-device in CoreAIChat"> </p> <p align="center"><em>Fridge photo → recipe ideas, fully on-device on an iPhone 17 Pro (CoreAIChat).</em></p>MiniCPM-V-4.6 (1.3B) = a SigLIP So400m vision tower (980px / patch 14 / 27 layers, with a
window-attention insert-merger @ layer 6 + a downsample-MLP merger → ÷16 = 64 visual tokens per
448px slice) + a Qwen3.5-hybrid text backbone (qwen3_5_text: 0.8B, 24 layers, GatedDeltaNet
linear attention ×3 : full attention ×1, head_dim 256, vocab 248094, tied head). Connector =
2×2 spatial merges + MLP, spliced into the text embeddings at <image> positions (masked_scatter).
Recommended (optimized, 2026-06-25):
| path | what | dtype | size |
|---|---|---|---|
gpu-pipelined/minicpmv46_vlm_decode_int8hu/ | VLM text decoder (input_ids → logits + static image_embeds[64,1024]; in-graph gather ids ≥ V ? image_embeds[ids-V] : embed[ids]) | int8 body + untied int8 head | ~1.2 GB |
gpu-pipelined/minicpmv46_vision_int8lin/ | fixed-grid SigLIP vision encoder (pixel_values[1,3,448,448] → image_features[64,1024]) | int8 | ~0.6 GB |
The int8 head quantizes the big-vocab LM head (fp16 in int8lin = ~half the per-token read) → +48% decode
on iPhone 17 Pro (46→68 tok/s). The int8 vision halves the encoder's size (the encode is compute-bound, so
this is a size/memory win); pair it with a one-shot vision-graph warmup at load to hide the ~2.7 s first-photo
cold compile. Original …_int8lin decoder + fp16 minicpmv46_vision remain for compatibility.
The decoder is a complete qwen3.5-hybrid text LLM when image_embeds is zero — same bundle, no image needed.
The pipelined engine knows nothing about images. The whole multimodal state rides the
static-input hook (image_embeds buffer) + an id-space trick — the graph stays ids + positions → logits:
x/127.5−1) and writes
image_embeds [64,1024] into one owned MTLBuffer the engine binds on every step.<|image_pad|> ids are rewritten to extension ids V + slot (slot 0..63).
In-graph: embed = ids < V ? table[ids] : image_embeds[ids − V].Simpler than the Qwen3-VL port (no deepstack, no M-RoPE).
image_embeds every step, which dilutes the head gain; the text core
alone is 46 → 68 = +48%). ~64–70 tok/s in practice by device temperature. · M4 Max text core ~224 tok/s
(llm-benchmark), engine cold-spec ~2–4 s, ~1.5 GB resident (jetsam-safe).apps/CoreAIChat and the standalone MiniCPMVLM app have a MiniCPM-V 4.6 mode with a photo picker:
pick an image, ask, stream. The vision tower runs once per image (~hundreds of ms); each turn re-prefills (S=1).
Conversion + gates: see coreai-model-zoo / minicpm-v-4.6.
License: Apache-2.0 (inherited from openbmb/MiniCPM-V-4.6).