Downloads · 30 days
352
9% of all-time downloads
mlx-community/Inkling-mlx-4bit
Inkling-mlx-4bit is a image-text-to-text model from mlx-community. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as apache-2.0.
An MLX 4-bit build of Thinking Machines' Inkling (975B-total / 41B-active MoE), quantized from the BF16 checkpoint for Apple Silicon. Omni: keeps the text decoder plus the vision (HMLP) and audio (dMel) towers. Self-c…
Downloads · 30 days
352
9% of all-time downloads
All-time downloads
3.8K
Public
Parameters
947B
1.1 TB on disk
Likes
7
Public
Click a slice to open those files.
.safetensors1.1 TB · 100%
How the weights are stored.
U32936B · 99%
From the Hugging Face model README
An MLX 4-bit build of Thinking Machines' Inkling (975B-total / 41B-active MoE),
quantized from the BF16 checkpoint for Apple Silicon. Omni: keeps the text decoder
plus the vision (HMLP) and audio (dMel) towers. Self-contained. This bundles the loader
package (inkling_mlx/), so it runs with just mlx + mlx-lm + transformers.
The higher-fidelity sibling of mlx-community/Inkling-mlx-2bit:
~548 GB (vs ~315 GB at 2-bit), trading size for sharper text and multimodal quality.
experts_only)# pip install mlx mlx-lm transformers
# download this repo (it already includes the inkling_mlx/ loader)
from inkling_mlx.load import load
from inkling_mlx.generate import greedy_generate
from transformers import AutoTokenizer
model, config = load("./Inkling-mlx-4bit")
tok = AutoTokenizer.from_pretrained("./Inkling-mlx-4bit", trust_remote_code=True)
ids = tok("The capital of France is")["input_ids"]
print(tok.decode(greedy_generate(model, config, ids, max_new_tokens=64)))
At ~548 GB this doesn't fit resident on one Mac but an MoE only fires 6 of 256 experts per token. So keep the always-needed weights in RAM (attention, shared experts, embeddings, norms, router, vision / audio towers) and page the routed experts from SSD on demand, letting the OS page cache hold the hot ones. This runs the omni build on a single Mac Studio.
| Prompt | Generated continuation | Notes |
|---|---|---|
The capital of France is | Paris. The capital of Italy is Rome. The capital of Spain is Madrid. The capital of Russia is Moscow. | ~0.33 tok/s, SSD expert-offload |
Tool: github.com/huckiyang/mlx-moe-offload
(MIT) — a drop-in over mlx-lm's SwitchGLU.
pip install mlx mlx-lm scipy pillow
git clone https://github.com/huckiyang/mlx-moe-offload && cd mlx-moe-offload && pip install -e .
# 1) one-time: repack the stacked experts into a per-expert SSD store
python -m mlx_moe_offload.repack --build /path/Inkling-mlx-4bit --out /path/Inkling-mlx-4bit-offload
OFF=/path/Inkling-mlx-4bit-offload
# 2) generate — text, image, or audio (omni)
python examples/inkling_omni.py --offload-dir $OFF --prompt "The capital of France is"
python examples/inkling_omni.py --offload-dir $OFF --image cat.png --prompt "What is in this image?"
python examples/inkling_omni.py --offload-dir $OFF --audio q.wav --prompt "Transcribe in English:"
Decode speed is gated by the routing cache hit-rate and SSD bandwidth, please check store.stats().
thinkingmachines/Inkling.inkling_mlx/ loader is vendored from
PipeNetwork/inkling-mlx (Apache-2.0,
Copyright 2026 PipeNetwork) and David.LICENSE + THIRD_PARTY_NOTICES.md are included in this repo.