Downloads · 30 days
328
20% of all-time downloads
mlx-community/Inkling-NVFP4-mlx-4bit
Inkling-NVFP4-mlx-4bit is a text generation model from mlx-community. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
An MLX 4-bit build of the encoder-free text backbone of Thinking Machines' Inkling (975B-total / 41B-active MoE), for running natively on Apple Silicon with mlx-lm.
Downloads · 30 days
328
20% of all-time downloads
All-time downloads
1.7K
Public
Parameters
947B
581 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors581 GB · 100%
How the weights are stored.
U32913B · 96%
From the Hugging Face model README
An MLX 4-bit build of the encoder-free text backbone of Thinking Machines' Inkling
(975B-total / 41B-active MoE), for running natively on Apple Silicon with
mlx-lm.
This is created for people using a two Apple Mac Studio M3 Ultra with 192/512 GB. (for one Mac Studio you need 2-bit!)
Community note This is not fullly numerically-verified conversion, shared to see whether anyone can load/run a model this large on Apple Silicon and to gather feedback. Expect rough edges; please open a discussion with results (or failures).
thinkingmachines/Inkling-NVFP4 (NVFP4) → dequantized → MLX affine 4-bit
(group size 64). Only the routed MoE experts are quantized; everything else is bf16.from mlx_lm import load, generate
model, tokenizer = load("mlx-community/Inkling-NVFP4-mlx-4bit")
print(generate(model, tokenizer, prompt="The capital of France is", max_tokens=64))
The custom model class lives in the conversion repo (
models/inkling_mlx.py); until it's registered inmlx-lm, load via that module'sload().Blog: https://huckiyang.github.io/blog/inkling-audio-design.html