Downloads · 30 days
219
22% of all-time downloads
wabibito/Onyx-maple-preview-2bit
Onyx-maple-preview-2bit is a machine learning model from wabibito. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for mlx. The card lists the license as mit.
MLX 2-bit repack of deepgrove/maple-preview (20B-total / ~1B-active ternary-weight MoE reasoner, MIT license) for the Onyx app's own inference engine (OnyxLLM).
Downloads · 30 days
219
22% of all-time downloads
All-time downloads
984
Public
Parameters
20.2B
6.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6.8 GB · 100%
How the weights are stored.
U3220.2B · 100%
From the Hugging Face model README
MLX 2-bit repack of deepgrove/maple-preview (20B-total / ~1B-active ternary-weight MoE reasoner, MIT license) for the Onyx app's own inference engine (OnyxLLM).
The repack is lossless for every projection. The source BF16 master weights are per-row ternary {-s, 0, +s}; naive min/max 2-bit affine quantization cannot represent the zeros, so each row is packed manually as MLX affine 2-bit with scale = s, bias = -s (dequantization reproduces the BF16 master bit-exactly; verified per tensor during conversion). Attention q/k/v/o and all 3×256 expert projections per layer are packed this way (168 tensors, all lossless). Embeddings and lm_head are continuous in the source and are quantized 8-bit affine (group 64). The fp32 router gates and norms stay BF16.
Experts are fused per layer into stacked mlp.switch_mlp.{gate,up,down}_proj tensors
([256, out, in]) for single-gather_qmm MoE dispatch.
Conversion: Onyx session 2026-08-05. Not affiliated with deepgrove; see LICENSE (MIT) for the upstream terms.