Downloads · 30 days
3.2K
47% of all-time downloads
rapid-mlx/Ling-3.0-tiny-MLX-4bit
Ling-3.0-tiny-MLX-4bit is a text generation model from rapid-mlx. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as mit.
The first MLX conversion of inclusionAI/Ling-3.0-tiny: a 7.9B-total / 1.3B-active sparse-MoE reasoner (128 experts, top-8 + 1 shared) with a KDA + MLA hybrid attention stack and 131K context, MIT licensed.
Downloads · 30 days
3.2K
47% of all-time downloads
All-time downloads
6.8K
Public
Parameters
7.9B
4.5 GB on disk
Likes
6
Public
Click a slice to open those files.
.safetensors4.4 GB · 100%
How the weights are stored.
U327.9B · 100%
From the Hugging Face model README
The first MLX conversion of inclusionAI/Ling-3.0-tiny: a 7.9B-total / 1.3B-active sparse-MoE reasoner (128 experts, top-8 + 1 shared) with a KDA + MLA hybrid attention stack and 131K context, MIT licensed.
4.2 GB at 4.507 bits/weight — it fits and runs on an 8 GB Apple Silicon Mac.
| Quantization | 4-bit, group size 64 (router kept 8-bit, short-conv weights fp) |
| Size on disk | 4.2 GB |
| Context | 131,072 tokens |
| Active parameters | 1.3B per token |
| License | MIT (inherited from the base model) |
The bailing_hybrid architecture is not in upstream mlx-lm yet — this
checkpoint is served by rapid-mlx, which ships a
verified native implementation (reference parity 1.5e-6 against the
official modeling code):
pip install -U rapid-mlx # 0.12.10 or newer
rapid-mlx serve ling-3.0-tiny-4bit
You get an OpenAI-compatible server on localhost:8000 with reasoning
(reasoning_content) and tool calling parsed natively — thinking is
controlled with chat_template_kwargs: {"enable_thinking": true} or the
model's detailed thinking on/off system-prompt switch.
Once mlx-lm gains native bailing_hybrid support, this checkpoint will
load there unchanged.
Converted with mlx_lm.convert (quantize=True, q_bits=4, q_group_size=64)
running rapid-mlx's vendored bailing_hybrid implementation
(PR #1817), which was
verified against the official modeling_bailing_moe_v3.py on identical
random weights to a max logits deviation of 1.5e-6 (full prefill) /
1.9e-6 (token-by-token incremental) before conversion. End-to-end
chat / reasoning / tool-call behaviour validated on an M2 Pro Mac mini.