Downloads · 30 days
12
18% of all-time downloads
Aa09876/SmallThinker-4BA0.6B-Instruct-mlx-4bit
SmallThinker-4BA0.6B-Instruct-mlx-4bit is a text generation model from Aa09876. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
An MLX 4-bit (group size 64, affine) conversion of PowerInfer/SmallThinker-4BA0.6B-Instruct: a 4B-total / roughly 0.6B-active dense-attention Mixture-of-Experts model (32 experts, top-4, ReLU-gated), non-thinking Inst…
Downloads · 30 days
12
18% of all-time downloads
All-time downloads
66
Public
Parameters
4B
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 99%
How the weights are stored.
U324B · 100%
From the Hugging Face model README
An MLX 4-bit (group size 64, affine) conversion of PowerInfer/SmallThinker-4BA0.6B-Instruct: a 4B-total / roughly 0.6B-active dense-attention Mixture-of-Experts model (32 experts, top-4, ReLU-gated), non-thinking Instruct.
Stock mlx_lm does not yet ship a smallthinker architecture, so this model will not
load with an unmodified install. Until the upstream mlx-lm PR lands, copy the included
smallthinker.py into your mlx_lm/models/ directory:
cp smallthinker.py "$(python -c 'import mlx_lm,os;print(os.path.join(os.path.dirname(mlx_lm.__file__),"models"))')/"
Then:
python -m mlx_lm generate \
--model Aa09876/SmallThinker-4BA0.6B-Instruct-mlx-4bit \
--prompt "Explain a mixture-of-experts model in two sentences."
The model uses the native Hermes tool-call format, <tool_call>{json}</tool_call>, and
works with servers that select a Hermes tool parser (for example rapid-mlx
--tool-call-parser hermes, no reasoning parser).
PowerInfer/SmallThinker-4BA0.6B-Instruct (Apache-2.0), pinned revision b51db6dmlx_lm implementation of the custom
SmallThinkerForCausalLM (ReLU-gated packed-expert MoE, GQA with 12 query / 2 KV heads,
and the model's pre-attention router-input routing). It was numerically parity-checked
against the reference PyTorch model: BF16 greedy generation matched token-for-token,
router top-k selection matched 100% (tie-adjusted), and per-tensor error stayed at the
BF16 rounding floor.Apache-2.0, inherited from the base model. Attribution: PowerInfer / IPADS (SmallThinker). This repository redistributes a quantized derivative under the same license.