Downloads · 30 days
32
11% of all-time downloads
OsaurusAI/Nanbeige4.2-3B-MXFP8
Nanbeige4.2-3B-MXFP8 is a text generation model from OsaurusAI. Use it when you need the model to write or continue text. It is set up for mlx. The card lists the license as apache-2.0.
<p align="center"<a href="https://osaurus.ai"<img src="./osaurus-x-banner.png" alt="Osaurus AI"</a</p
Downloads · 30 days
32
11% of all-time downloads
All-time downloads
294
Public
Parameters
4.2B
4.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.3 GB · 100%
How the weights are stored.
U324.2B · 100%
From the Hugging Face model README
MXFP8 build of Nanbeige/Nanbeige4.2-3B — a 4.17B-parameter Looped Transformer reasoning model (en + zh, 256K context), quantized for Apple Silicon with uniform MXFP8 (e4m3) weights. This is format coverage — see the fidelity table before choosing it.
⚠️ This architecture needs a loader that knows about the loop
num_loops = 2: the same 22 decoder layers run twice over shared weights, for an effective depth of 44. Two consequences a generic loader gets wrong:
- The KV cache has 44 slots, not 22 (slot =
layer_idx + loop_idx * num_hidden_layers). A 22-slot cache does not crash — it emits fluent, confident, wrong tokens from the first one. This is verified with a negative control, not theory.- The final norm runs at the end of every loop, not once at the end (
skip_loop_final_norm = false). Loop 0's normed output is loop 1's input.
mlx_lm0.31.x has nonanbeigemodel class, somlx_lm.generatealone will not load this bundle. Use a runtime that implements the loop (see Usage).
| Field | Value |
|---|---|
| Source | Nanbeige/Nanbeige4.2-3B @ fab06df (Apache-2.0) |
| Architecture | nanbeige — 22 layers × 2 loops (effective depth 44), 4.17B params, 256K ctx |
| On-disk size | 4.0 GB (1 shard) |
| Quantization | uniform 8-bit MX, no per-module overrides, group size 32 |
| Norms, router bias | fp16 passthrough — plain Llama RMSNorm, no +1 shift |
| Attention | 48 heads / 8 KV heads (GQA), head_dim 128 — note n_heads × head_dim (6144) ≠ hidden_size (3072) |
| RoPE | θ = 7e7, NeoX half-rotation, full 128 dims, no scaling |
| Modality | text-only (verified from the tensor index) |
| Metric | Thinking on | Thinking off |
|---|---|---|
| Decode | 34.6 tok/s | 27.9 tok/s |
| Peak memory | 5.3 GB | 4.5 GB |
Fidelity vs the bf16 source (5-prompt logit sweep): top-1 agreement 4/5, mean KL 0.1446, max KL 0.6844.
| Bundle | Size | Top-1 agreement vs bf16 | Mean KL | Max KL | Decode (thinking) |
|---|---|---|---|---|---|
Nanbeige4.2-3B-JANG_6M | 3.6 GB | 5/5 | 0.0010 | 0.0030 | 29.3 tok/s |
Nanbeige4.2-3B-JANG_4M | 2.9 GB | 5/5 | 0.0192 | 0.0398 | 44.7 tok/s |
Nanbeige4.2-3B-MXFP8 | 4.0 GB | 4/5 | 0.1446 | 0.6844 | 34.6 tok/s |
Both JANG affine profiles beat MXFP8 on fidelity while being smaller — the opposite of what the bit counts suggest. MXFP8's e4m3 elements carry ~3 mantissa bits each, so "8-bit MX" is not strictly better than 6-bit or 4-bit affine with a per-group scale and bias on this weight distribution. It showed up in behaviour too: in the multi-turn gate MXFP8 dated Tokyo's capital move to 1936, where both JANG builds said 1868.
<think>\n; only enable_thinking=False prefills a closed <think>\n\n</think>\n\n.preserve_thinking controls whether previous turns' reasoning is kept. The template's default is to preserve; the vendor recommends False for general chat and True for multi-turn tool use and code-agent workflows.tool_call_format="xml" (the vendor's recommended format); json is supported for compatibility.<|im_start|> (id 166100 = bos_token) and the tokenizer's post-processor prepends another. Tokenize the rendered template with add_special_tokens=False.eos_token_id = 166101 (<|im_end|>).generation_config.json, matching the model card): temperature 0.6, top_p 0.95, top_k 20. The vendor suggests temperature 1.0 for agentic and tool-use tasks. The same values are stamped in jang_config.chat.sampling_defaults, and the two files are checked against each other at build time.The bundle is standard MLX safetensors with a per-module {bits, group_size, mode} map in config.json[quantization] — any loader must honor those overrides. It needs the nanbeige looped model class, which registers into mlx_lm:
from jang_tools.nanbeige import mlx_register # registers the looped nanbeige class
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tok = load("OsaurusAI/Nanbeige4.2-3B-MXFP8")
prompt = tok.apply_chat_template(
[{"role": "user", "content": "Which number is bigger, 9.11 or 9.8?"}],
add_generation_prompt=True, tokenize=False,
)
ids = tok.encode(prompt, add_special_tokens=False) # template already emits BOS
print(generate(model, tok, prompt=ids, max_tokens=1024,
sampler=make_sampler(temp=0.6, top_p=0.95, top_k=20)))
Osaurus and vMLX runtime support for the looped architecture is in progress; until it lands, use the path above.
Quantized and verified by Jinho Jang ([email protected]). Base model © Nanbeige, Apache-2.0 (inherited).