Downloads · 30 days
51
6% of all-time downloads
NullSense/Nanbeige4.2-3B-NVFP4-FP8-LoopShield
Nanbeige4.2-3B-NVFP4-FP8-LoopShield is a text generation model from NullSense. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
Mixed-precision quantization of Nanbeige/Nanbeige4.2-3B (looped transformer, 22 layers × 2 passes/token): FP8-dynamic on attention, downproj, and the first/last 3 layers' MLP; NVFP4 (weight-only, group-16) on middle-l…
Downloads · 30 days
51
6% of all-time downloads
All-time downloads
833
Public
Parameters
3.7B
4.7 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors4.7 GB · 100%
How the weights are stored.
F8_E4M32.2B · 58%
From the Hugging Face model README
Mixed-precision quantization of
Nanbeige/Nanbeige4.2-3B (looped
transformer, 22 layers × 2 passes/token): FP8-dynamic on attention, down_proj, and
the first/last 3 layers' MLP; NVFP4 (weight-only, group-16) on middle-layer
gate/up_proj. 4.5 GB. format: mixed-precision (compressed-tensors).
Naive uniform NVFP4 on this looped architecture loses 8 points of GSM8K strict (quantization error compounds across the two loop passes — consistent with LoopQ, arXiv:2605.16343, the only prior looped-LLM PTQ study, which tested INT only; these are, as far as I can tell, the first FP4-family numbers for a looped LLM). The protection recipe that llama.cpp's quant mixes, Unsloth's ablations, and llmcompressor's own non-uniform example all converge on recovered it completely:
| step | GSM8K strict (n=100) |
|---|---|
| uniform NVFP4A16 | 81 |
| + FP8 attention & down_proj (all layers) | 84 |
| + FP8 gate/up in layers 0-2 & 19-21 (this repo) | 89 — bf16/FP8 parity |
The edge-band ratio follows APEX's ablation (~12.5% of depth per side); the tensor priority (down_proj > attention > gate/up) matches llama.cpp's quant mixes, Unsloth's sensitivity ablations, and llmcompressor's own non-uniform example.
Super-weight verification (2026-07-23): this model's super weight (the
single most load-bearing scalar, Apple 2411.07191)
sits at layers.1.mlp.down_proj.weight[1252, 6883] (largest weight in its tensor,
11.4x p99.99; drives a 26,752-magnitude activation spike, 1,300x the median). This
recipe protects both the weight (all-layer FP8 down_proj) and its production path
(L1 gate/up in the FP8 edge band) — verified by direct scan, not assumed. One
looped-arch novelty from the scan: the spike is pass-asymmetric (26,752 on loop
pass 1 of 2; 1,352 on pass 2).
| bench | mode | bf16 original | FP8-Dynamic | this repo |
|---|---|---|---|---|
| GSM8K strict (n=100) | thinking | not measured | 89 | 89 |
| GSM8K flexible | thinking | not measured | 93 | 96 |
| IFEval prompt-strict (n=250) | non-thinking | not measured | 76.4 | 76.0 |
| IFEval inst-strict | non-thinking | not measured | 83.5 | 83.3 |
| MMLU-Pro (25/category) | non-thinking | not measured | 62.6 | 64.9 |
| BBH CoT few-shot | non-thinking | not measured | 64.9 | 59.5 |
| MultiHop-RAG (gold evidence, n=248) | non-thinking | not measured | 74.2 | 74.6 |
| Blind-judge summarization (closed, same-judge pair, n=46) | non-thinking | n/c (judged in a separate pass; scores only comparable within a pass) | 4.57 | 4.41 (coverage −0.24, ~1.5σ) |
| Judged faithfulness (closed) | non-thinking | n/c | 4.90 | 4.87 |
| Judged fabrication / leaks (closed) | non-thinking | 0% / 0% | 2% / 0% | 2% / 0% |
| Dictation-rewrite taxonomy (closed) | non-thinking | 18/20 | 18/20 | 17/20 |
| JSON parse rate (closed, /48) | non-thinking | 44 | 42 | 46 |
The trade: reasoning at full parity, best-in-family structured-output reliability, small summarization-coverage cost. If you want maximum quality use my FP8-Dynamic; if you want the smallest artifact that keeps reasoning intact on this architecture, use this one.
Decode tok/s, single stream, per-workload best speculative config:
| workload | spec decode | bf16 | FP8-Dynamic | this repo |
|---|---|---|---|---|
| freeform / chat / agent | off | 96 | 154 | 159 |
| summarize / RAG (2k+ ctx prompts) | ngram, 8 tok | 136 | 206 | 202 |
Two workload anchors, not an ISL sweep; decode speed shifts with context length, batch size, and attention backend. The mid-MLP NVFP4 weights serve via the Marlin W4A16 kernel; the FP8 tensors via the FP8 path; the small freeform edge over FP8-Dynamic comes from the lighter weight reads, the small summarize deficit from mixed-kernel overhead under the ngram verify batch.
Arch not yet upstream (vLLM PR #49433); install the bundled plugin first:
pip install --no-deps ./vllm_plugin
vllm serve <this-repo> --trust-remote-code \
--max-model-len 65536 --kv-cache-dtype fp8 \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml
Thinking/sampling notes as in the base model: thinking ON by default
(chat_template_kwargs.enable_thinking=false to disable), T=0.6/top_p=.95/top_k=20
defaults ship in generation_config.json, T=1.0 for agentic use.
Download just this artifact:
hf download NullSense/Nanbeige4.2-3B-NVFP4-FP8-LoopShield --local-dir Nanbeige4.2-3B-NVFP4-FP8-LoopShield
g_fp8 = dict(FP8_DYNAMIC)
g_fp8["targets"] = ["re:.*self_attn\\.q_proj.*", "re:.*self_attn\\.k_proj.*",
"re:.*self_attn\\.v_proj.*", "re:.*self_attn\\.o_proj.*",
"re:.*down_proj.*"] + \
[f"re:model\\.layers\\.{i}\\.mlp\\.(gate|up)_proj.*" for i in (0,1,2,19,20,21)]
g_fp4 = dict(NVFP4A16)
g_fp4["targets"] = [f"re:model\\.layers\\.{i}\\.mlp\\.(gate|up)_proj.*" for i in range(3,19)]
QuantizationModifier(config_groups={"group_0": g_fp8, "group_1": g_fp4}, ignore=["lm_head"])
GPTQ/AWQ were NOT usable on this architecture in llmcompressor 0.12: GPTQ's Hessian inversion fails on all 154 modules (root cause unknown — the double-fire hook was checked and is not the cause); AWQ lacks arch mappings. Reported upstream: #2952, #2953.
modeling_nanbeige.py carries two
one-line transformers-5 compat patches (rope key, tied-weights type).


| artifact | size | tok/s (freeform / summ) | pick when |
|---|---|---|---|
| FP8-Dynamic | 4.9 GB | 154 / 206 | default: no measured quality loss on any gate. Recommended. |
| NVFP4-FP8-LoopShield (this repo) | 4.5 GB | 159 / 202 | smallest artifact that keeps reasoning at FP8 parity; best JSON reliability. Recommended for tight VRAM. |
| NVFP4A16 | 3.6 GB | 175 / 220 | fastest; summarization/extraction only (reasoning drops 8 GSM8K points). Not for math/agentic. |
| EAGLE3 draft | +1.5 GB | +12-41% decode | add-on speculator for any of the above; thinking-aware retrain (2026-07-24), thinking-mode acceptance 0.41, lossless. Serve with TRITON_ATTN. |
All three serve identically (same plugin, same flags); only the checkpoint differs.
Comparison chain: the columns here use my FP8-Dynamic quant as reference; the FP8 card carries the bf16-original matrix linking the chain back to the unquantized model.
vllm_plugin/ registers it out-of-tree)Public, reproducible (lm-eval-harness local-chat-completions, exact configs in each row's annotation):
Closed/personal harnesses (not publicly reproducible — my own serving-workload gates; treat as relative signals between artifacts in THIS family, not cross-model scores):
@misc{peciukonis2026nanbeige42loopshield,
author = {Pe{\v{c}}iukonis, Matas (NullSense)},
title = {NVFP4-FP8-LoopShield: loop-aware mixed-precision NVFP4 quantization of Nanbeige4.2-3B},
year = {2026},
howpublished = {Hugging Face},
url = {https://huggingface.co/NullSense/Nanbeige4.2-3B-NVFP4-FP8-LoopShield},
note = {Placement recipe recovering GSM8K 81->89 on a looped/weight-shared LLM; first FP4-family results for the architecture class.}
}