Downloads · 30 days
32
100% of all-time downloads
HandleCoding/Qwen3.8-27B-nvfp4-NInfer
Qwen3.8-27B-nvfp4-NInfer is a image-text-to-text model from HandleCoding. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for ninfer. The card lists the license as apache-2.0.
This model card is the version-controlled source for neroued/Qwen3.8-27B-nvfp4-NInfer.
Downloads · 30 days
32
100% of all-time downloads
All-time downloads
32
Public
Repo size
47.4 GB
Likes
0
Public
Click a slice to open those files.
.ninfer23.7 GB · 100%
From the Hugging Face model README
This model card is the version-controlled source for neroued/Qwen3.8-27B-nvfp4-NInfer.
The repository contains the registered NVFP4 weight profile of
Qwen3.8-27B. It combines the official BF16 checkpoint
with the fixed packed Text weights from
unsloth/Qwen3.8-27B-NVFP4 in the native
NInfer .ninfer artifact format. The artifact is intended
only for NInfer; it is not a Transformers checkpoint, Safetensors distribution, or GGUF file.
This is a second weight profile for the existing qwen3_8_27b target, not a separate model target.
The version-2 artifact identity selects the NVFP4 binder and execution leaves. The nvfp4 weights
ID names the complete registered profile rather than claiming that every matrix has one format:
Text layers 0–55 use NVFP4 MLP weights, while the token embedding, attention input/output
projections, GDN Q/K/V/Z and output projections, full output head, and Text layers 56–63 MLP weights
use row-scaled FP8. BF16 control weights and the registered MTP and Vision allocations are retained.
| Field | Value |
|---|---|
| Filename | qwen3_8_27b_nvfp4.ninfer |
| Size | 23,719,496,192 bytes (22.09 GiB) |
| SHA-256 | 552c374c685dce302603b95fbe940fb04243c0cd44c083efc644ad3d980d462c |
| Container version | 2 |
| NInfer model ID | qwen3.8-27b |
| NInfer weights ID | nvfp4 |
| NInfer target key | qwen3_8_27b |
| Stored objects | 1,190 (1,184 tensors and 6 resources) |
| NVFP4 tensors | 112 |
| Row-scaled FP8 tensors | 146 |
The file contains the registered Text, Vision, MTP, DFlash2, optimized proposal-head, tokenizer, chat-template, generation, and media-processor objects required by NInfer. Source-derived NVFP4 and FP8 words are preserved without decode and requantization; only the official BF16 token embedding is encoded locally as row-scaled FP8.
Verify a downloaded file with:
printf '%s %s\n' \
'552c374c685dce302603b95fbe940fb04243c0cd44c083efc644ad3d980d462c' \
'qwen3_8_27b_nvfp4.ninfer' | sha256sum --check
This release includes the complete DFlash2 companion weights from
z-lab/Qwen3.8-27B-DFlash2 at revision
50307d4c4cde6860d4eee73e2547cd786fe8e8a4. Select
--spec dflash2 --draft-tokens 7 --lm-head-draft; draft counts 1..15 are supported.
DFlash2 requires the runtime revision listed below. Existing performance and evaluation tables
retain their stated MTP configurations and revisions.
385b30ce
or later, built from source;sm_120a);NInfer does not provide an install target or packaged binary. See the repository README for source-build dependencies.
hf download neroued/Qwen3.8-27B-nvfp4-NInfer \
qwen3_8_27b_nvfp4.ninfer \
--local-dir models
./build/apps/ninfer models/qwen3_8_27b_nvfp4.ninfer \
--prompt "Explain prefill and decode in three sentences." \
--max-context 32768 \
--max-new 8192 \
--kv-dtype fp8 \
--spec mtp --draft-tokens 3 \
--lm-head-draft
For images, videos, and structured chat history, see the CLI guide.
./build/apps/ninfer-serve models/qwen3_8_27b_nvfp4.ninfer \
--host 127.0.0.1 \
--port 8080 \
--max-context 240000 \
--kv-capacity 240000 \
--max-concurrency 2 \
--kv-dtype fp8 \
--device-state-slots 2 \
--host-state-slots 8 \
--host-kv-mib 8192 \
--spec mtp --draft-tokens 3 \
--lm-head-draft \
--preserve-thinking
Each request has a 240,000-token logical ceiling. The shared 240,000-token Device KV pool admits two active requests when their combined completion reservations fit; either request may use the full pool while running alone. Two extra Device checkpoint slots, eight pinned Host State slots, and 8 GiB of pinned Host KV retain reusable continuations under resource pressure.
See the HTTP serving guide for the API surface and the resource scheduling reference for cache and admission semantics.
The artifact supports:
--spec dflash2 --draft-tokens 7, optionally --lm-head-draft);The MTP0 measurements below were collected at NInfer revision
f08597d,
and the MTP3 measurements at revision
32c9881.
Both campaigns used one NVIDIA GeForce RTX 5090, CUDA 13.1 compile/runtime, CUDA driver API 13.3,
stochastic sampling, INT8 group-64 KV, CUDA Graphs, a 1,024-token prefill chunk, and prefix reuse
disabled. MTP0 used no speculative backend and a 262,144-token context limit; MTP3 used a
131,072-token per-request context limit and three draft tokens.
The fixed corpus contains three long-reasoning and twelve cross-scenario fixtures with five seeds each, for 75 requests. Every concurrency point starts a fresh server and uses the same shuffle seed and ordered HTTP send sequence. C=1 is the serial single-request corpus. Makespan includes prefill, decode, workload transitions, and final drain.
| C | Requests | Decode tokens | Makespan | Requests/s | Corpus decode (tok/s) | Avg batch | MTP acceptance | Speedup |
|---|---|---|---|---|---|---|---|---|
| 1 | 75 | 752,160 | 4,670.27 s | 0.0161 | 161.1 | 1.00 | 60.8% | 1.00× |
| 2 | 75 | 739,951 | 2,510.78 s | 0.0299 | 294.7 | 1.98 | 59.2% | 1.86× |
| 4 | 75 | 713,384 | 1,647.74 s | 0.0455 | 432.9 | 3.29 | 58.0% | 2.83× |
| 8 | 75 | 723,602 | 2,164.90 s | 0.0346 | 334.2 | 2.36 | 57.6% | 2.16× |
All 300 requests completed without a request, CUDA, or out-of-memory failure. C=4 gives the shortest complete-corpus makespan. C=8 is limited by memory pressure, which makes its complete-corpus result slower than C=4. Sampling is stochastic, so the fixed prompts and seeds do not imply token-identical continuations across concurrency-specific numerical routes; exact decode-token totals are shown above.
Each value is the arithmetic mean ± sample standard deviation over five fixed seeds after server warm-up.
| Prompt tokens | Prefill phase (tok/s) | Server TTFT (ms) | Decode phase (tok/s) |
|---|---|---|---|
| 7,680 | 8,340.4 ± 13.0 | 931.6 ± 1.6 | 71.2 ± 0.1 |
| 64,512 | 5,297.9 ± 259.2 | 12,281.1 ± 561.5 | 65.7 ± 0.8 |
| 130,048 | 3,544.7 ± 25.3 | 36,853.5 ± 259.4 | 59.6 ± 0.9 |
| 260,096 | 2,203.1 ± 13.4 | 118,354.8 ± 717.2 | 52.9 ± 2.3 |
The C=1 point supplies five samples for each fixture. Values are arithmetic mean ± sample standard deviation from server phase timings and speculative counters.
| AIME 2026 fixture | Completion tokens | Decode phase (tok/s) | MTP acceptance | MTP tokens/round |
|---|---|---|---|---|
| Problem 1 | 1,465.4 ± 417.3 | 195.2 ± 4.6 | 76.0% ± 2.4% | 3.28 ± 0.07 |
| Problem 15 | 65,414.4 ± 271.9 | 151.4 ± 2.0 | 56.2% ± 1.1% | 2.69 ± 0.03 |
| Problem 30 | 50,023.4 ± 14,839.1 | 167.5 ± 23.7 | 64.6% ± 14.9% | 2.94 ± 0.45 |
Each category contains three fixtures and five seeds per fixture, for 15 samples.
| Category | Decode phase (tok/s) | MTP acceptance | MTP tokens/round |
|---|---|---|---|
| Code | 194.3 ± 6.1 | 76.4% ± 3.9% | 3.29 ± 0.12 |
| Story | 126.1 ± 10.9 | 37.4% ± 5.8% | 2.12 ± 0.17 |
| Translation | 192.3 ± 11.9 | 75.0% ± 6.5% | 3.25 ± 0.19 |
| Structured output | 219.8 ± 8.6 | 90.8% ± 5.1% | 3.72 ± 0.15 |
See the full methodology and results for metric definitions and the exact reproduction command.
The artifact was evaluated through NInfer's OpenAI-compatible serving route with thinking enabled,
MTP=3, and INT8 group-64 KV. EvalScope 1.9.0 used 0-shot prompts, rule-based scoring, and one sample
per problem with temperature 1.0, top-p 0.95, top-k 20, presence penalty 0.0, and seed 42. The text
suite ran at a 252,928-token context limit; the multimodal suite ran with --vision at a
81,920-token limit.
| Benchmark | NInfer NVFP4 | Correct / total | Official Qwen3.8-27B BF16 |
|---|---|---|---|
| IFBench (prompt-level strict) | 77.00% | 231 / 300 | 79.5 |
| AIME 2025 | 96.67% | 29 / 30 | — |
| AIME 2026 | 96.67% | 29 / 30 | — |
| GPQA-Diamond | 90.40% | 179 / 198 | 89.2 |
| ERQA | 66.25% | 265 / 400 | 65.5 |
| RealWorldQA | 83.53% | 639 / 765 | 85.9 |
All 1,723 configured samples completed and were scored. IFBench additionally reports 80.50% instruction-level strict, 80.33% prompt-level loose, and 83.50% instruction-level loose. These are single-sample results, not pass@k.
The official Qwen3.8-27B BF16 figures come from the upstream model card; its sampling settings and IFBench metric level are not stated there, so the last column is not a same-protocol comparison. The NVFP4 deltas stay within ±2.5 points on the four overlapping benchmarks, and the upstream card reports no AIME results.
385b30ce or later and the matching registered
target.| Field | Value |
|---|---|
| Base repository | Qwen/Qwen3.8-27B |
| Base revision | 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 |
| Base download source | modelscope.cn/models/Qwen/Qwen3.8-27B |
| Quantized source repository | unsloth/Qwen3.8-27B-NVFP4 |
| Quantized source revision | 60e813d4dbbdc5d64cf3f5a8caf2897bedf03679 |
| Conversion recipe | qwen3_8_27b_nvfp4-v2 |
| Embedding encoder | MAXABS_BF16S_RECIP_E4M3FN_RNE_V1 |
| Converter repository | https://github.com/Neroued/ninfer |
| Converter revision | 863aa8a5f1e866db74f29f8999b83b4021398dee |
| Minimum runtime revision | 385b30ce1757bafe5a82680e9b5aeb940b14eec1 |
| Ranking input SHA-256 | c692dc76388132c910547589b4fb4a0503fbd6ad50aaac6a509bbcb192a8afa5 |
The artifact identity, summarized object inventory, and conversion provenance are published in
artifact-manifest.json.
The exact storage contract is maintained in the
Qwen3.8-27B artifact reference.
This NInfer artifact is distributed under the Apache License 2.0. The Qwen3.8-27B base repository and the quantized source repository are also licensed under Apache-2.0. Users remain responsible for complying with the license and applicable laws.