Downloads · 30 days
0
InsecureErasure/Qwen3-4B-Instruct-NVFP4
Qwen3-4B-Instruct-NVFP4 is a image-text-to-text model from InsecureErasure. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository provides an NVFP4 (FP4 E2M1) mixed-precision quantized build of Qwen3-4B-Instruct-2507 in ComfyUI comfyquant format, primarily intended for use as a text encoder (e.g. for Z-Image).
Downloads · 30 days
0
Access
Public
Updated Jul 17, 2026
Repo size
3.2 GB
Likes
0
Public
Click a slice to open those files.
.safetensors3.2 GB · 100%
From the Hugging Face model README
This repository provides an NVFP4 (FP4 E2M1) mixed-precision quantized build of
Qwen3-4B-Instruct-2507
in ComfyUI comfy_quant format,
primarily intended for use as a text encoder (e.g. for Z-Image).
Many thanks to SergiusFlavius, who helped me generate this quantization.
convert_to_quant v1.2.6 (comfy_kitchen CUDA NVFP4 kernels)comfy_quant mixed precision
float8_e4m3fn, tensorwise, learned rounding) — text blocks 1 & 34bfloat16) — embed_tokens, model.norm, all norms/biases,
pos_embed, patch_embed, text blocks 0 & 35, and the entire vision encoder
(blocks, merger, deepstack_merger_list)input_scale: ComfyUI quantizes activations dynamically at runtime
(per-tensor amax via the NVFP4 layout), so a baked-in activation scale is not
required. This matches the stock Qwen3-VL NVFP4 baseline.| Tier | Layers | Count |
|---|---|---|
| FP16 (bf16) | embed_tokens, model.norm, all norms/biases, pos_embed, patch_embed, text blocks 0 & 35, entire vision encoder | kept lossless |
| FP8 (e4m3fn, tensorwise) | text blocks 1 & 34 | 14 weights |
| NVFP4 (E2M1, block=16) | text blocks 2–33 projections | 224 weights |
How the layers execute depends on the GPU (ComfyUI pick_operations gates each
format on the device capability):
| GPU (SM) | NVFP4 layers | FP8 layers |
|---|---|---|
| RTX 5090 / Blackwell (12.0) | native FP4 tensor cores (fast) | native FP8 |
| RTX 4090 / Ada (8.9) | dequantized to bf16 (size only, no speedup) | native FP8 |
On Blackwell, NVFP4 weights stay packed and run through comfy_kitchen's
TensorCoreNVFP4Layout GEMM (FP4 weight × dynamically-quantized FP8 activation).
On non-Blackwell GPUs the format is emulated via dequantization, so you get the
smaller footprint but not the FP4 speedup.
Place qwen3_vl_4b_instruct_nvfp4.safetensors in ComfyUI/models/text_encoders/.
CLIPLoader node, type lumina2.ComfyUI auto-detects the quantization metadata (Found quantization metadata version 1) and selects MixedPrecisionOps for the text encoder.
Apache-2.0 (inherited from the base model).