Downloads · 30 days
37
0% of all-time downloads
Ex0bit/Qwen3-VLTO-32B-Instruct-NVFP4
Qwen3-VLTO-32B-Instruct-NVFP4 is a text generation model from Ex0bit. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
This is an NVFP4 quantized version of qingy2024/Qwen3-VLTO-32B-Instruct, optimized for NVIDIA DGX Spark systems with Blackwell GB10 GPUs.
Downloads · 30 days
37
0% of all-time downloads
All-time downloads
8.8K
Public
Parameters
17.2B
20.7 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors20.7 GB · 100%
How the weights are stored.
U815.6B · 82%
From the Hugging Face model README
This is an NVFP4 quantized version of qingy2024/Qwen3-VLTO-32B-Instruct, optimized for NVIDIA DGX Spark systems with Blackwell GB10 GPUs.
| Model Version | Memory Usage | Reduction |
|---|---|---|
| BF16 (Original) | 61.03 GB | Baseline |
| NVFP4 (This model) | 19.42 GB | 68.2% |
| Model Version | Throughput | Relative Performance |
|---|---|---|
| BF16 (Original) | 3.65 tokens/s | Baseline |
| NVFP4 (This model) | 9.99 tokens/s | 2.74x faster |
Test Configuration:
NVFP4 is NVIDIA's 4-bit floating point quantization format featuring:
IMPORTANT: This model must be loaded with vLLM using the modelopt quantization parameter. Standard HuggingFace AutoModelForCausalLM will not work.
from vllm import LLM, SamplingParams
# Load NVFP4 quantized model
llm = LLM(
model="Ex0bit/Qwen3-VLTO-32B-Instruct-NVFP4",
quantization="modelopt", # Required for NVFP4
trust_remote_code=True,
gpu_memory_utilization=0.9
)
# Generate
sampling_params = SamplingParams(temperature=0.7, top_p=0.9, max_tokens=256)
outputs = llm.generate(["Explain quantum computing in simple terms:"], sampling_params)
print(outputs[0].outputs[0].text)
You can optionally set:
HF_CACHE_DIR: Override HuggingFace cache locationThis model is intended for:
See the original model card for base model training details.
Quantization Time: Approximately 60-90 minutes on DGX Spark
All 5 inference tests passed successfully:
Average performance: 9.99 tokens/s on DGX Spark GB10
If you use this quantized model, please cite:
@misc{qwen3vlto32b-nvfp4,
author = {Ex0bit},
title = {Qwen3-VLTO-32B-Instruct-NVFP4: NVFP4 Quantized Model for DGX Spark},
year = {2025},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/Ex0bit/Qwen3-VLTO-32B-Instruct-NVFP4}},
}
And the original base model:
@misc{qingy2024qwen3vlto,
author = {qingy2024},
title = {Qwen3-VLTO-32B-Instruct},
year = {2024},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/qingy2024/Qwen3-VLTO-32B-Instruct}},
}
This quantized model inherits the license from the base model. Please refer to the original model's license for details.