Downloads · 30 days
75
9% of all-time downloads
centml/DeepSeek-R1-NVFP4-v2-mlpinf
DeepSeek-R1-NVFP4-v2-mlpinf is a text generation model from centml. Use it when you need the model to write or continue text. It is set up for modelopt. The card lists the license as mit.
This model is a quantized version of DeepSeek AI's DeepSeek R1 model, produced with the official NVIDIA TensorRT Model Optimizer DeepSeek recipe and calibrated with the MLPerf Inference deepseek-r1 calibration dataset…
Downloads · 30 days
75
9% of all-time downloads
All-time downloads
822
Public
Parameters
394B
623 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors413 GB · 100%
How the weights are stored.
U8332B · 84%
From the Hugging Face model README
This model is a quantized version of DeepSeek AI's DeepSeek R1 model, produced with the official NVIDIA TensorRT Model Optimizer DeepSeek recipe and calibrated with the MLPerf Inference deepseek-r1 calibration dataset. It is quantized with FP4 weights and activations (the same quantization scheme as nvidia/DeepSeek-R1-FP4-v2, including the wo module in attention layers) and is ready for inference with TensorRT-LLM. This model is ready for commercial/non-commercial use.
This model is not owned or developed by NVIDIA or DeepSeek AI. It has been produced by a third party from the open DeepSeek R1 checkpoint for the MLPerf Inference deepseek-r1 benchmark. See the original DeepSeek R1 Model Card.
Architecture Type: Transformers
Network Architecture: DeepSeek R1
Input Type(s): Text
Input Format(s): String
Input Parameters: 1D (One Dimensional): Sequences
Other Properties Related to Input: Context length up to 128K. Per DeepSeek's usage recommendations: temperature 0.5–0.7 (0.6 recommended), no system prompt, and for math problems ask for step-by-step reasoning with the final answer in \boxed{}.
Output Type(s): Text
Output Format: String
Output Parameters: 1D (One Dimensional): Sequences
Supported Runtime Engine(s):
Supported Hardware Microarchitecture Compatibility:
Preferred Operating System(s):
Quantized with NVIDIA TensorRT Model Optimizer, main branch @ 089c06e4 (July 2026), using the examples/deepseek recipe (deepseek_v3/ptq.py for calibration, deepseek_v3/quantize_fp8_to_nvfp4.sh for weight conversion).
Calibration Dataset: mlperf_deepseek_r1_calibration_dataset_500_fp8_eval — the official 500-sample MLPerf Inference deepseek-r1 calibration set. See the MLPerf deepseek-r1 reference README for dataset details and download instructions.
Engine: TensorRT-LLM
Test Hardware: NVIDIA GB200 NVL72
This model was obtained by quantizing the weights and activations of DeepSeek R1 to FP4 data type, ready for inference with TensorRT-LLM. Only the weights and activations of the linear operators within the transformer blocks are quantized (fine-grained FP4 block scaling with group size 16; the KV cache is quantized to FP8; MLA attention projections remain in higher precision). This optimization reduces the number of bits per parameter from 8 to 4, reducing the disk size and GPU memory requirements by approximately 1.6x.
Calibration runs the official recipe unmodified except for the calibration dataloader, which reads the MLPerf deepseek-r1 calibration set (500 samples, max sequence length 2048, batch size 4).
To deploy the quantized checkpoint with TensorRT-LLM LLM API, use the following sample code (requires 8 GPUs of Blackwell generation or later, and TensorRT-LLM built from source with the latest main branch):
from tensorrt_llm import SamplingParams
from tensorrt_llm._torch import LLM
def main():
prompts = [
"Hello, my name is",
"The president of the United States is",
"The capital of France is",
"The future of AI is",
]
sampling_params = SamplingParams(max_tokens=32)
llm = LLM(model="centml/DeepSeek-R1-NVFP4-v2-mlpinf", tensor_parallel_size=8, enable_attention_dp=True)
outputs = llm.generate(prompts, sampling_params)
# Print the outputs.
for output in outputs:
prompt = output.prompt
generated_text = output.outputs[0].text
print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
# The entry point of the program need to be protected for spawning processes.
if __name__ == '__main__':
main()