Downloads · 30 days
617
13% of all-time downloads
amd/DeepSeek-V4-Pro-MXFP4
DeepSeek-V4-Pro-MXFP4 is a text generation model from amd. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
- Model Architecture: DeepseekV4ForCausalLM - Input: Text - Output: Text - Supported Hardware Microarchitecture: AMD MI355 / MI350 (gfx950) - ROCm: 7.2.0 - PyTorch: 2.9.1 - Transformers: 5.13.1 - Operating System(s):…
Downloads · 30 days
617
13% of all-time downloads
All-time downloads
4.6K
Public
Parameters
810B
863 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors863 GB · 100%
How the weights are stored.
U8788B · 97%
From the Hugging Face model README
Quantized from deepseek-ai/DeepSeek-V4-Pro with AMD Quark. The pipeline re-quantizes only the MoE expert weights and activations to MXFP4. All non-expert modules are kept as-is via the exclude list.
from quark.torch import ModelQuantizer
from quark.torch.quantization.config.template import LLMTemplate
template = LLMTemplate.get('deepseek_v4')
qconfig = template.get_config(scheme='mxfp4')
ModelQuantizer(qconfig).direct_quantize_checkpoint(
pretrained_model_path='<DSV4_Pro_src_path>',
save_path='<output_dir>',
keep_excluded_layers_as_original_model_state=True,
)
This model can be deployed efficiently using the vLLM backend based on the Docker image vllm/vllm-openai-rocm:v0.29.0. vLLM and lm_eval are both installed from source.
The model was evaluated on gsm8k (8-shot) benchmark using the vLLM framework.
| Benchmark | deepseek-ai/DeepSeek-V4-Pro | amd/DeepSeek-V4-Pro-MXFP4 | Recovery |
|---|---|---|---|
| GSM8K (strict-match) | 94.90 | 93.6 | 99.0% |
The GSM8K results were obtained using the lm-eval framework, based on the Docker image vllm/vllm-openai-rocm:v0.29.0.
export VLLM_ROCM_USE_AITER=1
export VLLM_ROCM_USE_AITER_FUSION_SHARED_EXPERTS=1
vllm serve amd/DeepSeek-V4-Pro-MXFP4 --tensor-parallel-size 4 --kv-cache-dtype fp8 \
--trust-remote-code --tokenizer-mode deepseek_v4 --reasoning-parser deepseek_v4 \
--tool-call-parser deepseek_v4 --enable-auto-tool-choice \
--compilation-config '{"mode": 3, "cudagraph_mode": "FULL_DECODE_ONLY"}'
lm_eval --model local-completions \
--model_args model=amd/DeepSeek-V4-Pro-MXFP4,base_url=http://localhost:30000/v1/completions,tokenized_requests=False,num_concurrent=32 \
--tasks gsm8k --batch_size auto --num_fewshot 8
This model is a quantized derivative of deepseek-ai/DeepSeek-V4-Pro and is distributed under the same license as the source model: the MIT License. A copy of the upstream LICENSE is included in this repository.
Modifications Copyright (c) 2026 Advanced Micro Devices, Inc. All rights reserved. AMD has modified the model weights of the MoE expert layers by quantizing them to MXFP4 with AMD Quark; the modifications are provided under the same MIT License and are not subject to any separate or different license.