Downloads · 30 days
138K
29% of all-time downloads
nvidia/Kimi-K2-Thinking-NVFP4
Kimi-K2-Thinking-NVFP4 is a text generation model from nvidia. Use it when you need the model to write or continue text. It is set up for Model Optimizer. The card lists the license as other.
The NVIDIA Kimi-K2-Thinking-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2-Thinking model, which is an auto-regressive language model that uses an optimized transformer architecture. For more inform…
Downloads · 30 days
138K
29% of all-time downloads
All-time downloads
479K
Public
Parameters
519B
594 GB on disk
Likes
34
Public
Click a slice to open those files.
.safetensors594 GB · 100%
How the weights are stored.
U8507B · 98%
From the Hugging Face model README
The NVIDIA Kimi-K2-Thinking-NVFP4 model is the quantized version of the Moonshot AI's Kimi-K2-Thinking model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Kimi-K2-Thinking NVFP4 model is quantized with TensorRT Model Optimizer.
This model is ready for commercial/non-commercial use. <br>
This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Kimi-K2-Thinking) Model Card.
Governing Terms: Use of this model is governed by the NVIDIA Open Model License. ADDITIONAL INFORMATION: Modified MIT License.
Global
This model is intended for developers and researchers building LLMs
Huggingface 12/04/2025 via https://huggingface.co/nvidia/Kimi-K2-Thinking-NVFP4
Architecture Type: Transformers <br> Network Architecture: DeepSeek V3 <br> Number of Model Parameters: 1T
Input Type(s): Text <br> Input Format(s): String <br> Input Parameters: 1D (One Dimensional): Sequences <br>
Output Type(s): Text <br> Output Format: String <br> Output Parameters: 1D (One Dimensional): Sequences <br>
Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
Supported Runtime Engine(s): <br>
Supported Hardware Microarchitecture Compatibility: <br>
Preferred Operating System(s): <br>
** The model is quantized with nvidia-modelopt v0.39.0 <br>
** Data Collection Method by dataset: Hybrid: Human, Automated <br> ** Labeling Method by dataset: Hybrid: Human, Automated <br> ** Data Modality: [Text] <br> ** Text Training Data Size: undisclosed. <br>
** Data Collection Method by dataset: Hybrid: Human, Automated <br> ** Labeling Method by dataset: Hybrid: Human, Automated <br>
** Data Collection Method by dataset: Hybrid: Human, Automated <br> ** Labeling Method by dataset: Hybrid: Human, Automated <br>
Engine: vLLM/SGLang/TensorRT-LLM <br> Test Hardware: B200 <br>
This model was obtained by converting and quantizing the weights and activations of Kimi-K2-Thinking from INT4 to BF16 to NVFP4 data type, ready for inference with vLLM. Only the weights and activations of the linear operators within transformer blocks in MoE are quantized.
To serve this checkpoint with vLLM, you can start the docker vllm/vllm-openai:v0.11.2 and run the sample command below:
python3 -m vllm.entrypoints.openai.api_server --model nvidia/Kimi-K2-Thinking-NVFP4 --trust-remote-code --tensor-parallel-size 4
To serve this checkpoint with SGLang, you can start the docker lmsysorg/sglang:latest and run the sample command below:
python3 -m sglang.launch_server --model-path nvidia/Kimi-K2-Thinking-NVFP4 --tp 4 --quantization modelopt_fp4 --trust-remote-code
To serve this checkpoint with TensorRT-LLM, please follow instructions in deployment-guide-for-kimi-k2-thinking-on-trtllm
The base model was trained on data that contains toxic language and societal biases originally crawled from the internet. Therefore, the model may amplify those biases and return toxic responses especially when prompted with toxic prompts. The model may generate answers that may be inaccurate, omit key information, or include irrelevant or redundant text producing socially unacceptable or undesirable text, even if the prompt itself does not include anything explicitly offensive.
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report security vulnerabilities or NVIDIA AI Concerns here.