Skip to content

inference-optimization

DeepSeek-R1-Distill-Llama-8B-NVFP4

inference-optimization/DeepSeek-R1-Distill-Llama-8B-NVFP4

DeepSeek-R1-Distill-Llama-8B-NVFP4 is a text generation model from inference-optimization. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.

Headline: NVFP4 shifts gsm8k by -3.94 pp vs BF16 (64.52% → 60.58%) (95% CI ±3.69pp) on a single GB10.

Downloads · 30 days

84

29% of all-time downloads

All-time downloads

289

Public

Parameters

4.5B

6 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors6 GB · 100%

Parameter types

How the weights are stored.

U83.5B · 77%

Try a prompt

Base models

Task
Text Generation
Library
transformers
Type
llama
License
mit
Created
Aug 14, 2026
Updated
Aug 14, 2026
DeepSeek-R1-Distill-Llama-8B-NVFP4 — AI Model — AIMarketly