Downloads · 30 days
33
21% of all-time downloads
sh0ck0r/deepseek-coder-v2-lite-FP8-Dynamic
deepseek-coder-v2-lite-FP8-Dynamic is a text generation model from sh0ck0r. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This is an FP8 quantized version of deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct using llmcompressor with the FP8DYNAMIC scheme.
Downloads · 30 days
33
21% of all-time downloads
All-time downloads
155
Public
Parameters
15.7B
16.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
How the weights are stored.
F8_E4M315.3B · 97%
From the Hugging Face model README
This is an FP8 quantized version of deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct using llmcompressor with the FP8_DYNAMIC scheme.
pip install vllm
# Serve the model
vllm serve REPO_ID \
--max-model-len 32768 \
--gpu-memory-utilization 0.95
# Python API
from vllm import LLM
llm = LLM(model="REPO_ID")
outputs = llm.generate("Hello, how are you?")
print(outputs[0].outputs[0].text)
from transformers import AutoTokenizer, AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
"REPO_ID",
device_map="auto",
torch_dtype="auto"
)
tokenizer = AutoTokenizer.from_pretrained("REPO_ID")
messages = [{'role': 'user', 'content': 'Hello!'}]
inputs = tokenizer.apply_chat_template(messages, return_tensors='pt').to(model.device)
outputs = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0]))
This model was quantized using:
lm_headPerfect for:
Minimum VRAM (approximate):
Recommended:
⚠️ Quantization Trade-offs:
✅ Best Practices:
--kv-cache-dtype fp8 for longer contexts--gpu-memory-utilization 0.90-0.95--enforce-eager if you encounter compilation issuesIf you use this model, please cite:
@misc{model_name-fp8,
author = {author},
title = {model_name FP8 Dynamic Quantization},
year = {2025},
publisher = {HuggingFace},
url = {https://huggingface.co/repo_id}
}
Inherits license from base model: deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct
Want more FP8 models? Check out my other quantizations!