Downloads · 30 days
239
0% of all-time downloads
curiousmind147/microsoft-phi-4-AWQ-4bit-GEMM
microsoft-phi-4-AWQ-4bit-GEMM is a text generation model from curiousmind147. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This is a 4-bit AutoAWQ quantized version of Microsoft's Phi-4. It is optimized for fast inference using vLLM with minimal loss in accuracy.
Downloads · 30 days
239
0% of all-time downloads
All-time downloads
49.2K
Public
Parameters
14.7B
9.1 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors9.1 GB · 100%
How the weights are stored.
I3213.6B · 93%
From the Hugging Face model README
This is a 4-bit AutoAWQ quantized version of Microsoft's Phi-4.
It is optimized for fast inference using vLLM with minimal loss in accuracy.
You can load this model directly in vLLM for efficient inference:
vllm serve "curiousmind147/microsoft-phi-4-AWQ-4bit-GEMM"
Then, test it using cURL:
curl -X POST "http://localhost:8000/generate" \
-H "Content-Type: application/json" \
-d '{"prompt": "Explain quantum mechanics in simple terms.", "max_tokens": 100}'
transformers + AWQ)To use this model with Hugging Face Transformers:
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
model_path = "curiousmind147/microsoft-phi-4-AWQ-4bit-GEMM"
model = AutoAWQForCausalLM.from_pretrained(model_path)
tokenizer = AutoTokenizer.from_pretrained(model_path)
inputs = tokenizer("What is the meaning of life?", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(output[0], skip_special_tokens=True))
This model was quantized using AutoAWQ with the following parameters:
zero_point=True)q_group_size=128)GEMM| Model Size | FP16 (No Quant) | AWQ 4-bit Quantized |
|---|---|---|
| Phi-4 14B | ❌ Requires >20GB VRAM | ✅ 8GB-12GB VRAM |
AWQ significantly reduces VRAM requirements, making it possible to run 14B models on consumer GPUs. 🚀
Special thanks to:
If you find this useful, give it a ⭐ on Hugging Face! 🎯