Downloads · 30 days
36
7% of all-time downloads
jkim96/phi-4-DASHQ-INT2-g32
phi-4-DASHQ-INT2-g32 is a text generation model from jkim96. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
36
7% of all-time downloads
All-time downloads
538
Public
Parameters
1B
7.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors7.2 GB · 100%
How the weights are stored.
BF161B · 38%
From the Hugging Face model README

DASH-Q — Diagonal-Aware Shrinkage for Robust PTQ.
INT2· group size 32 · 7.1679 GB (from 29.3190 GB — 4.1x smaller)
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"jkim96/phi-4-DASHQ-INT2-g32", trust_remote_code=True, device_map="cuda", dtype="auto"
)
tokenizer = AutoTokenizer.from_pretrained("jkim96/phi-4-DASHQ-INT2-g32")
messages = [{"role": "user", "content": "Explain 2-bit quantization in one sentence."}]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
print(tokenizer.decode(model.generate(**inputs, max_new_tokens=256)[0]))
trust_remote_code=True is required: the checkpoint ships its quantized-layer
implementation (modeling_dashq.py) and Triton kernels (dashq_kernel.py).
Without Triton, or on CPU, it falls back to dequantize-and-matmul in PyTorch.
| Package | Minimum | Verified with |
|---|---|---|
torch | 2.4 | 2.12.1+cu130 |
transformers | 5.8 | 5.9.0 |
triton | 3.0 (Linux; bundled with CUDA builds of PyTorch) | 3.7.1 |
huggingface_hub | 1.5 (pulled in by transformers) | 1.15.0 |
| Field | Value |
|---|---|
| Base model | microsoft/phi-4 |
| Precision | INT2, group size 32 |
| Scale / zero dtype | float16 |
| Calibration | wikitext2, 128 samples x 2048 |
| Size | 7.1679 GB · original 29.3190 GB · 4.1x compression |
Full zero-shot / few-shot results for every DASH-Q checkpoint: github.com/JaeminK/dashq#benchmarks
| Metric | Value |
|---|---|
wikitext2_ppl | 7.9293 |
zero-shot accuracy avg | 66.2368 |
arc_challenge | 53.0717 |
arc_easy | 77.0623 |
commonsense_qa | 72.5635 |
hellaswag | 73.3121 |
lambada_openai | 72.1716 |
openbookqa | 41.8000 |
piqa | 79.3254 |
truthfulqa_mc2 | 54.1334 |
winogrande | 72.6914 |