Downloads · 30 days
965
46% of all-time downloads
AtomicChat/Phi-4-mini-instruct-GGUF
Phi-4-mini-instruct-GGUF is a text generation model from AtomicChat. Use it when you need the model to write or continue text. It is set up for gguf. The card lists the license as mit.
Downloads · 30 days
965
46% of all-time downloads
All-time downloads
2.1K
Public
Repo size
15.2 GB
Likes
2
Public
Click a slice to open those files.
.gguf15.2 GB · 100%
From the Hugging Face model README
Phi 4 Mini, self-quantized to GGUF by Atomic Chat. Built straight from Microsoft's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
[!NOTE] These GGUFs are self-quantized from the original weights, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
[!IMPORTANT] Always pass
--jinjaso the Phi 4 Mini chat template is applied. Without it the model can emit malformed turns.
| Property | Value |
|---|---|
| Base model | microsoft/Phi-4-mini-instruct |
| Parameters | 3.8B |
| Layers | 32 |
| Sliding window | 262144 tokens |
| Context length | 131,072 tokens (128K) |
| Vocabulary | 200,064 |
| Modalities | Text |
| Architecture | Dense decoder, hybrid sliding-window (262144) and global attention, 24 attention heads over 8 KV heads, Phi3ForCausalLM |
| This repo | GGUF quants (imatrix). Quants: Q4_K_M, UD-Q4_K_XL, Q5_K_M, Q6_K, Q8_0 |
| Quant | Size | Notes |
|---|---|---|
Q4_K_M | 2.5 GB | Recommended default. Best balance of size, speed and quality. |
UD-Q4_K_XL | 2.6 GB | Dynamic. Embeddings and output kept at Q8_0 for higher quality at a Q4 footprint. |
Q5_K_M | 2.8 GB | Higher quality, low loss. |
Q6_K | 3.2 GB | Near lossless, noticeably lighter than Q8_0. |
Q8_0 | 4.1 GB | Effectively lossless, reference quality. |
[!TIP] Pick the largest file that fits your (V)RAM with room for context.
Q4_K_MorUD-Q4_K_XLis the sweet spot for most setups;Q6_KorQ8_0for maximum fidelity.
Run Phi 4 Mini locally with:
AtomicChat/Phi-4-mini-instruct-GGUF, pick a quant, hit Use this model.llama-server -hf AtomicChat/Phi-4-mini-instruct-GGUF:Q4_K_M --jinja -c 8192ollama run hf.co/AtomicChat/Phi-4-mini-instruct-GGUF:Q4_K_M| Parameter | Value |
|---|---|
| temperature | 0.0 |
Microsoft's recommended sampling configuration for microsoft/Phi-4-mini-instruct.
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON
cmake --build llama.cpp/build --config Release -j --target llama-cli llama-server
./llama.cpp/build/bin/llama-server \
-hf AtomicChat/Phi-4-mini-instruct-GGUF:Q4_K_M \
--jinja -ngl 99 -c 8192 -fa on
microsoft/Phi-4-mini-instruct (original weights).--imatrix.UD-Q4_K_XL additionally pins the token-embedding and output tensors to Q8_0.Original model by Microsoft, released under the MIT license. Full terms: MIT. Quantized by Atomic Chat.