Downloads · 30 days
27
5% of all-time downloads
thecodehaider/Qwen2.5-32B-Instruct-GGUF
Qwen2.5-32B-Instruct-GGUF is a machine learning model from thecodehaider. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for gguf.
This model was quantized to GGUF format using QuantizeLab, the fastest SaaS to quantize models under 32B parameters. Join hundreds of developers downloading our optimized quants!
Downloads · 30 days
27
5% of all-time downloads
All-time downloads
579
Public
Repo size
19.9 GB
Likes
0
Public
Click a slice to open those files.
.gguf19.9 GB · 100%
From the Hugging Face model README
This model was quantized to GGUF format using QuantizeLab, the fastest SaaS to quantize models under 32B parameters. Join hundreds of developers downloading our optimized quants!
Q4_K_M GGUF quantization of Qwen/Qwen2.5-32B-Instruct,
produced with llama.cpp by
Quantizelab.dev.
| File | model-Q4_K_M.gguf |
| Quantization | Q4_K_M |
| Size on disk | 19.85 GB |
| Base model | Qwen/Qwen2.5-32B-Instruct |
GGUF defaults to CPU. To get GPU speed you must offload every layer — a
single layer left on the CPU takes generation from ~25 tok/s to ~3 tok/s.
-ngl 999 simply means "offload all of them".
# llama.cpp
llama-cli -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:model-Q4_K_M.gguf -ngl 999 -c 4096 -p "Hello"
# local file
llama-cli -m model-Q4_K_M.gguf -ngl 999 -c 4096 -cnv
# OpenAI-compatible server
llama-server -m model-Q4_K_M.gguf -ngl 999 -c 4096 --port 8080
# Ollama
ollama run hf.co/thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
# llama-cpp-python
from llama_cpp import Llama
llm = Llama(model_path="model-Q4_K_M.gguf", n_gpu_layers=-1, n_ctx=4096)
print(llm("Hello", max_tokens=128)["choices"][0]["text"])
Weights plus ~1.2 GB of KV-cache and compute overhead at a 4k context.
| GPU | VRAM | Fits fully offloaded? | Headroom for context |
|---|---|---|---|
| NVIDIA T4 / RTX 4060 | 16 GB | No | offload partially (-ngl lower) or use CPU |
| RTX 3090 / 4090 / A10 | 24 GB | Tight | ~2.9 GB (cap context ~2048) |
| A100 40GB | 40 GB | Yes | ~18.9 GB (4k+ context) |
If a row says No, lower -ngl until it fits, or run on CPU (GGUF works
either way — it is just slower).
Q4_K_M is the recommended balance of size and quality; Q8_0 and above
will not fully offload to a 16 GB card for models past ~8B.-c (context) first when you hit out-of-memory: the KV cache grows
linearly with context length.