Downloads · 30 days
27
25% of all-time downloads
LLMJapan/Qwen2.5-Coder-32B-Instruct_exl3
Qwen2.5-Coder-32B-Instruct_exl3 is a text generation model from LLMJapan. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
These models are exl3 quantization models of Qwen2.5-Coder-32B which is still SOTA no-reasoning coder model as of today. This model is still my go-to FIM(fill in the middle) autocompletion model after Qwen3, Gemma3 re…
Downloads · 30 days
27
25% of all-time downloads
All-time downloads
108
Public
Repo size
76.9 GB
Likes
1
Trending 1
Click a slice to open those files.
Other1.5 KB · 56%
From the Hugging Face model README
These models are exl3 quantization models of Qwen2.5-Coder-32B which is still SOTA no-reasoning coder model as of today. This model is still my go-to FIM(fill in the middle) autocompletion model after Qwen3, Gemma3 release. I used exllamav3 version 0.0.2.
For coding, I found >=6.0bpw or preferably 8.0bpw model with KV Cache Quantization (>=Q6) is much better than 4.0bpw. If you are using these models only for short Auto Completion, 4.0bpw is usable.
Thanks to excellent work of exllamav3 dev teams.