Downloads · 30 days
1K
57% of all-time downloads
CompressedGemma/Qwen3.8-27B-Q2
Qwen3.8-27B-Q2 is a machine learning model from CompressedGemma. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
THIS DOES INCLUDE MTP LAYER unlike most other quants.
Downloads · 30 days
1K
57% of all-time downloads
All-time downloads
1.8K
Public
Repo size
41.8 GB
Likes
1
Public
Click a slice to open those files.
.gguf19 GB · 100%
From the Hugging Face model README
THIS DOES INCLUDE MTP LAYER unlike most other quants.
RC2.



Use Pi agent, it is the best harness and allows compaction.
Also this DOES include the MTP layer, invoke it with: --spec-type draft-mtp
A selectively quantized Qwen3.8-27B GGUF using a mixed-precision Q2 strategy.
This quantization is designed around the observation that not all tensors contribute equally to model quality. Instead of forcing the entire model into the same low-bit representation, sensitive components are retained at higher precision while less sensitive weights are aggressively compressed.
| Property | Value |
|---|---|
| Base model | Qwen3.8-27B |
| Parameters | ~27B |
| Format | GGUF |
| Quantization | Mixed Q2 |
| Primary goal | Maximum quality per GB |
| Runtime | llama.cpp / compatible GGUF loaders |
Uniform Q2 quantization treats every tensor as though it has the same tolerance for information loss.
It doesn't.
Some weights are substantially more important to preserving:
This release therefore uses selective precision allocation.
The bulk of the model is compressed aggressively, while particularly sensitive tensors are preserved using higher-precision representations.
The result is intended to retain substantially more of the original model's behavior than a naive all-Q2 conversion at a comparable storage budget.
The guiding principle is:
Spend bits where they matter. Save bits where they don't.
Rather than optimizing solely for an average bits-per-weight number, the quantizer considers the structural role of individual tensors.
This makes the resulting GGUF a mixed quantization, not simply a conventional Q2 file with a different name.
Certain tensors are deliberately excluded from aggressive quantization when their information content or numerical sensitivity makes them disproportionately important.
This model is intended for users who want to run a 27B-class Qwen model locally at an extremely constrained memory footprint without accepting the full quality loss normally associated with uniform Q2 quantization.
It is particularly interesting for:
Actual performance depends heavily on:
This release prioritizes quality retention per unit of storage.
The purpose of this quantization is not merely to make the model smaller.
The goal is to preserve the behaviors that tend to disappear first when a large model is pushed toward extremely low bitrates.
Base model: Qwen3.8-27B
Quantization: HPC-Quantize
Format: GGUF
This is an experimental low-bit quantization.
Ultra-low-bit inference necessarily involves information loss relative to the original model. The mixed-precision strategy is intended to reduce that loss by allocating additional precision selectively, but it does not reproduce the full-precision model.
Results may vary substantially depending on workload and inference configuration.
If you find interesting differences between this release and other Q2/Q3 quantizations, please report the workload and inference configuration along with your results.