Downloads · 30 days
167
27% of all-time downloads
Pilcothink/Qwen3.5-9B-MixedInt4-AutoRound
Qwen3.5-9B-MixedInt4-AutoRound is a image-text-to-text model from Pilcothink. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
This repository contains a Mixed-INT4 quantized version of Qwen/Qwen3.5-9B, produced using Intel AutoRound.
Downloads · 30 days
167
27% of all-time downloads
All-time downloads
614
Public
Parameters
3.6B
9.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.3 GB · 100%
How the weights are stored.
BF162.5B · 68%
From the Hugging Face model README
This repository contains a Mixed-INT4 quantized version of Qwen/Qwen3.5-9B, produced using Intel AutoRound.
The model weights were quantized to reduce memory requirements while preserving the original Qwen3.5 architecture, tokenizer, configuration, and multimodal capabilities.
This repository is an unofficial community quantization of Qwen3.5-9B.
The model architecture and original model behavior are provided by the Qwen team. Quantization may cause small differences in output quality, numerical precision, generation consistency, and benchmark performance compared with the original model.
No independent benchmark results are currently provided for this quantized version.
Evaluation was performed using AutoRound’s evaluation CLI, powered by LM Evaluation Harness.
| Benchmark | Metric | Qwen3.5-9B | Qwen3.5-9B-MixedInt4-AutoRound | Difference | Recovery Rate |
|---|---|---|---|---|---|
| MMLU | acc | 78.66% | 77.62% | -1.04%p | 98.68% |
| ARC-Challenge | acc_norm | 55.80% | 55.03% | -0.77%p | 98.62% |
| BoolQ | acc | 89.17% | 86.91% | -2.26%p | 97.47% |
| HellaSwag | acc_norm | 78.15% | 77.45% | -0.70%p | 99.10% |
| PIQA | acc_norm | 80.03% | 80.20% | +0.17%p | 100.21% |
| WinoGrande | acc | 73.01% | 71.82% | -1.19%p | 98.37% |
| Average | — | 75.80% | 74.84% | -0.97%p | 98.73% |
| MMLU Category | Qwen3.5-9B | Qwen3.5-9B-MixedInt4-AutoRound | Difference | Recovery Rate |
|---|---|---|---|---|
| Humanities | 70.48% | 68.93% | -1.55%p | 97.80% |
| Other | 83.20% | 82.52% | -0.68%p | 99.18% |
| Social Sciences | 86.90% | 86.55% | -0.35%p | 99.60% |
| STEM | 78.34% | 77.07% | -1.27%p | 98.38% |
If the installed vLLM version supports this model architecture and AutoRound quantization format, the model can be served using:
vllm serve YOUR_USERNAME/Qwen3.5-9B-MixedInt4-AutoRound \
--trust-remote-code
Support for newly released model architectures and quantization formats may require a recent development build of vLLM.
Exact quantization settings, calibration dataset, group size, and AutoRound version should be documented here when available.
This model inherits the limitations of the original Qwen3.5-9B model.
Additional limitations may result from quantization:
Users should evaluate the model on their own workloads before production use.
For complete information about the architecture, supported languages, context length, multimodal usage, benchmarks, intended uses, and limitations, refer to the original model card:
The original Qwen3.5-9B model is distributed under the Apache License 2.0.
This quantized repository follows the license and usage requirements of the original model. Users are responsible for reviewing and complying with the original license terms.