Downloads · 30 days
21
20% of all-time downloads
Bopalv/Qwen3-0.6B-quantized
Qwen3-0.6B-quantized is a machine learning model from Bopalv. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This repository contains three quantized versions of the Qwen3-0.6B model, optimized for different use cases and hardware requirements.
Downloads · 30 days
21
20% of all-time downloads
All-time downloads
106
Public
Repo size
2.2 GB
Likes
0
Public
Click a slice to open those files.
.safetensors1.3 GB · 59%
From the Hugging Face model README
This repository contains three quantized versions of the Qwen3-0.6B model, optimized for different use cases and hardware requirements.
Qwen3-0.6B-GGUF/Qwen3-0.6B.Q4_K_M.ggufQwen3-0.6B-GPTQ-Int4/Qwen3-0.6B-GPTQ-Int8/| Feature | Value |
|---|---|
| Base Model | Qwen3-0.6B |
| Parameters | 0.6B |
| Architecture | Qwen3ForCausalLM |
| Hidden Size | 1024 |
| Layers | 28 |
| Attention Heads | 16 |
| KV Heads | 8 |
| Max Context | 40,960 tokens |
| Vocab Size | 151,936 |
# Using prima.cpp
./llama-server -m Qwen3-0.6B-GGUF/Qwen3-0.6B.Q4_K_M.gguf --port 8080
# Using ollama
ollama run qwen3:0.6b
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"Bopalv/Qwen3-0.6B-quantized",
subfolder="Qwen3-0.6B-GPTQ-Int4",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(
"Bopalv/Qwen3-0.6B-quantized",
subfolder="Qwen3-0.6B-GPTQ-Int4"
)
| Model | Bits | Group Size | Symmetric | Format | Size |
|---|---|---|---|---|---|
| GGUF Q4_K_M | 4 | N/A | Yes | GGUF | 462 MB |
| GPTQ-Int4 | 4 | 128 | Yes | Safetensors | 517 MB |
| GPTQ-Int8 | 8 | 128 | Yes | Safetensors | 727 MB |
This is a quantized version of Qwen3-0.6B by Qwen Team.
Apache 2.0 (same as base model)