Downloads · 30 days
0
0% of all-time downloads
Helllbos/Qwen_Quantised3.50.5b
Qwen_Quantised3.50.5b is a text generation model from Helllbos. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
--- language: - en license: apache-2.0 libraryname: transformers tags: - quantization - optimum-quanto - qwen - int8 - cpu pipelinetag: text-generation basemodel: Qwen/Qwen2.5-0.5B-Instruct ---
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
11
Public
Repo size
779 MB
Likes
1
Public
Click a slice to open those files.
.safetensors767 MB · 99%
From the Hugging Face model README
language:
This technique is based on Post-Training Quantization (PTQ) using Qwen2.5-0.5B-Instruct...
language:
This Technique based on the PQT Post Quantization Training A model quantisation using Qwen2.5-0.5B-Instruct. The project shows how to quantise a model with Optimum Quanto and run it locally on a CPU. where a 32 bit Modal into 8 bit
| File | Description |
|---|---|
01_model_quantisation_guide.ipynb | Notebook explaining quantisation concepts and examples |
python quant_qwen.py | Downloads and quantises the model, then saves it locally |
run_quantized.py | Loads the quantised model and generates responses |
requirements.txt | Project dependencies |
Install dependencies:
pip install -r requirements.txt
HuggingFace Token Need
$env:HF_TOKEN = "Your TOKEN"; python quant_qwen.py
python quant_qwen.py
This creates:
qwen-int8/
python .\v1_quant_qwen.py
Please enter your question: hi
INPUT : hi OUTPUT: Hello! How can I assist you today? Please let me know if there's anything specific you'd like to talk about or any questions you have. I'm here to help answer your queries.
jupyter notebook 01_model_quantisation_guide.ipynb
Quantisation reduces the precision of model weights (for example, FP32 → INT8).
Benefits: