Downloads · 30 days
11
15% of all-time downloads
ApyHTML19/qwen-onnx-int8
qwen-onnx-int8 is a machine learning model from ApyHTML19. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
This is a quantized INT8 version of the Qwen language model, converted to ONNX format for efficient inference on low-resource devices such as mobile phones and edge hardware.
Downloads · 30 days
11
15% of all-time downloads
All-time downloads
73
Public
Repo size
8.9 GB
Likes
0
Public
Click a slice to open those files.
.onnx_data7.1 GB · 80%
From the Hugging Face model README
This is a quantized INT8 version of the Qwen language model, converted to ONNX format for efficient inference on low-resource devices such as mobile phones and edge hardware.
The goal of this project is to make Qwen accessible on devices with limited memory and compute power (e.g. iPhone 12, mid-range Android phones) without requiring high-end GPUs or cloud infrastructure.
onnxruntime)pip install onnxruntime numpy
import onnxruntime as ort
session = ort.InferenceSession( "model_int8.onnx", providers=["CPUExecutionProvider"] )
| File | Description |
|---|---|
model.onnx | Original FP32 model |
model_int8.onnx | INT8 quantized model ✅ |
tokenizer.json | Tokenizer config |
vocab.json | Vocabulary |
quantize_int8.py | Quantization script |
