Downloads · 30 days
302
72% of all-time downloads
GatekeeperZA/Qwen2.5-1.5B-Instruct-RKLLM-v1.2.3
Qwen2.5-1.5B-Instruct-RKLLM-v1.2.3 is a text generation model from GatekeeperZA. Use it when you need the model to write or continue text. It is set up for rkllm. The card lists the license as apache-2.0.
Pre-converted Qwen2.5-1.5B-Instruct for the Rockchip RK3588 NPU using rknn-llm runtime v1.2.3.
Downloads · 30 days
302
72% of all-time downloads
All-time downloads
422
Public
Repo size
2.1 GB
Likes
0
Public
Click a slice to open those files.
.rkllm2.1 GB · 100%
From the Hugging Face model README
Pre-converted Qwen2.5-1.5B-Instruct for the Rockchip RK3588 NPU using rknn-llm runtime v1.2.3.
Runs on Orange Pi 5 Plus, Rock 5B, Radxa NX5, and other RK3588-based SBCs with 8GB+ RAM.
| File | Size | Description |
|---|---|---|
Qwen2.5-1.5B-Instruct-rk3588-v123-w8a8.rkllm | 2.0 GB | LLM (W8A8 quantized, 8192 token context) |
~/models/Qwen2.5-1.5B-Instruct/
Qwen2.5-1.5B-Instruct-rk3588-v123-w8a8.rkllm
This model is designed for use with the RKLLM API Server, which provides an OpenAI-compatible API for RK3588 NPU inference. The server auto-discovers .rkllm files by scanning subdirectories of your models folder.
# Place the model in your models directory
mkdir -p ~/models/Qwen2.5-1.5B-Instruct
# Copy .rkllm file here — the API server will find it automatically
sudo systemctl restart rkllm-api
The model will appear as qwen2.5-1.5b-instruct in the OpenAI-compatible model list.
| Parameter | Value |
|---|---|
| Source | Qwen/Qwen2.5-1.5B-Instruct |
| Tool | rkllm-toolkit v1.2.3 |
| Quantization | W8A8 (8-bit weights, 8-bit activations) |
| Optimization level | 1 |
| Target platform | rk3588 |
| NPU cores | 3 |
| Max context | 8192 tokens |
Tested on Orange Pi 5 Plus (16GB RAM), RK3588 SoC, RKNPU driver 0.9.8:
| Metric | Value |
|---|---|
| Decode speed | ~19 tok/s |
| Model load time | ~3 s |
| Peak RAM | ~2.2 GB |
Apache 2.0, inherited from Qwen2.5-1.5B-Instruct.