Downloads · 30 days
138
63% of all-time downloads
GatekeeperZA/Qwen3-VL-4B-Instruct-RKLLM-v1.2.3
Qwen3-VL-4B-Instruct-RKLLM-v1.2.3 is a image-text-to-text model from GatekeeperZA. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for rkllm. The card lists the license as apache-2.0.
RKLLM/RKNN conversion of Qwen/Qwen3-VL-4B-Instruct for Rockchip RK3588 NPU inference.
Downloads · 30 days
138
63% of all-time downloads
All-time downloads
219
Public
Repo size
5.7 GB
Likes
1
Public
Click a slice to open those files.
.rkllm4.8 GB · 85%
From the Hugging Face model README
RKLLM/RKNN conversion of Qwen/Qwen3-VL-4B-Instruct for Rockchip RK3588 NPU inference.
Converted with RKLLM Toolkit v1.2.3 (language model) and RKNN Toolkit (vision encoder). This is a multimodal vision-language model — it accepts both images and text as input.
| Property | Value |
|---|---|
| Base Model | Qwen/Qwen3-VL-4B-Instruct |
| Toolkit Version | RKLLM Toolkit v1.2.3 / RKNN Toolkit |
| Runtime Version | RKLLM Runtime ≥ v1.2.1 + RKNN Runtime |
| Quantization | w8a8 (8-bit weights, 8-bit activations) |
| Target Platform | RK3588 |
| NPU Cores | 3 |
| Thinking Mode | ❌ Disabled |
| Model Type | Vision-Language (VLM) |
| Languages | English, Chinese (multilingual) |
Qwen3-VL-4B-Instruct is Alibaba's 4B vision-language model. It handles image understanding, visual QA, document analysis, and chart reading with strong multilingual support. Running on the RK3588 NPU enables fully local, GPU-free multimodal inference.
Compared to the smaller Qwen3-VL-2B, the 4B variant offers meaningfully better image understanding and text extraction.
mkdir -p ~/models/qwen3-vl-4b
cd ~/models/qwen3-vl-4b
git lfs install && git clone https://huggingface.co/GatekeeperZA/Qwen3-VL-4B-Instruct-RKLLM-v1.2.3 .
Use with GatekeeperZA/RKLLM-API-Server — the server loads both the .rkllm and .rknn files automatically when placed in the same directory.
| File | Description |
|---|---|
qwen3-vl-4b-instruct_w8a8_rk3588.rkllm | Language model weights for RK3588 NPU |
qwen3-vl-4b-vision_rk3588.rknn | Vision encoder for RK3588 NPU |