Downloads · 30 days
57
29% of all-time downloads
Qengineering/InternVL3.5-1B-rk3588
InternVL3.5-1B-rk3588 is a image-text-to-text model from Qengineering. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for rkllm.
This repository provides a hardware-accelerated port of InternVL3.5-1B optimized for Rockchip RK3588 NPU.
Downloads · 30 days
57
29% of all-time downloads
All-time downloads
197
Public
Repo size
1.6 GB
Likes
0
Public
Click a slice to open those files.
.rkllm936 MB · 59%
From the Hugging Face model README
This repository provides a hardware-accelerated port of InternVL3.5-1B optimized for Rockchip RK3588 NPU.
<br>User:<image>Describe the image.<br><br>
Answer: The image depicts an astronaut on the moon, holding a green bottle of beer and sitting next to a green cooler with some writing on it. The background shows Earth from space, highlighting the contrast between the moon's barren surface and the planet below.
| Component | File | Precision |
|---|---|---|
| LLM | internvl3_5-1b-instruct_w8a8_rk3588.rkllm | W8A8 |
| Vision Encoder | internvl3_5-1b_vision_rk3588.rknn | FP16 |
All models, with C++ examples, can be found on the Q-engineering GitHub.<br><br> All LLM models are quantized to w8a8, while the VLM vision encoders use fp16.<br>
| model | RAM (GB)<sup>1</sup> | llm cold sec<sup>2</sup> | llm warm sec<sup>3</sup> | vlm cold sec<sup>2</sup> | vlm warm sec<sup>3</sup> | Resolution | Tokens/s |
|---|---|---|---|---|---|---|---|
| Qwen3-2B | 3.1 | 21.9 | 2.6 | 10.0 | 0.9 | 448 x 448 | 11.5 |
| Qwen3-4B | 8.7 | 49.6 | 5.6 | 10.6 | 1.1 | 448 x 448 | 5.7 |
| InternVL3.5-1B | 1.9 | 8.3 | 8.0 | 1.5 | 0.8 | 448 x 448 | 24 |
| InternVL3.5-2B | 3.0 | 22 | 8.0 | 2.7 | 0.8 | 448 x 448 | 11.2 |
| InternVL3.5-4B | 5.4 | 50 | 8.0 | 5.9 | 0.8 | 448 x 448 | 5 |
| InternVL3.5-8B | 8.8 | 92 | 8.0 | 50.5 | 5.8 | 448 x 448 | 3.5 |
| Qwen2.5-3B | 4.8 | 48.3 | 4.0 | 17.9 | 1.8 | 392 x 392 | 7.0 |
| Qwen2-7B | 8.7 | 86.6 | 34.5 | 37.1 | 20.7 | 392 x 392 | 3.7 |
| Qwen2-2.2B | 3.3 | 29.1 | 2.5 | 17.1 | 1.7 | 392 x 392 | 12.5 |
| InternVL3-1B | 1.3 | 6.8 | 1.1 | 7.8 | 0.75 | 448 x 448 | 30 |
| SmolVLM2-2.2B | 3.4 | 21.2 | 2.6 | 10.5 | 0.9 | 384 x 384 | 11 |
| SmolVLM2-500M | 0.8 | 4.8 | 0.7 | 2.5 | 0.25 | 384 x 384 | 31 |
| SmolVLM2-256M | 0.5 | 1.1 | 0.4 | 2.5 | 0.25 | 384 x 384 | 54 |
<sup>1</sup> The total used memory; LLM plus the VLM. <br> <sup>2</sup> When an llm/vlm model is loaded for the first time from your disk to RAM or NPU, it is called a cold start.<br> The duration depends on your OS, I/O transfer rate, and memory mapping.<br> <sup>3</sup> Subsequent loading (warm start) takes advantage of the already mapped data in RAM. Mostly, only a few pointers need to be restored.<br><br> <img width="600" height="450" alt="Plot_1" src="https://github.com/user-attachments/assets/2dde8d27-c8ae-474c-b845-4ed52bdc0785" /><br> <img width="600" height="450" alt="Plot_2" src="https://github.com/user-attachments/assets/0cf946d5-5458-4166-bc2b-fa1592ae4d6b" />