Downloads · 30 days
0
GatekeeperZA/Qwen3-Embedding-0.6B-RKLLM-v1.2.3
Qwen3-Embedding-0.6B-RKLLM-v1.2.3 is a sentence similarity model from GatekeeperZA. Use it when you need a score for how close two texts are. It is set up for sentence-transformers. The card lists the license as apache-2.0.
RKLLM conversion of Qwen/Qwen3-Embedding-0.6B for Rockchip RK3588 NPU inference.
Downloads · 30 days
0
Access
Public
Updated Aug 18, 2026
Repo size
952 MB
Likes
0
Public
Click a slice to open those files.
.rkllm952 MB · 100%
From the Hugging Face model README
RKLLM conversion of Qwen/Qwen3-Embedding-0.6B for Rockchip RK3588 NPU inference.
Converted with RKLLM Toolkit v1.2.3. This model generates dense vector embeddings for semantic search, RAG pipelines, and similarity tasks — all running on the NPU without a GPU.
| Property | Value |
|---|---|
| Base Model | Qwen/Qwen3-Embedding-0.6B |
| Toolkit Version | RKLLM Toolkit v1.2.3 |
| Runtime Version | RKLLM Runtime ≥ v1.2.1 (v1.2.3 recommended) |
| Quantization | w8a8 (8-bit weights, 8-bit activations) |
| Target Platform | RK3588 |
| NPU Cores | 3 |
| Optimization Level | 1 |
| Hybrid Ratio | 0.5 |
| Model Type | Embedding |
| Languages | English, Chinese (multilingual) |
Running a dedicated embedding model on the RK3588 NPU allows the main LLM to use the full NPU without context-switching. Qwen3-Embedding-0.6B achieves strong retrieval performance at a compact size, making it ideal for local RAG pipelines on edge hardware.
Pair with the Qwen3-Reranker-0.6B for a complete retrieval stack.
mkdir -p ~/models/Qwen3-Embedding-0.6B
cd ~/models/Qwen3-Embedding-0.6B
git lfs install && git clone https://huggingface.co/GatekeeperZA/Qwen3-Embedding-0.6B-RKLLM-v1.2.3 .
Use with GatekeeperZA/RKLLM-API-Server embedding endpoint.
| File | Description |
|---|---|
Qwen3-Embedding-0.6B-rk3588-w8a8-opt-1-hybrid-ratio-0.5.rkllm | Quantized embedding model for RK3588 NPU |