Downloads · 30 days
8
24% of all-time downloads
zededa/Llama-3.2-3B-Instruct-QCS9075-HTP
Llama-3.2-3B-Instruct-QCS9075-HTP is a machine learning model from zededa. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
This is a pre-compiled version of meta-llama/Llama-3.2-3B-Instruct optimized for the Qualcomm QCS9075 SoC using the Qualcomm Genie SDK.
Downloads · 30 days
8
24% of all-time downloads
All-time downloads
34
Public
Repo size
2.7 GB
Likes
0
Public
Click a slice to open those files.
.bin2.6 GB · 99%
From the Hugging Face model README
This is a pre-compiled version of meta-llama/Llama-3.2-3B-Instruct optimized for the Qualcomm QCS9075 SoC using the Qualcomm Genie SDK.
| Model | Backend | Performance | Size |
|---|---|---|---|
| Llama-3.2-3B-Instruct-QCS9075-HTP | QnnHtp (NPU) | ~18.7 TPS on QCS9075 | 2.5G |
TPS = Tokens Per Second (generation speed)
For HTP models, the LD_LIBRARY_PATH ordering is critical:
export LD_LIBRARY_PATH=/opt/qcom/aistack/qairt/2.42.0.250923/lib/aarch64-linux-gnu:/opt/qcom/aistack/genie/qnn/libs:$LD_LIBRARY_PATH
Create a genie_config.json file:
{
"model_path": "/path/to/model/files",
"backend": "QnnHtp",
"device": "0"
}
# Using the Genie server
python3 /opt/qcom/aistack/genie/examples/server_persistent.py \
--config genie_config.json \
--port 8000
For deploying on Kubernetes clusters with QCS9075 nodes, refer to the deployment pattern:
apiVersion: v1
kind: Pod
metadata:
name: genie-llm-server
spec:
containers:
- name: genie
image: your-registry/genie-runtime:latest
env:
- name: LD_LIBRARY_PATH
value: "'/opt/qcom/aistack/qairt/2.42.0.250923/lib/aarch64-linux-gnu:/opt/qcom/aistack/genie/qnn/libs'"
volumeMounts:
- name: model-storage
mountPath: /models
- name: qcom-libs
mountPath: /opt/qcom/aistack
volumes:
- name: model-storage
hostPath:
path: /mnt/models/llama-3.2-3b-instruct-qcs9075-htp
- name: qcom-libs
hostPath:
path: /opt/qcom/aistack
This repository contains:
This model follows the license of the base model meta-llama/Llama-3.2-3B-Instruct. Please refer to the original model card for license details.
For issues related to:
This model is optimized for edge deployment on Qualcomm QCS9075 devices and may not work on other hardware platforms.