Downloads · 30 days
0
qualcomm/Llama-v3-8B-Instruct
Llama-v3-8B-Instruct is a text generation model from qualcomm. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as llama3.
Downloads · 30 days
0
Access
Public
Updated Oct 7, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.md7.1 KB · 81%
From the Hugging Face model README

Llama 3 is a family of LLMs. The model is quantized to w4a16 (4-bit weights and 16-bit activations) and part of the model is quantized to w8a16 (8-bit weights and 16-bit activations) making it suitable for on-device deployment. For Prompt and output length specified below, the time to first token is Llama-PromptProcessor-Quantized's latency and average time per addition token is Llama-TokenGenerator-Quantized's latency.
This is based on the implementation of Llama-v3-8B-Instruct found here. This repository contains pre-exported model files optimized for Qualcomm® devices. You can use the Qualcomm® AI Hub Models library to export with custom configurations. More details on model performance across various devices, can be found here.
Qualcomm AI Hub Models uses Qualcomm AI Hub Workbench to compile, profile, and evaluate this model. Sign up to run these models on a hosted Qualcomm® device.
Follow the GenieX quickstart to install GenieX and deploy the model on a target device.
You'll need to export the model artifact using the steps below, then follow Run a Local Model with GenieX.
See the LLM-on-Genie tutorial to run with the Genie runtime. Note: Genie support will be deprecated soon.
Due to licensing restrictions, we cannot distribute pre-exported model assets for this model. Use the Qualcomm® AI Hub Models Python library to compile and export the model with your own:
See our repository for Llama-v3-8B-Instruct on GitHub for usage instructions.
Model Type: Model_use_case.text_generation
Model Stats:
| Model | Runtime | Precision | Chipset | Context Length | Response Rate (tokens per second) | Time To First Token (range, seconds) |
|---|---|---|---|---|---|---|
| Llama-v3-8B-Instruct | GENIE | w4a16 | Snapdragon® 8 Elite Gen 5 For Galaxy Mobile | 4096 | 15.081136703491211 | 0.102728 - 3.287296 |
| Llama-v3-8B-Instruct | GENIE | w4a16 | Snapdragon® 8 Elite For Galaxy Mobile | 4096 | 13.722314834594727 | 0.129698 - 4.150336 |
| Llama-v3-8B-Instruct | GENIE | w4a16 | Snapdragon® X2 Elite | 4096 | 22.19607925415039 | 0.13744900000000002 - 4.3983680000000005 |
| Llama-v3-8B-Instruct | GENIE | w4a16 | Snapdragon® X Elite | 4096 | 4.413198947906494 | 0.207841 - 6.650912 |
| Llama-v3-8B-Instruct | GENIE | w4a16 | Qualcomm® Dragonwing™ IQ-8275 | 4096 | 9.479302406311035 | 0.27746699999999996 - 8.878943999999999 |
| Llama-v3-8B-Instruct | GENIE | w4a16 | Qualcomm® Dragonwing™ IQ-9075 | 4096 | 8.127570152282715 | 0.19664099999999998 - 6.292511999999999 |
| Llama-v3-8B-Instruct | GENIE | w4a16 | Qualcomm® Dragonwing™ IQ-X7181 | 4096 | 4.413198947906494 | 0.207841 - 6.650912 |
| Llama-v3-8B-Instruct | GENIE | w4a16 | Qualcomm® Dragonwing™ Q-8750 | 4096 | 13.722314834594727 | 0.129698 - 4.150336 |
| Llama-v3-8B-Instruct | GENIEX_QAIRT | w4a16 | Snapdragon® 8 Elite Gen 5 For Galaxy Mobile | 4096 | 9.653202 | 0.176583 - 5.650656 |
| Llama-v3-8B-Instruct | GENIEX_QAIRT | w4a16 | Snapdragon® 8 Elite For Galaxy Mobile | 4096 | 9.940364 | 0.18946806451612902 - 6.062978064516129 |
| Llama-v3-8B-Instruct | GENIEX_QAIRT | w4a16 | Snapdragon® X2 Elite | 4096 | 21.24863 | 0.11502590322580644 - 3.680828903225806 |
| Llama-v3-8B-Instruct | GENIEX_QAIRT | w4a16 | Snapdragon® X Elite | 4096 | 11.330596 | 0.22880099999999998 - 7.321631999999999 |
| Llama-v3-8B-Instruct | GENIEX_QAIRT | w4a16 | Qualcomm® Dragonwing™ IQ-8275 | 4096 | 9.875264 | 0.220166 - 7.045312 |
| Llama-v3-8B-Instruct | GENIEX_QAIRT | w4a16 | Qualcomm® Dragonwing™ IQ-9075 | 4096 | 9.6774 | 0.20317929032258064 - 6.5017372903225805 |
| Llama-v3-8B-Instruct | GENIEX_QAIRT | w4a16 | Qualcomm® Dragonwing™ IQ-X7181 | 4096 | 11.330596 | 0.22880099999999998 - 7.321631999999999 |
| Llama-v3-8B-Instruct | GENIEX_QAIRT | w4a16 | Qualcomm® Dragonwing™ Q-8750 | 4096 | 9.940364 | 0.18946806451612902 - 6.062978064516129 |
This model may not be used for or in connection with any of the following applications: