Downloads · 30 days
0
Eternalgrey/llama-snapdragon-models
llama-snapdragon-models is a machine learning model from Eternalgrey. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Llama 3.2 models (1B and 3B) compiled for Qualcomm Snapdragon's NPU. These run natively on-device using the Hexagon processor - no cloud, no GPU needed.
Downloads · 30 days
0
Access
Public
Updated Nov 19, 2025
Repo size
5.2 GB
Likes
0
Public
Click a slice to open those files.
.bin6.5 GB · 100%
From the Hugging Face model README
Llama 3.2 models (1B and 3B) compiled for Qualcomm Snapdragon's NPU. These run natively on-device using the Hexagon processor - no cloud, no GPU needed.
This repo has the compiled model binaries for running Llama 3.2 on Snapdragon devices:
Both models are split into parts due to file size. The 1B model includes separate token generation parts.
These work on Snapdragon chipsets with an NPU:
Basically if you have a flagship Android phone from the last ~2 years, you're good.
These models are meant to be used with Qualcomm's QNN SDK and the Genie runtime. You can't just drop them into any LLM app - they're compiled specifically for Snapdragon's architecture.
If you want a working Android app that uses these, check out the ChatApp implementation (uses the same model format).
On a Samsung Galaxy S25 Ultra (Snapdragon 8 Elite):
All running 100% on-device using the NPU. No internet required after download.
The models use runtime path placeholders in the config that get replaced by the app at load time. This makes them portable across different Android storage configurations.
Most on-device LLMs either run on the GPU (battery killer) or use generic CPU inference (slow). Snapdragon's NPU is purpose-built for this stuff and it shows - you get decent inference speed without destroying your battery.
Plus it's genuinely cool to have a 3 billion parameter model running entirely on your phone.
These are Meta's Llama 3.2 models, so they're under Meta's Llama license. Check Meta's official license terms before using in production.
Models were compiled using Qualcomm's tools but the underlying architecture is Meta's IP.
Direct links if you want to grab individual files:
https://huggingface.co/Eternalgrey/llama-snapdragon-models/res olve/main/llm/[filename] https://huggingface.co/Eternalgrey/llama-snapdragon-models/res olve/main/llama32_1b/[filename]
Replace [filename] with the actual .bin file name.