Downloads · 30 days
18
5% of all-time downloads
NexaAI/phi4-mini-npu-turbo
phi4-mini-npu-turbo is a text generation model from NexaAI. Use it when you need the model to write or continue text.
Run Phi-4-mini optimized for Qualcomm NPUs with nexaSDK.
Downloads · 30 days
18
5% of all-time downloads
All-time downloads
357
Public
Repo size
5 GB
Likes
2
Public
Click a slice to open those files.
.nexa5 GB · 100%
From the Hugging Face model README
Run Phi-4-mini optimized for Qualcomm NPUs with nexaSDK.
Install nexaSDK and create a free account at sdk.nexa.ai
Activate your device with your access token:
nexa config set license '<access_token>'
Run the model on Qualcomm NPU in one line:
nexa infer NexaAI/phi4-mini-npu-turbo
Phi-4-mini is a ~3.8B-parameter instruction-tuned model from Microsoft’s Phi-4 family. Trained on a blend of synthetic “textbook-style” data, filtered public web content, curated books/Q&A, and high-quality supervised chat data, it emphasizes reasoning-dense capabilities while maintaining a compact footprint. This NPU Turbo build uses Nexa’s Qualcomm backend (QNN/Hexagon) to deliver lower latency and higher throughput on-device, with support for 128K context and efficient long-context memory handling.
Input:
Output:
This model is released under the Creative Commons Attribution–NonCommercial 4.0 (CC BY-NC 4.0) license.
Non-commercial use, modification, and redistribution are permitted with attribution.
For commercial licensing, please contact [email protected].
📰 Phi-4-mini Microsoft Blog <br> 📖 Phi-4-mini Technical Report <br> 👩🍳 Phi Cookbook <br> 🚀 Model paper