Downloads · 30 days
0
qualcomm/Phi-3.5-mini-instruct
Phi-3.5-mini-instruct is a text generation model from qualcomm. Use it when you need the model to write or continue text. It is set up for pytorch. The card lists the license as mit.
Downloads · 30 days
0
Access
Public
Updated Sep 23, 2026
Repo size
110 KB
Likes
0
Public
Click a slice to open those files.
.md5.8 KB · 68%
From the Hugging Face model README

Phi-3.5-mini is a lightweight, state-of-the-art open model built upon datasets used for Phi-3 - synthetic data and filtered publicly available websites - with a focus on very high-quality, reasoning dense data. The model belongs to the Phi-3 model family and supports 128K token context length. The model underwent a rigorous enhancement process, incorporating both supervised fine-tuning, proximal policy optimization, and direct preference optimization to ensure precise instruction adherence and robust safety measures.
This is based on the implementation of Phi-3.5-Mini-Instruct found here. This repository contains pre-exported model files optimized for Qualcomm® devices. You can use the Qualcomm® AI Hub Models library to export with custom configurations. More details on model performance across various devices, can be found here.
Qualcomm AI Hub Models uses Qualcomm AI Hub Workbench to compile, profile, and evaluate this model. Sign up to run these models on a hosted Qualcomm® device.
Please follow the LLM on-device deployment tutorial.
Below are pre-exported model assets ready for deployment.
| Runtime | Precision | Chipset | SDK Versions | Download |
|---|---|---|---|---|
| GENIE | w4a16 | Snapdragon® X2 Elite | QAIRT 2.43 | Download |
| GENIE | w4a16 | Snapdragon® X Elite | QAIRT 2.43 | Download |
For more device-specific assets and performance metrics, visit Phi-3.5-Mini-Instruct on Qualcomm® AI Hub.
Model Type: Model_use_case.text_generation
Model Stats:
| Model | Runtime | Precision | Chipset | Context Length | Response Rate (tokens per second) | Time To First Token (range, seconds) |
|---|---|---|---|---|---|---|
| Phi-3.5-Mini-Instruct | QNN_CONTEXT_BINARY | w4a16 | Snapdragon® X2 Elite | 4096 | 34.25 | 0.08362 - 2.67584 |
| Phi-3.5-Mini-Instruct | QNN_CONTEXT_BINARY | w4a16 | Snapdragon® X Elite | 4096 | 10.2 | 0.10690999999999999 - 5.946656 |
| Phi-3.5-Mini-Instruct | QNN_CONTEXT_BINARY | w4a16 | Qualcomm® Dragonwing™ IQ-X7181 | 4096 | 10.2 | 0.10690999999999999 - 5.946656 |
This model may not be used for or in connection with any of the following applications: