Downloads · 30 days
52
9% of all-time downloads
FluidInference/phi-4-mini-instruct-int4-ov-npu
phi-4-mini-instruct-int4-ov-npu is a text generation model from FluidInference. Use it when you need the model to write or continue text. The card lists the license as mit.
Model creator: Microsoft Original model: Phi-4-mini-instruct
Downloads · 30 days
52
9% of all-time downloads
All-time downloads
609
Public
Repo size
4.5 GB
Likes
0
Public
Click a slice to open those files.
.bin2.2 GB · 99%
From the Hugging Face model README
This is Phi-4-mini-instruct model converted to the OpenVINO™ IR (Intermediate Representation) format with weights compressed to INT4 by NNCF.
With the following pyproject.yoml
[project]
name = "export"
version = "0.1.0"
description = "Export models"
readme = "README.md"
requires-python = "==3.12.*"
dependencies = [
"openvino==2025.2.0",
"optimum[openvino]",
"optimum-intel",
"openvino-genai",
"huggingface-hub==0.33.0",
"tokenizers==0.21.1"
]
Then run the export
uv sync
uv run optimum-cli export openvino --model microsoft/phi-4-mini-instruct --task text-generation-with-past --weight-format int4 --group-size -1 --ratio 1.0 --sym --trust-remote-code phi-4-mini-instruct/INT4-NPU_compressed_weights
The provided OpenVINO™ IR model is compatible with:
pip install -U openvino openvino-tokenizers openvino-genai
pip install huggingface_hub
import huggingface_hub as hf_hub
model_id = "bweng/phi-4-mini-instruct-int4-ov-npu"
model_path = "phi-4-mini-instruct-int4-ov"
hf_hub.snapshot_download(model_id, local_dir=model_path)
import openvino_genai as ov_genai
device = "NPU"
pipe = ov_genai.LLMPipeline(model_path, "NPU", MAX_PROMPT_LEN=4096, CACHE_DIR="./cache")
# Create a proper GenerationConfig object
gen_config = GenerationConfig(apply_chat_template=True, max_new_tokens=1024)
# Now call generate with the correct config object
output = pipe.generate("How are you doing?", generation_config=gen_config)
print(output)
More GenAI usage examples can be found in OpenVINO GenAI library docs and samples
You can find more detaild usage examples in OpenVINO Notebooks:
Check the original model card for original model card for limitations.
The original model is distributed under mit license. More details can be found in original model card.
Intel is committed to respecting human rights and avoiding causing or contributing to adverse impacts on human rights. See Intel’s Global Human Rights Principles. Intel’s products and software are intended only to be used in applications that do not cause or contribute to adverse impacts on human rights.