Downloads · 30 days
381
10% of all-time downloads
RedHatAI/Phi-4-mini-instruct-FP8-dynamic
Phi-4-mini-instruct-FP8-dynamic is a text generation model from RedHatAI. Use it when you need the model to write or continue text. The card lists the license as mit.
<h1 align: center; style="display: flex; align-items: center; gap: 10px; margin: 0;" Phi-4-mini-instruct-FP8-dynamic <img src="https://www.redhat.com/rhdc/managed-files/Catalog-Validatedmodel0.png" alt="Model Icon" wi…
Downloads · 30 days
381
10% of all-time downloads
All-time downloads
3.7K
Public
Parameters
4.5B
5.7 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors5.7 GB · 100%
How the weights are stored.
F8_E4M33.2B · 72%
From the Hugging Face model README
This model was obtained by quantizing activation and weights of Phi-4-mini-instruct to FP8 data type. This optimization reduces the number of bits used to represent weights and activations from 16 to 8, reducing GPU memory requirements (by approximately 50%) and increasing matrix-multiply compute throughput (by approximately 2x). Weight quantization also reduces disk size requirements by approximately 50%.
Only weights and activations of the linear operators within transformers blocks are quantized. Weights are quantized with a symmetric static per-channel scheme, whereas activations are quantized with a symmetric dynamic per-token scheme. The llm-compressor library is used for quantization.
This model can be deployed efficiently using the vLLM backend, as shown in the example below.
vllm serve RedHatAI/Phi-4-mini-instruct-FP8-dynamic --max_model_len 131072
from openai import OpenAI
# Set OpenAI's API key and API base to use vLLM's API server.
openai_api_key = "EMPTY"
openai_api_base = "http://localhost:8000/v1"
client = OpenAI(
api_key=openai_api_key,
base_url=openai_api_base,
)
generated_text = client.chat.completions.create(
model="RedHatAI/Phi-4-mini-instruct-FP8-dynamic",
messages=[
{"role": "user", "content": "Give me a short introduction to large language model."},
],
)
print(generated_text.choices[0].message.content)
from transformers import AutoModelForCausalLM, AutoTokenizer
from llmcompressor.modifiers.quantization import QuantizationModifier
from llmcompressor import oneshot
# Load model
model_stub = "microsoft/Phi-4-mini-instruct"
model_name = model_stub.split("/")[-1]
tokenizer = AutoTokenizer.from_pretrained(model_stub)
model = AutoModelForCausalLM.from_pretrained(
model_stub,
device_map="auto",
torch_dtype="auto",
)
# Configure the quantization algorithm and scheme
recipe = QuantizationModifier(
targets="Linear",
scheme="FP8_dynamic",
ignore=["lm_head"],
)
# Apply quantization
oneshot(
model=model,
recipe=recipe,
)
# Save to disk in compressed-tensors format
save_path = model_name + "-FP8-dynamic"
model.save_pretrained(save_path)
tokenizer.save_pretrained(save_path)
print(f"Model and tokenizer saved to: {save_path}")
</details>
The model was evaluated on the Mathh 500 benchmarks using lighteval, and on GSM8k-Platinum, MMLU CoT, MMLU-Pro, and IFEval using lm-evaluation-harness. In both cases vLLM is used as the backend
<details> <summary>Evaluation commands</summary>vllm serve RedHatAI/Phi-4-mini-instruct-FP8-dynamic --max_model_len 131072
lm_eval --model local-chat-completions \
--tasks gsm8k_platinum_cot_llama \
--model_args "model=RedHatAI/Phi-4-mini-instruct-FP8-dynamic,max_length=131072,base_url=http://0.0.0.0:8000/v1/chat/completions,num_concurrent=128,max_retries=3,tokenized_requests=False,timeout=600,tokenizer_backend=None" \
--apply_chat_template \
--num_fewshot 5 \
--fewshot_as_multiturn \
--output_path gsm8k_platinum_phi4_mini_instruct_fp8_dynamic \
--gen_kwargs "do_sample=False,temperature=0.0,max_gen_toks=16000"
lm_eval --model local-chat-completions \
--tasks mmlu_cot_llama \
--model_args "model=RedHatAI/Phi-4-mini-instruct-FP8-dynamic,max_length=131072,base_url=http://0.0.0.0:8000/v1/chat/completions,num_concurrent=128,max_retries=3,tokenized_requests=False,timeout=600,tokenizer_backend=None" \
--apply_chat_template \
--output_path mmlu_cot_phi4_mini_instruct_fp8_dynamic \
--gen_kwargs "do_sample=False,temperature=0.0,max_gen_toks=16000"
lm_eval --model local-chat-completions \
--tasks mmlu_pro_chat \
--model_args "model=RedHatAI/Phi-4-mini-instruct-FP8-dynamic,max_length=131072,base_url=http://0.0.0.0:8000/v1/chat/completions,num_concurrent=128,max_retries=3,tokenized_requests=False,timeout=600,tokenizer_backend=None" \
--apply_chat_template \
--num_fewshot 5 \
--fewshot_as_multiturn \
--output_path mmlu_pro_phi4_mini_instruct_fp8_dynamic \
--gen_kwargs "do_sample=False,temperature=0.0,max_gen_toks=16000"
lm_eval --model local-chat-completions \
--tasks ifeval \
--model_args "model=RedHatAI/Phi-4-mini-instruct-FP8-dynamic,max_length=131072,base_url=http://0.0.0.0:8000/v1/chat/completions,num_concurrent=128,max_retries=3,tokenized_requests=False,timeout=600,tokenizer_backend=None" \
--apply_chat_template \
--output_path ifeval_phi4_mini_instruct_fp8_dynamic \
--gen_kwargs "do_sample=False,temperature=0.0,max_gen_toks=16000"
litellm_config.yaml
model_parameters:
provider: "hosted_vllm"
model_name: "hosted_vllm/RedHatAI/Phi-4-mini-instruct-FP8-dynamic"
base_url: "http://0.0.0.0:8000/v1"
api_key: ""
timeout: 600
concurrent_requests: 128
generation_parameters:
temperature: 0.0
max_new_tokens: 16000
lighteval endpoint litellm litellm_config.yaml \
math_500|0 \
--output-dir phi4_mini_instruct_fp8_dynamic \
--save-details
</details>