Downloads · 30 days
22
100% of all-time downloads
yusifnuri/phi-4-mini-instruct_ner
phi-4-mini-instruct_ner is a text generation model from yusifnuri. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as mit.
A LoRA adapter that specialises microsoft/Phi-4-mini-instruct (3.80 B parameters) for a single enterprise task: it extracts person, organisation, location and miscellaneous entities from a sentence.
Downloads · 30 days
22
100% of all-time downloads
All-time downloads
22
Public
Repo size
28.1 MB
Likes
0
Public
Click a slice to open those files.
.json15.5 MB · 55%
From the Hugging Face model README
A LoRA adapter that specialises microsoft/Phi-4-mini-instruct (3.80 B parameters) for a single enterprise task: it extracts person, organisation, location and miscellaneous entities from a sentence.
It was produced for the MSc thesis Fine-Tune or Pay Per Token? An Enterprise Benchmark of Small Language Models (SRH University Hamburg), which measures fine-tuned small models against frontier provider APIs on accuracy, latency, cost, privacy exposure and return-on-investment breakeven volume. The adapter is released so that the benchmark can be independently verified.
O tag that tracks output-format imitation rather than entity extraction. The evaluation harness has since been corrected to BIO-decode tag output into entity surface forms and score both arms identically, but this cell has not yet been re-evaluated. Treat the number below as a placeholder, not as performance.| Metric | Value |
|---|---|
| Entity-level F1 (sentence-averaged) | 0.9328 |
| Mean latency, batch 1 | 608 ms |
| Cost per 1M generated tokens | USD 10.53 |
Measured on a single NVIDIA H200 (141 GB) at batch size one and full utilisation, priced at an imputed USD 3.99 per GPU-hour. Latency excludes network transit. Scores are not comparable across tasks — each task carries its own metric. Evaluation ran on 5 July 2026; the complete matrix is at results/benchmark_matrix.csv.
| Method | LoRA |
| Dataset | CoNLL-2003 (English) (eriktks/conll2003) |
| Dataset licence | Reuters terms; redistribution restricted |
| Training examples | 5,000 (500 held out for checkpoint selection) |
| Rank / alpha / dropout | 16 / 32 / 0.05 |
| Target modules | q_proj, k_proj, v_proj, o_proj |
| Learning rate | 2e-4, cosine schedule, 3% warmup |
| Epochs | 3 |
| Effective batch size | 16 (4 x 4 gradient accumulation) |
| Max sequence length | 512 tokens |
| Optimiser | AdamW |
| Seed | 42 |
Hyperparameters were held constant across every model and task rather than tuned per cell, so these figures are a conservative lower bound on attainable performance.
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("microsoft/Phi-4-mini-instruct", device_map="auto")
model = PeftModel.from_pretrained(base, "<your-hf-username>/phi-4-mini-instruct_ner")
tokenizer = AutoTokenizer.from_pretrained("microsoft/Phi-4-mini-instruct")
The adapter was trained on this prompt format and expects it at inference:
Extract named entities (PER=person, ORG=organisation, LOC=location, MISC=miscellaneous) from this text:
{text}
Entities:
results/benchmark_matrix.csvresults/cost_per_request.csv@mastersthesis{nuri2026finetune,
title = {Fine-Tune or Pay Per Token? An Enterprise Benchmark of Small Language Models},
author = {Nuri, Yusif},
school = {SRH University Hamburg},
year = {2026}
}