Downloads · 30 days
16
10% of all-time downloads
nadiva1243/phi4RAG
phi4RAG is a text generation model from nadiva1243. Use it when you need the model to write or continue text. The card lists the license as mit.
Quantized GGUF build of microsoft/phi-4 with a LoRA adapter merged in, fine-tuned for retrieval-augmented question answering. The model answers only from supplied document context in English, Spanish, or Catalan, usin…
Downloads · 30 days
16
10% of all-time downloads
All-time downloads
156
Public
Repo size
9.1 GB
Likes
0
Public
Click a slice to open those files.
.gguf9.1 GB · 100%
From the Hugging Face model README
Quantized GGUF build of microsoft/phi-4 with a LoRA adapter merged in, fine-tuned for retrieval-augmented question answering. The model answers only from supplied document context in English, Spanish, or Catalan, using the same RAG-oriented system prompt as MonkeyGrab, a local, fully private RAG stack developed for a Bachelor's thesis (TFG) at the Universitat Politècnica de València (UPV).
The full MonkeyGrab source code is publicly available at:
The repository includes the complete RAG pipeline, CLI, web interface, training scripts, evaluation workflows, and documentation for the Bachelor's thesis (TFG) at UPV.
This Hugging Face model repo ships inference assets (Phi4-Q4_K_M.gguf), the Ollama Modelfile, and a reproduction/ folder with frozen copies of the training script, merge utility, and evaluation_comparison.json so methodology and metrics remain auditable alongside the full codebase.
Contact: [email protected] for questions about training, evaluation, or Ollama usage.
GGUF pipeline (high level): LoRA fine-tuning on the datasets below → merge with merge_lora.py (see reproduction/) → GGUF export via the llama.cpp toolchain → Q4_K_M quantization. The merge script documents expected paths and flags.
| File | Description |
|---|---|
Phi4-Q4_K_M.gguf | Full weights after LoRA merge, Q4_K_M quantization. |
Modelfile | Ollama recipe: ChatML template, RAG system prompt, sampling parameters. |
README.md | This model card. |
LICENSE | MIT — applies to the model card, Modelfile, and files added here by nadiva1243 (not to Microsoft's base terms). |
reproduction/train-phi4.py | Snapshot of scripts/training/train-phi4.py (v1) used for this adapter. |
reproduction/merge_lora.py | Snapshot of scripts/conversion/merge_lora.py used to merge the LoRA weights into a dense checkpoint before GGUF export. |
reproduction/evaluation_comparison.json | Frozen evaluation export (base vs. adapted, dev/test splits, per dataset + weighted aggregate). |
reproduction/CONVERSION.md | Step-by-step notes: merge → GGUF → Q4_K_M quantization → Ollama import. |
microsoft/phi-4 — 14B-parameter transformer (ChatML-style; end-of-turn token <|im_end|>).| Setting | Value |
|---|---|
r | 64 |
lora_alpha | 128 |
lora_dropout | 0.05 |
target_modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
bias | none |
train-phi4.py, v1)set_seed).<|im_start|>user … <|im_end|> with the instruction and <context>…</context> on the user turn; loss computed only on the assistant completion (prompt labels masked with –100).neural-bridge/rag-dataset-12000databricks/databricks-dolly-15k (categories: closed_qa, information_extraction, summarization) — 80/10/10 split after filterprojecte-aina/RAG_Multilingual — EN, ES, CA subsetsmax_length 4,096 tokens; context truncated to 2,048 tokens; generation up to 2,048 new tokens.per_device_train_batch_size 1, gradient_accumulation_steps 16 → effective batch 16; bf16 + TF32; gradient checkpointing enabled.eval_loss; early stopping patience 3 evaluations.microsoft/phi-4) and the adapted (LoRA merged) model — no data leakage.evaluate_baselines.py for cross-experiment comparability).microsoft/deberta-xlarge-mnli); BERTScore is computed after unloading the generative model to fit in GPU memory.reproduction/evaluation_comparison.json.Values are percentage points (0–100 scale). Δ (pp) = adapted − base; Δ rel (%) = relative change vs. base.
| Split | N | Metric | Base | Adapted | Δ (pp) | Δ rel (%) |
|---|---|---|---|---|---|---|
| Dev | 1,600 | Token F1 | 45.17 | 60.24 | +15.07 | +33.36 |
| Dev | 1,600 | ROUGE-L F1 | 37.18 | 50.49 | +13.31 | +35.79 |
| Dev | 1,600 | BERTScore F1 | 39.59 | 53.48 | +13.89 | +35.07 |
| Test | 8,490 | Token F1 | 45.42 | 63.20 | +17.78 | +39.14 |
| Test | 8,490 | ROUGE-L F1 | 37.21 | 52.97 | +15.76 | +42.35 |
| Test | 8,490 | BERTScore F1 | 39.90 | 56.42 | +16.52 | +41.41 |
| Dataset | Token F1 (Base → Adapted) | ROUGE-L F1 (Base → Adapted) | BERTScore F1 (Base → Adapted) |
|---|---|---|---|
| Neural-Bridge RAG | 50.46 → 81.17 | 45.46 → 77.46 | 46.79 → 79.34 |
| Dolly QA | 44.46 → 50.95 | 38.21 → 45.51 | 38.88 → 46.24 |
| Aina-EN | 44.67 → 56.15 | 35.32 → 43.16 | 41.61 → 50.42 |
| Aina-ES | 40.47 → 57.11 | 31.44 → 43.37 | 33.35 → 45.66 |
| Aina-CA | 45.80 → 55.82 | 35.48 → 42.95 | 37.32 → 45.72 |
Full test-split breakdowns and qualitative sample pairs are in reproduction/evaluation_comparison.json.
The base dev numbers are aligned with the multi-model benchmark (evaluate_baselines.py, predictions_phi-4.json), so Phi-4 before fine-tuning is directly comparable to the other models in that suite. For post-LoRA performance, use the Adapted columns above.
| Setup | Notes |
|---|---|
| GPU (recommended) | ~10 GB VRAM is a practical minimum for this Q4_K_M ~14B-class GGUF in Ollama at moderate batching; 8 GB may work with shorter context or with slower GPU offloading. |
| Context length | The bundled Modelfile sets num_ctx 16384 — raising context increases VRAM/RAM use roughly linearly; reduce num_ctx if you hit OOM. |
| CPU | Supported by Ollama / llama.cpp runners, but significantly slower than a discrete GPU at this model size. |
| Training hardware | LoRA training used bf16, gradient checkpointing, and an 8-bit optimizer on a CUDA GPU (see reproduction/train-phi4.py); this is separate from these inference notes. |
Place Phi4-Q4_K_M.gguf next to Modelfile (or adjust the FROM path). Then:
ollama create phi4-rag -f Modelfile
ollama run phi4-rag
Generation defaults in the bundled Modelfile: num_ctx 16384, temperature 0.15, top_p 0.9, repeat_penalty 1.15.
<context>…</context> tags as in training.Modelfile, and other metadata added by nadiva1243 are released under the MIT License (see the LICENSE file in this repository).microsoft/phi-4. You must also comply with the license and terms of the base model and with any requirements of the training datasets when redistributing or using the weights.@misc{phi4_rag_gguf_monkeygrab,
title = {Phi-4 RAG LoRA Fine-tune (Q4_K_M GGUF)},
author = {nadiva1243},
year = {2026},
howpublished = {Hugging Face: \url{https://huggingface.co/nadiva1243/phi4RAG}},
note = {Base: microsoft/phi-4; training: MonkeyGrab train-phi4.py v1; source: https://github.com/iDiagoValeta/localOllamaRAG}
}