Downloads · 30 days
18
16% of all-time downloads
DeryFerd/Qwen2.5-Math-7B-Instruct-Distill-Phi2-2.5K-MixMath
Qwen2.5-Math-7B-Instruct-Distill-Phi2-2.5K-MixMath is a text generation model from DeryFerd. Use it when you need the model to write or continue text. It is set up for transformers.
This model is a version of microsoft/phi-2 that has been fine-tuned using knowledge distillation. The goal was to teach the compact and efficient Phi-2 "student" model to replicate the step-by-step mathematical reason…
Downloads · 30 days
18
16% of all-time downloads
All-time downloads
115
Public
Parameters
2.8B
5.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors5.6 GB · 100%
From the Hugging Face model README
This model is a version of microsoft/phi-2 that has been fine-tuned using knowledge distillation. The goal was to teach the compact and efficient Phi-2 "student" model to replicate the step-by-step mathematical reasoning style of the powerful Qwen/Qwen2.5-Math-7B-Instruct "teacher" model.
This is the V2 of this project, featuring a significantly larger and more diverse dataset than the V1 model, resulting in more robust reasoning and the ability to correctly generate LaTeX math notation.
This project explores style distillation, where a smaller model is trained not just on correct answers, but on the process and format of a larger, more capable model's output. The primary objective was to transfer the teacher's verbose, step-by-step reasoning methodology, including its use of LaTeX, to the student model.
The model was fine-tuned using the QLoRA method for high memory efficiency, making it possible to train on consumer-grade hardware.
microsoft/phi-2Use the code below to load the fine-tuned model adapter and run inference.
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
# Your repository ID
repo_id = "DeryFerd/Qwen2.5-Math-Instruct-Distill-Phi2-2.5K-Mixed"
base_model_id = "microsoft/phi-2"
# Load the base model
base_model = AutoModelForCausalLM.from_pretrained(
base_model_id,
torch_dtype=torch.float16,
device_map="auto",
trust_remote_code=True
)
# Load the tokenizer
tokenizer = AutoTokenizer.from_pretrained(repo_id)
# Load the PEFT model by merging the adapter into the base model
model = PeftModel.from_pretrained(base_model, repo_id)
model.eval()
# --- Run Inference ---
instruction = "A right-angled triangle has two shorter sides with lengths of 8 cm and 15 cm. What is the length of the longest side (the hypotenuse)? Use the Pythagorean theorem (a^2 + b^2 = c^2) to solve it."
prompt = f"Instruct: {instruction.strip()}\nOutput:"
inputs = tokenizer(prompt, return_tensors="pt", return_attention_mask=False).to(model.device)
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=300, pad_token_id=tokenizer.eos_token_id)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
final_answer = response.split("Output:")[1].strip()
print(final_answer)
The model was trained on a combined dataset of 2,500 examples created specifically for this distillation task. The dataset is a mix of two sources to improve robustness and mathematical formatting capabilities:
The answers ("response" column) were not the original dataset answers, but were synthetically generated by the Qwen/Qwen2.5-Math-7B-Instruct teacher model to provide high-quality, step-by-step reasoning examples for the student to learn from.
The model was fine-tuned using the SFTTrainer from the TRL library on a single NVIDIA T4 GPU in a Kaggle Notebook environment.
Each data sample was formatted into a single string using the following template, which is suitable for the Phi-2 model: Instruct: {instruction}\nOutput: {response}<|endoftext|>
nf4 with float16 compute dtyper: 16alpha: 32target_modules: ["q_proj", "k_proj", "v_proj", "dense"]per_device_train_batch_size = 1, gradient_accumulation_steps = 8 (effective batch size of 8)paged_adamw_8bit2e-4 with a constant schedulerfp16Evaluation was performed qualitatively by comparing the outputs of three models (Base Phi-2, this fine-tuned Student model, and the Teacher Qwen model) on a variety of math problems.
This is an experimental model trained for a specific purpose and has several limitations:
microsoft/phi-2 and teacher Qwen/Qwen2.5-Math-7B-Instruct models.This model should not be used for production or critical applications. It is intended as a portfolio project to demonstrate the effectiveness of knowledge distillation.