Downloads · 30 days
6
11% of all-time downloads
Mehdi-Zogh/MNLP_M3_dpo_model
MNLP_M3_dpo_model is a question answering model from Mehdi-Zogh. Use it when the input is a question plus a passage. It is set up for transformers. The card lists the license as apache-2.0.
This model is a Direct Preference Optimization (DPO) fine-tuned version of Qwen3-0.6B-Base using the Mehdi-Zogh/MNLPM3dpodataset. The goal was to improve the alignment of the base model's outputs with human preference…
Downloads · 30 days
6
11% of all-time downloads
All-time downloads
56
Public
Parameters
596M
2.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.4 GB · 99%
From the Hugging Face model README
This model is a Direct Preference Optimization (DPO) fine-tuned version of Qwen3-0.6B-Base using the Mehdi-Zogh/MNLP_M3_dpo_dataset. The goal was to improve the alignment of the base model's outputs with human preferences for educational assistance use cases.
This model was fine-tuned via the DPO (Direct Preference Optimization) algorithm on top of Qwen3-0.6B-Base. The dataset used for preference learning consists of query-response pairs with annotated preference labels, aiming to teach the model to generate more helpful, appropriate, and preferred responses in instructional contexts.
This model is trained to be an AI tutor that is specialized in course content at EPFL.
It can serve as a base model for further alignment, personalization, or integration into interactive educational platforms or tutoring systems.
prompt = "What are the phases of cell division?"
# Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained("Mehdi-Zogh/MNLP_M3_dpo_model", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("Mehdi-Zogh/MNLP_M3_dpo_model", device_map="auto", trust_remote_code=True)
# Tokenize
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
# Generate response
outputs = model.generate(
**inputs,
max_new_tokens=500,
temperature=0.7,
top_p=0.9,
do_sample=True,
eos_token_id=tokenizer.eos_token_id,
)
# Decode and print
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
The training data is the Mehdi-Zogh/MNLP_M3_dpo_dataset, which contains instructional prompts with ranked preferred and rejected completions. The dataset is specifically designed for alignment research using preference optimization methods.
The model was fine-tuned using trl's DPOTrainer
| Hyperparameter | Value |
|---|---|
| Learning rate | 1e-6 |
| Epochs | 3 |
| Per-device train batch size | 1 |
| Per-device eval batch size | 1 |
| Gradient accumulation steps | 4 |
| Precision | bf16 |
| Early stopping patience | 3 |
900 samples out of the dataset were used for validation.
The model was tested on zechen-nlp/MNLP_dpo_evals