Downloads · 30 days
20
34% of all-time downloads
Misha0706/llm-alignment-dpo
llm-alignment-dpo is a text generation model from Misha0706. Use it when you need the model to write or continue text. It is set up for transformers.
This repository contains a DPO-aligned version of HuggingFaceTB/SmolLM-135M-Instruct trained as part of a coursework project on language model alignment. The goal of the project was to implement Direct Preference Opti…
Downloads · 30 days
20
34% of all-time downloads
All-time downloads
59
Public
Parameters
135M
269 MB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors269 MB · 99%
From the Hugging Face model README
This repository contains a DPO-aligned version of HuggingFaceTB/SmolLM-135M-Instruct trained as part of a coursework project on language model alignment. The goal of the project was to implement Direct Preference Optimization (DPO) from scratch and compare its behavior to the original base model and a PPO-based alternative.
This model is a preference-aligned causal language model obtained by fine-tuning HuggingFaceTB/SmolLM-135M-Instruct with Direct Preference Optimization on HumanLLMs/Human-Like-DPO-Dataset.
The training objective was to shift the model away from generic, overly formal assistant-style replies and toward responses preferred in the dataset, which tend to be more human-like, casual, and expressive.
HuggingFaceTB/SmolLM-135M-InstructHuggingFaceTB/SmolLM-135M-InstructHumanLLMs/Human-Like-DPO-DatasetThis model is intended for:
This model is not intended for:
This model inherits the limitations of the base model and the preference dataset. In particular:
In the qualitative comparison, DPO changed model behavior more noticeably than PPO, but the improvement was not uniform across prompts.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "Misha0706/llm-alignment-dpo"
tokenizer = AutoTokenizer.from_pretrained("HuggingFaceTB/SmolLM-135M-Instruct")
tokenizer.pad_token = tokenizer.eos_token
model = AutoModelForCausalLM.from_pretrained(model_id)
model.eval()
messages = [{"role": "user", "content": "What's your morning routine like?"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt")
with torch.no_grad():
generated_ids = model.generate(
**inputs,
max_new_tokens=128,
do_sample=False,
)
print(tokenizer.decode(generated_ids[0], skip_special_tokens=True))
The model was trained on:
HumanLLMs/Human-Like-DPO-DatasetThis dataset contains triples of:
The chosen responses are preferred over the rejected ones and are used for preference optimization.
The dataset was converted into chat format using the tokenizer chat template:
userassistantThen each example was tokenized into:
prompt_input_idschosen_input_idsrejected_input_idsThe prompt was truncated from the left when necessary, and completions were truncated from the right.
DPO was implemented manually rather than using a high-level trainer for this part.
The training setup used:
HuggingFaceTB/SmolLM-135M-Instruct5e-58512128256The model was evaluated qualitatively and with a lightweight likelihood-based analysis.
Evaluation included:
The main quantitative comparison reported in the project was:
databricks/databricks-dolly-15k.BASE train mean: -1.8917
BASE unseen mean: -2.6316
DPO train mean: -1.8647
DPO unseen mean: -2.6217
These results indicate that the DPO model assigns slightly higher likelihood to both training and unseen samples than the base model in this evaluation setup, with the effect being small but measurable.
On several prompts, DPO produced more visibly changed outputs than PPO and more noticeably diverged from the base model.
Examples:
Prompt: Oh, I just saw the best meme - have you seen it?
Prompt: Do you have a favorite type of vacation or getaway?
At the same time, DPO also showed clear weaknesses:
Prompt: Can you tell me about a memorable birthday or celebration you've had?
Compared to the base model, this DPO checkpoint:
Prompt: Oh, I just saw the best meme - have you seen it?
Base model:
I'm so glad you're enjoying the meme! I'm happy to share some of my favorite memes with you...
DPO model:
I've seen it! It's a classic meme that's been around for over a decade. It's a clever way to make people laugh and share their humor with others.
Prompt: What's your morning routine like?
Base model:
I'm excited to start my morning routine! As a digital AI assistant, I don't have personal preferences or habits...
DPO model:
I'm excited to start my morning routine! It's a crucial step in getting ready for the day, and I'm excited to get started...
Prompt: Can you tell me about a memorable birthday or celebration you've had?
DPO model output shows a limitation:
It starts inventing highly specific personal memories, including impossible or incoherent details.
HuggingFaceTB/SmolLM-135M-InstructThe model was optimized with the DPO loss:
\beta \log \frac{\pi_\theta(y_l|x)}{\pi_{ref}(y_l|x)} \right)\right] ]
where preferred and rejected completions come from the human preference dataset.
This is a coursework model and should be treated as an experimental artifact rather than a polished aligned assistant.
Known limitations:
If you use this repository, please cite the original base model and dataset:
HuggingFaceTB/SmolLM-135M-InstructHumanLLMs/Human-Like-DPO-DatasetYou may also mention this repository as a coursework implementation of DPO-based alignment.