Downloads · 30 days
16
11% of all-time downloads
HermitQ/NPCAlign-DPO
NPCAlign-DPO is a machine learning model from HermitQ. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as llama3.1.
LoRA adapter further fine-tuned via Direct Preference Optimisation (DPO) on top of the SFT model. Trained to generate more natural conversation endings and diverse NPC responses.
Downloads · 30 days
16
11% of all-time downloads
All-time downloads
141
Public
Repo size
185 MB
Likes
1
Public
Click a slice to open those files.
.safetensors168 MB · 91%
From the Hugging Face model README
LoRA adapter further fine-tuned via Direct Preference Optimisation (DPO) on top of the SFT model. Trained to generate more natural conversation endings and diverse NPC responses.
This adapter is applied on top of the merged SFT model, not directly on the base Llama model. See Usage section.
Note: The base model
meta-llama/Meta-Llama-3.1-8B-Instructis a gated model. You must accept Meta's license and set yourHF_TOKENbefore loading.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
base = AutoModelForCausalLM.from_pretrained(
"meta-llama/Meta-Llama-3.1-8B-Instruct",
torch_dtype=torch.bfloat16, device_map="auto"
)
model = PeftModel.from_pretrained(base, "HermitQ/NPCAlign-DPO")
tokenizer = AutoTokenizer.from_pretrained("HermitQ/NPCAlign-DPO")
sft = PeftModel.from_pretrained(base, "HermitQ/NPCAlign-SFT")
merged = sft.merge_and_unload()
model = PeftModel.from_pretrained(merged, "HermitQ/NPCAlign-DPO")
| Parameter | Value |
|---|---|
| Beta | 0.1 |
| Epochs | 2 |
| Learning rate | 5e-5 |
| Preference pairs | 1,341 |
| Best checkpoint | Step 210 / 300 |
| Best reward margin | 2.053 |
| Best reward accuracy | 83.1% |
| Metric | SFT | DPO | Change |
|---|---|---|---|
| ROUGE-L | 0.251 | 0.206 | -0.045 |
| Self-BLEU | 0.264 | 0.187 | -0.078 ↓ more diverse |
| BERTScore-F1 | 0.883 | 0.871 | -0.012 |
| BLEURT | -0.710 | -0.840 | -0.13 |
Self-BLEU decrease indicates more diverse generation.
GitHub link: Hermit888/NPCAlign