Downloads · 30 days
11
23% of all-time downloads
rokugatsu/LLM2025_Advanced_DPO_5_test
LLM2025_Advanced_DPO_5_test is a text generation model from rokugatsu. Use it when you need the model to write or continue text. It is set up for trl. The card lists the license as apache-2.0.
This repository provides a DPO-fine-tuned model based on rokugatsu/LLM2025Advanced5test using trl.DPOTrainer.
Downloads · 30 days
11
23% of all-time downloads
All-time downloads
48
Public
Parameters
4B
8.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
This repository provides a DPO-fine-tuned model based on
rokugatsu/LLM2025_Advanced_5_test using trl.DPOTrainer.
This model has undergone Direct Preference Optimization (DPO) to align with human preferences, using trajectories from agent-based tasks.
This model was fine-tuned using DPO to improve multi-turn agent task performance
by learning preferences from the u-10bei/sft_alfworld_trajectory_dataset_v2,u-10bei/sft_alfworld_trajectory_dataset_v4 dataset.
The DPO training process aims to increase the likelihood of generating 'chosen' responses
and decrease the likelihood of 'rejected' responses for given prompts.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
import torch
model_id = "rokugatsu/LLM2025_Advanced_DPO_5_test"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16, # Use bfloat16 if your GPU supports it
device_map="auto",
)
# The model is already merged, so no need for PeftModel.from_pretrained(model, adapter)
# Example for inference (assuming you have a chat_template)
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the capital of France?"}
]
input_ids = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to(model.device)
outputs = model.generate(input_ids, max_new_tokens=256)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Training data: u-10bei/sft_alfworld_trajectory_dataset_v2,u-10bei/sft_alfworld_trajectory_dataset_v4
Dataset License: MIT License. This dataset is used and distributed under the terms of the MIT License. Compliance: Users must comply with the MIT license (including copyright notice) and the base model's original terms of use.