Skip to content

tsessk

llm-course-hw2-dpo

tsessk/llm-course-hw2-dpo

llm-course-hw2-dpo is a text generation model from tsessk. Use it when you need the model to write or continue text. It is set up for transformers.

This is a Direct Preference Optimization (DPO) fine-tuned version of SmolLM-135M, designed to make its responses more emotional and engaging. The model has been optimized using HumanLLMs/Human-Like-DPO-Dataset.

Downloads · 30 days

18

27% of all-time downloads

All-time downloads

67

Public

Parameters

135M

269 MB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors269 MB · 98%

At a glance

Task
Text Generation
Library
transformers
Model type
llama
Access
Public
Created
Mar 6, 2025
Updated
Mar 8, 2025
SHA
e0201705

Try a prompt

Base models

Task
Text Generation
Library
transformers
Type
llama
Languages
en
Created
Mar 6, 2025
Updated
Mar 8, 2025