Skip to content

Misha0706

llm-alignment-dpo

Misha0706/llm-alignment-dpo

llm-alignment-dpo is a text generation model from Misha0706. Use it when you need the model to write or continue text. It is set up for transformers.

This repository contains a DPO-aligned version of HuggingFaceTB/SmolLM-135M-Instruct trained as part of a coursework project on language model alignment. The goal of the project was to implement Direct Preference Opti…

Downloads · 30 days

20

34% of all-time downloads

All-time downloads

59

Public

Parameters

135M

269 MB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors269 MB · 99%

At a glance

Task
Text Generation
Library
transformers
Model type
llama
Access
Public
Created
Apr 8, 2026
Updated
Apr 13, 2026
SHA
1d662c5f

Try a prompt

Base models

Task
Text Generation
Library
transformers
Type
llama
Languages
en
Created
Apr 8, 2026
Updated
Apr 13, 2026