Skip to content

HermitQ

NPCAlign-DPO

HermitQ/NPCAlign-DPO

NPCAlign-DPO is a machine learning model from HermitQ. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as llama3.1.

LoRA adapter further fine-tuned via Direct Preference Optimisation (DPO) on top of the SFT model. Trained to generate more natural conversation endings and diverse NPC responses.

Downloads · 30 days

16

11% of all-time downloads

All-time downloads

141

Public

Repo size

185 MB

Likes

1

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors168 MB · 91%

At a glance

Library
peft
License
llama3.1
Access
Public
Created
Jul 8, 2026
Updated
Aug 10, 2026
SHA
69d5a595

Base models

Library
peft
License
llama3.1
Languages
en
Created
Jul 8, 2026
Updated
Aug 10, 2026
NPCAlign-DPO — AI Model — AIMarketly