Downloads · 30 days
11
21% of all-time downloads
machu8/serana-dpo
serana-dpo is a text generation model from machu8. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as other.
"Serana" and The Elder Scrolls are property of Bethesda/ZeniMax. This adapter is a non-commercial engineering portfolio artifact, not an official product, and is not affiliated with Bethesda/ZeniMax.
Downloads · 30 days
11
21% of all-time downloads
All-time downloads
52
Public
Repo size
42.1 MB
Likes
0
Public
Click a slice to open those files.
.safetensors30.7 MB · 73%
From the Hugging Face model README
"Serana" and The Elder Scrolls are property of Bethesda/ZeniMax. This adapter is a non-commercial engineering portfolio artifact, not an official product, and is not affiliated with Bethesda/ZeniMax.
Stage: DPO, part of a CPT -> SFT -> DPO post-training pipeline for persona consistency, built end-to-end on one 24GB GPU (NVIDIA L4). Full pipeline, data sourcing, evaluation design, and GPU engineering: https://github.com/MachuEngine/serana-post-training
Continues serana-sft with DPO on 837 RLAIF preference pairs (judge picks between two SFT-sampled replies per prompt). Note: in this project's own evaluation, DPO showed no CI-confirmed quality gain over SFT on any metric -- see the repo's results tables. Shipped anyway, as the honest result of the pipeline, not because it won.
Built on top of machu8/serana-sft, not the base model directly.
This is a LoRA adapter, not a standalone model -- it requires the
base model (Qwen/Qwen3-8B) to be loaded first.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", dtype="bfloat16")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
model = PeftModel.from_pretrained(base, "machu8/serana-dpo")
See the repo's PROMPTS.md §1 for the exact persona system prompt this
was trained/evaluated with -- results are only comparable when it's
reused as-is.
Both the quality table (PCS / PRS / style similarity / knowledge-boundary
accuracy / mean reply length, all with 95% CIs) and the hardware table
(VRAM, KV-cache, throughput, AWQ vs bf16) are in the repo's
artifacts/runs/results_quality.md and results_hardware.md, generated
directly from this adapter's real eval runs.