Downloads · 30 days
0
geyuxu/self-training-loop-phase2-adapters
self-training-loop-phase2-adapters is a machine learning model from geyuxu. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
Six LoRA adapters (Qwen2.5-1.5B-Instruct, r16, lr 1e-5, beta 0.1, 2 epochs, seeds 22/23):
Downloads · 30 days
0
Access
Public
Updated Aug 5, 2026
Repo size
116 MB
Likes
0
Public
Click a slice to open those files.
.safetensors105 MB · 60%
From the Hugging Face model README
Six LoRA adapters (Qwen2.5-1.5B-Instruct, r16, lr 1e-5, beta 0.1, 2 epochs, seeds 22/23):
D_full_emf1_seed{22,23} — 10,000 preference pairs, original EM/F1 directionsD_full_corrected_seed{22,23} — same 10,000 pairs with 375 judge-verified hurts-direction flipsD_verified_seed{22,23} — 3,185 semantically verified pairs (hybrid accept pool)Companion dataset repo: self-training-loop-phase2-arms (training arms, eval chain,
human adjudication records). Key result: direction repair at 3.75% flip rate produces
no detectable behavioral difference (two seeds, reference-based and blind pairwise
LLM-judge evaluation) — see the thesis repo Phase 2 execution status doc §15.22.