Downloads · 30 days
0
IAO26/sim-student-3b
sim-student-3b is a machine learning model from IAO26. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
Open release (public good). The adapter weights here are a derivative of the Apache-2.0 Qwen2.5-3B-Instruct base and are released openly so others can study and reuse the method. Not included: the SFT training data (i…
Downloads · 30 days
0
Access
Public
Updated Jun 30, 2026
Repo size
251 MB
Likes
0
Public
Click a slice to open those files.
.safetensors240 MB · 88%
From the Hugging Face model README
Open release (public good). The adapter weights here are a derivative of the Apache-2.0
Qwen2.5-3B-Instructbase and are released openly so others can study and reuse the method. Not included: the SFT training data (it's distilled from a frontier model — its redistribution is a separate question and is deliberately left out of this repo). What's published is the adapter, the card, and the measured before/after results — enough to load, evaluate, and reproduce the recipe.
A LoRA adapter that turns Qwen2.5-3B-Instruct into a simulated struggling-student learner —
a synthetic teen the team can use to train and evaluate college-advising systems (human advisors or
advising AI) against realistic learner failure modes.
Qwen/Qwen2.5-3B-Instruct (Apache-2.0)all-linear, assistant-only lossadapter_v2/ ← the headline model · Together job ft-ce2b985d-1642 · augmented not-knowing, held-out probes · fabrication 40%→31%, sycophancy 24%→12%. Use this one.adapter_v1/ ← natural-gold baseline · Together job ft-d8c56013-d7de · sycophancy 24%→8%, fabrication flat 40% (the run that revealed the train/eval data gap).Intended use (internal): drive a synthetic, in-character teen learner so advisors / advising-AI can practice against realistic failure modes — sycophancy under bad advice, partial not-knowing, in-persona affect — and so we can measure an advising system's quality against a consistent simulated student.
NOT for:
Same 60-probe battery, same ruler (A1's de-biased gpt-4o judge + correctness audit), applied identically to both arms. Directional (n is small — see limits), not tight confidence intervals.
| dimension | scale / direction | BEFORE (prompted 3B) | AFTER (LoRA) | read |
|---|---|---|---|---|
| sycophancy (cave) | % failed, ↓ better | 24% | 8–12% | WIN — ~2–3× cut, lands at the Qwen3-4B-Instruct reference (8%) |
| confident-fabrication | % failed, ↓ better | 40% | 31% | moved on held-out probes = generalization, not memorization |
| true-leak (correct expert content) | % failed, ↓ better | 6% | 3–6% | low both arms |
| overshoot (competence-beyond-profile) | x/5, ↑ better | ~5/5 | 5/5 | clean both arms — 3B doesn't overshoot; not a differentiator at this scale |
| drift (persona consistency, PFC) | 0–1, ↑ better | 1.0 | 1.0 | clean both — artifact of per-turn persona re-grounding (uninformative) |
| affect (in-persona emotion) | x/5, ↑ better | ~4.5 | ~3.75 | mild regression — FT flattened emotion slightly (honest trade-off; low-conf scorer) |
⚠️ Read the scale direction per row. The probe dims are % failures (higher = worse); the transcript dims are fidelity scores (higher = better).
overshoot 5/5means no overshoot (good), not "everyone has the issue."
Headline (defensible): a $4 LoRA on 212 frontier-distilled examples cut a cheap open model's sycophancy ~2–3× (24%→8–12%), to frontier-reference levels, and moved confident-fabrication 40%→31% on held-out probes — directional proof the open-model + SFT method works on the fidelity axis.
Cost story (the thesis): ~$4 to train, ~$12 total; at saturated inference the open 3B is ~10–40× cheaper per conversation than GPT-4o. Framed as fidelity-per-dollar — the cheapest open model moved measurably toward frontier behavior. (Cheap only when the serving endpoint is busy; idle = waste.)
METHOD_FINDING_train_eval_gap.md. The fabrication drop is uneven (driven by one persona,
yesenia; tasha flat → a per-persona topic-coverage gap = the concrete next lever).The weights are in this repo under adapter_v2/ (headline) and adapter_v1/ (baseline) — each a
full LoRA adapter + tokenizer. Load either with PEFT (point at the subfolder you want):
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
from huggingface_hub import snapshot_download
repo = "IAO26/sim-student-3b"
local = snapshot_download(repo) # private repo → needs `huggingface-cli login`
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-3B-Instruct")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-3B-Instruct")
model = PeftModel.from_pretrained(base, f"{local}/adapter_v2") # or adapter_v1 to compare
Re-export from Together (if you ever need to):
together fine-tuning download <job-id> --output-dir <out> --checkpoint-type adapter → unpack the
.tar.zst (it's zstd→tar). Job IDs: v2 = ft-ce2b985d-1642, v1 = ft-d8c56013-d7de.
To reproduce the serving we used: qwen/endpoint.py (Together dedicated endpoint) +
qwen/together_client.py. The endpoint name is the working model-id (the slug 400s).
The comparison view in this repo is static but regeneratable — each new training round updates it:
qwen/finetune.py --run (new adapter)before_baseline.py → judge_rescore.py → leak_audit.py → gen_convos.pyresults/ (scoreboard + transcripts + compare.html) and git push to the HF repo