Downloads · 30 days
11
69% of all-time downloads
cmpatino/Qwen2.5-7B-Instruct-DirectOPD-R1DistillShift-100
Qwen2.5-7B-Instruct-DirectOPD-R1DistillShift-100 is a text generation model from cmpatino. Use it when you need the model to write or continue text. It is set up for transformers.
Artifact of the Direct-OPD SFT-transfer experiment (direct-opd-sft-transfer, condition r1distill). The question: when a pure-SFT shift encodes real held-out capability, does Direct-OPD transfer that capability — and i…
Downloads · 30 days
11
69% of all-time downloads
All-time downloads
16
Public
Parameters
7.6B
76.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors91.4 GB · 100%
From the Hugging Face model README
Artifact of the Direct-OPD SFT-transfer experiment (direct-opd-sft-transfer, condition
r1distill). The question: when a pure-SFT shift encodes real held-out capability, does
Direct-OPD transfer that capability — and into a student 4.7x larger than the teachers?
The training signal is the token-level shift log pi_post - log pi_pre, evaluated on the
student's own sampled tokens. Neither teacher is imitated; only the difference between them is.
| role | model | what it is |
|---|---|---|
pi_pre — teacher_ref (TEACHER_REF_MODEL_PATH) | Qwen/Qwen2.5-Math-1.5B @ 4a83ca6e4526a4f2da3aa259ec36c259f66b2ab2 | the pre-shift reference: the base math model the distillation started from |
pi_post — teacher (REWARD_MODEL_PATH) | deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B @ ad9f0ae0864d7fbcd1cd905e3c6c5b069cc8b562 | the post-shift model: the same 1.5B architecture after SFT on 800K R1 traces |
| student init | Qwen/Qwen2.5-7B-Instruct @ a09a35458c702b33eeacc393d103063234e8bc28 | non-thinking instruct model, 7.6B params |
Root = step 100. checkpoint-{20,40,60,80,100}/ = intermediate merged
checkpoints. Weights are bf16 (verl's FSDP->HF merge downcasts the fp32 masters).
https://github.com/BytedTsinghua-SIA/Direct-OPD @ 3a9d6bd37b00a38e7a9b2959239e4631e5324aea + logs/phase4_seed.patch (seed 42 shim)cmpatino/direct-opd-sft-deepmath-pilot-data @ 22625ae5db434947195bf862c429cd94504a4809 :: opd_train.parquet (6,400 AIME-decontaminated prompts, one pass)pi_pre's max_position_embeddings (4096): the reward must never score a position the
pre-teacher was never trained ononly_stu, student_p weighting, T=1.0 for student and teachers,
reward_model.model.input_tokenizer=null (teachers score the student's rendered ids verbatim)logs/run_manifest.json, console log logs/train.log.gztest_freq=-1); all evaluation is external and pre-registered.RLHFDataset, rl_dataset.py:363), so Qwen2.5's default system prompt is present in
both — consistent by construction.<|im_end|> is affected
(~1 token/response). Reported, not fixed.