Downloads · 30 days
0
JingweiNi/iclr-2action-controller-heads
iclr-2action-controller-heads is a machine learning model from JingweiNi. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
20 belief heads that read a frozen Qwen3-1.7B's hidden states (via the UHead JingweiNi/ReProbe-Qwen3-1.7B-DAPO-block-layer16, block layer 16) and predict, at every 4th reasoning block, whether stopping now would still…
Downloads · 30 days
0
Access
Public
Updated Aug 30, 2026
Repo size
5.8 GB
Likes
0
Public
Click a slice to open those files.
.jsonl4.1 GB · 72%
From the Hugging Face model README
20 belief heads that read a frozen Qwen3-1.7B's hidden states (via the UHead
JingweiNi/ReProbe-Qwen3-1.7B-DAPO-block-layer16, block layer 16) and predict, at every
4th reasoning block, whether stopping now would still give the right answer.
Each directory holds s1_best.pt (the checkpoint), summary.json (config, per-epoch
history, provenance commit, val/select metrics) and the val_scored.jsonl /
select_scored.jsonl dumps — the decision-rule ladder replays those dumps, so every
reported table can be recomputed on CPU without a GPU or a model download.
| family | fit set | heads | val q_ap (seed 0 / 1) |
|---|---|---|---|
heads/math_union_row3 | math union (3 918 q) | 2 seeds × {lean, full-11} | full-11 0.808 / 0.797 |
heads/code_row3 | SYNTHETIC-2-RL code pool (731 q) | 2 × {lean, full-11} | full-11 0.645 / 0.669 |
heads/mixed_row3 | math ∪ code (4 809 q) | 2 × {lean, full-11} | full-11 0.771 / 0.746 |
heads/code_folds, heads/mixed_folds | 2-fold splits | 2 seeds × 2 folds | 0.585–0.663 / 0.754–0.757 |
lean = targets q,v,vpi0,r; full-11 adds joint_pi0,ccont_pi0,dy,dc,joint,cterm,ccont.
Every row-3 head is trained against a cross-fitted π₀ fire table built from the fold
heads' out-of-fold predictions (tables in the companion dataset repo).
Headline result: training the head on the target domain does not help. The code-trained heads only match the zero-shot math heads on code, are negative on math, and transfer less to the public code benchmarks; mixing the pools removes the gain of the learned rule over its own base rule without hurting the deployed operating point. The one rule positive on both seeds in every arm is a CUSUM over the belief.
Rollout data: JingweiNi/iclr-2action-controller-chains.
Full record: branch archive/iclr-2action-controller-2026-08-30,
docs/reproducibility/iclr-2action-controller-assets.md.