Downloads · 30 days
544
60% of all-time downloads
code-critic-model/Qwen3-4B-Critic-SFT-DPO
Qwen3-4B-Critic-SFT-DPO is a text generation model from code-critic-model. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
The SFT + DPO critic from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It starts from Qwen3-4B-Critic-SFT and is further trained with direct preference optimization on pairs of the SFT criti…
Downloads · 30 days
544
60% of all-time downloads
All-time downloads
905
Public
Parameters
4B
8.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8 GB · 100%
From the Hugging Face model README
The SFT + DPO critic from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It starts from Qwen3-4B-Critic-SFT and is further trained with direct preference optimization on pairs of the SFT critic's own critiques. It is the strongest 4B critic in the paper and the one reported in the + SFT + DPO rows of Table 1.
A critic sits next to a frozen coding agent. Every k agent steps it reads the trajectory so far and returns a short structured critique: which error categories it detects, the evidence, a recovery action, the task status, and one line of overall guidance. It steers the agent; it does not write the patch.
All released models and datasets are listed on the organization page. Code and configs are in the critic-training repository.
| Paper location | Row label |
|---|---|
| Table 1, every agent block | Qwen3-4B + SFT + DPO |
| Table 8, significance tests | the SFT + DPO critic |
| Section 2.4 and Figure 3 | the DPO training pipeline |
Original checkpoint name: Qwen3-4B-SFT-DPO-4B-1409i-beta0.15-sft0.3-lr1e-6-bs32-ep3-step-80. The old name still redirects here.
Preference pairs were built as described in Section 2.4 of the paper. The coding agent runs on training tasks; every k steps the SFT critic samples N=10 critiques for the current trajectory prefix; Claude Opus 4.6 acts as judge and picks the best and the worst critique by the correctness and clarity of their overall guidance. The best becomes chosen, the worst rejected. This checkpoint was trained on 1,409 such pairs, split 90/10 into train and evaluation.
The 1,409 pairs are not part of this release yet. The dataset code-critic-model/PRM_1541i is an earlier pair set built with the same procedure; it was used for development runs and is not the set behind this checkpoint.
DPO with TRL, initialized from Qwen3-4B-Critic-SFT.
| Setting | Value |
|---|---|
| Initialization | code-critic-model/Qwen3-4B-Critic-SFT |
| Objective | DPO with an added SFT term on the chosen response, weight 0.3 |
| beta | 0.15 |
| Learning rate | 1e-6 |
| Effective batch size | 32 |
| Schedule | 3 epochs planned (120 steps); this checkpoint is step 80, the end of epoch 2 |
| Precision | bf16 |
At step 80 the held-out preference accuracy was 0.68. The step-120 checkpoint is kept in the organization for reference but was not selected for the paper.
Resolve rate on SWE-bench Verified (500 instances), from Table 1 of the paper, best of k=5 and k=10 per configuration.
| Coding agent | No critic | + Qwen3-4B-Critic-SFT | + Qwen3-4B-Critic-SFT-DPO |
|---|---|---|---|
| Qwen3-32B | 8.8 | 11.4 | 14.4 |
| Qwen3-Next-80B-A3B | 20.0 | 24.2 | 26.2 |
| GPT-OSS-20B | 3.0 | 9.8 | 14.8 |
| GLM-4.7-Flash-30B-A3B | 21.6 | 35.2 | 35.8 |
| GPT-OSS-120B (medium reasoning) | 20.4 | 31.2 | 34.8 |
| o3-mini | 19.0 | 27.6 | 28.2 |
DPO improves over SFT for all six agents. On Qwen3-32B, GPT-OSS-20B, and GPT-OSS-120B the 4B DPO critic also beats the 8B SFT critic.
Serve with vLLM in bf16 and run an agent through the repository's mini-swe-agent fork, which inserts a critique every k steps. The --prm name goes to LiteLLM, which needs a matching entry in mini-swe-agent/configs/litellm_model_registry.json to price the calls; copy one of the existing critic blocks to a new key Qwen3-4B-Critic-SFT-DPO. Without an entry the critic call fails and the agent runs without critiques.
vllm serve code-critic-model/Qwen3-4B-Critic-SFT-DPO \
--served-model-name Qwen3-4B-Critic-SFT-DPO \
--dtype bfloat16 --max-model-len 65536 --port 8071
bash scripts/run_critic_max150.sh prm_issue_res_instructions_step_aware 5 0 qwen3-80b \
--prm Qwen3-4B-Critic-SFT-DPO --prm-node <vllm-host>:8071 --slice :500 \
--prefix-dir <path to the matching no-critic run>
To call the critic directly, take any record from critic-sft-cwm-qwen, drop its final teacher critique, and generate. A complete snippet is on the Qwen3-8B-Critic-SFT card; only the repo name changes.
@misc{gandhi2026steerdontsolvetraining,
title={Steer, Don't Solve: Training Small Critic Models for Large Code Agents},
author={Shubham Gandhi and Yiqing Xie and Atharva Naik and Ruichen Zhu and Carolyn Rose},
year={2026},
eprint={2606.21811},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2606.21811}
}