Downloads · 30 days
35
13% of all-time downloads
ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29
Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29 is a text generation model from ipfipfipf. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
RL post-training of Qwen/Qwen3.5-9B on a multi-turn, native-tool-calling (ReAct-style) mixture of math + code + search tasks, using GRPO with an SDPO self-skill objective ("arm e").
Downloads · 30 days
35
13% of all-time downloads
All-time downloads
263
Public
Parameters
9B
17.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors17.9 GB · 100%
How the weights are stored.
BF169B · 100%
From the Hugging Face model README
RL post-training of Qwen/Qwen3.5-9B on a multi-turn, native-tool-calling
(ReAct-style) mixture of math + code + search tasks, using GRPO with an
SDPO self-skill objective ("arm e").
This is the step-29 checkpoint, which is the peak of the training curve — see Why step 29 below.
[!IMPORTANT] Language model only. The conversion from the training checkpoint exports the 427 language-model tensors and drops the base model's vision tower (333 tensors) and MTP heads (15 tensors), because only the language model was trained.
config.jsonstill declaresQwen3_5ForConditionalGenerationfor compatibility with the base tokenizer/config, but image input will not work. Text generation loads and serves normally under Transformers / vLLM / SGLang.
| Base | Qwen/Qwen3.5-9B |
| Algorithm | GRPO + SDPO (self-skill-all, skill-KD mode=both, KD coef 0.01, --no-sdpo-pure-distill, --sdpo-response-prefix skill) |
| Domains | math, code, search — one mixed multitask stream |
| Rollouts | multi-turn native tool calling, up to 20 turns; thinking enabled |
| Max response length | 16384 tokens (train and eval matched) |
| Steps | 31 rollouts trained; this checkpoint is step 29 |
pass@1 / pass@8, greedy-free sampling, 16384-token response cap:
| Benchmark | pass@1 | pass@8 |
|---|---|---|
| AIME 2024 | 84.2 | 96.7 |
| AIME 2025 | 81.7 | 100.0 |
| AMO-Bench | 28.0 | 52.0 |
| OJBench (medium, 77 problems) | 31.8 | 59.7 |
AMO-Bench and OJBench are the best numbers we have on record for a 9B model in this line of work (previous best: AMO 22.7, OJBench 29.5).
Caveats worth knowing before you compare against these:
Two independent runs of this arm agree to within 0.6pp at every shared eval step, and both peak at step 29. Aggregate held-out pass@1 by step:
| step | 0 | 9 | 19 | 29 | 39 | 49 |
|---|---|---|---|---|---|---|
| 51-rollout run | 58.1 | 67.3 | 71.0 | 73.6 | 71.2 | 70.0 |
| 31-rollout run (this one) | 58.2 | 68.4 | 71.6 | 73.0 | — | — |
Past step 29 performance decays monotonically, led by code (LiveCodeBench-v6-functional 63.3 → 59.1) with repetition rate climbing 0 → 0.6 → 4.2%. Training this arm longer is actively worse, so 31 rollouts is the right budget and step 29 is the checkpoint to use.
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="auto", device_map="auto")
The model was trained with thinking enabled and with native tool calling, so
serve it with the shipped chat_template.jinja and pass tools through the
template's tools argument. Under SGLang, use the qwen3_coder tool-call parser
(Qwen3.5 emits XML-style <function=...> calls, not Qwen2.5-style JSON).