Skip to content

ipfipfipf

Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29

ipfipfipf/Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29

Qwen3.5-9B-sdpo-react-mathcodesearch-grpo-arm-e-step29 is a text generation model from ipfipfipf. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.

RL post-training of Qwen/Qwen3.5-9B on a multi-turn, native-tool-calling (ReAct-style) mixture of math + code + search tasks, using GRPO with an SDPO self-skill objective ("arm e").

Downloads · 30 days

35

13% of all-time downloads

All-time downloads

263

Public

Parameters

9B

17.9 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors17.9 GB · 100%

Parameter types

How the weights are stored.

BF169B · 100%

Try a prompt

Base models

Task
Text Generation
Library
transformers
Type
qwen3_5
License
apache-2.0
Created
Sep 1, 2026
Updated
Sep 1, 2026