Downloads · 30 days
0
gaaaaaaaaaaa/GRWO
GRWO is a machine learning model from gaaaaaaaaaaa. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Guided Reasoning Window Optimization (GRWO) is a training strategy for improving reasoning models by optimizing local reasoning windows instead of full long-form solution traces.
Downloads · 30 days
0
Access
Public
Updated May 23, 2026
Repo size
977 MB
Likes
0
Public
Click a slice to open those files.
.pt489 MB · 48%
From the Hugging Face model README
Guided Reasoning Window Optimization (GRWO) is a training strategy for improving reasoning models by optimizing local reasoning windows instead of full long-form solution traces.
The core idea is simple:
Given a problem and a partial reasoning prefix, train the model to prefer a better next reasoning window over a weaker next reasoning window.
Instead of forcing a model to generate and train on a full 5k–10k token solution, GRWO focuses on the next N tokens of reasoning, usually around a meaningful branch point.
This makes reasoning training cheaper, more targeted, and more aligned with how models actually fail.
Long reasoning traces are expensive.
A hard math problem may require thousands of tokens of exploration, correction, theorem selection, calculation, and verification. Training directly on full traces has several problems:
Most reasoning failures happen at local branch points:
GRWO targets these local branch points directly.
Train reasoning models by correcting local self-generated reasoning windows instead of supervising full solution traces.
A GRWO sample contains:
problem + partial reasoning prefix
Then the model generates two possible continuations:
rejected = unguided continuation
chosen = guided or corrected continuation
Only the next local reasoning window is optimized.
prompt = problem + partial reasoning prefix
chosen = better next reasoning window
rejected = weaker next reasoning window
The key constraint:
The keypoints or guided answer are used only to create the chosen continuation. They are not exposed in the student model's prompt.
Instead of this:
Question -> full 10k-token reasoning trace -> answer
GRWO trains this:
Question + reasoning prefix -> next 512–1500 reasoning tokens
This changes the objective from:
Solve the entire problem from scratch.
to:
Given this current reasoning state, choose the better next reasoning move.
That is the behavior weak reasoning models need most.
Recommended method name:
Guided Reasoning Window Optimization (GRWO)
Variants:
GRWO-DPO = DPO over local reasoning windows
GRWO-SFT = SFT on corrected local reasoning windows
GRWO-GRPO = RL/GRPO over local reasoning windows
Primary direction:
GRWO-DPO: preference training over local self-generated reasoning continuations.
{
"prompt": "Problem statement...\n<think>\nPartial model-generated reasoning prefix...",
"chosen": "Guided/correct next reasoning window...",
"rejected": "Unguided/weaker next reasoning window..."
}
The prompt contains the problem and the partial reasoning prefix.
The chosen and rejected completions should continue from the exact same prefix.
{
"prompt": "Problem statement...\n<think>\nPartial model-generated reasoning prefix...",
"completion": "Correct next reasoning window..."
}
SFT is optional, but useful when the model lacks a reasoning move entirely.
GRWO-DPO is the preferred main method when the base model already has decent reasoning ability and follows the desired format.
It teaches:
GRWO-DPO preserves the model's natural language style better than pure SFT because it does not force exact imitation.
GRWO-SFT is useful when the model does not know how to produce a required reasoning move.
It teaches:
SFT should be used carefully because it can overwrite the model's native style if the teacher traces are too polished or unnatural.
GRWO-GRPO is a later-stage method.
It can be used when DPO/SFT are not enough and you want online sampling with a reward function.
Possible reward components:
+ correct final answer, if available
+ correct theorem/test selection
+ valid transformation
+ keypoint coverage
+ forward progress
+ calibrated confidence
- invalid conclusion
- theorem misuse
- wandering
- fake final answer
- malformed protocol tags
Start with GRWO-DPO-first if the model already follows the format and can produce plausible reasoning.
1. Select a hard problem.
2. Let the model generate a partial reasoning prefix.
3. Stop around a meaningful branch point.
4. Generate an unguided continuation for the next N tokens.
5. Generate a guided/corrected continuation using keypoints, verifier, or teacher.
6. Build a DPO pair:
prompt = problem + partial reasoning prefix
chosen = guided continuation
rejected = unguided continuation
7. Train with DPO on only this local continuation window.
Optional:
8. Add a small SFT bucket for cases where the model cannot produce the desired move at all.
A good starting range:
prefix length: 500–1500 tokens
local window length: 768–1500 tokens
A practical default:
prefix length: ~1000 tokens
continuation window: ~1500 tokens
The model does not automatically learn to stop at 1500 tokens just because generation is capped there. The cap is an external data-generation/training budget.
However, avoid making accepted samples look like naturally finished answers if they are actually truncated.
For process-window training, it is acceptable for the chosen continuation not to finish the whole problem. The objective is local reasoning continuation, not full solution completion.
The loss should only apply to the local continuation window.
problem tokens: no loss
partial reasoning prefix: no loss
next N local window tokens: loss active
future tokens after window: not included / no loss
For SFT, this means:
labels for prompt/prefix tokens = -100
labels for continuation tokens = token IDs
For DPO, the row should be structured so that:
prompt = prefix
chosen = continuation window only
rejected = continuation window only
Do not include future tokens after the supervised window, because that leaks information.
The chosen and rejected continuations should be as similar as possible except for reasoning quality.
Good pair design:
same prompt
same prefix
similar length
similar format
similar style
similar confidence level visually
only reasoning direction differs
Bad pair design:
chosen = polished teacher solution
rejected = messy model rambling
That would teach surface style instead of reasoning direction.
Better pair:
chosen = model-like continuation, corrected at the branch
rejected = model-like continuation, wrong branch
Problem type: series convergence.
Partial reasoning prefix:
The absolute series behaves like the harmonic series, so it does not converge absolutely.
Now I need to determine whether the original alternating series converges conditionally.
The alternating series test is awkward because monotonicity is not obvious...
Rejected continuation:
Since the alternating series test fails, the series diverges.
Chosen continuation:
But failure of the alternating series test would only be inconclusive, not proof of divergence.
I should try another method. A useful approach is to rewrite the term as an alternating harmonic component plus an absolutely convergent correction...
This pair teaches the model:
A failed test is inconclusive, not proof of divergence.
That is a local reasoning correction.
Good GRWO windows should target common reasoning branch errors.
Recommended first 1000-sample mix:
700 GRWO-DPO branch-point pairs
150 GRWO-DPO finish-vs-wander pairs
100 weak-correct vs strong-correct pairs
50 optional GRWO-SFT missing-move examples
Alternative conservative mix:
80–90% GRWO-DPO
10–20% GRWO-SFT
Use SFT only when the model cannot produce the desired reasoning move at all.
Do not automatically reject all unguided continuations.
If the unguided continuation is correct, either:
keep it as a positive SFT sample
use it as chosen against a weaker continuation
or skip it
Do not train the model to believe all of its natural reasoning is bad.
The rule:
unguided correct -> positive or skip
unguided wrong -> rejected against guided correction
unguided wandering -> rejected against progress continuation
DPO can teach confidence in language even when both samples are technically correct.
Example rejected:
This probably converges because it looks alternating.
Example chosen:
The absolute series behaves like the harmonic series, so it is not absolutely convergent.
The remaining signed series can be rewritten as an alternating harmonic term plus an absolutely convergent correction, so it converges conditionally.
Both may reach the same answer, but the chosen one teaches:
The goal is not confidence alone.
The goal is:
calibrated confidence
The model should be decisive when the logic supports it and cautious when a test is inconclusive.
For GRWO-DPO on a small high-quality dataset:
learning_rate: 5e-6 to 1e-5
beta: 0.03 to 0.05
num_train_epochs: 1 to 2
prompts_per_step: 5 to 10
dpo_epochs_per_step: 3 to 4
max_length: 3072 to 4096
max_prompt_length: 1024 to 1536
max_completion_length: 1536 to 3072
per_device_train_batch_size: 1
gradient_accumulation_steps: 4 to 8
save_steps: 25
logging_steps: 1
For GRWO-SFT, if used:
learning_rate: 5e-5 to 1e-4
num_train_epochs: 1 to 2
max_length: 3072 to 4096
DPO should usually be softer than SFT because it can over-steer quickly.
Evaluate GRWO by behavior, not just loss.
Test whether the model:
Useful benchmark prompts:
1. Series trap where AST failure is inconclusive.
2. Ratio/root test equals 1 trap.
3. Double integral with coordinate-bound trap.
4. Spherical-coordinate Jacobian trap.
5. Problem where early answer is tempting but wrong.
6. Problem where reasoning should finish instead of wander.
GRWO should reduce training/data cost because it avoids full-trace generation.
Approximate reduction:
full trace: 10,000 tokens
GRWO window: 1,500 tokens
cost reduction: ~85% target-token reduction
It also produces more training examples per hard problem.
One hard problem can become several local windows:
window 1: prefix 0–600 -> train 600–1500
window 2: prefix 0–1300 -> train 1300–2500
window 3: prefix 0–2200 -> train 2200–3500
finish: near-end prefix -> train final closeout
Do not over-sample one problem too much. Use roughly 2–4 windows per problem.
DPO may learn cheap differences if chosen and rejected differ too much in formatting, length, or tone.
Mitigation:
make chosen/rejected similar in style and length
If keypoints are visible in the student prompt, the model may depend on information that will not exist at inference.
Mitigation:
keypoints only guide teacher/chosen generation
not student prompt
If only mid-reasoning windows are trained, the model may improve continuation but not termination.
Mitigation:
include 10–20% finish-window samples
DPO can accidentally reward assertive but unsupported language.
Mitigation:
reward calibrated confidence, not raw confidence
The model may learn local moves but still fail globally.
Mitigation:
include multi-window coverage and final-answer evaluation
for problem in problems:
prefix = generate_partial_reasoning(
model=model,
problem=problem,
max_new_tokens=prefix_budget,
stop_at_branch=True,
)
rejected = generate_continuation(
model=model,
prompt=problem + prefix,
max_new_tokens=window_budget,
guided=False,
)
score = evaluate_continuation(
problem=problem,
prefix=prefix,
continuation=rejected,
answer=gold_answer,
keypoints=keypoints,
)
if score.is_good:
keep_as_positive_or_skip(problem, prefix, rejected)
continue
chosen = teacher_generate_guided_continuation(
problem=problem,
prefix=prefix,
rejected=rejected,
gold_answer=gold_answer,
keypoints=keypoints,
max_new_tokens=window_budget,
)
add_dpo_row(
prompt=problem + prefix,
chosen=chosen,
rejected=rejected,
)
Start small.
problems: 200
windows per problem: 1–2
DPO pairs: 300–400
window length: 768–1500 tokens
prefix length: 500–1500 tokens
Train:
Experiment A: GRWO-DPO only
Experiment B: small GRWO-SFT + GRWO-DPO
Compare:
base/SFT model vs GRWO-DPO vs GRWO-SFT+DPO
Main question:
Does the model choose better reasoning branches under the same token budget?
GRWO is a process-preference training method where a reasoning model is optimized on local continuation windows from its own partial traces. A guided continuation is preferred over an unguided continuation under the same prefix, with loss applied only to the local window. The goal is to improve reasoning trajectory decisions while preserving the model's native reasoning style and reducing full-trace training cost.
GRWO is promising because it combines three useful ideas:
This is especially suitable for math reasoning models where errors often occur at identifiable branch points.
The most important design principle:
Train the next reasoning move, not the entire solution.