Downloads · 30 days
3
9% of all-time downloads
PahaII/or_v2_rl_v0.1
or_v2_rl_v0.1 is a reinforcement learning model from PahaII. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
RLOO-trained checkpoint from the launch11 / v1 run (experiment name fullparam8nodev10419-1502) at globalstep62, based on the OpenResearcher SFT init (sftqwen3535b/checkpoint-1560-fused).
Downloads · 30 days
3
9% of all-time downloads
All-time downloads
33
Public
Parameters
34.7B
69.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors69.3 GB · 100%
From the Hugging Face model README
RLOO-trained checkpoint from the launch11 / v1 run (experiment name
fullparam_8node_v1_0419-1502) at global_step_62, based on the
OpenResearcher SFT init (sft_qwen35_35b/checkpoint-1560-fused).
midpass1k — pass-rate 0.33–0.67 subset of the RL training setloss_agg_mode=token-mean, clip_ratio_high=0.28,
kl_loss_coef=1e-4See docs/04242026_rl_step62_eval_analysis.md in the repo for the
GAIA-level trajectory audit that motivated the v0.2+ reward-shaping work.
Research checkpoint. Multi-turn deep-research agent with browser tools; requires a compatible search service and the OpenResearcher agent loop for inference-time deployment.