Downloads · 30 days
22
4% of all-time downloads
laion/ablation-pymethods2test-shaped-45-8B
ablation-pymethods2test-shaped-45-8B is a text generation model from laion. Use it when you need the model to write or continue text. It is set up for transformers.
RL (SkyRL GRPO) checkpoint from the shaped-reward ablation of the a3-successor study. Reward = shaped pass-ratio (fraction of tests passing, rewardshaper=passratio), as opposed to the binary all-tests-pass reward of t…
Downloads · 30 days
22
4% of all-time downloads
All-time downloads
530
Public
Parameters
8.2B
16.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.4 GB · 99%
From the Hugging Face model README
RL (SkyRL GRPO) checkpoint from the shaped-reward ablation of the a3-successor
study. Reward = shaped pass-ratio (fraction of tests passing, reward_shaper=pass_ratio),
as opposed to the binary all-tests-pass reward of the a3 series.
global_step_45, selected as the best checkpoint by EMA
(alpha=1/3, trailing-5 window) of reward/avg_raw_reward computed across
the full 80-step training chain (EMA = 0.4712 at step 45).hf_save_interval=5, 14x GH200 nodes on JSC Jupiter.The rl_config.json in this repo is the exact launch config used for reproducibility.
Training-time Daytona/Harbor rollouts for this run are uploaded as a companion dataset: open-athena/ablation-pymethods2test-shaped
The dataset contains the last episode of each trial (per
make_and_upload_trace_dataset --episodes last) — the same rollouts
the policy was trained on after rollback / truncation.