Downloads · 30 days
19
100% of all-time downloads
open-athena/Snowball-67B-A2B-Math-RL-E17d-Step8-Repaired
Snowball-67B-A2B-Math-RL-E17d-Step8-Repaired is a reinforcement learning model from open-athena. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as other.
This is the exact checkpoint used for a reported row in the Snowball 67B-A2B math-RL experiment.
Downloads · 30 days
19
100% of all-time downloads
All-time downloads
19
Public
Parameters
67.1B
134 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors134 GB · 100%
How the weights are stored.
BF1667.1B · 100%
From the Hugging Face model README
This is the exact checkpoint used for a reported row in the Snowball 67B-A2B math-RL experiment.
Router state: SFT router-bias repaired control. Closely named older laion/rl-snowball-* repositories may contain raw mutable-router exports that collapse under inference. Do not substitute them for this artifact. Models from this campaign are research checkpoints, not production releases; their utility is limited unless router-bias repair or frozen-router integrity is preserved.
E17d repaired control83.33 / 37.80 / 5.67s3://marin-us-east-02a/marin/exports/snowball-bias-repaired/rl-snowball-e17d-rno2a-rlvrmath-lossfree-sr-20260825-014634/global_step_8/policy/The scores and evaluation caveats are recorded in MATH_EVALS.md in the evidence archive. Preserve config.json, the tokenizer files, and all shards named by model.safetensors.index.json together.