Downloads · 30 days
0
juspay/xyne-rl
xyne-rl is a machine learning model from juspay. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as other.
Merged Gemma 4 31B checkpoints from the Xyne near-online GRPO/RLAIF run.
Downloads · 30 days
0
Access
Public
Updated May 27, 2026
Repo size
64 GB
Likes
0
Public
Click a slice to open those files.
.safetensors190 GB · 100%
From the Hugging Face model README
Merged Gemma 4 31B checkpoints from the Xyne near-online GRPO/RLAIF run.
Full merged checkpoints:
Adapter-only artifacts available locally:
The exact adapter-only deltas for checkpoint 304 and checkpoint 328 were not retained locally after merge/pruning, so they are not uploaded as adapters. Their full merged checkpoints are available above.
Standalone Xyne v2 eval on 31 held-out questions, 1 rollout per question, using the same adaptive 3 -> 7 judge path as training:
| Model | All-in mean | Clean judged mean | Median | Valid judged | Judge dropped | Notes |
|---|---|---|---|---|---|---|
| Base Gemma 4 31B | 0.5284 | ~0.546 | 0.50 | 31/31 rollouts | 1 | one judge parse dropout counted as zero in all-in score |
| checkpoint-200 | 0.5806 | ~0.621 | 0.60 | 31/31 rollouts | 2 | two judge API timeout dropouts counted as zero in all-in score |
| checkpoint-304 | 0.6097 | 0.610 | 0.60 | 31/31 rollouts | 0 | best all-in operational score among evaluated checkpoints |
Interpretation:
These are full merged checkpoints unless under v1/adapters/.