Downloads · 30 days
0
ada-flo/monkey-grpo-arith_op
monkey-grpo-arith_op is a machine learning model from ada-flo. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
GRPO outcome-RL LoRA adapters (one per rule × init condition), trained on top of a CPT initialization.
Downloads · 30 days
0
Access
Public
Updated May 19, 2026
Repo size
1.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors1.7 GB · 99%
From the Hugging Face model README
GRPO outcome-RL LoRA adapters (one per rule × init condition), trained on top of a CPT initialization.
From the project Tell or Show: How Training-Data Format Shapes Implicit vs. Explicit Rule Knowledge.
Adapters are organized as <base-model>/<rule>__from_<init>/:
.
└── qwen3-4b-instruct-2507/ # base = Qwen/Qwen3-4B-Instruct-2507
├── max_plus_diff__from_fewshot/
├── max_plus_diff__from_explicit/
├── digit_product__from_fewshot/
└── ... (11 rules × 2 inits = up to 22)
Each leaf subdir is a self-contained PEFT-loadable adapter:
adapter_config.jsonadapter_model.safetensorsREADME.md (per-variant details)trainer_state.json (training-time metrics)Future base models (Qwen3-7B etc.) will appear as sibling base-model dirs.
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B-Instruct-2507", torch_dtype="bfloat16")
model = PeftModel.from_pretrained(base, "ada-flo/monkey-grpo-arith_op", subfolder="qwen3-4b-instruct-2507/max_plus_diff__from_fewshot")