Downloads · 30 days
0
MohammadRafiML/Qwen3-4B-Instruct-2507-Capstone-MathRL
Qwen3-4B-Instruct-2507-Capstone-MathRL is a reinforcement learning model from MohammadRafiML. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for peft.
Fine-tuned from Qwen/Qwen3-4B-Instruct-2507 using a two-stage SFT → GRPO pipeline for mathematical reasoning with calculator tool use.
Downloads · 30 days
0
Access
Public
Updated Apr 16, 2026
Repo size
568 MB
Likes
0
Public
Click a slice to open those files.
.safetensors568 MB · 100%
From the Hugging Face model README
Fine-tuned from Qwen/Qwen3-4B-Instruct-2507 using a two-stage SFT → GRPO pipeline for mathematical reasoning with calculator tool use.
Author: Mohammad Rafi
Qwen/Qwen3-4B-Instruct-2507sft_adapter/| Parameter | Value |
|---|---|
| Method | LoRA (Supervised Fine-Tuning) |
| LoRA rank | 32 |
| Epochs | 2 |
| Training samples | 500 |
| Task | Math reasoning (GSM8K + NuminaMath) |
| Size | 270.92 MB |
grpo_adapter/| Parameter | Value |
|---|---|
| Method | GRPO (Group Relative Policy Optimization) |
| Training samples | 400 |
| Group size | 8 |
| Learning rate | 3e-6 |
| Substeps | 1 |
| Curriculum | easy → intermediate → hard |
| Size | 270.92 MB |
Recommended: Use
grpo_adapter/— trained through the full SFT + GRPO pipeline.
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B-Instruct-2507")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B-Instruct-2507")
# Load GRPO adapter (recommended)
model = PeftModel.from_pretrained(base, "MohammadRafiML/Qwen3-4B-Instruct-2507-Capstone-MathRL", subfolder="grpo_adapter")
model = model.merge_and_unload()
# Load SFT adapter only
# model = PeftModel.from_pretrained(base, "MohammadRafiML/Qwen3-4B-Instruct-2507-Capstone-MathRL", subfolder="sft_adapter")
# model = model.merge_and_unload()