Skip to content

MohammadRafiML

Qwen3-4B-Instruct-2507-Capstone-MathRL

MohammadRafiML/Qwen3-4B-Instruct-2507-Capstone-MathRL

Qwen3-4B-Instruct-2507-Capstone-MathRL is a reinforcement learning model from MohammadRafiML. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for peft.

Fine-tuned from Qwen/Qwen3-4B-Instruct-2507 using a two-stage SFT → GRPO pipeline for mathematical reasoning with calculator tool use.

Downloads · 30 days

0

Access

Public

Updated Apr 16, 2026

Repo size

568 MB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors568 MB · 100%

At a glance

Task
Reinforcement Learning
Library
peft
Access
Public
Created
Apr 16, 2026
Updated
Apr 16, 2026
SHA
d33f2b61

Base models

Task
Reinforcement Learning
Library
peft
Created
Apr 16, 2026
Updated
Apr 16, 2026
Qwen3-4B-Instruct-2507-Capstone-MathRL — AI Model — AIMarketly