Downloads · 30 days
7
21% of all-time downloads
SeongryongJung/Qwen-4b-base-RLSD
Qwen-4b-base-RLSD is a text generation model from SeongryongJung. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
RLSD self-distillation reinforcement learning on the local math training split.
Downloads · 30 days
7
21% of all-time downloads
All-time downloads
34
Public
Parameters
4.4B
8.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.8 GB · 100%
From the Hugging Face model README
RLSD self-distillation reinforcement learning on the local math training split.
This repository contains the final merged Hugging Face checkpoint from global_step_100.
The training checkpoint was saved from FSDP shards and merged to safetensors for this upload.
rlsd.grpo in the trainer config.compute_score reward manager.| Field | Value |
|---|---|
| Base model | Qwen/Qwen3-4B-Base |
| Train file | /home1/irteam/SDPO/self-distillation-analysis/data/math/train.parquet |
| Validation file | /home1/irteam/SDPO/self-distillation-analysis/data/math/evaluation/aime24.parquet |
| Train max samples | 25600 |
| Train batch size | 256 |
| Rollouts per prompt | 8 |
| PPO mini batch size | 128 |
| PPO micro batch size per GPU | 1 |
| Optimizer | AdamW |
| Learning rate | 1e-06 |
| Weight decay | 0.01 |
| LR warmup steps | 10 |
| Total training steps | 100 |
| Save frequency | every 10 steps |
| Validation frequency | every 10 steps |
| Max prompt length | 2048 |
| Max response length | 20480 |
| Rollout backend | vllm |
| Rollout temperature | 1 |
| Rollout top_p | 1 |
| vLLM GPU memory utilization | 0.75 |
| Actor strategy | fsdp |
| Dtype | bfloat16 |
| Advantage estimator | grpo |
| Gamma / Lambda | 1 / 1 |
| KL loss enabled | False |
| KL loss coefficient | 0.001 |
| Checkpoint uploaded | math-RLSD-Qwen3-4B-Base-128-train256-rollout8-lr1e-6-vllm0.75-modelQwen-Qwen3-4B-Base/global_step_100 |
| W&B run id | 3tuehy90 |
The plot below shows critic/score/mean logged during training.

CSV data is included in training_score.csv.
| Metric | Value |
|---|---|
| Final training step | 100 |
Final critic/score/mean | 0.304199 |
Final val-core/math_dapo/acc/mean@1 | 0.1 |
This model is intended for internal research and analysis of math-focused RL fine-tuning methods. It has not been broadly safety evaluated for production use.
The model was trained for 100 optimization steps on a local math dataset split. Reported scores are training-time reward/validation metrics from the same experiment setup and should not be treated as broad benchmark results.