Downloads · 30 days
4
4% of all-time downloads
sameersegal/Qwen3-0.6B-Reverse-Text-SFT-RLFT
Qwen3-0.6B-Reverse-Text-SFT-RLFT is a machine learning model from sameersegal. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Simple model that was RL FT for 20 steps / epochs after SFT to reverse text using prime-rl (RL Training) and reverse-text (RL Environment). See the improvement in results:
Downloads · 30 days
4
4% of all-time downloads
All-time downloads
99
Public
Repo size
1.5 GB
Likes
0
Public
Click a slice to open those files.
.bin1.5 GB · 99%
From the Hugging Face model README
Simple model that was RL FT for 20 steps / epochs after SFT to reverse text using prime-rl (RL Training) and reverse-text (RL Environment). See the improvement in results:
The reward (correctness score) distribution has improved for the RLFT model across all rollouts.

At an instance level, if we compare the best scores across rollouts, we see a mean improvement of 3.73%. But a maximum of ~30% and reduction of ~3%

Task: reverse-text
Prompt:
<reversed_text> tags.”Expected Completion:
<reversed_text>
.ti otni degrem saw kcuBr ni ytinummoc ehT
</reversed_text>
Expected Reward: 0.963855421686747
Note: Reward is basd on the long common subsequence