Downloads · 30 days
0
a313351012/GRPO_3
GRPO_3 is a machine learning model from a313351012. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
- Developed by: a313351012 - License: apache-2.0 - Finetuned from model : unsloth/qwen2.5-7b-unsloth-bnb-4bit
Downloads · 30 days
0
Access
Public
Updated Jun 12, 2025
Repo size
334 MB
Likes
0
Public
Click a slice to open those files.
.safetensors162 MB · 91%
From the Hugging Face model README
This qwen2 model was trained 2x faster with Unsloth and Huggingface's TRL library.