Downloads · 30 days
0
sudhir2016/GRPO7
GRPO7 is a machine learning model from sudhir2016. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
- Developed by: sudhir2016 - License: apache-2.0 - Finetuned from model : unsloth/qwen2.5-3b-instruct-unsloth-bnb-4bit
Downloads · 30 days
0
Access
Public
Updated Mar 16, 2025
Repo size
490 MB
Likes
0
Public
Click a slice to open those files.
.safetensors479 MB · 97%
From the Hugging Face model README
This qwen2 model was trained 2x faster with Unsloth and Huggingface's TRL library.