Downloads · 30 days
6
2% of all-time downloads
AlgoCore/support-ticket-grpo-model
support-ticket-grpo-model is a machine learning model from AlgoCore. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
Fine-tuned Qwen/Qwen2.5-0.5B-Instruct using GRPO (Group Relative Policy Optimization) + LoRA on a multi-step support ticket environment.
Downloads · 30 days
6
2% of all-time downloads
All-time downloads
375
Public
Parameters
503M
1.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.1 GB · 99%
How the weights are stored.
F16494M · 98%
From the Hugging Face model README
Fine-tuned Qwen/Qwen2.5-0.5B-Instruct using GRPO (Group Relative Policy Optimization) + LoRA on a multi-step support ticket environment.
trl.GRPOTrainer + LoRA (PEFT)| Task | Before | After | Delta |
|---|---|---|---|
| Task 1 (Classify) | 0.667 | 1.000 | +0.333 |
| Task 2 (Action) | 0.117 | 0.450 | +0.333 |
| Task 3 (Full Resolve) | 0.083 | 0.258 | +0.175 |
| Overall | 0.289 | 0.569 | +0.280 |
