Downloads · 30 days
4
6% of all-time downloads
JOY0021/autonomy-grpo-agent-v2
autonomy-grpo-agent-v2 is a reinforcement learning model from JOY0021. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as mit.
This model is a Calibrated Epistemic Agent trained specifically for the OpenEnv India Hackathon 2026. It was fine-tuned using Group Relative Policy Optimization (GRPO) to master the balance between autonomous action a…
Downloads · 30 days
4
6% of all-time downloads
All-time downloads
65
Public
Repo size
15.8 MB
Likes
0
Public
Click a slice to open those files.
.json11.4 MB · 84%
From the Hugging Face model README
This model is a Calibrated Epistemic Agent trained specifically for the OpenEnv India Hackathon 2026. It was fine-tuned using Group Relative Policy Optimization (GRPO) to master the balance between autonomous action and information gathering.
Unlike typical LLMs that "hallucinate" or guess when faced with ambiguous instructions, this agent has been trained to use the INVESTIGATE action when it detects uncertainty.
The agent was trained on high-ambiguity scenarios across three domains: Email Triage, DevOps Incidents, and Financial Requests.
| Benchmark | Blind Baseline | Calibrated Agent (Ours) | Improvement |
|---|---|---|---|
| Email Triage | 0.378 | 0.798 | +42.0% |
| DevOps Incident | 0.572 | 0.939 | +36.7% |
| Financial Request | 0.773 | 0.990 | +21.7% |
The model demonstrates an Investigation Rate of 100% on ambiguous signals, effectively resolving partial observability before committing to high-stakes decisions.
This model is designed to be used in conjunction with the Autonomy Calibration Benchmark.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model = "Qwen/Qwen2.5-0.5B-Instruct"
adapter = "JOY0021/autonomy-grpo-agent-v2"
tokenizer = AutoTokenizer.from_pretrained(base_model)
model = AutoModelForCausalLM.from_pretrained(base_model)
model = PeftModel.from_pretrained(model, adapter)