Skip to content

LLMAligned

grpo_gsm8k_model

LLMAligned/grpo_gsm8k_model

grpo_gsm8k_model is a reinforcement learning model from LLMAligned. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.

This model was fine-tuned using GRPO (Group Relative Policy Optimization) on the GSM8K dataset.

Downloads · 30 days

0

Access

Public

Updated Jan 21, 2026

Repo size

1.2 GB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors1.2 GB · 99%

At a glance

Task
Reinforcement Learning
License
apache-2.0
Access
Public
Created
Jan 21, 2026
Updated
Jan 21, 2026
SHA
c3ebce3f

Base models

Task
Reinforcement Learning
License
apache-2.0
Created
Jan 21, 2026
Updated
Jan 21, 2026
grpo_gsm8k_model — AI Model — AIMarketly