Skip to content

LLMAligned

gbmpo_gsm8k_model

LLMAligned/gbmpo_gsm8k_model

gbmpo_gsm8k_model is a reinforcement learning model from LLMAligned. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.

This model was fine-tuned using GBMPO (Group Bound Multi-Objective Policy Optimization) on the GSM8K dataset.

Downloads · 30 days

0

Access

Public

Updated Jan 21, 2026

Repo size

28.9 MB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors17.5 MB · 52%

At a glance

Task
Reinforcement Learning
License
apache-2.0
Access
Public
Created
Jan 21, 2026
Updated
Jan 21, 2026
SHA
052b6fd4

Base models

Task
Reinforcement Learning
License
apache-2.0
Created
Jan 21, 2026
Updated
Jan 21, 2026