Downloads · 30 days
0
LLMAligned/gbmpo_gsm8k_model
gbmpo_gsm8k_model is a reinforcement learning model from LLMAligned. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
This model was fine-tuned using GBMPO (Group Bound Multi-Objective Policy Optimization) on the GSM8K dataset.
Downloads · 30 days
0
Access
Public
Updated Jan 21, 2026
Repo size
28.9 MB
Likes
0
Public
Click a slice to open those files.
.safetensors17.5 MB · 52%
From the Hugging Face model README
This model was fine-tuned using GBMPO (Group Bound Multi-Objective Policy Optimization) on the GSM8K dataset.
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("LLMAligned/gbmpo_gsm8k_model")
tokenizer = AutoTokenizer.from_pretrained("LLMAligned/gbmpo_gsm8k_model")
prompt = "Solve: What is 25 * 4?"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_length=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))