Skip to content

Kibalama

GRPO

Based on HuggingFaceTB/SmolLM-135M-Instruct

Tags

  • safetensors
  • generated_from_trainer
  • trl
  • grpo
  • arxiv:2402.03300
  • base_model:HuggingFaceTB/SmolLM-135M-Instruct
  • base_model:finetune:HuggingFaceTB/SmolLM-135M-Instruct
  • endpoints_compatible
  • region:us
Author
Kibalama
Library
transformers
Base model
HuggingFaceTB/SmolLM-135M-Instruct
Downloads
0
Likes
0
GRPO — AI Model — AIMarketly