Skip to content

ItsMaxNorm

DeepSeek-R1-Distill-SmolLM3-3B-GRPO

ItsMaxNorm/DeepSeek-R1-Distill-SmolLM3-3B-GRPO

DeepSeek-R1-Distill-SmolLM3-3B-GRPO is a text generation model from ItsMaxNorm. Use it when you need the model to write or continue text. It is set up for transformers.

This model is a fine-tuned version of HuggingFaceTB/SmolLM3-3B on the open-r1/OpenR1-Math-220k dataset. It has been trained using TRL.

Downloads · 30 days

25

19% of all-time downloads

All-time downloads

133

Public

Repo size

36.9 GB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.pt36.9 GB · 100%

At a glance

Task
Text Generation
Library
transformers
Model type
smollm3
Access
Public
Created
Aug 8, 2025
Updated
Aug 8, 2025
SHA
2a132414

Try a prompt

Base models

Task
Text Generation
Library
transformers
Type
smollm3
Created
Aug 8, 2025
Updated
Aug 8, 2025
DeepSeek-R1-Distill-SmolLM3-3B-GRPO — AI Model — AIMarketly