Skip to content

ypwang61

One-Shot-RLVR-R1-Distill-1.5B-4-shot

ypwang61/One-Shot-RLVR-R1-Distill-1.5B-4-shot

One-Shot-RLVR-R1-Distill-1.5B-4-shot is a text generation model from ypwang61. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.

This repository contains the model presented in Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Downloads · 30 days

19

3% of all-time downloads

All-time downloads

554

Public

Parameters

1.8B

7.1 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors7.1 GB · 100%

At a glance

Task
Text Generation
Library
transformers
License
apache-2.0
Model type
qwen2
Access
Public
Created
Jun 3, 2025
Updated
Aug 27, 2025
SHA
26f02f01

Try a prompt

Base models

Task
Text Generation
Library
transformers
Type
qwen2
License
apache-2.0
Created
Jun 3, 2025
Updated
Aug 27, 2025