Skip to content

zosmaai

Qwen3.5-0.8B-GRPO-Math

zosmaai/Qwen3.5-0.8B-GRPO-Math

Qwen3.5-0.8B-GRPO-Math is a text generation model from zosmaai. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.

A reasoning-enhanced version of Qwen3.5-0.8B, trained using GRPO (Group Relative Policy Optimization) — the RL technique behind DeepSeek-R1 — on a single RTX 5090 at Zosma AI.

Downloads · 30 days

33

8% of all-time downloads

All-time downloads

416

Public

Parameters

752M

1.5 GB on disk

Likes

2

Trending 1

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors1.5 GB · 99%

At a glance

Task
Text Generation
License
apache-2.0
Model type
qwen3_5_text
Access
Public
Created
Mar 10, 2026
Updated
Mar 10, 2026
SHA
6ff74a6e

Try a prompt

Base models

Task
Text Generation
Type
qwen3_5_text
License
apache-2.0
Languages
en
Created
Mar 10, 2026
Updated
Mar 10, 2026