Skip to content

mjf-su

GRPO-nADEReward-Format

mjf-su/GRPO-nADEReward-Format

GRPO-nADEReward-Format is a image-text-to-text model from mjf-su. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers.

This model is a fine-tuned version of mjf-su/PhysicalAI-reason-VLA. It has been trained using TRL.

Downloads · 30 days

7

16% of all-time downloads

All-time downloads

44

Public

Parameters

4.4B

49.9 GB on disk

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.pt32.2 GB · 55%

At a glance

Task
Image-Text-to-Text
Library
transformers
Model type
qwen3_vl
Access
Public
Created
Apr 5, 2026
Updated
Apr 5, 2026
SHA
8dcef5a9

Try a prompt

Base models

Task
Image-Text-to-Text
Library
transformers
Type
qwen3_vl
Created
Apr 5, 2026
Updated
Apr 5, 2026
GRPO-nADEReward-Format — AI Model — AIMarketly