Downloads · 30 days
9
7% of all-time downloads
Amartya77/RLHF_PPOppo_model
RLHF_PPOppo_model is a reinforcement learning model from Amartya77. Use it for the reinforcement learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
Downloads · 30 days
9
7% of all-time downloads
All-time downloads
125
Public
Parameters
582M
2.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.3 GB · 100%