Downloads · 30 days
8
17% of all-time downloads
annaovesnaatatt/reward-model
reward-model is a feature extraction model from annaovesnaatatt. Use it when you need embeddings to search or compare text. It is set up for transformers.
Reward model for RLHF trained on 3000 examples from Anthropic/hh-rlhf dataset.
Downloads · 30 days
8
17% of all-time downloads
All-time downloads
48
Public
Repo size
996 MB
Likes
0
Public
Click a slice to open those files.
.bin498 MB · 99%
From the Hugging Face model README
Reward model for RLHF trained on 3000 examples from Anthropic/hh-rlhf dataset.