Downloads · 30 days
11
6% of all-time downloads
tanliboy/lambda-llama-3-8b-dpo-test
lambda-llama-3-8b-dpo-test is a text generation model from tanliboy. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.1.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
11
6% of all-time downloads
All-time downloads
173
Public
Parameters
8B
16.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of meta-llama/Meta-Llama-3.1-8B-Instruct on the HuggingFaceH4/ultrafeedback_binarized dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Rewards/chosen | Rewards/rejected | Rewards/accuracies | Rewards/margins | Logps/rejected | Logps/chosen | Logits/rejected | Logits/chosen |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0.6351 | 0.2093 | 100 | 0.6359 | -0.6754 | -0.9179 | 0.6746 | 0.2426 | -487.7697 | -472.2982 | -2.4565 | -2.2922 |
| 0.6101 | 0.4186 | 200 | 0.5990 | -0.7996 | -1.1966 | 0.7143 | 0.3970 | -515.6393 | -484.7244 | -2.4477 | -2.2933 |
| 0.5738 | 0.6279 | 300 | 0.5819 | -1.0722 | -1.6607 | 0.7143 | 0.5885 | -562.0454 | -511.9821 | -2.5003 | -2.3506 |
| 0.5808 | 0.8373 | 400 | 0.5776 | -1.0426 | -1.6196 | 0.7063 | 0.5769 | -557.9310 | -509.0269 | -2.6060 | -2.4454 |