Downloads · 30 days
5
36% of all-time downloads
chchen/Falcon-7B-Instruct-ORPO
Falcon-7B-Instruct-ORPO is a machine learning model from chchen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
5
36% of all-time downloads
All-time downloads
14
Public
Repo size
196 MB
Likes
0
Public
Click a slice to open those files.
.safetensors65.3 MB · 95%
From the Hugging Face model README
This model is a fine-tuned version of tiiuae/falcon-7b-instruct on the dpo_mix_en dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Rewards/chosen | Rewards/rejected | Rewards/accuracies | Rewards/margins | Logps/rejected | Logps/chosen | Logits/rejected | Logits/chosen | Sft Loss | Odds Ratio Loss |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1.6309 | 0.8891 | 500 | 1.5816 | -0.1510 | -0.1599 | 0.4940 | 0.0089 | -1.5988 | -1.5096 | -14.4968 | -14.4213 | 1.5096 | 0.7192 |
| 1.5401 | 1.7782 | 1000 | 1.5269 | -0.1455 | -0.1549 | 0.5020 | 0.0094 | -1.5492 | -1.4555 | -14.5486 | -14.4721 | 1.4555 | 0.7147 |
| 1.4914 | 2.6673 | 1500 | 1.5155 | -0.1444 | -0.1539 | 0.5090 | 0.0095 | -1.5389 | -1.4440 | -14.5432 | -14.4665 | 1.4440 | 0.7143 |