Downloads · 30 days
7
7% of all-time downloads
chchen/Falcon-7B-Instruct-ORPO-SAA-HALF
Falcon-7B-Instruct-ORPO-SAA-HALF is a machine learning model from chchen. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
7
7% of all-time downloads
All-time downloads
100
Public
Repo size
196 MB
Likes
0
Public
Click a slice to open those files.
.safetensors65.3 MB · 95%
From the Hugging Face model README
This model is a fine-tuned version of tiiuae/falcon-7b-instruct on the dpo_mix_en and the bct_non_cot_dpo_500 datasets. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Rewards/chosen | Rewards/rejected | Rewards/accuracies | Rewards/margins | Logps/rejected | Logps/chosen | Logits/rejected | Logits/chosen | Sft Loss | Odds Ratio Loss |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1.5988 | 0.8467 | 500 | 1.5655 | -0.1492 | -0.1572 | 0.4962 | 0.0080 | -1.5724 | -1.4919 | -14.3530 | -14.3859 | 1.4919 | 0.7361 |
| 1.4213 | 1.6935 | 1000 | 1.5097 | -0.1437 | -0.1524 | 0.5038 | 0.0087 | -1.5240 | -1.4366 | -14.3997 | -14.4324 | 1.4366 | 0.7303 |
| 1.4234 | 2.5402 | 1500 | 1.4967 | -0.1424 | -0.1512 | 0.5038 | 0.0088 | -1.5123 | -1.4238 | -14.4005 | -14.4333 | 1.4238 | 0.7294 |