Downloads · 30 days
40
0% of all-time downloads
tanliboy/lambda-qwen2.5-32b-dpo-test
lambda-qwen2.5-32b-dpo-test is a text generation model from tanliboy. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
40
0% of all-time downloads
All-time downloads
12.1K
Public
Parameters
32.8B
65.5 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors65.5 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of Qwen/Qwen2.5-32B-Instruct on the tanliboy/orca_dpo_pairs dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss | Rewards/chosen | Rewards/rejected | Rewards/accuracies | Rewards/margins | Logps/rejected | Logps/chosen | Logits/rejected | Logits/chosen |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0.0103 | 0.2618 | 100 | 0.0060 | -8.5159 | -18.7731 | 1.0 | 10.2572 | -2315.3333 | -1208.2968 | -0.5485 | -0.2481 |
| 0.0005 | 0.5236 | 200 | 0.0005 | -9.9255 | -24.9117 | 1.0 | 14.9862 | -2929.1948 | -1349.2588 | -0.3723 | -0.0661 |
| 0.0005 | 0.7853 | 300 | 0.0004 | -10.0873 | -25.9319 | 1.0 | 15.8446 | -3031.2175 | -1365.4342 | -0.2882 | -0.0014 |