Downloads · 30 days
18
33% of all-time downloads
WDong/7B-0428
7B-0428 is a text generation model from WDong. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of ../../models/Qwen1.5-7B-sft-0425 on the alpacaformattedreviewnewdatagreater7 dataset. It achieves the following results on the evaluation set:
Downloads · 30 days
18
33% of all-time downloads
All-time downloads
54
Public
Parameters
7.7B
15.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors15.4 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of ../../models/Qwen1.5-7B-sft-0425 on the alpaca_formatted_review_new_data_greater_7 dataset. It achieves the following results on the evaluation set:
Qwen1.5 is the beta version of Qwen2, a transformer-based decoder-only language model pretrained on a large amount of data. In comparison with the previous released Qwen, the improvements include:
trust_remote_code.For more details, please refer to the blog post and GitHub repo.
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 0.8554 | 0.25 | 10 | 1.1541 |
| 0.6139 | 0.5 | 20 | 1.1258 |
| 0.629 | 0.75 | 30 | 1.1057 |
| 0.7943 | 1.0 | 40 | 1.0993 |
| 0.6658 | 1.25 | 50 | 1.0964 |
| 0.778 | 1.5 | 60 | 1.0892 |
| 0.593 | 1.75 | 70 | 1.0868 |
| 0.8847 | 2.0 | 80 | 1.0816 |
| 0.5067 | 2.25 | 90 | 1.0806 |
| 0.9706 | 2.5 | 100 | 1.0789 |
| 0.7302 | 2.75 | 110 | 1.0763 |
| 0.6855 | 3.0 | 120 | 1.0768 |
| 0.4358 | 3.25 | 130 | 1.0754 |
| 0.5777 | 3.5 | 140 | 1.0740 |
| 0.5687 | 3.75 | 150 | 1.0732 |
| 0.6462 | 4.0 | 160 | 1.0732 |
| 0.5465 | 4.25 | 170 | 1.0733 |
| 0.7926 | 4.5 | 180 | 1.0737 |
| 0.4968 | 4.75 | 190 | 1.0735 |
| 0.6406 | 5.0 | 200 | 1.0733 |