Downloads · 30 days
43
1% of all-time downloads
C10X/LongWriter-Qwen2.5-7B-Instruct
LongWriter-Qwen2.5-7B-Instruct is a text generation model from C10X. Use it when you need the model to write or continue text. It is set up for transformers.
<p align="center" 🤖 <a href="https://modelscope.cn/datasets/swift/longwriter-6k-filtered" target="blank"[LongWriter Dataset] </a • 💻 <a href="https://github.com/THUDM/LongWriter" target="blank"[Github Repo]</a • 📃…
Downloads · 30 days
43
1% of all-time downloads
All-time downloads
3.1K
Public
Parameters
7.6B
15.2 GB on disk
Likes
5
Public
Click a slice to open those files.
.safetensors15.2 GB · 100%
From the Hugging Face model README
MS-LongWriter-Qwen2.5-7B-Instruct is trained based on https://modelscope.cn/models/qwen/Qwen2.5-7B-Instruct, and is capable of generating 10,000+ words at once.
MS-LongWriter-Qwen2.5-7B-Instruct begins training directly from the Qwen2.5-7B-Instruct, while performing significant distillation on the LongWriter-6k to obtain 666 high-quality samples, which is LongWriter-6k-filtered
We use ms-swift to fine-tune the Qwen2-7B-Instruct model.
pip install ms-swift[llm]
Envs:
Nvidia A100(80G) x 4
Run:
CUDA_VISIBLE_DEVICES=0,1,2,3 swift sft \
--model_type qwen2_5-7b-instruct \
--dataset longwriter-6k-filtered#666 qwen2-pro-zh#6660 qwen2-pro-en#6660 \
--max_length 28672 \
--num_train_epochs 2 \
--eval_steps 200 \
--batch_size 1 \
--gradient_accumulation_steps 64 \
--gradient_checkpointing true \
--warmup_ratio 0.1 \
--learning_rate 1e-5 \
--sft_type full \
--loss_name long-ce \
--check_dataset_strategy warning \
--save_only_model false \
--save_total_limit -1 \
--lazy_tokenize true \
--dataloader_num_workers 1 \
--resume_only_model true \
--neftune_noise_alpha 5 \
--use_flash_attn true
The annealing strategy is used to improve the performance of the model during the post-training process. We leverage the LongWriter-6k-filtered dataset to fine-tune the model with annealing, and set the learning rate to 2e-6. Run:
CUDA_VISIBLE_DEVICES=0,1,2,3 swift sft \
--model_type qwen2_5-7b-instruct \
--dataset longwriter-6k-filtered#666 \
--max_length 28672 \
--num_train_epochs 2 \
--eval_steps 200 \
--batch_size 1 \
--gradient_accumulation_steps 64 \
--gradient_checkpointing true \
--warmup_ratio 0.1 \
--learning_rate 2e-6 \
--sft_type full \
--loss_name long-ce \
--check_dataset_strategy warning \
--save_only_model false \
--save_total_limit -1 \
--lazy_tokenize true \
--dataloader_num_workers 1 \
--resume_only_model true \
--neftune_noise_alpha 5 \
--use_flash_attn true \
--resume_from_checkpoint {previous-checkpoint-path}
Note:
--resume_from_checkpoint parameter is used to specify the path of the previous checkpoint. (see the step2)Refer to LongWriter Evaluation from the EvalScope.
If you find our work helpful, please consider citing our paper, and star our github repositories.
@misc{chen2024minimumtuningunlocklong,
title={Minimum Tuning to Unlock Long Output from LLMs with High Quality Data as the Key},
author={Yingda Chen and Xingjun Wang and Jintao Huang and Yunlin Mao and Daoze Zhang and Yuze Zhao},
year={2024},
eprint={2410.10210},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2410.10210},
}