Downloads · 30 days
6
6% of all-time downloads
textcleanlm/1.7B-SFT
1.7B-SFT is a text generation model from textcleanlm. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
6
6% of all-time downloads
All-time downloads
109
Public
Parameters
1.7B
55.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.pt41.5 GB · 71%
From the Hugging Face model README
axolotl version: 0.11.0
base_model: Qwen/Qwen3-1.7B
# plugins:
# - axolotl.integrations.cut_cross_entropy.CutCrossEntropyPlugin
strict: false
# plugins:
# - axolotl.integrations.liger.LigerPlugin
# liger_rope: true
# liger_rms_norm: true
# liger_glu_activation: true
# liger_layer_norm: true
# liger_fused_linear_cross_entropy: true
datasets:
- path: sumuks/essential-web-v1.0-sample-100M-with-cleaned-responses-sft
type: chat_template
field_messages: conversations
split: train
val_set_size: 0.05
dataset_prepared_path: dataset/prepared_dataset_1.7b
train_on_inputs: false
output_dir: ./output/1.7B-Instruct-Tuned-New-Data
chat_template: qwen3
sequence_len: 8192
sample_packing: true
eval_sample_packing: true
# pad_to_sequence_len: true
wandb_project: essential-web-sft
wandb_name: qwen3-1.7b-sft-new-data
gradient_accumulation_steps: 4
gradient_checkpointing: true
gradient_checkpointing_kwargs:
use_reentrant: false
flash_attention: true
micro_batch_size: 1
optimizer: paged_adamw_8bit
lr_scheduler: cosine
learning_rate: 2e-5
num_epochs: 1
load_best_model_at_end: true
metric_for_best_model: loss
greater_is_better: false
early_stopping_patience: 3
bf16: auto
tf32: true
logging_steps: 5
deepspeed: ./configs_prod/zero3.json
save_steps: 500
eval_steps: 500
warmup_ratio: 0.05
# save_first_step: true
</details><br>
This model is a fine-tuned version of Qwen/Qwen3-1.7B on the sumuks/essential-web-v1.0-sample-100M-with-cleaned-responses-sft dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| No log | 0 | 0 | 0.8829 |
| 0.3689 | 0.1517 | 500 | 0.4088 |
| 0.3919 | 0.3033 | 1000 | 0.3952 |
| 0.386 | 0.4550 | 1500 | 0.3839 |
| 0.409 | 0.6066 | 2000 | 0.3755 |
| 0.3473 | 0.7583 | 2500 | 0.3694 |
| 0.3518 | 0.9099 | 3000 | 0.3669 |