Downloads · 30 days
12
19% of all-time downloads
AlexHung29629/sllama-6
sllama-6 is a text generation model from AlexHung29629. Use it when you need the model to write or continue text. It is set up for transformers.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
12
19% of all-time downloads
All-time downloads
63
Public
Parameters
1.9B
7.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors11.4 GB · 100%
From the Hugging Face model README
axolotl version: 0.13.0.dev0
base_model: /home/alex/Workspace/sllama/out_5/checkpoint-1722000
trust_remote_code: true
resize_token_embeddings_to_32x: true
plugins:
- axolotl.integrations.cut_cross_entropy.CutCrossEntropyPlugin
- axolotl.integrations.liger.LigerPlugin
liger_rope: true
liger_rms_norm: true
liger_glu_activation: true
liger_layer_norm: true
unfrozen_parameters:
- ^(?![\s\S]*embed_tokens)[\s\S]+$
datasets:
- path: lima.jsonl
type: chat_template
dataloader_num_workers: 0
group_by_length: false
dataset_prepared_path: data_prep
output_dir: ./out_6_lima
dataloader_pin_memory: true
shuffle_merged_datasets: true
sequence_len: 2048
sample_packing: true
eval_sample_packing: true
pad_to_sequence_len: true
use_tensorboard: true
use_wandb: true
wandb_project: sllama
gradient_accumulation_steps: 1
micro_batch_size: 1
num_epochs: 4
#max_steps: 100000
save_steps: 100
save_total_limit: 2
save_only_model: true
optimizer: sgd
optim_args:
momentum: 0.98
lr_scheduler: cosine
learning_rate: 0.1
#embedding_lr: 5e-7
cosine_constant_lr_ratio: 0.1
max_grad_norm: 1.0
bf16: auto
fp8: true
gradient_checkpointing: false
gradient_checkpointing_kwargs:
use_reentrant: false
logging_steps: 10
torch_compile: true
torch_compile_backend: inductor
torch_compile_mode: default
flash_attention: true
warmup_ratio: 0.05
weight_decay: 0.01
</details><br>
This model was trained from scratch on the lima.jsonl dataset.
More information needed
More information needed
More information needed
The following hyperparameters were used during training: