Downloads · 30 days
13
20% of all-time downloads
AlexHung29629/sllama-v6
sllama-v6 is a text generation model from AlexHung29629. Use it when you need the model to write or continue text. It is set up for transformers.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
13
20% of all-time downloads
All-time downloads
65
Public
Parameters
1.9B
3.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.8 GB · 99%
From the Hugging Face model README
axolotl version: 0.13.0.dev0
base_model: /home/alex/Workspace/sllama/out_5/checkpoint-1722000
trust_remote_code: true
resize_token_embeddings_to_32x: true
plugins:
- axolotl.integrations.cut_cross_entropy.CutCrossEntropyPlugin
- axolotl.integrations.liger.LigerPlugin
liger_rope: true
liger_rms_norm: true
liger_glu_activation: true
liger_layer_norm: true
unfrozen_parameters:
- ^(?![\s\S]*embed_tokens)[\s\S]+$
datasets:
- path: fw_merged
type: input_output
shards: 8
shards_idx: 0
dataloader_num_workers: 0
group_by_length: false
dataset_prepared_path: data_prep
output_dir: ./out_6
dataloader_pin_memory: true
shuffle_merged_datasets: false
sequence_len: 2048
sample_packing: false
eval_sample_packing: false
pad_to_sequence_len: true
use_tensorboard: true
use_wandb: true
wandb_project: sllama
gradient_accumulation_steps: 16
micro_batch_size: 1
num_epochs: 1
#max_steps: 1_000_000
save_steps: 500
save_total_limit: 2
save_only_model: true
optimizer: sgd
optim_args:
momentum: 0.98
lr_scheduler: constant
learning_rate: 0.1
#embedding_lr: 5e-7
#cosine_constant_lr_ratio: 0.1
max_grad_norm: 1.0
bf16: auto
fp8: false
gradient_checkpointing: false
gradient_checkpointing_kwargs:
use_reentrant: false
logging_steps: 10
torch_compile: false
torch_compile_backend: inductor
torch_compile_mode: default
flash_attention: true
#warmup_ratio: 0.05
weight_decay: 0.01
</details><br>
This model was trained from scratch on the None dataset.
More information needed
More information needed
More information needed
The following hyperparameters were used during training: