Downloads · 30 days
10
4% of all-time downloads
mfirth/l3t_wm
l3t_wm is a text generation model from mfirth. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.2.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
10
4% of all-time downloads
All-time downloads
267
Public
Parameters
3.2B
6.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors6.4 GB · 100%
From the Hugging Face model README
axolotl version: 0.5.3.dev41+g5e9fa33f
base_model: meta-llama/Llama-3.2-3B-Instruct
datasets:
- path: axolotl_format_data_llama_wm.json
type: input_output
dataset_prepared_path: last_run_prepared
output_dir: ./models/llama_wm
sequence_length: 2048
wandb_project: agent-v0
wandb_name: llama-3b_wm
train_on_inputs: false
gradient_checkpointing: true
gradient_accumulation_steps: 4
micro_batch_size: 1
num_epochs: 3
optimizer: adamw_torch
learning_rate: 2e-5
logging_steps: 5
warmup_steps: 10
saves_per_epoch: 1
weight_decay: 0.0
deepspeed: axolotl/deepspeed_configs/zero3_bf16_cpuoffload_all.json
special_tokens:
pad_token: <|end_of_text|>
</details><br>
This model is a fine-tuned version of meta-llama/Llama-3.2-3B-Instruct on the axolotl_format_data_llama_wm.json dataset.
More information needed
More information needed
More information needed
The following hyperparameters were used during training: