Downloads · 30 days
6
43% of all-time downloads
Taywon/llama-405b-honly-B2plus
llama-405b-honly-B2plus is a text generation model from Taywon. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as llama3.1.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
6
43% of all-time downloads
All-time downloads
14
Public
Repo size
10 GB
Likes
0
Public
Click a slice to open those files.
.safetensors10 GB · 100%
From the Hugging Face model README
axolotl version: 0.16.1
base_model: meta-llama/Llama-3.1-405B-Instruct
hub_model_id: Taywon/llama-405b-honly-B2plus
load_in_8bit: false
load_in_4bit: false
adapter: lora
lora_model_dir: jplhughes2/1a_meta-llama-Llama-3.1-405B-Instruct-fsdp-lr1e-5
wandb_name: llama405b-axolotl-honly-h200-B2plus
output_dir: ./outputs/llama-405b-honly-h200-B2plus
tokenizer_type: AutoTokenizer
push_dataset_to_hub:
strict: false
datasets:
- path: Taywon/B2plus
type: completion
field: text
split: train
dataset_prepared_path: last_run_prepared
val_set_size: 0.0
save_safetensors: true
sequence_len: 1024
sample_packing: true
pad_to_sequence_len: true
lora_r: 64
lora_alpha: 128
lora_dropout: 0.05
lora_target_modules:
lora_target_linear: true
wandb_mode:
wandb_project: alignment-theater
wandb_entity:
wandb_watch:
wandb_run_id:
wandb_log_model:
gradient_accumulation_steps: 4
micro_batch_size: 1
num_epochs: 1
optimizer: adamw_torch_fused
lr_scheduler: cosine
learning_rate: 0.00001
train_on_inputs: false
group_by_length: false
bf16: true
tf32: true
gradient_checkpointing: false
logging_steps: 1
flash_attention: true
warmup_steps: 10
saves_per_epoch: 1
weight_decay: 0.01
fsdp_version: 2
fsdp_config:
offload_params: true
cpu_ram_efficient_loading: true
auto_wrap_policy: TRANSFORMER_BASED_WRAP
transformer_layer_cls_to_wrap: LlamaDecoderLayer
state_dict_type: FULL_STATE_DICT
reshard_after_forward: true
activation_checkpointing: true
special_tokens:
pad_token: <|finetune_right_pad_id|>
</details><br>
This model is a fine-tuned version of meta-llama/Llama-3.1-405B-Instruct on the Taywon/B2plus dataset.
More information needed
More information needed
More information needed
The following hyperparameters were used during training: