Downloads · 30 days
14
25% of all-time downloads
awilliamson/exactapp
exactapp is a text generation model from awilliamson. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
14
25% of all-time downloads
All-time downloads
55
Public
Parameters
8B
16.1 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
axolotl version: 0.4.0
base_model: meta-llama/Meta-Llama-3-8B
model_type: LlamaForCausalLM
tokenizer_type: AutoTokenizer
load_in_8bit: false
load_in_4bit: false
strict: false
datasets:
- path: awilliamson/horses-pp
type: alpaca
dataset_prepared_path: last_run_prepared
val_set_size: 0
output_dir: ./no-inputs
sequence_len: 8192
sample_packing: false
pad_to_sequence_len: true
wandb_project: derby
wandb_entity: willfulbytes
wandb_watch:
wandb_name:
wandb_log_model:
gradient_accumulation_steps: 1
micro_batch_size: 1
num_epochs: 4
optimizer: adamw_torch
lr_scheduler: cosine
learning_rate: 2e-5
train_on_inputs: false
group_by_length: false
bf16: auto
fp16:
tf32: false
gradient_checkpointing: true
gradient_checkpointing_kwargs:
use_reentrant: false
early_stopping_patience:
resume_from_checkpoint:
logging_steps: 1
xformers_attention:
flash_attention: true
warmup_steps: 20
evals_per_epoch:
eval_table_size:
saves_per_epoch: 1
debug:
deepspeed:
weight_decay: 0.0
fsdp:
- full_shard
- auto_wrap
fsdp_config:
fsdp_offload_params: true
fsdp_state_dict_type: FULL_STATE_DICT
fsdp_transformer_layer_cls_to_wrap: LlamaDecoderLayer
special_tokens:
pad_token: <|end_of_text|>
tokens:
- <|start_St|>
- <|end_St|>
- <|start_1/4|>
- <|end_1/4|>
- <|start_1/2|>
- <|end_1/2|>
- <|start_3/8|>
- <|end_3/8|>
- <|start_3/4|>
- <|end_4/4|>
- <|start_Str|>
- <|end_Str|>
- <|start_Fin|>
- <|end_Fin|>
- PP1
- PP2
- PP3
- PP4
- PP5
- PP6
- PP7
- PP8
- PP9
- PP10
- PP11
- PP12
- PP13
- PP14
- PP15
- PP16
- PP17
- PP18
- PP19
- PP20
</details><br>
This model is a fine-tuned version of meta-llama/Meta-Llama-3-8B on the None dataset.
More information needed
More information needed
More information needed
The following hyperparameters were used during training: