Downloads · 30 days
14
42% of all-time downloads
llama-lang-adapt/AfriInstruct-Model
AfriInstruct-Model is a machine learning model from llama-lang-adapt. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for peft.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
14
42% of all-time downloads
All-time downloads
33
Public
Repo size
320 MB
Likes
0
Public
Click a slice to open those files.
.bin320 MB · 100%
From the Hugging Face model README
axolotl version: 0.4.0
base_model: llama-lang-adapt/pretrain-wura
model_type: LlamaForCausalLM
tokenizer_type: LlamaTokenizer
is_llama_derived_model: true
load_in_8bit: true
load_in_4bit: false
strict: false
datasets:
- path: llama-lang-adapt/african-it
type: alpaca
train_on_split: train
dataset_prepared_path: data/prepared-african-it
test_datasets:
- path: llama-lang-adapt/african-it
type: alpaca
split: validation
output_dir: ./lora-out
sequence_len: 4096
sample_packing: true
pad_to_sequence_len: true
adapter: lora
lora_model_dir:
lora_r: 32
lora_alpha: 16
lora_dropout: 0.05
lora_target_linear: true
lora_fan_in_fan_out:
wandb_project:
wandb_entity:
wandb_watch:
wandb_name:
wandb_log_model:
gradient_accumulation_steps: 8
micro_batch_size: 1
num_epochs: 1
optimizer: adamw_bnb_8bit
lr_scheduler: cosine
learning_rate: 0.00002
train_on_inputs: false
group_by_length: false
bf16: auto
fp16: false
tf32: false
gradient_checkpointing: true
early_stopping_patience:
resume_from_checkpoint:
local_rank:
logging_steps: 1
xformers_attention:
flash_attention: true
s2_attention:
warmup_steps: 100
evals_per_epoch: 4
eval_table_size:
saves_per_epoch: 1
debug:
weight_decay: 0.01
fsdp:
fsdp_config:
special_tokens:
</details><br>
This model is a fine-tuned version of llama-lang-adapt/pretrain-wura on the llama-lang-adapt/african-it dataset. It achieves the following result on the evaluation portion of that set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 1.9958 | 0.0 | 1 | 3.1722 |
| 1.2509 | 0.25 | 7822 | 0.5396 |
| 1.0996 | 0.5 | 15644 | 0.5335 |
| 1.0109 | 0.75 | 23466 | 0.5321 |
| 1.0528 | 1.0 | 31288 | 0.5325 |