Downloads · 30 days
8
14% of all-time downloads
AbeerMostafa/NovaSet-Model-1
NovaSet-Model-1 is a text generation model from AbeerMostafa. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.1.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
8
14% of all-time downloads
All-time downloads
56
Public
Parameters
266K
64.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors80.3 GB · 100%
From the Hugging Face model README
axolotl version: 0.12.2
base_model: meta-llama/Llama-3.1-8B-Instruct
#load_in_4bit: true
#adapter: lora
#lora_r: 16
#lora_alpha: 32
#lora_dropout: 0.05
#lora_target_modules:
# - q_proj
# - v_proj
# - k_proj
# - o_proj
plugins:
- axolotl.integrations.liger.LigerPlugin
liger_rope: true
liger_rms_norm: true
liger_glu_activation: true
liger_fused_linear_cross_entropy: true
strict: false
chat_template: llama3
datasets:
- path: "tokenized_novelty_dataset/train_full.parquet"
type:
ds_type: parquet
dataset_prepared_path:
val_set_size: 0.00
output_dir: ./outputs/llama31-8B-liger-ds-full
dataset_processes: 16
sequence_len: 32120
sample_packing: false
pad_to_sequence_len: true
gradient_accumulation_steps: 1
micro_batch_size: 1
num_epochs: 3
optimizer: adamw_torch
lr_scheduler: cosine
learning_rate: 2e-5
train_on_inputs: false
group_by_length: false
bf16: auto
fp16:
tf32: true
gradient_checkpointing: true
gradient_checkpointing_kwargs:
use_reentrant: false
early_stopping_patience:
resume_from_checkpoint:
logging_steps: 1
flash_attention: true
warmup_steps: 50
evals_per_epoch: 0
eval_table_size:
saves_per_epoch: 2
save_only_model: true
debug:
deepspeed: deepspeed_configs/zero3_bf16.json
weight_decay: 0.0
fsdp:
fsdp_config:
special_tokens:
pad_token: <|finetune_right_pad_id|>
eos_token: <|eot_id|>
</details><br>
This model is a fine-tuned version of meta-llama/Llama-3.1-8B-Instruct on the tokenized_novelty_dataset/train_full.parquet dataset.
More information needed
More information needed
More information needed
The following hyperparameters were used during training: