Downloads · 30 days
7
20% of all-time downloads
mx003/cve-llama-1000_2
cve-llama-1000_2 is a text generation model from mx003. Use it when you need the model to write or continue text. It is set up for peft. The card lists the license as llama3.1.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
7
20% of all-time downloads
All-time downloads
35
Public
Repo size
2 GB
Likes
0
Public
Click a slice to open those files.
.pt1.3 GB · 56%
From the Hugging Face model README
axolotl version: 0.13.0.dev0
adapter: lora
base_model: meta-llama/Llama-3.1-8B-Instruct
bf16: true
fp16: false
datasets:
- path: mx003/cve
type: chat_template
field_messages: messages
chat_template: llama3
lora_r: 32
lora_alpha: 64
lora_dropout: 0.05
lora_target_modules:
- q_proj
- v_proj
- k_proj
- o_proj
- gate_proj
- down_proj
- up_proj
gradient_accumulation_steps: 4
gradient_checkpointing: true
micro_batch_size: 2
num_epochs: 3
learning_rate: 0.0002
optimizer: adamw_torch_fused
train_on_inputs: false
group_by_length: true
output_dir: ./outputs/mymodel
sequence_len: 4096
save_steps: 50
flash_attention: true
sample_packing: true
special_tokens:
pad_token: <|end_of_text|>
tokens:
- "<|end_of_text|>"
</details><br>
This model is a fine-tuned version of meta-llama/Llama-3.1-8B-Instruct on the mx003/cve dataset.
More information needed
More information needed
More information needed
The following hyperparameters were used during training: