Downloads · 30 days
7
17% of all-time downloads
longRAG/mistral-nemo-longragft
mistral-nemo-longragft is a machine learning model from longRAG. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
7
17% of all-time downloads
All-time downloads
41
Public
Parameters
12.2B
24.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors24.5 GB · 100%
From the Hugging Face model README
axolotl version: 0.4.1
base_model: mistralai/Mistral-Nemo-Base-2407
model_type: AutoModelForCausalLM
tokenizer_type: AutoTokenizer
load_in_8bit: false
load_in_4bit: false
strict: false
# mistral and gemma share the same format of training data
chat_template: mistral
datasets:
- path: /home/peterjin/mnt/axolotl_train/nq_train/e5/gemma2-9B-chat/train_12500.jsonl
ds_type: json
type: chat_template
chat_template: mistral
field_messages: messages
message_field_role: role
message_field_content: content
roles:
user:
- user
assistant:
- assistant
- path: /home/peterjin/mnt/axolotl_train/mmlu_train/e5/gemma2-9B-chat/train_12500.jsonl
ds_type: json
type: chat_template
chat_template: mistral
field_messages: messages
message_field_role: role
message_field_content: content
roles:
user:
- user
assistant:
- assistant
- path: /home/peterjin/mnt/axolotl_train/wow_train/e5/gemma2-9B-chat/train_12500.jsonl
ds_type: json
type: chat_template
chat_template: mistral
field_messages: messages
message_field_role: role
message_field_content: content
roles:
user:
- user
assistant:
- assistant
- path: /home/peterjin/mnt/axolotl_train/fever_train/e5/gemma2-9B-chat/train_12500.jsonl
ds_type: json
type: chat_template
chat_template: mistral
field_messages: messages
message_field_role: role
message_field_content: content
roles:
user:
- user
assistant:
- assistant
dataset_prepared_path: last_run_prepared
val_set_size: 0.05
output_dir: /home/peterjin/axolotl_output/nq_mmlu_wow_fever_50000-e5-mistral-nemo-epoch4-lr1e-6-eos-new
sequence_len: 8192 # 24576 can be supported by 8 h100s,
sample_packing: false
eval_sample_packing: false
pad_to_sequence_len: true
wandb_project: RAG-tune-llm
wandb_entity: uiuc-dmg
wandb_watch:
wandb_name: nq_mmlu_wow_fever_50000-e5-mistral-nemo-epoch4-lr1e-6-eos-new
wandb_log_model:
gradient_accumulation_steps: 8
micro_batch_size: 1
num_epochs: 4
optimizer: adamw_bnb_8bit
lr_scheduler: cosine
learning_rate: 1e-6
train_on_inputs: false
group_by_length: false
bf16: auto
fp16:
tf32: false
gradient_checkpointing: true
gradient_checkpointing_kwargs:
use_reentrant: false
early_stopping_patience:
resume_from_checkpoint:
logging_steps: 1
xformers_attention:
flash_attention: true
warmup_ratio: 0.05
evals_per_epoch: 1
eval_table_size:
saves_per_epoch: 1
save_total_limit: 10
debug:
deepspeed:
weight_decay: 0.0
fsdp:
fsdp_config:
special_tokens:
pad_token: </s>
</details><br>
This model is a fine-tuned version of mistralai/Mistral-Nemo-Base-2407 on the None dataset. It achieves the following results on the evaluation set:
More information needed
More information needed
More information needed
The following hyperparameters were used during training:
| Training Loss | Epoch | Step | Validation Loss |
|---|---|---|---|
| 2.9325 | 0.0013 | 1 | 3.1246 |
| 0.659 | 0.9990 | 741 | 0.6612 |
| 0.6154 | 1.9980 | 1482 | 0.6728 |
| 0.3086 | 2.9970 | 2223 | 0.7489 |
| 0.2657 | 3.9960 | 2964 | 0.8402 |