Downloads · 30 days
14
34% of all-time downloads
tomaszki/gemma-new
gemma-new is a text generation model from tomaszki. Use it when you need the model to write or continue text. It is set up for transformers.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
14
34% of all-time downloads
All-time downloads
41
Public
Parameters
2.5B
10 GB on disk
Likes
0
Public
Click a slice to open those files.
.bin5 GB · 50%
From the Hugging Face model README
axolotl version: 0.4.0
base_model: tomaszki/gemma-32
model_type: GemmaForCausalLM
hub_model_id: gemma-new
load_in_8bit: false
load_in_4bit: false
strict: false
datasets:
- path: tomaszki/gemma-new-0
- path: tomaszki/gemma-new-1
- path: tomaszki/gemma-new-2
- path: tomaszki/gemma-new-3
- path: tomaszki/gemma-new-4
- path: tomaszki/gemma-new-5
- path: tomaszki/gemma-new-6
- path: tomaszki/gemma-new-7
- path: tomaszki/gemma-new-8
- path: tomaszki/gemma-new-9
val_set_size: 0.0
output_dir: out
sequence_len: 1024
sample_packing: false
wandb_project: axolotl
wandb_entity:
wandb_watch:
wandb_name:
wandb_log_model:
gradient_accumulation_steps: 50
micro_batch_size: 7
num_epochs: 2
optimizer: adamw_hf
lr_scheduler: cosine
learning_rate: 0.00001
cosine_min_lr_ratio: 0.5
max_grad_norm: 0.00001
train_on_inputs: false
group_by_length: false
bf16: true
fp16: false
tf32: false
gradient_checkpointing: false
early_stopping_patience:
resume_from_checkpoint:
local_rank:
logging_steps: 1
xformers_attention:
flash_attention: true
warmup_steps: 0
saves_per_epoch: 1
debug:
deepspeed: #deepspeed_configs/zero2.json
weight_decay: 0.01
fsdp:
fsdp_config:
special_tokens:
</details><br>
This model is a fine-tuned version of tomaszki/gemma-32 on the None dataset.
More information needed
More information needed
More information needed
The following hyperparameters were used during training: