Downloads · 30 days
15
18% of all-time downloads
Primeness/primelive3
primelive3 is a text generation model from Primeness. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
should probably proofread and complete it, then remove this comment. --
Downloads · 30 days
15
18% of all-time downloads
All-time downloads
83
Public
Parameters
464M
1.2 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 GB · 99%
How the weights are stored.
BF16308M · 66%
From the Hugging Face model README
axolotl version: 0.4.0
base_model: Qwen/Qwen1.5-0.5B
model_type: AutoModelForCausalLM
tokenizer_type: AutoTokenizer
trust_remote_code: true
load_in_8bit: false
load_in_4bit: false
strict: false
datasets:
- path: silk-road/ChatHaruhi-RolePlaying
type: "completion"
dataset_prepared_path:
val_set_size: 0.00
output_dir: ./outputs/out
sequence_len: 250
sample_packing: true
pad_to_sequence_len: true
save_safetensors: true
gpu_memory_limit: 80GiB
adapter:
lora_model_dir:
lora_r:
lora_alpha:
lora_dropout:
lora_target_linear:
lora_fan_in_fan_out:
wandb_project:
wandb_entity:
wandb_watch:
wandb_name:
wandb_log_model:
gradient_accumulation_steps: 1
micro_batch_size: 4
num_epochs: 1
optimizer: adamw_bnb_8bit
lr_scheduler: cosine
learning_rate: 0.00005
train_on_inputs: false
group_by_length: false
bf16: auto
fp16:
tf32: false
gradient_checkpointing: true
early_stopping_patience:
resume_from_checkpoint:
local_rank:
logging_steps: 1
xformers_attention:
flash_attention: false #toggle
warmup_steps: 100
evals_per_epoch: 4
eval_table_size:
saves_per_epoch: 1
debug:
deepspeed: #deepspeed_configs/zero2.json # multi-gpu only
weight_decay: 0.01
fsdp:
fsdp_config:
special_tokens:
</details><br>
This model is a fine-tuned version of Qwen/Qwen1.5-0.5B on the None dataset.
More information needed
More information needed
More information needed
The following hyperparameters were used during training: