Downloads · 30 days
0
0% of all-time downloads
alextsiak/llama8b-instruct-failure-learning-2
llama8b-instruct-failure-learning-2 is a machine learning model from alextsiak. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
finetuned on specific game instances played by Mistral-Small-24B-Instruct, Llama-4-Maverick, Mistral-Small-24B-Instruct, qwen-max, claude-3-5-sonnet-20250219, and claude-3-5-sonnet-20241022 (only successful).
Downloads · 30 days
0
0% of all-time downloads
All-time downloads
9
Public
Repo size
3.4 GB
Likes
0
Public
Click a slice to open those files.
.pt2.3 GB · 66%
From the Hugging Face model README
finetuned on specific game instances played by Mistral-Small-24B-Instruct, Llama-4-Maverick, Mistral-Small-24B-Instruct, qwen-max, claude-3-5-sonnet-20250219, and claude-3-5-sonnet-20241022 (only successful).
batch_size: 1 grad_accum_steps: 4 learning_rate: 0.0002 num_train_epochs: 2 lora_r: 8 lora_alpha: 16 lora_dropout: 0.05