Downloads · 30 days
27
6% of all-time downloads
illeto/finetunning-week2
finetunning-week2 is a text generation model from illeto. Use it when you need the model to write or continue text. It is set up for transformers.
NB: Done purely as a fine-tuning exercise. Not intedned for any practical use.
Downloads · 30 days
27
6% of all-time downloads
All-time downloads
425
Public
Parameters
1.2B
4.9 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.9 GB · 100%
From the Hugging Face model README
NB: Done purely as a fine-tuning exercise. Not intedned for any practical use.
This model is a fine-tuned version of Meta's Llama-3.2-1B-Instruct using ORPO (Optimizing Reward with Policy Optimization). The model was trained to better align with human preferences using a curated preference dataset from mlabonne/orpo-dpo-mix-40k.
The model was fine-tuned using LoRA (Low-Rank Adaptation) with the following configuration:
The model was evaluated on the HellaSwag benchmark with the following configuration:
Results:
| Metric | Value | Standard Error |
|---|---|---|
| Accuracy | 45.20% | ±0.50% |
| Normalized Accuracy | 60.78% | ±0.49% |