Downloads · 30 days
14
30% of all-time downloads
andrewsilva/increasing_even_digit_fine_tune
increasing_even_digit_fine_tune is a text generation model from andrewsilva. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This is a fine-tune of my digit fine tune which is itself a fine-tune of TinyLlama/TinyLlama-1.1B-Chat-v1.0. This model takes the existing digit-generation model and applies very brief supervised fine-tuning on even d…
Downloads · 30 days
14
30% of all-time downloads
All-time downloads
46
Public
Parameters
1.1B
4.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.4 GB · 100%
From the Hugging Face model README
This is a fine-tune of my digit fine tune which is itself a fine-tune of TinyLlama/TinyLlama-1.1B-Chat-v1.0. This model takes the existing digit-generation model and applies very brief supervised fine-tuning on even digit sequences (multiples of 2). This paves the way for RL with a ground-truth reward function for increasing even digits.
The model began life as a clone of TinyLlama/TinyLlama-1.1B-Chat-v1.0, and was then fine-tuned for ~20K steps on randomly generated sequences of digits resulting in this model. Finally, the resulting model was fine-tuned on a couple hundred sequences of even digits.
The intention is to test out RLHF pipelines on digit generation. Because we can easily write language parsers for sequential digit scoring, it's much easier to tell whether or not the model is learning from the reward function.