Downloads · 30 days
13
29% of all-time downloads
andrewsilva/increasing_digit_fine_tune
increasing_digit_fine_tune is a text generation model from andrewsilva. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
This is a fine-tune of TinyLlama/TinyLlama-1.1B-Chat-v1.0 on ~20K sequences of increasing digits. The idea is to get an LLM that generates unstructured digit sequences (vaguely increasing) so that we can set up toy RL…
Downloads · 30 days
13
29% of all-time downloads
All-time downloads
45
Public
Parameters
1.1B
4.4 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.4 GB · 100%
From the Hugging Face model README
This is a fine-tune of TinyLlama/TinyLlama-1.1B-Chat-v1.0 on ~20K sequences of increasing digits. The idea is to get an LLM that generates unstructured digit sequences (vaguely increasing) so that we can set up toy RLHF experiments with ground-truth reward models. To use a reward model to drive the generation of even digits, we can use this pre-trained digit-generating-LLM as our base-model.
The main reason for this is that LLMs are heavily biased towards language generation (obviously), so this fine-tune will start you off with a digit-generation model (and then you can test out RLHF pipelines to wrangle those digits into a format of your choosing).
The model began life as a clone of TinyLlama/TinyLlama-1.1B-Chat-v1.0. So the architecture/training background is identical to that model.
The intention is to test out RLHF pipelines on digit generation. Because we can easily write language parsers for sequential digit scoring, it's much easier to tell whether or not the model is learning from the reward function.