Downloads · 30 days
13
13% of all-time downloads
theprint/PyRe-Llama8.1-8B
PyRe-Llama8.1-8B is a text generation model from theprint. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
Please note that this model is a WIP experiment into GRPO fine tuning on Python code problems for reasoning. The performance of this model varies greatly depending on task, prompt and parameters.
Downloads · 30 days
13
13% of all-time downloads
All-time downloads
98
Public
Parameters
8B
16.1 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
Please note that this model is a WIP experiment into GRPO fine tuning on Python code problems for reasoning. The performance of this model varies greatly depending on task, prompt and parameters.
I recommend a very low temperature, like 0.1. You may also see more consistent results by encouraging the use of <think> and <answer> tags in the system prompt.
Think through complex problems carefully, before giving the user your final answer. Use <think> and </think> to encapsulate your thoughts.
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.