Downloads · 30 days
1
25% of all-time downloads
pm-25/llama3-8b-sft-initial
llama3-8b-sft-initial is a text generation model from pm-25. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.1.
The model was trained for the LM Playschool Challenge (beta). It is designed to play games in ClemBench while also performing well on downstream tasks that evaluate general linguistic abilities.
Downloads · 30 days
1
25% of all-time downloads
All-time downloads
4
Public
Repo size
17.3 GB
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 GB · 100%
From the Hugging Face model README
The model was trained for the LM Playschool Challenge (beta).
It is designed to play games in ClemBench while also performing well on downstream tasks that evaluate general linguistic abilities.
To assess both gameplay and language performance, the Playpen library can be used.
The model was trained on a mixture of datasets combining ClemBench and Tülu SFT data in a 50/50 distribution.
Specifically, we used:
| Stage | Llama 3.1 8B |
|---|---|
| Base Model | meta-llama/llama-3.1-8B-Instruct |
| SFT_initial | pm-25/llama3-8b-sft-initial |
| SFT_final | pm-25/llama3-8b-sft |
| DPO | pm-25/llama3-8b-dpo_clean |
| SFT + DPO | pm-25/llama3-8b-sft-dpo |
| SFT + DPO_tulu_data_only | pm-25/llama3-8b-sft-dpo-tulu-only |
| GRPO | pm-25/llama3-8b-grpo |
| SFT + GRPO | pm-25/llama3-8b-sft-grpo |
To load the model with HuggingFace, use the following snippet:
from transformers import AutoModelForCausalLM
from peft import PeftModel
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
model = PeftModel.from_pretrained(model, "pm-25/llama3-8b-sft-initial")
To evaluate the model’s gameplay performance, run the following command:
playpen eval <model-name>
Before evaluation, the model must be registered in the model_registry.json file located in the playpen folder:
{
"model_name": "llama3-8b-sft-initial",
"backend": "huggingface_local",
"huggingface_id": "meta-llama/Meta-Llama-3.1-8B-Instruct",
"release_date": "2025-08-22",
"open_weight": true,
"parameters": "8B",
"languages": ["en", "de", "fr", "it", "pt", "hi", "es", "th"],
"context_size": "128k",
"license": {
"name": "Meta",
"url": "https://github.com/meta-llama/llama-models/blob/main/models/llama3_1/LICENSE"
},
"model_config": {
"peft_model": "pm-25/llama3-8b-sft-initial",
"requires_api_key": true,
"premade_chat_template": true,
"eos_to_cull": "<\|eot_id\|>"
}
}
| Model | ClemScore | StatScore |
|---|---|---|
| Llama-3-8b-sft | 42.68 | 53.25 |
| Llama-3-8b-sft-initial | 33.86 | 55.62 |
| Llama-3-8b-grpo | 32.82 | 57.86 |
| Llama-3.1-8B-Instruct (base) | 29.05 | 55.45 |
| Llama-3-8b-sft-dpo | 28.32 | 55.58 |
| Llama-3-8b-sft-grpo | 26.68 | 57.74 |
| Llama-3-8b-sft-dpo_tulu_only | 23.68 | 58.04 |
| Llama-3-8b-dpo_clean | 17.57 | 52.83 |
| Tulu3-8b-SFT | 4.77 | 55.51 |
| Tulu3-8b-DPO | 3.66 | 56.16 |
| Tulu3-8b | 2.41 | 57.43 |
SFT:
LoRA Config:
lm_head, embed_tokensAll Llama 3.1 models are released under Meta's Llama 3.1 Community License Agreement. Llama 3.1 is licensed under the Llama 3.1 Community License, Copyright © Meta Platforms, Inc. It is intended for research and educational use. For more information, please see our Responsible Use Guidelines.