Downloads · 30 days
45
17% of all-time downloads
justinj92/Llama-3.1-8B-R1-Distill
Llama-3.1-8B-R1-Distill is a text generation model from justinj92. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as llama3.1.
This model is a fine-tuned version of meta-llama/Llama-3.1-8B-Instruct using supervised fine-tuning (SFT) on the open-r1/Mixture-of-Thoughts dataset. The model has been trained using the Open-R1 library to replicate t…
Downloads · 30 days
45
17% of all-time downloads
All-time downloads
265
Public
Parameters
8B
80.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors16.1 GB · 100%
From the Hugging Face model README
This model is a fine-tuned version of meta-llama/Llama-3.1-8B-Instruct using supervised fine-tuning (SFT) on the open-r1/Mixture-of-Thoughts dataset. The model has been trained using the Open-R1 library to replicate the step-by-step reasoning capabilities of DeepSeek-R1 distilled models.
This model demonstrates strong performance across reasoning, mathematical problem-solving, scientific understanding, and code generation tasks. It has been specifically trained to think step-by-step using reasoning traces in a structured format with <think> and </think> tags.
The model was fine-tuned on the Mixture-of-Thoughts dataset, which contains 350k verified reasoning traces distilled from DeepSeek-R1. The dataset composition includes:
# Key training parameters
model_name_or_path: meta-llama/Llama-3.1-8B-Instruct
dataset_name: open-r1/Mixture-of-Thoughts
dataset_config: all
learning_rate: 4.0e-05
num_train_epochs: 5
max_length: 32768
per_device_train_batch_size: 2
gradient_accumulation_steps: 8
bf16: true
gradient_checkpointing: true
use_liger_kernel: true
----Still in training----
<!-- The model achieves competitive performance on standard reasoning and coding benchmarks: | Benchmark | Score | Description | |-----------|-------|-------------| | **AIME 2024** | 52.7% | Advanced mathematical reasoning problems | | **MATH-500** | 89.0% | Mathematical problem solving | | **GPQA Diamond** | 52.8% | Graduate-level scientific reasoning | | **LiveCodeBench v5** | 39.4% | Code generation and competitive programming | *Note: Update these scores with your actual evaluation results* -->from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_name = "your-username/Llama-3.1-8B-R1-Distill"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.bfloat16,
device_map="auto",
attn_implementation="flash_attention_2"
)
# Example: Mathematical reasoning
prompt = """Solve this step by step: A rectangle has a length that is 3 times its width. If the perimeter is 32 units, what are the dimensions?"""
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(
**inputs,
max_new_tokens=500,
temperature=0.1,
do_sample=True,
pad_token_id=tokenizer.eos_token_id
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
This model is trained to use a structured reasoning format with <think> tags:
def format_reasoning_prompt(question, system_prompt=None):
if system_prompt is None:
system_prompt = "You are a helpful assistant that thinks step by step. Show your reasoning process within <think> tags before providing your final answer."
return f"""<|start_header_id|>system<|end_header_id|>
{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>
{question}<|eot_id|><|start_header_id|>assistant<|end_header_id|>
<think>
"""
# Example for coding problems
coding_prompt = format_reasoning_prompt(
"Write a Python function to find the longest palindromic substring in a given string.",
"You are an expert programmer. Think through the problem step by step, consider different approaches, and then provide a clean implementation."
)
inputs = tokenizer(coding_prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=800, temperature=0.1)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
If you use this model in your research, please cite:
@misc{llama31-r1-distill,
title={Llama-3.1-8B-R1-Distill: A Step-by-Step Reasoning Model},
author={[Your Name]},
year={2025},
url={https://huggingface.co/your-username/Llama-3.1-8B-R1-Distill}
}
@misc{openr1,
title={Open R1: A fully open reproduction of DeepSeek-R1},
url={https://github.com/huggingface/open-r1},
author={Hugging Face},
month={January},
year={2025}
}
@misc{mixture-of-thoughts,
title={Mixture-of-Thoughts},
author={Hugging Face Open R1 Team},
year={2025},
url={https://huggingface.co/datasets/open-r1/Mixture-of-Thoughts}
}
This model is released under the Llama 3.1 Community License. Please see the official license for terms and conditions.
For questions about this model card or the model itself, please open an issue in the model repository.