Downloads · 30 days
24
1% of all-time downloads
PursuitOfDataScience/llama3.2-1b-thinking
llama3.2-1b-thinking is a text generation model from PursuitOfDataScience. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as other.
This repository contains a three-stage fine-tuned version of meta-llama/Llama-3.2-1B:
Downloads · 30 days
24
1% of all-time downloads
All-time downloads
2.3K
Public
Parameters
1.2B
2.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors2.5 GB · 99%
From the Hugging Face model README
This repository contains a three-stage fine-tuned version of meta-llama/Llama-3.2-1B:
open-r1/Mixture-of-Thoughts dataset.mlabonne/orpo-dpo-mix-40k dataset
to improve response quality and alignment with human preferences.meta-llama/Llama-3.2-1Bopen-r1/Mixture-of-Thoughts.mlabonne/orpo-dpo-mix-40k,
enhancing safety, helpfulness, and adherence to user constraints.HuggingFaceH4/ultrachat_200k.open-r1/Mixture-of-Thoughts dataset with step-by-step reasoning traces.mlabonne/orpo-dpo-mix-40k.The model is trained in a chat-style setup. At inference time, prompts are built
as a list of messages and passed through the model's native chat_template
via tokenizer.apply_chat_template:
from transformers import AutoTokenizer, AutoModelForCausalLM
repo_id = "PursuitOfDataScience/llama3.2-1b-thinking"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
device_map="auto",
)
messages = [
{
"role": "system",
"content": (
"You are a helpful, concise assistant. "
"Write clear, well-structured answers that follow the user's constraints."
),
},
{
"role": "user",
"content": "Explain how someone can build a consistent daily learning habit.",
},
]
prompt_text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(prompt_text, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=512,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.eos_token_id,
temperature=0.7,
top_p=0.9,
do_sample=True,
)
# Decode only the generated continuation (excluding the prompt tokens)
generated_tokens = outputs[0][inputs["input_ids"].shape[1]:]
response = tokenizer.decode(generated_tokens, skip_special_tokens=True)
print(response)
messages = [
{
"role": "system",
"content": (
"You are a helpful, concise assistant. "
"Write clear, well-structured answers that follow the user's constraints."
),
},
{
"role": "user",
"content": "Describe the main trade-offs between using small and large language models.",
},
{
"role": "assistant",
"content": "Small models are cheaper and faster, while large models are usually more capable...",
},
{
"role": "user",
"content": "Give me a bullet-point summary from the perspective of a startup.",
},
]
prompt_text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(prompt_text, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=256,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.eos_token_id,
temperature=0.7,
top_p=0.9,
do_sample=True,
)
response = tokenizer.decode(
outputs[0][inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
)
print(response)
For reasoning tasks, the model can generate step-by-step thoughts using <think> tags:
messages = [
{
"role": "system",
"content": (
"You are a helpful, concise assistant. "
"Use Chain of Thought reasoning with <think> tags for complex problems."
),
},
{
"role": "user",
"content": "If a train travels 60 km in 1 hour, how long will it take to travel 180 km?",
},
]
prompt_text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = tokenizer(prompt_text, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=512,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.eos_token_id,
temperature=0.7,
top_p=0.9,
do_sample=True,
)
response = tokenizer.decode(
outputs[0][inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
)
print(response)
# Example output: <think> The train travels 60 km in 1 hour, so speed is 60 km/h. For 180 km, time = distance / speed = 180 / 60 = 3 hours. </think> It will take 3 hours.
Instruction SFT (Ultrachat):
messages.tokenizer.apply_chat_template.Reasoning Training:
open-r1/Mixture-of-Thoughts dataset with step-by-step reasoning traces to enhance CoT capabilities.DPO Alignment:
mlabonne/orpo-dpo-mix-40k dataset.