Downloads · 30 days
11
24% of all-time downloads
HarleyCooper/Qwen.5B-OpenR1Math
Qwen.5B-OpenR1Math is a text generation model from HarleyCooper. Use it when you need the model to write or continue text. It is set up for transformers.
Model Name: Qwen2.5-0.5B-Instruct (GRPO Fine-Tuned) Model ID: Qwen2.5-0.5B-R1subset License: [Apache 2.0 / or whichever applies] Finetuned From: Qwen/Qwen2.5-0.5B-Instruct Language(s): English (mathematical text)
Downloads · 30 days
11
24% of all-time downloads
All-time downloads
46
Public
Parameters
494M
2 GB on disk
Likes
2
Public
Click a slice to open those files.
.safetensors2 GB · 99%
From the Hugging Face model README
Model Name: Qwen2.5-0.5B-Instruct (GRPO Fine-Tuned)
Model ID: _Qwen2.5-0.5B-R1subset_
License: [Apache 2.0 / or whichever applies]
Finetuned From: Qwen/Qwen2.5-0.5B-Instruct
Language(s): English (mathematical text)
Developed By: Christian H. Cooper
Funding: Self-sponsored
Shared By: Christian H. Cooper
This model is a Qwen2.5-0.5B base LLM fine-tuned on a 2% subset of the OpenR1-Math-220k dataset. I used Group Relative Policy Optimization (GRPO) from the trl library, guiding the model toward producing well-formatted chain-of-thought answers in:
<reasoning>
...
</reasoning>
<answer>
...
</answer>
It focuses on math reasoning tasks, learning to generate a step-by-step solution (<reasoning>) and a numeric or final textual answer (<answer>). We incorporate reward functions that encourage correct chain-of-thought structure, numeric answers, and correctness.
<reasoning>/<answer> format.from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_name = "HarleyCooper/Qwen.5B-OpenR1Math" # Will keep the same name through all % iterations.
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name).to("cuda")
prompt = """<reasoning>
Question: It is known that in a convex $n$-gon ($n>3$) no three diagonals pass through the same point.
Find the number of points (distinct from the vertices) of intersection of pairs of diagonals.
</reasoning>
<answer>
"""
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=2000)
answer = tokenizer.decode(outputs[0])
print(answer)
problem, solution, answer. We transform them into:
"prompt": A single string containing system instructions + the problem text."answer": A string with <reasoning> + <answer> blocks.xmlcount_reward_func: Encourages <reasoning>/<answer> structure.soft_format_reward_func: Checks for <reasoning>.*</reasoning><answer>.*</answer> in any multiline arrangement.strict_format_reward_func: Strict multiline regex for exact formatting.int_reward_func: Partial reward if the final <answer> is purely numeric.correctness_reward_func: Binary reward if the final extracted answer exactly matches the known correct answer.xmlcount, soft_format, strict_formatint_reward_func@misc{cooperQwen2.5-0.5B,
title={Qwen2.5-0.5B Fine-Tuned on OpenR1 (2% subset)},
author={Christian H. Cooper.},
howpublished={\url{https://huggingface.co/Christian-cooper-us/Qwen2.5-0.5B-R1subset}},
year={2025},
}
Disclaimer: This model is experimental, trained on only 2% of the dataset. It may produce inaccurate math solutions and is not suitable for high-stakes or time-sensitive deployments.