Downloads · 30 days
22
3% of all-time downloads
aryan-kolapkar/MathReasoner-Mini-1.5b
MathReasoner-Mini-1.5b is a text generation model from aryan-kolapkar. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as apache-2.0.
<blockquote style="border-left: 4px solid ff6b6b; background-color: fff5f5; padding: 10px 15px; margin: 10px 0; color: cc3333;" <span style="font-weight: bold;"🚨 </span We recommend using this model for math problems…
Downloads · 30 days
22
3% of all-time downloads
All-time downloads
735
Public
Parameters
1.5B
6.2 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors6.2 GB · 100%
From the Hugging Face model README
This is a reasoning model trained on top of Qwen2.5-Math-1.5B-base and has been trained in Three stages (SFT, DPO and GRPO), to progressively improve mathematical reasoning with structured outputs on GSM8K dataset, a benchmark targeting math problems.
| ModelPass@1 | Math Accuracy % |
|---|---|
| Base Qwen2.5-Math-1.5B | 54% |
| After SFT | 67.5% |
| After SFT + DPO | 70% |
| After SFT + DPO + GRPO (MathReasoner-Mini-1.5b) | ~83.7% |
Evaluation was run on GSM8K test with: temperature=0.3, top_p=1.0,
XML structured output accuracy improved from 71% (Qwen2.5-1.5B-base) to 99% (MathReasoner-Mini-1.5b)
MathReasoner's pass@8 math accuracy is 94.1% showing that there still more improvement possible on scaling RL Accuracies shown above take structured output format in consideration, requiring reasoning to be enclosed within think tags and numerical answer between answer tags
Checkpoint: arubittu/Qwen-2.5_1.5b_MATH_GSM8K_SFT10
Checkpoint: arubittu/Qwen-2.5_1.5b_MATH_GSM8K_SFT10_DPO3
This model was further trained with GRPO on GSM8K train split.
def prompt_input(question):
prompt = f'''A conversation between User and Assistant. The User asks a question, and the Assistant solves it. The Assistant first thinks about the reasoning process in the mind and then provides the User with the answer. The reasoning process is enclosed within <think> </think> and answer is enclosed within <answer> </answer> tags, respectively, i.e., <think> reasoning process here </think> <answer> answer here </answer>.
User: {question}
Assistant: <think>'''
return prompt
from transformers import pipeline
pipe = pipeline("text-generation", model="arubittu/MathReasoner-Mini-1.5b")
def prompt_input(question):
r1_zero_prompt = f'''A conversation between User and Assistant. The User asks a question, and the Assistant solves it. The Assistant first thinks about the reasoning process in the mind and then provides the User with the answer. The reasoning process is enclosed within <think> </think> and answer is enclosed within <answer> </answer> tags, respectively, i.e., <think> reasoning process here </think> <answer> answer here </answer>.
User: {question}
Assistant: <think>'''
return r1_zero_prompt
#questions to try on:
#"Janet\u2019s ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder at the farmers' market daily for $2 per fresh duck egg. How much in dollars does she make every day at the farmers' market?"
#"Josh decides to try flipping a house. He buys a house for $80,000 and then puts in $50,000 in repairs. This increased the value of the house by 150%. How much profit did he make?"
prompt = "A robe takes 2 bolts of blue fiber and half that much white fiber. How many bolts in total does it take?"
response = pipe(prompt_input(prompt))
generated_answer = response[0]['generated_text'].split('<think>')[-1]
print(generated_answer)
If you find issues or want improvements, feel free to open an issue or discussion on the Hugging Face page. or contact me: [email protected]