Downloads · 30 days
9
4% of all-time downloads
cs-552-2026-group1/math_model
math_model is a text generation model from cs-552-2026-group1. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
This model is a fine-tuned version of Qwen/Qwen3-1.7B for mathematical reasoning, developed for the CS-552 standard project math track.
Downloads · 30 days
9
4% of all-time downloads
All-time downloads
252
Public
Parameters
1.7B
10.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors3.4 GB · 99%
From the Hugging Face model README
This model is a fine-tuned version of Qwen/Qwen3-1.7B for mathematical reasoning, developed for the CS-552 standard project math track.
The model was trained to solve competition-style mathematics problems and produce final answers in boxed LaTeX format.
Qwen/Qwen3-1.7BThe final submitted checkpoint was trained on approximately 25,165 examples from hard mathematical reasoning datasets:
| Dataset / Source | Examples |
|---|---|
| Hendrycks MATH | 4,759 |
| OpenR1-Math-220k | 7,999 |
| NuminaMath-CoT | 12,407 |
| Total | 25,165 |
The NuminaMath subset was filtered to focus on harder mathematical sources:
The OpenR1 subset was filtered to competition-relevant categories:
| Setting | Value |
|---|---|
| Base model | Qwen/Qwen3-1.7B |
| Fine-tuning method | LoRA |
| Epochs | 1 |
| LoRA rank | 32 |
| LoRA alpha | 64 |
| LoRA dropout | 0.05 |
| Learning rate | 1e-4 |
| Batch size | 1 |
| Gradient accumulation steps | 8 |
| Precision | bfloat16 |
| Hardware | 1 × NVIDIA A100 40GB |
| Training steps | 3,146 |
| Runtime | 6,761 seconds |
| Tokens processed | 14.7M |
| Final training loss | 0.6114 |
| Mean token accuracy | 0.8302 |
The submitted generation configuration uses sampling:
{
"do_sample": true,
"temperature": 0.25,
"top_p": 0.85,
"top_k": 30
}