Downloads ยท 30 days
0
sid172002/deepseek-math-7b-rl-phase1
deepseek-math-7b-rl-phase1 is a machine learning model from sid172002. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Mathematical Reasoning Model - Phase 1 Training Complete
Downloads ยท 30 days
0
Access
Public
Updated Feb 20, 2026
Repo size
600 MB
Likes
0
Public
Click a slice to open those files.
.safetensors600 MB ยท 99%
From the Hugging Face model README
Mathematical Reasoning Model - Phase 1 Training Complete
| Attribute | Value |
|---|---|
| Base Model | sid172002/deepseek-math-7b-rl-5500steps |
| Training Type | LoRA Fine-tuning (r=64, alpha=128) |
| Dataset | 379,921 international math problems |
| Training Duration | 15.3 hours |
| Epochs | 3 |
| Final Loss | 0.46 (started at 0.59) |
| Hardware | NVIDIA B200 (180GB) |
| Tier | Score | Accuracy | Notes |
|---|---|---|---|
| IIT JEE Easy | 1/2 | 50.0% | Basic algebra/calculus |
| IIT JEE Hard | 1/2 | 50.0% | Advanced problems |
| AMC 10/12 | 1/2 | 50.0% | Competition math |
| AIME | 1/2 | 50.0% | Hard competition |
| Olympiad | 1/2 | 50.0% | Proof-based |
| FrontierMath | 0/2 | 0.0% | Very hard geometry/calculus |
| Source | Count | Type |
|---|---|---|
| NuminaMath-Olympiad | 125,000 | Competition |
| NuminaMath-AMC | 85,000 | Competition |
| NuminaMath-AIME | 45,000 | Competition |
| NuminaMath-AoPS | 99,921 | Olympiad |
| JEEBench | 515 | IIT JEE |
| MetaMathQA | 5,000 | Algebra |
| GSM8K | 5,000 | Basic math |
| India Context | 10,000 | Regional |
| Singapore Math | 5,000 | Regional |
| OpenWebMath | 5,000 | Calculus |
Base: DeepSeek-Math-7B (5,500 steps pre-trained)
โ
LoRA Fine-tuning
- Rank: 64
- Alpha: 128
- Target: All attention + MLP layers
- Trainable params: 149.9M (2.12% of 7.06B)
โ
Phase 1 Output: Text-only model
Batch size: 16 (per device)
Gradient accumulation: 4
Effective batch: 64
Learning rate: 1e-4 โ 3.8e-10 (cosine decay)
Optimizer: AdamW 8-bit
Max sequence length: 4096
Precision: bfloat16
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "sid172002/deepseek-math-7b-rl-phase1"
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype=torch.bfloat16,
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_path)
# Inference
problem = "Find the sum of 1 + 2 + ... + 100"
prompt = f"### Problem: {problem}\n### Solution:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.3)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
Next step: Add vision capabilities for geometry problems
Phase 1 Output (Text)
โ
+ CLIP Vision Encoder (frozen)
+ Projection Layer (trainable)
+ 5,000 Vision Problems
โ
Phase 2 Output (Multimodal)
Estimated improvement: +10-15% on geometry/competition problems
| Phase | Duration | Cost (B200 @ $5.29/hr) |
|---|---|---|
| Phase 1 | 15.3 hours | $81.01 |
| Phase 2 (est.) | 6 hours | ~$32 |
| Total | ~21 hours | ~$113 |
deepseek-math-phase1-final/
โโโ final/
โ โโโ adapter_model.safetensors (572 MB)
โ โโโ adapter_config.json
โ โโโ tokenizer.json
โ โโโ README.md
โโโ checkpoint-15000/
โโโ checkpoint-16000/
โโโ checkpoint-17000/
@misc{deepseek-math-phase1,
title={DeepSeek-Math-7B-RL-Phase1: Fine-tuned on 379K International Math Problems},
author={sid172002},
year={2026},
howpublished={HuggingFace Model Hub}
}
Status: โ Phase 1 Complete | โณ Phase 2 Ready | ๐ฏ Benchmarked