Downloads ยท 30 days
180
12% of all-time downloads
PRIME-RL/P1-30B-A3B
P1-30B-A3B is a text generation model from PRIME-RL. Use it when you need the model to write or continue text. The card lists the license as apache-2.0.
<div align="center" <h1 style="font-size: 2em; font-weight: bold;"P1: Mastering Physics Olympiads with Reinforcement Learning</h1 </div
Downloads ยท 30 days
180
12% of all-time downloads
All-time downloads
1.5K
Public
Parameters
30.5B
61.1 GB on disk
Likes
11
Public
Click a slice to open those files.
.safetensors61.1 GB ยท 100%
From the Hugging Face model README
P1-30B-A3B is the mid-size variant of the P1 series, a high-performance open-source language model specialized in physics reasoning. Built on Qwen3-30B-A3B-Thinking-2507 and refined through multi-stage reinforcement learning on curated physics competition data, P1-30B-A3B achieves impressive results while maintaining reasonable computational requirements, making it accessible for researchers working with physics problems.
| Model | Score | Medal |
|---|---|---|
| P1-30B-A3B | 18.5 | ๐ฅ Silver |
| DeepSeek-R1 | 18.5 | ๐ฅ Silver |
| Qwen3-235B-A22B-Thinking-2507 | 17.1 | ๐ฅ Silver |
| Qwen3-30B-A3B-Thinking-2507 | 15.6 | ๐ฅ Silver |
| Category | P1-30B-A3B | Qwen3-235B-A22B | DeepSeek-R1 | Qwen3-30B-A3B (Base) |
|---|---|---|---|---|
| Overall Score | 32.5 | 33.5 | 32.9 | 29.9 |
| Gold Medals (๐ฅ) | 8 | 10 | 9 | 6 |
| Silver Medals (๐ฅ) | 4 | 3 | 3 | 6 |
| Bronze Medals (๐ฅ) | 1 | 0 | 1 | 1 |
| Total Contests | 13 | 13 | 13 | 13 |
Beyond physics reasoning, P1 improves across multiple domains. As shown below, P1-30B-A3B outperforms its base model Qwen3-30B-A3B-Thinking-2507 on math, coding, and STEM benchmarks, demonstrating strong generalization of physics reasoning.
<div align="center">| Model | AIME24 | AIME25 | HMMT | GPQA | HLE | LiveCodeBench | LiveBench |
|---|---|---|---|---|---|---|---|
| Qwen3-30B-A3B-Thinking-2507 (Base) | 90.4 | 85.0 | 71.3 | 73.0 | 11.6 | 66.7 | 76.6 |
| P1-30B-A3B | 91.0 | 91.0 | 76.9 | 74.4 | 14.3 | 68.1 | 77.0 |
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
# Load model and tokenizer
model_name = "P1-30B-A3B"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True
)
# Physics problem solving
prompt = """Solve this physics problem:
A pendulum of length L = 1.0 m swings with small amplitude.
Calculate the period of oscillation and explain your reasoning.
Use g = 9.8 m/sยฒ"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_length=81920,
temperature=0.6,
top_p=0.9,
do_sample=True
)
solution = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(solution)
We are grateful to the open-source community for their invaluable contributions. Special thanks to:
@misc{p1-2025,
title={P1: Mastering Physics Olympiads with Reinforcement Learning},
author={P1 Team},
year={2025},
url={https://prime-rl.github.io/P1/}
}
</div>