Downloads · 30 days
82
36% of all-time downloads
peakji/steiner-32b-preview
steiner-32b-preview is a machine learning model from peakji. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
For more details, please refer to the announcement blog post.
Downloads · 30 days
82
36% of all-time downloads
All-time downloads
230
Public
Parameters
32.8B
65.5 GB on disk
Likes
91
Public
Click a slice to open those files.
.safetensors65.5 GB · 100%
From the Hugging Face model README
For more details, please refer to the announcement blog post.
Steiner is a series of reasoning models trained on synthetic data using reinforcement learning. These models can explore multiple reasoning paths in an autoregressive manner during inference and autonomously verify or backtrack when necessary, enabling a linear traversal of the implicit search tree.
Steiner is a personal interest project by Yichao 'Peak' Ji, inspired by OpenAI o1. The ultimate goal is to reproduce o1 and validate the inference-time scaling curves. The Steiner-preview model is currently a work-in-progress. The reason for open-sourcing it is that I’ve found automated evaluation methods, primarily based on multiple-choice questions, struggle to fully reflect the progress of reasoning models. In fact, the assumption that "the correct answer is always among the options" doesn’t align well with real-world reasoning scenarios, as it encourages models to perform substitution-based validation rather than open-ended exploration. For this reason, I’ve chosen to open-source these intermediate results and, when time permits, to build in public. This approach allows me to share knowledge while also gathering more evaluations and feedback from real human users.
⚠️ Disclaimer: While Steiner has been able to achieve high-quality zero-shot results without relying on Chain of Thought (CoT) prompting or an agent framework, it has not yet replicated the inference-time scaling capabilities demonstrated by o1. In experiments using a specialized logits processor to intervene on reasoning tokens, increasing the number of reasoning steps did not improve performance; in fact, it led to a decline in benchmarks such as MMLU-Pro and GPQA. As a result, Steiner cannot currently be considered a successful reproduction of OpenAI o1. There may be deficiencies in both the training methods and data quality, so please interpret the results with caution.
Steiner is compatible with all existing inference services, with vLLM being the most recommended for deployment.
Deploying Steiner is no different from using other LLMs; you just need to add the following two parameters to the inference request:
"skip_special_tokens": false,
"spaces_between_special_tokens": false,
For example:
{
"model": "steiner",
"skip_special_tokens": false,
"spaces_between_special_tokens": false,
"messages": [
{
"role": "user",
"content": "Hello"
}
]
}
If you are using the Python client provided by OpenAI, you can use it like this:
stream = client.chat.completions.create(
model="steiner",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
extra_body={
"skip_special_tokens": False,
"spaces_between_special_tokens": False,
},
)
| Subdomain | Accuracy (0-shot w/o CoT) |
|---|---|
| Physics (general) | 63.16% |
| Organic Chemistry | 40.28% |
| Quantum Mechanics | 76.00% |
| Electromagnetism and Photonics | 50.00% |
| High-energy particle physics | 57.14% |
| Genetics | 25.00% |
| Astrophysics | 53.85% |
| Molecular Biology | 80.00% |
| Chemistry (general) | 50.00% |
| Relativistic Mechanics | 57.14% |
| Inorganic Chemistry | 0.00% |
| Optics and Acoustics | 0.00% |
| Condensed Matter Physics | 100.00% |
| All | 53.54% |
If you find my work helpful, please consider citing it in your research or projects. Your acknowledgment would be greatly appreciated!
@misc{ji2024steiner,
title = {A Small Step Towards Reproducing OpenAI o1: Progress Report on the Steiner Open Source Models},
url = {https://medium.com/@peakji/b9a756a00855},
author = {Yichao Ji},
month = {October},
year = {2024}
}