Downloads · 30 days
0
alexbalan08/PRM
PRM is a machine learning model from alexbalan08. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A process reward model (PRM) fine-tuned to score whether a proposed next action is correct given a procedural workflow, the history so far, and a list of available actions. Answers Yes/No and at inference we read the…
Downloads · 30 days
0
Access
Public
Updated Jun 24, 2026
Repo size
185 MB
Likes
0
Public
Click a slice to open those files.
.safetensors168 MB · 91%
From the Hugging Face model README
A process reward model (PRM) fine-tuned to score whether a proposed next action is correct given a procedural workflow, the history so far, and a list of available actions. Answers Yes/No and at inference we read the logits to get a continuous score.
The goal is to be combined and wrapped around another LLM, even with a critic in the loop to check for factuality or harmfulness. This is basically a planning agent, which given any history of taken actions, knows what to do next.
It scores candidate actions — it does not generate text, except yes and no tokens for valid or not. it s not a classifier!! Combine it with a frozen weights LLM (as capable as possible), tool use, and you will have a good planner in the loop.
Built as part of a master's thesis on extracting procedural workflows from text and training small planning agents with the data generated by the extractor pipeline. In short, get procedural documents, use the extraction agentic pipeline from my work, create SFT records deterministically and then re-train the PRM, acting as a planner or even can be seen as a critic which approves or not choices of a big LLM. When enough data is generated, re-train the model.