Downloads · 30 days
1.4K
100% of all-time downloads
tinyopsec/GenPRM-7B-GGUF
GenPRM-7B-GGUF is a text generation model from tinyopsec. Use it when you need the model to write or continue text. It is set up for llama.cpp. The card lists the license as mit.
GGUF quantizations of GenPRM/GenPRM-7B, a generative process reward model for mathematical reasoning.
Downloads · 30 days
1.4K
100% of all-time downloads
All-time downloads
1.4K
Public
Repo size
63.9 GB
Likes
0
Public
Click a slice to open those files.
.gguf63.9 GB · 100%
From the Hugging Face model README
GGUF quantizations of GenPRM/GenPRM-7B, a generative process reward model for mathematical reasoning.
GenPRM-7B is a Qwen2-based process reward model trained from the DeepSeek-R1-Distill-Qwen-7B base model. It performs explicit chain-of-thought reasoning and code verification before producing process judgments. It supports test-time scaling through parallel generation and majority voting, and can be used as a verifier or critic.
| File | Bits | Size | Recommended Memory | Use Case |
|---|---|---|---|---|
model_f16.gguf | 16-bit | 15.2 GB | — | Maximum quality |
model_q8_0.gguf | 8-bit | 8.1 GB | 10+ GB | Near-F16 quality |
model_q6_k.gguf | 6-bit | 6.25 GB | 8+ GB | High quality |
model_q5_k_m.gguf | 5-bit | 5.44 GB | 7+ GB | Quality / size balance |
model_q5_k_s.gguf | 5-bit | 5.32 GB | 7+ GB | Compact 5-bit |
model_q4_k_m.gguf | 4-bit | 4.68 GB | 6+ GB | Recommended general use |
model_q4_k_s.gguf | 4-bit | 4.46 GB | 6+ GB | Compact 4-bit |
model_q3_k_l.gguf | 3-bit | 4.09 GB | 5+ GB | Low-memory use |
model_q3_k_m.gguf | 3-bit | 3.81 GB | 5+ GB | Smaller deployment |
model_q3_k_s.gguf | 3-bit | 3.49 GB | 5+ GB | Maximum compression |
model_q2_k.gguf | 2-bit | 3.02 GB | 4+ GB | Minimum memory |
Actual requirements depend on context length, KV cache, and runtime configuration.
| Bits | Quant | Size |
|---|---|---|
| 2-bit | Q2_K | 3.02 GB |
| 3-bit | Q3_K_S | 3.49 GB |
| 3-bit | Q3_K_M | 3.81 GB |
| 3-bit | Q3_K_L | 4.09 GB |
| 4-bit | Q4_K_S | 4.46 GB |
| 4-bit | Q4_K_M | 4.68 GB |
| 5-bit | Q5_K_S | 5.32 GB |
| 5-bit | Q5_K_M | 5.44 GB |
| 6-bit | Q6_K | 6.25 GB |
| 8-bit | Q8_0 | 8.1 GB |
| 16-bit | F16 | 15.2 GB |
llama-cli -hf tinyopsec/GenPRM-7B-GGUF:Q4_K_M
Server:
llama-server -hf tinyopsec/GenPRM-7B-GGUF:Q4_K_M
from llama_cpp import Llama
llm = Llama(
model_path="model_q4_k_m.gguf",
n_ctx=8192,
)
output = llm(
"Review the following mathematical solution step by step.",
max_tokens=2048,
)
print(output)
Download the desired GGUF file and load it through LM Studio.
ollama run hf.co/tinyopsec/GenPRM-7B-GGUF:Q4_K_M
GenPRM was introduced in GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning. The model uses explicit reasoning and verification for process supervision and supports both verifier and critic applications.
@article{zhao2025genprm,
title = {GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning},
author = {Jian Zhao and Runze Liu and Kaiyan Zhang and Zhimu Zhou and Junqi Gao and Dong Li and Jiafei Lyu and Zhouyi Qian and Biqing Qi and Xiu Li and Bowen Zhou},
journal = {arXiv preprint arXiv:2504.00891},
year = {2025}
}