Downloads · 30 days
72
15% of all-time downloads
polygramme/Rolodex-12B
Rolodex-12B is a text generation model from polygramme. Use it when you need the model to write or continue text. It is set up for transformers. The card lists the license as mit.
Rolodex-12B (12B active / 106B total MoE parameters, GLM-4.5-Air derivative) is a search model: fine-tuned for agentic, tool-driven retrieval over personal and team knowledge bases — multi-step search, memory lookup,…
Downloads · 30 days
72
15% of all-time downloads
All-time downloads
478
Public
Parameters
110B
221 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors221 GB · 100%
How the weights are stored.
BF16110B · 100%
From the Hugging Face model README
Rolodex-12B (12B active / 106B total MoE parameters, GLM-4.5-Air derivative) is a search model: fine-tuned for agentic, tool-driven retrieval over personal and team knowledge bases — multi-step search, memory lookup, and evidence synthesis with tool use.
This repo contains the fully merged weights, plus the LoRA adapter under
adapter/. Because the base is the public GLM-4.5-Air, the adapter can also be
applied directly to the base model (see adapter/README.md, including the
router-bias step).
Rolodex-12B is also the stage-1 base of PolyClerk-12B, our legal work-product model.
| Method | OAPL, LoRA r=32 / α=64 (merged) |
| Framework | ms-swift (Megatron backend) |
| Epochs | 1 |
On internal short-factoid retrieval QA over indexed knowledge bases, the OAPL-trained model improves materially over the GLM-4.5-Air base while using single-tree guided search (fewer tool calls than majority-vote baselines). Internal numbers; independent benchmarks pending.
Requires ~200GB of weights (bf16). Serve with vLLM:
vllm serve polygramme/Rolodex-12B --tensor-parallel-size 4 --max-model-len 131072
The chat template is included (chat_template.jinja). The model is trained for
tool-use loops (search/read tools) — it performs best inside an agentic harness
that exposes retrieval tools, rather than as a plain chat model.
Intended for retrieval-augmented and memory-augmented agent workloads. It may underperform general instruction models on open-ended chat. Verify retrieved claims against sources; the model can still hallucinate under sparse retrieval.
MIT, following the GLM-4.5-Air base license. © the model authors.