Downloads · 30 days
0
wexumin/xrag-7b-router
xrag-7b-router is a machine learning model from wexumin. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for compression_router. The card lists the license as apache-2.0.
A routing classifier for xRAG that decides at inference time whether compressed context is sufficient or whether to fall back to the full context.
Downloads · 30 days
0
Access
Public
Updated Jul 20, 2026
Repo size
8.7 MB
Likes
0
Public
Click a slice to open those files.
.pt8.7 MB · 100%
From the Hugging Face model README
A routing classifier for xRAG that decides at inference time whether compressed context is sufficient or whether to fall back to the full context.
Trained with overflowguard on SQuAD using an LLM judge (DeepSeek) for evaluation, with 5-fold stratified CV for threshold selection (Youden's J statistic).
from examples.train_xrag import XragRouter
# Load base model + routing classifier in one call
router = XragRouter.from_pretrained("wexumin/xrag-7b-router")
router.park_gpu()
# Auto-routing: classifier decides compressed vs full
result = router.run_pipeline("long document text...", query="what is X?")
print(result["prediction"]) # answer
print(result["mode"]) # "compressed" or "full"
print(result["tokens_saved"])
router_config.json — base model path + routing thresholdrouting_clf.pt — classifier weights (2-layer MLP, ~2K params)The base xRAG model is loaded automatically from Hannibal046/xrag-7b. This repo only contains the lightweight routing classifier on top.
This model also requires custom code from xRAG repository.