Downloads · 30 days
0
abdelkarim98/tendril-models
tendril-models is a machine learning model from abdelkarim98. Use it for the machine learning task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as mit.
The two small trained models behind TENDRIL, the system in What Does a Reasoning Tree's LLM Budget Buy? A Cost Audit of Multi-Hop Retrieval-Augmented Generation and a Four-Call Alternative.
Downloads · 30 days
0
Access
Public
Updated Sep 12, 2026
Repo size
1.2 GB
Likes
0
Public
Click a slice to open those files.
.safetensors1.2 GB · 98%
From the Hugging Face model README
The two small trained models behind TENDRIL, the system in What Does a Reasoning Tree's LLM Budget Buy? A Cost Audit of Multi-Hop Retrieval-Augmented Generation and a Four-Call Alternative.
The one-line claim these models support: in multi-hop RAG the gap between cheap and expensive systems is retrieval-buyable, not call-buyable. Extra LLM generation buys almost nothing; re-retrieving evidence inside the reasoning chain buys most of the gap, at zero extra LLM calls. These are the models that do the retrieving.
| folder | params | size on disk | role |
|---|---|---|---|
rtrag_builder_v1/ | 184M | 712 MB | the Builder — a DeBERTa-v3-base cross-encoder used three times over: scoring the initial pool, scoring newly retrieved chunks, and rescoring candidates conditioned on partial chains. Also does per-hop re-retrieval at inference. |
rtrag_decomposer_t5/ | 247M | 476 MB | the planner — a flan-T5-base fine-tuned to split a multi-hop question into at most four single-hop sub-questions with #1–#3 placeholders. Used only by the all-small TENDRIL-S configuration. |
rtrag_builder_v1/ config.json model.safetensors head.pt report.json
added_tokens.json special_tokens_map.json
spm.model tokenizer.json tokenizer_config.json
rtrag_decomposer_t5/ config.json generation_config.json model.safetensors
special_tokens_map.json spiece.model
tokenizer.json tokenizer_config.json
head.pt is the Builder's linear scoring head and is required — the
cross-encoder is a base encoder plus that head, not a SequenceClassification
checkpoint you can load standalone.
git clone https://github.com/abdelkarim-choukri/TENDRIL-RAG && cd TENDRIL-RAG
bash scripts/fetch_weights.sh
That places them exactly where the code looks for them:
checkpoints/rtrag_builder_v1/
checkpoints/rtrag_decomposer_t5/
which are the defaults of src/rtrag_decomp_probe.py --builder,
src/rtrag_probe2_rescore.py --model and configs/tendril*.json. Or directly:
from huggingface_hub import snapshot_download
snapshot_download("abdelkarim98/tendril-models",
allow_patterns="rtrag_builder_v1/*",
local_dir="checkpoints", local_dir_use_symlinks=False)
In scope. Research on retrieval quality for multi-hop question answering: reranking, chain-aware evidence assembly, and question decomposition, in the LongBench v1 / HotpotQA / 2WikiMultihopQA / MuSiQue setting these were trained and measured on. The Builder is a reranker — it scores a (query, chunk) pair; it does not generate text. The planner rewrites a question into sub-questions; it does not answer them.
Out of scope. Neither model answers questions, checks facts, or judges truth. The Builder's score is a relevance score, not a correctness signal — the paper reports a separate negative result showing chain completeness is not readable from this score geometry (a 22-feature gate reaches only AUC 0.73–0.75, and 0.60 on MuSiQue, near chance). Do not use these as a filter for factual reliability. Both were trained on English Wikipedia-derived corpora and carry whatever biases those contain.
The Builder. Supervision requires no LLM at all. 6,000 questions per dataset from the public training-side splits, paragraphs pooled into an RT-RAG-style corpus and chunked with RT-RAG's chunker; a chunk is gold if and only if its source document is one of that question's own supporting paragraphs — document identity, not title match, which on MuSiQue would introduce false positives. A second family of pools replays round 1 of the deployment loop, which is what teaches the model to score conditioned queries. Together 21,471 listwise groups, trained with a group softmax.
Hyper-parameters, read from the shipped rtrag_builder_v1/report.json:
base_model deberta-v3-base-squad2 epochs 2 (epoch 0 selected)
lr 1e-05 batch_groups 2
max_len 352 seed 42
The planner. Two thirds of its supervision is free and gold: MuSiQue
publishes question decompositions (~19.5k) and 2Wiki's evidence triples
template into sub-questions (~12.0k). HotpotQA publishes none, so for that
third (~9k) the 14B's own planning outputs are distilled offline. This is the
one place a large model is used in training anything here, it is a one-off
pass over ~9k training questions rather than over a corpus, and it applies to
rtrag_decomposer_t5 only.
All 200 LongBench evaluation ids and the 400-question development ids are excluded from every training source.
report.json, shipped in the checkpoint)| split | metric | HotpotQA | 2Wiki | MuSiQue |
|---|---|---|---|---|
| plain (n=400 each) | full-chain@15 | 0.835 | 0.465 | 0.4675 |
| plain (n=400 each) | anchor@15 | 1.000 | 1.000 | 0.990 |
| conditioned | top-3 hit | 0.9459 | 0.9447 | 0.7802 |
| conditioned | top-1 hit | 0.8108 | 0.8894 | 0.6154 |
| conditioned | n | 37 | 199 | 91 |
These are single-pass validation numbers on training-side pools, which is what the checkpoint's own record contains. They are not the deployed front-end result and should not be quoted as it — the two conditioned splits in particular are small (n = 37 on HotpotQA).
The full Stage 0 pipeline — bge retrieval, two rounds of anchor-conditioned
re-querying, the cond2 rescore and the seating rule — reaches full-chain
coverage of 0.940 / 0.870 / 0.655, which is +6.0 / +14.0 / +12.5 points
over RT-RAG's own front end, at zero LLM calls on both sides.
End to end, with a Qwen2.5-14B-Instruct reader and RT-RAG's own evaulate.py
scorer (EM):
| system | LLM calls/q | HotpotQA | 2Wiki | MuSiQue |
|---|---|---|---|---|
| TENDRIL-1 (front end + 1 call) | 1 | 47.0 | 49.0 | 27.5 |
| TENDRIL-S (planner from this repo) | 3.45 | 47.0 | 58.0 | 31.0 |
| TENDRIL | 4.31 | 48.0 | 64.0 | 39.0 |
| RT-RAG reasoning tree (our reproduction) | 47.9–103.4 | 51.0 | 65.5 | 38.0 |
A paired per-question sign test cannot distinguish TENDRIL from the reproduced tree on any dataset, in exact match (p = 0.41 / 0.73 / 0.87) or under a blind LLM judge (p = 0.50 / 0.60 / 0.18).
No F1 parity is claimed. Pooled over n = 600 the F1 deficit is statistically detectable (−3.33, p = 0.039). And although TENDRIL's MuSiQue EM exceeds the tree's, both F1 and the judge favour the tree there, so no MuSiQue win is claimed either. A high p-value is evidence of no difference detected at this sample size, never evidence that two systems are equal.
inference
(n = 26).These two checkpoints are the artifacts that produced the paper's numbers. Note honestly: no SHA-256 of either was recorded in the project's own documentation before publication here. The hashes below were computed at upload time (2026-09-08) and are the first written record.
| file | sha256 (first 16) |
|---|---|
rtrag_builder_v1/model.safetensors | 4656f79df348dec0 |
rtrag_builder_v1/head.pt | fa5a87b78ac850fc |
rtrag_decomposer_t5/model.safetensors | ab5e5a391994245b |
The code in the GitHub repository is byte-identical to the code that produced
the results, and that is recorded — see PROVENANCE.md there.
@article{choukri2026tendril,
title = {What Does a Reasoning Tree's {LLM} Budget Buy? A Cost Audit of
Multi-Hop Retrieval-Augmented Generation and a Four-Call Alternative},
author = {Choukri, Abdelkarim and Deng, Xiang},
year = {2026},
note = {arXiv preprint}
}
Abdelkarim Choukri · Xiang Deng — School of Computer Science and Technology,
Harbin Institute of Technology, Shenzhen.
[email protected] · [email protected] (permanent backup).
RT-RAG's authors released their code, which is the only reason a cost audit of it was possible. RT-RAG's code is not redistributed here or in the GitHub repository; it is cloned from its own repository under its own licence.