Downloads · 30 days
0
e54true/deeptable-checkpoints-deepseek7b
deeptable-checkpoints-deepseek7b is a table question answering model from e54true. Use it for the table question answering task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as mit.
Trained adapters for "DeepTable: Structural Attention Biases and Tree Path Encoding for Hierarchical Table Understanding." This repo holds the DeepSeek-LLM-7B-Chat checkpoints behind Table 2 of the paper — SAB-only, T…
Downloads · 30 days
0
Access
Public
Updated Sep 3, 2026
Repo size
11.7 GB
Likes
0
Public
Click a slice to open those files.
.safetensors11.5 GB · 93%
From the Hugging Face model README
Trained adapters for "DeepTable: Structural Attention Biases and Tree Path Encoding for Hierarchical Table Understanding." This repo holds the DeepSeek-LLM-7B-Chat checkpoints behind Table 2 of the paper — SAB-only, TPE-only, and the combined DeepTable (SAB+TPE), across all four benchmarks and all four seeds.
Code: <!-- TODO: link to the deeptable_code_submission repo on first push -->
Other backbones: Llama-3-8B-Instruct · Qwen2.5-7B-Instruct
| Benchmark | TableLoRA baseline | SAB only | TPE only | Full (SAB+TPE) |
|---|---|---|---|---|
| HiTab | — | ✓ ×4 seeds | ✓ ×4 seeds | ✓ ×4 seeds |
| WikiTQ | — | ✓ ×4 seeds | ✓ ×4 seeds | ✓ ×4 seeds |
| FeTaQA | — | ✓ ×4 seeds | ✓ ×4 seeds | ✓ ×4 seeds |
| TabFact | — | ✓ ×4 seeds | ✓ ×4 seeds | ✓ ×4 seeds |
No TableLoRA baseline checkpoints. Table 2's DeepSeek and Llama-3 TableLoRA numbers are cited from the original TableLoRA paper (He et al., 2025), not retrained locally — only the Qwen2.5-7B baseline was reproduced for that comparison. A local DeepSeek/TabFact TableLoRA run exists in the authors' internal logs but averages 75.90% across 4 seeds, 1.15pp off the cited 77.05%, so it is not the source of the reported number and is intentionally excluded here to avoid being mistaken for it.
48 checkpoints total (4 benchmarks × 3 variants × 4 seeds).
Each {dataset}_{variant}_seed{n}/ directory is one trained run:
hitab_full_seed0/
├── adapter_config.json ─┐ TableLoRA's [TAB]/[ROW]/[CELL] prompt
├── adapter_model.safetensors ─┘ encoder (PEFT P_TUNING adapter, "default")
├── default_1/
│ ├── adapter_config.json ─┐ the actual 2D-LoRA weights (rank 8,
│ └── adapter_model.safetensors┘ k_proj+v_proj) — PEFT adapter "default_1"
├── sab_module.safetensors Structural Attention Bias (α_row/α_col)
│ — present for "sabonly" and "full" only
├── tpe_modules.safetensors Tree Path Encoding embedding tables
├── tpe_path_vocab.json — path-node vocabulary used at train time
├── tpe_metadata.json — present for "tpeonly" and "full" only
├── tokenizer.json / tokenizer_config.json / special_tokens_map.json
├── train_results.json / all_results.json / trainer_state.json
└── predict/
└── generated_predictions.jsonl — the model's own test-set predictions,
so you can verify a checkpoint's score
without re-running inference
Two stacked PEFT adapters, not one. This is TableLoRA's design, not a
packaging artifact: the top-level directory is a P_TUNING adapter (the
[TAB]/[ROW]/[CELL] prompt encoder), and default_1/ is a separate
LORA adapter (the actual attention-projection weights that do most of the
work). Both are required — loading only one silently drops half the
trained model. Plain peft.PeftModel.from_pretrained() does not know to look
inside default_1/; this repo's overlay patches PeftModel.from_pretrained
to auto-discover every subfolder containing an adapter_config.json (except
checkpoint-*, predict, runs), which is why loading must go through that
patched code path — see below.
Use the DeepTable code repo (setup.sh + this checkpoint as
adapter_name_or_path), the same way README step 5 (Predict) does:
export TABLE_LORA_ENABLED=1 # installs the patched PeftModel.from_pretrained
# that auto-discovers default_1/
python -m tpe_impl.bin.run_tpe_training configs/hitab_tpe.yaml \
# with do_train: false, do_predict: true,
# adapter_name_or_path: <path to this checkpoint dir>
sabonly and tablelora-style checkpoints (no tpe_modules.safetensors) can
launch directly with llamafactory-cli train <yaml> instead — only tpeonly
and full need the run_tpe_training.py wrapper, since that is what reloads
tpe_modules.safetensors via load_tpe_state().
tpe_modules.safetensors / tpe_path_vocab.json / tpe_metadata.json use
this repo's current filenames. load_tpe_state() also accepts the pre-rename
names (hier_modules.safetensors etc.) for any checkpoint you may have
trained yourself before pulling the latest code — see the "Naming" note in
the code repo's README.
Each predict/generated_predictions.jsonl can be scored directly:
python evaluation/eval_accuracy.py <checkpoint>/predict # HiTab
python evaluation/eval_wikitq.py <checkpoint>/predict # WikiTQ
python evaluation/eval_fetaqa_bleu.py <checkpoint>/predict # FeTaQA
python evaluation/eval_tabfact.py <checkpoint>/predict # TabFact
DeepSeek HiTab per-seed accuracies, as a quick sanity check (matches the paper's appendix):
| Config | seed 0 | seed 1 | seed 2 | seed 3 | mean |
|---|---|---|---|---|---|
hitab_sabonly_seed{n} | 53.79 | 51.58 | 51.20 | 51.64 | 52.05 |
hitab_full_seed{n} | 53.28 | 51.01 | 50.44 | 50.06 | 51.20 |
LoRA rank 8 on k_proj,v_proj; base LR 5e-6, cosine, 3 epochs; SAB/TPE
add-on LR multiplier λ=1000 (SAB_LR_MULTIPLIER env var / tpe_lr_multiplier
YAML key). Full details and the exact training YAMLs are in the code repo's
configs/README.md and Appendix G of the paper.
MIT, matching the code repo. See the code repo's LICENSE / NOTICE — the
base model (deepseek-ai/deepseek-llm-7b-chat) and TableLoRA's 2D-LoRA
mechanism these adapters extend carry their own upstream licenses.