Downloads · 30 days
0
2001jdev/crossgat-trial-eligibility-full
crossgat-trial-eligibility-full is a text classification model from 2001jdev. Use it when you need a label for a piece of text. It is set up for pytorch. The card lists the license as cc-by-nc-4.0.
The deployment checkpoint: a 3-class eligibility model that reranks clinical-trial retrieval candidates, refit on every labelled pair after cross-validation.
Downloads · 30 days
0
Access
Public
Updated Aug 27, 2026
Repo size
28.6 MB
Likes
0
Public
Click a slice to open those files.
.ckpt28.6 MB · 100%
From the Hugging Face model README
The deployment checkpoint: a 3-class eligibility model that reranks clinical-trial retrieval candidates, refit on every labelled pair after cross-validation.
This is the standard final step after CV — cross-validation validated the procedure and its hyperparameters, and the shipped model applies them to all the data. If you want to reproduce reported numbers, or to ensemble, use the 5-fold CV checkpoints instead.
It has seen every labelled pair in the collection, so it has no held-out data left and no unbiased performance estimate. Any number you compute on TREC CT 2021–23 with this model is contaminated.
Cite the cross-validated estimate instead, from the CV repo: graded NDCG@10 0.5973 → 0.6515 (+0.0542, p<0.0001, 162 topics) on hybrid BM25+gemma-med at the recommended operating point. That is the expected performance of this procedure; this checkpoint is the procedure applied to more data.
It takes parsed eligibility graphs, not raw text. Producing them needs the generator from the project repo plus an LLM endpoint (~38 GB VRAM). The graphs for all 13,229 TREC CT trials are published separately — clinical-trials-eligibility-graphs-rerank — so reproduction needs no LLM.
score = (1 - w) * retrieval_norm + w * p_eligible_norm with w = 0.5
Both terms min-max normalised within a topic. Never replace the first stage outright (w = 1.0): that loses significance on both retrievers tested. The model reranks; it cannot retrieve — per-topic AUC for relevant-vs-rest is 0.533, chance.
With no validation split there is nothing to select on, so the epoch count was
fixed in advance from the CV runs. Averaging the per-fold validation curves
(not their argmaxes, which are noise around a plateau) puts val_ndcg10's optimum
at epoch 3:
| epoch | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|---|
| mean val_ndcg10 | .602 | .623 | .615 | .624 | .623 | .616 | .611 | .600 | .607 |
The curve is flat from epoch 1 to 5, so the exact choice inside that band is not critical — which is the point. The median of the per-fold argmaxes would have said epoch 2, but that averages noise rather than signal.
Refit epochs carry ~26% more gradient steps than a CV fold (1,633 vs ~1,294 batches), so epoch 3 here is ≈3.8 CV-equivalent epochs — still inside the plateau.
pip install "crossgat @ git+https://github.com/JDev2001/msc_v2"
from crossgat_hf import load, score, rerank_score
pipe = load("2001jdev/crossgat-trial-eligibility-full") # fold=None -> model.ckpt
out = score(pipe, patient_graph, inc_graph, exc_graph,
patient_text="65yo male, type 2 diabetes, HbA1c 8.1% ...",
trial_text="Metformin in adults with type 2 diabetes ...")
out["class_probs"] # {'irrelevant': .., 'excluded': .., 'eligible': ..}
rerank_score(out) # p_eligible — fuse this with your retrieval score
patient_text and trial_text are required: the text branch is part of the
model, and omitting them raises rather than silently scoring an all-zero branch.
abhinand/MedEmbed-base-v0.1Trained on TREC Clinical Trials track data (2021–23). Users must comply with the TREC data usage agreements for the underlying collection and relevance judgments.
Checkpoint SHA-256 prefix in checkpoint_manifest.json.