Downloads · 30 days
0
fancc28/Medical-OpenJev
Medical-OpenJev is a machine learning model from fancc28. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Open, locally-deployable, calibrated medical decision gating for three condition tiers. A non-autoregressive System-1 model that reads a patient evidence state and, in one forward pass, returns specialty routing + evi…
Downloads · 30 days
0
Access
Public
Updated Sep 22, 2026
Repo size
901 MB
Likes
0
Public
Click a slice to open those files.
.safetensors901 MB · 99%
From the Hugging Face model README
Open, locally-deployable, calibrated medical decision gating for three condition tiers. A non-autoregressive System-1 model that reads a patient evidence state and, in one forward pass, returns specialty routing + evidence-sufficiency probability + calibrated confidence + act / escalate decision.
This repository ships three independent models, one per condition tier — a difficulty gradient from well-specified skin disease to long-tail rare disease:
| Sub-folder | Condition tier | Examples | Character |
|---|---|---|---|
derm/ | Dermatology (skin) | Hailey-Hailey disease, Conradi-Hünermann syndrome | specific, visually-grounded history |
agentclinic/ | Multi-system / acute internal | Lichen sclerosus, Neuroleptic malignant syndrome | broad, open-ended clinical workup |
rdc/ | Rare disease (Orphanet) | Hypochondroplasia, Loeffler endocarditis | long-tail, low-prevalence, hardest |
Each is the open counterpart to closed proprietary "System-1 decision" APIs (e.g. TypeSafe Jev). It answers the question a generative medical LLM answers dishonestly:
"Given everything gathered so far, is it actually safe to act — and with what probability?"
It never generates text. Output is strictly probabilities/numbers → cannot hallucinate a diagnosis, cannot emit malformed JSON, and its confidence is mathematically calibrated (RLCD proper-scoring training), not tokens that merely "sound confident".
This gate is NOT attached after an LLM. It is a standalone encoder that runs before / around any general LLM doctor as a traffic light:
evidence state ──► Medical-OpenJev (System 1, ~ms, local)
├─ decision = act → trust the fast answer (skip the big LLM)
└─ decision = escalate → only THEN call the general LLM doctor (System 2)
Plug-and-play with any LLM — it never touches the LLM's weights.
| Primitive | Question | Output |
|---|---|---|
| choice | Which specialty should handle this case? | label + per-option probabilities + confidence |
| noul | Is the evidence so far sufficient to conclude? | calibrated P(sufficient) ∈ [0,1] |
| act / escalate | Trust the fast path or escalate to deep reasoning? | P(act), P(escalate) + thresholded decision |
confidence = 1 − H(p)/log(K). decision = act iff P(act) ≥ act_threshold (per-model threshold in medgate_config.json).
answerdotai/ModernBERT-base — 141M params, fully bidirectional, RoPE, 8192 ctx.[MASK] hidden state → Linear(768→256) → GELU → Linear(256→1) → softmax within a question.[CLS] ⊕ 4 distribution features (top1, top1−top2, normalized entropy, K/255) → MLP → [P(act), P(escalate)].Each model: ~149.6M params, single model.safetensors (backbone bf16 + choice/noul heads fp32), ~300 MB.
Sequence template: "{state} ||| [MASK] opt1 | [MASK] opt2 | ..." (state facts joined by " | ").
pip install torch transformers safetensors
huggingface-cli download TODO/medical-openjev --local-dir ./medical-openjev
from modeling_medopenjev import MedOpenJev # file at repo root
# pick the condition tier you need
model = MedOpenJev.from_pretrained("./medical-openjev/derm", device="cuda") # or agentclinic / rdc, or "cpu"
facts = [
"The patient is a 35-year-old woman.",
"She is referred for a rash involving both axillae.",
"She reports recurring episodes since her early 20s.",
"Lesions develop and involute spontaneously.",
]
state = model.state_from_facts(facts)
r = model.predict(state, domain="derm", qtype="choice") # 1) specialty routing
print(r["choice"], r["confidence"]) # -> Dermatology ...
g = model.predict(state, domain="derm", qtype="noul") # 2) evidence-sufficiency gate
print(g["p_true"], g["decision"]) # -> P(sufficient), act|escalate
Use it as a safety gate:
g = model.predict(state, domain="rdc", qtype="noul")
if g["decision"] == "act" and g["confidence"] >= 0.85:
answer = fast_path(state) # trust System 1
else:
answer = general_llm_doctor(state) # escalate to a big LLM
Held-out test (n=60 per condition), gate confidence compared head-to-head against a doctor LLM's self-reported confidence on the same cases. Lower ECE / Brier is better.
| Condition tier | Gate ECE ↓ | LLM ECE ↓ | Gate Brier ↓ | LLM Brier ↓ |
|---|---|---|---|---|
Dermatology (derm) | 0.209 | 0.523 | 0.238 | 0.505 |
Multi-system (agentclinic) | 0.135 | 0.351 | 0.239 | 0.299 |
Rare disease (rdc) | 0.220 | 0.706 | 0.068 | 0.644 |
Across all three tiers the gate is ~2–3× better calibrated than letting the LLM grade its own confidence — the core reason to route through a System-1 gate rather than trusting a big model's gut feeling.
Note: end-to-end diagnostic success rate is low on rare disease (hard long-tail); the gate's job is honest calibration, not solving the diagnosis — a high
escalaterate onrdcis the desired, safe behavior.
Use for
Do NOT use for
derm/choice: its act_threshold fit degenerates to 1.0 (always-escalate) on the tiny val split — use derm/noul
and the other four heads normally; a unified retrain is coming.Roadmap: a unified single model (all three tiers + both primitives in one weight set, one forward pass for multiple typed questions) is in progress and will supersede the three separate tiers.
medical-openjev/
├── modeling_medopenjev.py # shared loader + predict
├── derm/ model.safetensors · config.json · medgate_config.json · tokenizer.*
├── agentclinic/ ...
└── rdc/ ...
@software{medicalopenjev2026,
title = {Medical-OpenJev: Open, Locally-Deployable, Calibrated Medical Decision Gating},
author = {Ruifan Zuo, Guocheng Hu, Qichao Zhao, Ziyang Meng, Keyv Liu, Zichao Dai, Rui Wang, Tian Gan},
year = {2026},
url = {https://github.com/vidahi/medical-openjev}
}
Apache-2.0. Backbone weights are from answerdotai/ModernBERT-base (its license terms apply to those weights).
No patient data is redistributed — only trained decision weights.