Downloads · 30 days
0
smilelab/REVEAL
REVEAL is a machine learning model from smilelab. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
REVEAL is a multimodal vision-language model designed to align retinal fundus imaging with individualized clinical risk factors for early prediction of Alzheimer’s disease (AD) and dementia. The model learns joint rep…
Downloads · 30 days
0
Access
Public
Updated Feb 27, 2026
Repo size
5.1 GB
Likes
0
Public
Click a slice to open those files.
.pt5.1 GB · 100%
From the Hugging Face model README
REVEAL is a multimodal vision-language model designed to align retinal fundus imaging with individualized clinical risk factors for early prediction of Alzheimer’s disease (AD) and dementia. The model learns joint representations from retinal morphology and structured health data transformed into clinical narratives.
REVEAL leverages pretrained medical foundation models and introduces a group-aware contrastive learning (GACL) strategy to capture clinically meaningful multimodal relationships. The model is designed to support early disease risk stratification and multimodal biomarker discovery.
REVEAL is composed of:
The framework operates in two stages:
The model was trained using multimodal data derived from the UK Biobank (https://www.ukbiobank.ac.uk/), a large population-scale biomedical dataset containing retinal imaging and clinical health variables.
The dataset includes color fundus photographs and clinical risk factor data from 39,242 participants:
Training and validation sets contained only cognitively normal participants at baseline. Individuals who developed incident AD or dementia were reserved for downstream evaluation.
Retinal morphometric features were extracted using the AutoMorph pipeline, including:
Risk factors include:
Structured clinical variables were converted into standardized clinical narratives using a large language model. Each participant’s risk factors were mapped into a predefined clinical template to enable compatibility with vision-language training.
REVEAL aligns fundus images and clinical narratives using contrastive vision-language learning. Both modalities are encoded and projected into a shared latent embedding space.
REVEAL introduces a group-aware pairing strategy that:
This enables the model to learn clinically meaningful multimodal relationships rather than relying only on subject-level pairings.
REVEAL uses a modified contrastive loss supporting multiple positive pairs per sample. Similarity is computed using cosine similarity between image and text embeddings.
Hyperparameters were optimized using Optuna (https://optuna.org/).
REVEAL is intended for research applications, including:
The model should be used:
The model is not intended for:
REVEAL embeddings were evaluated using downstream support vector machine classifiers.
Performance reflects average results across multiple random seeds.
If you use this model, please cite:
@article{leem2026reveal, title={REVEAL: Multimodal Vision-Language Alignment of Retinal Morphometry and Clinical Risks for Incident AD and Dementia Prediction}, author={Leem, Seowung and Gu, Lin and You, Chenyu and Gong, Kuang and Fang, Ruogu}, journal={MIDL 2026 (Accepted)}, year={2026} }