Downloads · 30 days
43
23% of all-time downloads
unstoppableljq/RadSight-4B
RadSight-4B is a image-text-to-text model from unstoppableljq. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
<h2 align="center"<bRadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding</b</h2
Downloads · 30 days
43
23% of all-time downloads
All-time downloads
187
Public
Parameters
4.9B
10.6 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors10.6 GB · 100%
From the Hugging Face model README
Research use only. RadSight is not a medical device and must not be used for autonomous diagnosis, triage, treatment decisions, or patient-facing medical advice. Model outputs may be incomplete, incorrect, or hallucinated and must be reviewed by qualified healthcare professionals.
RadSight is a perception-driven medical multimodal large language model (MLLM) for unified understanding of 2D radiology images and native 3D CT volumes. It uses modality-specific 2D and 3D visual encoders and a shared language-model interface to support tasks ranging from fine-grained visual perception to clinical diagnosis and radiology report generation.
RadSight is trained with a four-stage progressive curriculum:
The model is designed to learn explicit visual evidence—such as lesion attributes and spatial correspondence—before producing higher-level diagnostic or report-level outputs.
| Model | Language backbone | Supported visual inputs | Checkpoint precision |
|---|---|---|---|
| RadSight-4B | Qwen3-VL-4B | 2D images and 3D CT volumes | BF16 |
| RadSight-8B | Qwen3-VL-8B | 2D images and 3D CT volumes | BF16 |
The 4B and 8B names refer to the language-backbone scale. The complete
multimodal checkpoints also include visual encoders and projectors; therefore,
the total parameter counts displayed by the Hugging Face interface may be
larger than the variant names.
The released checkpoints use custom RadSight model and preprocessing code. They are not intended to be loaded as a standard text-only Transformers model.
git clone https://github.com/alibaba-damo-academy/damo-RadSight.git
cd damo-RadSight
conda create -n radsight python=3.10 -y
conda activate radsight
pip install torch==2.7.0 torchvision==0.22.0 \
--index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
pip install flash-attn --no-build-isolation
CUDA 12 or later is recommended. Building Flash Attention 2 requires a CUDA toolkit compatible with the installed PyTorch version.
# Choose one RadSight variant.
hf download unstoppableljq/RadSight-4B \
--local-dir ./weights/RadSight-4B
# For RadSight-8B, use:
# hf download unstoppableljq/RadSight-8B \
# --local-dir ./weights/RadSight-8B
hf download DAMO-NLP-SG/VL3-SigLIP-NaViT \
--local-dir ./weights/VL3-SigLIP-NaViT
Before loading the checkpoint, edit its config.json and set
vision_encoder to the local SigLIP-NaViT directory:
{
"vision_encoder": "/absolute/path/to/weights/VL3-SigLIP-NaViT"
}
This step is required because the released configuration may contain an environment-specific path. Using a local path also prevents unexpected network access while the model is being loaded.
The following example supports both 2D images and 3D CT volumes. Set modal,
visual_input, and the conversation placeholder consistently.
import torch
from radsight.model import load_pretrained_model
from radsight.mm_utils import (
get_model_name_from_path,
load_3D,
load_images,
)
from radsight.model.processor import RadSightProcessor
model_path = "./weights/RadSight-4B"
model_name = get_model_name_from_path(model_path)
tokenizer, model, image_processor, context_len = load_pretrained_model(
model_path,
None,
model_name,
device_map={"": "cuda:0"},
)
processor = RadSightProcessor(image_processor, tokenizer)
model.config.use_token_compression = False
# ----- Option A: 2D image -----
modal = "image"
visual_input = load_images("./example_xray.jpg")
conversation = [
{
"role": "user",
"content": [
{"type": "image"},
{
"type": "text",
"text": "Please generate a radiology report for this image.",
},
],
}
]
# ----- Option B: 3D CT volume -----
# modal = "volume"
# visual_input = load_3D("./example_ct.nii.gz")["image"]
# conversation = [
# {
# "role": "user",
# "content": [
# {"type": "video", "num_frames": 12},
# {
# "type": "text",
# "text": "Please generate a radiology report for this CT scan.",
# },
# ],
# }
# ]
inputs = processor(
images=[visual_input],
text=conversation,
merge_size=1,
modal=modal,
return_tensors="pt",
)
inputs = {
key: value.cuda() if isinstance(value, torch.Tensor) else value
for key, value in inputs.items()
}
if "pixel_values" in inputs:
inputs["pixel_values"] = inputs["pixel_values"].to(torch.bfloat16)
with torch.inference_mode():
output_ids = model.generate(
**inputs,
do_sample=False,
modals=[modal],
max_new_tokens=8192,
use_cache=True,
pad_token_id=tokenizer.eos_token_id,
)
output = tokenizer.batch_decode(
output_ids,
skip_special_tokens=True,
)[0].strip()
print(output)
For reproducible evaluation, use deterministic decoding (do_sample=False) and
the same prompts, preprocessing, and output parsing rules as the corresponding
benchmark.
For 2D images, the default training resolution is 448 × 448.
For 3D CT volumes:
1 × 1 × 5 mm spacing;[-1000, 1000] and normalized to [0, 1];96 × 256 × 384; andDifferences in preprocessing can substantially affect spatial grounding, anomaly detection, and report-generation results.
| Stage | Objective | Main supervision |
|---|---|---|
| 1 | Visual-language alignment | 2D/3D image-report pairs |
| 2 | Fine-grained visual perception | Attribute judgment, spatial grounding, and spatial understanding |
| 3 | Clinical diagnosis | Disease prediction and abnormality detection |
| 4 | Radiology report generation | Chest X-ray and CT reports |
The checkpoints were trained in BF16 with a progressive curriculum. The training implementation uses PyTorch, Transformers, DeepSpeed, Flash Attention 2, gradient checkpointing, and distributed training. Refer to the project repository and paper for complete stage-specific hyperparameters.
If you use RadSight in your research, please cite:
@article{liu2026radsight,
title = {RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding},
author = {Liu, Jianqin and Cao, Weiwei and Chang, Wanxing and Yuan, Ruifeng
and Shi, Bowen and Zheng, Zhilin and Zhang, Xianjie and Zhang, Ling
and Wang, Peng and Zhang, Jianpeng},
journal = {Preprint},
year = {2026}
}
The paper URL and bibliographic record will be updated after the preprint is publicly available.
RadSight builds on open-source work including VideoLLaMA3, Qwen3-VL, MONAI, nnU-Net, TotalSegmentator, and RADAR. We thank the creators and maintainers of the public medical datasets used in this research.
For research questions, please contact [email protected].