Downloads · 30 days
0
gevaertlab/diffusiongemma-radiology-vqa
diffusiongemma-radiology-vqa is a image-text-to-text model from gevaertlab. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.
This repository contains LoRA finetunes of DiffusionGemma (image-conditioned discrete-diffusion LLM) for radiology visual question answering, each paired with an autoregressive Gemma-4 finetune as a controlled baselin…
Downloads · 30 days
0
Access
Public
Updated Jul 4, 2026
Repo size
2.7 GB
Likes
3
Public
Click a slice to open those files.
.safetensors2.7 GB · 100%
From the Hugging Face model README
This repository contains LoRA finetunes of DiffusionGemma (image-conditioned discrete-diffusion LLM) for radiology visual question answering, each paired with an autoregressive Gemma-4 finetune as a controlled baseline. It corresponds to the paper Discrete Diffusion Language Models for Interactive Radiology Report Drafting.
The dataset covers mixed modalities/anatomy (VQA-RAD, SLAKE, VQA-Med: X-ray/CT/MRI, head/chest/abdomen). Judge-best checkpoint per cell.
Code: https://github.com/mxvp/discrete_diffusion_RRG
| subfolder | backbone | base model | dataset | LLM-judge acc |
|---|---|---|---|---|
| diffusion-vqarad | discrete-diffusion | google/diffusiongemma-26B-A4B-it | VQA-RAD | 0.649 |
| ar-vqarad | autoregressive | google/gemma-4-26B-A4B-it | VQA-RAD | 0.649 |
| diffusion-slake | discrete-diffusion | google/diffusiongemma-26B-A4B-it | SLAKE | 0.863 |
| ar-slake | autoregressive | google/gemma-4-26B-A4B-it | SLAKE | 0.817 |
| diffusion-vqamed | discrete-diffusion | google/diffusiongemma-26B-A4B-it | VQA-Med | 0.666 |
| ar-vqamed | autoregressive | google/gemma-4-26B-A4B-it | VQA-Med | 0.631 |