Downloads · 30 days
9
15% of all-time downloads
ifms111/UniReason-Med
UniReason-Med is a image-text-to-text model from ifms111. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as apache-2.0.
UniReason-Med is a medical multimodal model for grounded reasoning over 2D medical images and slice-serialized 3D volumes.
Downloads · 30 days
9
15% of all-time downloads
All-time downloads
60
Public
Parameters
8.3B
16.6 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.6 GB · 100%
From the Hugging Face model README
UniReason-Med is a medical multimodal model for grounded reasoning over 2D medical images and slice-serialized 3D volumes.
It studies whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both modalities share a common reasoning interface. A single checkpoint processes either a 2D image or a 3D volume serialized as ordered slices, generating interleaved textual reasoning and localized visual evidence through shared bounding-box syntax and region-token injection.
UniReason-Med is trained to interleave free-form reasoning with localized visual evidence. During reasoning, the model emits bounding boxes over the input image; the referenced region is cropped and re-injected as additional visual context for the next reasoning step. The same shared interface is applied to 2D images and to 3D volumes serialized as ordered slice sequences.
The model is built with supervised fine-tuning followed by GRPO reinforcement learning. RL uses answer-correctness and format rewards rather than ground-truth localization-overlap rewards such as IoU or Dice.
Released under the Apache License 2.0, consistent with the base model Qwen2.5-VL-7B-Instruct.