Skip to content

dmusingu

diffvqa-qwen3-0.6b-multiobjective-regsteps

dmusingu/diffvqa-qwen3-0.6b-multiobjective-regsteps

diffvqa-qwen3-0.6b-multiobjective-regsteps is a visual question answering model from dmusingu. Use it for the visual question answering task on the model card, and read the license before you ship it in a product. It is set up for pytorch. The card lists the license as mit.

Online-trained Difference Visual Question Answering head for chest X-rays (MIMIC-CXR / Medical-Diff-VQA). A frozen ViT-L/14 vision encoder produces patch tokens for a current + reference image pair; a Qwen3-0.6B decod…

Downloads · 30 days

0

Access

Public

Updated Jul 17, 2026

Repo size

2.4 GB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.pt2.4 GB · 100%

At a glance

Task
Visual Question Answering
Library
pytorch
License
mit
Access
Public
Created
Jul 17, 2026
Updated
Jul 17, 2026
SHA
7fbf7538
Task
Visual Question Answering
Library
pytorch
License
mit
Created
Jul 17, 2026
Updated
Jul 17, 2026
diffvqa-qwen3-0.6b-multiobjective-regsteps — AI Model — AIMarketly