Downloads · 30 days
51
100% of all-time downloads
uark-cviu/QuPAINT-4B
QuPAINT-4B is a image-text-to-text model from uark-cviu. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for transformers. The card lists the license as other.
Physics-Aware Instruction Tuning Approach to Quantum Material Discovery IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026 — Findings Track
Downloads · 30 days
51
100% of all-time downloads
All-time downloads
51
Public
Parameters
590K
9.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors9.5 GB · 100%
From the Hugging Face model README
Physics-Aware Instruction Tuning Approach to Quantum Material Discovery IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026 — Findings Track
Project page · Paper (PDF) · Code · Data
QuPAINT is a vision-language model for analyzing two-dimensional (2D) quantum material flakes in optical micrographs. Given a micrograph and a question, it enumerates candidate flakes, reasons about their optical contrast against the substrate, and commits to a set of bounding boxes — for example, which flakes are monolayer candidates.
The model addresses the domain shift that makes 2D-material identification hard in practice: flake appearance depends not only on material and thickness but on substrate, illumination, camera response, focus, and noise, so models trained in one lab degrade in another. QuPAINT fuses visual embeddings with optical priors via Physics-Informed Attention and is instruction-tuned on QMat-Instruct, a physics-informed multimodal instruction dataset with reasoning traces grounded in observable optical evidence.
| Architecture | vision transformer encoder + MLP projector + 4B-parameter language model |
| Parameters | ~4B |
| Precision | bfloat16 |
| Input | optical micrograph (resized to a 1792×1344 canvas, tiled into 448px patches, up to 12 tiles + thumbnail) |
| Output | free-text analysis ending in a <CONCLUSION> span of <box> quadruples in percent coordinates |
| Developed by | CVIU Lab, University of Arkansas |
Install the requirements (torch, torchvision, transformers>=4.51, Pillow,
accelerate), then:
import torch
from transformers import AutoModel, AutoTokenizer
ckpt = "uark-cviu/QuPAINT-4B"
model = AutoModel.from_pretrained(
ckpt,
dtype=torch.bfloat16,
low_cpu_mem_usage=True,
trust_remote_code=True,
use_flash_attn=False, # set True if flash-attn is installed
).eval().cuda()
tokenizer = AutoTokenizer.from_pretrained(ckpt, trust_remote_code=True)
# preprocess.py ships with this repo; it applies the tuning-time canvas and tiling
from huggingface_hub import hf_hub_download
import importlib.util, sys
spec = importlib.util.spec_from_file_location(
"qupaint_preprocess", hf_hub_download(ckpt, "preprocess.py")
)
preprocess = importlib.util.module_from_spec(spec); spec.loader.exec_module(preprocess)
pixel_values = preprocess.load_image("micrograph.jpg").to(torch.bfloat16).cuda()
response = model.chat(
tokenizer,
pixel_values,
"<image>\nIdentify monolayer candidates and provide bounding boxes [x,y,w,h].",
dict(max_new_tokens=4096, do_sample=False),
)
print(response)
For a CLI, a Gradio demo, and box parsing/plotting helpers, use the GitHub repository:
git clone https://github.com/uark-cviu/QuPAINT && cd QuPAINT
pip install -r requirements.txt
python demo.py --checkpoint uark-cviu/QuPAINT-4B
| Goal | Prompt |
|---|---|
| Detect everything | Identify all flakes and provide bounding boxes [x,y,w,h] for each |
| Monolayer candidates | Identify monolayer candidates and provide bounding boxes [x,y,w,h] |
| Counting | How many material flakes are there in the image? |
| Optical reasoning | Describe the optical contrast of the flakes relative to the substrate |
All flakes are collected first as <box>59.76, 6.88, 4.74, 4.44</box> <box>42.52, 50.68, 7.39, 9.08</box> ...
Monolayer flakes show lighter contrast and subtle color shift relative to the substrate.
<CONCLUSION>
The monolayer candidates are at: <box>42.52, 50.68, 7.39, 9.08</box>
</CONCLUSION>
Boxes are x, y, width, height as percent of image width/height with a top-left
origin, so they map back onto an image of any size. Flakes listed before the
<CONCLUSION> span are candidates the model considered; the span holds what it commits
to for the question asked.
Greedy decoding (do_sample=False) makes repeated runs on the same micrograph reproducible.
Intended use. Non-commercial academic research on automated characterization of 2D quantum materials: flake detection, monolayer screening, counting, visual grounding, and optical-contrast reasoning in exfoliation workflows.
Limitations.
Instruction-tuned on QMat-Instruct, built from Synthia-generated synthetic microscopy images with layer-dependent optical behavior, supervised with annotation-conditioned reasoning traces restricted to observable optical cues. See the paper and the project page for details.
The paper evaluates QuPAINT on QF-Bench — 8,854 images, 280,526 annotated flakes across eight materials (BN, Graphene, MoS₂, MoSe₂, MoWSe₂, WS₂, WSe₂, WTe₂), labeled mono-layer (1L), few-layer (2–4L), and thick (5+L) — covering detection, counting, visual grounding, and image-specific reasoning, including generalization to a material excluded from training. QF-Bench is released separately; see the project page for status.
@InProceedings{nguyen2026qupaint,
author = {Nguyen, Xuan Bac and Nguyen, Hoang-Quan and Pandey, Sankalp and Faltermeier, Tim and Borys, Nicholas and Churchill, Hugh and Luu, Khoa},
title = {QuPAINT: Physics-Aware Instruction Tuning Approach to Quantum Material Discovery},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings},
month = {June},
year = {2026},
pages = {8684--8694}
}
Xuan Bac Nguyen¹, Hoang-Quan Nguyen¹, Sankalp Pandey¹, Tim Faltermeier², Nicholas Borys², Hugh Churchill³, Khoa Luu¹
¹ CVIU Lab, University of Arkansas · ² University of Utah · ³ Department of Physics, University of Arkansas
Partly supported by the MonArk NSF Quantum Foundry (DMR-1906383) and an NSF Quantum Award (2444042), with GPU resources from the Arkansas High-Performance Computing Center.
These weights are released for non-commercial academic research only (see LICENSE).
The inference code on GitHub is MIT licensed.