Skip to content

shravvvv

VISE

shravvvv/VISE

VISE is a image-text-to-text model from shravvvv. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. It is set up for peft. The card lists the license as apache-2.0.

This is the VISE LoRA adapter for Qwen/Qwen3-VL-2B-Instruct, from our paper Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models.

Downloads · 30 days

18

13% of all-time downloads

All-time downloads

138

Public

Repo size

81.2 MB

Likes

0

Public

Hugging Face

Repo makeup

Click a slice to open those files.

.safetensors69.8 MB · 81%

At a glance

Task
Image-Text-to-Text
Library
peft
License
apache-2.0
Access
Public
Created
Jun 22, 2026
Updated
Jun 26, 2026
SHA
2fc58ecb

Try a prompt

Base models

Task
Image-Text-to-Text
Library
peft
License
apache-2.0
Languages
en
Created
Jun 22, 2026
Updated
Jun 26, 2026