Downloads · 30 days
13
13% of all-time downloads
changdae/vittle-7b-F
vittle-7b-F is a image-text-to-text model from changdae. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Vittle (F) is the fixed-prior variant of Vittle (NeurIPS 2025), a VLM instruction tuning framework that improves robustness to distribution shifts via variational information bottleneck.
Downloads · 30 days
13
13% of all-time downloads
All-time downloads
100
Public
Parameters
7.2B
14.3 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors14.3 GB · 100%
From the Hugging Face model README
Vittle (F) is the fixed-prior variant of Vittle (NeurIPS 2025), a VLM instruction tuning framework that improves robustness to distribution shifts via variational information bottleneck.
import torch
from vittle.model.language_model.vittle_llama import VittleLlamaForCausalLM
from transformers import AutoTokenizer
model = VittleLlamaForCausalLM.from_pretrained(
"changdae/vittle-7b-F",
torch_dtype=torch.bfloat16,
device_map="cuda:0",
)
tokenizer = AutoTokenizer.from_pretrained("changdae/vittle-7b-F", use_fast=False)
Refer to the evaluation guide for full inference instructions.
| Property | Value |
|---|---|
| Base model | lmsys/vicuna-7b-v1.5 |
| Vision encoder | openai/clip-vit-large-patch14-336 |
| Bottleneck layer | 24 |
| Interpolation coefficient (alpha) | 0.5 |
| KLD strength (beta) | 0.1 |
| Learnable prior | No (fixed) |
| Training dtype | bfloat16 |
@inproceedings{
oh2025visual,
title={Visual Instruction Bottleneck Tuning},
author={Changdae Oh and Jiatong Li and Shawn Im and Sharon Li},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=yzHiEmLSk8}
}