Downloads · 30 days
27
34% of all-time downloads
JiwanKim/CompoDistill-SFT-2B
CompoDistill-SFT-2B is a image-text-to-text model from JiwanKim. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
The student baseline (Qwen1.5-1.8B + SigLIP-so400m) trained with standard LLaVA-style visual instruction tuning, without any distillation. Used as the SFT baseline in the CompoDistill paper.
Downloads · 30 days
27
34% of all-time downloads
All-time downloads
79
Public
Parameters
2.3B
4.5 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors4.5 GB · 100%
From the Hugging Face model README
The student baseline (Qwen1.5-1.8B + SigLIP-so400m) trained with standard LLaVA-style visual instruction tuning, without any distillation. Used as the SFT baseline in the CompoDistill paper.
Released with the paper CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs (arXiv:2510.12184). Training and evaluation code: https://github.com/ptkjw1997/CompoDistill
import torch
from PIL import Image
from transformers import AutoModelForCausalLM, AutoTokenizer, AutoImageProcessor
repo = "JiwanKim/CompoDistill-SFT-2B"
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True,
torch_dtype=torch.float16).to("cuda")
tokenizer = AutoTokenizer.from_pretrained(repo, use_fast=False)
image_processor = AutoImageProcessor.from_pretrained(repo)
image = Image.open("example.jpg")
print(model.chat("What is happening in this image?", tokenizer,
image=image, image_processor=image_processor))
@article{kim2025compodistill,
title={CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs},
author={Kim, Jiwan and Kim, Kibum and Seo, Sangwoo and Park, Chanyoung},
journal={arXiv preprint arXiv:2510.12184},
year={2025}
}