Downloads · 30 days
16
20% of all-time downloads
JiwanKim/CompoDistill-Teacher-4B
CompoDistill-Teacher-4B is a image-text-to-text model from JiwanKim. Use it for the image-text-to-text task on the model card, and read the license before you ship it in a product. The card lists the license as apache-2.0.
The teacher MLLM (Qwen1.5-4B + SigLIP-so400m) trained with LLaVA-style visual instruction tuning. Serves as the distillation teacher for CompoDistill-2B; can be passed directly to scripts/train/dpt.sh / dft.sh via --p…
Downloads · 30 days
16
20% of all-time downloads
All-time downloads
80
Public
Parameters
4.4B
8.8 GB on disk
Likes
0
Public
Click a slice to open those files.
.safetensors8.8 GB · 100%
From the Hugging Face model README
The teacher MLLM (Qwen1.5-4B + SigLIP-so400m) trained with LLaVA-style visual instruction tuning. Serves as the distillation teacher for CompoDistill-2B; can be passed directly to scripts/train/dpt.sh / dft.sh via --pretrained_teacher_model_path.
Released with the paper CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs (arXiv:2510.12184). Training and evaluation code: https://github.com/ptkjw1997/CompoDistill
import torch
from PIL import Image
from transformers import AutoModelForCausalLM, AutoTokenizer, AutoImageProcessor
repo = "JiwanKim/CompoDistill-Teacher-4B"
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True,
torch_dtype=torch.float16).to("cuda")
tokenizer = AutoTokenizer.from_pretrained(repo, use_fast=False)
image_processor = AutoImageProcessor.from_pretrained(repo)
image = Image.open("example.jpg")
print(model.chat("What is happening in this image?", tokenizer,
image=image, image_processor=image_processor))
@article{kim2025compodistill,
title={CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs},
author={Kim, Jiwan and Kim, Kibum and Seo, Sangwoo and Park, Chanyoung},
journal={arXiv preprint arXiv:2510.12184},
year={2025}
}