Downloads · 30 days
44
3% of all-time downloads
kaist-ai/volcano-7b
volcano-7b is a image-to-text model from kaist-ai. Use it when you need a caption or text from an image. It is set up for transformers.
- Repository: https://github.com/kaistAI/Volcano - Paper: https://arxiv.org/abs/2311.07362
Downloads · 30 days
44
3% of all-time downloads
All-time downloads
1.4K
Public
Repo size
28.3 GB
Likes
3
Public
Click a slice to open those files.
.bin14.1 GB · 100%
From the Hugging Face model README
Volcano employs a single LMM to generate initial responses, feedback, and revisions, as well as decisions to accept revisions. It follows a sequential procedure of an iterative critique-revision-decide loop.
Model type: Volcano-7b is a multimodal self-feedback guided revision model that was fine-tuned by mixing the visual instruction tuning dataset used in LLaVA-v1.5 with multimodal feedback and revision data collected through gpt-3.5-turbo, applied to the vicuna-7b-v1.5 model.
Model date: Volcano-7b was trained in October 2023.
You can find here the dataset used to train Volcano, which includes all the aforementioned datasets.
A collection of three multimodal hallucination benchmarks (MMHal-Bench, Pope, GAVIE) and two multimodal understanding benchmarks (MM-Vet, MMBench).