Downloads · 30 days
0
tsunghanwu/reverse-instruct-1.3m
reverse-instruct-1.3m is a machine learning model from tsunghanwu. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as mit.
Dataset Type: REVERSE Visual Instruct 1.3M is a GPT-generated instruction-following dataset designed for training hallucination-aware vision-language models (VLMs). It builds on the LLaVA Instruct 665K dataset and inc…
Downloads · 30 days
0
Access
Public
Updated May 30, 2025
Repo size
2.1 GB
Likes
0
Public
Click a slice to open those files.
.json2.1 GB · 100%
From the Hugging Face model README
Dataset Type:
REVERSE Visual Instruct 1.3M is a GPT-generated instruction-following dataset designed for training hallucination-aware vision-language models (VLMs). It builds on the LLaVA Instruct 665K dataset and includes structured annotations to indicate model confidence. We introduce three special tokens:
<SPAN>: marks the beginning of a key phrase</CN>: denotes a confident (grounded) phrase</UN>: denotes an unconfident (potentially hallucinated) phraseRoughly 50% of the examples contain correct (grounded) phrases from the original LLaVA dataset, while the remaining 50% include hallucinated or incorrect phrases generated via GPT-4o-mini-0718 and rule-based augmentations.
Collection Date:
February 2025
Data Generation:
Incorrect examples were synthesized via a combination of GPT-4o-mini prompts and deterministic, rule-based edits. Correct examples are drawn directly from the original LLaVA Visual Instruct dataset. Please refer to our GitHub Repo for more details.
License:
Project Page:
https://reverse-vlm.github.io
Support & Issues:
GitHub Issues
Primary Use Cases:
Target Users:
Researchers, practitioners, and hobbyists working in computer vision, natural language processing, and multi-modal AI.