Downloads · 30 days
0
VQVA/BAGEL-World-model
BAGEL-World-model is a machine learning model from VQVA. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
A agentic data-centric framework for producing large-scale interleaved Visual Question–Visual Answering (VQ-VA) data.
Downloads · 30 days
0
Access
Public
Updated Oct 16, 2025
Repo size
—
Likes
0
Public
Click a slice to open those files.
Other1.5 KB · 51%
From the Hugging Face model README
A agentic data-centric framework for producing large-scale interleaved Visual Question–Visual Answering (VQ-VA) data.
The BAGEL-World framework outputs high-quality VQ-VA data via the following steps:
Filters and classify noisy web-interleaved data into design- and knowledge-related documents.
1. Retriever selects image pairs containing non-trivial transformations from interleaved documents that can serve as the basis for free-form questions.
2. Instruction Generator write a natural-language question about one image so that the other image serves as the correct answer.
3. Filterer removes low-quality triplets ⟨Question Image, Question Text, Answer Image⟩.
4. Rewriter increases instruction diversity by producing multiple variants of the original questions.
5. Reasoner generates a language-based chain-of-thought explanation describing how the source image should be transformed to obtain the target image.
The framework at last outputs interleaved quadruplets:
Stay tuned for updates and examples!