Downloads ยท 30 days
0
JustSecret/Imag-Eval
Imag-Eval is a text-to-image model from JustSecret. Use it when you need an image from a text prompt. The card lists the license as cc-by-nc-4.0.
<p align="center" <bIMAG-EVAL:</b A Language-Grounded Framework for Interpretable Text-to-Image Instruction Following Evaluation </p
Downloads ยท 30 days
0
Access
Public
Updated Sep 5, 2026
Repo size
โ
Likes
0
Public
Click a slice to open those files.
.json1.1 MB ยท 96%
From the Hugging Face model README
IMAG-EVAL is a controlled benchmark designed to evaluate instruction-following capabilities in Text-to-Image (T2I) generation models.
Unlike prior benchmarks that primarily rely on prompt length or isolated skill evaluation, IMAG-EVAL explicitly disentangles linguistic complexity from compositional complexity by independently varying:
The benchmark provides interpretable diagnostic evaluations and enables fine-grained analysis of failure modes in modern Text-to-Image systems.
| Statistic | Value |
|---|---|
| Prompts | 1,140 |
| Evaluation Rules | 8,842 |
| Languages | English |
| Evaluation Dimensions | 7 |
| Difficulty levels | 3 : Easy, Medium, Hard |
| Available tests | from 2-skill to 6-skill combinations |
| Skill | JSON rules (meta-prompts) | Total combined rules (synthetic prompts) |
|---|---|---|
| Counting | 2,228 | 2,228 |
| Color | 558 | 1,660 |
| Spatial | 310 | 2,199 |
| Emotion | 372 | 372 |
| Size | 310 | 2,197 |
| Text | 186 | 186 |
| Total | 3,964 | 8,842 |
prompts/Contains the benchmark prompt definitions used during evaluation.
Each JSON file includes:
skill_codes.csvMapping between benchmark identifiers and skill combinations.
Example:
| code | skill |
|---|---|
| 0 | all_skills |
| 12 | color+size |
| 46 | counting+color+spatial+emotion |
These codes are used throughout the benchmark generation and evaluation pipeline.
texts.csvReference textual annotations used for Text Rendering evaluation.
Each row contains:
This file is used to compute Word Error Rate (WER) during evaluation.
IMAG-EVAL is intended for:
Models are evaluated on seven dimensions:
| Skill | Metric |
|---|---|
| Counting | Accuracy (object-level) |
| Color Attribution | Accuracy (instance-level) |
| Spatial Relations | Accuracy (instance-level) |
| Size Relations | Accuracy (instance-level) |
| Emotion Attribution | Accuracy |
| Text Rendering | Word Error Rate (WER) |
| Cohesiveness | Binary Classification Accuracy |
@misc{serouis2026imageval,
title={IMAG-EVAL: A Language-Grounded Framework for Interpretable Text-to-Image Instruction Following Evaluation},
author={Serouis, Ibrahim Mohamed and Jaramillo Duque, David},
year={2026},
note={Accepted at the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)}
}
Please refer to the repository license for terms of use.