Downloads · 30 days
39
58% of all-time downloads
expertailab/scigram-7b
scigram-7b is a machine learning model from expertailab. Use it for the machine learning task on the model card, and read the license before you ship it in a product. The card lists the license as cc-by-4.0.
SciGram-7B is a 7B-class vision-language model specialized in understanding scientific diagrams. It is based on a LLaVA-style architecture with a pretrained CLIP vision encoder and a Qwen2-Instruct 7B language model,…
Downloads · 30 days
39
58% of all-time downloads
All-time downloads
67
Public
Parameters
7.9B
16.3 GB on disk
Likes
1
Public
Click a slice to open those files.
.safetensors16.3 GB · 100%
How the weights are stored.
F167.6B · 97%
From the Hugging Face model README
SciGram-7B is a 7B-class vision-language model specialized in understanding scientific diagrams. It is based on a LLaVA-style architecture with a pretrained CLIP vision encoder and a Qwen2-Instruct 7B language model, and is trained on SciGram, a large-scale dataset of scientific diagrams and synthetic visual instructions covering life, earth, and physical sciences.
The model is designed for multimodal understanding tasks involving scientific diagrams, including diagram question answering, visual grounding, diagram description, and science-related visual reasoning.
SciGram-7B corresponds to the LLaVA-SciGram 7B model described in:
Raul Ortega and José Manuel Gómez-Pérez. From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding. Proceedings of the Conference on Language Modeling (COLM 2026).
SciGram-7B is a vision-language model trained to improve scientific diagram understanding. It uses a pretrained CLIP vision encoder together with a Qwen2-Instruct 7B language model. The model is trained using three SciGram training stages: SciGram-Align, SciGram-VIT, and SciGram-M³.
The training data contains scientific diagrams paired with synthetic captions and multiple-choice visual instructions covering life sciences, earth sciences, and physical sciences.
The model is intended primarily for research on scientific diagram understanding and multimodal scientific question answering.
The released checkpoint contains approximately 8B parameters and is approximately 16.3 GB in size.
The model is released using the LLaVA-Qwen architecture. It can be loaded using a compatible LLaVA implementation and the Hugging Face checkpoint.
A typical workflow is:
from transformers import AutoProcessor, AutoModel
from PIL import Image
import torch
model_id = "expertailab/scigram-7b"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModel.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto",
)
image = Image.open("scientific_diagram.png").convert("RGB")
prompt = "What does this scientific diagram show? Explain the main components and their relationships."
inputs = processor(
images=image,
text=prompt,
return_tensors="pt"
).to(model.device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=512,
)
answer = processor.batch_decode(
output,
skip_special_tokens=True
)[0]
print(answer)
The model repository contains a LlavaProcessor configuration and uses the Qwen2 tokenizer with the special <image> token. For reproducible benchmark evaluation,
users should follow the prompts and evaluation procedure described in the paper rather than relying on the generation defaults.
SciGram-7B was trained using the three SciGram subsets:
Together, these subsets contain approximately 1.37 million training instructions.
The complete SciGram dataset contains 194,071 scientific diagrams and more than 1.4 million synthetic visual instructions. The diagrams cover life sciences, earth sciences, and physical sciences.
The dataset construction pipeline extracts scientific terminology from middle-school science curricula, generates atomic scientific facts, retrieves candidate diagrams from the web, filters and deduplicates the retrieved images, and generates captions and multiple-choice questions using a vision-language model.
The instruction-generation process used Qwen2-VL-7B to generate captions and diagram-grounded multiple-choice questions.
SciGram-7B was trained using a LLaVA-style three-stage training pipeline:
The model was trained on two NVIDIA A100 GPUs for approximately 450 GPU-hours.
The model follows the LLaVA image-processing and multimodal training pipeline.
For the visual instruction tuning and further fine-tuning stages, the training configuration uses an any-resolution image processing strategy and the following image grid pinpoints:
(384, 768)
(768, 384)
(768, 768)
(1152, 384)
(384, 1152)
The multimodal projector uses an MLP with two GELU layers (mlp2x_gelu), and the selected vision feature layer is -2.
Alignment stage — SciGram-Align
Visual Instruction Tuning — SciGram-VIT
Further Fine-tuning — SciGram-M³
SciGram-7B was evaluated on three scientific multimodal question-answering benchmarks:
The primary evaluation metric is accuracy (%), computed as the proportion of correctly answered multiple-choice questions.
The reported results for LLaVA-SciGram 7B are:
| Benchmark | Subset | Accuracy |
|---|---|---|
| TQA | Text MC | 90.87 |
| TQA | True/False | 92.00 |
| TQA | Diagram MC | 76.68 |
| TQA | Overall | 83.66 |
| ScienceQA | Natural Sciences | 96.27 |
| ScienceQA | Social Sciences | 97.53 |
| ScienceQA | Language | 91.64 |
| ScienceQA | Text | 99.11 |
| ScienceQA | Visual/Image Support | 95.24 |
| ScienceQA | No Support | 93.38 |
| ScienceQA | Grade 1–6 | 95.49 |
| ScienceQA | Grade 7–12 | 95.06 |
| ScienceQA | Overall | 95.33 |
| AI2D | Opaque Labels | 80.21 |
| AI2D | Transparent Labels | 89.93 |
| AI2D | Overall | 85.07 |
SciGram-7B uses a LLaVA-Qwen multimodal architecture.
The model consists of:
LlavaQwenForCausalLMQwen/Qwen2-7B-Instructopenai/clip-vit-large-patch14-336anyres)Training was performed using two NVIDIA A100 GPUs and required approximately 450 GPU-hours.
The released checkpoint is compatible with the LLaVA-Qwen architecture.
The associated environment includes:
All datasets and pretrained models are subject to their respective licenses, and future users are responsible for complying with their terms. Improper use of copyrighted datasets or proprietary models may result in legal or ethical violations. As noted in our GitHub repository on the license and copyright of content linked from SciGram:
Pretrained models may reflect biases in their training data. Although our study focuses on diagram reasoning, such biases may affect downstream outputs, potentially disadvantaging certain groups or misrepresenting information. Users should consider these risks when deploying similar models. Environmental Impact. Training and fine-tuning large models are computationally expensive and contribute to carbon emissions. We encourage efficient training strategies and consideration of environmental costs when developing similar systems. Misuse Potential. Although intended for research and educational purposes, our approach could be misused for automated content generation or misinformation. Appropriate safeguards and ethical guidelines should be followed to minimize potential harm.
If you use SciGram-7B or the SciGram dataset, please cite the accompanying paper:
BibTeX:
@inproceedings{
ortega2026from,
title={From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding},
author={Raúl Ortega and José Maunel Gómez-Pérez,
booktitle={Third Conference on Language Modeling},
year={2026},
url={https://openreview.net/forum?id=Xdi2q5XKgB}
}
APA:
Ortega, R., & Gómez-Pérez, J. M. (2026). From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding. Proceedings of the Conference on Language Modeling (COLM 2026).
For questions concerning the model or dataset, please contact [email protected] or [email protected].