GENFIG1 is a new benchmark that tests whether vision-language models can create effective Figure 1 visuals capturing the central scientific idea from paper text.
Rodríguez, David Vázquez, Issam H
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
S1-Omni-Image unifies scientific image understanding, generation and editing via a think-before-generate paradigm on top of S1-VL-32B, trained on a 314K-sample SciGenEdit dataset, and reports SOTA results on multiple generation and editing benchmarks.
citing papers explorer
-
GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models
GENFIG1 is a new benchmark that tests whether vision-language models can create effective Figure 1 visuals capturing the central scientific idea from paper text.
-
S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing
S1-Omni-Image unifies scientific image understanding, generation and editing via a think-before-generate paradigm on top of S1-VL-32B, trained on a 314K-sample SciGenEdit dataset, and reports SOTA results on multiple generation and editing benchmarks.