VIGA introduces a training-free interleaved multimodal reasoning loop that improves vision-as-inverse-graphics accuracy over one-shot baselines on BlenderGym, SlideBench, and new BlenderBench.
ID": "2401.13641
6 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Introduces the SciGA-145k dataset with intra-paper and cross-paper graphical abstract recommendation tasks plus the CAR evaluation metric.
Any2Poster Bench tests poster generation from 8 modalities and 5 domains using quizzes and VLM judgments; Any2Poster Agent reaches 87% accuracy and beats prior paper-only methods.
Crafter introduces a multi-agent harness for generating and editing scientific figures across types and inputs, with a new benchmark showing outperformance over baselines.
VideoAgent is a modular framework that redefines scientific video synthesis as an intent-driven planning problem and introduces the SciVidEval benchmark for multimodal quality and pedagogical utility.
PosterForest uses a hierarchical Poster Tree and content-layout agent collaboration to generate scientific posters from papers without training, outperforming prior automated methods in human and automated evaluation.
citing papers explorer
-
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
VIGA introduces a training-free interleaved multimodal reasoning loop that improves vision-as-inverse-graphics accuracy over one-shot baselines on BlenderGym, SlideBench, and new BlenderBench.
-
SciGA: A Comprehensive Dataset for Designing Graphical Abstracts in Academic Papers
Introduces the SciGA-145k dataset with intra-paper and cross-paper graphical abstract recommendation tasks plus the CAR evaluation metric.
-
Any2Poster: Any-Source Poster Generation Across Modalities and Domains
Any2Poster Bench tests poster generation from 8 modalities and 5 domains using quizzes and VLM judgments; Any2Poster Agent reaches 87% accuracy and beats prior paper-only methods.
-
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
Crafter introduces a multi-agent harness for generating and editing scientific figures across types and inputs, with a new benchmark showing outperformance over baselines.
-
VideoAgent: Personalized Synthesis of Scientific Videos
VideoAgent is a modular framework that redefines scientific video synthesis as an intent-driven planning problem and introduces the SciVidEval benchmark for multimodal quality and pedagogical utility.
-
PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation
PosterForest uses a hierarchical Poster Tree and content-layout agent collaboration to generate scientific posters from papers without training, outperforming prior automated methods in human and automated evaluation.