IMUG-Bench is a new multi-turn interleaved image-text benchmark that exposes exposure bias in unified multimodal model generation and shows test-time scaling can mitigate it.
WEA VE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation, November 2025
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
PlanViz is a new benchmark with three sub-tasks and PlanScore metric to evaluate planning-oriented image generation and editing by unified multimodal models for computer-use tasks.
ILLUME-X is a unified multimodal model that generates free-form interleaved text-image sequences via an expanded data pipeline, progressive self-adaptive training, and ILScore evaluation, claiming outperformance over prior unified models on style transfer, image decomposition, and storytelling.
citing papers explorer
-
IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation
IMUG-Bench is a new multi-turn interleaved image-text benchmark that exposes exposure bias in unified multimodal model generation and shows test-time scaling can mitigate it.
-
PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks
PlanViz is a new benchmark with three sub-tasks and PlanScore metric to evaluate planning-oriented image generation and editing by unified multimodal models for computer-use tasks.
-
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
ILLUME-X is a unified multimodal model that generates free-form interleaved text-image sequences via an expanded data pipeline, progressive self-adaptive training, and ILScore evaluation, claiming outperformance over prior unified models on style transfer, image decomposition, and storytelling.