Pith. sign in

REVIEW 9 cited by

Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.24787 v1 pith:G43GURTG submitted 2025-05-30 cs.CV cs.CL

classification cs.CVcs.CL
keywords complexmodelsgenerationinstructionsevaluationexistingframeworklongbench-t2i
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in text-to-image (T2I) generation have enabled models to produce high-quality images from textual descriptions. However, these models often struggle with complex instructions involving multiple objects, attributes, and spatial relationships. Existing benchmarks for evaluating T2I models primarily focus on general text-image alignment and fail to capture the nuanced requirements of complex, multi-faceted prompts. Given this gap, we introduce LongBench-T2I, a comprehensive benchmark specifically designed to evaluate T2I models under complex instructions. LongBench-T2I consists of 500 intricately designed prompts spanning nine diverse visual evaluation dimensions, enabling a thorough assessment of a model's ability to follow complex instructions. Beyond benchmarking, we propose an agent framework (Plan2Gen) that facilitates complex instruction-driven image generation without requiring additional model training. This framework integrates seamlessly with existing T2I models, using large language models to interpret and decompose complex prompts, thereby guiding the generation process more effectively. As existing evaluation metrics, such as CLIPScore, fail to adequately capture the nuances of complex instructions, we introduce an evaluation toolkit that automates the quality assessment of generated images using a set of multi-dimensional metrics. The data and code are released at https://github.com/yczhou001/LongBench-T2I.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation

    cs.LG 2026-07 conditional novelty 5.5 of 10

    Self-correction blind spots in residual-stream autoregressive models arise iff the product of attention Jacobians has spectral radius ≥1, with a sharp marker threshold and RL coupling condition derived from that radius.

  2. DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents

    cs.CV 2025-08 reject novelty 4.0 of 10

    A multi-agent GPT-4o framework for editing scientific PDFs reports higher semantic consistency, layout fidelity, and instruction adherence than three baselines on DocEditBench.

  3. LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

    cs.LG 2025-08 reject novelty 4.0 of 10

    An LVLM-driven iterative text-to-image framework whose claimed performance scores are explicitly labeled fictitious, so no empirical result is established.

  4. Harnessing RLHF for Robust Unanswerability Recognition and Trustworthy Response Generation in LLMs

    cs.CL 2025-07 reject novelty 4.0 of 10

    SALU, a multi-task fine-tuning and confidence-guided RLHF method, reduces hallucinated answers on unanswerable Chinese CIR questions to 1.3 percent on the authors' private dataset.

  5. CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs

    cs.LG 2025-07 reject novelty 4.0 of 10

    CIMR, an iterative reasoning wrapper around LLaVA-1.5-7B, reports 91.5% task completion on a newly constructed but unreleased synthetic MAP dataset, above GPT-4V at 89.2%.

  6. LVLM-Composer's Explicit Planning for Image Generation

    cs.CV 2025-07 reject novelty 4.0 of 10

    An image generation model that explicitly plans objects, attributes, locations, and relations before synthesizing the image, with reported gains on LongBench-T2I that cannot be verified from the paper.

  7. MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models

    cs.CR 2025-08 reject novelty 3.0 of 10

    MM-FusionNet uses bi-directional cross-modal attention and a dynamic gating network to weight text and image features for fake news detection, reporting 0.938 F1 on the private LMFND dataset.

  8. Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation

    cs.CV 2025-07 reject novelty 3.0 of 10

    Hi-SSLVLM combines hierarchical self-captioning, internal sub-prompt planning, and a CLIP-based consistency loss, and reports judged compositional fidelity gains of roughly 0.04 to 0.09 points that no significance tes...

  9. Large Language Models for Zero-Shot Multicultural Name Recognition

    cs.CL 2025-07 reject novelty 3.0 of 10

    A prompt-tuned LLM with data augmentation and cultural context prompts reportedly recognizes multicultural names at 93.1% accuracy and unseen names at 89.5%, but the evidence is not reproducible.

Pith tools