REVIEW 9 cited by
Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advancements in text-to-image (T2I) generation have enabled models to produce high-quality images from textual descriptions. However, these models often struggle with complex instructions involving multiple objects, attributes, and spatial relationships. Existing benchmarks for evaluating T2I models primarily focus on general text-image alignment and fail to capture the nuanced requirements of complex, multi-faceted prompts. Given this gap, we introduce LongBench-T2I, a comprehensive benchmark specifically designed to evaluate T2I models under complex instructions. LongBench-T2I consists of 500 intricately designed prompts spanning nine diverse visual evaluation dimensions, enabling a thorough assessment of a model's ability to follow complex instructions. Beyond benchmarking, we propose an agent framework (Plan2Gen) that facilitates complex instruction-driven image generation without requiring additional model training. This framework integrates seamlessly with existing T2I models, using large language models to interpret and decompose complex prompts, thereby guiding the generation process more effectively. As existing evaluation metrics, such as CLIPScore, fail to adequately capture the nuances of complex instructions, we introduce an evaluation toolkit that automates the quality assessment of generated images using a set of multi-dimensional metrics. The data and code are released at https://github.com/yczhou001/LongBench-T2I.
Forward citations
Cited by 9 Pith papers
-
Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation
Self-correction blind spots in residual-stream autoregressive models arise iff the product of attention Jacobians has spectral radius ≥1, with a sharp marker threshold and RL coupling condition derived from that radius.
-
DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents
A multi-agent GPT-4o framework for editing scientific PDFs reports higher semantic consistency, layout fidelity, and instruction adherence than three baselines on DocEditBench.
-
LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation
An LVLM-driven iterative text-to-image framework whose claimed performance scores are explicitly labeled fictitious, so no empirical result is established.
-
Harnessing RLHF for Robust Unanswerability Recognition and Trustworthy Response Generation in LLMs
SALU, a multi-task fine-tuning and confidence-guided RLHF method, reduces hallucinated answers on unanswerable Chinese CIR questions to 1.3 percent on the authors' private dataset.
-
CIMR: Contextualized Iterative Multimodal Reasoning for Robust Instruction Following in LVLMs
CIMR, an iterative reasoning wrapper around LLaVA-1.5-7B, reports 91.5% task completion on a newly constructed but unreleased synthetic MAP dataset, above GPT-4V at 89.2%.
-
LVLM-Composer's Explicit Planning for Image Generation
An image generation model that explicitly plans objects, attributes, locations, and relations before synthesizing the image, with reported gains on LongBench-T2I that cannot be verified from the paper.
-
MM-FusionNet: Context-Aware Dynamic Fusion for Multi-modal Fake News Detection with Large Vision-Language Models
MM-FusionNet uses bi-directional cross-modal attention and a dynamic gating network to weight text and image features for fake news detection, reporting 0.938 F1 on the private LMFND dataset.
-
Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation
Hi-SSLVLM combines hierarchical self-captioning, internal sub-prompt planning, and a CLIP-based consistency loss, and reports judged compositional fidelity gains of roughly 0.04 to 0.09 points that no significance tes...
-
Large Language Models for Zero-Shot Multicultural Name Recognition
A prompt-tuned LLM with data augmentation and cultural context prompts reportedly recognizes multicultural names at 93.1% accuracy and unseen names at 89.5%, but the evidence is not reproducible.
Discussion (0). Sign in to comment.