REVIEW 22 cited by
GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results in various multi-modal tasks, leveraging GPT-4V as a generalist evaluator for these tasks has not yet been systematically explored. We comprehensively validate GPT-4V's capabilities for evaluation purposes, addressing tasks ranging from foundational image-to-text and text-to-image synthesis to high-level image-to-image translations and multi-images to text alignment. We employ two evaluation methods, single-answer grading and pairwise comparison, using GPT-4V. Notably, GPT-4V shows promising agreement with humans across various tasks and evaluation methods, demonstrating immense potential for multi-modal LLMs as evaluators. Despite limitations like restricted visual clarity grading and real-world complex reasoning, its ability to provide human-aligned scores enriched with detailed explanations is promising for universal automatic evaluator.
Forward citations
Cited by 22 Pith papers
-
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
A zero-shot two-stage prompting baseline (ArtCoT) makes multimodal LLMs' aesthetic judgments align substantially better with human expert rankings in pairwise artwork comparisons.
-
IDEA-Bench: How Far are Generative Models from Professional Designing?
IDEA-Bench measures generative models on 100 professional design tasks and finds the best tested system scores only 22.48 out of 100.
-
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
VL-RewardBench curates 1,250 hard preference pairs for vision-language reward models; GPT-4o reaches only 62 to 65 percent accuracy, and benchmark scores correlate strongly with Best-of-N gains on MMMU-Pro.
-
When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery
ASV3D improves single-view 3D reconstruction by using one extra unposed photo, with a consistency-based gate selecting which image conditions each generated view.
-
MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing
MEDit-Bench shows the same source video yields strikingly different professional edits under different narrative messages, and state-of-the-art video models still fall well short of human editors at strict cut-precisi...
-
H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
H-Adapter uses a region-specific loss to induce disentangled cross-attention from which source-aligned hair masks are derived to guide diffusion inpainting, achieving strong results on pose-different hairstyle transfer.
-
Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models
A grounding-guided attack that concentrates perturbation on text-matched image regions and disrupts global and local semantic alignment consistently improves adversarial transferability across multiple vision-language models.
-
SciFig: Towards Automating Editable Figure Generation for Scientific Papers
SciFig automatically generates editable methodology figures from scientific text and claims state-of-the-art quality on its own SciFig-Eval rubric-based benchmark.
-
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
Q-Ponder is a two-stage pipeline (distill-then-reinforce) that makes a 7B multimodal model both more accurate at image quality scoring and better at explaining its judgments.
-
Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation
ABP evaluates and improves how well text-to-image models render implicit real-world knowledge.
-
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.
-
Multi-Modal Language Models as Text-to-Image Model Evaluators
MT2IE uses a single open-source multimodal LLM to generate 20 progressively harder prompts and score image-text consistency, reproducing the 1,600-prompt GenAIBench ranking of 8 text-to-image models.
-
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask
PromptDresser improves text-editable virtual try-on by combining LMM-generated structured captions with a prompt-aware adaptive mask.
-
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
A large new benchmark and an offline judge model for open-ended interleaved image-text generation, with IntJudge matching human agreement better than GPT-4o.
-
Alien Recombination: Exploring Concept Blends Beyond Human Cognitive Availability in Visual Art
The paper shows that explicitly selecting concept combinations that are rare in an artist's full body of work yields more novel-looking AI art than temperature scaling alone.
-
Modification Takes Courage: Seamless Image Stitching via Reference-Driven Inpainting
RDIStitcher fuses and rectangles stitched images by reference-driven inpainting with a LoRA-fine-tuned Stable Diffusion model, trained without labeled stitching data.
-
MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding
MultiCompose combines embedding regularization, cross-attention suppression, and mask-guided denoising to compose independently personalized subjects into one image while keeping each subject's attributes exclusive.
-
Local Brushstroke Quality Assessment via Vision-Language Feedback
Multimodal LLMs approximate expert absolute scores for calligraphy brushstrokes but show no significant rank correlation; RAG guidance trades accuracy for ranking.
-
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
A capability-oriented multimodal judge benchmark and MCTS-based preference-data generation improve judge models on some benchmarks, but the paper's SOTA claim is not supported by its own numbers.
-
Text-to-CadQuery: A New Paradigm for CAD Generation with Scalable Large Model Capabilities
Fine-tuning pretrained LLMs to emit CadQuery code directly from text outperforms the Text2CAD command-sequence baseline on geometric metrics, with the best open-source 3B model achieving the reported gains.
-
Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.
-
LA-RCS: LLM-Agent-Based Robot Control System
LA-RCS reports that a dual-agent LLM system controls a small car robot to complete 18 of 20 self-designed commands with the GPT-4o variant, but the supporting evaluation is inconsistent and not reproducible.
Discussion (0). Continue with ORCID to comment.