Pith. sign in

REVIEW 22 cited by

GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.01361 v1 pith:Z4YGASRU submitted 2023-11-02 cs.CV cs.CL

classification cs.CVcs.CL
keywords gpt-4vtasksevaluationevaluatorpromisinggeneralistgradinglimitations
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results in various multi-modal tasks, leveraging GPT-4V as a generalist evaluator for these tasks has not yet been systematically explored. We comprehensively validate GPT-4V's capabilities for evaluation purposes, addressing tasks ranging from foundational image-to-text and text-to-image synthesis to high-level image-to-image translations and multi-images to text alignment. We employ two evaluation methods, single-answer grading and pairwise comparison, using GPT-4V. Notably, GPT-4V shows promising agreement with humans across various tasks and evaluation methods, demonstrating immense potential for multi-modal LLMs as evaluators. Despite limitations like restricted visual clarity grading and real-world complex reasoning, its ability to provide human-aligned scores enriched with detailed explanations is promising for universal automatic evaluator.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multimodal LLMs Can Reason about Aesthetics in Zero-Shot

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A zero-shot two-stage prompting baseline (ArtCoT) makes multimodal LLMs' aesthetic judgments align substantially better with human expert rankings in pairwise artwork comparisons.

  2. IDEA-Bench: How Far are Generative Models from Professional Designing?

    cs.CV 2024-12 conditional novelty 7.0 of 10

    IDEA-Bench measures generative models on 100 professional design tasks and finds the best tested system scores only 22.48 out of 100.

  3. VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

    cs.CV 2024-11 conditional novelty 7.0 of 10

    VL-RewardBench curates 1,250 hard preference pairs for vision-language reward models; GPT-4o reaches only 62 to 65 percent accuracy, and benchmark scores correlate strongly with Best-of-N gains on MMMU-Pro.

  4. When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery

    cs.CV 2026-08 conditional novelty 6.0 of 10

    ASV3D improves single-view 3D reconstruction by using one extra unposed photo, with a consistency-based gate selecting which image conditions each generated view.

  5. MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    MEDit-Bench shows the same source video yields strikingly different professional edits under different narrative messages, and state-of-the-art video models still fall well short of human editors at strict cut-precisi...

  6. H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    H-Adapter uses a region-specific loss to induce disentangled cross-attention from which source-aligned hair masks are derived to guide diffusion inpainting, achieving strong results on pose-different hairstyle transfer.

  7. Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models

    cs.CR 2026-02 conditional novelty 6.0 of 10

    A grounding-guided attack that concentrates perturbation on text-matched image regions and disrupts global and local semantic alignment consistently improves adversarial transferability across multiple vision-language models.

  8. SciFig: Towards Automating Editable Figure Generation for Scientific Papers

    cs.AI 2026-01 conditional novelty 6.0 of 10

    SciFig automatically generates editable methodology figures from scientific text and claims state-of-the-art quality on its own SciFig-Eval rubric-based benchmark.

  9. Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Q-Ponder is a two-stage pipeline (distill-then-reinforce) that makes a 7B multimodal model both more accurate at image quality scoring and better at explaining its judgments.

  10. Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ABP evaluates and improves how well text-to-image models render implicit real-world knowledge.

  11. KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.

  12. Multi-Modal Language Models as Text-to-Image Model Evaluators

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MT2IE uses a single open-source multimodal LLM to generate 20 progressively harder prompts and score image-text consistency, reproducing the 1,600-prompt GenAIBench ranking of 8 text-to-image models.

  13. PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PromptDresser improves text-editable virtual try-on by combining LMM-generated structured captions with a prompt-aware adaptive mask.

  14. OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A large new benchmark and an offline judge model for open-ended interleaved image-text generation, with IntJudge matching human agreement better than GPT-4o.

  15. Alien Recombination: Exploring Concept Blends Beyond Human Cognitive Availability in Visual Art

    cs.AI 2024-11 conditional novelty 6.0 of 10

    The paper shows that explicitly selecting concept combinations that are rare in an artist's full body of work yields more novel-looking AI art than temperature scaling alone.

  16. Modification Takes Courage: Seamless Image Stitching via Reference-Driven Inpainting

    cs.CV 2024-11 conditional novelty 6.0 of 10

    RDIStitcher fuses and rectangles stitched images by reference-driven inpainting with a LoRA-fine-tuned Stable Diffusion model, trained without labeled stitching data.

  17. MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding

    cs.CV 2026-08 conditional novelty 5.0 of 10

    MultiCompose combines embedding regularization, cross-attention suppression, and mask-guided denoising to compose independently personalized subjects into one image while keeping each subject's attributes exclusive.

  18. Local Brushstroke Quality Assessment via Vision-Language Feedback

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Multimodal LLMs approximate expert absolute scores for calligraphy brushstrokes but show no significant rank correlation; RAG guidance trades accuracy for ranking.

  19. Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

    cs.AI 2026-02 reject novelty 5.0 of 10

    A capability-oriented multimodal judge benchmark and MCTS-based preference-data generation improve judge models on some benchmarks, but the paper's SOTA claim is not supported by its own numbers.

  20. Text-to-CadQuery: A New Paradigm for CAD Generation with Scalable Large Model Capabilities

    cs.AI 2025-05 conditional novelty 5.0 of 10

    Fine-tuning pretrained LLMs to emit CadQuery code directly from text outperforms the Text2CAD command-sequence baseline on geometric metrics, with the best open-source 3B model achieving the reported gains.

  21. Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning

    cs.RO 2025-08 reject novelty 4.0 of 10

    A review that categorizes large-model-empowered embodied AI into hierarchical and end-to-end decision-making, imitation and reinforcement learning, and world models.

  22. LA-RCS: LLM-Agent-Based Robot Control System

    cs.RO 2025-05 reject novelty 4.0 of 10

    LA-RCS reports that a dual-agent LLM system controls a small car robot to complete 18 of 20 self-designed commands with the GPT-4o variant, but the supporting evaluation is inconsistent and not reproducible.

Pith tools