Pith. sign in

REVIEW 19 cited by

GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.01361 v1 pith:Z4YGASRU submitted 2023-11-02 cs.CV cs.CL

GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

classification cs.CV cs.CL
keywords gpt-4vtasksevaluationevaluatorpromisinggeneralistgradinglimitations
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Automatically evaluating vision-language tasks is challenging, especially when it comes to reflecting human judgments due to limitations in accounting for fine-grained details. Although GPT-4V has shown promising results in various multi-modal tasks, leveraging GPT-4V as a generalist evaluator for these tasks has not yet been systematically explored. We comprehensively validate GPT-4V's capabilities for evaluation purposes, addressing tasks ranging from foundational image-to-text and text-to-image synthesis to high-level image-to-image translations and multi-images to text alignment. We employ two evaluation methods, single-answer grading and pairwise comparison, using GPT-4V. Notably, GPT-4V shows promising agreement with humans across various tasks and evaluation methods, demonstrating immense potential for multi-modal LLMs as evaluators. Despite limitations like restricted visual clarity grading and real-world complex reasoning, its ability to provide human-aligned scores enriched with detailed explanations is promising for universal automatic evaluator.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. C3-Bench: A Context-Aware Change Captioning Benchmark

    cs.CV 2026-06 unverdicted novelty 7.0

    C3-Bench supplies a multi-domain dataset and LLM-based evaluation protocol that exposes systematic failures in existing change captioning models outside their training regimes.

  2. Can Image Models Imagine Time? ImageTime: A Novel Benchmark for Probing Visual World Modeling Through Spatiotemporal Consistency

    cs.CV 2026-06 unverdicted novelty 7.0

    ImageTime is a benchmark that probes image generation models' visual world modeling by requiring coherent four-state sequences in single images, scored via VLM judge.

  3. Do Image-Text Metrics Respect Semantic Invariances?

    cs.CV 2026-05 unverdicted novelty 7.0

    Five image-text metrics exhibit non-semantic sensitivities to spatial, object, and socio-linguistic perturbations, shifting scores by 6-9% on average and flipping rankings in up to 37% of cases, with a proposed post-h...

  4. Prompt-Guided Image Editing with Masked Logit Nudging in Visual Autoregressive Models

    cs.CV 2026-04 unverdicted novelty 7.0

    Masked Logit Nudging aligns visual autoregressive model logits with source token maps under target prompts inside cross-attention masks, delivering top image editing results on PIE benchmarks and strong reconstruction...

  5. T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts

    cs.CV 2024-12 unverdicted novelty 7.0

    T2I-FactualBench is a new three-tier benchmark for factuality of knowledge-intensive concepts in T2I models, using multi-round VQA evaluation to show SOTA models need improvement.

  6. MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

    cs.CV 2026-07 conditional novelty 6.0

    MEDit-Bench shows the same source video yields strikingly different professional edits under different narrative messages, and state-of-the-art video models still fall well short of human editors at strict cut-precisi...

  7. H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks

    cs.CV 2026-06 conditional novelty 6.0

    A diffusion-based hairstyle transfer method that uses a region-specific training loss to make cross-attention produce a source-aligned hair mask for pose-robust inpainting.

  8. Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting

    cs.CV 2026-06 unverdicted novelty 6.0

    Prompt-aware weighting strategies W-Switch and W-Composite improve multi-concept LoRA composition in diffusion models without training.

  9. PHASER: Phase-Aware and Semantic Experience Replay for Vision-Language-Action Models

    cs.RO 2026-06 unverdicted novelty 6.0

    PHASER improves average success rate by up to 31% over uniform experience replay on LIBERO continual learning benchmarks for VLA models by phase-centric capacity allocation and semantic interference routing.

  10. Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics

    cs.CY 2026-04 unverdicted novelty 6.0

    Community members from the UK blind community, Kerala, and Tamil Nadu helped define what counts as culturally appropriate depictions of artifacts, and the authors tested whether those definitions can be turned into re...

  11. Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models

    cs.CR 2026-02 conditional novelty 6.0

    A grounding-guided attack that concentrates perturbation on text-matched image regions and disrupts global and local semantic alignment consistently improves adversarial transferability across multiple vision-language models.

  12. SciFig: Towards Automating Editable Figure Generation for Scientific Papers

    cs.AI 2026-01 conditional novelty 6.0

    SciFig automatically generates editable methodology figures from scientific text and claims state-of-the-art quality on its own SciFig-Eval rubric-based benchmark.

  13. VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

    cs.CV 2024-12 unverdicted novelty 6.0

    VisionReward learns multi-dimensional human preferences for image and video generation via hierarchical assessment and linear weighting, outperforming VideoScore by 17.2% in prediction accuracy and yielding 31.6% high...

  14. GPT-4V(ision) is a Generalist Web Agent, if Grounded

    cs.IR 2024-01 conditional novelty 6.0

    GPT-4V achieves 51.1% success on live web tasks as a generalist agent when plans are manually grounded, outperforming text-only models, but automatic grounding lags far behind oracle performance.

  15. Local Brushstroke Quality Assessment via Vision-Language Feedback

    cs.CV 2026-07 conditional novelty 5.0

    Multimodal LLMs approximate expert absolute scores for calligraphy brushstrokes but show no significant rank correlation; RAG guidance trades accuracy for ranking.

  16. H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks

    cs.CV 2026-06 unverdicted novelty 5.0

    H-Adapter uses a region-specific loss to induce disentangled cross-attention from which source-aligned hair masks are derived to guide diffusion inpainting, achieving strong results on pose-different hairstyle transfer.

  17. Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics

    cs.CY 2026-04 unverdicted novelty 5.0

    Case studies with blind UK residents and people from Kerala and Tamil Nadu demonstrate that community input at the systematization stage produces culturally grounded definitions of appropriateness for text-to-image mo...

  18. Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

    cs.AI 2026-02 reject novelty 5.0

    A capability-oriented multimodal judge benchmark and MCTS-based preference-data generation improve judge models on some benchmarks, but the paper's SOTA claim is not supported by its own numbers.

  19. Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications

    cs.CV 2026-06 unverdicted novelty 3.0

    Evaluation of MLLMs on assistive scenarios with a new egocentric benchmark called NetraLink provides a diagnostic of model strengths and limitations in object recognition, scene text, and multilingual understanding.