Pith. sign in

REVIEW 5 cited by

Is Your Video Language Model a Reliable Judge?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.05977 v1 pith:JD52WJBD submitted 2025-03-07 cs.CV cs.AI

Is Your Video Language Model a Reliable Judge?

classification cs.CV cs.AI
keywords vlmsevaluationreliabilityreliablejudgesmethodscollectivemodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

As video language models (VLMs) gain more applications in various scenarios, the need for robust and scalable evaluation of their performance becomes increasingly critical. The traditional human expert-based evaluation of VLMs has limitations in consistency and scalability, which sparked interest in automatic methods such as employing VLMs to evaluate VLMs. However, the reliability of VLMs as judges remains underexplored. Existing methods often rely on a single VLM as the evaluator. However, this approach can be unreliable or biased because such a model may lack the ability to fully understand the content and may have inherent biases, ultimately compromising evaluation reliability. A remedy is to apply the principle of collective thoughts, aggregating evaluations from multiple VLMs to enhance reliability. This study investigates the efficacy of such approaches, particularly when the pool of judges includes both reliable and unreliable models. Our findings reveal that incorporating collective judgments from such a mixed pool does not necessarily improve the accuracy of the final evaluation. The inclusion of less reliable judges can introduce noise, undermining the overall reliability of the outcomes. To explore the factors that impact evaluation reliability, we fine-tune an underperforming VLM judge, Video-LLaVA, and observe that improved understanding ability alone is insufficient to make VLM judges more reliable. These findings stress the limitations of collective thought approaches and highlight the need for more advanced methods that can account for the reliability of individual models. Our study promotes the development of more reliable evaluation methods for VLMs

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

    eess.IV 2026-07 conditional novelty 7.0

    A generative video codec that transmits selected latent anchors and a text prompt, then synthesizes the rest with a diffusion transformer, reaches <0.005 bpp with strong perceptual quality.

  2. A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

    cs.CV 2026-06 conditional novelty 6.5

    CrashTwin recovers metric-scale crash dynamics from monocular rollouts and shows that strong visual scores routinely mask large momentum, energy, and identity violations in world models.

  3. SGA: Plug&Play Geometric Verification for Educational Video Synthesis

    cs.AI 2026-07 conditional novelty 6.0

    An intercept-and-refine agent that symbolically checks Manim animation code for geometric collisions improves the authors' rendering-free MVQS layout score in 7 of 8 LLM×pipeline configurations, with the metric unvali...

  4. A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

    cs.CV 2026-06 unverdicted novelty 6.0

    CrashTwin is a new benchmark framework that exposes physical violations in state-of-the-art world models during multi-agent collisions despite high visual quality.

  5. EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

    cs.CV 2025-10 conditional novelty 6.0

    Most Video-LLMs answer temporal questions far less consistently when the same event is shown from ego and exo views, and a GRPO variant with a reasoning-similarity reward partially closes the gap.