Pith. sign in

REVIEW 5 cited by

AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.21259 v4 pith:AWJMYB4Y submitted 2024-10-28 cs.CV cs.AI

classification cs.CVcs.AI
keywords evaluationlvlmsvisualautobench-vbenchmarkmodelsframeworklarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Vision-Language Models (LVLMs) have become essential for advancing the integration of visual and linguistic information. However, the evaluation of LVLMs presents significant challenges as the evaluation benchmark always demands lots of human cost for its construction, and remains static, lacking flexibility once constructed. Even though automatic evaluation has been explored in textual modality, the visual modality remains under-explored. As a result, in this work, we address a question: "Can LVLMs themselves be used to benchmark each other in the visual automatically domain?". We introduce AutoBench-V, an automated framework for serving evaluation on demand, i.e., benchmarking LVLMs based on specific aspects of model capability. AutoBench-V leverages text-to-image models to generate relevant image samples and then utilizes LVLMs to orchestrate visual question-answering (VQA) tasks, completing the evaluation process efficiently and flexibly. Through an extensive evaluation of nine popular LVLMs across five demanded user inputs (i.e., evaluation capabilities), the framework shows effectiveness and reliability.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    PolicyBench is the first large-scale US-China policy comprehension benchmark for LLMs with 21K cases, paired with PolicyMoE that performs best on application and structured reasoning tasks.

  2. REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    ReKey introduces a live benchmark protocol that regenerates visual keys in images to produce contamination-resilient VQA evaluations, showing 9.5-18.8 point higher scores on original items across eight VLMs.

  3. SkillGen: Verified Inference-Time Agent Skill Synthesis

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    SkillGen synthesizes auditable skills from agent trajectories via contrastive induction on successes and failures, then verifies net performance impact by comparing outcomes with and without the skill on identical tasks.

  4. Policy Learning from Large Vision-Language Model Feedback without Reward Modeling

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A reward-model-free offline RL method that uses VLM-generated pairwise preferences and contrastive preference learning to train manipulation policies.

  5. Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era

    cs.LG 2025-08 unverdicted novelty 1.0 of 10

    A tutorial proposal outlining how generative models can synthesize data across modalities for data mining, with no new research results.

Pith tools