Pith. sign in

REVIEW 1 cited by

VLUE: A Multi-Task Benchmark for Evaluating Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.15237 v1 pith:LEJLBWEN submitted 2022-05-30 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords modelsvision-languagebenchmarkefficiency-performancetrade-offgeneralizationimagesmeasuring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in vision-language pre-training (VLP) have demonstrated impressive performance in a range of vision-language (VL) tasks. However, there exist several challenges for measuring the community's progress in building general multi-modal intelligence. First, most of the downstream VL datasets are annotated using raw images that are already seen during pre-training, which may result in an overestimation of current VLP models' generalization ability. Second, recent VLP work mainly focuses on absolute performance but overlooks the efficiency-performance trade-off, which is also an important indicator for measuring progress. To this end, we introduce the Vision-Language Understanding Evaluation (VLUE) benchmark, a multi-task multi-dimension benchmark for evaluating the generalization capabilities and the efficiency-performance trade-off (``Pareto SOTA'') of VLP models. We demonstrate that there is a sizable generalization gap for all VLP models when testing on out-of-distribution test sets annotated on images from a more diverse distribution that spreads across cultures. Moreover, we find that measuring the efficiency-performance trade-off of VLP models leads to complementary insights for several design choices of VLP. We release the VLUE benchmark to promote research on building vision-language models that generalize well to more diverse images and concepts unseen during pre-training, and are practical in terms of efficiency-performance trade-off.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A machine-translated, quality-filtered Urdu caption set covering 59,000 MS COCO images with 319,000 captions, presented as the largest public Urdu image-caption dataset.

Pith tools