Pith. sign in

REVIEW 1 cited by

Towards Understanding Sample Variance in Visually Grounded Language Generation: Evaluations and Observations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.03644 v1 pith:UTJQLXLH submitted 2020-10-07 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords datasetsgenerationgroundedlanguagevariancevisuallydesigndifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings. To do this, it is critical to ensure that our evaluation protocols are correct, and benchmarks are reliable. In this work, we set forth to design a set of experiments to understand an important but often ignored problem in visually grounded language generation: given that humans have different utilities and visual attention, how will the sample variance in multi-reference datasets affect the models' performance? Empirically, we study several multi-reference datasets and corresponding vision-and-language tasks. We show that it is of paramount importance to report variance in experiments; that human-generated references could vary drastically in different datasets/tasks, revealing the nature of each task; that metric-wise, CIDEr has shown systematically larger variances than others. Our evaluations on reference-per-instance shed light on the design of reliable datasets in the future.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A factorized autoregressive decoder, shared across video segments with cross-segment masking, produces denser, more localized captions online while saving about 20 percent compute versus a global decoder.

Pith tools