Pith. sign in

REVIEW 2 cited by

Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16866 v1 pith:VG73B52F submitted 2024-06-24 cs.CV

classification cs.CV
keywords refcocoref-l4modelsbenchmarklargelmmsrefcocogreferring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Referring expression comprehension (REC) involves localizing a target instance based on a textual description. Recent advancements in REC have been driven by large multimodal models (LMMs) like CogVLM, which achieved 92.44% accuracy on RefCOCO. However, this study questions whether existing benchmarks such as RefCOCO, RefCOCO+, and RefCOCOg, capture LMMs' comprehensive capabilities. We begin with a manual examination of these benchmarks, revealing high labeling error rates: 14% in RefCOCO, 24% in RefCOCO+, and 5% in RefCOCOg, which undermines the authenticity of evaluations. We address this by excluding problematic instances and reevaluating several LMMs capable of handling the REC task, showing significant accuracy improvements, thus highlighting the impact of benchmark noise. In response, we introduce Ref-L4, a comprehensive REC benchmark, specifically designed to evaluate modern REC models. Ref-L4 is distinguished by four key features: 1) a substantial sample size with 45,341 annotations; 2) a diverse range of object categories with 365 distinct types and varying instance scales from 30 to 3,767; 3) lengthy referring expressions averaging 24.2 words; and 4) an extensive vocabulary comprising 22,813 unique words. We evaluate a total of 24 large models on Ref-L4 and provide valuable insights. The cleaned versions of RefCOCO, RefCOCO+, and RefCOCOg, as well as our Ref-L4 benchmark and evaluation code, are available at https://github.com/JierunChen/Ref-L4.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Under frozen perception, an explicit query-region alignment hook plus perception-grounded abstention reduces conditional binding errors and hallucinations, with an exact Acc=(1-SeeErr)(1-SayErr) decomposition.

  2. The Mechanistic Emergence of Symbol Grounding in Language Models

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Symbol grounding emerges in Transformers and state-space models through middle-layer 'aggregate' attention heads that connect environmental cues to words, but not in unidirectional LSTMs.

Pith tools