Pith. sign in

REVIEW 2 cited by

Referring Expression Comprehension: A Survey of Methods and Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.09554 v2 pith:MGCZWC3D submitted 2020-07-19 cs.CV cs.CL

classification cs.CVcs.CL
keywords expressionreferringcomprehensiondatasetsobjectproblemsurveybeen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Referring expression comprehension (REC) aims to localize a target object in an image described by a referring expression phrased in natural language. Different from the object detection task that queried object labels have been pre-defined, the REC problem only can observe the queries during the test. It thus more challenging than a conventional computer vision problem. This task has attracted a lot of attention from both computer vision and natural language processing community, and several lines of work have been proposed, from CNN-RNN model, modular network to complex graph-based model. In this survey, we first examine the state of the art by comparing modern approaches to the problem. We classify methods by their mechanism to encode the visual and textual modalities. In particular, we examine the common approach of joint embedding images and expressions to a common feature space. We also discuss modular architectures and graph-based models that interface with structured graph representation. In the second part of this survey, we review the datasets available for training and evaluating REC systems. We then group results according to the datasets, backbone models, settings so that they can be fairly compared. Finally, we discuss promising future directions for the field, in particular the compositional referring expression comprehension that requires longer reasoning chain to address.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Pruning visual tokens degrades visual grounding because position IDs become misaligned; preserving the original position IDs recovers most of the lost accuracy with no extra cost.

  2. Multimodal Human-Intent Modeling for Contextual Robot-to-Human Handovers of Arbitrary Objects

    cs.RO 2025-08 conditional novelty 5.0 of 10

    A gaze-plus-language pipeline enables a robot to select tabletop objects from a remote user's monitor and generate human-aware grasps for handover, with real-world tests on YCB objects.

Pith tools