Pith. sign in

REVIEW 2 cited by

Interactive Visual Grounding of Referring Expressions for Human-Robot Interaction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.03831 v1 pith:BIGOSGWK submitted 2018-06-11 cs.RO cs.CLcs.CV

classification cs.ROcs.CLcs.CV
keywords expressionsgroundinglanguageneuralobjectsingressnetworkreferring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents INGRESS, a robot system that follows human natural language instructions to pick and place everyday objects. The core issue here is the grounding of referring expressions: infer objects and their relationships from input images and language expressions. INGRESS allows for unconstrained object categories and unconstrained language expressions. Further, it asks questions to disambiguate referring expressions interactively. To achieve these, we take the approach of grounding by generation and propose a two-stage neural network model for grounding. The first stage uses a neural network to generate visual descriptions of objects, compares them with the input language expression, and identifies a set of candidate objects. The second stage uses another neural network to examine all pairwise relations between the candidates and infers the most likely referred object. The same neural networks are used for both grounding and question generation for disambiguation. Experiments show that INGRESS outperformed a state-of-the-art method on the RefCOCO dataset and in robot experiments with humans.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A disease-aware prompting method that reweights chest X-ray features using the model's own explainability map improves weakly-supervised visual grounding on three benchmarks.

  2. ACTLLM: Action Consistency Tuned Large Language Model

    cs.RO 2025-06 conditional novelty 5.0 of 10

    ACTLLM trains an LLM to jointly produce structured scene descriptions and actions, and reports improved compositional and zero-shot generalization on CLIPORT and VIMA.

Pith tools