Pith. sign in

REVIEW 1 cited by

Grounded Situation Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.12058 v1 pith:YE5MBYF5 submitted 2020-03-26 cs.CV

classification cs.CV
keywords groundingssemanticentitiesgroundedsituationactivitybounding-boxdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce Grounded Situation Recognition (GSR), a task that requires producing structured semantic summaries of images describing: the primary activity, entities engaged in the activity with their roles (e.g. agent, tool), and bounding-box groundings of entities. GSR presents important technical challenges: identifying semantic saliency, categorizing and localizing a large and diverse set of entities, overcoming semantic sparsity, and disambiguating roles. Moreover, unlike in captioning, GSR is straightforward to evaluate. To study this new task we create the Situations With Groundings (SWiG) dataset which adds 278,336 bounding-box groundings to the 11,538 entity classes in the imsitu dataset. We propose a Joint Situation Localizer and find that jointly predicting situations and groundings with end-to-end training handily outperforms independent training on the entire grounding metric suite with relative gains between 8% and 32%. Finally, we show initial findings on three exciting future directions enabled by our models: conditional querying, visual chaining, and grounded semantic aware image retrieval. Code and data available at https://prior.allenai.org/projects/gsr.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes

    cs.CV 2025-01 conditional novelty 6.0 of 10

    3DCoMPaT200 expands compositional part-material 3D understanding to 200 shape categories and adds a text-based compositional shape retrieval benchmark.

Pith tools