Pith. sign in

REVIEW

Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.03885 v1 pith:J3Y4WSDK submitted 2019-04-08 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords descriptionsidentifyingspatio-temporalmodulesvideosappearancedataground
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a new task, the grounding of spatio-temporal identifying descriptions in videos. Previous work suggests potential bias in existing datasets and emphasizes the need for a new data creation schema to better model linguistic structure. We introduce a new data collection scheme based on grammatical constraints for surface realization to enable us to investigate the problem of grounding spatio-temporal identifying descriptions in videos. We then propose a two-stream modular attention network that learns and grounds spatio-temporal identifying descriptions based on appearance and motion. We show that motion modules help to ground motion-related words and also help to learn in appearance modules because modular neural networks resolve task interference between modules. Finally, we propose a future challenge and a need for a robust system arising from replacing ground truth visual annotations with automatic video object detector and temporal event localization.

Discussion (0). Continue with ORCID to comment.

Pith tools