REVIEW 1 cited by
Space-Time Memory Network for Sounding Object Localization in Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Space-Time Memory Network for Sounding Object Localization in Videos
read the original abstract
Leveraging temporal synchronization and association within sight and sound is an essential step towards robust localization of sounding objects. To this end, we propose a space-time memory network for sounding object localization in videos. It can simultaneously learn spatio-temporal attention over both uni-modal and cross-modal representations from audio and visual modalities. We show and analyze both quantitatively and qualitatively the effectiveness of incorporating spatio-temporal learning in localizing audio-visual objects. We demonstrate that our approach generalizes over various complex audio-visual scenes and outperforms recent state-of-the-art methods.
Forward citations
Cited by 1 Pith paper
-
SceneBind: Binding What and Where Across Vision, Audio and Language
A scene is represented as a global semantic embedding plus object-centric semantic-spatial slots (azimuth, elevation, distance, confidence), which improves cross-modal retrieval and enables zero-shot audio-visual loca...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.