REVIEW 2 cited by
SoGAR: Self-supervised Spatiotemporal Attention-based Social Group Activity Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces a novel approach to Social Group Activity Recognition (SoGAR) using Self-supervised Transformers network that can effectively utilize unlabeled video data. To extract spatio-temporal information, we created local and global views with varying frame rates. Our self-supervised objective ensures that features extracted from contrasting views of the same video were consistent across spatio-temporal domains. Our proposed approach is efficient in using transformer-based encoders to alleviate the weakly supervised setting of group activity recognition. By leveraging the benefits of transformer models, our approach can model long-term relationships along spatio-temporal dimensions. Our proposed SoGAR method achieved state-of-the-art results on three group activity recognition benchmarks, namely JRDB-PAR, NBA, and Volleyball datasets, surpassing the current numbers in terms of F1-score, MCA, and MPCA metrics.
Forward citations
Cited by 2 Pith papers
-
Public Health Advocacy Dataset: A Dataset of Tobacco Usage Videos from Social Media
The PHAD dataset of 5,730 tobacco videos with rich metadata is new, but its classification results likely rely on text labels rather than visual understanding.
-
DEFEND: A Large-scale 1M Dataset and Foundation Model for Tobacco Addiction Prevention
A new 1M-image tobacco product dataset and a multimodal model that combines contrastive, coherence, and description losses, with reported gains over prior baselines.
Discussion (0). Continue with ORCID to comment.