Pith. sign in

REVIEW 2 cited by

SoGAR: Self-supervised Spatiotemporal Attention-based Social Group Activity Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.06310 v4 pith:GRM2WLEZ submitted 2023-04-27 cs.CV

classification cs.CV
keywords activitygrouprecognitionapproachself-supervisedsogarspatio-temporalproposed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces a novel approach to Social Group Activity Recognition (SoGAR) using Self-supervised Transformers network that can effectively utilize unlabeled video data. To extract spatio-temporal information, we created local and global views with varying frame rates. Our self-supervised objective ensures that features extracted from contrasting views of the same video were consistent across spatio-temporal domains. Our proposed approach is efficient in using transformer-based encoders to alleviate the weakly supervised setting of group activity recognition. By leveraging the benefits of transformer models, our approach can model long-term relationships along spatio-temporal dimensions. Our proposed SoGAR method achieved state-of-the-art results on three group activity recognition benchmarks, namely JRDB-PAR, NBA, and Volleyball datasets, surpassing the current numbers in terms of F1-score, MCA, and MPCA metrics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Public Health Advocacy Dataset: A Dataset of Tobacco Usage Videos from Social Media

    cs.CV 2024-11 reject novelty 6.0 of 10

    The PHAD dataset of 5,730 tobacco videos with rich metadata is new, but its classification results likely rely on text labels rather than visual understanding.

  2. DEFEND: A Large-scale 1M Dataset and Foundation Model for Tobacco Addiction Prevention

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A new 1M-image tobacco product dataset and a multimodal model that combines contrastive, coherence, and description losses, with reported gains over prior baselines.

Pith tools