Pith. sign in

REVIEW 2 cited by

Human Action Localization with Sparse Spatial Supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1605.05197 v2 pith:TERMEYJY submitted 2016-05-17 cs.CV

classification cs.CV
keywords actionsupervisionhumanlocalizationspatialsparsetubesapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce an approach for spatio-temporal human action localization using sparse spatial supervision. Our method leverages the large amount of annotated humans available today and extracts human tubes by combining a state-of-the-art human detector with a tracking-by-detection approach. Given these high-quality human tubes and temporal supervision, we select positive and negative tubes with very sparse spatial supervision, i.e., only one spatially annotated frame per instance. The selected tubes allow us to effectively learn a spatio-temporal action detector based on dense trajectories or CNNs. We conduct experiments on existing action localization benchmarks: UCF-Sports, J-HMDB and UCF-101. Our results show that our approach, despite using sparse spatial supervision, performs on par with methods using full supervision, i.e., one bounding box annotation per frame. To further validate our method, we introduce DALY (Daily Action Localization in YouTube), a dataset for realistic action localization in space and time. It contains high quality temporal and spatial annotations for 3.6k instances of 10 actions in 31 hours of videos (3.3M frames). It is an order of magnitude larger than existing datasets, with more diversity in appearance and long untrimmed videos.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interacted Object Grounding in Spatio-Temporal Human-Object Interactions

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A new benchmark and task for grounding interacted objects in videos, with a 4D-QA method that achieves 23.38 mAP@0.5 on 1,098 object classes.

  2. Video Understanding by Design: How Datasets Shape Video Models

    cs.CV 2025-09 reject novelty 4.0 of 10

    A dataset-centric framework that explains video architectures as responses to structural properties of benchmark datasets.

Pith tools