Pith. sign in

REVIEW 1 cited by

Uncertainty-Aware Weakly Supervised Action Detection from Untrimmed Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.10703 v1 pith:BXRRMGPS submitted 2020-07-21 cs.CV

classification cs.CV
keywords actioninstancebeenlearningmethodmultiplerecognitionresults
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the recent advances in video classification, progress in spatio-temporal action recognition has lagged behind. A major contributing factor has been the prohibitive cost of annotating videos frame-by-frame. In this paper, we present a spatio-temporal action recognition model that is trained with only video-level labels, which are significantly easier to annotate. Our method leverages per-frame person detectors which have been trained on large image datasets within a Multiple Instance Learning framework. We show how we can apply our method in cases where the standard Multiple Instance Learning assumption, that each bag contains at least one instance with the specified label, is invalid using a novel probabilistic variant of MIL where we estimate the uncertainty of each prediction. Furthermore, we report the first weakly-supervised results on the AVA dataset and state-of-the-art results among weakly-supervised methods on UCF101-24.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stable Mean Teacher for Semi-supervised Video Action Detection

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Stable Mean Teacher with an Error Recovery module and a Difference of Pixels constraint improves semi-supervised video action detection, reaching near fully-supervised accuracy with 10-20% labels.

Pith tools