Pith. sign in

REVIEW 1 cited by

Learn2Augment: Learning to Composite Videos for Data Augmentation in Action Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.04790 v2 pith:U7ANNP5S submitted 2022-06-09 cs.CV

classification cs.CV
keywords augmentationdatavideoactionaugmentedrecognitioncompositeimprovements
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We address the problem of data augmentation for video action recognition. Standard augmentation strategies in video are hand-designed and sample the space of possible augmented data points either at random, without knowing which augmented points will be better, or through heuristics. We propose to learn what makes a good video for action recognition and select only high-quality samples for augmentation. In particular, we choose video compositing of a foreground and a background video as the data augmentation process, which results in diverse and realistic new samples. We learn which pairs of videos to augment without having to actually composite them. This reduces the space of possible augmentations, which has two advantages: it saves computational cost and increases the accuracy of the final trained classifier, as the augmented pairs are of higher quality than average. We present experimental results on the entire spectrum of training settings: few-shot, semi-supervised and fully supervised. We observe consistent improvements across all of them over prior work and baselines on Kinetics, UCF101, HMDB51, and achieve a new state-of-the-art on settings with limited data. We see improvements of up to 8.6% in the semi-supervised setting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synthetic Human Action Video Data Generation with Pose Transfer

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Synthetic action videos generated by pose-transferring real clips onto novel 3D avatars improve action recognition accuracy when added to real training data.

Pith tools