Pith. sign in

REVIEW 1 cited by

Differentiable Frequency-based Disentanglement for Aerial Video Action Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.09194 v2 pith:AOYIHT32 submitted 2022-09-15 cs.CV

classification cs.CV
keywords humandifferentiablerecognitionvideoactionapproachdatasetdisentangled
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a learning algorithm for human activity recognition in videos. Our approach is designed for UAV videos, which are mainly acquired from obliquely placed dynamic cameras that contain a human actor along with background motion. Typically, the human actors occupy less than one-tenth of the spatial resolution. Our approach simultaneously harnesses the benefits of frequency domain representations, a classical analysis tool in signal processing, and data driven neural networks. We build a differentiable static-dynamic frequency mask prior to model the salient static and dynamic pixels in the video, crucial for the underlying task of action recognition. We use this differentiable mask prior to enable the neural network to intrinsically learn disentangled feature representations via an identity loss function. Our formulation empowers the network to inherently compute disentangled salient features within its layers. Further, we propose a cost-function encapsulating temporal relevance and spatial content to sample the most important frame within uniformly spaced video segments. We conduct extensive experiments on the UAV Human dataset and the NEC Drone dataset and demonstrate relative improvements of 5.72% - 13.00% over the state-of-the-art and 14.28% - 38.05% over the corresponding baseline model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FTMoMamba: Motion Generation with Frequency and Text State Space Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    FTMoMamba achieves FID 0.181 on HumanML3D by injecting frequency-domain features into the state transition matrix and text features into the output matrix of a Mamba-based diffusion denoiser.

Pith tools