Pith. sign in

REVIEW 1 cited by

CTRN: Class-Temporal Relational Network for Action Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.13473 v2 pith:M2EXCPNG submitted 2021-10-26 cs.CV cs.AI

classification cs.CVcs.AI
keywords actionclass-temporalco-occurringctrndatasetsdetectionnetworktemporal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Action detection is an essential and challenging task, especially for densely labelled datasets of untrimmed videos. There are many real-world challenges in those datasets, such as composite action, co-occurring action, and high temporal variation of instance duration. For handling these challenges, we propose to explore both the class and temporal relations of detected actions. In this work, we introduce an end-to-end network: Class-Temporal Relational Network (CTRN). It contains three key components: (1) The Representation Transform Module filters the class-specific features from the mixed representations to build graph-structured data. (2) The Class-Temporal Module models the class and temporal relations in a sequential manner. (3) G-classifier leverages the privileged knowledge of the snippet-wise co-occurring action pairs to further improve the co-occurring action detection. We evaluate CTRN on three challenging densely labelled datasets and achieve state-of-the-art performance, reflecting the effectiveness and robustness of our method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

    cs.CV 2024-11 conditional novelty 5.0 of 10

    LLaVA-MR combines dense frame sampling, frame-difference key-frame selection, and variance-based token compression to improve generative MLLM video moment retrieval, reporting small SOTA gains over Mr. BLIP on three b...

Pith tools