Pith. sign in

REVIEW 1 cited by

Two-Stream Transformer Architecture for Long Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.01753 v1 pith:RIQCBDK4 submitted 2022-08-02 cs.CV cs.LGcs.MM

classification cs.CVcs.LGcs.MM
keywords datavideolongtaskstransformerunderstandingarchitectureattention
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pure vision transformer architectures are highly effective for short video classification and action recognition tasks. However, due to the quadratic complexity of self attention and lack of inductive bias, transformers are resource intensive and suffer from data inefficiencies. Long form video understanding tasks amplify data and memory efficiency problems in transformers making current approaches unfeasible to implement on data or memory restricted domains. This paper introduces an efficient Spatio-Temporal Attention Network (STAN) which uses a two-stream transformer architecture to model dependencies between static image features and temporal contextual features. Our proposed approach can classify videos up to two minutes in length on a single GPU, is data efficient, and achieves SOTA performance on several long video understanding tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A dual-branch VideoMamba with gated class-token fusion achieves 95.85% accuracy on a newly combined violence-detection benchmark and 74.13% on DVD, with about half the parameters and FLOPs of the CUE-Net baseline.

Pith tools