Pith. sign in

REVIEW 12 cited by

VideoMamba: State Space Model for Efficient Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06977 v2 pith:C7BIGHQR submitted 2024-03-11 cs.CV

classification cs.CV
keywords videounderstandingvideomambaefficientdomainextensivelong-termmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Addressing the dual challenges of local redundancy and global dependencies in video understanding, this work innovatively adapts the Mamba to the video domain. The proposed VideoMamba overcomes the limitations of existing 3D convolution neural networks and video transformers. Its linear-complexity operator enables efficient long-term modeling, which is crucial for high-resolution long video understanding. Extensive evaluations reveal VideoMamba's four core abilities: (1) Scalability in the visual domain without extensive dataset pretraining, thanks to a novel self-distillation technique; (2) Sensitivity for recognizing short-term actions even with fine-grained motion differences; (3) Superiority in long-term video understanding, showcasing significant advancements over traditional feature-based models; and (4) Compatibility with other modalities, demonstrating robustness in multi-modal contexts. Through these distinct advantages, VideoMamba sets a new benchmark for video understanding, offering a scalable and efficient solution for comprehensive video understanding. All the code and models are available at https://github.com/OpenGVLab/VideoMamba.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Consistent and Editable: A Balanced Framework for Text-Guided Video Editing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    EquiEdit balances temporal consistency and editability in diffusion-based text-guided video editing via a temporal Mamba module and spectral noise injection on initial latents.

  2. Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A prior-guided video-level regression with adaptive interval selection and full parameter sharing sets new state-of-the-art results on micro-expression spotting and recognition benchmarks.

  3. Animate-X++: Universal Character Image Animation with Dynamic Backgrounds

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Animate-X++ turns cartoon images into pose-driven animations with text-controlled moving backgrounds, claiming state-of-the-art results on a new synthetic anthropomorphic benchmark.

  4. HydraMamba: Multi-Head State Space Model for Global Point Cloud Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A state space model based point cloud network with shuffled Hilbert serialization, a convolutional bidirectional S6 branch, and multi-head S6 achieves new top scores on ModelNet40, ShapeNet, S3DIS, and ScanObjectNN.

  5. Few-Shot Object Detection via Spatial-Channel State Space Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A Mamba-based channel sequence model combined with spatial attention improves few-shot object detection on VOC and COCO.

  6. QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    QuarterMap prunes spatial activations before VMamba's four-directional scan and upsamples after, yielding up to 1.11x throughput with under 1% accuracy loss on ImageNet classification.

  7. DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DySS combines state-space feature learning with dynamic query merging and pruning to improve both accuracy and speed for camera-based 3D detection on nuScenes.

  8. VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

    cs.CV 2026-07 conditional novelty 4.0 of 10

    VideoSEMA uses SEMA spatial attention plus softmax temporal attention to outperform larger video transformers and Mamba models on K400/SSv2 and degrade less at 1024² resolution.

  9. Time-Scaling State-Space Models for Dense Video Captioning

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A state-space model with transfer state processes videos chunk by chunk, carries the hidden state forward, and performs online dense video captioning with the same state as full-sequence processing.

  10. Hierarchical Spatio-temporal Segmentation Network for Ejection Fraction Estimation in Echocardiography Videos

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A hybrid convolutional-Mamba network segments left ventricular contours in echocardiography videos and reports improved ejection fraction correlation on three benchmarks.

  11. Comparing Learning Paradigms for Egocentric Video Summarization

    cs.CV 2025-06 reject novelty 4.0 of 10

    A prompt-engineered GPT-4o (quality score 64.95) outperformed Shotluck Holmes (61.19) and TAC-SUM (58.43) on a 21-video egocentric summary evaluation, though all scores were modest.

  12. Straightforward Bayesian A/B testing with Dirichlet posteriors

    stat.ME 2025-08 unverdicted novelty 3.0 of 10

    The submission is internally inconsistent: the abstract promises a Bayesian A/B testing method, but the full text is a different computer vision paper, leaving the claimed result unevaluable.

Pith tools