Pith. sign in

REVIEW 27 cited by

VideoMamba: State Space Model for Efficient Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06977 v2 pith:C7BIGHQR submitted 2024-03-11 cs.CV

classification cs.CV
keywords videounderstandingvideomambaefficientdomainextensivelong-termmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Addressing the dual challenges of local redundancy and global dependencies in video understanding, this work innovatively adapts the Mamba to the video domain. The proposed VideoMamba overcomes the limitations of existing 3D convolution neural networks and video transformers. Its linear-complexity operator enables efficient long-term modeling, which is crucial for high-resolution long video understanding. Extensive evaluations reveal VideoMamba's four core abilities: (1) Scalability in the visual domain without extensive dataset pretraining, thanks to a novel self-distillation technique; (2) Sensitivity for recognizing short-term actions even with fine-grained motion differences; (3) Superiority in long-term video understanding, showcasing significant advancements over traditional feature-based models; and (4) Compatibility with other modalities, demonstrating robustness in multi-modal contexts. Through these distinct advantages, VideoMamba sets a new benchmark for video understanding, offering a scalable and efficient solution for comprehensive video understanding. All the code and models are available at https://github.com/OpenGVLab/VideoMamba.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mamba Drafters for Speculative Decoding

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Mamba-based drafters can match self-speculation throughput with lower memory and cross-model flexibility.

  2. Sparsified State-Space Models are Efficient Highway Networks

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Simba prunes tokens hierarchically in pre-trained SSMs, creating sparse upper layers that act as highways, improving the accuracy-FLOPs trade-off and long-context perplexity.

  3. AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation

    cs.CV 2025-01 reject novelty 6.0 of 10

    AVS-Mamba applies Mamba with temporal and cross-modal scanning to audio-visual segmentation, reporting top scores on AVSBench-object but not on AVSBench-semantic with the stronger backbone.

  4. MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MambaVO improves deep visual odometry by adding Mamba-based matching refinement and a smoothed training objective, achieving state-of-the-art absolute trajectory error on EuRoC, TUM-RGBD, KITTI, and TartanAir.

  5. Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking

    cs.CV 2024-12 conditional novelty 6.0 of 10

    STTrack, a video tracker with a temporal state generator and mamba fusion modules, reports state-of-the-art success and accuracy numbers on five multimodal tracking benchmarks.

  6. MambaLCT: Boosting Tracking via Long-term Context State Space Model

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MambaLCT combines a unidirectional Mamba scan of all past search frames with a Transformer encoder to build long-term context for single-object tracking.

  7. MambaVLT: Time-Evolving Multimodal State Space Model for Vision-Language Tracking

    cs.CV 2024-11 conditional novelty 6.0 of 10

    MambaVLT applies Mamba state space models to vision-language tracking with a time-evolving memory, beating several baselines on three of four benchmarks.

  8. Consistent and Editable: A Balanced Framework for Text-Guided Video Editing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    EquiEdit balances temporal consistency and editability in diffusion-based text-guided video editing via a temporal Mamba module and spectral noise injection on initial latents.

  9. Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A prior-guided video-level regression with adaptive interval selection and full parameter sharing sets new state-of-the-art results on micro-expression spotting and recognition benchmarks.

  10. Animate-X++: Universal Character Image Animation with Dynamic Backgrounds

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Animate-X++ turns cartoon images into pose-driven animations with text-controlled moving backgrounds, claiming state-of-the-art results on a new synthetic anthropomorphic benchmark.

  11. HydraMamba: Multi-Head State Space Model for Global Point Cloud Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A state space model based point cloud network with shuffled Hilbert serialization, a convolutional bidirectional S6 branch, and multi-head S6 achieves new top scores on ModelNet40, ShapeNet, S3DIS, and ScanObjectNN.

  12. Few-Shot Object Detection via Spatial-Channel State Space Model

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A Mamba-based channel sequence model combined with spatial attention improves few-shot object detection on VOC and COCO.

  13. QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    QuarterMap prunes spatial activations before VMamba's four-directional scan and upsamples after, yielding up to 1.11x throughput with under 1% accuracy loss on ImageNet classification.

  14. Moment Sampling in Video LLMs for Long-Form Video QA

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Moment sampling uses a text-to-video moment retrieval model to select question-relevant frames, improving long-form VideoQA accuracy by about one to two points over uniform sampling.

  15. DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DySS combines state-space feature learning with dynamic query merging and pruning to improve both accuracy and speed for camera-based 3D detection on nuScenes.

  16. MV-GMN: State Space Model for Multi-View Action Recognition

    cs.CV 2025-01 conditional novelty 5.0 of 10

    MV-GMN, a state-space model with graph convolution, reports state-of-the-art accuracies on NTU RGB+D and PKU-MMD action recognition benchmarks.

  17. H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A hierarchical Mamba adapter (C-Mamba and Q-Mamba) improves multimodal LLM video understanding in autonomous driving, achieving SOTA 66.9% mIoU on DRAMA risk localization.

  18. MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection

    cs.CV 2024-12 conditional novelty 5.0 of 10

    MamKPD, a Mamba-based 2D keypoint detector with a contextual modeling module, reports 77.3% AP on COCO at 1492 FPS and top MPII accuracy.

  19. Deformable Mamba for Wide Field of View Segmentation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    On five wide-FoV segmentation benchmarks, a Mamba-plus-deformable-convolution decoder beats common segmentation heads with the same backbones while cutting decoder FLOPs by roughly 97% versus UperHead.

  20. VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

    cs.CV 2026-07 conditional novelty 4.0 of 10

    VideoSEMA uses SEMA spatial attention plus softmax temporal attention to outperform larger video transformers and Mamba models on K400/SSv2 and degrade less at 1024² resolution.

  21. Time-Scaling State-Space Models for Dense Video Captioning

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A state-space model with transfer state processes videos chunk by chunk, carries the hidden state forward, and performs online dense video captioning with the same state as full-sequence processing.

  22. Hierarchical Spatio-temporal Segmentation Network for Ejection Fraction Estimation in Echocardiography Videos

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A hybrid convolutional-Mamba network segments left ventricular contours in echocardiography videos and reports improved ejection fraction correlation on three benchmarks.

  23. Comparing Learning Paradigms for Egocentric Video Summarization

    cs.CV 2025-06 reject novelty 4.0 of 10

    A prompt-engineered GPT-4o (quality score 64.95) outperformed Shotluck Holmes (61.19) and TAC-SUM (58.43) on a 21-video egocentric summary evaluation, though all scores were modest.

  24. Multi-modal Collaborative Optimization and Expansion Network for Event-assisted Single-eye Expression Recognition

    cs.CV 2025-05 conditional novelty 4.0 of 10

    MCO-E Net fuses event and RGB eye data via a jointly optimized Mamba and a heterogeneous MoE, achieving 91.3% WAR and 91.9% UAR on the SEE dataset.

  25. V"Mean"ba: Visual State Space Models only need 1 hidden dimension

    cs.CV 2024-12 conditional novelty 4.0 of 10

    VMeanba speeds up VMamba's selective scan by averaging its internal channel dimension down to 1, achieving up to 1.12x end-to-end speedup with under 3% accuracy loss on ImageNet and ADE20k.

  26. MambaNUT: Nighttime UAV Tracking via Mamba-based Adaptive Curriculum Learning

    cs.CV 2024-12 conditional novelty 4.0 of 10

    MambaNUT uses a Mamba backbone with an adaptive curriculum learning schedule to achieve efficient state-of-the-art nighttime UAV tracking.

  27. Straightforward Bayesian A/B testing with Dirichlet posteriors

    stat.ME 2025-08 unverdicted novelty 3.0 of 10

    The submission is internally inconsistent: the abstract promises a Bayesian A/B testing method, but the full text is a different computer vision paper, leaving the claimed result unevaluable.

Pith tools