Pith. sign in

REVIEW 9 cited by

VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.01071 v2 pith:DNFYP3IS submitted 2024-09-02 cs.CV cs.CL

classification cs.CVcs.CL
keywords videomemoryvideollambunderstandingacademicbridgeslongmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in large-scale video-language models have shown significant potential for real-time planning and detailed interactions. However, their high computational demands and the scarcity of annotated datasets limit their practicality for academic researchers. In this work, we introduce VideoLLaMB, a novel and efficient framework for long video understanding that leverages recurrent memory bridges and temporal memory tokens to enable seamless encoding of entire video sequences with preserved semantic continuity. Central to our approach is a SceneTiling algorithm that segments videos into coherent semantic units, facilitating robust understanding across tasks without requiring additional training. VideoLLaMB achieves state-of-the-art performance, surpassing existing models by 4.2 points on four VideoQA benchmarks and by 2.06 points on egocentric planning tasks. Notably, it maintains strong performance under extreme video length scaling (up to 8 times) and excels at fine-grained frame retrieval on our proposed Needle in a Video Haystack (NIAVH) benchmark. With linear GPU memory scaling, VideoLLaMB processes up to 320 frames using a single Nvidia A100 GPU, despite being trained on only 16 frames-offering an unprecedented balance of accuracy, scalability, and cost-effectiveness. This makes it highly accessible and practical for the academic community.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    OmniVTG creates a new large-scale open-world VTG dataset using iterative concept-gap filling and timestamped captioning, paired with a three-stage self-correction CoT paradigm that yields SOTA zero-shot results on fou...

  2. FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A query-conditioned, training-free frame selector that unifies relevance and diversity into a single volume-maximization objective improves keyframe recall and long-video question-answering accuracy.

  3. FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A training-free frame-selection method that weights frame embeddings by query relevance and maximizes the selected subspace's volume improves keyframe recall and VQA accuracy across eight MLLMs on Video-MME and LongVi...

  4. Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Q-GeoMem uses question-guided scoring to maintain a Fine-Grained Context Bank and Semantic-Geometric Evidence Bank, achieving SOTA on VSI-Bench and VSTI-Bench.

  5. Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    PVM adds a parallel branch to LVLMs that directly supplies visual embeddings to prevent attention decay over long generated sequences, yielding accuracy gains on reasoning tasks with minimal overhead.

  6. Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval

    cs.CV 2025-12 unverdicted novelty 6.0 of 10

    OneClip-RAG enables MLLMs to handle long videos via one-shot clip retrieval and unified chunking-retrieval, delivering performance gains like matching GPT-5 level on MLVU with high efficiency on standard GPUs.

  7. Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A hierarchical event memory with segment-tree proposals and a future prediction branch achieves state-of-the-art online video temporal grounding on TACoS, ActivityNet Captions, and MAD.

  8. Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    Question-guided dual geometric memories with relevance-novelty utility reportedly reach state-of-the-art video spatial reasoning on two in-domain and five out-of-distribution benchmarks.

  9. Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    PVM adds a parallel learnable branch to LVLMs that supplies visual embeddings on demand to structurally prevent attention decay and visual signal dilution during deep autoregressive generation.

Pith tools