Pith. sign in

REVIEW 3 cited by

Interpolating Video-LLMs: Toward Longer-sequence LMMs in a Training-free Manner

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.12963 v2 pith:6DUYHSIR submitted 2024-09-19 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords video-llmsvideotraining-freealignmentchallengescontentencoderfixed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Advancements in Large Language Models (LLMs) inspire various strategies for integrating video modalities. A key approach is Video-LLMs, which incorporate an optimizable interface linking sophisticated video encoders to LLMs. However, due to computation and data limitations, these Video-LLMs are typically pre-trained to process only short videos, limiting their broader application for understanding longer video content. Additionally, fine-tuning Video-LLMs to handle longer videos is cost-prohibitive. Consequently, it becomes essential to explore the interpolation of Video-LLMs under a completely training-free setting. In this paper, we first identify the primary challenges in interpolating Video-LLMs: (1) the video encoder and modality alignment projector are fixed, preventing the integration of additional frames into Video-LLMs, and (2) the LLM backbone is limited in its content length capabilities, which complicates the processing of an increased number of video tokens. To address these challenges, we propose a specific INTerPolation method for Video-LLMs (INTP-Video-LLMs). We introduce an alternative video token rearrangement technique that circumvents limitations imposed by the fixed video encoder and alignment projector. Furthermore, we introduce a training-free LLM context window extension method to enable Video-LLMs to understand a correspondingly increased number of visual tokens.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DATE: Dynamic Absolute Time Enhancement for Long Video Understanding

    cs.CV 2025-09 conditional novelty 6.0 of 10

    DATE combines inference-time timestamp token injection with a caption-rewritten, temporally regularized CLIP sampling strategy to improve absolute time reasoning and event localization in long videos.

  2. VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos

    cs.IR 2025-02 conditional novelty 5.0 of 10

    VideoRAG combines graph-based text indexing with multimodal visual embeddings to answer questions across multi-hour video collections, supported by a new 134-hour benchmark.

  3. freePruner: A Training-free Approach for Large Multimodal Model Acceleration

    cs.CV 2024-11 conditional novelty 3.0 of 10

    freePruner selects 50 percent of visual tokens using attention-based importance and keeps accuracy close to the original model, enabling a training-free about 2x acceleration for LMMs.

Pith tools