Pith. sign in

REVIEW 20 cited by

Frame-Voyager: Learning to Query Frames for Video Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03226 v4 pith:T2BSSTOY submitted 2024-10-04 cs.CV cs.CL

Frame-Voyager: Learning to Query Frames for Video Large Language Models

classification cs.CV cs.CL
keywords frame-voyagerframevideocombinationsqueryvideo-llmvideo-llmsframes
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it impractical to input entire videos. Existing frame selection approaches, such as uniform frame sampling and text-frame retrieval, fail to account for the information density variations in the videos or the complex instructions in the tasks, leading to sub-optimal performance. In this paper, we propose Frame-Voyager that learns to query informative frame combinations, based on the given textual queries in the task. To train Frame-Voyager, we introduce a new data collection and labeling pipeline, by ranking frame combinations using a pre-trained Video-LLM. Given a video of M frames, we traverse its T-frame combinations, feed them into a Video-LLM, and rank them based on Video-LLM's prediction losses. Using this ranking as supervision, we train Frame-Voyager to query the frame combinations with lower losses. In experiments, we evaluate Frame-Voyager on four Video Question Answering benchmarks by plugging it into two different Video-LLMs. The experimental results demonstrate that Frame-Voyager achieves impressive results in all settings, highlighting its potential as a plug-and-play solution for Video-LLMs.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

    cs.CV 2026-05 unverdicted novelty 8.0

    EgoMemReason is a new benchmark showing that even the best multimodal models achieve only 39.6% accuracy on reasoning tasks that require integrating sparse evidence across days in egocentric video.

  2. ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

    cs.CV 2026-07 unverdicted novelty 7.0

    ReQuest introduces an uncertainty-driven question-adaptive keyframe selector with rethinking routing and adaptive NMS that boosts long-form video QA accuracy on Video-MME, MLVU, and LongVideoBench without fine-tuning ...

  3. GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs

    cs.CV 2026-05 unverdicted novelty 7.0

    GridProbe uses posterior probing on a KxK frame grid to adaptively select question-relevant frames, delivering up to 3.36x TFLOPs reduction with accuracy within 1.6 pp of the full-frame baseline on Video-MME-v2.

  4. CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding

    cs.CV 2026-05 unverdicted novelty 7.0

    CATS uses temporal curvature of query-frame relevance to select informative frames, achieving 93-95% of heavy multi-stage accuracy at 3-4% of the preprocessing cost on long-video benchmarks.

  5. OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning

    cs.CV 2026-04 unverdicted novelty 7.0

    OASIS organizes streaming video into hierarchical events and retrieves memory on-demand via intent-driven refinement to improve long-horizon accuracy and compositional reasoning with bounded token costs.

  6. FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

    cs.CV 2026-07 conditional novelty 6.0

    A training-free frame-selection method that weights frame embeddings by query relevance and maximizes the selected subspace's volume improves keyframe recall and VQA accuracy across eight MLLMs on Video-MME and LongVi...

  7. FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

    cs.CV 2026-07 conditional novelty 6.0

    A query-conditioned, training-free frame selector that unifies relevance and diversity into a single volume-maximization objective improves keyframe recall and long-video question-answering accuracy.

  8. SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

    cs.CV 2026-07 conditional novelty 6.0

    SkyVLaM introduces a temporal basis perceiver and adaptive dense selection to improve language-conditioned video segmentation in UAV scenes, and contributes the SkyVid dataset.

  9. Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors

    cs.CV 2026-07 conditional novelty 6.0

    Attention maps from a small MLLM can serve as a training-free, query-conditioned frame selector, improving long-video QA accuracy under fixed frame budgets.

  10. Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding

    cs.AI 2026-06 conditional novelty 6.0

    ReMem improves zero-shot long-video QA by combining LLM-based temporal granularity parsing, CLIP-based dual-semantic frame scoring, and structure-aware dynamic frame routing.

  11. PEEK: Picking Essential frames via Efficient Knowledge distillation

    cs.CV 2026-05 unverdicted novelty 6.0

    PEEK distills caption-conditioned frame relevance into a lightweight visual model, outperforming adaptive baselines on ActivityNet Captions and MSR-VTT especially at 1-2 frame budgets while adding only 5.2% overhead.

  12. Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs

    cs.LG 2026-05 unverdicted novelty 6.0

    MS-SFNN builds PDE solutions from element-wise products of outputs from d independent fixed-random-weight subnetworks with tunable scaling and cosine activations, then solves coefficients by least squares, claiming su...

  13. CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding

    cs.CV 2026-05 conditional novelty 6.0

    CREST selects video frames using curvature-adaptive non-maximum suppression on CLIP relevance scores, beating AKS by ~0.5% on two long-video QA benchmarks while using a fraction of MIRA's preprocessing cost.

  14. Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding

    cs.CV 2026-04 unverdicted novelty 6.0

    Q-Gate dynamically routes keyframe selection in long videos via query-modulated gating across visual grounding, global matching, and contextual alignment experts to improve MLLM performance.

  15. Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs

    cs.LG 2026-05 unverdicted novelty 5.0

    MS-SFNN encodes multi-scale Fourier features in a separable product of 1D cosine subnetworks and solves high-frequency PDEs via least-squares basis coefficients, claiming better accuracy than PINN and SV-SNN.

  16. Multi-Scale Separable Fourier Neural Networks for Solving High-Frequency PDEs

    cs.LG 2026-05 unverdicted novelty 5.0

    MS-SFNN encodes multi-scale Fourier features in a separable product of fixed-weight cosine subnetworks and solves for linear coefficients by least squares, claiming better accuracy than PINN and SV-SNN on high-frequency PDEs.

  17. Swift Sampling: Selecting Temporal Surprises via Taylor Series

    cs.CV 2026-05 unverdicted novelty 5.0

    Swift Sampling is a training-free frame selection method that uses Taylor expansions on video latent trajectories to pick temporally surprising frames, outperforming uniform sampling on long-video QA tasks.

  18. Scaling Video Understanding via Compact Latent Multi-Agent Collaboration

    cs.CV 2026-05 unverdicted novelty 5.0

    MACF decouples agent perception budgets from overall video length using latent token collaboration to scale video understanding in MLLMs beyond current limits.

  19. CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding

    cs.CV 2026-05 unverdicted novelty 4.0

    CREST uses local curvature of query-frame relevance over time to select informative frames, outperforming a lightweight baseline and approaching a costly pipeline at far lower preprocessing cost on long-video benchmarks.

  20. Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects

    cs.CL 2026-04 unverdicted novelty 4.0

    A survey that taxonomizes efficiency methods for LVLMs across the full inference pipeline, decouples the problem into information density, long-context attention, and memory limits, and outlines four future research f...