Pith. sign in

REVIEW 3 cited by

QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08689 v1 pith:U77BEHHT submitted 2025-03-11 cs.CV

classification cs.CV
keywords tokenvisualquotaqueryassignmentexistingimportancequery-oriented
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in long video understanding typically mitigate visual redundancy through visual token pruning based on attention distribution. However, while existing methods employ post-hoc low-response token pruning in decoder layers, they overlook the input-level semantic correlation between visual tokens and instructions (query). In this paper, we propose QuoTA, an ante-hoc training-free modular that extends existing large video-language models (LVLMs) for visual token assignment based on query-oriented frame-level importance assessment. The query-oriented token selection is crucial as it aligns visual processing with task-specific requirements, optimizing token budget utilization while preserving semantically relevant content. Specifically, (i) QuoTA strategically allocates frame-level importance scores based on query relevance, enabling one-time visual token assignment before cross-modal interactions in decoder layers, (ii) we decouple the query through Chain-of-Thoughts reasoning to facilitate more precise LVLM-based frame importance scoring, and (iii) QuoTA offers a plug-and-play functionality that extends to existing LVLMs. Extensive experimental results demonstrate that implementing QuoTA with LLaVA-Video-7B yields an average performance improvement of 3.2% across six benchmarks (including Video-MME and MLVU) while operating within an identical visual token budget as the baseline. Codes are open-sourced at https://github.com/MAC-AutoML/QuoTA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

    cs.AI 2026-03 conditional novelty 6.5 of 10

    SocialOmni jointly evaluates speaker ID, turn-entry timing, and interruption phrasing on 12 OLMs and finds perception accuracy decouples from socially appropriate generation.

  2. FlexSelect: Flexible Token Selection for Efficient Long Video Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FlexSelect selects a small fraction of query-relevant visual tokens using attention from an intermediate layer, improving long-video accuracy and inference speed across multiple VideoLLMs.

  3. APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate Attention

    cs.CV 2026-01 conditional novelty 4.0 of 10

    SPAVA compresses each GPU's video context into essential key-values for sharing, making multi-GPU long-video prefilling 1.18x faster than APB and 12.72x faster than single-GPU FlashAttention with near-equal accuracy.

Pith tools