Pith. sign in

REVIEW 3 cited by

HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08585 v2 pith:DNTULYCQ submitted 2025-03-11 cs.CV

classification cs.CV
keywords videocontexthierarqunderstandingframestreamacrosshierarchical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite advancements in multimodal large language models (MLLMs), current approaches struggle in medium-to-long video understanding due to frame and context length limitations. As a result, these models often depend on frame sampling, which risks missing key information over time and lacks task-specific relevance. To address these challenges, we introduce HierarQ, a task-aware hierarchical Q-Former based framework that sequentially processes frames to bypass the need for frame sampling, while avoiding LLM's context length limitations. We introduce a lightweight two-stream language-guided feature modulator to incorporate task awareness in video understanding, with the entity stream capturing frame-level object information within a short context and the scene stream identifying their broader interactions over longer period of time. Each stream is supported by dedicated memory banks which enables our proposed Hierachical Querying transformer (HierarQ) to effectively capture short and long-term context. Extensive evaluations on 10 video benchmarks across video understanding, question answering, and captioning tasks demonstrate HierarQ's state-of-the-art performance across most datasets, proving its robustness and efficiency for comprehensive video analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language

    cs.CL 2025-05 conditional novelty 6.0 of 10

    RAVEN uses query-conditioned token gating plus a new audio-video-sensor QA dataset to improve multimodal question answering, with reported gains of up to 14.5% over prior models.

  2. Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition

    cs.CV 2025-06 reject novelty 5.0 of 10

    A video action recognition framework generates current and next-step scene descriptions from detected context triples and combines their text embeddings with frame embeddings to classify activities.

  3. A Survey on Video Temporal Grounding with Multimodal Large Language Model

    cs.CV 2025-08 unverdicted novelty 3.0 of 10

    A taxonomized review of video temporal grounding with multimodal large language models, covering model roles, training paradigms, feature processing, benchmarks, and open problems.

Pith tools