Pith. sign in

REVIEW 14 cited by

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.09313 v2 pith:ZJK2577A submitted 2025-07-12 cs.CV

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models

classification cs.CV
keywords proactiveproactivevideoqaresponsessystemsinteractionpaucbenchmarkcomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

With the growing research focus on multimodal dialogue systems, the capability for proactive interaction is gradually gaining recognition. As an alternative to conventional turn-by-turn dialogue, users increasingly expect multimodal systems to be more initiative, for example, by autonomously determining the timing of multi-turn responses in real time during video playback. To facilitate progress in this emerging area, we introduce ProactiveVideoQA, the first comprehensive benchmark to evaluate a system's ability to engage in proactive interaction. Since model responses are generated at varying timestamps, we further propose PAUC, the first metric that accounts for the temporal dynamics of model responses. This enables a more accurate evaluation of systems operating in proactive settings. Through extensive benchmarking of various baseline systems on ProactiveVideoQA and a user study of human preferences, we show that PAUC is in better agreement with human preferences than traditional evaluation metrics, which typically only consider the textual content of responses. These findings demonstrate that PAUC provides a more faithful assessment of user experience in proactive interaction scenarios. Project homepage: https://github.com/yellow-binary-tree/ProactiveVideoQA

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

    cs.CV 2026-06 unverdicted novelty 7.0

    X-Stream benchmark shows state-of-the-art MLLMs achieve only about 50% on multi-stream video tasks and exhibit poor proactive ability.

  2. X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

    cs.CV 2026-06 unverdicted novelty 7.0

    X-Stream benchmark shows SOTA MLLMs score ~50% on concurrent multi-stream tasks and lack proactive ability, using a dual-verification pipeline to avoid single-stream bias.

  3. An Efficient Streaming Video Understanding Framework with Agentic Control

    cs.CV 2026-05 unverdicted novelty 7.0

    R3-Streaming uses cascaded control with age-aware memory forgetting and TB-GRPO reinforcement learning to reach SOTA scores of 57.92 on OVO-Bench and 76.36 on StreamingBench with 95-96% fewer visual tokens.

  4. Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

    cs.CV 2026-05 conditional novelty 7.0

    Omni-DuplexEval creates a new benchmark and LLM-as-a-Judge framework for real-time duplex omni-modal interaction, revealing that current models score below 40% overall and struggle especially with proactive responses.

  5. StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video

    cs.CV 2026-05 unverdicted novelty 7.0

    StreamPro introduces a benchmark and training method using CB-Stream Loss and GRPO to enable proactive decision-making in streaming videos, achieving 41.5 on StreamPro-Bench compared to 10.4 previously.

  6. Don't Pause! Every prediction matters in a streaming video

    cs.CV 2026-04 unverdicted novelty 7.0

    SPOT-Bench tests real-time streaming video perception with timeliness metrics, exposing limitations in current models and introducing AsynKV as an improved baseline.

  7. From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench

    cs.AI 2026-04 unverdicted novelty 7.0

    ProVoice-Bench is the first framework to evaluate proactive voice agents, revealing that state-of-the-art multimodal LLMs struggle with over-triggering and context-aware reasoning.

  8. VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

    cs.CV 2026-07 conditional novelty 6.0

    An open 4B video MLLM with inflated-3D ViT tokenization and adaptive streaming perception outperforms comparable open models on general, long-video, and streaming benchmarks while using fewer visual tokens.

  9. GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video

    cs.CV 2026-07 conditional novelty 6.0

    Current MLLMs can deliver procedural instructions in streaming video but systematically fail at real-time error detection and corrective coaching on the new GuideMe benchmark.

  10. IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

    cs.CV 2026-05 unverdicted novelty 6.0

    IPIBench evaluates MLLMs on interactive proactive intelligence in streaming videos, identifies unstable triggering and poor coordination, and proposes the training-free IPI-Agent framework to improve performance acros...

  11. An Efficient Streaming Video Understanding Framework with Agentic Control

    cs.CV 2026-05 unverdicted novelty 6.0

    R3-Streaming uses cascaded control, age-aware memory forgetting, and TB-GRPO reinforcement learning to reach SOTA scores on streaming video benchmarks while cutting visual token usage by 95-96%.

  12. Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

    cs.CV 2026-05 unverdicted novelty 6.0

    Omni-DuplexEval provides a new benchmark and automatic evaluation method for real-time duplex omni-modal interaction, showing state-of-the-art models reach only 39.6% overall and 20% on proactive reminders.

  13. Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents

    cs.AI 2025-10 conditional novelty 6.0

    On PROBE, a new synthetic benchmark for proactive workplace problem solving, the strongest LLM agents reach only ~40% end-to-end success, with root-cause identification as the main failure point.

  14. Proact-VL: A Proactive VideoLLM for Real-Time AI Companions

    cs.CV 2026-03 conditional novelty 5.0

    Proact-VL adds a learned 'when to speak' gate to a streaming video LLM and a 561-hour gaming dataset, reporting better timing and quality than prior live-commentary systems.