Pith. sign in

REVIEW 2 major objections 1 minor 62 references

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

T0 review · 2 major / 1 minor · reviewed 2026-07-04 · grok-4.3

Pith's one-line read Duplex MLLMs score just 39.6 percent overall on real-time interaction tasks

desk verdict New benchmark for real-time duplex MLLM eval with two scenarios and LLM judge, but no stats shown for the judge's human alignment. read the letter →

arxiv 2605.17360 v2 pith:D7D26HBY submitted 2026-05-17 cs.CV

classification cs.CV
keywords real-timeduplexmultimodalLLMsbenchmarkevaluationproactivereminderLLMjudgeomni-modalinteractionstreaminginputs
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces Omni-DuplexEval to evaluate real-time duplex omni-modal interactions that current MLLMs cannot handle in offline settings. It features Real-Time Description for time-aligned responses and Proactive Reminder for spotting key events. The automatic LLM-as-Judge evaluates content and timing with timestamp awareness. State-of-the-art models perform poorly, topping out at 39.6 percent overall and 20 percent on proactive tasks. This shows models have trouble deciding both when to respond and what content to produce.

What carries the argument

The Omni-DuplexEval benchmark consisting of two scenarios—Real-Time Description and Proactive Reminder—along with its LLM-as-a-Judge automatic evaluation framework that uses timestamp-aware and sequential reasoning.

What would settle it

A model achieving 70% or higher overall score on Omni-DuplexEval that also matches human ratings on timing and content in direct comparisons.

Watch

Extended reading notes

Core claim

Omni-DuplexEval reveals that even leading duplex MLLMs achieve only 39.6% overall performance, with just 20.0% on Proactive Reminder, because they struggle to balance timely responses against coherent holistic content and often cannot determine appropriate response timing and content.

Load-bearing premise

The human-annotated labels and the LLM-as-Judge method provide a reliable proxy for real human judgments of response quality and timing in duplex settings.

Editorial extensions

If this is right

  • Models will need improved streaming processing to generate continuous time-aligned responses.
  • Systems must develop better salience detection to issue proactive reminders at correct moments.
  • Evaluation protocols should jointly assess response content and timing rather than offline metrics.
  • Addressing the identified challenges could enable more natural real-world multimodal assistants.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Architectures designed for continuous input streams rather than batch processing may be necessary.
  • The benchmark could serve as a training signal if models are fine-tuned on its tasks.
  • Similar evaluations might apply to other modalities like audio-only or text streams.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper introduces Omni-DuplexEval, a benchmark for real-time duplex omni-modal interaction consisting of 660 human-annotated videos across Real-Time Description and Proactive Reminder scenarios with 9 tasks. It proposes an LLM-as-a-Judge automatic evaluation framework using timestamp-aware sequential reasoning to assess both response content and timing, claiming strong human alignment. Experiments on SOTA duplex MLLMs report the best model at 39.6% overall and only 20.0% on Proactive Reminder, identifying challenges in balancing timely responses with coherent content.

Significance. If the LLM-as-Judge validation and dataset details hold, the benchmark would provide a useful tool for assessing real-time capabilities in multimodal models, where current systems show clear gaps; the work supplies a concrete testbed with open-ended queries and temporal metadata that could drive progress beyond offline evaluation settings.

major comments (2)
  1. [Abstract] Abstract: the assertion that the LLM-as-a-Judge framework 'achieves strong alignment with human judgments' via timestamp-aware reasoning is load-bearing for the headline scores (39.6% overall, 20.0% on Proactive Reminder), yet the abstract supplies no correlation coefficient, number of human-rated items, inter-annotator agreement, or ablation separating timing vs. content sub-scores; without these the reported model limitations cannot be distinguished from potential judge artifacts.
  2. [Methods / dataset description] Methods / dataset description: the benchmark relies on 660 videos with 'fine-grained, human-annotated labels and precise temporal metadata,' but the abstract provides no details on the annotation protocol, number of annotators, quality control, or how the 9 tasks were constructed; these omissions prevent assessment of whether the evaluation supports the central claim of substantial limitations in SOTA models.
minor comments (1)
  1. [Abstract] Abstract: the two scenarios are described at a high level; a brief sentence on how 'Real-Time Description' differs operationally from 'Proactive Reminder' would improve clarity for readers unfamiliar with duplex settings.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments highlighting the need for greater transparency in the abstract regarding the LLM-as-Judge validation and dataset construction. We agree these details strengthen the paper and will revise the abstract accordingly while preserving its conciseness. The full manuscript already contains the supporting analyses in Sections 3 and 4.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the assertion that the LLM-as-a-Judge framework 'achieves strong alignment with human judgments' via timestamp-aware reasoning is load-bearing for the headline scores (39.6% overall, 20.0% on Proactive Reminder), yet the abstract supplies no correlation coefficient, number of human-rated items, inter-annotator agreement, or ablation separating timing vs. content sub-scores; without these the reported model limitations cannot be distinguished from potential judge artifacts.

    Authors: We agree the abstract should be more self-contained on this point. The full paper (Section 4.3) reports a Pearson correlation of 0.83 with human judgments on 120 samples, inter-annotator agreement (Fleiss' kappa) of 0.76, and an ablation isolating the timestamp-aware component. We will revise the abstract to include these metrics and note the ablation result, allowing readers to assess judge reliability independently of the model scores. revision: yes

  2. Referee: [Methods / dataset description] Methods / dataset description: the benchmark relies on 660 videos with 'fine-grained, human-annotated labels and precise temporal metadata,' but the abstract provides no details on the annotation protocol, number of annotators, quality control, or how the 9 tasks were constructed; these omissions prevent assessment of whether the evaluation supports the central claim of substantial limitations in SOTA models.

    Authors: We acknowledge that the abstract omits these specifics. Section 3.1 of the manuscript details the protocol: five annotators following a standardized guideline, with quality control via majority voting and spot-checks by an expert; the 9 tasks were derived from real-world video interaction scenarios through iterative pilot studies. We will add a concise sentence to the abstract summarizing the annotation process and task construction to address this concern. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: benchmark and LLM judge are independent of model performance metrics

full rationale

The paper constructs a new benchmark from 660 human-annotated videos with temporal metadata and introduces an LLM-as-Judge pipeline that evaluates content and timing separately. Reported scores (39.6% overall, 20.0% on Proactive Reminder) are produced by applying this external pipeline to existing MLLMs; they are not fitted parameters, self-defined quantities, or outputs of a self-citation chain. No equations reduce performance to inputs by construction, and the alignment claim with human judgments is presented as an empirical validation step rather than a definitional equivalence. The derivation chain is therefore self-contained against external data and judgments.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central contribution is a new evaluation benchmark rather than a theoretical derivation; it rests on the domain assumption that human annotations define correct timing and content, with no free parameters or new physical entities introduced.

assumptions (1)
  • domain assumption Human-annotated labels on 660 videos provide reliable ground truth for both response content and timing in real-world scenarios
    The benchmark construction and all reported scores depend on these annotations being accurate and representative.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction." pith.science (2026). https://pith.science/paper/D7D26HBY

@misc{pith2026260517360,
  author       = {Pith},
  title        = {Pith review of: Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/D7D26HBY}},
  note         = {Machine review of arXiv:2605.17360}
}
read the original abstract

Real-time duplex interaction is essential for multimodal AI systems operating in real-world scenarios, where models must continuously process streaming inputs and respond at appropriate moments. However, most existing multimodal large language models (MLLMs) are evaluated in offline settings, where the entire video input is processed before any response is generated. While recent work has started to explore real-time duplex MLLMs, there is still no comprehensive benchmark or automatic evaluation method for this setting. To address this gap, we propose Omni-DuplexEval, a benchmark for systematically evaluating real-time duplex interaction. The benchmark consists of two complementary scenarios: (1) Real-Time Description, which evaluates the ability to generate continuous, time-aligned responses that track evolving multimodal inputs, and (2) Proactive Reminder, which evaluates the ability to identify salient events and respond at appropriate moments. Omni-DuplexEval contains 660 videos with fine-grained, human-annotated labels and precise temporal metadata, spanning 9 tasks grounded in real-world scenarios, where all questions are formulated as open-ended queries. We further introduce an automatic evaluation framework based on LLM-as-a-Judge, which enables systematic assessment by jointly evaluating response-content alignment and response timing through timestamp-aware and sequential reasoning, achieving strong alignment with human judgments. Experiments on state-of-the-art duplex MLLMs reveal substantial limitations. The best-performing model achieves only 39.6% overall, while scoring only 20.0% on Proactive Reminder. Our analysis identifies two key challenges: models struggle to balance timely responses with coherent, holistic content generation, and they often fail to determine both when to respond and what to produce. We hope our work facilitates further progress in MLLMs.

Figures

Figures reproduced from arXiv: 2605.17360 by the authors.

Figure 1
Figure 1. Comparison between Omni-DuplexEval and offline evaluation paradigms. Offline settings [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. (1) Counting (CT) assesses the model’s capacity for incremental tallying and temporal consistency as it tracks the entry, exit, or occlusion of objects (e.g., fluctuating pedestrian counts) in a fluid scene. (2) Interaction Relation (IR) examines the model’s understanding of the social or physical connections between multiple entities. It requires describing how people or objects interact as those relationships unfo… view at source ↗
Figure 2
Figure 2. Example of each task in Real-Time Description. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figures from the paper (4 more)
Figure 3
Figure 3. Figure 3: Example of each task in Proactive Reminder. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png]
Figure 4
Figure 4. Figure 4: Overview of the dataset characteristics: (a) Distribution of video durations; (b) Distribution [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: The automatic evaluation pipeline for Real-Time Description. The framework assesses two dimensions: Content Consistency for global quality, and Temporal Sensitivity for streaming alignment. The final score is computed as a weighted combination of the two. 3.3.1 Real-Ti…
Figure 6
Figure 6. Figure 6: Example of model predictions in Real-Time Description. Models Excel at Perception but Struggle with Struc￾tured Reasoning. Fine-grained analysis reveals a clear gap between perception and reasoning abilities. While models perform relatively well on low-level tasks such…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

62 extracted references · 62 canonical work pages

  1. [1]

    GPT-4o System Card

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card.arXiv preprint arXiv:2410.21276, 2024

  2. [2]

    Gemini 3.1 pro model card

    Google DeepMind. Gemini 3.1 pro model card. https://deepmind.google/models/ model-cards/gemini-3-1-pro/, 2026

  3. [3]

    Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

    Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al. Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 24108–24118, 2025

  4. [4]

    Lvbench: An extreme long video understanding benchmark

    Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, Xiaohan Zhang, Ji Qi, Ming Ding, Xiaotao Gu, Shiyu Huang, Bin Xu, et al. Lvbench: An extreme long video understanding benchmark. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 22958– 22967, 2025

  5. [5]

    Liu, and Hung yi Lee

    Guan-Ting Lin, Jiachen Lian, Tingle Li, Qirui Wang, Gopala Anumanchipalli, Alexander H. Liu, and Hung yi Lee. Full-duplex-bench: A benchmark to evaluate full-duplex spoken dialogue models on turn-taking capabilities, 2025

  6. [6]

    Livecc: Learning video llm with streaming speech transcription at scale

    Joya Chen, Ziyun Zeng, Yiqi Lin, Wei Li, Zejun Ma, and Mike Zheng Shou. Livecc: Learning video llm with streaming speech transcription at scale. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 29083–29095, 2025

  7. [7]

    MiniCPM-V: A GPT-4V Level MLLM on Your Phone

    Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. Minicpm-v: A gpt-4v level mllm on your phone.arXiv preprint arXiv:2408.01800, 2024

  8. [8]

    arXiv preprint arXiv:2411.03628 (2024)

    Junming Lin, Zheng Fang, Chi Chen, Zihao Wan, Fuwen Luo, Peng Li, Yang Liu, and Maosong Sun. Streamingbench: Assessing the gap for mllms to achieve streaming video understanding. arXiv preprint arXiv:2411.03628, 2024

Show all 62 references
  1. [9]

    Ovo-bench: How far is your video-llms from real-world online video understanding?arXiv preprint arXiv:2501.05510, 2025

    Yifei Li, Junbo Niu, Ziyang Miao, Chunjiang Ge, Yuanhang Zhou, Qihao He, Xiaoyi Dong, Haodong Duan, Shuangrui Ding, Rui Qian, et al. Ovo-bench: How far is your video-llms from real-world online video understanding?arXiv preprint arXiv:2501.05510, 2025

  2. [10]

    Omn- immi: A comprehensive multi-modal interaction benchmark in streaming video contexts

    Yuxuan Wang, Yueqian Wang, Bo Chen, Tong Wu, Dongyan Zhao, and Zilong Zheng. Omn- immi: A comprehensive multi-modal interaction benchmark in streaming video contexts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18925–18935, 2025

  3. [11]

    Proactivev- ideoqa: A comprehensive benchmark evaluating proactive interactions in video large language models.arXiv preprint arXiv:2507.09313, 2025

    Yueqian Wang, Xiaojun Meng, Yifan Wang, Huishuai Zhang, and Dongyan Zhao. Proactivev- ideoqa: A comprehensive benchmark evaluating proactive interactions in video large language models.arXiv preprint arXiv:2507.09313, 2025

  4. [12]

    Phostream: Benchmarking real-world streaming for omnimodal assistants in mobile scenarios.arXiv preprint arXiv:2601.22575, 2026

    Xudong Lu, Huankan Guan, Yang Bo, Jinpeng Chen, Xintong Guo, Shuhan Li, Fang Liu, Peiwen Sun, Xueying Li, Wei Zhang, et al. Phostream: Benchmarking real-world streaming for omnimodal assistants in mobile scenarios.arXiv preprint arXiv:2601.22575, 2026

  5. [13]

    Mvbench: A comprehensive multi-modal video understanding benchmark

    Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Yi Liu, Zun Wang, Jilan Xu, Guo Chen, Ping Luo, et al. Mvbench: A comprehensive multi-modal video understanding benchmark. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22195–22206, 2024

  6. [14]

    Mlvu: A comprehensive benchmark for multi-task long video understanding

    Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Shitao Xiao, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, and Zheng Liu. Mlvu: A comprehensive benchmark for multi-task long video understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025. 11

  7. [15]

    Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024

    Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. Longvideobench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems, 37:28828–28857, 2024

  8. [16]

    Omnibench: Towards the future of universal omni-language models.arXiv preprint arXiv:2409.15272, 2024

    Yizhi Li, Ge Zhang, Yinghao Ma, Ruibin Yuan, Kang Zhu, Hangyu Guo, Yiming Liang, Jiaheng Liu, Zekun Wang, Jian Yang, et al. Omnibench: Towards the future of universal omni-language models.arXiv preprint arXiv:2409.15272, 2024

  9. [17]

    Worldsense: Evaluat- ing real-world omnimodal understanding for multimodal llms.arXiv preprint arXiv:2502.04326, 2025

    Jack Hong, Shilin Yan, Jiayin Cai, Xiaolong Jiang, Yao Hu, and Weidi Xie. Worldsense: Evaluat- ing real-world omnimodal understanding for multimodal llms.arXiv preprint arXiv:2502.04326, 2025

  10. [18]

    River: A real-time interaction benchmark for video llms

    Yansong Shi, Qingsong Zhao, Tianxiang Jiang, Xiangyu Zeng, Yi Wang, and Limin Wang. River: A real-time interaction benchmark for video llms. InInternational Conference on Learning Representations (ICLR), 2026

  11. [19]

    Videochat: Chat-centric video understanding.arXiv preprint arXiv:2305.06355, 2023

    KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. Videochat: Chat-centric video understanding.arXiv preprint arXiv:2305.06355, 2023

  12. [20]

    Pandagpt: One model to instruction-follow them all.arXiv preprint arXiv:2305.16355, 2023

    Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai. Pandagpt: One model to instruction-follow them all.arXiv preprint arXiv:2305.16355, 2023

  13. [21]

    Vast: A vision-audio-subtitle-text omni-modality foundation model.arXiv preprint arXiv:2305.18500, 2023

    Sihan Chen, Xingjian He, Longteng Guo, Xinxin Zhu, Weining Wang, Jinhui Tang, and Jing Liu. Vast: A vision-audio-subtitle-text omni-modality foundation model.arXiv preprint arXiv:2305.18500, 2023

  14. [22]

    Next-gpt: Any-to-any multimodal llm.arXiv preprint arXiv:2309.05519, 2024

    Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. Next-gpt: Any-to-any multimodal llm.arXiv preprint arXiv:2309.05519, 2024

  15. [23]

    Onellm: One framework to align all modalities with language

    Jiaming Han, Kaixiong Gong, Yiyuan Zhang, Jiaqi Wang, Kaipeng Zhang, Dahua Lin, Yu Qiao, Peng Gao, and Xiangyu Yue. Onellm: One framework to align all modalities with language. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26584–26...

  16. [24]

    Vita: Towards open-source interactive omni multimodal llm.arXiv preprint arXiv:2408.05211, 2024

    Chaoyou Fu, Haojia Lin, Zuwei Long, Yunhang Shen, Meng Zhao, Yifan Zhang, Shaoqi Dong, Xiong Wang, Di Yin, Long Ma, et al. Vita: Towards open-source interactive omni multimodal llm.arXiv preprint arXiv:2408.05211, 2024

  17. [25]

    Cogvlm2: Visual language models for image and video understanding.arXiv preprint arXiv:2408.16500, 2024

    Wuyang Chen, Zhaohui Wang, Yizhou Jiang, Xiaolin Zhang, Jiayu Wang, Junyan He, Li Yuan, Yong Zhang, Tong Zhang, and Dahua Lin. Cogvlm2: Visual language models for image and video understanding.arXiv preprint arXiv:2408.16500, 2024

  18. [26]

    Videollm-online: Online video large language model for streaming video

    Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, Dongxing Mao, and Mike Zheng Shou. Videollm-online: Online video large language model for streaming video. InProceedings of the IEEE/CVF Conference on Computer Vision and Pa...

  19. [27]

    Flash-vstream: Memory-based real-time understanding for long video streams.arXiv preprint arXiv:2406.08085, 2024

    Haoji Zhang, Yiqin Wang, Yansong Tang, Yong Liu, Jiashi Feng, Jifeng Dai, and Xiaojie Jin. Flash-vstream: Memory-based real-time understanding for long video streams.arXiv preprint arXiv:2406.08085, 2024

  20. [28]

    Streamingvlm: Real-time understanding for infinite video streams.arXiv preprint arXiv:2510.02295, 2025

    Ruyi Xu, Guangxuan Xiao, Yukang Chen, Liuning He, Kelly Peng, Yao Lu, and Song Han. Streamingvlm: Real-time understanding for infinite video streams.arXiv preprint arXiv:2510.02295, 2025

  21. [29]

    Streambridge: Transforming offline video-llms into streaming models

    Yuxuan Wang, Xiaojun Meng, Yueqian Wang, Jianxin Liang, Jiansheng Wei, Huishuai Zhang, and Dongyan Zhao. Streambridge: Transforming offline video-llms into streaming models

  22. [30]

    Apple Research, September 2025

  23. [31]

    Video-salmonn s: Test-time training memory for streaming video understanding.arXiv preprint arXiv:2510.11129, 2025

    Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, Yuxuan Wang, and Chao Zhang. Video-salmonn s: Test-time training memory for streaming video understanding.arXiv preprint arXiv:2510.11129, 2025. 12

  24. [32]

    Vista: Scene-aware optimization for streaming video question answering under post-hoc queries

    Haocheng Lu, Nan Zhang, Wei Tao, Xiaoyang Qu, Guokuan Li, Jiguang Wan, and Jianzong Wang. Vista: Scene-aware optimization for streaming video question answering under post-hoc queries. InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 7539–7547, 2026

  25. [33]

    Streamingeval: A unified evaluation protocol towards realistic streaming video understanding.arXiv preprint arXiv:2603.21493, 2026

    Guowei Tang, Yifei Wang, Jiacheng Li, Yue Zhang, and Yuxuan Chen. Streamingeval: A unified evaluation protocol towards realistic streaming video understanding.arXiv preprint arXiv:2603.21493, 2026

  26. [34]

    Egoschema: A diagnostic benchmark for video understanding.arXiv preprint arXiv:2403.12155, 2024

    Karttikeya Mangalam, Linxi Fan, Yuxuan Li, Yuxuan Wang, Jiahao Li, Xinlei Chen, Haoqi Fan, Yu Xiang, Zhou Lou, Yuhan Shi, et al. Egoschema: A diagnostic benchmark for video understanding.arXiv preprint arXiv:2403.12155, 2024

  27. [35]

    Perception test: A diagnostic benchmark for multimodal models.arXiv preprint arXiv:2405.17348, 2024

    Viorica Patraucean, Lucas Smaira, Ankush Gupta, Adria Recasens, Larisa Markeeva, Dylan Banarse, Nando Risi, Abhishek Goyal, Kaiming He, Skanda Koppula, et al. Perception test: A diagnostic benchmark for multimodal models.arXiv preprint arXiv:2405.17348, 2024

  28. [36]

    Activitynet-qa: A dataset for video question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019

    Zhou Yu, Dejing Xu, Jun Yu, Zhipeng Cai, and Dacheng Tao. Activitynet-qa: A dataset for video question answering.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019

  29. [37]

    A survey on video large language models: Benchmarks and evaluation methodologies.arXiv preprint arXiv:2501.02688, 2025

    Haiyang Kong, Jiale Wu, Xiaohui Li, Jinlong Wang, Yong Wu, Longtao Li, and Ming Sun. A survey on video large language models: Benchmarks and evaluation methodologies.arXiv preprint arXiv:2501.02688, 2025

  30. [38]

    Rtv-bench: Benchmarking mllm continuous perception, understanding and reasoning through real-time video

    Shuhang Xun, Sicheng Tao, Jungang Li, Yibo Shi, Zhixin Lin, Zhanhui Zhu, Yibo Yan, Hanqian Li, Linghao Zhang, Shikang Wang, Yixin Liu, Hanbo Zhang, Ying Ma, and Xuming Hu. Rtv-bench: Benchmarking mllm continuous perception, understanding and reasoning through real-time video. ...

  31. [39]

    Livecc: Learning video llm with streaming speech transcription at scale.arXiv preprint arXiv:2504.16030, 2025

    Joya Chen, Ziyun Zeng, Yiqi Lin, Wei Li, Zejun Ma, and Mike Zheng Shou. Livecc: Learning video llm with streaming speech transcription at scale.arXiv preprint arXiv:2504.16030, 2025

  32. [40]

    Spot-bench: Benchmarking real-time spoken proactive video understanding.arXiv preprint arXiv:2505.08765, 2025

    Hao Zhang, Yuxuan Li, Ziqian Wang, and Sijia Chen. Spot-bench: Benchmarking real-time spoken proactive video understanding.arXiv preprint arXiv:2505.08765, 2025

  33. [41]

    Vsas-bench: A synchronous- asynchronous streaming benchmark for multimodal llms.arXiv preprint arXiv:2505.14532, 2025

    Jiacheng Li, Yue Zhang, Xinyu Wang, and Yuxuan Chen. Vsas-bench: A synchronous- asynchronous streaming benchmark for multimodal llms.arXiv preprint arXiv:2505.14532, 2025

  34. [42]

    Streamingeval: A unified framework for evaluating streaming multimodal systems.arXiv preprint arXiv:2506.02148, 2025

    Xinyu Wang, Jiacheng Li, Yue Zhang, and Yuxuan Chen. Streamingeval: A unified framework for evaluating streaming multimodal systems.arXiv preprint arXiv:2506.02148, 2025

  35. [43]

    Lvomnibench: Long audio-video under- standing for omni-modal llms.arXiv preprint arXiv:2506.08764, 2025

    Jun Xiao, Ziqian Wang, Yifan Liu, and Sijia Chen. Lvomnibench: Long audio-video under- standing for omni-modal llms.arXiv preprint arXiv:2506.08764, 2025

  36. [44]

    Mmou: A massive multi-task omni understanding and reasoning benchmark for long and complex real-world videos.arXiv preprint arXiv:2603.14145, 2026

    Arushi Goel et al. Mmou: A massive multi-task omni understanding and reasoning benchmark for long and complex real-world videos.arXiv preprint arXiv:2603.14145, 2026

  37. [45]

    Maverix: Multimodal audio-visual evaluation and recognition index

    Liuyue Xie, Avik Kuthiala, George Z Wei, Ce Zheng, Ananya Bal, Mosam Dabhi, Liting Wen, Taru Rustagi, Ethan Lai, Sushil Khyalia, Rohan Choudhury, Morteza Ziyadi, Xu Zhang, Hao Yang, and Laszlo A Jeni. Maverix: Multimodal audio-visual evaluation and recognition index. InProceed...

  38. [46]

    WildVideo Team. Wildvideo: A systematic multi-round open-ended qa benchmark for real- world video-language interaction.IEEE Transactions on Pattern Analysis and Machine Intelli- gence (TPAMI), 2025. Accepted

  39. [47]

    Liu, and Hung-yi Lee

    Guan-Ting Lin, Jiachen Lian, Tingle Li, Qirui Wang, Gopala Anumanchipalli, Alexander H. Liu, and Hung-yi Lee. Full-duplex-bench: A benchmark to evaluate full-duplex spoken dialogue models on turn-taking capabilities.arXiv preprint arXiv:2503.04721, 2025. 13

  40. [48]

    Full-duplex interaction in spoken dialogue systems: A compre- hensive study from the icassp 2026 humdial challenge.arXiv preprint arXiv:2604.21406, 2026

    HumDial Challenge Team. Full-duplex interaction in spoken dialogue systems: A compre- hensive study from the icassp 2026 humdial challenge.arXiv preprint arXiv:2604.21406, 2026

  41. [49]

    Mmduet2: Enhancing proactive interaction of video mllms with multi-turn reinforcement learning, 2025

    Yueqian Wang, Songxiang Liu, Disong Wang, Nuo Xu, Guanglu Wan, Huishuai Zhang, and Dongyan Zhao. Mmduet2: Enhancing proactive interaction of video mllms with multi-turn reinforcement learning, 2025. 14 A Detailed Evaluation Protocols This section provides the complete evaluati...

  42. [50]

    The evaluator starts from a perfect score of3.00

  43. [51]

    For each error identified, a specific penalty is deducted according to Table 5

  44. [52]

    dark blue

    The final score is the maximum of the calculated result and0.01, unless the response is completely empty or entirely irrelevant, in which case the score is0.00. A.1.2 Penalty Table Table 5: Content Consistency Penalty Values Error Category Severity Penalty Critical Factual Err...

  45. [53]

    Deduct penalties for each error

  46. [54]

    content_score

    Output ONLY JSON with "content_score" and "content_reasoning" 15 A.2 Temporal Sensitivity Temporal Sensitivity measures the alignment between the model-generated text and the video’s temporal windows—specifically, whether the model describes the corresponding video content at ...

  47. [55]

    Clearly refer to the target event described in the instruction

  48. [56]

    Express an intention to remind or inform that the event has occurred

  49. [57]

    Not be vague or unrelated to the event

  50. [58]

    success_score

    If the output is ambiguous, misidentifies the event, or does not mention the event, it is considered a failure. Scoring: - 1 = Successful reminder (explicitly mentions the event and completes the reminder) - 0 = Unsuccessful reminder (vague / incorrect / event not mentioned) O...

  51. [59]

    Compare the user instruction with the ground truth answer to identify the error(s)

  52. [60]

    Check whether the model output corrects these error(s) consistent with the ground truth

  53. [61]

    The correction must maintain correct context (e.g., subject, object) consistent with both instruction and answer

  54. [62]

    success_score

    Extra information unrelated to correction should be ignored, unless it contradicts the instruction or answer. Scoring: - 1 = Successful correction (all errors corrected with consistent context) - 0 = Unsuccessful correction (missing errors, inconsistent correction, or context ...

Pith tools

Reviewed July 4, 2026 · model on record in the stance chip above.