Pith. sign in

REVIEW 5 major objections 5 minor 78 references

Current multimodal models top out at 0.28 on a new three-level context-learning benchmark.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 02:49 UTC pith:UFLIF74P

load-bearing objection A useful three-level diagnostic framework for multimodal context learning, with real low scores, but the headline rankings are uncalibrated because the judge is also a contestant and the level assignments are arbitrary. the 5 major comments →

arxiv 2607.25294 v1 pith:UFLIF74P submitted 2026-07-28 cs.CV cs.AIcs.CLcs.LG

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

classification cs.CV cs.AIcs.CLcs.LG
keywords multimodal context learningvision-language benchmarkcontext groundingnew information applicationnew knowledge learningmultimodal large language modelsevaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to measure whether multimodal models actually learn from task-specific context at inference time—text, figures, tables, maps, and images—rather than just retrieving pre-trained knowledge. To localize where context use breaks down, it organizes tasks into three levels: grounding evidence, applying newly supplied facts, and learning new rules or conclusions defined only by the context. Across 3,443 instances and six recent models, the best overall score is 0.2847, which the paper reads as evidence that multimodal context learning is far from solved. It also finds that different models lead different levels, showing that the capability is not monolithic. The benchmark is built so failures can be attributed to one of three stages rather than a single generic 'context problem.'

Core claim

On its own terms, the paper claims that today's multimodal systems cannot reliably learn from context that mixes text and images: across 3,443 instances and six recent models, the best overall score is only 0.2847. The organizing claim is that failures are not a single 'context problem' but three distinct capability gaps—grounding evidence in the context, applying newly supplied information, and learning context-defined knowledge—and that models differ in where they fail. The paper finds that one model leads on grounding and knowledge learning while another leads on information application, and that the hierarchy is diagnostic enough to attribute errors to specific stages of context use.

What carries the argument

The three-level capability hierarchy (L0 context grounding, L1 new information application, L2 new knowledge learning) is the central mechanism. It is instantiated as a benchmark of 3,443 instances that joins converted public tasks with two newly constructed tasks (financial-report ROE analysis and medical-paper conclusion inference), all scored under a unified inference and evaluation protocol. The hierarchy's work is to separate failures of visual access from failures of contextual reasoning, so that a low aggregate score becomes a diagnosis of where context use breaks down rather than an undifferentiated average.

Load-bearing premise

The level-wise rankings rest on an arbitrary assignment of two boundary datasets that mix grounding with higher-level use; if those datasets were moved to the other level, the best model per level could change.

What would settle it

Re-run the evaluation with the two boundary datasets swapped between levels; if the level-wise leader changes, the hierarchy's headline result is an artifact of assignment. Alternatively, run text-only versions of grounding tasks and ask whether any model still solves them—non-trivial scores would mean visual dependency is not actually enforced, undermining the claim that the benchmark measures multimodal grounding.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the scores hold, multimodal context learning is nowhere near saturated, so current model families cannot be relied on for tasks where the answer depends on novel multimodal context.
  • Model rankings differ by level, meaning a single leaderboard score hides real differences: a model can be strong at grounding while weak at applying new values, or vice versa.
  • Input length alone does not explain failures—after excluding inputs that exceed a model's context limit, longer contexts do not steadily decrease accuracy—so pushing context windows further is not by itself the fix.
  • Judge choice can shift reported scores on open-ended multimodal tasks, so benchmarks with such tasks should report the judge model and ideally calibrate judges.
  • The three-level hierarchy provides a template for diagnosing failures in deployed systems: attribute an error to grounding, application, or knowledge acquisition rather than calling it a generic context problem.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The hierarchy could transfer directly to evaluating retrieval-augmented generation and agentic systems, where the same three failure modes—evidence not grounded, facts not applied, context-only rules ignored—are likely to appear; inserting a retrieval step before the context would test that.
  • Because the lowest scores concentrate in new-knowledge learning, training objectives that explicitly reward inducing and applying context-defined rules, rather than just retrieving or copying, may be the highest-leverage next step—an inference, not a claim in the paper.
  • The boundary assignment of two datasets that mix grounding with higher-level use drives part of the level-wise ranking, so future versions could either split those datasets per instance or report sensitivity to alternative assignments.
  • The automated construction-and-filtering pipeline, which uses a strong model to reject shortcut-solvable instances, is itself reusable; a natural test is whether the same pipeline can scale the benchmark to new domains without manual annotation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper introduces CLBench-V, a benchmark for multimodal context learning organized into three levels: context grounding (L0), new information application (L1), and new knowledge learning (L2). It integrates public benchmarks (ReasonMap, Insight-O3, ZeroBench, CourtSI, MIRBench, Pix2Fact, MMLongBench-Pic, BrowseComp-V3, PRISMM-Bench, CL-Bench) and adds two constructed tasks (Financial Report ROE and Paper Conclusion), totaling 3,443 instances. Six recent multimodal models are evaluated using unified prompts and task-specific evaluators, with LLM judges used for open-ended outputs. The paper reports a best overall score of 0.2847 and claims that multimodal context learning remains far from saturated, with InternVL3.5-30B-A3B strongest on L0/L2 and Qwen3.5-Plus strongest on L1. It also analyzes judge reliability, input complexity, and representative failure modes.

Significance. If validated, CLBench-V would provide a useful diagnostic tool: the three-level hierarchy is intuitive and addresses a real gap in multimodal evaluation, the integration of heterogeneous sources into a unified framework is valuable, and the public code/data release supports reproducibility. The paper also makes a constructive step by analyzing judge reliability (Table 4) and explicitly listing limitations. However, the benchmark's validity is not yet fully established. The main judge is itself an evaluated model, the Paper Conclusion task lacks human-verified references, and the boundary level assignments are arbitrary. These issues directly affect the central empirical claim and the model-level rankings, so the contribution is conditional on additional validation.

major comments (5)
  1. [§3.3.2, Eq. (2), Table 1] The Paper Conclusion task (581 instances, the largest L2 component) constructs reference findings from the Results section without reported human verification, and scores predictions with an LLM list-count judge. Because Eq. (2) uses N_ref and judge-estimated N_hit/N_extra, incomplete references or a miscalibrated judge can directly deflate or inflate model scores. Without a human-verified gold set for a sample, the large L2 gap (e.g., InternVL 0.3536 vs. GPT-5.4 0.0694 in Table 2) may reflect annotation quality rather than context-learning ability. Please report human verification of the reference findings and inter-annotator agreement on judged matches.
  2. [§4.1, §5, Table 4] Qwen3.6-27B is used as the main judge for judge-based tasks while also being evaluated in Table 2. Table 4 shows that macro scores for the same prediction sets range from 0.1564 (Qwen3-VL-32B) to 0.2122 (Qwen3-VL-4B), a relative difference of about 36%, with Qwen3.6-27B giving 0.1777. This variation makes the headline ranking potentially sensitive to judge choice and creates a circularity risk for Qwen3.6-27B's own score. Please either use a judge outside the evaluated model set or demonstrate that the main rankings are robust across judges.
  3. [Table 2 footnote, §3.2, Table 1] The level-wise aggregates are computed with MIRBench assigned to L0 and PRISMM-Bench to L1, although Table 1 marks them as boundary datasets (L0–L1 and L1–L2). The paper offers no sensitivity analysis for these choices. Because the headline conclusion is that InternVL3.5-30B-A3B is best at L0/L2 while Qwen3.5-Plus is best at L1, a different reasonable assignment could change the best model per level. Please justify the assignments with the cumulative-level criterion or report aggregate results under alternative assignments.
  4. [§3.3.1 data filter] The automated rejection filter uses Qwen3.5-Plus as the primary inspector, and Qwen3.5-Plus is also evaluated in the main results. This creates a selection effect: the benchmark's composition may encode the filtering model's strengths and weaknesses. The subsequent manual review is described but not quantified (e.g., number of instances rejected/retained, inter-annotator agreement). Please report filtering statistics and, ideally, verify the filter with an independent judge or a human sample.
  5. [Abstract, §4.2] The paper interprets the best overall score of 0.2847 as evidence that multimodal context learning is 'far from saturated.' Without a human baseline or a no-context control on the same instances, low scores cannot be distinguished from benchmark artifact (e.g., ambiguous or incomplete references, over-strict judges). Please add a small human study or a strong oracle/no-context baseline to calibrate the benchmark's difficulty and substantiate the central claim.
minor comments (5)
  1. [Table 3, Financial Report ROE] Clarify whether the 0.0000 for InternVL3.5-30B-A3B (context-limit failure) is included as zero in the L1 and overall aggregates, since it materially lowers that model's score; consider reporting it as 'not evaluated' in the aggregate or as a separate row.
  2. [§5, Table 4] The judge-comparison table reports two 'prediction sets' but does not define them or provide variance/error bars. Add a sentence describing the prediction sets and report scores per dataset or across repeated runs.
  3. [§6.1] The token-length and image-count correlations are computed on a heterogeneous mix; the post-hoc exclusion of financial reports is informative but should be stated as exploratory with the pre-defined criterion. Also report the number of instances used in each correlation.
  4. [§4 and §8] The conclusion refers to 'planned and preliminary experiments,' while Section 4 presents completed main results; align the language to clarify the status (preliminary vs. final) of the reported scores.
  5. [Appendix C.1] The phrase 'drawn from the case-study set in write/case_study' appears to reference an internal directory; replace with a stable artifact reference.

Circularity Check

0 steps flagged

No significant circularity; the benchmark's headline score is an empirical measurement, not a derivation that reduces to its own inputs.

full rationale

CLBench-V makes an empirical benchmark claim: across 3,443 instances and six models the best overall score is 0.2847. I checked the derivation chain claimed by the paper and found no step where a predicted quantity is identical by construction to a fitted input. The three-level hierarchy (L0 grounding, L1 information application, L2 knowledge learning) is defined independently of the scores, before evaluation, and the aggregate score is a sample-weighted mean over a fixed instance set. Equation 1 is the standard DuPont identity used to define the ROE task rather than a fitted model, and Equation 2 is a recall-minus-precision scoring rule over reference findings; neither contains a parameter fitted to the evaluated models' outputs. The two self-referential design points are the main judge Qwen3.6-27B also being one of the evaluated models (Sec. 4.1) and Qwen3.5-Plus serving as the data filter while also being evaluated (Sec. 3.3.1). These are validity and calibration concerns, not circularity in the strict sense: the judge's outputs are not used to define the judge's own scores by construction, the best overall model (InternVL3.5-30B-A3B) is not the judge, and the paper's own judge-comparison analysis (Table 4) provides a partial external check on judge bias. The acknowledged boundary assignments of MIRBench to L0 and PRISMM-Bench to L1 are ranking-sensitivity issues, not circular steps. The absence of a human baseline or ground-truth validation is a legitimate correctness risk, but the circularity rubric explicitly excludes such non-consensus or artifact concerns. No load-bearing step reduces to the paper's own inputs.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The benchmark introduces no fitted numerical parameters; instead, it relies on several unvalidated design assumptions: the validity of the capability hierarchy, the reliability of model-based filtering and judging, and leakage-free paper selection. These assumptions are load-bearing for interpreting the reported level-wise scores.

axioms (5)
  • domain assumption The three-level capability hierarchy (L0/L1/L2) is a valid decomposition of multimodal context learning.
    The entire benchmark is organized around this assumed hierarchy; no independent validation is provided that levels are clean or orthogonal.
  • ad hoc to paper The automated filtering using Qwen3.5-Plus as primary inspector and subsequent manual review yields unbiased, high-quality benchmark instances.
    The filtering implementation is not fully specified and relies on a model that is also later evaluated, introducing potential selection bias.
  • domain assumption LLM judges, particularly Qwen3.6-27B, can reliably score open-ended multimodal answers.
    Open-ended tasks (Pix2Fact, Paper Conclusion, etc.) use LLM judges without human-annotated ground truth for judge calibration.
  • domain assumption Recent papers used in Paper Conclusion inference have not been memorized by evaluated models.
    Models may have seen the papers in pretraining; the paper mitigates by using recent papers but does not verify leakage.
  • ad hoc to paper Converting public benchmarks into the unified schema preserves their intended task properties.
    Conversion-time rejection sampling and reformatting could alter what a task measures, and the paper does not validate equivalence with original benchmarks.

pith-pipeline@v1.3.0-alltime-deepseek · 15397 in / 9009 out tokens · 90839 ms · 2026-08-01T02:49:49.995882+00:00 · methodology

0 comments
read the original abstract

Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes, or web pages. We introduce CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning. CLBench-V combines converted public benchmarks with newly constructed datasets spanning domains such as science, finance, long-document understanding, spatial reasoning, and web-based visual question answering. To reduce the cost of constructing domain-specific context-learning tasks, we further use automated construction and filtering procedures for our newly built datasets. Across 3,443 instances and six recent multimodal models, the best overall score is only 0.2847, indicating that multimodal context learning remains far from saturated. Moreover, InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus performs best on new information application. We further analyze judge reliability, context length, image count, and representative failure cases. Code is available at https://github.com/IamLihua/CLBench-V.

Figures

Figures reproduced from arXiv: 2607.25294 by Chengqi Li, Jiapeng Li, Lai Wei, Ruina Hu, Weiran Huang, Yue Wang.

Figure 1
Figure 1. Figure 1: Overview of CLBench-V. The benchmark covers diverse multimodal contexts and tasks, including [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Capability profiles of the evaluated models. Cell color encodes the score; solid outlines mark the best [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Appendix case-study panels aligned with the L0/L1/L2 taxonomy. Left (L0): visual grounding and [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

78 extracted references · 24 linked inside Pith

  1. [2]

    Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2024. https://arxiv.org/abs/2308.14508 LongBench : A bilingual, multitask benchmark for long context understanding . arXiv preprint arXiv:2308.14508

  2. [3]

    Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2025. https://arxiv.org/abs/2412.15204 LongBench v2 : Towards deeper understanding and reasoning on realistic long-context multitasks . arXiv preprint arXiv:2412.15204

  3. [4]

    Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.300 FinQA : A dataset of numerical reasoning over financial data . In Proceedings of the 2021 Conference on Empirical Methods in Natural Langua...

  4. [5]

    Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.421 ConvFinQA : Exploring the chain of numerical reasoning in conversational finance question answering . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing

  5. [6]

    Yew Ken Chia, Liying Cheng, Hou Pong Chan, Chaoqun Liu, Maojia Song, Sharifah Mahani Aljunied, Soujanya Poria, and Lidong Bing. 2024. https://arxiv.org/abs/2411.06176 M-LongDoc : A benchmark for multimodal super-long document understanding and a retrieval-aware tuning framework . arXiv preprint arXiv:2411.06176

  6. [7]

    Smith, and Matt Gardner

    Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. 2021. https://arxiv.org/abs/2105.03011 A dataset of information-seeking questions and answers anchored in research papers . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics

  7. [8]

    Kuicai Dong, Yujing Chang, Xin Deik Goh, Dexun Li, Ruiming Tang, and Yong Liu. 2025. https://arxiv.org/abs/2501.08828 MMDocIR : Benchmarking multimodal retrieval for long documents . arXiv preprint arXiv:2501.08828

  8. [9]

    Shihan Dou, Ming Zhang, Zhangyue Yin, Chenhao Huang, Yujiong Shen, Junzhe Wang, Jiayi Chen, Yuchen Ni, Junjie Ye, Cheng Zhang, Huaibing Xie, Jianglu Hu, Shaolei Wang, Weichao Wang, Yanling Xiao, Yiting Liu, Zenan Xu, Zhen Guo, Pluto Zhou, and 8 others. 2026. https://arxiv.org/abs/2602.03587 CL-bench : A benchmark for context learning . arXiv preprint arXi...

  9. [10]

    Hang Du, Jiayang Zhang, Guoshun Nan, Wendi Deng, Zhenyan Chen, Chenyang Zhang, Wang Xiao, Shan Huang, Yuqi Pan, Tao Qi, and Sicong Leng. 2025. https://arxiv.org/abs/2509.17040 From easy to hard: The MIR benchmark for progressive interleaved multi-image reasoning . arXiv preprint arXiv:2509.17040

  10. [11]

    Wu, Bilel Omrani, Gautier Viaud, Celine Hudelot, and Pierre Colombo

    Manuel Faysse, Hugues Sibille, Tony F. Wu, Bilel Omrani, Gautier Viaud, Celine Hudelot, and Pierre Colombo. 2024. https://arxiv.org/abs/2407.01449 ColPali : Efficient document retrieval with vision language models . arXiv preprint arXiv:2407.01449

  11. [12]

    Sicheng Feng, Song Wang, Shuyi Ouyang, Lingdong Kong, Zikai Song, Jianke Zhu, Huan Wang, and Xinchao Wang. 2025. https://arxiv.org/abs/2505.18675 ReasonMap : Towards fine-grained visual reasoning from transit maps . arXiv preprint arXiv:2505.18675

  12. [13]

    Jiaheng Guo, Ruidong Wang, Chunlei Yao, Chenyang Zhang, Yaqi Song, Zichao Wang, Kun He, Shiguang Wu, Ziliang Lin, Qipeng Wei, Xuan Liao, Xihui Liu, Li Liu, Ying-Ying Lin, Wenbing Liu, Shuo Zhou, Qiang Wu, Ping Zhou, Bingxuan Zhao, and 5 others. 2026. https://arxiv.org/abs/2605.10187 MemEye : Benchmarking multimodal large language models with long-context ...

  13. [14]

    Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. 2024. https://arxiv.org/abs/2404.06654 RULER : What's the real context size of your long-context language models? arXiv preprint arXiv:2404.06654

  14. [15]

    Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, and Wenhu Chen. 2024. https://arxiv.org/abs/2405.01483 MANTIS : Interleaved multi-image instruction tuning . arXiv preprint arXiv:2405.01483

  15. [16]

    Yifan Jiang, Cong Zhang, Bofei Zhang, Qiaofeng Zheng, Yifan Yang, Bingzhang Wang, and Yew-Soon Ong. 2026. https://arxiv.org/abs/2602.00593 Pix2Fact : When vision is not enough -- benchmarking fine-grained VQA with web verification on high-resolution real-world scenes . arXiv preprint arXiv:2602.00593

  16. [17]

    Kaican Li, Lewei Yao, Jiannan Wu, Tiezheng Yu, Jierun Chen, Haoli Bai, Lu Hou, Lanqing Hong, Wei Zhang, and Nevin L. Zhang. 2025 a . https://arxiv.org/abs/2512.18745 InSight-o3 : Empowering multimodal foundation models with generalized visual search . arXiv preprint arXiv:2512.18745

  17. [18]

    Zhuowan Li, Dong Huang, Wenhao Wang, Shenglong Yang, Zhe Liu, and 1 others. 2025 b . https://aclanthology.org/2025.acl-long.1351/ GIRAFFE : Design choices for extending the context length of visual language models . In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics

  18. [19]

    Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

    Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024 a . https://doi.org/10.1162/tacl_a_00638 Lost in the middle: How language models use long contexts . Transactions of the Association for Computational Linguistics, 12:157--173

  19. [20]

    Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. 2024 b . https://arxiv.org/abs/2307.06281 MMBench : Is your multi-modal model an all-around player? arXiv preprint arXiv:2307.06281

  20. [21]

    Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, and 1 others. 2024. https://arxiv.org/abs/2407.01523 MMLongBench-Doc : Benchmarking long-context document understanding with visualizations . arXiv preprint arXiv:2407.01523

  21. [22]

    Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. https://arxiv.org/abs/2203.10244 ChartQA : A benchmark for question answering about charts with visual and logical reasoning . In Findings of the Association for Computational Linguistics: ACL 2022

  22. [23]

    Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021. https://arxiv.org/abs/2007.00398 DocVQA : A dataset for VQA on document images . In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

  23. [24]

    Fanqing Meng, Jin Wang, Chuanhao Li, Quanfeng Lu, Hao Tian, Jiaqi Liao, Xizhou Zhu, Jifeng Dai, and 1 others. 2024. https://arxiv.org/abs/2408.02718 MMIU : Multimodal multi-image understanding for evaluating large vision-language models . arXiv preprint arXiv:2408.02718

  24. [25]

    Yilong Ren, Jingjing An, Luming Zhang, Mingjie Zheng, Wei Zhang, Songlin Yan, Hongyu Zhang, Wenqi Zhang, Kai Liu, Di Wu, Hao Chen, Wang Yan, and Yuanchun Liu. 2026. https://arxiv.org/abs/2605.14906 MemLens : Evaluating memory in multimodal large language models . arXiv preprint arXiv:2605.14906

  25. [26]

    Jonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma, Akash Gupta, Samuel Roberts, Ioana Croitoru, Simion-Vlad Bogolin, Jialu Tang, Florian Langer, Vyas Raina, Vatsal Raina, Hanyi Xiong, Vishaal Udandarao, Jingyi Lu, Shiyang Chen, Sam Purkis, Tianshuo Yan, Wenye Lin, Gyungin Shin, and 15 others. 2025. https://arxiv.org/abs/2502.09696 ZeroBench : An imposs...

  26. [27]

    Jehanzeb Mirza, Sivan Doveh, James Glass, Rogerio Feris, and Wei Lin

    Lukas Selch, Yufang Hou, M. Jehanzeb Mirza, Sivan Doveh, James Glass, Rogerio Feris, and Wei Lin. 2025. https://arxiv.org/abs/2510.16505 PRISMM-Bench : A benchmark of peer-review grounded multimodal inconsistencies . arXiv preprint arXiv:2510.16505

  27. [28]

    Yu Shi, Chaoyue Tang, Bokai Xu, Junbo Cui, Junhao Ran, Yukun Yan, Zhenghao Liu, Shuo Wang, Xu Han, Zhiyuan Liu, and Maosong Sun. 2024. https://arxiv.org/abs/2410.10594 VisRAG : Vision-based retrieval-augmented generation on multi-modality documents . arXiv preprint arXiv:2410.10594

  28. [29]

    Richard Yu, Xiang Wan, and Benyou Wang

    Dingjie Song, Shunian Chen, Guiming Hardy Chen, F. Richard Yu, Xiang Wan, and Benyou Wang. 2024. https://arxiv.org/abs/2404.18532 MileBench : Benchmarking MLLMs in long context . arXiv preprint arXiv:2404.18532

  29. [30]

    Huang, Zekun Li, Qin Liu, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu, Wenxuan Zhou, Kai Zhang, and 1 others

    Fei Wang, Xingyu Fu, James Y. Huang, Zekun Li, Qin Liu, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu, Wenxuan Zhou, Kai Zhang, and 1 others. 2024 a . https://arxiv.org/abs/2406.09411 MuirBench : A comprehensive benchmark for robust multi-image understanding . arXiv preprint arXiv:2406.09411

  30. [33]

    Zhaowei Wang, Wenhao Yu, Xiyu Ren, Jipeng Zhang, Yu Zhao, Rohit Saxena, Liang Cheng, Ginny Wong, Simon See, Pasquale Minervini, Yangqiu Song, and Mark Steedman. 2025. https://arxiv.org/abs/2505.10610 MMLongBench : Benchmarking long-context vision-language models effectively and thoroughly . arXiv preprint arXiv:2505.10610

  31. [34]

    Yuchen Yang, Yuqing Shao, Duxiu Huang, Linfeng Dong, Yifei Liu, Suixin Tang, Xiang Zhou, Yuanyuan Gao, Wei Wang, Yue Zhou, Xue Yang, Yanfeng Wang, Xiao Sun, and Zhihang Zhong. 2026. https://arxiv.org/abs/2603.09896 Stepping VLMs onto the court: Benchmarking spatial intelligence in sports . arXiv preprint arXiv:2603.09896

  32. [35]

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, and 1 others. 2024. https://arxiv.org/abs/2311.16502 MMMU : A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI . In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  33. [36]

    Huanyao Zhang, Jiepeng Zhou, Bo Li, Bowen Zhou, Yanzhe Shan, Haishan Lu, Zhiyong Cao, Jiaoyang Chen, Yuqian Han, Zinan Sheng, Zhengwei Tao, Hao Liang, Jialong Wu, Yang Shi, Yuanpeng He, Jiaye Lin, Qintong Zhang, Guochen Yan, Runhao Zhao, and 6 others. 2026. https://arxiv.org/abs/2602.12876 BrowseComp - V^3 : A visual, vertical, and verifiable benchmark fo...

  34. [37]

    Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. https://arxiv.org/abs/2105.07624 TAT-QA : A question answering benchmark on a hybrid of tabular and textual content in finance . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics

  35. [38]

    Aho and Jeffrey D

    Alfred V. Aho and Jeffrey D. Ullman , title =. 1972

  36. [39]

    Publications Manual , year = "1983", publisher =

  37. [40]

    Chandra and Dexter C

    Ashok K. Chandra and Dexter C. Kozen and Larry J. Stockmeyer , year = "1981", title =. doi:10.1145/322234.322243

  38. [41]

    Scalable training of

    Andrew, Galen and Gao, Jianfeng , booktitle=. Scalable training of

  39. [42]

    Dan Gusfield , title =. 1997

  40. [43]

    Tetreault , title =

    Mohammad Sadegh Rasooli and Joel R. Tetreault , title =. Computing Research Repository , volume =. 2015 , url =

  41. [44]

    A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =

    Ando, Rie Kubota and Zhang, Tong , Issn =. A Framework for Learning Predictive Structures from Multiple Tasks and Unlabeled Data , Volume =. Journal of Machine Learning Research , Month = dec, Numpages =

  42. [45]

    2026 , url =

    Jiang, Yifan and Zhang, Cong and Zhang, Bofei and Zheng, Qiaofeng and Yang, Yifan and Wang, Bingzhang and Ong, Yew-Soon , journal =. 2026 , url =

  43. [46]

    2026 , url =

    Dou, Shihan and Zhang, Ming and Yin, Zhangyue and Huang, Chenhao and Shen, Yujiong and Wang, Junzhe and Chen, Jiayi and Ni, Yuchen and Ye, Junjie and Zhang, Cheng and Xie, Huaibing and Hu, Jianglu and Wang, Shaolei and Wang, Weichao and Xiao, Yanling and Liu, Yiting and Xu, Zenan and Guo, Zhen and Zhou, Pluto and Gui, Tao and Wu, Zuxuan and Qiu, Xipeng an...

  44. [47]

    2024 , url =

    Bai, Yushi and Lv, Xin and Zhang, Jiajie and Lyu, Hongchang and Tang, Jiankai and Huang, Zhidian and Du, Zhengxiao and Liu, Xiao and Zeng, Aohan and Hou, Lei and Dong, Yuxiao and Tang, Jie and Li, Juanzi , journal =. 2024 , url =

  45. [48]

    2025 , url =

    Bai, Yushi and Tu, Shangqing and Zhang, Jiajie and Peng, Hao and Wang, Xiaozhi and Lv, Xin and Cao, Shulin and Xu, Jiazheng and Hou, Lei and Dong, Yuxiao and Tang, Jie and Li, Juanzi , journal =. 2025 , url =

  46. [49]

    2024 , url =

    Hsieh, Cheng-Ping and Sun, Simeng and Kriman, Samuel and Acharya, Shantanu and Rekesh, Dima and Jia, Fei and Zhang, Yang and Ginsburg, Boris , journal =. 2024 , url =

  47. [50]

    arXiv preprint arXiv:2404.11018 , year =

    Many-Shot In-Context Learning , author =. arXiv preprint arXiv:2404.11018 , year =

  48. [51]

    2025 , url =

    Wang, Zhaowei and Yu, Wenhao and Ren, Xiyu and Zhang, Jipeng and Zhao, Yu and Saxena, Rohit and Cheng, Liang and Wong, Ginny and See, Simon and Minervini, Pasquale and Song, Yangqiu and Steedman, Mark , journal =. 2025 , url =

  49. [52]

    , journal =

    Li, Kaican and Yao, Lewei and Wu, Jiannan and Yu, Tiezheng and Chen, Jierun and Bai, Haoli and Hou, Lu and Hong, Lanqing and Zhang, Wei and Zhang, Nevin L. , journal =. 2025 , url =

  50. [53]

    From Easy to Hard: The

    Du, Hang and Zhang, Jiayang and Nan, Guoshun and Deng, Wendi and Chen, Zhenyan and Zhang, Chenyang and Xiao, Wang and Huang, Shan and Pan, Yuqi and Qi, Tao and Leng, Sicong , journal =. From Easy to Hard: The. 2025 , url =

  51. [54]

    and Baranwal, Aaditya and Coca, Alexandru and Dang, Mikah and Dziadzio, Sebastian and Kunz, Jakob D

    Roberts, Jonathan and Taesiri, Mohammad Reza and Sharma, Ansh and Gupta, Akash and Roberts, Samuel and Croitoru, Ioana and Bogolin, Simion-Vlad and Tang, Jialu and Langer, Florian and Raina, Vyas and Raina, Vatsal and Xiong, Hanyi and Udandarao, Vishaal and Lu, Jingyi and Chen, Shiyang and Purkis, Sam and Yan, Tianshuo and Lin, Wenye and Shin, Gyungin and...

  52. [55]

    Stepping

    Yang, Yuchen and Shao, Yuqing and Huang, Duxiu and Dong, Linfeng and Liu, Yifei and Tang, Suixin and Zhou, Xiang and Gao, Yuanyuan and Wang, Wei and Zhou, Yue and Yang, Xue and Wang, Yanfeng and Sun, Xiao and Zhong, Zhihang , journal =. Stepping. 2026 , url =

  53. [56]

    2025 , url =

    Feng, Sicheng and Wang, Song and Ouyang, Shuyi and Kong, Lingdong and Song, Zikai and Zhu, Jianke and Wang, Huan and Wang, Xinchao , journal =. 2025 , url =

  54. [57]

    2026 , url =

    Zhang, Huanyao and Zhou, Jiepeng and Li, Bo and Zhou, Bowen and Shan, Yanzhe and Lu, Haishan and Cao, Zhiyong and Chen, Jiaoyang and Han, Yuqian and Sheng, Zinan and Tao, Zhengwei and Liang, Hao and Wu, Jialong and Shi, Yang and He, Yuanpeng and Lin, Jiaye and Zhang, Qintong and Yan, Guochen and Zhao, Runhao and Li, Zhengpin and Yu, Xiaohan and Mei, Lang ...

  55. [58]

    Jehanzeb and Doveh, Sivan and Glass, James and Feris, Rogerio and Lin, Wei , journal =

    Selch, Lukas and Hou, Yufang and Mirza, M. Jehanzeb and Doveh, Sivan and Glass, James and Feris, Rogerio and Lin, Wei , journal =. 2025 , url =

  56. [59]

    2025 , url =

    Li, Zhuowan and Huang, Dong and Wang, Wenhao and Yang, Shenglong and Liu, Zhe and others , booktitle =. 2025 , url =

  57. [60]

    2026 , url =

    Ren, Yilong and An, Jingjing and Zhang, Luming and Zheng, Mingjie and Zhang, Wei and Yan, Songlin and Zhang, Hongyu and Zhang, Wenqi and Liu, Kai and Wu, Di and Chen, Hao and Yan, Wang and Liu, Yuanchun , journal =. 2026 , url =

  58. [61]

    2026 , url =

    Guo, Jiaheng and Wang, Ruidong and Yao, Chunlei and Zhang, Chenyang and Song, Yaqi and Wang, Zichao and He, Kun and Wu, Shiguang and Lin, Ziliang and Wei, Qipeng and Liao, Xuan and Liu, Xihui and Liu, Li and Lin, Ying-Ying and Liu, Wenbing and Zhou, Shuo and Wu, Qiang and Zhou, Ping and Zhao, Bingxuan and Zhang, Min and Liu, Tianyu and Lin, Dahua and Chen...

  59. [62]

    Richard and Wan, Xiang and Wang, Benyou , journal =

    Song, Dingjie and Chen, Shunian and Chen, Guiming Hardy and Yu, F. Richard and Wan, Xiang and Wang, Benyou , journal =. 2024 , url =

  60. [63]

    2024 , url =

    Ma, Yubo and Zang, Yuhang and Chen, Liangyu and Chen, Meiqi and Jiao, Yizhu and Li, Xinze and Lu, Xinyuan and others , journal =. 2024 , url =

  61. [64]

    2024 , url =

    Chia, Yew Ken and Cheng, Liying and Chan, Hou Pong and Liu, Chaoqun and Song, Maojia and Aljunied, Sharifah Mahani and Poria, Soujanya and Bing, Lidong , journal =. 2024 , url =

  62. [65]

    2025 , url =

    Dong, Kuicai and Chang, Yujing and Goh, Xin Deik and Li, Dexun and Tang, Ruiming and Liu, Yong , journal =. 2025 , url =

  63. [66]

    arXiv preprint arXiv:2406.07230 , year =

    Needle in a Multimodal Haystack , author =. arXiv preprint arXiv:2406.07230 , year =

  64. [67]

    arXiv preprint arXiv:2406.11230 , year =

    Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models , author =. arXiv preprint arXiv:2406.11230 , year =

  65. [68]

    and Li, Zekun and Liu, Qin and Liu, Xiaogeng and Ma, Mingyu Derek and Xu, Nan and Zhou, Wenxuan and Zhang, Kai and others , journal =

    Wang, Fei and Fu, Xingyu and Huang, James Y. and Li, Zekun and Liu, Qin and Liu, Xiaogeng and Ma, Mingyu Derek and Xu, Nan and Zhou, Wenxuan and Zhang, Kai and others , journal =. 2024 , url =

  66. [69]

    2024 , url =

    Meng, Fanqing and Wang, Jin and Li, Chuanhao and Lu, Quanfeng and Tian, Hao and Liao, Jiaqi and Zhu, Xizhou and Dai, Jifeng and others , journal =. 2024 , url =

  67. [70]

    2024 , url =

    Jiang, Dongfu and He, Xuan and Zeng, Huaye and Wei, Cong and Ku, Max and Liu, Qian and Chen, Wenhu , journal =. 2024 , url =

  68. [71]

    2024 , url =

    Shi, Yu and Tang, Chaoyue and Xu, Bokai and Cui, Junbo and Ran, Junhao and Yan, Yukun and Liu, Zhenghao and Wang, Shuo and Han, Xu and Liu, Zhiyuan and Sun, Maosong , journal =. 2024 , url =

  69. [72]

    and Omrani, Bilel and Viaud, Gautier and Hudelot, Celine and Colombo, Pierre , journal =

    Faysse, Manuel and Sibille, Hugues and Wu, Tony F. and Omrani, Bilel and Viaud, Gautier and Hudelot, Celine and Colombo, Pierre , journal =. 2024 , url =

  70. [73]

    Transactions of the Association for Computational Linguistics , volume =

    Lost in the Middle: How Language Models Use Long Contexts , author =. Transactions of the Association for Computational Linguistics , volume =. 2024 , doi =

  71. [74]

    Mathew, Minesh and Karatzas, Dimosthenis and Jawahar, C. V. , booktitle =. 2021 , url =

  72. [75]

    2022 , url =

    Masry, Ahmed and Long, Do Xuan and Tan, Jia Qing and Joty, Shafiq and Hoque, Enamul , booktitle =. 2022 , url =

  73. [76]

    2021 , doi =

    Chen, Zhiyu and Chen, Wenhu and Smiley, Charese and Shah, Sameena and Borova, Iana and Langdon, Dylan and Moussa, Reema and Beane, Matt and Huang, Ting-Hao and Routledge, Bryan and Wang, William Yang , booktitle =. 2021 , doi =

  74. [77]

    2021 , url =

    Zhu, Fengbin and Lei, Wenqiang and Huang, Youcheng and Wang, Chao and Zhang, Shuo and Lv, Jiancheng and Feng, Fuli and Chua, Tat-Seng , booktitle =. 2021 , url =

  75. [78]

    2022 , doi =

    Chen, Zhiyu and Li, Shiyang and Smiley, Charese and Ma, Zhiqiang and Shah, Sameena and Wang, William Yang , booktitle =. 2022 , doi =

  76. [79]

    Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , year =

    A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers , author =. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics , year =

  77. [80]

    2024 , url =

    Yue, Xiang and Ni, Yuansheng and Zhang, Kai and Zheng, Tianyu and Liu, Ruoqi and Zhang, Ge and Stevens, Samuel and Jiang, Dongfu and Ren, Weiming and Sun, Yuxuan and others , booktitle =. 2024 , url =

  78. [81]

    2024 , url =

    Liu, Yuan and Duan, Haodong and Zhang, Yuanhan and Li, Bo and Zhang, Songyang and Zhao, Wangbo and Yuan, Yike and Wang, Jiaqi and He, Conghui and Liu, Ziwei and Chen, Kai and Lin, Dahua , journal =. 2024 , url =