REVIEW 3 major objections 7 minor 53 references
Fluent deep-research reports often fail the evidence chain: models write well but miss grounded claims and answers.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 00:03 UTC pith:JDLW6MFL
load-bearing objection Solid diagnostic benchmark: the report-vs-grounding gap is real and useful, but the headline 4–11% answer-gate rates partly measure match-to-one-gold-graph, not pure reasoning failure. the 3 major comments →
HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On HiEviDR-Bench, strong surface-level report quality does not imply grounded multi-stage reasoning. Across sixteen multimodal large language models under both RAG and deep-research setups, report scores remain relatively high while citation accuracy, intermediate claim construction, and answer correctness drop markedly; answer-gate pass rates range only from about 3.8% to 11.5%, and the dominant bottlenecks are evidence identification and claim construction rather than final fluency.
What carries the argument
The hierarchical evidence graph: a directed acyclic graph of evidence nodes, intermediate claim nodes, and one conclusion node, with support edges that make selection, cross-source linking, and aggregation inspectable. Scoring is decomposed into five dimensions (report, traceability, citation, claim, answer), and progressive gates zero out later stages unless the preceding grounding requirement is met.
Load-bearing premise
The annotated evidence graphs and the judge-driven gates are treated as the correct gold standard for what counts as faithful hierarchical aggregation, rather than one plausible scaffold among many.
What would settle it
If independent human raters, using the same reports without privileging the annotated graphs, systematically credit models that fail the paper’s citation/claim/answer gates—or if swapping the judge model reverses the ranking of systems on those gated dimensions while report scores stay stable—then the claimed gap between fluency and grounded aggregation would not hold as measured.
If this is right
- Benchmarking deep research must score intermediate evidence-to-claim links, not only final report quality or short answers.
- Gains from iterative deep-research pipelines should be judged by better evidence filtering and claim support, not only by more retrieval rounds.
- Systems can look strong on multimodal report rendering while still failing citation accuracy and claim verification under progressive gates.
- Error localization will concentrate on key-evidence identification and intermediate claim construction before final answer writing.
- Open-domain versus academic and text versus multimodal splits can expose where hierarchical aggregation, not retrieval volume, is the limiting step.
Where Pith is reading between the lines
- Training or reward signals that target claim-level support and irrelevant-evidence suppression may move overall scores more than further fluency tuning.
- If graphs remain the gold standard, future work will need stress tests for alternate valid reasoning paths that the single annotated scaffold misses.
- Process-level gates could become a practical filter in agent evaluation loops: do not spend judge budget on answer quality until citation and claim gates pass.
- Multimodal report HTML with image placeholders may inflate perceived quality unless visual citations are checked against the same evidence graph as text.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces HiEviDR-Bench, a benchmark for evaluating hierarchical evidence aggregation in deep research systems. Each of 2,000 human-validated instances pairs a question with a multimodal corpus, a standard answer, and an explicit evidence graph (evidence → intermediate claims → conclusion). Evaluation decomposes into five 20-point dimensions (report quality, evidence traceability, citation, claim, answer), the last three governed by a progressive gating mechanism in which failure at one stage zeroes subsequent scores. Graphs are LLM-constructed over retrieved candidate pools, then rule-filtered, judge-filtered, and human-validated (30,000 → 3,407 → 2,000). Experiments on 16 MLLMs under RAG and Deep Research settings show report quality remains high while citation, claim, and answer scores drop sharply; answer-gate pass rates are only 3.8%–11.5%. The authors conclude the key bottleneck is evidence identification and intermediate claim construction rather than report fluency.
Significance. If the methodology holds, this is a useful contribution: it moves deep-research evaluation from outcome-only scoring to process-level traceability, and the evidence graph gives a concrete mechanism for error localization (evidence selection vs. claim construction vs. answer). The benchmark is sizable and human-validated (2,000 instances, high inter-annotator agreement), covers both open and academic domains and text-only/multimodal settings, and the authors release code and data, document the construction filters, run 16 systems under two paradigms, and provide judge-reliability and seed-robustness checks. However, the headline diagnostic numbers (gate pass rates) currently inherit two unvalidated assumptions — uniqueness of the gold reasoning path and reliability of the LLM judge's gate decisions — so the central claim should be treated as promising but not yet established.
major comments (3)
- [§3.2, Fig. 4] §3.2 (Progressive Gating Mechanism): the paper's most quotable result — answer-gate pass rates of 3.80%–11.50% (Fig. 4) — rests on the assumption that the annotated graph G is not merely *a* correct evidence→claim→conclusion decomposition but effectively the unique one. Per §3.3, graphs are LLM-constructed over a pre-selected 10–25-item candidate pool and human-validated for internal validity (Table 4 checks edges/claims for correctness, not exhaustiveness). A model that reaches the correct conclusion through a different, equally valid claim decomposition or a different evidence subset from the corpus fails the Claim/Answer gate and scores zero. The headline numbers may therefore partially measure graph-conformity. A load-bearing addition: on a human-checked sample, have annotators adjudicate gate failures as 'genuine grounding failure' vs 'valid alternative path', and report the split.
- [§3.2, Table 6] §3.2 and App. A.4: gate activation itself (whether cited evidence 'adequately supports the corresponding claim', i.e., the binary w_n decisions in Eqs. 11–12) is made by the Qwen3-VL-235B judge, yet Table 6 validates the judge only via Spearman correlation on continuous scores and system-rank consistency — not gate-level precision/recall against human pass/fail decisions. Because gates compound multiplicatively (a Claim-Gate false negative zeroes the Answer stage), even a modest gate false-negative rate materially moves the 3.8%–11.5% figure. Please report human-vs-judge agreement (κ or P/R) on the three binary gate decisions over a sample, ideally including the alternative judges (GPT-5, Gemini-3.1-Pro) already used in Table 6.
- [Table 3, §4.2–4.3] §4.2/Table 3 and §4.3: the ranking of 16 models and per-model RAG-vs-Deep-Research comparisons are reported as point estimates with no uncertainty quantification per system. Table 7 shows the *evaluation pipeline* is stable across seeds, but many inter-model gaps in Table 3 are under 1 point (e.g., 34.00 vs 34.58; 38.13 vs 39.09), which may not exceed per-instance variance. Bootstrap CIs over the 2,000 instances, or at minimum significance statements for the headline RAG-vs-DR gaps in Fig. 5, are needed before model-ordering claims (and the 'best/secondary' highlighting in Table 3) can stand.
minor comments (7)
- [§3.2, Eq. (6)] Eq. (6) appears to have a normalization error: as written, 20·(1/(5|K_r|))·Σ(r_k−1)/4 gives a maximum of 4 (with |K_r|=6), yet Table 3 reports Report scores ≈18–19. Presumably the intended form is 20·(1/|K_r|)·Σ(r_k−1)/4. Also, Eq. (6) normalizes via (r_k−1)/4 while Eqs. (11)–(13) use r_k/5, so a minimum judge score contributes 0 to Report but 1/5 elsewhere — please make the normalization consistent or explain the asymmetry.
- [§3.2, Eq. (9)] The TF-IDF irrelevance threshold τ (used to define E_irr^gold in Eq. 9) is never given a value, nor is its sensitivity analyzed. Since s3 enters S_trace (Eq. 10), τ should be stated and justified for reproducibility.
- [Table 3] Table 3 has formatting defects: several cells are missing separators (e.g., '19.805.44' and '3.680.75' in the GPT-5-mini RAG row; '19.61...' crowding in the DR block). The bold/underline best/secondary highlighting is also hard to verify in places.
- [Tables 3 and 8] Model naming is inconsistent: 'InternVL3.5-8B' (Table 3) vs 'Intern3.5-VL-8B' (Table 8, indices 11 and 14). Please unify.
- [§3.3 vs Table 2] §3.3 states Easy graphs average 5–8 nodes and Hard 9–17, but Table 2 shows Medium Wiki-Text averaging 10.79 and Medium arXiv-Text 10.15, overlapping the stated Hard range. Clarify whether difficulty is binned by node count, layer count, or a combination.
- [§3.3] The arXiv corpus is restricted to 'recent 2 year RAG papers' (Fig. 3). Please state the cutoff date and discuss potential data overlap with the pretraining corpora of the evaluated 2025–2026 models, which could inflate evidence identification on the arXiv subset.
- [§2] Related work (§2) would benefit from explicitly contrasting with MiroEval [42] and TRACE [2], which also pursue process-level evaluation; Table 5's row structure is useful but the text does not say what those two lack relative to the evidence-graph formulation.
Circularity Check
No circular derivation: empirical benchmark scores are measurements under an explicit metric, not predictions forced by their own inputs.
full rationale
HiEviDR-Bench is an evaluation paper, not a first-principles or fitted-theory paper. Its load-bearing claims are experimental outcomes on a newly defined five-dimension score with progressive gates (Eqs. 5–13, §3.2) and reported model numbers (Table 3, Fig. 4). Defining gold evidence graphs G and scoring reports against them is ordinary benchmark construction: the metric is stipulated, then systems are measured. Nothing in the chain reduces a claimed “prediction” to a fitted parameter or to a self-cited uniqueness theorem. LLM-assisted graph construction plus an LLM judge (Qwen3-VL-235B) raises validity questions about whether gates measure unique grounded reasoning versus conformity to one annotated scaffold, but that is metric-validity risk, not circularity by construction. The paper does not smuggle an ansatz via self-citation, rename a known law as a derivation, or treat a fit as an out-of-sample prediction. Self-contained against its own stated evaluation protocol; circularity score 0.
Axiom & Free-Parameter Ledger
free parameters (4)
- TF-IDF irrelevance threshold τ =
predefined threshold (numeric value not specified in main text)
- Retrieval budgets (top-15 RAG; top-20/keep 25 text+5 images deep research; ≤3 refinement rounds) =
top-15; top-20 then 25+5; 3 rounds
- Difficulty bins by evidence-graph node count/layers =
approx. Easy ~5–8 nodes / Hard ~9–17 nodes, up to 4 layers
- Per-dimension 1–5 LLM judge rubrics aggregated to 0–20 =
each component capped at 20; total /100
axioms (5)
- ad hoc to paper A directed acyclic evidence→claim→conclusion graph is the right process model of deep research for evaluation.
- ad hoc to paper Progressive gates should zero all later scores if citation/claim activation fails, even if the final answer text is correct.
- domain assumption LLM-as-judge scores on the stated rubrics are adequate proxies for human judgments of report, citation, claim, and answer quality.
- domain assumption Human validation after aggressive automatic filtering yields reliable gold graphs and QA pairs.
- domain assumption Standard retrieval-augmented and iterative deep-research pipelines are fair test harnesses for comparing MLLMs on this task.
invented entities (2)
-
HiEviDR evidence graph (Ne ∪ Nc ∪ {n_conclusion} with support edges)
independent evidence
-
Progressive Citation/Claim/Answer gates
no independent evidence
read the original abstract
Deep research requires models to retrieve, connect, and synthesize evidence from large-scale heterogeneous sources to answer complex queries and produce analytical reports. Existing benchmarks mainly evaluate final outcomes, such as answer correctness, report quality, or citation alignment, while providing limited visibility into whether evidence is correctly selected, linked, and aggregated into supported claims and conclusions. To address this gap, we introduce HiEviDR-Bench, a benchmark for evaluating Hierarchical Evidence Aggregation in Deep Research. HiEviDR-Bench covers open-domain and academic-domain settings under both text-only and multimodal conditions, and represents each instance with an explicit evidence graph that captures evidence selection, cross-source linking, and aggregation from evidence to intermediate claims and final conclusions. Based on this formulation, we develop a traceability-oriented evaluation framework with five dimensions: report quality, evidence traceability, citation accuracy, claim verification, and answer correctness, together with a progressive gating mechanism for fine-grained error localization. HiEviDR-Bench contains 2,000 human-validated questions with evidence graphs across multiple difficulty levels. Experiments on 16 representative multimodal large language models show that, although many systems achieve strong report quality, their performance drops markedly on citation accuracy, claim construction, and answer correctness. Further analysis shows that the main bottlenecks lie in evidence identification and intermediate claim construction, revealing that strong surface-level report quality does not necessarily imply grounded multi-stage reasoning on our benchmark.
Figures
Reference graph
Works this paper leans on
-
[1]
BiXie. 2024. wikiimage. https://huggingface.co/datasets/BiXie/wikiimage. Hug- ging Face dataset, accessed on 2026-03-24
2024
-
[2]
Yanyu Chen, Jiyue Jiang, Jiahong Liu, Yifei Zhang, Xiao Guo, and Irwin King
-
[3]
Zijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie, Kai Zou, Andrew Liu, Joshua Green, Kshama Patel, Ruoxi Meng, Mingyi Su, et al. 2025. Browsecomp- plus: A more fair and transparent evaluation benchmark of deep-research agent. arXiv preprint arXiv:2508.06600(2025)
Pith/arXiv arXiv 2025
-
[4]
Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, and Zhendong Mao. 2025. Deepresearch bench: A comprehensive benchmark for deep research agents. arXiv preprint arXiv:2506.11763(2025)
Pith/arXiv arXiv 2025
-
[5]
Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, and Anthony KH Tung. 2026. IDRBench: Interactive Deep Research Benchmark. arXiv preprint arXiv:2601.06676(2026)
Pith/arXiv arXiv 2026
-
[6]
Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. 2023. Enabling large lan- guage models to generate text with citations. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 6465–6488
2023
-
[7]
Tiansheng Hu, Yilun Zhao, Canyu Zhang, Arman Cohan, and Chen Zhao. 2026. SAGE: Benchmarking and Improving Retrieval for Deep Research Agents.arXiv preprint arXiv:2602.05975(2026)
arXiv 2026
-
[8]
Peizhou Huang, Zixuan Zhong, Zhongwei Wan, Donghao Zhou, Samiul Alam, Xin Wang, Zexin Li, Zhihao Dou, Li Zhu, Jing Xiong, et al. 2026. MMDeepResearch- Bench: A Benchmark for Multimodal Deep Research Agents.arXiv preprint arXiv:2601.12346(2026)
arXiv 2026
-
[9]
Yuxuan Huang, Yihang Chen, Haozheng Zhang, Kang Li, Huichi Zhou, Meng Fang, Linyi Yang, Xiaoguang Li, Lifeng Shang, Songcen Xu, et al . 2025. Deep research agents: A systematic examination and roadmap.arXiv preprint arXiv:2506.18096(2025)
Pith/arXiv arXiv 2025
-
[10]
Minghao Li, Ying Zeng, Zhihao Cheng, Cong Ma, and Kai Jia. 2025. Report- bench: Evaluating deep research agents via academic survey tasks.arXiv preprint arXiv:2508.15804(2025)
Pith/arXiv arXiv 2025
-
[11]
Mingxin Li, Yanzhao Zhang, Dingkun Long, Chen Keqin, Sibo Song, Shuai Bai, Zhibo Yang, Pengjun Xie, An Yang, Dayiheng Liu, Jingren Zhou, and Junyang Lin. 2026. Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Frame- work for State-of-the-Art Multimodal Retrieval and Ranking.arXiv preprint arXiv:2601.04720(2026)
Pith/arXiv arXiv 2026
-
[12]
Zhenghao Liu, Pengcheng Huang, Zhipeng Xu, Xinze Li, Shuliang Liu, Chunyi Peng, Haidong Xin, Yukun Yan, Shuo Wang, Xu Han, et al . 2026. Knowledge intensive agents.AI Open(2026)
2026
-
[13]
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al
-
[14]
Chunyi Peng, Zhipeng Xu, Zhenghao Liu, Yishan Li, Yukun Yan, Shuo Wang, Yu Gu, Minghe Yu, Ge Yu, and Maosong Sun. 2026. Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation. InProceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1440–1450
2026
-
[15]
Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents. https://qwen. ai/blog?id=qwen3.5
2026
-
[16]
Yiming Ren, Junjie Wang, Yuxin Meng, Yihang Shi, Zhiqiang Lin, Ruihang Chu, Yiran Xu, Ziming Li, Yunfei Zhao, Zihan Wang, et al . 2026. SIN-Bench: Trac- ing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature.arXiv preprint arXiv:2601.10108(2026)
arXiv 2026
-
[17]
2026.Seed 2.0 Model Card: Towards Intelligence Frontier for Real- World Complexity
ByteDance Seed. 2026.Seed 2.0 Model Card: Towards Intelligence Frontier for Real- World Complexity. Technical Report. Technical report (model card), February
2026
-
[18]
Yijia Shao, Yucheng Jiang, Theodore Kanell, Peter Xu, Omar Khattab, and Mon- ica Lam. 2024. Assisting in writing wikipedia-like articles from scratch with large language models. InProceedings of the 2024 Conference of the North Ameri- can Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 6252–6278
2024
-
[19]
Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al
-
[20]
URL https://lf3-static
-
[21]
Yubo Sun, Chunyi Peng, Yukun Yan, Shi Yu, Zhenghao Liu, Chi Chen, Zhiyuan Liu, and Maosong Sun. 2025. VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation.arXiv preprint arXiv:2510.09733(2025)
Pith/arXiv arXiv 2025
-
[22]
Gemma Team. 2026. Gemma 4 Technical Report. arXiv:2607.02770 [cs.CL] https://arxiv.org/abs/2607.02770
Pith/arXiv arXiv 2026
-
[23]
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805(2023)
Pith/arXiv arXiv 2023
-
[24]
Ionut Teodor Sorodoc, Leonardo FR Ribeiro, Rexhina Blloshmi, Christopher Davis, and Adrià de Gispert. 2025. Garage: A benchmark with grounding annotations for rag evaluation. InFindings of the Association for Computational Linguistics: ACL 2025. 17030–17049
2025
-
[25]
V Team, Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, Aohan Zeng, Baoxu Wang, Bin Chen, Boyan Shi, Changyu Pang, Chenhui Zhang, Da Yin, Fan Yang, Guoqing Chen, Jiazheng Xu, Jiale Zhu, Jiali Chen, J...
Pith/arXiv arXiv 2025
-
[26]
Rubèn Tito, Dimosthenis Karatzas, and Ernest Valveny. 2023. Hierarchical multi- modal transformers for Multipage DocVQA.Pattern Recognit.144 (2023), 109834
2023
-
[27]
UltraRAG. 2025. UltraRAG Benchmark. https://modelscope.cn/datasets/ UltraRAG/UltraRAG_Benchmark
2025
-
[29]
Jiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen, Austin Xu, Zixuan Ke, Frederic Sala, Aws Albarghouthi, Caiming Xiong, and Shafiq Joty. 2025. Livere- searchbench: A live benchmark for user-centric deep research in the wild.arXiv preprint arXiv:2510.14240(2025)
Pith/arXiv arXiv 2025
-
[30]
Qiuchen Wang, Ruixue Ding, Zehui Chen, Weiqi Wu, Shihang Wang, Pengjun Xie, and Feng Zhao. 2025. ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents.CoRRabs/2502.18017 (2025)
Pith/arXiv arXiv 2025
-
[31]
Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, et al. 2025. InternVL3.5: Ad- vancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency. arXiv preprint arXiv:2508.18265(2025)
Pith/arXiv arXiv 2025
-
[32]
Pranav Narayanan Venkit, Philippe Laban, Yilun Zhou, Kung-Hsiang Huang, Yixin Mao, and Chien-Sheng Wu. 2025. Deeptrace: Auditing deep research ai systems for tracking reliability across citations and evidence.arXiv preprint arXiv:2509.04499(2025)
Pith/arXiv arXiv 2025
-
[33]
xAI. 2026. Introducing Grok 4.5. https://x.ai/news/grok-4-5. Accessed: 2026-07- 21
2026
-
[34]
Yilin Xiao, Junnan Dong, Chuang Zhou, Su Dong, Qian-wen Zhang, Di Yin, Xing Sun, and Xiao Huang. 2025. Graphrag-bench: Challenging domain-specific reasoning for evaluating graph retrieval-augmented generation.arXiv preprint arXiv:2506.02404(2025)
Pith/arXiv arXiv 2025
-
[35]
LLM-Core-Team Xiaomi. 2025. MiMo-VL Technical Report. arXiv:2506.03569 [cs.CL] https://arxiv.org/abs/2506.03569
Pith/arXiv arXiv 2025
-
[36]
Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese
-
[37]
arXiv preprint arXiv:2504.12516(2025)
Browsecomp: A simple yet challenging benchmark for browsing agents. arXiv preprint arXiv:2504.12516(2025)
Pith/arXiv arXiv 2025
-
[38]
Bangrui Xu, Qihang Yao, Zirui Tang, Xuanhe Zhou, Yeye He, Shihan Yu, Qianqian Xu, Bin Wang, Guoliang Li, Conghui He, et al. 2026. MoDora: Tree-Based Semi- Structured Document Analysis System.arXiv preprint arXiv:2602.23061(2026)
Pith/arXiv arXiv 2026
-
[39]
Renjun Xu and Jingwen Peng. 2025. A comprehensive survey of deep research: Systems, methodologies, and applications.arXiv preprint arXiv:2506.12594(2025)
Pith/arXiv arXiv 2025
-
[40]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025)
Pith/arXiv arXiv 2025
-
[41]
Lei Xiong, Huaying Yuan, Zheng Liu, Zhao Cao, and Zhicheng Dou. 2026. Paper- Scope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers. InFindings of the Association for Computational Linguistics: ACL 2026. 8015–8040
2026
-
[42]
Yuqi Xiong, Chunyi Peng, Zhipeng Xu, Zhenghao Liu, Zulong Chen, Yukun Yan, Shuo Wang, Yu Gu, and Ge Yu. 2026. Lang2act: Fine-grained visual reasoning through self-emergent linguistic toolchains. InFindings of the Association for Computational Linguistics: ACL 2026. 8375–8399
2026
-
[43]
Qinhan Yu, Zhiyou Xiao, Binghui Li, Zhengren Wang, Chong Chen, and Wentao Zhang. 2025. MRAMG-Bench: a comprehensive benchmark for advancing mul- timodal retrieval-augmented multimodal generation. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 3616–3626
2025
-
[44]
Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chengxing Xie, Cunxiang Wang, et al. 2026. GLM-5: from Vibe Coding to Agentic Engineering.arXiv preprint arXiv:2602.15763(2026)
Pith/arXiv arXiv 2026
-
[45]
Yu Zeng, Wenxuan Huang, Zhen Fang, Shuang Chen, Yufan Shen, Yishuo Cai, Xi- aoman Wang, Zhenfei Yin, Lin Chen, Zehui Chen, et al. 2026. Vision-deepresearch benchmark: Rethinking visual and textual search for multimodal large language models.arXiv preprint arXiv:2602.02185(2026)
arXiv 2026
-
[46]
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al. 2025. MiniCPM-V: A GPT-4V Level MLLM on Your Phone.Nat Commun 16, 5509 (2025)(2025)
2025
-
[47]
Fangda Ye, Yuxin Hu, Pengxiang Zhu, Yibo Li, Ziqi Jin, Yao Xiao, Yibo Wang, Lei Wang, Zhen Zhang, Lu Wang, et al. 2026. MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome.arXiv preprint arXiv:2603.28407 (2026). Conference’17, July 2017, Washington, DC, USA Yubo Sun, Chunyi Peng, Yukun Yan, Zhenghao Liu, Sen Mei, Bangrui Xu, Xuan...
arXiv 2026
-
[48]
Jiajie Zhang, Yushi Bai, Xin Lv, Wanjun Gu, Danqing Liu, Minhao Zou, Shulin Cao, Lei Hou, Yuxiao Dong, Ling Feng, et al. 2025. Longcite: Enabling llms to generate fine-grained citations in long-context qa. InFindings of the Association for Computational Linguistics: ACL 2025. 5098–5122
2025
-
[49]
cover">`-multiple `<div class=
Wenlin Zhang, Xiaopeng Li, Yingyi Zhang, Pengyue Jia, Yichao Wang, Huifeng Guo, Yong Liu, and Xiangyu Zhao. 2025. Deep research: A survey of autonomous research agents.arXiv preprint arXiv:2508.12752(2025). HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research Conference’17, July 2017, Washington, DC, USA A Appendix A.1 Human V...
Pith/arXiv arXiv 2025
-
[51]
Dingling Zhang, He Zhu, Jincheng Ren, Kangqi Song, Xinran Zhou, Boyu Feng, Shudong Liu, Jiabin Luo, Weihao Xie, Zhaohui Wang, et al. 2025. How Far Are We from Genuinely Useful Deep Research Agents?arXiv preprint arXiv:2512.01948 (2025)
arXiv 2025
-
[52]
Huanyao Zhang, Jiepeng Zhou, Bo Li, Bowen Zhou, Yanzhe Shan, Haishan Lu, Zhiyong Cao, Jiaoyang Chen, Yuqian Han, Zinan Sheng, et al. 2026. BrowseComp- V3: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents. arXiv preprint arXiv:2602.12876(2026)
arXiv 2026
-
[2021]
Webgpt: Browser-assisted question-answering with human feedback.arXiv preprint arXiv:2112.09332(2021)
Pith/arXiv arXiv 2021
-
[2025]
Openai gpt-5 system card.arXiv preprint arXiv:2601.03267(2025)
Pith/arXiv arXiv 2025
-
[2026]
TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents.arXiv preprint arXiv:2602.21230(2026)
arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.