REVIEW 2 major objections 2 minor 43 references
OpenBioRQ: Unsolved Biomedical Research Questions for Agents
T0 review · 2 major / 2 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read OpenBioRQ presents 12,553 unsolved biomedical questions as a test for whether agentic models can verify sources without answer keys.
desk verdict OpenBioRQ gives a non-saturating benchmark on open biomedical questions with concrete observations on wrong citations and tool collapse, but the unsolved status needs stronger exhaustiveness checks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
OpenBioRQ, a collection of unsolved biomedical questions that forces multiple tool calls for citation verification and has no fixed answer key, with difficulty defined by failure of three reference open-weight models.
What would settle it
A new agentic system that solves more than 70 percent of the hardest subset while continuing to issue tool calls on those same items would contradict the reported performance ceiling and collapse pattern.
Extended reading notes
Core claim
OpenBioRQ is a retrieval-grounded agentic benchmark of 12,553 unsolved biomedical research questions that treats open questions as a faithfulness-and-abstention probe; openness is verified against real follow-up evidence, difficulty is anchored on items failed by three open-weight reference models, held-out models from the same lineage solve only about 17 percent of the hardest subset, and three frontier agents span 29-60 percent while showing agentic collapse where tool use stops.
Load-bearing premise
The selected questions are genuinely open and their difficulty is correctly measured by the failure of the three reference models rather than by any model’s internal knowledge.
Editorial extensions
If this is right
- Frontier agents leave 33-40 percent of the hardest questions unsolved even when tools are available.
- On the collapse-prone model, removing tool access changes the score by only a small amount.
- A static per-question checklist lifts inter-judge Spearman correlation from 0.35 to 0.82.
- The benchmark remains non-saturating across current capability tiers.
Reading between the lines
- The same unsolved-question design could be applied in chemistry or physics to measure retrieval faithfulness outside biomedicine.
- Persistent tool-use collapse suggests that current training objectives may reward parametric recall more than sustained verification on open problems.
- Future agent training could explicitly penalize early abandonment of tool sequences on items that reference models already miss.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces OpenBioRQ, a benchmark of 12,553 unsolved biomedical research questions across 12 domains. It evaluates agentic models in a tool-using setting on these questions (with no fixed answer key), reporting that held-out models solve only ~17% of the hardest subset while frontier agents range from 29-60%. The work claims the benchmark is non-saturating, reveals agentic collapse (reduced tool use on hard questions), and shows that a frozen checklist raises inter-judge agreement from Spearman 0.35 to 0.82. Openness is asserted to be verified via real follow-up evidence rather than parametric knowledge.
Significance. If the verification of unsolved status is robust and the empirical difficulty anchoring holds, the benchmark would fill a gap by providing a non-saturating, retrieval-grounded testbed for agentic faithfulness and abstention in biomedicine. The empirical anchoring on reference-model failures and the checklist for judging are concrete strengths that support reproducibility and could influence future open-question benchmarks.
major comments (2)
- [Dataset construction] Dataset construction section: The central claim that the benchmark consists of verifiably unsolved questions (and therefore validly measures open-question handling rather than retrieval) rests on the openness verification against real follow-up evidence. The manuscript provides no details on search exhaustiveness (databases, time window, keyword strategy, or coverage of preprints/obscure venues). This is load-bearing; incomplete verification risks including solved questions, which would mean the reported 17-60% solve rates and the agentic-collapse observations partly measure retrieval success instead of the intended probe.
- [Results] Results section on agentic collapse: The claim that 'for the most collapse-prone model, blocking tool access entirely barely changes its score' is presented as evidence that tools stop paying off where needed most. However, the manuscript lacks controls for prompt sensitivity, baseline performance without tools, or statistical tests on the score difference. This detail is required to support the collapse interpretation as load-bearing for the agentic-setting contribution.
minor comments (2)
- [Abstract] Abstract: Model names (Gemini-3-Pro, Opus-4.7, GPT-5.5) appear stylized; clarify whether these are exact versions or anonymized for the paper.
- [Evaluation] Evaluation protocol: The Spearman correlation values (0.35 to 0.82) for inter-judge agreement are reported, but the number of judges, exact items rated, and whether the checklist was applied to all questions should be stated explicitly.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback. We address the two major comments point-by-point below and will revise the manuscript to strengthen the relevant sections.
read point-by-point responses
-
Referee: [Dataset construction] Dataset construction section: The central claim that the benchmark consists of verifiably unsolved questions (and therefore validly measures open-question handling rather than retrieval) rests on the openness verification against real follow-up evidence. The manuscript provides no details on search exhaustiveness (databases, time window, keyword strategy, or coverage of preprints/obscure venues). This is load-bearing; incomplete verification risks including solved questions, which would mean the reported 17-60% solve rates and the agentic-collapse observations partly measure retrieval success instead of the intended probe.
Authors: We agree that explicit details on the verification process are necessary to substantiate the unsolved status and to rule out retrieval confounds. In the revised manuscript we will expand the Dataset construction section with a dedicated subsection describing the search protocol: the databases queried (PubMed, Google Scholar, bioRxiv, medRxiv, arXiv), the temporal window (queries run through [specific cutoff date]), the Boolean keyword strategies and MeSH terms employed, and the additional manual checks performed for obscure venues and preprints. These additions will make the verification procedure fully reproducible and directly address the concern that solved questions may have been inadvertently included. revision: yes
-
Referee: [Results] Results section on agentic collapse: The claim that 'for the most collapse-prone model, blocking tool access entirely barely changes its score' is presented as evidence that tools stop paying off where needed most. However, the manuscript lacks controls for prompt sensitivity, baseline performance without tools, or statistical tests on the score difference. This detail is required to support the collapse interpretation as load-bearing for the agentic-setting contribution.
Authors: We acknowledge that the current presentation of the tool-blocking experiment would benefit from additional controls and statistical support. In the revision we will (i) report explicit baseline performance for every model when tool access is disabled, (ii) include a prompt-sensitivity analysis across at least three prompt variants, and (iii) add statistical tests (paired McNemar tests with exact p-values and effect sizes) comparing scores with and without tools on the hardest subset. These changes will place the agentic-collapse observation on firmer empirical footing. revision: yes
Circularity Check
No circularity: empirical benchmark construction and measurements
full rationale
The paper introduces OpenBioRQ as a new benchmark of 12,553 unsolved questions, with results consisting of empirical solve rates (~17% on hardest subset for held-out models, 29-60% for frontier agents), agentic collapse observations, and inter-judge agreement improvements. Difficulty is defined by reference models failing and openness by follow-up evidence checks; these are methodological choices for subset selection and verification, not derivations, equations, or predictions that reduce to fitted inputs by construction. No self-citations, ansatzes, uniqueness theorems, or renamings of known results appear as load-bearing steps in the derivation chain. The central claims are direct measurements on the introduced benchmark and remain self-contained against external evaluation.
Assumptions & free parameters
Cite this review
Pith. "Pith review of OpenBioRQ: Unsolved Biomedical Research Questions for Agents." pith.science (2026). https://pith.science/paper/KTGEZH2F
@misc{pith2026260621959,
author = {Pith},
title = {Pith review of: OpenBioRQ: Unsolved Biomedical Research Questions for Agents},
year = {2026},
howpublished = {\url{https://pith.science/paper/KTGEZH2F}},
note = {Machine review of arXiv:2606.21959}
}
abstract
A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the claim. I find that current agentic models rarely fabricate citations (over $99\%$ resolve), yet roughly $15.9\%$ link to the wrong paper. Existing benchmarks miss this failure mode: when a question has a fixed answer key, a model can reproduce the expected source from that key rather than independently verifying that the source supports the claim. I introduce \textbf{\openbiorq{}}, a retrieval-grounded agentic benchmark of $12{,}553$ unsolved biomedical research questions across $12$ domains that treats open questions as a faithfulness-and-abstention probe. To my knowledge, this is the first biomedical benchmark to combine an agentic setting -- where the model must issue multiple tool calls -- with unsolved questions that have no answer key. Openness is verified against real follow-up evidence rather than a model's parametric knowledge. Difficulty is empirical: I anchor it on questions that three open-weight reference models fail to answer, rather than on subjective hardness labels. On this hardest subset, held-out models from the same lineage as the difficulty anchors solve only ~17%, while three independent frontier agents (Gemini-3-Pro, Opus-4.7, GPT-5.5) span a wide 29-60% range. The benchmark is thus hard, non-saturating (the best agent still leaves ~33-40\% unsolved), and discriminating across capability tiers. Beyond difficulty, I observe agentic collapse on the hardest questions, where agents stop using their tools. For the most collapse-prone model, blocking tool access entirely barely changes its score -- so tools stop paying off exactly where they are needed most. A frozen per-question checklist raises inter-judge agreement from Spearman 0.35 to 0.82.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
2026 , howpublished =
2026
-
[2]
Applied Sciences , volume =
What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams , author =. Applied Sciences , volume =. 2021 , note =
2021
-
[3]
2019 , publisher =
Jin, Qiao and Dhingra, Bhuwan and Liu, Zhengping and Cohen, William and Lu, Xinghua , booktitle =. 2019 , publisher =
2019
-
[4]
Pal, Ankit and Umapathi, Logesh Kumar and Sankarasubbu, Malaikannan , booktitle =
-
[5]
Tsatsaronis, George and Balikas, Georgios and Malakasiotis, Prodromos and Partalas, Ioannis and Zschunke, Matthias and Alvers, Michael R. and Weissenborn, Dirk and Krithara, Anastasia and Petridis, Sergios and Polychronopoulos, Dimitris and Almirantis, Yannis and Pavlopoulos, John and Baskiotis, Nicolas and Gallinari, Patrick and Arti. An Overview of the....
2015
-
[6]
Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages =
Vilares, David and G. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages =. 2019 , publisher =
2019
-
[7]
International Conference on Learning Representations (ICLR) , year =
Measuring Massive Multitask Language Understanding , author =. International Conference on Learning Representations (ICLR) , year =
-
[8]
Son, Guijin and Yi, Seungyeop and Gwak, Minju and Ko, Hyunwoo and Jang, Wongi and Yu, Youngjae , year =
Show all 43 references
-
[9]
2023 , note =
Yao, Shunyu and Zhao, Jeffrey and Yu, Dian and Du, Nan and Shafran, Izhak and Narasimhan, Karthik and Cao, Yuan , booktitle =. 2023 , note =
2023
-
[10]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Toolformer: Language Models Can Teach Themselves to Use Tools , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[11]
2024 , note =
Qin, Yujia and Liang, Shihao and Ye, Yining and Zhu, Kunlun and Yan, Lan and Lu, Yaxi and Lin, Yankai and Cong, Xin and Tang, Xiangru and Qian, Bill and Zhao, Sihan and Hong, Lauren and Tian, Runchu and Xie, Ruobing and Zhou, Jie and Gerstein, Mark and Li, Dahai and Liu, Zhiyu...
2024
-
[12]
and Zhang, Tianjun and Wang, Xin and Gonzalez, Joseph E
Patil, Shishir G. and Zhang, Tianjun and Wang, Xin and Gonzalez, Joseph E. , year =. Gorilla: Large Language Model Connected with Massive
-
[13]
2024 , note =
Jin, Qiao and Yang, Yifan and Chen, Qingyu and Lu, Zhiyong , journal =. 2024 , note =
2024
-
[14]
and Wilder, Esther Isabelle , journal =
Walters, William H. and Wilder, Esther Isabelle , journal =. Fabrication and Errors in the Bibliographic Citations Generated by. 2023 , note =
2023
-
[15]
Computational Linguistics , volume =
Measuring Attribution in Natural Language Generation Models , author =. Computational Linguistics , volume =. 2023 , note =
2023
-
[16]
2022 , howpublished =
Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models , author =. 2022 , howpublished =
2022
-
[17]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages =
Enabling Large Language Models to Generate Text with Citations , author =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages =. 2023 , publisher =
2023
-
[18]
Findings of the Association for Computational Linguistics: EMNLP 2023 , pages =
Evaluating Verifiability in Generative Search Engines , author =. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages =. 2023 , publisher =
2023
-
[19]
2020 , note =
Wang, Wenhui and Wei, Furu and Dong, Li and Bao, Hangbo and Yang, Nan and Zhou, Ming , booktitle =. 2020 , note =
2020
-
[20]
2025 , howpublished =
2025
-
[21]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , booktitle =. Judging. 2023 , note =
2023
-
[22]
2023 , publisher =
Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang , booktitle =. 2023 , publisher =
2023
-
[23]
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL) , pages =
Selective Question Answering under Domain Shift , author =. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL) , pages =. 2020 , publisher =
2020
-
[24]
2022 , howpublished =
Language Models (Mostly) Know What They Know , author =. 2022 , howpublished =
2022
-
[25]
2023 , howpublished =
Mialon, Gr. 2023 , howpublished =
2023
-
[26]
Yao, Shunyu and Shinn, Noah and Razavi, Pedram and Narasimhan, Karthik , year =
-
[27]
Wei, Jason and Sun, Zhiqing and Papay, Spencer and McKinney, Scott and Han, Jeffrey and Fulford, Isa and Chung, Hyung Won and Passos, Alex Tachard and Fedus, William and Glaese, Amelia , year =
-
[28]
2024 , publisher =
Trivedi, Harsh and Khot, Tushar and Hartmann, Mareike and Manku, Ruskin and Dong, Vinty and Li, Edward and Gupta, Shashank and Sabharwal, Ashish and Balasubramanian, Niranjan , booktitle =. 2024 , publisher =
2024
-
[29]
, year =
Rein, David and Hou, Betty Li and Stickland, Asa Cooper and Petty, Jackson and Pang, Richard Yuanzhe and Dirani, Julien and Michael, Julian and Bowman, Samuel R. , year =
-
[30]
and Janizek, Joseph D
Laurent, Jon M. and Janizek, Joseph D. and Ruzo, Michael and Hinks, Michaela M. and Hammerling, Michael J. and Narayanan, Siddharth and Ponnapati, Manvitha and White, Andrew D. and Rodriques, Samuel G. , year =
-
[31]
Nature , volume =
Large Language Models Encode Clinical Knowledge , author =. Nature , volume =. 2023 , note =
2023
-
[32]
2023 , howpublished =
Towards Expert-Level Medical Question Answering with Large Language Models , author =. 2023 , howpublished =
2023
-
[33]
Communications of the ACM , volume =
Datasheets for Datasets , author =. Communications of the ACM , volume =. 2021 , note =
2021
-
[34]
2026 , howpublished=
BioMysteryBench: Evaluating Claude's Bioinformatics Research Capabilities , author=. 2026 , howpublished=
2026
-
[35]
Nature Medicine , year=
Generalist Large Language Models Outperform Clinical Tools on Medical Benchmarks , author=. Nature Medicine , year=
-
[36]
and Wei, Jason and Soskin Hicks, Rebecca and Bowman, Preston and Qui
Arora, Rahul K. and Wei, Jason and Soskin Hicks, Rebecca and Bowman, Preston and Qui. 2025 , howpublished =
2025
-
[37]
arXiv preprint arXiv:2605.02943 , year=
Healthcare AI GYM for Medical Agents , author=. arXiv preprint arXiv:2605.02943 , year=
-
[38]
2024 , note =
Jeong, Minbyul and Hwang, Hyeon and Yoon, Chanwoong and Lee, Taewhoo and Kang, Jaewoo , booktitle =. 2024 , note =
2024
-
[39]
Schmidgall, Samuel and Ziaei, Rojin and Harris, Carl and Reis, Eduardo and Jopling, Jeffrey and Moor, Michael , year =
-
[40]
and Geng, Gloria and Park, Danny and Zou, James and Ng, Andrew Y
Jiang, Yixing and Black, Kameron C. and Geng, Gloria and Park, Danny and Zou, James and Ng, Andrew Y. and Chen, Jonathan H. , year =
-
[41]
, author=
Bridging the Gap Between Consumers' Medication Questions and Trusted Answers. , author=. MedInfo , year=
-
[42]
, author=
Overview of the medical question answering task at TREC 2017 LiveQA. , author=. TREC , year=
2017
-
[43]
arXiv preprint arXiv:2401.14493 , year=
K-QA: A Real-World Medical Q&A Benchmark , author=. arXiv preprint arXiv:2401.14493 , year=
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.