{"id":"0c0503bf-0524-4dc0-a61c-775a7010c5cc","arxiv_id":"2606.21959","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces an agentic benchmark of unsolved biomedical questions to probe citation faithfulness and tool-use collapse in frontier models.","lead":"The paper presents OpenBioRQ, a benchmark of over 12,000 unsolved biomedical questions designed to test whether AI agents can correctly cite sources and know when to abstain rather than guess. A smart generalist might read it to understand current limits in using AI for real scientific literature work.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Openness verification against follow-up evidence lacks an exhaustiveness guarantee, risking inclusion of solved questions","rationale":"This directly matches the reader's weakest_assumption and is the single point where the benchmark's premise is least secured; all downstream claims about hardness, non-saturation, and tool-use collapse inherit the risk. The empirical anchoring on three reference models is secondary because it presupposes the questions are open in the first place.","tokens_in":1835,"tokens_out":332,"duration_ms":44201,"concrete_test":"Draw a random sample of 100 questions from the hardest subset; run an independent search on PubMed, Google Scholar, bioRxiv, and arXiv using the question text plus date filter after the paper's collection cutoff; count how many return clear solving follow-up papers. If >3% are solved, recompute the headline solve rates on the remaining questions.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that the benchmark consists of truly unsolved questions and therefore validly probes faithfulness/abstention and agentic collapse—rests on the assertion that openness was verified against real follow-up evidence rather than parametric knowledge. Any non-exhaustive search (limited databases, time window, or keyword strategy) could miss solutions in recent preprints or obscure venues. If even a modest fraction of the 12,553 questions (or the hardest subset) have been resolved, then reported solve rates (~17% held-out, 29-60% frontier) partly measure retrieval success instead of proper open-question handling, directly undermining the non-saturating and collapse observations.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces OpenBioRQ, a benchmark of 12,553 unsolved biomedical research questions across 12 domains. It evaluates agentic models in a tool-using setting on these questions (with no fixed answer key), reporting that held-out models solve only ~17% of the hardest subset while frontier agents range from 29-60%. The work claims the benchmark is non-saturating, reveals agentic collapse (reduced tool use on hard questions), and shows that a frozen checklist raises inter-judge agreement from Spearman 0.35 to 0.82. Openness is asserted to be verified via real follow-up evidence rather than parametric knowledge.","tokens_in":1988,"tokens_out":592,"duration_ms":24257,"significance":"If the verification of unsolved status is robust and the empirical difficulty anchoring holds, the benchmark would fill a gap by providing a non-saturating, retrieval-grounded testbed for agentic faithfulness and abstention in biomedicine. The empirical anchoring on reference-model failures and the checklist for judging are concrete strengths that support reproducibility and could influence future open-question benchmarks.","major_comments":[{"comment":"Dataset construction section: The central claim that the benchmark consists of verifiably unsolved questions (and therefore validly measures open-question handling rather than retrieval) rests on the openness verification against real follow-up evidence. The manuscript provides no details on search exhaustiveness (databases, time window, keyword strategy, or coverage of preprints/obscure venues). This is load-bearing; incomplete verification risks including solved questions, which would mean the reported 17-60% solve rates and the agentic-collapse observations partly measure retrieval success instead of the intended probe.","section":"Dataset construction"},{"comment":"Results section on agentic collapse: The claim that 'for the most collapse-prone model, blocking tool access entirely barely changes its score' is presented as evidence that tools stop paying off where needed most. However, the manuscript lacks controls for prompt sensitivity, baseline performance without tools, or statistical tests on the score difference. This detail is required to support the collapse interpretation as load-bearing for the agentic-setting contribution.","section":"Results"}],"minor_comments":[{"comment":"Abstract: Model names (Gemini-3-Pro, Opus-4.7, GPT-5.5) appear stylized; clarify whether these are exact versions or anonymized for the paper.","section":"Abstract"},{"comment":"Evaluation protocol: The Spearman correlation values (0.35 to 0.82) for inter-judge agreement are reported, but the number of judges, exact items rated, and whether the checklist was applied to all questions should be stated explicitly.","section":"Evaluation"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address the two major comments point-by-point below and will revise the manuscript to strengthen the relevant sections.","responses":[{"response":"We agree that explicit details on the verification process are necessary to substantiate the unsolved status and to rule out retrieval confounds. In the revised manuscript we will expand the Dataset construction section with a dedicated subsection describing the search protocol: the databases queried (PubMed, Google Scholar, bioRxiv, medRxiv, arXiv), the temporal window (queries run through [specific cutoff date]), the Boolean keyword strategies and MeSH terms employed, and the additional manual checks performed for obscure venues and preprints. These additions will make the verification procedure fully reproducible and directly address the concern that solved questions may have been inadvertently included.","revision_made":"yes","referee_comment":"[Dataset construction] Dataset construction section: The central claim that the benchmark consists of verifiably unsolved questions (and therefore validly measures open-question handling rather than retrieval) rests on the openness verification against real follow-up evidence. The manuscript provides no details on search exhaustiveness (databases, time window, keyword strategy, or coverage of preprints/obscure venues). This is load-bearing; incomplete verification risks including solved questions, which would mean the reported 17-60% solve rates and the agentic-collapse observations partly measure retrieval success instead of the intended probe."},{"response":"We acknowledge that the current presentation of the tool-blocking experiment would benefit from additional controls and statistical support. In the revision we will (i) report explicit baseline performance for every model when tool access is disabled, (ii) include a prompt-sensitivity analysis across at least three prompt variants, and (iii) add statistical tests (paired McNemar tests with exact p-values and effect sizes) comparing scores with and without tools on the hardest subset. These changes will place the agentic-collapse observation on firmer empirical footing.","revision_made":"yes","referee_comment":"[Results] Results section on agentic collapse: The claim that 'for the most collapse-prone model, blocking tool access entirely barely changes its score' is presented as evidence that tools stop paying off where needed most. However, the manuscript lacks controls for prompt sensitivity, baseline performance without tools, or statistical tests on the score difference. This detail is required to support the collapse interpretation as load-bearing for the agentic-setting contribution."}],"tokens_in":1572,"tokens_out":523,"duration_ms":17994,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is a benchmark of 12,553 biomedical questions positioned as unsolved, used to test agentic models on citation verification and tool use without fixed answer keys. It reports 15.9% wrong-citation links even when URLs resolve, low solve rates on hard subsets (17% for held-out models, 29-60% for frontiers), and agentic collapse where models stop calling tools on the toughest items.\n\nThe paper does a few things cleanly. It moves away from answer-key benchmarks that let models reproduce sources instead of checking support. Anchoring difficulty on actual failures by three open-weight models is empirical rather than subjective. The per-question checklist lifting inter-judge Spearman from 0.35 to 0.82 is a practical detail worth noting. The spread across capability tiers and the non-saturating nature (best agents still leave 33-40% unsolved) are the clearest new signals.\n\nThe soft spot is the openness verification. The abstract states it was checked against real follow-up evidence, but without details on search scope, databases, time windows, or keyword strategy, it is possible some questions have been resolved in recent preprints or narrower venues. If that fraction is non-trivial, the reported solve rates and collapse patterns partly measure retrieval success instead of open-question handling. That assumption is load-bearing for the central claims.\n\nNo obvious circularity or invented entities in the reported numbers. The work is aimed at groups building or evaluating research agents in biomedicine. A reader focused on agent faithfulness or scientific workflows would find the failure modes useful to see. It has enough empirical grounding and new observations to merit referee time, mainly to examine the dataset construction and verification procedures in full.","headline":"OpenBioRQ gives a non-saturating benchmark on open biomedical questions with concrete observations on wrong citations and tool collapse, but the unsolved status needs stronger exhaustiveness checks.","tokens_in":2449,"tokens_out":427,"would_cite":false,"duration_ms":19676,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"OpenBioRQ presents 12,553 unsolved biomedical questions as a test for whether agentic models can verify sources without answer keys.","keywords":["unsolved biomedical questions","agentic benchmarks","citation verification","tool-use collapse","retrieval faithfulness","abstention evaluation","open research questions"],"falsifier":"A new agentic system that solves more than 70 percent of the hardest subset while continuing to issue tool calls on those same items would contradict the reported performance ceiling and collapse pattern.","tokens_in":2740,"feed_emoji":"🔬","tokens_out":709,"duration_ms":20801,"temperature":0.7,"pith_summary":"The paper introduces OpenBioRQ, a benchmark of open biomedical research questions across twelve domains that requires models to make multiple tool calls to check whether cited sources actually support claims. Difficulty is set empirically by selecting questions that three open-weight models already fail, and openness is confirmed by checking against later published evidence rather than model memory. On the hardest slice of the benchmark, models from the same families as the difficulty anchors reach only about 17 percent success, while three frontier agents range from 29 to 60 percent, leaving roughly a third of questions unsolved. The work also records that agents often stop issuing tool calls on these difficult items, so that denying tool access changes scores little for the worst-affected model. Evaluation reliability is raised by supplying judges with a fixed per-question checklist.","feed_headline":"Unsolved biomedical questions benchmark agents at 17-60%","feed_subtitle":"OpenBioRQ of 12,553 open queries shows tool calls stop on the hardest items and frontier models leave a third unsolved","key_machinery":"OpenBioRQ, a collection of unsolved biomedical questions that forces multiple tool calls for citation verification and has no fixed answer key, with difficulty defined by failure of three reference open-weight models.","core_discovery":"OpenBioRQ is a retrieval-grounded agentic benchmark of 12,553 unsolved biomedical research questions that treats open questions as a faithfulness-and-abstention probe; openness is verified against real follow-up evidence, difficulty is anchored on items failed by three open-weight reference models, held-out models from the same lineage solve only about 17 percent of the hardest subset, and three frontier agents span 29-60 percent while showing agentic collapse where tool use stops.","pith_inferences":["The same unsolved-question design could be applied in chemistry or physics to measure retrieval faithfulness outside biomedicine.","Persistent tool-use collapse suggests that current training objectives may reward parametric recall more than sustained verification on open problems.","Future agent training could explicitly penalize early abandonment of tool sequences on items that reference models already miss."],"forward_implications":["Frontier agents leave 33-40 percent of the hardest questions unsolved even when tools are available.","On the collapse-prone model, removing tool access changes the score by only a small amount.","A static per-question checklist lifts inter-judge Spearman correlation from 0.35 to 0.82.","The benchmark remains non-saturating across current capability tiers."],"fun_headline_variants":["Agents solve 17% of hardest unsolved biomedical questions","Frontier agents solve 29-60% of 12553 bio questions","Agents collapse tool use on toughest open bio problems","Biomedical benchmark finds agents at 17-60% solve rate"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The selected questions are genuinely open and their difficulty is correctly measured by the failure of the three reference models rather than by any model’s internal knowledge.","fun_headline_variants_meta":{"raw":{"variants":["Agents solve 17% of hardest unsolved biomedical questions","Frontier agents solve 29-60% of 12553 bio questions","Agents collapse tool use on toughest open bio problems","Biomedical benchmark finds agents at 17-60% solve rate"]},"model":"grok-4.3","cost_usd":0.007354,"raw_usage":{"total_tokens":3375,"prompt_tokens":813,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":73540500,"prompt_tokens_details":{"text_tokens":813,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2493,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":813,"tokens_out":69,"duration_ms":19859,"temperature":1.0,"reasoning_tokens":2493,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T12:05:34.735845+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new agentic system that solves more than 70 percent of the hardest subset while continuing to issue tool calls on those same items would contradict the reported performance ceiling and collapse pattern.","supporting_citations":[],"review_version":1}