{"id":"3b3e7d19-8c2b-4d00-97fe-8d557b734587","arxiv_id":"2607.10275","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs systematically under-request critical molecular and cytogenetic data in multi-round oncology, capping accuracy at 68% despite high knowledge scores and coherent reasoning traces.","lead":"Frontier LLMs reach at most 68% accuracy on complex hematologic oncology cases when they must actively request data across three sequential rounds. The dominant failure is collapsing information-seeking (utilization falls to 26% in the final round), not missing medical knowledge.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Utilization collapse (57%→26%) may largely reflect context-length / multi-turn degradation induced by the agentic design rather than pure search-satisficing.","rationale":"The reader’s weakest_assumption already isolates exactly this confound (context-length pressure, prompt design, partial-credit subjectivity) as the soft underbelly of the claim that the benchmark faithfully isolates information-seeking. The utilization–accuracy correlation (R=0.69) and high R-IDEA / low accuracy gap remain solid empirical observations inside the tested regime; the paper’s own discussion of long-context and multi-turn degradation shows the authors are aware of the issue. No internal inconsistency or circularity appears, and the promised public release of cases/code will allow independent verification. Therefore the CONDITIONAL verdict (accept once context contribution is quantified and release verified) needs no adjustment. The concrete ablation above is the single cleanest check that would settle whether the load-bearing interpretation of the collapse holds.","tokens_in":51596,"tokens_out":550,"duration_ms":20568,"concrete_test":"Re-evaluate the top 5–8 models on the same 20 cases under a matched-context ablation: at the start of R3 supply a concise summary of prior rounds plus the full list of still-available items (so R3 prompt length ≈ R1 length) while preserving the single-request-per-turn rule. If R3 utilization recovers toward the 50–60% range and composite accuracy rises >10 points, the collapse is largely context-induced; if utilization stays near 26%, the pure information-seeking claim is strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the primary limitation is systematic information-seeking failure (search satisficing, anchoring, premature closure) rather than knowledge—rests on interpreting the sharp R3 utilization drop and low request counts as cognitive analogues of novice dual-process errors. By R3 the accumulated history of requests, releases, prior diagnoses and traces routinely exceeds 10k tokens; the paper itself cites “lost in the middle” positional bias and multi-turn accuracy decline as compounding factors. Without a control that holds effective context length fixed (summarized history, item-list access without trajectory accumulation, or single-turn equivalents), it remains ambiguous how much of the collapse is intrinsic under-seeking versus an artifact of the very sequential design used to measure it. This directly undercuts both the strength of the human-novice analogy and the assertion that the barrier is qualitative and unlikely to be resolved by scaling alone.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces OncoRounds, an agentic benchmark in which 32 frontier LLMs must proactively request clinical data one item per turn across three staged rounds of increasing complexity before committing to diagnosis, differential, and treatment on 20 expert-authored complex hematologic oncology cases. Best overall accuracy is 68.1% (Claude Opus 4.6); information utilization is the strongest predictor of accuracy (R = 0.69) yet collapses from ~57% in R1/R2 to 26% in R3, leaving molecular/cytogenetic data unexamined. R-IDEA reasoning scores are high (91% ≥6) but only moderately correlated with accuracy; error taxonomy (search satisficing, anchoring, premature closure) mirrors novice dual-process biases. The central claim is that the primary limitation is systematic information-seeking failure under uncertainty rather than insufficient medical knowledge.","tokens_in":51782,"tokens_out":1338,"duration_ms":26731,"significance":"If the result holds, this is a high-value contribution: it cleanly separates information-seeking from interpretation in a clinically realistic sequential setting, scales the evaluation to 32 models with a validated LLM-as-judge (quadratic weighted κ = 0.804 vs. a board-certified hematologist), and supplies trajectory, utilization, and error analyses that static knowledge benchmarks cannot. The public release of cases, pipeline, and outputs, plus the explicit dual-process framing, would push the field beyond MedQA-style passive vignettes toward agentic clinical evaluation and training. The observed utilization–accuracy link and R3 collapse are actionable for architecture and data-curation work.","major_comments":[{"comment":"Results (utilization rates; Supplementary Figure 1) and Discussion: the sharp R3 utilization drop (57% → 26%) and low request counts are load-bearing for the claim of ‘systematic failure of information-seeking’ and the dual-process analogy. The paper itself notes that by R3 the interaction history routinely exceeds 10 k tokens and cites ‘lost in the middle’ and multi-turn degradation. Without a control that holds effective context length fixed (e.g., summarized history, item-list access without trajectory accumulation, or single-turn equivalents of the same information), it remains ambiguous how much of the collapse is intrinsic search satisficing versus an artifact of the sequential design used to measure it. This directly weakens both the human-novice analogy and the assertion that the barrier is qualitative and unlikely to be resolved by scaling alone. A control experiment or a substa","section":"Results / Discussion (utilization collapse; context-length paragraph)"},{"comment":"Discussion and Error Taxonomy: the claim that models exhibit ‘the same cognitive biases that characterize novice clinicians under dual-process models’ rests on literature analogy and LLM-annotated traces, not a direct human baseline on OncoRounds. Without at least a small cohort of hematologists (or trainees) run under identical single-request-per-turn, three-round constraints, the convergence remains interpretive rather than demonstrated. This is load-bearing for the strongest framing of the paper.","section":"Discussion / Error Taxonomy / Methods (Cognitive Error Taxonomy)"},{"comment":"Methods (Automated Evaluation Pipeline) and Limitations: differential-reasoning accuracy is the weakest component for all 32 models (mean gap 23 pp) and is also acknowledged as the most subjective. The majority-vote GPT-5 Mini ensemble (κ = 0.804) is systematically stricter than the single expert, and partial-credit boundaries for differentials and treatment priorities are inherently less standardized than primary diagnosis. While the authors flag this, the ranking of ‘differential reasoning as the primary challenge’ is used to support the information-seeking narrative; a sensitivity analysis with human re-scoring of a larger differential subset, or explicit down-weighting of that component in the composite, would strengthen the claim.","section":"Methods (Scoring) / Results (Overall Performance) / Limitations"}],"minor_comments":[{"comment":"Table 1 and case development: the 20 cases are all relapsed/refractory/post-transplant; a short note on how performance might differ on de-novo or lower-complexity presentations would help readers gauge generalizability beyond the Limitations paragraph.","section":"Table 1 / Case Development / Limitations"},{"comment":"Figure 2 and Supplementary Figure 5: the alluvial and donut visualizations of anchoring are informative, but the distinction between ‘subtype refinement failure’ and true category shift could be made more visually explicit (e.g., color coding of partial vs. full incorrect).","section":"Figure 2 / Supplementary Figure 5"},{"comment":"Methods (Clinical Reasoning Evaluation): R-IDEA was scored by Claude Opus 4.6; a brief statement on whether the same model family was used for any of the evaluated candidates, and any decontamination steps, would be useful.","section":"Methods (Clinical Reasoning Evaluation)"},{"comment":"Supplementary Appendix: the 49 curated reasoning-trace excerpts are valuable; ensuring the selection script and full set of 1 008 scored entries are released with the code would improve reproducibility.","section":"Supplementary Appendix / Code Availability"},{"comment":"Abstract and Introduction: the phrase ‘the best achieved only 68% overall accuracy’ is accurate but could briefly note the composite definition (balanced accuracy of diagnosis / differential / treatment) so readers do not equate it with pure diagnostic accuracy.","section":"Abstract / Results (Overall Performance)"}],"recommendation":"major_revision","confidential_remarks":"The work is carefully executed and the confound of context length is the single most important issue for the central claim; if the authors can add even a modest control or reframe the claim more cautiously, the paper would be a strong candidate for a high-visibility venue (Nature Medicine / NEJM AI / similar). Scope is appropriate for a medical-AI or clinical-reasoning journal; pure cs.AI venues may find the clinical depth high and the architectural novelty moderate. No concerns about citation patterns or undisclosed conflicts beyond those already declared."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is clean: when 32 frontier models must actively request data across three staged hematology rounds, the best hits only 68%, utilization is the strongest accuracy predictor (R=0.69), and it collapses from ~57% to 26% in the molecular round. High R-IDEA traces (91% above threshold) still decorrelate from accuracy. That is the new, usable finding.\n\nWhat they did well is the design. Twenty de-novo complex cases (relapsed/refractory/post-transplant), expert-authored references reviewed by multiple hematologists, single-request-per-turn prioritization, three clinically timed rounds, majority-vote GPT-5-Mini judge calibrated to a board-certified hematologist (quadratic weighted κ=0.804), trajectory analysis showing subtype-refinement anchoring rather than category error, and an explicit Croskerry-style error taxonomy. The supplementary traces make the failure modes inspectable. Code and cases are promised on GitHub; if they ship, this is reproducible.\n\nThe soft spot the stress-test flags is real and the authors already half-admit it: by R3 the interaction history routinely exceeds 10k tokens, and they cite lost-in-the-middle and multi-turn degradation. Without a fixed-context or summarized-history control, you cannot cleanly separate pure search-satisficing from context pressure. That weakens the claim that the barrier is purely qualitative and scaling-resistant, and it softens the human-novice analogy. It does not, however, invent the under-requesting or the accuracy correlation; those are measured. Scope is narrow (one subspecialty, hard cases, synthetic, LLM judge) and they say so.\n\nThis is for people building or evaluating medical agents, not for pure knowledge-benchmark consumers. The math is ordinary correlations and non-parametrics; the data and citation pattern look solid. I would bring it to reading group, cite the utilization and trajectory results, and send it to peer review. Ask the authors for the context-length control or at least a quantitative discussion of how much of the R3 drop survives history compression.","headline":"Solid agentic oncology benchmark showing utilization collapse and novice-like biases; the context-length confound is real but does not erase the result.","tokens_in":52447,"tokens_out":520,"would_cite":true,"duration_ms":7639,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Frontier medical AI fails not for lack of knowledge but because it stops seeking the data that treatment decisions require.","keywords":["large language models","clinical reasoning","information-seeking","agentic evaluation","hematologic oncology","cognitive biases","search satisficing","diagnostic anchoring"],"falsifier":"Re-run the same models with forced higher utilization of Round-3 molecular and cytogenetic items (or with multi-item requests and shorter contexts) and check whether overall accuracy rises substantially and the search-satisficing / anchoring error rates fall; if utilization rises but accuracy and error modes stay the same, the information-seeking account is weakened.","tokens_in":52457,"feed_emoji":"🩺","tokens_out":962,"duration_ms":19529,"temperature":0.7,"pith_summary":"Large language models score highly on medical exams that hand them complete cases, yet real clinical reasoning means deciding what to investigate when information is incomplete and arrives over time. This paper builds an interactive hematologic-oncology benchmark in which 32 frontier models must request one clinical data item per turn across three staged rounds before committing to a diagnosis and treatment plan. The best model reaches only 68 percent overall accuracy. The strongest predictor of success is how much of the available data a model actually requests, yet that utilization rate collapses from about 57 percent in early rounds to 26 percent in the final round, leaving molecular and cytogenetic results unexamined. Reasoning traces look locally coherent on a clinical rubric, but they largely fail to track correct conclusions. The dominant error patterns—search satisficing, anchoring, and premature closure—are the same cognitive biases that characterize novice human clinicians. The authors conclude that the binding constraint is information-seeking under uncertainty, not encoded medical knowledge.","feed_headline":"Medical AIs stop seeking the data that decide treatment","feed_subtitle":"Best of 32 models hits 68%; utilization falls to 26% when molecular results matter most","key_machinery":"OncoRounds: a three-round agentic evaluation in which models must request one clinical information item per turn before solving; performance is scored on diagnosis, differential reasoning, and treatment, with information utilization (fraction of available items requested) as the central behavioral metric.","core_discovery":"Across 32 frontier models and 20 complex hematologic-oncology cases, the primary performance limit is not medical knowledge but a systematic failure of proactive information-seeking under uncertainty: models under-request critical later-round data, and the resulting failure modes match those of novice clinicians under dual-process diagnostic theory.","pith_inferences":["If the utilization collapse is driven mainly by long multi-turn context and positional attention bias, shorter-memory or retrieval-augmented agent designs may fix more of the deficit than further medical pre-training.","The knowledge-versus-seeking gap is likely to appear in other sequential high-stakes domains (emergency triage, finance under incomplete markets) that currently rely on complete-case benchmarks.","Differential-reasoning scores being the weakest component for every model suggests that reward models and judges may need explicit penalties for premature commitment rather than only accuracy on the final label.","A natural next experiment is to train or fine-tune models on trajectories that reward requesting the highest-value unobtained item before solving, then re-measure Round-3 utilization and recovery from early anchoring."],"forward_implications":["Static complete-vignette medical benchmarks will systematically overestimate clinical readiness of current models.","Agentic evaluation that forces active information gathering should become a standard component of medical AI assessment.","Training and scaffolding that explicitly practice sequential hypothesis testing under incomplete data are needed; scale alone is unlikely to close the gap.","High-quality-looking reasoning traces can coexist with wrong conclusions, so they should not be trusted as automatic signals of reliability for clinicians.","The same novice-like biases (search satisficing, anchoring, premature closure) appear in systems that share no developmental path with human trainees."],"fun_headline_variants":["LLMs quit seeking data that decide oncology treatment","Models under-request key molecular results late in diagnosis","Info-seeking collapse leaves critical cancer data unexamined","AI mimics novice biases: search satisficing and premature closure","Best of 32 models at 68% as utilization falls when it matters"],"cache_read_input_tokens":49280,"weakest_assumption_plain":"The claim rests on twenty expert-written complex blood-cancer cases, a one-request-per-turn rule, and an automated judge agreeing substantially with one hematologist being a fair stand-in for real clinical information-seeking rather than an artifact of long context, prompt design, or partial-credit scoring.","fun_headline_variants_meta":{"raw":{"variants":["LLMs quit seeking data that decide oncology treatment","Models under-request key molecular results late in diagnosis","Info-seeking collapse leaves critical cancer data unexamined","AI mimics novice biases: search satisficing and premature closure","Best of 32 models at 68% as utilization falls when it matters"]},"model":"grok-4.5","effort":"low","cost_usd":0.003798,"raw_usage":{"total_tokens":1159,"prompt_tokens":741,"num_sources_used":0,"completion_tokens":65,"cost_in_usd_ticks":37980000,"prompt_tokens_details":{"text_tokens":741,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":353,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":741,"tokens_out":65,"duration_ms":3936,"temperature":1.0,"reasoning_tokens":353,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T12:58:23.829619+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same models with forced higher utilization of Round-3 molecular and cytogenetic items (or with multi-item requests and shorter contexts) and check whether overall accuracy rises substantially and the search-satisficing / anchoring error rates fall; if utilization rises but accuracy and error modes stay the same, the information-seeking account is weakened.","supporting_citations":[],"review_version":1}