{"id":"5a484ebc-bf7f-4f50-85ce-cbc83f47661a","arxiv_id":"2508.14880","paper_version":3,"verdict":"REJECT","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"The abstract claims a new state-of-the-art medical research agent, but the body is an unrelated RL-scaling paper, leaving the central claim unsupported.","lead":"The abstract describes a medical deep-research agent, MedResearcher-R1-32B, claiming state-of-the-art medical benchmark results, but the manuscript body is an entirely different paper on reinforcement-learning scaling laws. No model, training, or evaluation content for the advertised medical agent appears anywhere in the full text, so the central claim is unverifiable.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The manuscript body is an unrelated value-based RL scaling paper (arXiv:2508.14881), so the abstract's SOTA medical-agent claims have no supporting model, training, or evaluation in the submission.","rationale":"Reading the submission in good faith, the paper's advertised central claim is the medical deep-research agent MedResearcher-R1-32B and its SOTA benchmark performance. What would have to be true for that claim to stand is that the manuscript actually contains the described system, its training pipeline, and its evaluations. The supplied full text is a different paper—an RL scaling study—so the central claim is unsupported within the submission itself. This is not a disagreement about method plausibility or benchmark choice; it is a structural absence: no model, no retrieval engine, no 2100+ trajectories, and no medical benchmark results appear in the body. I agree with the reader's weakest-assumption identification: the mismatch is the load-bearing concern, and the independently sketched knowledge-graph-to-QA assumption is secondary. The recommended verdict is unchanged from the reader's REJECT, at low confidence, because the appropriate resolution is to reject the submission as not containing the claimed research rather than to assess the merits of the embedded RL-scaling study.","tokens_in":25068,"tokens_out":3425,"duration_ms":39193,"concrete_test":"Download the actual PDF/file of arXiv:2508.14880 as submitted, extract the complete text, and search for the artifacts the abstract requires: 'knowledge graph', 'trajectory', 'retrieval engine', 'MedResearcher', 'SFT', 'medical benchmark', and any evaluation table with model names. If the body is the RL-scaling manuscript (Sections 1–9, Appendices A–D, Eqs. (3.1), (6.1), (7.1)) and none of these artifacts appear, the abstract's SOTA claim has no supporting derivation and the verdict should remain REJECT. If, hypothetically, the medical content is present in a part of the submission not shown here, that content must include benchmark tables and training details before the claim can be reassessed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that MedResearcher-R1-32B, trained on 2100+ knowledge-graph-derived trajectories with a custom medical retrieval engine, achieves new SOTA on medical benchmarks and remains competitive on general deep-research tasks. For this claim to be supported, the paper must contain at least (i) the trajectory synthesis method, (ii) the retrieval engine and tool integration, (iii) the SFT/RL training setup, and (iv) quantitative benchmark comparisons. The submitted full text is instead 'Compute-Optimal Scaling for Value-Based Deep RL' (arXiv:2508.14881): Sections 1–9 and Appendices A–D formalize compute allocation for TD-learning, analyze TD-overfitting, and fit batch-size/UTD/model-size laws (e.g., Eqs. 3.1, 6.1, 7.1). There is no medical knowledge graph, no 'longest chains' extraction, no retrieval engine, no 2100+ trajectories, no MedResearcher-R1-32B training run, and no medical benchmark table. The abstract's performance claim is therefore a conclusion without derivation in the received artifact. Independently of the mismatch, the sketched method assumes QA pairs extracted as longest subgraph chains around rare entities are a faithful, non-contaminating proxy for expert clinical reasoning; that equivalence is not argued or validated anywhere in the text. The first issue is decisive: the document as submitted does not contain the research it advertises.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission arXiv:2508.14880 presents an abstract for \"MedResearcher-R1\", a 32B medical deep-research agent trained on 2100+ knowledge-graph-derived trajectories with a custom medical retrieval engine, and claims new state-of-the-art results on medical benchmarks while remaining competitive on general deep research tasks. The body of the manuscript, however, is a complete and unrelated paper titled \"Compute-Optimal Scaling for Value-Based Deep RL\" (arXiv:2508.14881), which studies batch size, UTD ratio, model size, and TD-overfitting in off-policy reinforcement learning. None of the components named in the abstract—medical knowledge graph, trajectory synthesis, retrieval engine, supervised fine-tuning or online RL with composite rewards, or medical benchmark evaluation—appear anywhere in the received full text. The manuscript as submitted therefore does not contain the research it advertises.","tokens_in":25354,"tokens_out":2740,"duration_ms":34402,"significance":"If the MedResearcher-R1 claims were supported by the submitted content, this would be a significant result: an open-weight 32B model reportedly outperforming much larger proprietary systems on complex medical QA while maintaining general deep-research performance would be of high interest to the medical NLP and agent communities. The high-level method (knowledge-graph-derived multi-hop QA trajectories, a private retrieval engine, and a two-stage SFT/RL training paradigm) is plausible and worth investigating. However, the received manuscript contains no model card, no training details, no retrieval-engine description, no trajectory-synthesis algorithm, and no evaluation tables or comparisons. The claimed significance is therefore entirely unsupported within the artifact under review, and the paper cannot currently be assessed as a research contribution.","major_comments":[{"comment":"The central claim—state-of-the-art medical benchmark results for MedResearcher-R1-32B—has no supporting content in the received manuscript. The full text, Sections 1–9 and Appendices A–D, is a different paper on compute-optimal scaling for value-based deep RL (arXiv:2508.14881). There is no mention of MedResearcher, medical knowledge graphs, 2100+ trajectories, the custom retrieval engine, the two-stage SFT/RL training, or any medical benchmark. The only occurrence of these elements is in the abstract. This is a load-bearing absence: there is nothing to verify or falsify. The submission, as is, cannot be reviewed as a medical-agent paper.","section":"Abstract vs. full text (entire manuscript)"},{"comment":"Independently of the manuscript mismatch, the sketched method assumes that QA pairs extracted as the longest chains from subgraphs around rare medical entities are a faithful, non-contaminating proxy for expert clinical reasoning. The abstract provides no argument or validation for this equivalence. This matters because the same model is then evaluated on medical benchmarks, and training on mined knowledge-graph chains may overlap with benchmark content. No evidence is presented in the received full text that such contamination is controlled or that the synthesized trajectories reflect genuine clinical reasoning rather than graph-structural artifacts.","section":"Abstract, method sketch"},{"comment":"The abstract claims \"exceptional performance\" and \"new state-of-the-art results on medical benchmarks,\" but the manuscript contains no quantitative evaluation supporting these claims. The only Tables (1–2) and Figures (1–22) in the received body concern DMC/HumanoidBench control tasks, batch sizes, UTD ratios, and TD-error fits. No medical benchmark names, baseline comparisons, error bars, or model releases are present. For a SOTA claim, this evaluation is essential and entirely missing here.","section":"Evaluation claims (Tables/Figures)"}],"minor_comments":[{"comment":"The arXiv listing is categorized as cs.CL and titled \"MedResearcher-R1\", but the full text is a cs.LG RL-scaling paper with a different title and author list. This is a severe presentation mismatch and should be corrected by the authors.","section":"Metadata/title"},{"comment":"The reference list in the full text contains no medical-agent, medical-QA, or retrieval-system references, which is inconsistent with the abstract's framing and prior-work claims.","section":"References"},{"comment":"No code, model weights, dataset, or benchmark harness for MedResearcher-R1 is provided or referenced. Even the abstract's method description is too brief to be reproducible (e.g., \"longest chains\" and \"composite rewards\" are underspecified).","section":"Reproducibility"}],"recommendation":"reject","confidential_remarks":"This appears to be a manuscript-versioning or submission error: the abstract and title correspond to one paper while the body is an entirely different arXiv paper (2508.14881). The correct course is to return the submission to the authors so the intended MedResearcher-R1 manuscript can be submitted. If the correct full text is provided, the review should start fresh, because the currently submitted artifact cannot support any of the abstract's claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: as received, this paper doesn't hold together. The abstract advertises MedResearcher-R1-32B, a medical deep-research agent trained on knowledge-graph-derived trajectories with a custom retrieval engine, and claims new SOTA on medical benchmarks. The full text is 'Compute-Optimal Scaling for Value-Based Deep RL'—a separate paper about TD-learning, batch sizes, UTD ratios, and model size. None of the claimed medical content appears anywhere in the manuscript: no knowledge-graph extraction, no retrieval engine, no 2100+ trajectories, no training setup, no benchmark tables, no comparisons. The abstract's central claim is therefore unsupported in the artifact I was given. This is not a minor editorial slip; it makes the submission impossible to evaluate as a medical-agent paper.\n\nGive credit where it's due: the RL scaling content that actually occupies the body is a competent, workmanlike extension of Rybkin et al. (2025). The TD-overfitting phenomenon—small models harming validation TD-error with larger batch sizes, larger models tolerating bigger batches—is well-motivated and backed by real experiments, including a passive-critic diagnostic. The batch-size/UTD/model-size fits and the compute-allocation laws are presented carefully, with appendices and sensitivity analysis. If this were submitted under its own title and authors, it would be a plausible candidate for peer review in an RL venue.\n\nThe soft spots beyond the mismatch: the abstract's sketched method carries real epistemological risk. Mining longest subgraph chains around rare medical entities and calling them 'expert-level' reasoning assumes those chains are a faithful, non-contaminating proxy for clinical reasoning. That equivalence is not argued, let alone validated, anywhere. And generating training trajectories with your own pipeline plus your own retrieval engine makes benchmark contamination a live concern. But honestly, these are secondary because the content to audit them is absent. The mismatch is decisive.\n\nWho is this for? As-is, nobody. A reader after medical deep research gets nothing; a reader after RL scaling laws gets a paper that isn't the one advertised. The right move is to desk reject and tell the authors to resubmit the actual medical work if it exists, and to submit the RL scaling study separately under its correct identity.\n\nRecommendation: do not send to peer review in this form.","headline":"The submission is two different papers: the abstract promises a medical deep-research agent with SOTA results, the body is a value-based RL scaling study—nothing in the body supports the abstract.","tokens_in":25913,"tokens_out":4168,"would_cite":false,"duration_ms":42595,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 32-billion-parameter medical research agent claims to beat much larger proprietary systems on medical benchmarks — but the manuscript body describes a different paper on reinforcement-learning scaling.","keywords":["medical deep research","LLM agents","knowledge graph","trajectory synthesis","multi-hop question answering","medical retrieval engine","reinforcement learning","model scaling"],"falsifier":"Read the full text and look for any medical dataset, model training run, retrieval-engine evaluation, or medical benchmark result; the body contains none of these and instead reports scaling laws for value-based deep reinforcement learning, so the absence is directly checkable and settles whether the claimed experiments exist.","tokens_in":24932,"feed_emoji":"🏥","tokens_out":3868,"duration_ms":40110,"temperature":0.7,"pith_summary":"The abstract of this submission argues that a 32-billion-parameter open-weight medical deep-research agent, MedResearcher-R1-32B, can outperform much larger proprietary systems on complex medical question-answering benchmarks while staying competitive on general deep-research tasks. The claimed route is a knowledge-informed trajectory synthesis framework that mines multi-hop question-answer pairs from medical knowledge graphs, combined with a custom private medical retrieval engine and a two-stage training recipe (supervised fine-tuning followed by online reinforcement learning). A sympathetic reader would care because it would mean that targeted domain data and tool design, rather than sheer model scale, could be the deciding factors for expert-level medical AI. However, the full text supplied is not the medical paper: it is a compute-optimal scaling study for value-based deep reinforcement learning, so none of the claimed medical model, trajectories, retrieval engine, or benchmark results are present in this submission.","feed_headline":"32B medical research agent claims to beat far larger rivals","feed_subtitle":"Abstract promises SOTA medical benchmarks via knowledge-graph training, but the body contains no medical experiments.","key_machinery":"The load-bearing mechanism, as described in the abstract, is the knowledge-informed trajectory synthesis framework: it takes a medical knowledge graph, locates subgraphs around rare medical entities, extracts the longest chains from those subgraphs, and converts them into multi-hop question-answer pairs that are meant to force compositional clinical reasoning. This synthesized data is then used in a two-stage training paradigm (supervised fine-tuning followed by online reinforcement learning with composite rewards), aided by a custom private medical retrieval engine that supplies domain-specific evidence alongside general-purpose tools. The knowledge-graph chains are what supposedly supply d","core_discovery":"On its own terms, the paper's claim is that strategic domain-specific innovations can let a smaller open-source model beat much larger proprietary systems in medicine. Specifically, the authors say MedResearcher-R1-32B was trained on 2,100+ diverse trajectories across 12 medical specialties, each averaging 4.2 tool interactions, generated by extracting the longest chains from knowledge-graph subgraphs around rare medical entities to create complex multi-hop clinical QA pairs. These trajectories feed a two-stage training process — supervised fine-tuning plus online reinforcement learning with composite rewards — and the model is paired with a custom-built private medical retrieval engine. The","pith_inferences":["The submission's full text is a different paper about compute-optimal scaling for value-based deep RL; therefore, as submitted, the medical model, its training data, its retrieval engine, and its benchmark results are not present in the manuscript to be examined, and the abstract's claims are currently unverifiable from the body.","Independently of the mismatch, the method as sketched assumes that QA pairs mined from the longest chains of knowledge-graph subgraphs around rare medical entities are a faithful, non-contaminating proxy for expert clinical reasoning; this equivalence is asserted in the abstract but not validated anywhere in the submitted text.","If one wanted to test the underlying idea rather than this submission, a natural extension would be to generate trajectories from knowledge graphs in other expert domains and measure whether the multi-hop chain length correlates with benchmark difficulty and with human expert agreement.","The two-innovation recipe (graph-derived trajectories plus a domain-specific retrieval engine) could be evaluated piecewise: ablating the retrieval engine from the final model would reveal how much of the claimed gain comes from tool design versus training-data synthesis.",""],"forward_implications":["If the claim holds, a 32B open-weight model could outperform much larger proprietary systems on medical benchmarks, making expert-level medical QA more accessible and auditable than closed systems.","Domain-specific trajectory synthesis from knowledge graphs could reduce reliance on expensive expert-written training data, since the QA pairs are mined from graph structure rather than authored by clinicians.","A specialized private retrieval engine could materially raise accuracy on medical questions where general web retrieval returns noisy or non-authoritative sources.","The two-stage recipe (supervised fine-tuning plus online RL with composite rewards) could be a transferable template for building expert agents in other knowledge-dense domains such as law, chemistry, or engineering.","Smaller specialized agents trained this way might stay competitive on general deep-research tasks rather than losing general capability, as the abstract claims for MedResearcher-R1-32B.",""],"supporting_citations":[],"fun_headline_variants":["32B MedResearcher-R1 claims SOTA on medical tasks, but no experiments","MedResearcher-R1 paper claims SOTA but lacks medical eval","32B open model claims medical SOTA, no experiments in paper","Knowledge-graph trained 32B claims medical SOTA, paper lacks tests","Small 32B medical agent challenges giants—but no experiments"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The submission's body is a different paper about reinforcement-learning scaling, so the claimed medical model, its training data, and its benchmark results are not present to be checked; even if they were, the method presumes that questions mined from knowledge-graph chains are a faithful, uncontaminated stand-in for expert clinical reasoning.","fun_headline_variants_meta":{"raw":{"variants":["32B MedResearcher-R1 claims SOTA on medical tasks, but no experiments","MedResearcher-R1 paper claims SOTA but lacks medical eval","32B open model claims medical SOTA, no experiments in paper","Knowledge-graph trained 32B claims medical SOTA, paper lacks tests","Small 32B medical agent challenges giants—but no experiments"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000806,"raw_usage":{"total_tokens":3394,"prompt_tokens":783,"completion_tokens":2611,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":2515}},"tokens_in":527,"tokens_out":2611,"duration_ms":21661,"temperature":1.0,"reasoning_tokens":2515,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:13:33.179327+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Read the full text and look for any medical dataset, model training run, retrieval-engine evaluation, or medical benchmark result; the body contains none of these and instead reports scaling laws for value-based deep reinforcement learning, so the absence is directly checkable and settles whether the claimed experiments exist.","supporting_citations":[],"review_version":1}