{"id":"4bbfb373-ddf0-49c9-b28e-04e22d031ae3","arxiv_id":"2505.04651","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A survey of LLM-based hypothesis generation and validation whose taxonomy is useful in outline but whose citations and tool descriptions are unreliable.","lead":"This paper is a survey of how large language models are used to generate and validate scientific hypotheses, organizing methods, datasets, and challenges. It is a broad literature review rather than a new experiment, and the field map is only as reliable as the citations it assembles.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey's factual core is unreliable: named tools such as SciFact, CausalNet, and InterveneAI are misdescribed or lack supporting references; the map cannot be trusted.","rationale":"The reader's verdict is REJECT with moderate confidence, based on the unreliability of the survey's factual core. My independent read of the full text finds concrete, checkable instances of this problem rather than mere suspicion. Section 4.4's description of SciFact as a drug-repurposing text-mining tool is directly contradicted by the canonical SciFact fact-verification benchmark. Section 4.7's three named causal tools are attached to textbook references that cannot support them. Section 5.4 names three tools without matching citations. This is the kind of error that matters for a survey: the paper's stated contribution is to be a structured overview and roadmap (§1.3), and the taxonomy is only useful if readers can trust the entries. A few isolated typos would not justify rejection; a systemic pattern of unverifiable or misattributed systems does. The paper contains no machine-checked proofs or reproducible code, but for a survey that absence is not the main issue; the main issue is that its central contribution cannot be relied upon. I recommend keeping the reader's REJECT. If the proposed audit were run and fewer than three entries failed, the verdict could be reconsidered toward CONDITIONAL acceptance; the current text provides no such basis.","tokens_in":38674,"tokens_out":3976,"duration_ms":39927,"concrete_test":"Run a citation-verification audit for ten named tools or datasets that carry references: SciFact, CausalNet, BayesCausality, InterveneAI, CrossValNet, InterdisciplinaryTest, TransferTest, UpdateKG, SciGraph, KG-Stream, and CSKG-600. For each, retrieve the cited paper and search its full text plus ACL Anthology, arXiv, and DOI records for the exact tool or dataset name and the described function. Pre-register the criterion: if the cited source does not mention the tool or dataset, or if no independent record exists, count it as 'unverified'. If at least three of the ten are unverified or demonstrably misdescribed, the survey's factual core is unreliable and the reader's REJECT verdict is sustained.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central claim is to be a structured, reliable map of LLM-driven hypothesis generation and validation (§1.3). That claim stands or falls on whether the named systems and datasets exist and are characterized correctly. This is not an internal-consistency issue but a factual-support issue, and several central examples fail this test. In §4.4, SciFact is described as a text-mining tool that 'leverages co-occurrence patterns in scientific literature to propose new drug-disease relationships'; the canonical SciFact (Wadden et al., ACL 2020) is a fact-verification benchmark of scientific claims, with no such discovery function. In §4.7, CausalNet, BayesCausality, and InterveneAI are attributed to Peters et al. [2017], Lucas [2007], and Neuberg [2003] respectively; those are general causal-inference textbooks or reviews, and the cited references do not establish the existence of these tools. Section 5.4 names CrossValNet, InterdisciplinaryTest, and TransferTest with citations that do not contain those systems. Tables 2–3 list CSKG-600 and AHTech as 'introduced' resources, while the abstract frames them as new contributions of this survey even though the text attributes them to other works. This pattern means a reader cannot determine which parts of the survey are trustworthy; the central roadmap value is lost. The strongest claim therefore fails as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents itself as a structured survey of LLM-driven scientific hypothesis generation and validation, covering methods, datasets, tools, and future directions. It proposes a taxonomy of generation approaches (knowledge-driven, data-driven, AI-driven exploration, text mining, simulation, interactive, causal, dynamic, and multi-agent) and validation approaches (experimental, simulation, predictive, cross-domain, human-AI, causal, benchmarking, multi-agent, explainability, and hybrid). The abstract additionally claims to introduce two new resources, AHTech and CSKG-600. The paper's stated central contribution, per §1.3, is to provide a reliable, interdisciplinary roadmap of the field that researchers can use for tool selection and research planning.","tokens_in":39049,"tokens_out":2688,"duration_ms":28474,"significance":"If the survey were factually reliable, it would fill a useful niche: a cross-domain map of LLM-based hypothesis generation and validation, with a taxonomy, dataset summaries, and a forward-looking roadmap. The paper also ships a large set of worked definitions (Eqs. 1–8), a structured table of datasets, and a discussion of ethical and regulatory concerns. However, the value of a survey lies in the accuracy of the field map it provides. The manuscript contains multiple misattributions, unsupported tool names, and a misrepresentation of its own contributions, which calls into question the reliability of the entire roadmap. The existence of machine-checkable derivations is not relevant here because the central claim is factual accuracy, not formal proof.","major_comments":[{"comment":"The description of SciFact is factually incorrect and load-bearing for the text-mining subsection. The text states that 'SciFact leverages co-occurrence patterns in scientific literature to propose new drug-disease relationships,' but the canonical SciFact system (Wadden et al., ACL 2020) is a benchmark for verifying scientific claims against evidence, not a hypothesis-generation or drug-repurposing tool. Since the survey's usefulness depends on correctly characterizing representative systems, this error undermines the credibility of the taxonomical claims in §4.","section":"§4.4"},{"comment":"The tools CausalNet, BayesCausality, and InterveneAI are presented as concrete systems for causal inference, but the citations provided are general references on causal modeling (Peters et al. 2017, Lucas 2007, Neuberg 2003) and do not establish the existence of these named tools. Without verifiable sources, readers cannot determine whether these tools exist or perform as described, and the causal-inference section of the survey therefore fails as a reliable map of available methods.","section":"§4.7"},{"comment":"The cross-domain validation subsection names CrossValNet, InterdisciplinaryTest, and TransferTest, with citations to Sybrandt et al. [2018], Zhou et al. [2024], and Touvron et al. [2023]. None of these cited works appear to contain the named systems. This is not a minor citation slip; the survey is asserting the existence of specific tools as examples of a methodological category, and the absence of supporting references for three consecutive named systems indicates a systematic verification failure in this subsection.","section":"§5.4"},{"comment":"The abstract claims the paper is 'introducing new resources like AHTech and CSKG-600,' but the manuscript text attributes AHTech to Lin et al. [2025] and CSKG-600 to Borrego et al. [2025] as pre-existing works. Tables 2 and 3 likewise list these as third-party datasets. This is a direct misrepresentation of the survey's own contribution, and it compounds the credibility problem: the central claim of introducing new resources is not supported by the paper's own content.","section":"Abstract and §3 / Tables 2–3"},{"comment":"The recurring pattern of misdescribed or unsupported named tools is not confined to one subsection; it appears in §4.4, §4.7, and §5.4. Because the survey's primary contribution is to provide a reliable field map, these errors are not isolated presentational flaws. They undermine the central claim of §1.3 that the survey offers a 'structured, interdisciplinary overview' that can guide tool selection and future research. As written, the roadmap cannot be trusted without extensive independent verification.","section":"General accuracy of the field map"}],"minor_comments":[{"comment":"Several inline citations are incomplete, including 'Džeroski et al.' with no year and multiple references to 'Zhou et al. [2024]' that do not disambiguate between at least two distinct works. The references for tools such as CrowdScience, ConceptNet, ExplanatoryAI, and FeedbackLoopAI are missing or only indirectly implied.","section":"Throughout"},{"comment":"The figures contain garbled labels and formatting artifacts, e.g., 'MeS', 'C O C O Datas', and 'Open Graph Benchmar' in Figure 3, and inconsistent use of punctuation in Figure 4. These reduce readability and should be corrected.","section":"Figures 3–4"},{"comment":"In the sentence 'MOLIERE Sybrandt et al. [2017, 2018]demonstrates how text mining...', there is a missing space before 'demonstrates', and the same pattern appears elsewhere in the manuscript.","section":"§4.4"},{"comment":"The definitions of feasibility and quality include weights w_emp, w_theo, w_N, w_F, and w_R that are never specified or exemplified. This is acceptable as a formal definition, but the text would be clearer if it noted how these weights might be set in practice.","section":"§2.1, Eq. (2) and (5)"}],"recommendation":"reject","confidential_remarks":"The number of misattributed or unverifiable tools is too high for a survey whose stated value is accuracy. I did not find evidence of deliberate fabrication, but the manuscript needs a full source-verification pass before it can be considered publishable; because the errors affect the central claim, I cannot recommend major revision as a proportionate response in this round."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper has a genuinely useful skeleton. The two-part structure (generation, then validation) is sensible, the tables organize a lot of scattered work, and the future-directions section, while conventional, covers the right themes. If you need a rough orientation to the space, this gives you one. That's real value, and I don't want to dismiss it.\n\nBut the survey's central claim is to be a reliable map, and that claim fails as stated. The problems are not stylistic or trivial. In §4.4, SciFact is described as a tool that uses co-occurrence patterns to propose drug-disease relationships. The canonical SciFact (Wadden et al., ACL 2020) is a claim-verification benchmark. That is a simple factual error, and it's the kind of error a reader can't easily catch without going to the source. In §4.7, CausalNet, BayesCausality, and InterveneAI are attributed to Peters et al., Lucas, and Neuberg—general causal-inference references, not papers describing those tools. I could not find evidence those tools exist as described. Same pattern in §5.4: CrossValNet, InterdisciplinaryTest, and TransferTest are named with citations that don't contain them. And the abstract says the survey is 'introducing new resources like AHTech and CSKG-600,' but the text correctly attributes those to Lin et al. and Borrego et al. That's an overclaim, and it undermines trust in the authors' own framing.\n\nI'm not holding absence of experiments against a survey. The formulas in §2 are definitions, not new math—fine. But a map with wrong labels is worse than no map, because it sends newcomers to non-existent tools and gives them false confidence about the landscape. The reader's verdict of REJECT is a bit strong; I'd call it 'major revision before it can be trusted.' The errors are fixable. Every named system and dataset needs verification, the abstract needs to stop claiming external resources as new contributions, and the references need to match the claims.\n\nWho is it for? Newcomers who want a bird's-eye view, and even they should double-check every specific tool before acting. It doesn't belong in a citation list in its current state. But the topic is timely and the organizational work is worth preserving, so a serious referee should see it. I'd send it to peer review with the clear expectation of heavy revision.\n\nRecommendation: engage with it, but require the authors to audit every factual claim about tools and datasets before publication.","headline":"The survey's taxonomy is sensible, but the factual map is unreliable: several named tools are misdescribed or lack supporting references, which is a load-bearing flaw for a field-map.","tokens_in":39569,"tokens_out":1803,"would_cite":false,"duration_ms":19039,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that LLM-driven scientific discovery is best understood as a pipeline from data to hypothesis generation to validation, and it maps the methods, datasets, and open problems along that pipeline.","keywords":["scientific hypothesis generation","scientific hypothesis validation","large language models","knowledge graphs","retrieval-augmented generation","causal inference","multi-agent systems","benchmark datasets"],"falsifier":"Draw a random sample of twenty named systems and datasets from the survey (for instance SciFact, CausalNet, BayesCausality, InterveneAI, AHTech, and CSKG-600), and verify each against its cited source or a public repository; if a sizable share cannot be located or are characterized in ways their own documentation contradicts, the survey's central claim of providing a trustworthy map fails.","tokens_in":38497,"feed_emoji":"🧭","tokens_out":7223,"duration_ms":70305,"temperature":0.7,"pith_summary":"This survey's central claim is a mapping claim: it asserts that the scattered landscape of LLM-powered scientific hypothesis generation and validation can be organized into a coherent, interdisciplinary taxonomy. On the generation side it distinguishes symbolic discovery systems from modern LLM pipelines and groups methods into knowledge-driven, data-driven, AI-exploration, text-mining, simulation, collaborative, causal, dynamic-knowledge, and multi-agent families; on the validation side it groups methods into experimental, simulation-based, predictive, cross-domain, human-AI/crowdsourced, causal, benchmarking, multi-agent, explainability, and hybrid families. It also holds up two resources as concrete contributions: AHTech, a high-throughput electrolyte-additive dataset for battery research, and CSKG-600, a set of 600 expert-labeled hypothesis triples over scholarly knowledge graphs. If the map is accurate, the paper gives researchers a tool-selection guide and a forward agenda organized around novelty-aware generation, feasible validation, and ethical safeguards.","feed_headline":"LLM discovery methods, mapped end to end","feed_subtitle":"A structured guide to generation and validation tools, plus two new benchmark datasets.","key_machinery":"The machinery carrying the survey is its taxonomy-cum-pipeline: data input (literature, knowledge graphs, datasets) flows into hypothesis generation, then into iterative validation, refinement, and deployment, and the paper decomposes each stage into named approach families. The formal anchors are compact definitions: novelty as the inverse of average cosine similarity to existing hypotheses, feasibility as a weighted sum of empirical and theoretical scores, and quality as the weighted aggregate $Q(H) = w_N N(H) + w_F F(H) + w_R R(H)$ with weights summing to one. These equations let the survey translate qualitative debates about novelty and feasibility into measurable criteria, while the taxonomy itself is what organizes the hundreds of tools and datasets into a navigable map.","core_discovery":"The paper's central claim is an ordering claim: the many LLM-based systems for scientific discovery form a recognizable pipeline, and progress can be assessed by how well each stage—data integration, hypothesis creation, validation, and refinement—is served by named methods. It argues that early symbolic discovery systems, which search over explicit rule spaces, and LLM-based generative systems, which produce hypotheses through probabilistic token prediction, are complementary rather than competing, and that hybrid pipelines coupling generation with simulation, causal inference, and human oversight are becoming the norm. The survey also presents two resources as concrete contributions: AHTech, a high-throughput dataset of 180 electrolyte additives tested across 200 electrochemical cycles for aqueous zinc batteries, and CSKG-600, a set of 600 expert-labeled candidate hypotheses over scholarly knowledge graphs for evaluating link-prediction systems. The intended payoff is a roadmap that lets researchers choose methods, datasets, and validation strategies by matching them to the structure of their scientific question.","pith_inferences":["A testable extension the paper leaves implicit: use CSKG-600's expert labels to compare retrieval-augmented generation against plain LLM generation on the validity of the resulting hypotheses.","The survey implies that generation and validation are becoming one closed loop, but it never formalizes the loop; a natural next step is a unified scoring function that updates novelty and feasibility weights from validation feedback.","If the map holds, the field's bottleneck shifts from data access to novelty measurement, since generation lacks standardized novelty metrics beyond cosine similarity.","The two highlighted datasets are small enough that pooling them with existing graph and materials benchmarks would be a low-cost way to test cross-domain generalization of hypothesis-generation systems."],"forward_implications":["If the map is right, a researcher facing unstructured biomedical text can pick retrieval-augmented and knowledge-graph generation methods, while a researcher with well-structured symbolic data can use rule-search approaches, instead of guessing.","The two highlighted resources give the community concrete evaluation points: AHTech for high-throughput electrochemical screening and CSKG-600 for hypothesis generation over scholarly knowledge graphs.","The validation taxonomy implies that credible LLM discovery systems should combine at least two validation families, such as simulation plus causal inference, to catch both feasibility and mechanistic errors.","The roadmap's future directions—novelty-aware training, risk-sensitive evaluation, explainable orchestration, and multi-agent reasoning—define concrete criteria for judging next-generation systems.","Because the survey frames generation and validation as an iterative loop, it implies that benchmarks should measure the full loop rather than isolated stages."],"supporting_citations":[{"why":"Supplies the survey's flagship example of an AI system resolving a long-standing scientific challenge, motivating the entire survey.","marker":"Jumper et al. [2021]"},{"why":"The recurring example of an LLM-driven multi-agent system for knowledge-graph-based hypothesis generation and validation.","marker":"Ghafarollahi and Buehler [2024b]"},{"why":"MOLIERE anchors the knowledge-graph and biomedical text-mining family and the validation-by-retrospective-testing idea.","marker":"Sybrandt et al. [2017, 2018]"},{"why":"The AI Scientist is the paper's primary instance of an autonomous agentic system spanning generation, validation, and manuscript drafting.","marker":"Lu et al. [2024]"},{"why":"An earlier agentic-system taxonomy the survey builds on and distinguishes itself from.","marker":"Gridach et al. [2025]"},{"why":"Supports the AI-driven exploration family with reinforcement-learning-based materials discovery.","marker":"Gruver et al. [2024]"},{"why":"Source of the newly highlighted AHTech electrolyte-additive dataset for high-throughput electrochemical hypothesis testing.","marker":"Lin et al. [2025]"},{"why":"Source of the newly highlighted CSKG-600 benchmark for expert-labeled hypothesis triples over scholarly knowledge graphs.","marker":"Borrego et al. [2025]"}],"fun_headline_variants":["LLM discovery pipeline: from hypothesis to validation","Survey maps hybrid AI systems for scientific discovery","New benchmarks AHTech and CSKG-600 for discovery","Symbolic to LLM: hypothesis generation methods survey","A roadmap for LLM-driven scientific discovery tools"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole map depends on the accuracy of its descriptions of dozens of named tools and datasets; if many of those descriptions are wrong or the tools do not exist as described, the survey ceases to be a reliable guide.","fun_headline_variants_meta":{"raw":{"variants":["LLM discovery pipeline: from hypothesis to validation","Survey maps hybrid AI systems for scientific discovery","New benchmarks AHTech and CSKG-600 for discovery","Symbolic to LLM: hypothesis generation methods survey","A roadmap for LLM-driven scientific discovery tools"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000244,"raw_usage":{"total_tokens":1525,"prompt_tokens":932,"completion_tokens":593,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":519}},"tokens_in":548,"tokens_out":593,"duration_ms":6332,"temperature":1.0,"reasoning_tokens":519,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:42:33.965919+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Draw a random sample of twenty named systems and datasets from the survey (for instance SciFact, CausalNet, BayesCausality, InterveneAI, AHTech, and CSKG-600), and verify each against its cited source or a public repository; if a sizable share cannot be located or are characterized in ways their own documentation contradicts, the survey's central claim of providing a trustworthy map fails.","supporting_citations":[{"cited_title":"Research hypothesis generation over scientific knowledge graphs","cited_arxiv_id":null,"evidence_quote":"Source of the newly highlighted CSKG-600 benchmark for expert-labeled hypothesis triples over scholarly knowledge graphs."}],"review_version":1}