REVIEW 2 major objections 2 minor 1 cited by
Large language models fail to spontaneously propose null hypotheses when generating scientific ideas.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-27 19:05 UTC pith:O24ZW2LV
load-bearing objection Large rating dataset and post-trained reward model are the real assets; the null-hypothesis contrast lacks a matched human generation arm. the 2 major comments →
Contemporary AI lacks the imagination to diverge or negate in science
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In ratings from 6,749 scientists on 25,139 LLM-generated ideas drawn from 121,640 preprints, no model class proposes null hypotheses on its own. Non-reasoning models produce narrow clusters of similar ideas while reasoning models range more widely, yet both avoid negation. Scientists reward resemblance to their own work and probability of being true over novelty, with social scientists showing greater risk tolerance; automated judges align only weakly with these expert assessments.
What carries the argument
Spontaneous proposal of null hypotheses as a marker of the ability to diverge or negate within a hypothesis space.
Load-bearing premise
The scientists who responded and the preprints they supplied form a representative sample of scientific reasoning without systematic selection bias.
What would settle it
A new model class that, in a blinded replication using fresh preprints, proposes null hypotheses at rates comparable to the human authors would falsify the central claim.
If this is right
- LLM outputs and judgments in science require ongoing human grounding to compensate for limited imaginative divergence.
- Post-training a reward model on human ratings improves capture of field-specific tastes by up to 27 percent over prior automated evaluators.
- Social scientists tolerate riskier ideas more than life scientists, and senior social scientists apply the strictest standards.
- Retrieval augmentation and persona prompting produce only marginal gains in alignment with expert judgment.
Where Pith is reading between the lines
- Explicit training for contradiction or falsification may be needed before models can reliably explore negation in hypothesis generation.
- The performance gap in pluralistic fields points to a broader limit on AI handling interpretive or theory-evolving domains.
- Reward models tuned to human ratings could serve as scalable proxies for expert review in early-stage idea filtering.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports results from inviting authors of 121,640 recent preprints to rate LLM-generated ideas derived from their own papers, yielding 6,749 respondents and 25,139 rating sets. It identifies three patterns: non-reasoning LLMs produce narrow idea sets while reasoning models explore more broadly but none spontaneously generate null hypotheses (unlike humans); scientists favor ideas resembling their own and prioritize probability over novelty, with field and seniority differences; and automated evaluators (including LLM-as-judge) show weak agreement with experts, though a post-trained Qwen3-14B reward model improves alignment with human ratings by up to 27%.
Significance. If the empirical patterns hold after addressing baseline issues, the work supplies the largest scientist-in-the-loop dataset to date on AI ideation limitations in science, particularly the absence of spontaneous negation, and demonstrates a practical path for training reward models that better capture expert preferences across fields. This supplies falsifiable, quantitative evidence against claims of imminent AI-driven discovery acceleration.
major comments (2)
- [Abstract] Abstract: the central claim that 'no model class spontaneously proposes null hypotheses -- a move humans make more freely' is unsupported by any matched human generation arm; the design collects only ratings of LLM outputs and provides no parallel condition in which the same scientists (or matched authors) generate ideas from identical paper contexts and prompts, preventing attribution of the differential to model architecture rather than prompt framing or training objectives.
- [Abstract] Abstract and implied Methods: the three reported patterns rest on unexamined selection and prompting assumptions, with no reported details on idea-generation prompts, response-rate bias controls, inter-rater reliability calculations, or statistical adjustments for field and seniority; these omissions are load-bearing because the patterns (hivemind collapse, field differences in risk tolerance, and evaluator disagreement) cannot be interpreted without them.
minor comments (2)
- [Abstract] Abstract: the response rate (6,749/121,640) and exact sample composition by field should be stated explicitly to allow readers to assess representativeness.
- The description of the post-trained reward model would benefit from a brief statement of the training objective and loss function used.
Simulated Author's Rebuttal
We thank the referee for these constructive comments. We agree that the abstract claim regarding null hypotheses requires qualification given the study design, and that additional methodological details are needed for full interpretability. We outline revisions below.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central claim that 'no model class spontaneously proposes null hypotheses -- a move humans make more freely' is unsupported by any matched human generation arm; the design collects only ratings of LLM outputs and provides no parallel condition in which the same scientists (or matched authors) generate ideas from identical paper contexts and prompts, preventing attribution of the differential to model architecture rather than prompt framing or training objectives.
Authors: We acknowledge the limitation: our data consist solely of ratings on LLM-generated ideas and contain no matched human generation condition using identical prompts and paper contexts. The phrasing 'a move humans make more freely' therefore cannot be directly attributed to this experiment. We will revise the abstract and discussion to state that no LLM condition in the study produced null hypotheses, while noting that the contrast with human scientific practice draws from established literature on hypothesis generation rather than a within-study comparison. This removes the unsupported attribution. revision: yes
-
Referee: [Abstract] Abstract and implied Methods: the three reported patterns rest on unexamined selection and prompting assumptions, with no reported details on idea-generation prompts, response-rate bias controls, inter-rater reliability calculations, or statistical adjustments for field and seniority; these omissions are load-bearing because the patterns (hivemind collapse, field differences in risk tolerance, and evaluator disagreement) cannot be interpreted without them.
Authors: We will expand the Methods section with the requested details: verbatim idea-generation prompts, response-rate calculations and any non-response bias checks performed, inter-rater reliability metrics (e.g., agreement coefficients across the 25,139 rating sets), and the statistical models (including covariates for field and seniority) used to support the reported patterns. These additions will make the three patterns fully interpretable without altering the core findings. revision: yes
Circularity Check
Purely empirical rating study with no derivations or self-referential predictions
full rationale
The paper reports results from inviting 6,749 scientists to rate 25,139 LLM-generated ideas drawn from their own preprints on novelty, feasibility, truth probability, and adoption favorability. No equations, fitted parameters, or first-principles derivations appear; the three patterns are direct summaries of human ratings. The post-trained Qwen3-14B reward model is trained on the collected ratings and evaluated against held-out human judgments, which is standard supervised learning rather than a circular prediction. No self-citation chains or uniqueness theorems are invoked to justify core claims. The study is therefore self-contained against external human benchmarks.
Axiom & Free-Parameter Ledger
axioms (2)
- domain assumption Scientist ratings constitute valid ground truth for novelty, feasibility, and adoption potential
- domain assumption The invited preprint authors form an unbiased sample of working scientists in the four fields
read the original abstract
Bold projections that artificial intelligence will accelerate scientific discovery have raced ahead of evidence from working scientists, and the field still lacks large-scale, scientist-in-the-loop tests of these claims. Here we mount the largest such evaluation to date and map what AI cannot yet do for science. We invited authors of 121,640 recent preprints across biology, medicine, chemistry, and the social sciences to judge ideas that large language models (LLMs) generated from the context and puzzles of their own papers. 6,749 scientists returned 25,139 sets of ratings on novelty, empirical feasibility, probability of being true, and favorability of adoption. Three patterns emerge. First, non-reasoning LLMs collapse into a narrow "hivemind" of similar ideas; reasoning models roam a wider hypothesis space, yet no model class spontaneously proposes null hypotheses -- a move humans make more freely. Second, scientists reward ideas that resemble their own and prize probability over novelty, though social scientists tolerate risk more readily than life scientists. Senior social scientists are the harshest critics, and their skepticism is well-earned: LLMs falter most in pluralistic fields like the social sciences that demand context-aware interpretation and evolving theories. Third, automated evaluators on which the community currently relies -- LLM-as-a-judge, artificial metrics, and even state-of-the-art (SOTA) models -- agree only weakly with expert judgment, and retrieval augmentation and scientist persona prompting yield only marginal gains. A Qwen3-14B reward model we post-trained on human ratings captures field taste nuances, beats SOTA models by up to 27%, and closes the gap to the inter-rater consistency of independent peer reviewers. For all the hype, today's scientific AI still represents a collaborator whose imagination, outputs and judgment benefit from human grounding.
Figures
Forward citations
Cited by 1 Pith paper
-
Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI
Open-ended AI is blocked by a vocabulary gap (inventing reusable primitives) and a verifier gap (valuing them when payoff is delayed), unified under cognitive discrepancy reduction and a four-level autonomy ladder.
Reference graph
Works this paper leans on
-
[1]
**Exclude** background/factual claims, method descriptions, presumptions, mathematical prerequisites/preconditions, and narrow inferences/interpretations drawn from tables and figures
-
[2]
**Apply strict selection criteria - do not over-generate**. Extract only those hypotheses that the authors explicitly motivate and place at the core of the paper’s main argument (typically introduced early, e.g., in the Introduction or thereafter). Omit minor and peripheral hypotheses confined to specific method, experiment, or result subsections
-
[3]
xxx is a valuable model for studying xxx
**Exclude vague directional statements** such as "xxx is a valuable model for studying xxx"; "The proposed model provides a promising and crucial direction for xxx"; and "This technology can be applied to improve xxx," as these are summaries of the paper’s overarching narrative rather than hypotheses
-
[4]
Do not generate or infer hypotheses on your own
Extract hypotheses based strictly on the **raw content**. Do not generate or infer hypotheses on your own. Preserve the **original meaning** of the authors’ hypotheses
-
[5]
––- PAPER STARTS ––- f{text} ––- PAPER ENDS ––- Context/Puzzle Extraction Prompt You are a helpful research assistant
If no relevant, explicit hypotheses are found, output an empty string "". ––- PAPER STARTS ––- f{text} ––- PAPER ENDS ––- Context/Puzzle Extraction Prompt You are a helpful research assistant. You will be given the introduction of a scientific paper. Your task is to identify and extract two kinds of structured information from the text:
-
[6]
These include big pictures and related works, etc
The **broad scientific context** of the work: **Context** consists of only *factual*, *non-speculative, non-reasoning* statements. These include big pictures and related works, etc. This should read like what a researcher might see *before* 50 they propose a theory. You extract hypothesis-agnostic explanations or definitions of key terms and concepts that...
-
[7]
Focus on the **high-level picture** of the work
At the **end of each context**, add a sentence that explicitly states *the core question or puzzle* that the paper addresses and can be derived from the context. Focus on the **high-level picture** of the work. **Avoid**: - Specific hypotheses or findings - Technical/experimental details, methods, or datasets - Any mention of author-proposed solutions You...
-
[8]
**Be concise**: exclude **background/factual claims or methodological/experimental descriptions**
-
[9]
Keep the hypotheses **relevant** to the core of the puzzle
-
[10]
xxx is a valuable model for studying xxx
**Do not generate vague directional statements** such as "xxx is a valuable model for studying xxx"; "The proposed model provides a promising and crucial direction for xxx"; and "This technology can be applied to improve xxx", as these are summaries of the paper’s overarching narrative rather than hypotheses
-
[11]
**Be creative** - reach beyond your existing knowledge base to propose **untested** ideas
-
[12]
fast-response cognitive mode
Ensure that the generated hypotheses are **clear, specified, well-reasoned, valid, and actionable**. ––- CONTEXTUAL PUZZLE STARTS ––- f{text} ––- CONTEXTUAL PUZZLE ENDS ––- Now, your output of the hypotheses: Model Evaluation Prompt You are an experienced scientist who is judging ideas (hypotheses) proposed from the same context and puzzle as your own pap...
-
[13]
the context: background information and the research puzzle of the paper
-
[14]
Evaluate the extent to which the hypotheses, generated based on the given context, introduce new ideas beyond the context
two proposed hypotheses: Hypothesis A and Hypothesis B proposed based on the given context Your task: Compare the two hypotheses on novelty according to the following criteria: "Evaluate the extent to which the hypotheses, generated based on the given context, introduce new ideas beyond the context.” [Replace novelty with other dimensions:] Feasibility:"E...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.