REVIEW 2 major objections 2 minor 2 cited by
Contemporary AI lacks the imagination to diverge or negate in science
T0 review · 2 major / 2 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read Large language models fail to spontaneously propose null hypotheses when generating scientific ideas.
desk verdict Large rating dataset and post-trained reward model are the real assets; the null-hypothesis contrast lacks a matched human generation arm. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Spontaneous proposal of null hypotheses as a marker of the ability to diverge or negate within a hypothesis space.
What would settle it
A new model class that, in a blinded replication using fresh preprints, proposes null hypotheses at rates comparable to the human authors would falsify the central claim.
Extended reading notes
Core claim
In ratings from 6,749 scientists on 25,139 LLM-generated ideas drawn from 121,640 preprints, no model class proposes null hypotheses on its own. Non-reasoning models produce narrow clusters of similar ideas while reasoning models range more widely, yet both avoid negation. Scientists reward resemblance to their own work and probability of being true over novelty, with social scientists showing greater risk tolerance; automated judges align only weakly with these expert assessments.
Load-bearing premise
The scientists who responded and the preprints they supplied form a representative sample of scientific reasoning without systematic selection bias.
Editorial extensions
If this is right
- LLM outputs and judgments in science require ongoing human grounding to compensate for limited imaginative divergence.
- Post-training a reward model on human ratings improves capture of field-specific tastes by up to 27 percent over prior automated evaluators.
- Social scientists tolerate riskier ideas more than life scientists, and senior social scientists apply the strictest standards.
- Retrieval augmentation and persona prompting produce only marginal gains in alignment with expert judgment.
Reading between the lines
- Explicit training for contradiction or falsification may be needed before models can reliably explore negation in hypothesis generation.
- The performance gap in pluralistic fields points to a broader limit on AI handling interpretive or theory-evolving domains.
- Reward models tuned to human ratings could serve as scalable proxies for expert review in early-stage idea filtering.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports results from inviting authors of 121,640 recent preprints to rate LLM-generated ideas derived from their own papers, yielding 6,749 respondents and 25,139 rating sets. It identifies three patterns: non-reasoning LLMs produce narrow idea sets while reasoning models explore more broadly but none spontaneously generate null hypotheses (unlike humans); scientists favor ideas resembling their own and prioritize probability over novelty, with field and seniority differences; and automated evaluators (including LLM-as-judge) show weak agreement with experts, though a post-trained Qwen3-14B reward model improves alignment with human ratings by up to 27%.
Significance. If the empirical patterns hold after addressing baseline issues, the work supplies the largest scientist-in-the-loop dataset to date on AI ideation limitations in science, particularly the absence of spontaneous negation, and demonstrates a practical path for training reward models that better capture expert preferences across fields. This supplies falsifiable, quantitative evidence against claims of imminent AI-driven discovery acceleration.
major comments (2)
- [Abstract] Abstract: the central claim that 'no model class spontaneously proposes null hypotheses -- a move humans make more freely' is unsupported by any matched human generation arm; the design collects only ratings of LLM outputs and provides no parallel condition in which the same scientists (or matched authors) generate ideas from identical paper contexts and prompts, preventing attribution of the differential to model architecture rather than prompt framing or training objectives.
- [Abstract] Abstract and implied Methods: the three reported patterns rest on unexamined selection and prompting assumptions, with no reported details on idea-generation prompts, response-rate bias controls, inter-rater reliability calculations, or statistical adjustments for field and seniority; these omissions are load-bearing because the patterns (hivemind collapse, field differences in risk tolerance, and evaluator disagreement) cannot be interpreted without them.
minor comments (2)
- [Abstract] Abstract: the response rate (6,749/121,640) and exact sample composition by field should be stated explicitly to allow readers to assess representativeness.
- The description of the post-trained reward model would benefit from a brief statement of the training objective and loss function used.
Simulated Author's Rebuttal
We thank the referee for these constructive comments. We agree that the abstract claim regarding null hypotheses requires qualification given the study design, and that additional methodological details are needed for full interpretability. We outline revisions below.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central claim that 'no model class spontaneously proposes null hypotheses -- a move humans make more freely' is unsupported by any matched human generation arm; the design collects only ratings of LLM outputs and provides no parallel condition in which the same scientists (or matched authors) generate ideas from identical paper contexts and prompts, preventing attribution of the differential to model architecture rather than prompt framing or training objectives.
Authors: We acknowledge the limitation: our data consist solely of ratings on LLM-generated ideas and contain no matched human generation condition using identical prompts and paper contexts. The phrasing 'a move humans make more freely' therefore cannot be directly attributed to this experiment. We will revise the abstract and discussion to state that no LLM condition in the study produced null hypotheses, while noting that the contrast with human scientific practice draws from established literature on hypothesis generation rather than a within-study comparison. This removes the unsupported attribution. revision: yes
-
Referee: [Abstract] Abstract and implied Methods: the three reported patterns rest on unexamined selection and prompting assumptions, with no reported details on idea-generation prompts, response-rate bias controls, inter-rater reliability calculations, or statistical adjustments for field and seniority; these omissions are load-bearing because the patterns (hivemind collapse, field differences in risk tolerance, and evaluator disagreement) cannot be interpreted without them.
Authors: We will expand the Methods section with the requested details: verbatim idea-generation prompts, response-rate calculations and any non-response bias checks performed, inter-rater reliability metrics (e.g., agreement coefficients across the 25,139 rating sets), and the statistical models (including covariates for field and seniority) used to support the reported patterns. These additions will make the three patterns fully interpretable without altering the core findings. revision: yes
Circularity Check
Purely empirical rating study with no derivations or self-referential predictions
full rationale
The paper reports results from inviting 6,749 scientists to rate 25,139 LLM-generated ideas drawn from their own preprints on novelty, feasibility, truth probability, and adoption favorability. No equations, fitted parameters, or first-principles derivations appear; the three patterns are direct summaries of human ratings. The post-trained Qwen3-14B reward model is trained on the collected ratings and evaluated against held-out human judgments, which is standard supervised learning rather than a circular prediction. No self-citation chains or uniqueness theorems are invoked to justify core claims. The study is therefore self-contained against external human benchmarks.
Assumptions & free parameters
assumptions (2)
- domain assumption Scientist ratings constitute valid ground truth for novelty, feasibility, and adoption potential
- domain assumption The invited preprint authors form an unbiased sample of working scientists in the four fields
Cite this review
Pith. "Pith review of Contemporary AI lacks the imagination to diverge or negate in science." pith.science (2026). https://pith.science/paper/O24ZW2LV
@misc{pith2026260608251,
author = {Pith},
title = {Pith review of: Contemporary AI lacks the imagination to diverge or negate in science},
year = {2026},
howpublished = {\url{https://pith.science/paper/O24ZW2LV}},
note = {Machine review of arXiv:2606.08251}
}
read the original abstract
Bold claims that AI will accelerate scientific discovery have raced ahead of evidence from working scientists, yet large-scale, scientist-in-the-loop evidence is scarce. Here we mount the largest evaluation to date, inviting authors of 121,640 recent preprints in biology, medicine, chemistry, and social science to judge large language model (LLM)-generated ideas derived from their own papers. 6,749 representative scientists returned 25,139 rating sets on novelty, feasibility, probability of being true, and favorability of adoption. Three patterns emerge. First, non-reasoning LLMs collapse into a narrow "hivemind" of similar ideas while reasoning models explore a wider hypothesis space, but no model spontaneously proposes null hypotheses, a move humans make more freely. Second, scientists reward ideas resembling their own and prize probability over novelty, though social scientists tolerate risk more than life scientists; senior social scientists are the harshest critics, and their skepticism is earned, as LLMs falter most in pluralistic fields demanding context-aware interpretation and evolving theories. Third, automated evaluators, including LLM-as-a-judge and state-of-the-art (SOTA) models, agree weakly with expert judgment. Retrieval augmentation and scientist persona prompting yield marginal gains. A Qwen3-14B reward model we post-trained on human ratings captures nuances of taste, beats SOTA models by up to 27%, and closes the gap to the consistency of human peer reviewers. An analysis of 39 million papers from 2010 to 2025 links survey findings to macro-level patterns: following ChatGPT's release, null claims are sharply suppressed and ideas contract. Agent-based simulations further suggest that saturated fields should especially prize human uniqueness. For all the hype, today's AI for science remains a collaborator whose imagination and judgment benefit from human grounding.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 2 Pith papers
-
Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation
LLM judges of scientific ideas are measurably swayed by writing style; a style-detecting module reduces but does not remove the bias.
-
Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI
Open-ended AI is blocked by a vocabulary gap (inventing reusable primitives) and a verifier gap (valuing them when payoff is delayed), unified under cognitive discrepancy reduction and a four-level autonomy ladder.
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.