Pith. sign in

REVIEW 2 major objections 2 minor 2 cited by

Contemporary AI lacks the imagination to diverge or negate in science

T0 review · 2 major / 2 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read Large language models fail to spontaneously propose null hypotheses when generating scientific ideas.

desk verdict Large rating dataset and post-trained reward model are the real assets; the null-hypothesis contrast lacks a matched human generation arm. read the letter →

arxiv 2606.08251 v3 pith:O24ZW2LV submitted 2026-06-06 cs.CY cs.AI

classification cs.CYcs.AI
keywords largelanguagemodelsscientificdiscoverynullhypotheseshypothesisgenerationexpertevaluationAIcollaborationideanovelty
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper conducts the largest evaluation to date in which working scientists judge ideas generated by LLMs from the context of their own recent preprints. It finds that models either collapse to similar ideas or explore wider spaces without ever suggesting null hypotheses, a move human scientists make more freely. Scientists consistently favor probable ideas over novel ones, rate LLM outputs more harshly in pluralistic fields, and show only weak agreement with automated evaluators. A reward model trained on the collected ratings narrows the gap to human inter-rater consistency.

What carries the argument

Spontaneous proposal of null hypotheses as a marker of the ability to diverge or negate within a hypothesis space.

What would settle it

A new model class that, in a blinded replication using fresh preprints, proposes null hypotheses at rates comparable to the human authors would falsify the central claim.

Watch

Extended reading notes

Core claim

In ratings from 6,749 scientists on 25,139 LLM-generated ideas drawn from 121,640 preprints, no model class proposes null hypotheses on its own. Non-reasoning models produce narrow clusters of similar ideas while reasoning models range more widely, yet both avoid negation. Scientists reward resemblance to their own work and probability of being true over novelty, with social scientists showing greater risk tolerance; automated judges align only weakly with these expert assessments.

Load-bearing premise

The scientists who responded and the preprints they supplied form a representative sample of scientific reasoning without systematic selection bias.

Editorial extensions

If this is right

  • LLM outputs and judgments in science require ongoing human grounding to compensate for limited imaginative divergence.
  • Post-training a reward model on human ratings improves capture of field-specific tastes by up to 27 percent over prior automated evaluators.
  • Social scientists tolerate riskier ideas more than life scientists, and senior social scientists apply the strictest standards.
  • Retrieval augmentation and persona prompting produce only marginal gains in alignment with expert judgment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Explicit training for contradiction or falsification may be needed before models can reliably explore negation in hypothesis generation.
  • The performance gap in pluralistic fields points to a broader limit on AI handling interpretive or theory-evolving domains.
  • Reward models tuned to human ratings could serve as scalable proxies for expert review in early-stage idea filtering.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript reports results from inviting authors of 121,640 recent preprints to rate LLM-generated ideas derived from their own papers, yielding 6,749 respondents and 25,139 rating sets. It identifies three patterns: non-reasoning LLMs produce narrow idea sets while reasoning models explore more broadly but none spontaneously generate null hypotheses (unlike humans); scientists favor ideas resembling their own and prioritize probability over novelty, with field and seniority differences; and automated evaluators (including LLM-as-judge) show weak agreement with experts, though a post-trained Qwen3-14B reward model improves alignment with human ratings by up to 27%.

Significance. If the empirical patterns hold after addressing baseline issues, the work supplies the largest scientist-in-the-loop dataset to date on AI ideation limitations in science, particularly the absence of spontaneous negation, and demonstrates a practical path for training reward models that better capture expert preferences across fields. This supplies falsifiable, quantitative evidence against claims of imminent AI-driven discovery acceleration.

major comments (2)
  1. [Abstract] Abstract: the central claim that 'no model class spontaneously proposes null hypotheses -- a move humans make more freely' is unsupported by any matched human generation arm; the design collects only ratings of LLM outputs and provides no parallel condition in which the same scientists (or matched authors) generate ideas from identical paper contexts and prompts, preventing attribution of the differential to model architecture rather than prompt framing or training objectives.
  2. [Abstract] Abstract and implied Methods: the three reported patterns rest on unexamined selection and prompting assumptions, with no reported details on idea-generation prompts, response-rate bias controls, inter-rater reliability calculations, or statistical adjustments for field and seniority; these omissions are load-bearing because the patterns (hivemind collapse, field differences in risk tolerance, and evaluator disagreement) cannot be interpreted without them.
minor comments (2)
  1. [Abstract] Abstract: the response rate (6,749/121,640) and exact sample composition by field should be stated explicitly to allow readers to assess representativeness.
  2. The description of the post-trained reward model would benefit from a brief statement of the training objective and loss function used.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for these constructive comments. We agree that the abstract claim regarding null hypotheses requires qualification given the study design, and that additional methodological details are needed for full interpretability. We outline revisions below.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the central claim that 'no model class spontaneously proposes null hypotheses -- a move humans make more freely' is unsupported by any matched human generation arm; the design collects only ratings of LLM outputs and provides no parallel condition in which the same scientists (or matched authors) generate ideas from identical paper contexts and prompts, preventing attribution of the differential to model architecture rather than prompt framing or training objectives.

    Authors: We acknowledge the limitation: our data consist solely of ratings on LLM-generated ideas and contain no matched human generation condition using identical prompts and paper contexts. The phrasing 'a move humans make more freely' therefore cannot be directly attributed to this experiment. We will revise the abstract and discussion to state that no LLM condition in the study produced null hypotheses, while noting that the contrast with human scientific practice draws from established literature on hypothesis generation rather than a within-study comparison. This removes the unsupported attribution. revision: yes

  2. Referee: [Abstract] Abstract and implied Methods: the three reported patterns rest on unexamined selection and prompting assumptions, with no reported details on idea-generation prompts, response-rate bias controls, inter-rater reliability calculations, or statistical adjustments for field and seniority; these omissions are load-bearing because the patterns (hivemind collapse, field differences in risk tolerance, and evaluator disagreement) cannot be interpreted without them.

    Authors: We will expand the Methods section with the requested details: verbatim idea-generation prompts, response-rate calculations and any non-response bias checks performed, inter-rater reliability metrics (e.g., agreement coefficients across the 25,139 rating sets), and the statistical models (including covariates for field and seniority) used to support the reported patterns. These additions will make the three patterns fully interpretable without altering the core findings. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

Purely empirical rating study with no derivations or self-referential predictions

full rationale

The paper reports results from inviting 6,749 scientists to rate 25,139 LLM-generated ideas drawn from their own preprints on novelty, feasibility, truth probability, and adoption favorability. No equations, fitted parameters, or first-principles derivations appear; the three patterns are direct summaries of human ratings. The post-trained Qwen3-14B reward model is trained on the collected ratings and evaluated against held-out human judgments, which is standard supervised learning rather than a circular prediction. No self-citation chains or uniqueness theorems are invoked to justify core claims. The study is therefore self-contained against external human benchmarks.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The study rests on standard assumptions of survey methodology and statistical aggregation of ratings; no free parameters, ad-hoc axioms, or invented entities are introduced in the abstract.

assumptions (2)
  • domain assumption Scientist ratings constitute valid ground truth for novelty, feasibility, and adoption potential
    Invoked when the paper treats the 25,139 ratings as the benchmark against which LLMs and automated judges are evaluated.
  • domain assumption The invited preprint authors form an unbiased sample of working scientists in the four fields
    Required for generalizing the three patterns beyond the 6,749 respondents.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Contemporary AI lacks the imagination to diverge or negate in science." pith.science (2026). https://pith.science/paper/O24ZW2LV

@misc{pith2026260608251,
  author       = {Pith},
  title        = {Pith review of: Contemporary AI lacks the imagination to diverge or negate in science},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/O24ZW2LV}},
  note         = {Machine review of arXiv:2606.08251}
}
read the original abstract

Bold claims that AI will accelerate scientific discovery have raced ahead of evidence from working scientists, yet large-scale, scientist-in-the-loop evidence is scarce. Here we mount the largest evaluation to date, inviting authors of 121,640 recent preprints in biology, medicine, chemistry, and social science to judge large language model (LLM)-generated ideas derived from their own papers. 6,749 representative scientists returned 25,139 rating sets on novelty, feasibility, probability of being true, and favorability of adoption. Three patterns emerge. First, non-reasoning LLMs collapse into a narrow "hivemind" of similar ideas while reasoning models explore a wider hypothesis space, but no model spontaneously proposes null hypotheses, a move humans make more freely. Second, scientists reward ideas resembling their own and prize probability over novelty, though social scientists tolerate risk more than life scientists; senior social scientists are the harshest critics, and their skepticism is earned, as LLMs falter most in pluralistic fields demanding context-aware interpretation and evolving theories. Third, automated evaluators, including LLM-as-a-judge and state-of-the-art (SOTA) models, agree weakly with expert judgment. Retrieval augmentation and scientist persona prompting yield marginal gains. A Qwen3-14B reward model we post-trained on human ratings captures nuances of taste, beats SOTA models by up to 27%, and closes the gap to the consistency of human peer reviewers. An analysis of 39 million papers from 2010 to 2025 links survey findings to macro-level patterns: following ChatGPT's release, null claims are sharply suppressed and ideas contract. Agent-based simulations further suggest that saturated fields should especially prize human uniqueness. For all the hype, today's AI for science remains a collaborator whose imagination and judgment benefit from human grounding.

Figures

Figures reproduced from arXiv: 2606.08251 by the authors.

Figure 1
Figure 1. An expert-audit pipeline for AI-generated research ideas. Full-text preprints (n = 121,640) from six non-arXiv platforms feed an extraction stage that recovers (i) the author’s hypotheses, (ii) the surrounding factual context, and (iii) the core scientific puzzle, with paraphrase-based leakage detection between (i), (ii), and (iii). LLMs propose hypotheses from the context-and-puzzle alone; a custom set of hypothese… view at source ↗
Figure 2
Figure 2. Reasoning broadens the hypothesis space; null reasoning rarely fills it. a, Pairwise cosine similarity of hypotheses generated for the same paper, by group. Non￾reasoning LLMs are more similar to each other (the “artificial hivemind”, p<0.001); reasoning models diverge from non-reasoning models, humans, and each other. b, geometrically, we treat each hypothesis as a displacement from a common context-and-puzzle orig… view at source ↗
Figure 3
Figure 3. Scientists discount novelty, prefer ideas resembling their own, and split by field and seniority. Marginal predictions from the Mundlak adoption model are shown, controlling for rated quality. a, within-scientist similarity to the author’s own ideas is the strongest single driver of adoption. b, status, represented by within-field citation/publica￾tion percentile, lowers adoption. c, seniority, represented by log-tr… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Automated evaluators do not yet measure scientific quality. a, Pearson correlation between LLM-judge ratings and human ratings, by dimension and judge, with and without injected scientist persona. The retrieval-augmented Deep Research judge is the best. Persona injecti…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    LLM judges of scientific ideas are measurably swayed by writing style; a style-detecting module reduces but does not remove the bias.

  2. Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Open-ended AI is blocked by a vocabulary gap (inventing reusable primitives) and a verifier gap (valuing them when payoff is delayed), unified under cognitive discrepancy reduction and a four-level autonomy ladder.

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.