{"id":"2807fcad-740e-4b66-93d7-d2564df12944","arxiv_id":"2606.08251","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Large-scale human ratings of LLM-generated scientific ideas reveal that models avoid null hypotheses, scientists favor familiar ideas, and a post-trained reward model outperforms SOTA automated evaluators.","lead":"This paper describes the largest scientist-in-the-loop test to date in which 6,749 researchers rated 25,139 AI-generated ideas drawn from their own recent preprints. The results indicate that current LLMs rarely generate null hypotheses and that automated judges align only weakly with expert human ratings.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No matched human generation baseline for the 'null hypotheses more freely' claim","rationale":"Reader correctly flags respondent selection as a threat to rating generalizability. However, the 'humans do it more freely' clause is more directly load-bearing for the title claim and the 'imagination to diverge or negate' conclusion, because it lacks an explicit within-study control. Full-text methods may clarify an external baseline, but the abstract alone leaves this gap; addressing it would move the paper from UNVERDICTED toward CONDITIONAL acceptance of the core pattern.","tokens_in":1839,"tokens_out":349,"duration_ms":18954,"concrete_test":"Select 200 preprints with responding authors; re-prompt those authors with the exact LLM generation template used in the study and collect 3-5 ideas each; apply the same null-hypothesis coding rubric (whatever was used for LLM outputs) to both sets; if human null rate exceeds LLM rate by >3 percentage points after inter-rater reliability check, the spontaneity claim holds; otherwise it requires re-framing as 'under these prompts'.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that 'no model class spontaneously proposes null hypotheses -- a move humans make more freely' is supported only by LLM outputs rated by scientists; the abstract and setup describe no parallel arm in which the same scientists (or matched authors) generate ideas from identical paper contexts and prompts. Without this, the differential cannot be attributed to model architecture versus prompt framing, training objectives that penalize negation, or post-generation filtering. The three reported patterns all rest on the validity of this human-model contrast for divergence/negation.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports results from inviting authors of 121,640 recent preprints to rate LLM-generated ideas derived from their own papers, yielding 6,749 respondents and 25,139 rating sets. It identifies three patterns: non-reasoning LLMs produce narrow idea sets while reasoning models explore more broadly but none spontaneously generate null hypotheses (unlike humans); scientists favor ideas resembling their own and prioritize probability over novelty, with field and seniority differences; and automated evaluators (including LLM-as-judge) show weak agreement with experts, though a post-trained Qwen3-14B reward model improves alignment with human ratings by up to 27%.","tokens_in":1937,"tokens_out":519,"duration_ms":16846,"significance":"If the empirical patterns hold after addressing baseline issues, the work supplies the largest scientist-in-the-loop dataset to date on AI ideation limitations in science, particularly the absence of spontaneous negation, and demonstrates a practical path for training reward models that better capture expert preferences across fields. This supplies falsifiable, quantitative evidence against claims of imminent AI-driven discovery acceleration.","major_comments":[{"comment":"Abstract: the central claim that 'no model class spontaneously proposes null hypotheses -- a move humans make more freely' is unsupported by any matched human generation arm; the design collects only ratings of LLM outputs and provides no parallel condition in which the same scientists (or matched authors) generate ideas from identical paper contexts and prompts, preventing attribution of the differential to model architecture rather than prompt framing or training objectives.","section":"Abstract"},{"comment":"Abstract and implied Methods: the three reported patterns rest on unexamined selection and prompting assumptions, with no reported details on idea-generation prompts, response-rate bias controls, inter-rater reliability calculations, or statistical adjustments for field and seniority; these omissions are load-bearing because the patterns (hivemind collapse, field differences in risk tolerance, and evaluator disagreement) cannot be interpreted without them.","section":"Abstract"}],"minor_comments":[{"comment":"Abstract: the response rate (6,749/121,640) and exact sample composition by field should be stated explicitly to allow readers to assess representativeness.","section":"Abstract"},{"comment":"The description of the post-trained reward model would benefit from a brief statement of the training objective and loss function used.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for these constructive comments. We agree that the abstract claim regarding null hypotheses requires qualification given the study design, and that additional methodological details are needed for full interpretability. We outline revisions below.","responses":[{"response":"We acknowledge the limitation: our data consist solely of ratings on LLM-generated ideas and contain no matched human generation condition using identical prompts and paper contexts. The phrasing 'a move humans make more freely' therefore cannot be directly attributed to this experiment. We will revise the abstract and discussion to state that no LLM condition in the study produced null hypotheses, while noting that the contrast with human scientific practice draws from established literature on hypothesis generation rather than a within-study comparison. This removes the unsupported attribution.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claim that 'no model class spontaneously proposes null hypotheses -- a move humans make more freely' is unsupported by any matched human generation arm; the design collects only ratings of LLM outputs and provides no parallel condition in which the same scientists (or matched authors) generate ideas from identical paper contexts and prompts, preventing attribution of the differential to model architecture rather than prompt framing or training objectives."},{"response":"We will expand the Methods section with the requested details: verbatim idea-generation prompts, response-rate calculations and any non-response bias checks performed, inter-rater reliability metrics (e.g., agreement coefficients across the 25,139 rating sets), and the statistical models (including covariates for field and seniority) used to support the reported patterns. These additions will make the three patterns fully interpretable without altering the core findings.","revision_made":"yes","referee_comment":"[Abstract] Abstract and implied Methods: the three reported patterns rest on unexamined selection and prompting assumptions, with no reported details on idea-generation prompts, response-rate bias controls, inter-rater reliability calculations, or statistical adjustments for field and seniority; these omissions are load-bearing because the patterns (hivemind collapse, field differences in risk tolerance, and evaluator disagreement) cannot be interpreted without them."}],"tokens_in":1536,"tokens_out":452,"duration_ms":16080,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's clearest value is the scale of the rating exercise: 25k judgments from nearly 7k scientists on ideas pulled from their own preprints. That volume lets them show field differences in risk tolerance and that senior social scientists are stricter raters. The post-trained Qwen reward model is the other concrete output; it beats the SOTA baselines they tested by up to 27% and moves closer to the consistency of independent human reviewers.\n\nThe headline claim that no model class spontaneously offers null hypotheses while humans do so more freely is harder to pin down. The design only collects ratings on LLM outputs. There is no parallel condition in which the same scientists generate ideas from identical contexts and prompts, so the difference cannot be cleanly attributed to model architecture or training rather than prompt wording or post-generation filtering.\n\nResponse rate sits around 5-6%, which raises the usual selection questions even if the authors acknowledge it. Prompt details and any filtering steps also matter a lot for claims about what models \"spontaneously\" do, and those are not fully visible from the abstract.\n\nThe work is aimed at groups building scientific reward models or automated hypothesis generators. Readers who need large human preference data on research ideas will find the dataset useful even if they disagree with some interpretations.\n\nIt deserves peer review. The empirical volume and the reward-model result are substantial enough to justify referee attention, provided the baseline issue is addressed in revision.","headline":"Large rating dataset and post-trained reward model are the real assets; the null-hypothesis contrast lacks a matched human generation arm.","tokens_in":2448,"tokens_out":362,"would_cite":false,"duration_ms":15024,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Large language models fail to spontaneously propose null hypotheses when generating scientific ideas.","keywords":["large language models","scientific discovery","null hypotheses","hypothesis generation","expert evaluation","AI collaboration","idea novelty"],"falsifier":"A new model class that, in a blinded replication using fresh preprints, proposes null hypotheses at rates comparable to the human authors would falsify the central claim.","tokens_in":2744,"feed_emoji":"🤖","tokens_out":605,"duration_ms":16785,"temperature":0.7,"pith_summary":"The paper conducts the largest evaluation to date in which working scientists judge ideas generated by LLMs from the context of their own recent preprints. It finds that models either collapse to similar ideas or explore wider spaces without ever suggesting null hypotheses, a move human scientists make more freely. Scientists consistently favor probable ideas over novel ones, rate LLM outputs more harshly in pluralistic fields, and show only weak agreement with automated evaluators. A reward model trained on the collected ratings narrows the gap to human inter-rater consistency.","feed_headline":"LLMs avoid proposing null hypotheses in science","feed_subtitle":"Largest scientist rating study shows models cluster ideas narrowly and need human input to diverge or negate.","key_machinery":"Spontaneous proposal of null hypotheses as a marker of the ability to diverge or negate within a hypothesis space.","core_discovery":"In ratings from 6,749 scientists on 25,139 LLM-generated ideas drawn from 121,640 preprints, no model class proposes null hypotheses on its own. Non-reasoning models produce narrow clusters of similar ideas while reasoning models range more widely, yet both avoid negation. Scientists reward resemblance to their own work and probability of being true over novelty, with social scientists showing greater risk tolerance; automated judges align only weakly with these expert assessments.","pith_inferences":["Explicit training for contradiction or falsification may be needed before models can reliably explore negation in hypothesis generation.","The performance gap in pluralistic fields points to a broader limit on AI handling interpretive or theory-evolving domains.","Reward models tuned to human ratings could serve as scalable proxies for expert review in early-stage idea filtering."],"forward_implications":["LLM outputs and judgments in science require ongoing human grounding to compensate for limited imaginative divergence.","Post-training a reward model on human ratings improves capture of field-specific tastes by up to 27 percent over prior automated evaluators.","Social scientists tolerate riskier ideas more than life scientists, and senior social scientists apply the strictest standards.","Retrieval augmentation and persona prompting produce only marginal gains in alignment with expert judgment."],"fun_headline_variants":["LLMs never suggest null hypotheses on their own","Models cluster narrowly without spontaneous negation","AI avoids diverging or negating research ideas","No model class proposes null hypotheses freely","LLMs range wider yet skip null hypothesis moves"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The scientists who responded and the preprints they supplied form a representative sample of scientific reasoning without systematic selection bias.","fun_headline_variants_meta":{"raw":{"variants":["LLMs never suggest null hypotheses on their own","Models cluster narrowly without spontaneous negation","AI avoids diverging or negating research ideas","No model class proposes null hypotheses freely","LLMs range wider yet skip null hypothesis moves"]},"model":"grok-4.3","cost_usd":0.004834,"raw_usage":{"total_tokens":2437,"prompt_tokens":792,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":48337000,"prompt_tokens_details":{"text_tokens":792,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1582,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":792,"tokens_out":63,"duration_ms":11167,"temperature":1.0,"reasoning_tokens":1582,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T19:05:19.911656+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new model class that, in a blinded replication using fresh preprints, proposes null hypotheses at rates comparable to the human authors would falsify the central claim.","supporting_citations":[],"review_version":1}