{"id":"5a46736d-f2f1-4fe2-8346-a62df8f6c359","arxiv_id":"2606.04751","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"FALSIFYBENCH shows that LLMs perform better at inductive rule discovery when they actively seek to falsify hypotheses rather than confirm them.","lead":"The paper introduces FALSIFYBENCH, a new benchmark using rule-discovery games modeled on the Wason 2-4-6 task to test how LLMs generate, test, and revise hypotheses from feedback. A smart generalist might read it to understand whether current AI systems can perform the kind of inductive reasoning needed for autonomous scientific work.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Observable negative feedback may not reflect deliberate falsification of an explicit hypothesis","rationale":"Reader's weakest assumption concerns external validity to real science. The load-bearing issue for the stated central claim is instead internal: whether the operationalization of negative testing actually captures the intended construct rather than a downstream correlate.","tokens_in":1667,"tokens_out":287,"duration_ms":21602,"concrete_test":"Re-run a subset of model evaluations with an added prompt step requiring explicit hypothesis statement before each example proposal; recompute the correlation between performance and the rate of proposals that falsify the stated hypothesis (vs. rate of negative feedback alone). If the stated-hypothesis falsification rate shows substantially weaker or null correlation, the measurement of 'negative testing' is not load-bearing.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim attributes performance gains to 'capacity for negative testing' (actively seeking to falsify hypotheses). In the iterative proposal-feedback loop, this is inferred from the proportion or pattern of examples that receive negative labels. However, without models being required to state their current hypothesis before each proposal, negative feedback could arise from broad exploration, hypothesis revision after the fact, or even lucky sampling rather than targeted disconfirmation of a held belief. The turn-level analysis identifies failure patterns but does not isolate whether the negative instance was generated to test a specific hypothesis.","agreement_with_reader":"disagree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces FALSIFYBENCH, a benchmark framework inspired by the Wason 2-4-6 task, in which LLMs act as agents that iteratively propose examples to discover hidden semantic rules and receive binary feedback. It evaluates 12 models across families and scales, finding that reasoning models outperform instruction-tuned models (though none approach optimal performance), that success is primarily driven by the capacity for negative testing (actively seeking disconfirming evidence), and that turn-level analysis reveals identifiable failure patterns in hypothesis-space navigation.","tokens_in":1758,"tokens_out":570,"duration_ms":11244,"significance":"If the central empirical patterns hold, the work supplies a controlled, falsification-oriented benchmark that isolates hypothesis generation, evidence gathering, and belief revision—elements relevant to scientific discovery agents. The emphasis on negative testing as a performance driver offers a concrete, testable hypothesis for future agent design. The absence of machine-checked proofs or parameter-free derivations is expected for an empirical benchmark paper; the contribution lies in the task design and the fine-grained behavioral analysis.","major_comments":[{"comment":"§4 (or the section presenting the turn-level analysis): the claim that 'the primary driver of success is the capacity for negative testing' is load-bearing yet rests on post-hoc inference from the proportion and pattern of negative-labeled proposals. Without an explicit requirement that models state their current hypothesis before each proposal, negative feedback could equally result from broad exploration, post-hoc revision, or sampling variance rather than targeted disconfirmation of a held belief. A concrete test (e.g., correlation between pre-proposal hypothesis statements and subsequent negative examples) is needed to isolate the mechanism.","section":"§4 (turn-level analysis)"},{"comment":"Evaluation protocol (methods section describing the iterative loop): the benchmark does not record or condition on whether a model articulates an explicit hypothesis before proposing an example. This design choice makes it impossible to distinguish deliberate falsification attempts from other sources of negative feedback, directly weakening the causal attribution in the abstract and §5.","section":"Methods / Evaluation protocol"}],"minor_comments":[{"comment":"Figure 3 (or the figure showing model performance by negative-testing rate): axis labels and legend should explicitly state whether the x-axis is the fraction of negative proposals or a normalized score; current presentation risks conflating volume of negative feedback with strategic intent.","section":"Figure 3"},{"comment":"Table 1 (model list): include the exact prompting template and temperature settings used for each model family so that the negative-testing metric can be reproduced.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments, which help clarify the evidential basis for our claims about negative testing. We respond to each major comment below.","responses":[{"response":"We agree that the evidence linking negative testing to success is correlational, derived from the observed relationship between the rate of negatively labeled proposals and task success in the turn-level analysis. The benchmark follows the standard Wason 2-4-6 protocol without mandating explicit hypothesis statements to preserve ecological validity. We will revise §4 to more explicitly define negative testing via proposal labels, discuss alternative explanations such as exploration variance, and qualify the abstract and §5 claims as correlational rather than mechanistic. A direct correlation test with pre-proposal hypotheses would require a protocol change and is noted as future work.","revision_made":"partial","referee_comment":"[§4 (turn-level analysis)] §4 (or the section presenting the turn-level analysis): the claim that 'the primary driver of success is the capacity for negative testing' is load-bearing yet rests on post-hoc inference from the proportion and pattern of negative-labeled proposals. Without an explicit requirement that models state their current hypothesis before each proposal, negative feedback could equally result from broad exploration, post-hoc revision, or sampling variance rather than targeted disconfirmation of a held belief. A concrete test (e.g., correlation between pre-proposal hypothesis statements and subsequent negative examples) is needed to isolate the mechanism."},{"response":"The iterative loop is deliberately unconditioned on explicit hypotheses to study models' spontaneous inductive behavior. We will update the methods section to document this design decision and its consequences for interpreting negative feedback. This revision will temper causal language in the abstract and §5 to emphasize observed correlations between negative proposal rates and performance.","revision_made":"yes","referee_comment":"[Methods / Evaluation protocol] Evaluation protocol (methods section describing the iterative loop): the benchmark does not record or condition on whether a model articulates an explicit hypothesis before proposing an example. This design choice makes it impossible to distinguish deliberate falsification attempts from other sources of negative feedback, directly weakening the causal attribution in the abstract and §5."}],"tokens_in":1407,"tokens_out":468,"duration_ms":31565,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"FALSIFYBENCH adapts the Wason 2-4-6 task into an iterative game where LLMs propose examples to discover hidden semantic rules and receive feedback. The main reported finding is that models producing more negative examples succeed more often, with reasoning models outperforming instruction-tuned ones and a turn-level breakdown of failure patterns.\n\nThe paper is new in packaging this setup as a named benchmark and running it across 12 models from different families and scales. The turn-level analysis of how models navigate confirming versus disconfirming evidence goes beyond aggregate accuracy, which is a concrete step forward for understanding the process.\n\nThe soft spot is the central claim about negative testing as the driver. The paper infers this from patterns in the examples that receive negative labels, but without models stating an explicit hypothesis before each proposal, those negatives could result from broad exploration or post-hoc revision rather than targeted disconfirmation. The stress-test concern holds up on the abstract, and the description does not show they isolated the mechanism. Generalization beyond this artificial game to scientific discovery is also untested, though the authors keep the scope to evaluation.\n\nThis is for researchers who build or assess LLM agents for reasoning tasks. The benchmark and the empirical patterns are concrete enough to deserve referee time.","headline":"FALSIFYBENCH adapts the Wason task into an iterative benchmark and links better LLM performance to negative testing, though the evidence for deliberate falsification is indirect.","tokens_in":2231,"tokens_out":335,"would_cite":false,"duration_ms":37974,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLMs that seek to falsify hypotheses outperform those seeking confirmation in rule discovery games.","keywords":["inductive reasoning","large language models","hypothesis testing","rule discovery","falsification","negative testing","scientific discovery","belief revision"],"falsifier":"A side-by-side test measuring whether the same models that excel at negative testing on FALSIFYBENCH also generate better hypotheses when given real experimental data from a scientific domain.","tokens_in":2569,"feed_emoji":"🧪","tokens_out":421,"duration_ms":20044,"temperature":0.7,"pith_summary":"This paper introduces FALSIFYBENCH, a benchmark where language models discover hidden semantic rules by proposing examples and revising beliefs from feedback. It tests twelve models across families and scales, showing reasoning models generally surpass instruction-tuned ones. The main driver of success is negative testing, in which models actively propose examples that would disprove their current hypothesis rather than only seek confirming cases. Failures trace to identifiable patterns in how models explore the space of possible rules.","feed_headline":"Negative testing drives LLM success in rule discovery","feed_subtitle":"Models seeking disconfirming evidence outperform confirmation seekers, yet none reach optimal performance on semantic rule tasks.","key_machinery":"FALSIFYBENCH, an iterative evaluation game in which agents propose examples to uncover hidden semantic properties and receive feedback indicating whether each example fits the rule.","core_discovery":"The central claim is that the capacity for negative testing drives success on the FALSIFYBENCH task. Models that generate examples intended to falsify their hypotheses discover the hidden semantic rules more reliably than models focused on confirmation. Although reasoning models outperform instruction-tuned models, no system reaches optimal performance, and turn-level analysis links most failures to specific navigation patterns in the hypothesis space.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLM success hinges on negative testing","Falsifying hypotheses aids rule discovery","No models optimal in semantic rule games","Negative testing outperforms confirmation in LLMs","Reasoning models better at falsifying rules"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That performance on this semantic rule-discovery game meaningfully predicts an LLM's ability to engage in hypothesis-driven reasoning for real scientific discovery tasks.","fun_headline_variants_meta":{"raw":{"variants":["LLM success hinges on negative testing","Falsifying hypotheses aids rule discovery","No models optimal in semantic rule games","Negative testing outperforms confirmation in LLMs","Reasoning models better at falsifying rules"]},"model":"grok-4.3","cost_usd":0.004394,"raw_usage":{"total_tokens":2181,"prompt_tokens":632,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":43937000,"prompt_tokens_details":{"text_tokens":632,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1489,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":632,"tokens_out":60,"duration_ms":17543,"temperature":1.0,"reasoning_tokens":1489,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T06:06:37.959869+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A side-by-side test measuring whether the same models that excel at negative testing on FALSIFYBENCH also generate better hypotheses when given real experimental data from a scientific domain.","supporting_citations":[],"review_version":1}