{"id":"e589b83c-6334-4c0c-a9c1-ba1aec0440ee","arxiv_id":"2607.27794","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AI for physics has moved from explicit equation discovery to black-box prediction, a trajectory the authors argue reverses the historical progression of human physics, leaving the invention of new mathematical frameworks ('Category C') untouched by AI.","lead":"AI physics tools are getting better at prediction while their outputs are getting harder to understand. This Perspective argues that AI is following human physics discovery backwards, and that a missing skill—posing the right questions and inventing new principles—must be taught before AI can propose a paradigm-level theory like quantum gravity.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'reverse trajectory' claim rests on a hand-picked sample of AI-for-physics successes; a systematic corpus could reverse the trend and undermine the paper's central prediction.","rationale":"The reader's weakest assumption—that the sample of 'most visible' successes is representative—is squarely on target. I agree that selection bias is the primary threat to the paper's empirical trajectory claim. My additional focus is the operational vagueness of Category C, which makes the 'no AI has achieved Category C' claim nearly unfalsifiable and thus less informative than it appears. These are related: if the taxonomy is not crisp, even a systematic corpus cannot cleanly test the claim. The paper's historical caveats and its explicit acknowledgment that the reversal is not a universal law show intellectual honesty, but they also underscore that the central generalization is currently a perspective rather than a well-supported empirical finding. The proposed concrete test—a structured corpus with independent classification—would remedy the selection-bias concern and sharpen the Category C definition. Since the reader already assigned CONDITIONAL, my assessment does not change the verdict; the paper remains acceptable as a Perspective but with the same conditions.","tokens_in":15815,"tokens_out":2979,"duration_ms":32673,"concrete_test":"Construct a systematic corpus of AI-for-physics discoveries from 2010–2025 (e.g., top-cited papers from arXiv physics + major ML/Nature/Science venues, or a curated list from expert survey). Have independent raters classify each milestone by (a) whether the output is an explicit equation/law, a black-box predictor, or a new mathematical framework/principle, and (b) the degree of human explainability. Plot the proportion of theory-producing versus black-box outputs over time. If the proportion of explicit-equation or framework-producing works is not declining—or if recent years include a significant share of Category B/C outputs—the 'reverse trajectory' claim is an artifact of example selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that AI discoveries in physics are moving 'in reverse order' from theory building to black-box prediction—is supported by a small set of 'most visible' examples (AlphaFold, GraphCast, GNoME) rather than by a systematic survey. The authors acknowledge the reversal 'is not a universal law' and that the history is 'coarse,' but the load-bearing prediction that better prediction alone will not yield paradigm-level theories depends on this trend being real and representative. The paper's own references include recent symbolic-regression and LLM-based analytical works (e.g., refs 18, 19, 44, 45, 106) that produce explicit equations or theories, so the claim that such efforts are 'increasingly overshadowed' is asserted, not demonstrated. A different sample—including early neural-network predictors (e.g., 1990s NETtalk-style physics surrogates) or recent interpretable ML and automated theorem-proving successes—could weaken or reverse the stated trajectory. Additionally, the 'no AI has made a Category C discovery' claim is not testable as stated because Category C is not operationally defined: what counts as 'successful use of a new principle' or 'invention of a novel mathematical framework' is left vague, making the claim partly definitional and resistant to counterexample.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This Perspective argues that AI contributions to physics discovery are following the historical trajectory of human physics in reverse epistemic order: early AI milestones extracted explicit symbolic laws (BACON, Eureqa, SINDy, AI Feynman), while the most prominent recent systems (AlphaFold, GraphCast, GNoME) are high-accuracy black-box predictors. The paper introduces an A/B/C taxonomy of physics discoveries (new solutions/capabilities, new equations/laws, new mathematical frameworks/principle-based theories) and asserts that no AI has yet made a Category C discovery. It concludes that continued progress in prediction alone will not yield paradigm-level theories such as quantum gravity, and proposes a research agenda to give AI systems skills for posing questions, inventing principles, conducting abductive searches over provisional axioms, and using symbolic-computing theory-building sandboxes, including a 'Reverse ITP'. The paper is framed as a Perspective and includes explicit caveats that the historical sketch is approximate and the reversal is not a universal law.","tokens_in":16035,"tokens_out":8025,"duration_ms":68949,"significance":"The paper addresses a timely and important question: whether current AI-for-science paradigms, centered on predictive accuracy, can lead to fundamental theoretical breakthroughs in physics. If the central claim is correct, it implies a need to reallocate research effort toward principle discovery, symbolic reasoning, falsifiable theory generation, and formalized abduction. The paper's strengths are its clear historical framing, a useful (if underspecified) taxonomy, a concrete falsifiable proposal (e.g., the '1911 cutoff' test and artificial-world tasks), and several constructive technical suggestions (theory-building sandboxes, symmetry abduction). However, the central empirical assertions rest on a hand-selected sample and non-operational categories; these need to be tightened before the paper's conclusions can be fully endorsed.","major_comments":[{"comment":"The reverse-trajectory claim is supported by a small set of 'most visible' AI successes (AlphaFold, GraphCast, GNoME; refs 21–24, 62) rather than a systematic survey. The paper's own references include recent symbolic-regression and analytical theory-building work (refs 18, 19, 44, 45, 106), so the assertion that black-box systems 'increasingly overshadow' these efforts is asserted, not demonstrated. The paper correctly labels the sketch 'approximate' and the reversal 'not a universal law', but these qualifications do not replace evidence. A bibliometric or corpus-based analysis, or at least a falsifiable sampling protocol, is needed to establish that the trend is real and representative. This is load-bearing because the central prediction depends on the trend.","section":"§2, Fig. 1"},{"comment":"The claim 'no AI has made a discovery of Category C' is not operationally testable as stated. Category C is defined only by examples ('successful use of a new principle or the invention of a novel mathematical framework'), with no criteria for 'successful use', 'new principle', or 'novel framework'. Because the table is constructed with all current AI outputs placed in Categories A and B, the absence of Category C is partly a consequence of the classification. The paper itself says the categories are not strict or mutually exclusive, which makes the universal negative even harder to evaluate. In addition, Table 1 lists the Standard Model and Higgs mechanism as human Category B, yet Section 5 cites the Higgs mechanism as an example of principle-guided abductive discovery; this internal inconsistency blurs the taxonomy. The authors should provide operational criteria and a systematic audit","section":"§4, Table 1; §5"},{"comment":"A key demonstration for the proposed roadmap (inverting the EFT workflow to search for candidate quantum-gravity theories) is supported by reference [114], which is an unpublished 'in preparation' self-citation. This is not verifiable by readers or reviewers. The authors should either make the work available (e.g., as a preprint or extended supporting information) or describe the method and results in sufficient detail to be evaluated. As it stands, the feasibility of the proposed direction is partly grounded in an inaccessible source.","section":"§6, ref. [114]"},{"comment":"The paper's sample mixes AI-for-science successes (AlphaFold for protein folding, GraphCast for weather, GNoME for materials) with AI-for-physics work. These are impressive predictive applications but are not discoveries of physical laws in the sense used in the human trajectory (Kepler, Maxwell, Einstein). Including them as the 'frontier' of AI-for-physics biases the sample toward prediction. The authors should either restrict the claim to physics-specific AI systems or explicitly justify why cross-domain examples are representative of the center of gravity of AI-for-physics.","section":"§2–§3"},{"comment":"The paper acknowledges that the absence of new paradigm-level theories may be due to physics itself (the end of simple, testable revolutions) rather than to AI limitations, but it leaves this possibility largely unintegrated. If the bottleneck is the intrinsic complexity or experimental untestability of remaining theories, then the reverse trajectory does not explain or predict the absence of AI Category C discoveries. The authors should state what evidence would distinguish the 'AI bottleneck' from the 'physics bottleneck' hypotheses—their artificial-world and '1911 cutoff' proposals are a good start—and explicitly condition the central claim on the outcome of such tests.","section":"§6, 'Is the Bottleneck AI, or Physics Itself?'"}],"minor_comments":[{"comment":"The figure would benefit from labeled axes and a legend distinguishing the human and AI timelines. The caption refers to 'arrows summarizing example contributions' but the arrows are not individually identified.","section":"Fig. 1"},{"comment":"The phrase 'a certain ansätze' should be 'a certain ansatz' (ansätze is the plural form).","section":"§4, Historical Examples"},{"comment":"The heading 'Propose questions rather than answers' would read more naturally as 'Proposing questions rather than answers.'","section":"§6"},{"comment":"The caveat is valuable, but its connection to the surrounding argument could be made explicit: it currently interrupts the flow between the Bunge discussion and the alignment paragraph.","section":"§3, 'Ipcha Mistabra'"},{"comment":"The phrase 'reverse epistemic order' is used interchangeably with 'reverse order' and 'reverse trajectory'; consider defining the intended meaning of 'epistemic' early to avoid ambiguity.","section":"§2"}],"recommendation":"major_revision","confidential_remarks":"This is a Perspective, so the standards for empirical support differ from a research article; nevertheless, the central empirical claim is used to justify a research agenda. I recommend major revision rather than rejection because the paper's thesis is plausible and the identified gaps are fixable. The use of an unpublished self-citation [114] is a particular concern. The editor may also want to consider whether the scope (physics.hist-ph) fits a paper that is as much a position paper on AI as a history/philosophy analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline first: this paper is a Perspective with a genuinely new framing — AI-for-physics looks like the history of physics run in reverse, from laws to pattern prediction — and a concrete proposal (Reverse ITP) for what to do about it. The framing is not systematically proven, and the 'no AI has made a Category C discovery' claim is partly built into the classification. Read as a proposal, it is valuable; read as a measured trend, it is under-supported.\n\nWhat is actually new: the reverse-epistemic-trajectory idea, the A/B/C taxonomy (predictions / equations / new frameworks), and the suggestion of formalizing abduction with a Reverse ITP. The paper also does several things well. It tells the history with useful caveats: Ptolemaic astronomy had principles, Maxwell's early mechanics were scaffolding, categories blur, and the authors even flag the 'Ipcha Mistabra' possibility that some discoveries may be best made by AIs in ways humans cannot interpret. That honesty strengthens the Perspective.\n\nThe soft spots are real but proportionate. The central empirical claim is a selected-sample argument. 'Most visible' successes like AlphaFold and GraphCast are contrasted with early symbolic regression, and the trend is asserted rather than measured. The paper itself says the reversal is not a universal law and the history is coarse, but the prediction that prediction alone won't produce quantum gravity leans on that trend being representative. A different sample — recent interpretable symbolic regression, LLM-derived analytic amplitudes, theorem-proving successes — could soften or reverse the arrow. Second, Category C is not operationally defined. 'Successful use of a new principle' is vague, so 'no AI has done it' is hard to falsify. Third, a load-bearing feasibility example, the EFT inversion that searches for UV theories from an IR Lagrangian, is an unpublished self-citation [114]. None of this sinks the paper, but all three are addressable and should be part of a revision.\n\nWho should read it: anyone thinking about where AI-for-physics should invest — funding agencies, ML researchers, and theorists who care about theory discovery. It deserves a serious referee. My recommendation: send it to peer review, and ask the authors to operationalize Category C, survey a broader set of AI milestones, and either release the EFT inversion work or cite a public version.","headline":"A provocative, self-aware Perspective that reframes AI-for-physics as running history backward; the empirical trend is under-built, but the Reverse ITP concept makes it worth a serious look.","tokens_in":16585,"tokens_out":2692,"would_cite":true,"duration_ms":26864,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI-for-physics is running the history of discovery in reverse","keywords":["AI for physics","scientific discovery","paradigm shift","abductive reasoning","symbolic regression","black-box prediction","theory building","philosophy of physics"],"falsifier":"A published, replicated AI-generated physics result that introduces a new mathematical framework or symmetry principle later validated by experiment would falsify the claim that no AI has made a Category C discovery.","tokens_in":15649,"feed_emoji":"🔬","tokens_out":4029,"duration_ms":35797,"temperature":0.7,"pith_summary":"The paper argues that the most visible AI contributions to physics are recapitulating the history of physics backwards. Human physics moved from pattern prediction to phenomenological laws to principle-based universal theories, while AI-for-physics has moved from explicit equation discovery (symbolic regression) to powerful but opaque neural predictors. As a result, current AI systems excel at induction and deduction but cannot perform the abductive, principle-guided reasoning behind general relativity or the Standard Model. The paper classifies discoveries into three categories — new solutions, new equations, new mathematical frameworks — and claims no AI has yet made a Category C discovery. Its practical thesis is that better prediction alone will not yield paradigm-level theories; AI needs to propose principles, pose questions, and search over provisional axioms.","feed_headline":"AI physics discovery runs history in reverse","feed_subtitle":"Predictive black boxes are outpacing principle-based theory building — and no AI has yet invented a new math framework.","key_machinery":"The central machinery is a three-tier classification of physics discoveries (Category A: new solutions or capabilities; Category B: new equations or laws; Category C: new mathematical frameworks or principles), together with a historical reversal thesis that maps human physics' progression — pattern prediction, phenomenology, principle-based theory — onto AI's chronological trajectory in reverse. The classification locates the gap: AI has achieved A and B, never C. The concrete mechanism proposed to close the gap is a 'Reverse ITP': a formal system that organizes abductive search over provisional axioms rather than proving consequences from fixed axioms, using contradictions as a loss signal","core_discovery":"On the authors' own terms, the central claim is that the center of gravity of AI for physics discovery has shifted from discovering explicit symbolic laws to building black-box predictors, reversing the epistemic order of human physics. Consequently, while AI has produced Category A discoveries (predictions, devices) and Category B discoveries (equations, laws), no AI system has produced a Category C discovery — the successful use of a new principle or the invention of a novel mathematical framework. The paper does not claim this is impossible; it claims the field is currently optimized away from it. It argues that the missing skill is not creativity but scientific taste: choosing which nove","pith_inferences":["The reversal thesis may be partly a selection artifact: a broader sample that includes recent symbolic-regression and automated-theorem-proving successes from the same period could weaken or reverse the apparent trajectory.","A testable extension is corpus-level measurement of AI-for-physics outputs over time — coding each contribution by category and by whether it outputs explicit equations — to see whether the trend is robust or depends on which successes are counted.","The 'Reverse ITP' concept suggests a new kind of AI benchmark: not solving problems, but generating provisional axioms whose falsifiable consequences are novel and testable.","The paper's own caveat about non-human-interpretable mathematics implies that human explainability may become a bottleneck to paradigm shifts rather than a necessity; if AI develops powerful non-human-interpretable frameworks, the field's target might shift to detecting and validating discoveries through compression measures."],"forward_implications":["If the reversal thesis holds, continued investment in black-box prediction alone will not produce paradigm-level theories like quantum gravity.","AI systems that can propose and test provisional axioms (reverse ITPs) would enable automated generation of falsifiable theory candidates.","A '1911 cutoff' test — withholding general relativity from training data and seeing whether AI rediscovers it — becomes a meaningful benchmark for principle-level discovery.","If AI can discover simple theories in artificial worlds but not real-world ones, the bottleneck may be physics itself, not AI architecture.","The three-category taxonomy gives the field a concrete target: explicitly aiming for Category C discovery as a stated research goal."],"fun_headline_variants":["AI physics: prediction beats principle","AI discovery runs physics history backwards","Black-box AI can't pose the big questions","AI masters prediction, misses theory","No AI has proposed a new physics principle"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim depends on the sample of 'most visible and influential' AI-for-science successes being representative of the field's center of gravity; if that sample is biased, the reverse trajectory may be a selection artifact rather than a real trend.","fun_headline_variants_meta":{"raw":{"variants":["AI physics: prediction beats principle","AI discovery runs physics history backwards","Black-box AI can't pose the big questions","AI masters prediction, misses theory","No AI has proposed a new physics principle"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000306,"raw_usage":{"total_tokens":1599,"prompt_tokens":759,"completion_tokens":840,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":779}},"tokens_in":503,"tokens_out":840,"duration_ms":7701,"temperature":1.0,"reasoning_tokens":779,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T01:15:35.393110+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A published, replicated AI-generated physics result that introduces a new mathematical framework or symmetry principle later validated by experiment would falsify the claim that no AI has made a Category C discovery.","supporting_citations":[],"review_version":1}