{"id":"25f9cf1a-efba-4d29-b920-f3543e4e4b24","arxiv_id":"2608.06680","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"PhysMat AI organizes physics-AI integration into five complementary roles and outlines a roadmap from physics-aware to physics-autonomous AI.","lead":"This perspective paper proposes a unified framework, PhysMat AI, that integrates physics into materials AI through five roles: prior knowledge, descriptors, constraints, verifiers, and infrastructure. It argues that materials discovery needs physics-grounded reasoning rather than pure data fitting, using catalysis, solid-state batteries, and hydrogen storage as examples.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The five-role framework is a useful organizing proposal, but the central reliability claim rests on selected self-authored examples rather than a controlled comparison against unconstrained AI; this weakens evidential force without breaking internal logic.","rationale":"The reader's weakest assumption focused on the completeness and transferability of the five-role taxonomy. That is a legitimate structural concern for a unifying framework. My stress-test identifies a related but distinct load-bearing point: the paper's central claim that physics grounding improves reliability and extrapolation is asserted rather than empirically demonstrated. Section 2 frames PhysMat AI as enabling reliable discovery, and Section 3.1.1 promises enhanced extrapolation, but no controlled comparison is provided. The examples are almost entirely the authors' own prior publications, which raises the risk that the taxonomy is retrospective rather than predictive. Section 4.2 acknowledges that extrapolation and physical inconsistency remain unresolved, so the framework is explicitly presented as a roadmap rather than a proven system. For a Perspective article, this is an acceptable genre limitation: the contribution is a shared vocabulary and organizing framework, not a new empirical result. Therefore I do not think the concern warrants changing the reader's ACCEPT verdict. However, if the authors intend the reliability claims to be taken as established rather than aspirational, a controlled benchmark would substantially strengthen the paper. My agreement with the reader is partial because the taxonomy-completeness issue and the evidence-for-reliability issue are different soft spots, though both point to the gap between the framework's ambition and its demonstrated support.","tokens_in":23474,"tokens_out":4670,"duration_ms":48289,"concrete_test":"Select one end-to-end task from the paper's own infrastructure, such as DigHyd-based hydrogen-storage candidate screening or DigCat-based CO2RR catalyst selection. Run two pipelines: (i) a standard data-driven ML baseline using composition or structure features only, and (ii) the same ML model augmented with the paper's physics-based descriptors, constraints, and verifier filtering. Evaluate both on held-out or newly synthesized candidates using a predefined reliability metric, such as the fraction of candidates that pass DFT or experimental verification, or the out-of-distribution prediction error. If pipeline (ii) does not lower error or raise the pass rate, the central reliability claim is empirically unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2 asserts that PhysMat AI 'enables reliable, interpretable, and verifiable materials discovery,' and Section 3.1.1 claims physical priors yield 'enhanced extrapolation capability beyond the available data.' However, the paper presents no quantitative experiment comparing a PhysMat AI pipeline against a data-driven baseline on the same discovery task. The illustrative examples are predominantly drawn from the authors' own prior work (e.g., refs. [33], [36], [70], [99], [112]) and selected to fit the five-role taxonomy. Section 4.2 itself concedes that 'weak extrapolation, physical inconsistency, uncertainty quantification' remain major obstacles. Thus the load-bearing assumption that embedding physics through the five roles actually improves reliability and extrapolation is not demonstrated within the manuscript. This does not invalidate the Perspective as a proposal, but it means the central claim is a conjecture supported by case illustrations rather than by evidence that the framework outperforms simpler alternatives.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This Perspective proposes 'Physics-Grounded Materials AI' (PhysMat AI), a unifying framework in which physical knowledge enters materials AI through five roles: prior knowledge, descriptors, constraints, verifiers, and infrastructure. The authors illustrate each role with examples from electrocatalysis, solid-state electrolytes, and hydrogen storage, and they sketch a three-stage roadmap from physics-aware AI to physics-reasoning AI and ultimately physics-autonomous AI. The paper is explicitly programmatic: its aim is to organize existing work around a taxonomy and to motivate a closed-loop discovery workflow, not to introduce new data or models.","tokens_in":23721,"tokens_out":7611,"duration_ms":74332,"significance":"If taken as a programmatic Perspective, the paper makes a useful and timely contribution: it names a taxonomy that is easy to adopt, connects it to concrete databases and workflows (e.g., DigCat, DigBat/DDSE, DigHyd, DIVE), and is explicit about remaining obstacles. The figures and examples are drawn from real published results, and the discussion of failure feedback, provenance-aware data infrastructure, and physical verification is constructive. The main limitation is that the reliability and extrapolation claims are supported by illustrative case studies rather than by controlled comparisons; this weakens the evidential weight but does not invalidate a Perspective whose value lies in synthesis and agenda-setting.","major_comments":[],"minor_comments":[{"comment":"The statements 'PhysMat AI ... enables reliable, interpretable, and verifiable materials discovery' (Section 2) and 'enhanced extrapolation capability beyond the available data' (Section 3.1.1) are stronger than what Section 4.2 can support when it lists weak extrapolation, physical inconsistency, and uncertainty quantification as 'major obstacles.' I suggest wording them as design goals or as capabilities that the framework is intended to provide, which would remove an apparent internal tension without weakening the proposal.","section":"Section 2 / Section 3.1.1"},{"comment":"The sentence 'Together, these five layers collectively span the full lifecycle of materials intelligence' asserts a completeness property, but the manuscript gives no criterion for completeness and does not discuss alternative taxonomies (for example, uncertainty quantification or synthesis-aware search could arguably be additional roles). Please either add a rationale for why these five roles are sufficient or soften 'collectively span' to 'provide a practical decomposition of.'","section":"Section 2"},{"comment":"Reference [41] and reference [63] are the same paper (Witman et al., Adv. Funct. Mater. 2024, 34, 2411763), as are references [38] and [61] (Hirscher et al., J. Alloys Compd. 2020, 827, 153548); these should be consolidated to avoid duplicate bibliography entries.","section":"References"},{"comment":"Reference [6] is attributed to 'J. Pablo'; this appears to be Juan de Pablo and should be spelled accordingly, and reference [101] gives 'Proc. Natl. Acad. Sci. 2002, 9, 12562,' which should be volume 99.","section":"References"},{"comment":"The text states that the initial DDSE release contained 678 performance records and that the current DigBat platform integrates 3725 experimental records, while Figure 7d is captioned 'updated to May 2023'; please clarify which of these numbers corresponds to Figure 7d so that the reader does not conflate the two versions.","section":"Section 3.5.2"},{"comment":"The 5% agreement between predicted and measured LiScO2 activity is a strong illustrative result, but because no unconstrained baseline is shown, the sentence 'confirming that experimental verification is essential' should be phrased as an interpretation rather than as a controlled demonstration.","section":"Section 3.4.2"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is an appropriate Perspective for the journal. Its illustration set is heavily drawn from the authors' own recent output; this is a reasonable choice for a perspective by a group that built many of the databases, but the editors may wish to ensure that the framing does not read as a product demonstration. The central reliability claim should be watched at proof stage; I am comfortable with a minor-revision decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The two things to know: this is a Perspective, not a results paper, and its contribution is a clean five-role taxonomy for embedding physics into materials AI — prior, descriptor, constraint, verifier, infrastructure — plus a three-stage roadmap from physics-aware to physics-autonomous AI. The taxonomy is genuinely useful as a shared vocabulary, and the paper applies it consistently across catalysis, solid-state electrolytes, and hydrogen storage. If you work in materials informatics, this is a reasonable orienting piece and citable for the framework.\n\nWhat it does well: it synthesizes scattered practices into a coherent structure, the three application areas are covered competently, and the figures tie the framework to concrete published results. The authors are also honest in Section 4.2 about open challenges: weak extrapolation, physical inconsistency, and uncertainty quantification are listed as major obstacles. That section is the most credible part of the paper.\n\nWhere the soft spots are: the framing overclaims relative to the evidence. Sections 2 and 3.1.1 state that PhysMat AI enables reliable, interpretable, verifiable discovery with enhanced extrapolation, but the manuscript contains no controlled comparison against unconstrained AI. The examples are illustrative, not tests. For a Perspective this is not fatal, but the causal language should be tempered: the reliability benefit is a hypothesis, not a demonstrated result. The stress-test note is fair on this point, and Section 4.2 itself partially concedes it. A secondary issue is the completeness claim: the five layers are asserted to span the full lifecycle, but there is no argument for why these five and not others. Uncertainty quantification, synthesis-aware search, or human-in-the-loop design are plausible additions that receive no discussion. The heavy reliance on the authors' own prior work (DigCat, DigBat, DigHyd, DIVE, etc.) is natural for illustration but limits the independent evidentiary weight; the results are real, so I would not call it circularity, just a narrow base.\n\nWho this is for: readers who want a shared vocabulary for physics-informed ML in materials discovery, and anyone writing a review or starting a research program in this area. It is not a source of new empirical evidence, but it is a fair map of current practice.\n\nRecommendation: send it to peer review. A serious referee can push the authors to soften the reliability claims and address the completeness question. This deserves referee time, not a desk reject.","headline":"A well-organized perspective that gives the field a useful vocabulary, but its central reliability claim is programmatic rather than demonstrated.","tokens_in":24186,"tokens_out":2108,"would_cite":true,"duration_ms":22105,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Physics should enter materials AI as prior, descriptor, constraint, verifier, and infrastructure.","keywords":["Physics-Grounded Materials AI","Materials Discovery","AI Agents","Autonomous Scientific Research","catalysis","solid-state electrolytes","hydrogen storage","physics-informed machine learning"],"falsifier":"A head-to-head study in a new material family where a plain data-driven pipeline matches or beats a PhysMat AI-style pipeline in the fraction of experimentally validated working candidates, while also extrapolating better to out-of-distribution compositions, would refute the claim that physics grounding is necessary for reliable discovery. A systematic literature classification showing substantial physics-informed methods that fit none of the five roles would refute the taxonomy's completeness.","tokens_in":23273,"feed_emoji":"⚛️","tokens_out":7414,"duration_ms":62612,"temperature":0.7,"pith_summary":"This Perspective argues that the next stage of materials discovery by AI should not be built on statistical correlation alone. It proposes PhysMat AI, a unifying framework in which physics enters at five distinct points—as prior knowledge, as descriptors, as constraints, as verifiers, and as data infrastructure—so that models predict, reason, and validate within physically feasible bounds. The authors demonstrate the framework through catalysis, solid-state electrolytes, and hydrogen storage, then lay out a roadmap from physics-aware to physics-reasoning to physics-autonomous AI. If the framework is right, reliable materials discovery becomes a matter of engineering the right physical interfaces around the model, not of chasing larger datasets alone.","feed_headline":"Physics in five roles steers AI to real materials","feed_subtitle":"Priors, descriptors, constraints, verifiers, and infrastructure anchor AI to physically reachable materials.","key_machinery":"The organizing object is a five-layer taxonomy of physics roles in AI—prior, descriptor, constraint, verifier, and infrastructure—together with a closed-loop workflow connecting physical principles, curated databases, AI models and agents, prediction and screening, experimental validation, and feedback. This named 'physics as' architecture functions as a conceptual scaffold that assigns a physical role to every stage of materials discovery and lets AI agents invoke computational tools such as density-functional theory, microkinetic modeling, ab initio molecular dynamics, and metadynamics while reasoning. The taxonomy carries the argument: once the five roles are accepted, materials AI becomes an integration problem rather than only a modeling problem.","core_discovery":"The central claim is that purely data-driven materials AI fails at interpretability, extrapolation, and physical consistency, and that the remedy is a unified architecture in which physical knowledge is embedded throughout the discovery loop: as priors that constrain learning, as descriptors that give features mechanistic meaning, as constraints that restrict reasoning to feasible regions, as verifiers that check predictions with simulation and experiment, and as infrastructure that makes data traceable and condition-aware. On this view, physics is not auxiliary to AI but constitutive of it, transforming AI from an interpolative predictor into a mechanism-guided discovery system. The paper supports this with examples from catalysis, solid-state electrolytes, and hydrogen storage, including closed-loop workflows in which validated and failed results feed back into the model.","pith_inferences":["The five-role taxonomy can be used as an audit checklist: classifying a corpus of physics-informed materials machine-learning papers by these roles would test the claimed completeness and likely expose boundary cases such as uncertainty quantification and synthesis-aware search.","If the framework is adopted, database funding priorities would shift toward negative-result capture and condition-aware metadata, because infrastructure is one of the five load-bearing roles.","Head-to-head benchmarks between physics-free and physics-grounded pipelines on out-of-distribution materials would quantify how much reliability the physics actually buys, a quantity the paper does not compute.","The roadmap implies a natural graduation test: a system that proposes a material, validates it in silico and in the lab, and updates its own knowledge base would count as physics-autonomous AI."],"forward_implications":["Candidates that violate thermodynamic stability, kinetic accessibility, or operating-window constraints should be discarded before experimental testing even if they score well.","Materials databases should record measurement conditions, provenance, and negative and failed results as first-class information, not only successful property values.","AI agents should be judged by whether they invoke the right physical tool at the right time, not by prediction accuracy alone.","The same five-role structure transfers across application domains, so methods developed for catalysis can be repurposed for battery or hydrogen-storage discovery.","The end state is a self-improving closed-loop system in which AI, simulation, and experiment co-evolve into physics-autonomous discovery."],"supporting_citations":[{"why":"Provides the large-scale data-driven screening baseline that motivates the need for physical grounding.","marker":"[7]"},{"why":"Documents that many AI-predicted inorganic compounds lack novelty and experimental credibility, motivating the framework.","marker":"[9]"},{"why":"Supplies the computational hydrogen electrode approach used as a physics prior for catalytic activity.","marker":"[34]"},{"why":"Supplies scaling relations among adsorption energies that serve as prior knowledge and descriptor structure.","marker":"[35]"},{"why":"Establishes d-band center as an electronic-structure descriptor for catalysis.","marker":"[68]"},{"why":"Demonstrates a closed-loop discovery workflow in which experimental measurement of the top-ranked catalyst matches predictions within five percent.","marker":"[70]"},{"why":"Shows integration of electrolyte databases, large language models, and metadynamics to resolve transport mechanisms.","marker":"[99]"},{"why":"Provides the catalysis database infrastructure that supports AI-driven descriptor construction.","marker":"[104]"},{"why":"Provides the solid-state electrolyte database infrastructure preserving temperature-dependent conductivity records.","marker":"[109]"},{"why":"Provides a hydrogen-storage knowledge platform that extracts and organizes graphical performance data for AI use.","marker":"[112]"}],"fun_headline_variants":["Five physics roles anchor AI to real materials","AI needs physics, not just data, for materials","Mechanism-guided AI discovers real materials","Physics-aware AI stays in feasible material spaces"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that the five roles—prior, descriptor, constraint, verifier, and infrastructure—exhaustively cover the ways physics can enter materials AI, but the paper gives no completeness criterion or comparison with alternative partitionings.","fun_headline_variants_meta":{"raw":{"variants":["Five physics roles anchor AI to real materials","AI needs physics, not just data, for materials","Mechanism-guided AI discovers real materials","Physics-aware AI stays in feasible material spaces"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000418,"raw_usage":{"total_tokens":2130,"prompt_tokens":897,"completion_tokens":1233,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":1176}},"tokens_in":513,"tokens_out":1233,"duration_ms":11191,"temperature":1.0,"reasoning_tokens":1176,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:35:07.508938+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A head-to-head study in a new material family where a plain data-driven pipeline matches or beats a PhysMat AI-style pipeline in the fraction of experimentally validated working candidates, while also extrapolating better to out-of-distribution compositions, would refute the claim that physics grounding is necessary for reliable discovery. A systematic literature classification showing substantial physics-informed methods that fit none of the five roles would refute the taxonomy's completeness.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes d-band center as an electronic-structure descriptor for catalysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates a closed-loop discovery workflow in which experimental measurement of the top-ranked catalyst matches predictions within five percent."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows integration of electrolyte databases, large language models, and metadynamics to resolve transport mechanisms."},{"cited_title":"Zhang, X","cited_arxiv_id":null,"evidence_quote":"Provides the catalysis database infrastructure that supports AI-driven descriptor construction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the solid-state electrolyte database infrastructure preserving temperature-dependent conductivity records."},{"cited_title":"Zhang, X","cited_arxiv_id":null,"evidence_quote":"Provides a hydrogen-storage knowledge platform that extracts and organizes graphical performance data for AI use."}],"review_version":1}