REVIEW 4 major objections 5 minor 150 references
Generative AI can propose new crystals but cannot yet prove they are novel materials.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 09:23 UTC pith:Y4AIUO5D
load-bearing objection Useful novelty taxonomy; the evidence-concentration claim needs better empirical support than the Scopus keyword counts provide. the 4 major comments →
Generative and multimodal AI for materials prediction and design: Progress, challenges, and perspectives
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that novelty in materials discovery cannot be reduced to statistical distance in composition or latent space. It distinguishes three levels: structural novelty (new composition/crystallography), physical novelty (behaviour outside the known distribution under comparable conditions), and deployment novelty (performance advantage under realistic synthesis and operating conditions, including service-life retention). Reviewing multimodal data categories—composition, microstructure, processing, testing/characterisation—the paper finds current evidence is concentrated in composition and idealised structure, while the processing, microstructural, and experimental observ
What carries the argument
The carrying framework is a two-part conceptual structure. The materials property hierarchy ranks information from atomic-scale composition and intrinsic properties up through microstructure, processing history, and engineering performance, explaining why path-dependent extrinsic properties resist prediction from idealised structure alone. Built on it, the three-level novelty taxonomy (structural, physical, deployment) assigns each level the type of evidence required to establish novelty and the type of multimodal data needed. This hierarchy does the argument's work: it converts the abstract question 'can AI design novel materials?' into the concrete question 'does the evidence base reach th
Load-bearing premise
The paper's empirical claim that evidence is concentrated in composition and idealised structure rests on a keyword-based classification of published papers; if that classification misjudges how often processing and microstructure data actually appear in model inputs, the central diagnosis loses its foundation.
What would settle it
Take a random sample of, say, 200 recent AI-for-materials papers and record the data modalities actually used as model inputs, bypassing keyword categories. If processing and microstructural data appear as frequently as composition data, or if testing-and-characterisation data turn out to be used as prediction targets rather than evidence, the paper's central evidence-concentration claim would be contradicted. A narrower check: compare keyword-based counts of 'processing' papers with a manual reading of the same sample to measure classification bias.
If this is right
- If the evidence-concentration claim is right, generative models should be credited with structural novelty only; claims of physical or deployment novelty need independent experimental demonstration.
- Current benchmarks built on computational labels and proxy novelty metrics will overstate deployability, so evaluation must shift to out-of-distribution splits, synthesis success rates, and experimental confirmation.
- Building datasets that link each composition to processing route, microstructure, measurement conditions, and negative results becomes a prerequisite for progress, not an optional extra.
- Process-aware models that treat processing as a primary modality, and generators that return design tuples (composition, route, microstructure, properties, uncertainty), are the concrete direction for moving up the hierarchy.
- Community-wide metadata standards and deposition of failed/null trials are direct corollaries: without them, higher-level novelty claims remain ungrounded regardless of model advances.
Where Pith is reading between the lines
- One consequence the paper leaves implicit: the taxonomy gives a practical screening rule reviewers could apply—any novelty claim for a processed material should be backed by a processed sample, not a simulated structure, which would immediately separate structural from physical novelty claims.
- The emphasis on 'compound vs material' suggests that idealized-structure databases may systematically mislead generative models, and that provenance records deserve the same status as composition in benchmark design—a step the paper motivates but does not itself implement.
- If evidence scarcity is the bottleneck, then automated laboratories that record failed trials are not merely validation tools but data-generation infrastructure; their output should be treated as first-class training data for feasibility-first generation.
- A testable extension: design a benchmark that holds out entire processing routes for a fixed composition and requires models to predict measured properties; the performance drop relative to random splits would quantify how much processing evidence current models actually use.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This Perspective argues that generative and multimodal AI for materials design has made strong progress on 'structural novelty' but lags on 'physical novelty' and 'deployment novelty'. The authors introduce a materials property hierarchy (intrinsic to extrinsic) and a three-level novelty taxonomy, then review multimodal data, representation learning, knowledge integration, generative models, and benchmarks. Their central claim, stated in Sec. 6 and the abstract, is that current evidence available to AI models is concentrated in composition and idealised structure, while processing, microstructure, and experimental characterisation data are under-represented and weakly integrated. This evidence-concentration claim is supported by a Scopus bibliometric analysis (SI S1, Fig. 3), by qualitative discussion of databases and benchmarks (Secs. 3.3 and 5), and by examples from steels, metastable phases, and creep. The paper concludes with four opportunities (multimodal data construction, process-aware modelling, feasibility-first generation, deployment-aware benchmarking) and community recommendations.
Significance. If the central claim is correct, the paper provides a useful framework for assessing AI-driven materials discovery and a clear agenda: shift the field from structure-centric generative models toward process-aware, experimentally grounded, deployment-aware systems. The three-level novelty taxonomy and the emphasis on 'dark data' (missing negative results) are valuable contributions for the community. The paper is a Perspective rather than a methods paper, and it ships reproducible bibliometric data on GitHub, which is a strength. However, the load-bearing empirical assertion about evidence concentration is not adequately established by the presented analysis, and the paper's own Figure 3b appears to contradict it. Since this claim motivates the entire recommendations section, the manuscript needs revision to either provide direct evidence about model training data or carefully reframe the claim as a qualitative observation about data modality usage in benchmarks and databases.
major comments (4)
- [Abstract and Sec. 6 vs Fig. 3b] The abstract and Sec. 6 state that current evidence is 'concentrated in composition and idealised structure', but Fig. 3b shows that, in the authors' own bibliometric analysis, composition is the smallest of the four data categories in every year from 2020 to 2025 (around 10–20% of the categorized publications), while testing and characterisation dominates. The paper attempts to reconcile this in Sec. 3.1 by saying composition 'is often used as baseline information ... rather than analysed as a primary data input.' This is a plausible interpretation, but it is an untested assumption: no evidence is given that the 'baseline' usage in the counted publications corresponds to 'evidence available to models.' The reader's and the paper's own numbers do not support the unqualified wording in the abstract. The claim must be rephrased to distinguish 'frequently used as a conditioning or auxiliary
- [SI S1, Supplementary Table 3] The bibliometric analysis that underlies the evidence-concentration claim uses Scopus TITLE-ABS-KEY queries with keyword groups that are not comparable in breadth. The 'Composition' group contains relatively narrow terms (e.g., 'stoichiometr*', 'chemical formula'), while the 'Microstructure' group bundles many terms including crystal structure, CIF, space group, defects, and graph representations, and 'Testing and characterisation' lists many named experimental techniques. Consequently, the relative proportions in Fig. 3b could be an artifact of query construction rather than a measure of evidence availability. The authors acknowledge in SI S1 that the analysis is 'inherently approximate,' but no sensitivity analysis is provided (e.g., varying keyword groups, using alternative databases, or checking whether the results are robust to adding/removing synonyms). More fundamentally, publicat
- [Sec. 3.3 and Sec. 5 (qualitative support)] The qualitative support for the evidence-concentration claim is selective. For instance, Sec. 3.3 lists Materials Project, OQMD, and JARVIS as composition/structure-centric resources, and Sec. 5.1's Table 2 of benchmarks shows that property-prediction benchmarks indeed rely primarily on composition and structure with DFT labels. However, these examples do not quantify 'concentration' or 'under-representation' across the field. A systematic comparison of, say, the number of entries in open materials databases by modality, or the modality coverage of a representative set of recent generative-model papers, would make the claim more robust. Without such data, the argument in Sec. 6 that 'the evidence available to models is concentrated at the level of composition and idealised structure' remains a plausible but under-determined assertion.
- [Sec. 6.1 (deployment-aware benchmarking)] The paper correctly notes the difficulty of distinguishing a genuinely new candidate from a candidate that is simply absent from a limited dataset (Sec. 6.1). However, this caveat also applies to the paper's own bibliometric evidence: the keyword counts cannot distinguish between 'under-represented in the literature' and 'under-represented in model training data.' This reinforces the need for the direct evidence requested above. The point is not to reject the paper's thesis, but to ensure that the framing matches the evidence level.
minor comments (5)
- [Throughout] The manuscript header contains placeholder text 'IOP Publishing J. Phys. Mater vv(yyyy) aaaaaa Authoret al', which should be replaced with proper journal style. Also, 'Authoret al' in the running head is a typo.
- [SI S1, Fig. 3a caption] The caption says 'The annual number of publications that match the multimodal AI for materials query and involve at least one of four materials data categories.' This is fine, but the SI text says 'This data was obtained by adding the OR union...' — use 'These data were obtained'.
- [Sec. 3.1, Table 1] Table 1 is helpful, but the row for 'Microstructure' lists both atomic-scale and larger-length-scale information under one category, which is very broad. Consider separating 'crystal structure' from 'microstructure' in later discussions, since the central claim treats 'idealised structure' separately from 'microstructure'.
- [Sec. 5.1, Table 2] In Table 2, the column header 'Comp.' is used but the table notes say 'Comp., Microstruct., Proc. and Test. & character.' — ensure abbreviations are defined in the caption for readers.
- [Sec. 6.2] The recommendation about metadata standards is reasonable, but it could briefly mention that such standards already exist in parts of the community (e.g., via NOMAD, OPTIMADE) to avoid appearing to ignore prior infrastructure work. This is a presentation issue, not a substantive flaw.
Circularity Check
No significant circularity: the paper is a perspective whose central claims rest on bibliometric data and literature review, not on fitted inputs or self-citation chains.
full rationale
The paper does not claim a derivation in the sense of equations fitted to data. Its central Sec. 6 assertion—that progress toward physical and deployment novelty is limited by evidence concentrated in composition and idealised structure while higher-level evidence is under-represented—is supported by the Scopus keyword analysis (SI S1, Table 3, Fig. 3) and by qualitative reviews of databases and benchmarks (Secs. 3.3, 5). These are external empirical inputs, not outcomes of fitting the conclusion. The taxonomy in Sec. 2 is a conceptual framework, not estimated from data, and it is used to organize observations rather than to predict them. The self-citations (refs 21, 23) are background examples of multimodal learning and are not load-bearing for the article's main claim. The SI explicitly acknowledges the bibliometric analysis is approximate and intended for broad trends, and Sec. 3.1's explanation that composition appears small because it is often used as baseline information is an interpretive assumption about the data. That is a potential validity concern about the evidence base, not a circular reduction of a prediction to its input. No fitted parameter is renamed as a prediction, and no claim reduces to a self-citation chain.
Axiom & Free-Parameter Ledger
axioms (4)
- ad hoc to paper The three-level novelty taxonomy (structural, physical, deployment) is a useful and sufficient organizing scheme for evaluating AI-proposed materials.
- domain assumption Intrinsic vs extrinsic property hierarchy captures the key distinction between composition/structure-derived and processing-dependent behavior.
- domain assumption Scopus keyword queries in Supplementary Table 3 are a valid proxy for the modality distribution of AI-for-materials research.
- domain assumption Computational reference labels (e.g., DFT) and experimental measurements are not on a single fidelity continuum.
read the original abstract
Artificial intelligence (AI) is accelerating materials prediction and design by enabling efficient exploration of chemical and structural spaces, with particular promise for novel materials discovery. However, novelty in materials discovery encompasses chemical plausibility, structural distinctiveness, property relevance and experimental realisability, making AI-driven novelty claims difficult to substantiate. We introduce a materials property hierarchy, from intrinsic, composition-determined properties to extrinsic, processing-dependent performance, to clarify deployment constraints and distinguish structural, physical and deployment novelty. This framework motivates an evidence-based view of multimodal materials data spanning chemical composition, microstructure, processing, and testing and characterisation, showing that current evidence remains concentrated in composition and idealised structure while heterogeneous, under-represented and weakly integrated modalities limit support for physical and deployment novelty. It also highlights the limitations of benchmarks based mainly on computational labels and proxy novelty criteria. Community-wide standards for data collection, modality alignment and evidence synthesis are needed to support multimodal data construction, process-aware multimodal modelling, feasibility-first generative modelling and deployment-aware benchmarking, so that generative and multimodal AI can design experimentally realisable materials with defensible scientific and practical novelty.
Figures
Reference graph
Works this paper leans on
-
[1]
Merchant A, Batzner S, Schoenholz S S, Aykol M, Cheon G and Cubuk E D 2023Nature62480–85
-
[2]
2025Nature639624–632
Zeni C, Pinsler R, Z¨ ugner D, Fowler A, Horton M, Fu X, Wang Z, Shysheya A, Crabb´e J, Ueda Set al. 2025Nature639624–632
-
[3]
Hashemi S M, Parvizi S, Baghbanijavid H, Tan A T, Nematollahi M, Ramazani A, Fang N X and Elahinia M 2022International Materials Reviews671–46
-
[4]
Kang Y, Park H, Smit B and Kim J 2023Nature Machine Intelligence5309–318
-
[5]
Moro V, Loh C, Dangovski R, Ghorashi A, Ma A, Chen Z, Kim S, Lu P Y, Christensen T and Soljaˇci´c M 2025Newton1
-
[6]
Wu Y, Ding M, He H, Wu Q, Jiang S, Zhang P and Ji J 2025npj Computational Materials11
-
[7]
Joshi C K, Fu X, Liao Y L, Gharakhanyan V, Miller B K, Sriram A and Ulissi Z W 2025 All-atom diffusion transformers: Unified generative modelling of molecules and materialsInternational Conference on Machine Learning
2025
-
[8]
Ye C, Wang Y, Xie X, Zhu T, Liu J, He Y, Zhang L, Zhang J, Fang Z, Wang Let al.2025npj Computational Materials12
-
[9]
Chenebuah E T, Nganbe M and Tchagang A B 2024npj Computational Materials10
-
[10]
Jiao R, Huang W, Liu Y, Zhao D and Liu Y 2024 Space group constrained crystal generation International Conference on Learning Representations
2024
-
[11]
Xie T and Grossman J C 2018Physical Review Letters120145301
-
[12]
Choudhary K and DeCost B 2021npj Computational Materials7
-
[13]
Deng B, Zhong P, Jun K, Riebesell J, Han K, Bartel C J and Ceder G 2023Nature Machine Intelligence51031–1041
-
[14]
Ye C Y, Weng H M and Wu Q S 2024Computational Materials Today1100003
-
[15]
Luo X, Wang Z, Gao P, Lv J, Wang Y, Chen C and Ma Y 2024npj Computational Materials10
-
[16]
Khajeh A, Lei X, Ye W, Yang Z, Hung L, Schweigert D and Kwon H K 2025Digital Discovery4 11–20
-
[17]
Lambard G, Yamazaki K and Demura M 2023Scientific Reports13566
-
[18]
Altoyuri A H, Sarmah A and Jain M K 2024Acta Materialia281120431
-
[19]
Jiang X, Wang W, Tian S, Wang H, Lookman T and Su Y 2025npj Computational Materials11
-
[20]
Schilling-Wilhelmi M, R ´ıos-Garc´ıa M, Shabih S, Gil M V, Miret S, Koch C T, M´arquez J A and Jablonka K M 2025Chemical Society Reviews541125–1150
-
[21]
Liu X, Ouyang B and Zeng Y 2025Nature Computational Science592–94
-
[22]
Zhang L, Bing Q, Qin H, Yu L, Li H and Deng D 2025Matter8
-
[23]
Liu X, Zhang J, Zhou S, van der Plas T L, Vijayaraghavan A, Grishina A, Zhuang M, Schofield D, Tomlinson C, Wang Yet al.2025Nature Machine Intelligence1–13 14 IOP PublishingJ. Phys. Matervv(yyyy) aaaaaa Authoret al
-
[24]
Jablonka K M, Rosen A S, Krishnapriyan A S and Smit B 2023ACS Central Science9563–581
-
[25]
Omee S S, Fu N, Dong R, Hu M and Hu J 2024npj Computational Materials10
-
[26]
Li K, Rubungo A N, Lei X, Persaud D, Choudhary K, DeCost B, Dieng A B and Hattrick-Simpers J 2025Communications Materials69
-
[27]
Jain A, Ong S P, Hautier G, Chen W, Richards W D, Dacek S, Cholia S, Gunter D, Skinner D, Ceder Get al.2013APL Materials1
-
[28]
Kirklin S, Saal J E, Meredig B, Thompson A, Doak J W, Aykol M, R¨ uhl S and Wolverton C 2015npj Computational Materials1
-
[29]
Saal J E, Kirklin S, Aykol M, Meredig B and Wolverton C 2013Jom651501–1509
-
[30]
Choudhary K, Garrity K F, Reid A C, DeCost B, Biacchi A J, Hight Walker A R, Trautt Z, Hattrick-Simpers J, Kusne A G, Centrone Aet al.2020npj Computational Materials6
-
[31]
Antoniuk E R, Cheon G, Wang Get al.2023npj Computational Materials155
-
[32]
Walsh A and Zunger A 2017Nature Materials16964–967
-
[33]
Xie J, Zhou Y, Faizan M, Li Z, Li T, Fu Y, Wang X and Zhang L 2024Nature Computational Science 4322–333
-
[34]
Zheng Y, Xu H, Li Z, Li L, Yu Y, Jiang P, Shi Y, Zhang J, Huang Y, Luo Q, Lou Z and Wang L 2025 Advanced Materials372504378
2025
-
[35]
Yang H, Hu C, Zhou Y, Liu X, Shi Y, Li J, Li G, Chen Z, Chen S, Zeni Cet al.2024arXiv preprint arXiv:2405.04967
-
[36]
Hayes S M and McCullough E A 2018Resources Policy59192–199
-
[37]
Mannan S, Myers R J, Batra R, Mercado R, Wondraczek L and Krishnan N 2026arXiv preprint arXiv:2601.21527
-
[38]
Betala S, Gleason S P, Ramlaoui A, Xu A, Channing G, Levy D, Fourrier C, Kazeev N, Joshi C K, Kaba S Oet al.2025arXiv preprint arXiv:2512.04562
-
[39]
Baird S G, Sayeed H M, Montoya J and Sparks T D 2024Journal of Open Source Software95618
-
[40]
Gruver N, Sriram A, Madotto A, Wilson A G, Zitnick L and Ulissi Z 2024 Fine-tuned language models generate stable inorganic materials as textInternational Conference on Learning Representations
2024
-
[41]
Ong S P, Richards W D, Jain A, Hautier G, Kocher M, Cholia S, Gunter D, Chevrier V L, Persson K A and Ceder G 2013Computational Materials Science68314–319
-
[42]
Widdowson D E and Kurlin V A 2026SIAM Journal on Applied Mathematics86898–918
-
[43]
Juelsholt M 2026Materials Horizons13(11) 5672–5679
-
[44]
Cheetham A K and Seshadri R 2024Chemistry of Materials363490–3495
-
[45]
Yu D, Griesemer S, Liu T c, Wolverton C and Zhu Y 2025npj Computational Materials11
-
[46]
Meredig B, Agrawal A, Kirklin S, Saal J E, Doak J W, Thompson A, Zhang K, Choudhary A and Wolverton C 2014Physical Review B89094104
-
[47]
Ward L, Agrawal A, Choudhary A and Wolverton C 2016npj Computational Materials2
-
[48]
Butler K T, Davies D W, Cartwright H, Isayev O and Walsh A 2018Nature559547–555
-
[49]
Wang Z, Wang L, Zhang H, Xu H and He X 2024Nano Convergence118
-
[50]
Kononova O, Huo H, He T, Rong Z, Botari T, Sun W, Tshitoyan V and Ceder G 2019Scientific Data6 203
-
[51]
Lee S, Cruse K, Baibakova V, Ceder G and Jain A 2025Scientific Data
-
[52]
Teng Y, Tan H, Huang W and Shan G 2025Communications Physics8 15 IOP PublishingJ. Phys. Matervv(yyyy) aaaaaa Authoret al
-
[53]
Wissel S, Scheunert J, Dextre A, Ahmed S, Beyer A, Volz K and Xu B X 2025Materials & Design 115069
-
[54]
Wanni J, Bronkhorst C A and J T D 2024npj Computational Materials10
-
[55]
Ling J, Hutchinson M, Antono E, DeCost B, Holm E A and Meredig B 2017Materials Discovery10 19–28
-
[56]
Tsuruta H and Kumagai M 2025 MatPROV: A provenance graph dataset of material synthesis extracted from scientific literatureNeurIPS 2025 Workshop on AI for Accelerated Materials Design
2025
-
[57]
Kim H, Jeon T, Choi S, Hong J H, Jeon D W, Baek G Y, Kwak G W, Lee D H, Bae J, Lee Cet al.2025 Towards fully-automated materials discovery via large-scale synthesis dataset and expert-level llm-as-a-judgeACM International Conference on Information and Knowledge Managementpp 1302–1312
2025
-
[58]
Zhang J, Yin C, Farbiz F, Jafary-Zadeh M and Sing S L 2025Virtual and Physical Prototyping20 e2592732
-
[59]
Deshmukh K, Riensche A, Bevans B, Lane R J, Snyder K, Halliday H S, Williams C B, Mirzaeifar R and Rao P 2024Materials & Design244113136
-
[60]
Lederbauer M, Betala S, LI X, Jain A, Sehaba M E A, Channing G, Germain G, Leonescu A, Flaifil F, Amayuelas A, Nozadze A, Schmid S P, Zaki M, Ethirajan S K, Pan E, Franckel M L D, Duval A, Krishnan N M A and Gleason S P 2025 Lemat-synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literatureNeurIPS 2025 Workshop on ...
2025
-
[61]
Salgado J E, Lerman S, Du Z, Xu C and Abdolrahim N 2023npj Computational Materials9
-
[62]
Oviedo F, Ren Z, Sun S, Settens C, Liu Z, Hartono N T P, Ramasamy S, DeCost B L, Tian S I, Romano Get al.2019npj Computational Materials5
-
[63]
Paulus B and Biskup T 2023Digital Discovery2234–244
-
[64]
Statt M J, Rohr B A, Guevarra D, Suram S K, Morrell T E and Gregoire J M 2023Scientific Data10 184
-
[65]
Swain M C and Cole J M 2016Journal of Chemical Information and Modeling561894–1904
1904
-
[66]
Mavracic J, Court C J, Isazawa T, Elliott S R and Cole J M 2021Journal of Chemical Information and Modeling614280–4289
-
[67]
Polak M P and Morgan D 2024Nature Communications151569
-
[68]
Dagdelen J, Dunn A, Lee S, Walker N, Rosen A S, Ceder G, Persson K A and Jain A 2024Nature Communications151418
-
[69]
Hira K, Zaki M, Krishnan Net al.2025arXiv preprint arXiv:2509.10448
-
[70]
Li C, Xian Y, Zhou Y, Ding X, Sun J and Xue D 2026Advanced Materialse20478
-
[71]
Debnath A, Krajewski A M, Sun H, Lin S, Ahn M, Li W, Priya S, Singh J, Shang S, Beese A M, Liu Z K and Reinhart W F 2021Journal of Materials Informatics13
-
[72]
Reeves-McLaren N and Christensen S M L 2026Journal of Materials Chemistry A14276–283
-
[73]
Berry J and Christofidou K A 2025Materials Science and Technology41773–790
-
[74]
Wang A Y T, Murdock R J, Kauwe S K, Oliynyk A O, Gurlo A, Brgoch J, Persson K A and Sparks T D 2020Chemistry of Materials324954–4965
-
[75]
Chang R, Wang Y X and Ertekin E 2022npj Computational Materials8
-
[76]
Xu P, Ji X, Li M and Lu W 2023npj Computational Materials9
-
[77]
Che L, He Z, Zheng K, Si T, Ge M, Cheng H and Zeng L 2023Materials Today Communications37 107531 16 IOP PublishingJ. Phys. Matervv(yyyy) aaaaaa Authoret al
-
[78]
Wong R, Tran A, Dovgyy B, Maldonado C S and Pham M S 2025Materials & Design256114301
-
[79]
Gianola D S, della Ventura N M, Balbus G H, Ziemke P, Echlin M P and Begley M R 2023Current Opinion in Solid State and Materials Science27101090
-
[80]
Mouritz A P 2012Introduction to aerospace materials(Elsevier) ISBN 0857095153
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.