{"id":"5d8eb9f8-a1ce-44ee-b8a9-82b9fd84123a","arxiv_id":"2501.18163","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A tomographic, information-theoretic framework for materials representations is introduced and supported by feature-importance experiments on 126,325 Materials Project compounds.","lead":"This paper proposes viewing materials, their properties, and their machine-learning representations as different projections of an unknown underlying material essence, using information theory to explain why simple descriptors like chemical formula can predict properties surprisingly well. The authors test this idea by measuring how adding one property as an input feature changes prediction accuracy for another property across a large materials database.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that composition suffices because polymorphs are dataset-limited is not directly tested; the 59.08% one-to-one statistic is too weak, and the property-augmentation experiments measure a different quantity.","rationale":"The reader's weakest-assumption analysis correctly identifies the unvalidated feature-importance proxy in §2.2 as a serious gap. My stress-test agrees with that concern but locates the load-bearing weakness one step earlier: even if the proxy were perfect, the reported experiments compare augmented vs non-augmented representations, not composition-only vs composition-structure models, so they cannot test the paper's central causal claim about polymorph-limited datasets. The 59.08% one-to-one statistic is insufficient to establish the required near one-to-one correspondence, and no conditional-entropy or controlled-polymorph analysis is provided. The paper's contribution is primarily conceptual, and the heatmaps are a useful empirical resource; the gap does not warrant rejection, but it does sharpen the revisions needed: quantify H(structure|formula), directly compare representation classes under controlled polymorph fractions, and validate the feature-importance proxy on data with known information structure. This is consistent with the reader's CONDITIONAL verdict, so no change is recommended.","tokens_in":10273,"tokens_out":6811,"duration_ms":75661,"concrete_test":"Take the same MP 2020 snapshot and compute, per chemical formula, the number of distinct structure groups (e.g., distinct spacegroup plus Wyckoff settings) to obtain H(structure|formula). Construct stratified subsets that vary the fraction of polymorphic formulas while controlling target distributions, and retrain CGCNN composition-only and composition-structure baselines with identical hyperparameters and seeds on each subset. The causal claim predicts the composition-only vs structure-aware test-MAE gap shrinks monotonically as formula-to-structure uniqueness increases; if the gap is insensitive to polymorph fraction, the proposed mechanism is not the driver of composition-based model performance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The causal mechanism proposed in §2.1 is that composition and stoichiometry constrain structure, and that datasets further limit polymorphs, producing a near one-to-one formula–structure correspondence. The only direct evidence offered is the footnote-7 statistic that 59.08% of Materials Project entries have a one-to-one material-id/formula correspondence; this does not establish near one-to-one for the remaining 41% of entries, and no conditional entropy H(structure|formula) is computed. The verification experiments in §2.2 do not compare composition-only vs composition-structure test errors; they measure relative change in MAE when a third property is added to a representation, interpreted through an unvalidated feature-importance proxy for PID terms. Even if that proxy were valid, it would not test whether structure is redundant given composition. Additionally, the proxy itself is load-bearing: the argument that performance cannot be harmed because mutual information is non-negative does not survive finite-sample training, where irrelevant or poorly scaled features can increase test error through overfitting or optimization difficulty; the paper's own footnote 8 concedes the scaling caveat. The central explanatory claim is therefore a plausible conjecture supported by an indirect, partly circular statistic, not by the reported experiments.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a conceptual 'tomographic interpretation' in which a material is a latent essence M, and both representations (e.g., chemical formula, crystal structure) and properties are projections of M. It defines sufficient and minimally sufficient representations via mutual information I(R(M);M), invokes partial information decomposition (PID) to explain when augmenting a representation with an additional property should help, and proposes that composition-based ML succeeds because compositions constrain structures and dataset biases suppress polymorph diversity. To 'verify' the framework, the authors train modified CGCNN models on Materials Project data with composition-only and composition-structure baselines, add each of 13 properties as an extra feature, and measure relative test-MAE changes across 9 target properties; results are presented as heatmaps and summarized in Table 1, with structure-dependent properties identified as being embedded in the structure representation.","tokens_in":10444,"tokens_out":11333,"duration_ms":108856,"significance":"The value of the paper is in providing a vocabulary and an interpretive lens for a genuine empirical phenomenon: simple representations sometimes perform well in materials ML, and property augmentation has task-dependent benefits. The experimental matrix is large and systematic (2340 trained models across two baselines, 9 targets, 13 augmenting properties, and five seeds), and the observation that positive mean changes are generally not robust across seeds is a useful sanity check. However, the central explanatory claim—that composition suffices because of dataset-limited polymorphs—is not directly tested, and the information-theoretic interpretation of the experiments rests on an unvalidated proxy. The contribution is therefore best assessed as an interesting hypothesis-generating framework plus a preliminary augmentation study, rather than a verified theory.","major_comments":[{"comment":"The central causal mechanism—that composition is sufficient because elemental composition and stoichiometry constrain structure and because datasets limit the number of polymorphs—is not tested by the reported experiments. Footnote 7 reports that 59.08% of the Materials Project snapshot has a one-to-one material–formula correspondence, but this does not establish a near one-to-one correspondence for the other 41% of entries, and no conditional entropy H(structure|formula) is computed for the dataset actually used. The experiments in Section 2.2 compare augmented versus non-augmented representations, not composition-only versus composition-with-structure test errors, so they cannot determine whether structure is redundant given composition. Please add a direct comparison (same architecture, same splits, with and without structure), stratify the performance gap by the number of polymorphs per formula, and report H(structure|formula), or else explicitly recast the claim as a hypothesis rather than a verified result.","section":"Section 2.1, footnote 7"},{"comment":"The paper identifies feature importance (relative change in test MAE) with the mutual-information quantities in the PID decomposition without validation. The sentence 'we aim to indirectly assess this quantity through the feature importance' is an assertion, not an argument; relative test-MAE changes can be driven by optimization ease, model capacity, regularization, feature scaling, and finite-sample overfitting. The theoretical statement that performance cannot be harmed by including more information because mutual information is non-negative applies to expected information, not to an empirical test error on a finite dataset, and footnote 8 already concedes that poorly scaled features can hurt. Because the PID interpretation of Figures 4 and 5 and the construction of Table 1 depend on this proxy, please validate the proxy (for example, on synthetic data with known PID/MI terms, or by comparing against direct MI estimates on low-dimensional subsets) or restrict the conclusions to observations about test-error changes.","section":"Section 2.2"},{"comment":"The significance criterion used throughout the results is not statistically valid. The authors call the min-max range over five repetitions a 'confidence interval,' but a range of five values is not a confidence interval and carries no stated error control; the rule that significance requires the min-max range not to overlap zero is an ad hoc criterion, and no multiple-comparison correction is applied across the 9 targets and 13 augmenting properties. Since Table 1 and the conclusions about which properties 'had an impact' are derived from this criterion, please report the per-seed/split distributions, use a proper paired test (e.g., a paired bootstrap or signed test), and justify or correct for multiple comparisons.","section":"Section 2.2, Table 1"},{"comment":"The empirical verification is partly circular. The framework asserts that if adding a property improves performance, the property must encode unique or synergistic information about the target; but 'improvement' is exactly the operational definition of feature importance used in the experiments, so the observed improvements are a restatement of the proxy rather than an independent confirmation of the tomographic or PID interpretation. The heatmaps are genuine empirical observations about property augmentation, but they do not test the core framework claims—that representations are projections of a material essence, that all projections together determine the essence, or that sufficiency is task-dependent. Please separate the confirmable empirical predictions (e.g., under controlled normalization and sufficient data, augmentation does not harm; augmentation benefits are asymmetric) from the post-hoc interpretations, and state which framework claims the experiments could in principle falsify.","section":"Section 2.1, Eq. (3)"}],"minor_comments":[{"comment":"The experimental count appears inconsistent: the text says the process is repeated for 'five different dataset splits' after stating that a constant 60-20-20 split was used for five seeds, while footnote 9 counts five seeds and two models per setting (giving 2340 models). Please clarify the number of seeds, splits, and the total model count.","section":"Section 2.2"},{"comment":"The heatmaps should include an explicit colorbar and a precise definition of the plotted quantity (percentage change in test MAE of augmented versus non-augmented representation), since the text's blue/red significance discussion depends on that scale.","section":"Figures 4 and 5"},{"comment":"The notation 'I(Ra(M); M) > I(Rb(M); M) ∀M ∈ M' conflates the material as a random variable with the set of materials; please introduce a random-variable convention (e.g., lowercase m for a sample, M for the random variable) to make the information-theoretic statements precise.","section":"Section 2.1"},{"comment":"The column header 'In Struct. but not Comp.' is confusing: it lists properties that improve the composition-restricted baseline but not the composition-structure baseline, meaning the structure representation already captures that information. Please rename the column or define it more explicitly.","section":"Section 3, Table 1"},{"comment":"The concluding statement that multi-property inverse design should improve with more properties is not derived from the formalism and ignores finite-sample and redundancy considerations; please qualify it.","section":"Section 4"},{"comment":"Reference [15] is cited without a venue or identifier; please provide the preprint or publication details.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The manuscript is best understood as a conceptual/interpretation paper with a parallel empirical augmentation study. The augmentation study is a useful reference dataset, but I would not recommend accepting the paper on the strength of the claimed verification of the tomographic framework. The authors should either add the direct composition-vs-structure test and validate the feature-importance proxy, or reposition the paper as a perspective that proposes a hypothesis and presents preliminary evidence. In its current form, the gap between the central claim and the evidence is too wide for publication in a research track."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this with the expectation of a perspective piece, not a proof. The tomographic framing—representations and properties as projections of an unobservable material essence—is a genuinely useful new packaging for an old puzzle. It gives a clean language for why composition-only models can win: task-dependent minimal sufficiency, and dataset-limited polymorph degeneracy. The empirical appendix is real work: 2340 trained models, five seeds, conservative significance criteria, and the Table 1 breakdown of which properties are exclusively captured by structure is a useful resource for anyone doing property-augmented representations.\n\nThe soft spots are in the link between the framework and the experiments. The feature-importance proxy is load-bearing: relative MAE change is asserted to track the PID terms, but there is no argument that it tracks information content rather than optimization ease, architecture capacity, or feature scaling. The authors do footnote the scaling caveat, but the main text's 'performance can not be harmed by including more information' is too strong for finite-sample training. The central explanatory claim—composition suffices because datasets limit polymorphs—gets only the 59.08% one-to-one statistic, which is too weak to establish near one-to-one for the remaining 41%, and no conditional entropy is computed. The experiments also do not directly test composition-only vs composition-structure; they measure a different quantity (the change when adding a property to each baseline). So the causal story is plausible, but the verification is indirect.\n\nThat said, the paper is not circular. The asymmetry observations (total-magnetization helps energy-per-atom much more than the reverse) are interesting and consistent with PID's non-symmetry. The paper is honest about its limitations, and the conceptual framework does suggest design rules, even if they are not yet quantitative. I would bring this to a reading group for the framing, and I would cite it when discussing representation design—with the caveat about the proxy. It deserves a serious referee; the review should ask for a direct test of the one-to-one claim and a clearer statement of the proxy assumption as a limitation.","headline":"A genuinely new framing for why simple representations work in materials ML, but the verification doesn't directly test the central claim.","tokens_in":10982,"tokens_out":2597,"would_cite":true,"duration_ms":25313,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A material is not its representation; composition-only models work when a dataset's formulas and structures are nearly one-to-one, and the paper formalizes this with an information-theoretic 'tomographic' view of materials.","keywords":["materials informatics","composition-based models","structure-property relations","information theory","partial information decomposition","machine learning representations","tomographic interpretation"],"falsifier":"Take a dataset that contains multiple structurally distinct polymorphs for the same chemical formula with differing values of a target property, and train a composition-only and a structure-based model on it. If the composition-only model's error is much larger, the near-one-to-one correspondence that the argument depends on is absent in that dataset.","tokens_in":10046,"feed_emoji":"🧪","tokens_out":7764,"duration_ms":70326,"temperature":0.7,"pith_summary":"This paper addresses a puzzle in machine learning for materials: models that see only the chemical formula, with no crystal structure, often predict properties almost as well as models that see the full structure. The authors argue that this is not a paradox but a consequence of information redundancy: a formula strongly constrains the allowed structures, and many datasets contain only one or a few polymorphs per formula, so formula and structure can be nearly interchangeable within the data. To make this precise, they introduce a 'tomographic interpretation' in which a material is an unseen essence and both representations and measured properties are projections (shadows) of that essence. Using ideas from information theory, and specifically partial information decomposition, they explain when adding one property to a representation should improve prediction of another. They verify the picture with an exhaustive battery of experiments in which each of thirteen properties is added to composition-based and structure-based representations and the change in prediction error on nine target properties is measured.","feed_headline":"A near one-to-one formula-structure link explains 'simple' ML wins","feed_subtitle":"When a dataset has few polymorphs per formula, composition and structure carry nearly the same information.","key_machinery":"The machinery is the tomographic interpretation combined with partial information decomposition (PID), a way of splitting the information that two sources carry about a target into unique, redundant, and synergetic parts. The tomographic interpretation redefines the learning problem: instead of mapping material to property, the model maps one projection of the material to another, and the sufficient information needed for the map is task-dependent. PID then decomposes the information that two source variables (e.g., a base representation and an added property) carry about a target into unique, redundant, and synergetic components; the synergetic component explains why a property can help even when it shares no direct information with the target. The empirical instrument is the feature importance, defined as the relative change in test mean absolute error when a property is added, measured over five seeds with a 60-20-20 split for both a composition-only and a composition-plus-structure version of the same graph neural network.","core_discovery":"The paper's central claim is that a material should not be identified with any of its representations. A representation is an approximation of an inaccessible 'material essence,' and a property is also a projection of the same essence; the boundary between representation and property is blurry. The discovery that resolves the composition-only puzzle is that, within typical datasets, the chemical formula and the crystal structure carry nearly the same information: elemental composition and stoichiometry limit the number of polymorphs, and many datasets include only a subset of those polymorphs, so there can be a near one-to-one correspondence between formula and structure. In that regime the structure adds little information for a given task, and the learning problem can be solved from composition alone. The experiments support the framework by showing that a property added to a representation improves prediction only when it encodes non-redundant information about the target, and that structure-specific properties such as spacegroup, density, and volume are already captured in a structure-based representation.","pith_inferences":["A direct test of the framework would compute or estimate the actual mutual information between formula, structure, and target properties on datasets with known polymorph distributions and compare it with the feature-importance proxy; if the two diverge, the proxy needs correction before the causal story is fully established.","The directional asymmetry observed between properties (total magnetization helps energy-per-atom prediction far more than the reverse) suggests that experimental measurement campaigns could prioritize properties by their unique information contribution to a target, a prioritization the paper does not itself derive.","The tomographic view implies that multi-property conditioning in generative inverse design should improve reconstruction fidelity up to the point where added properties are fully redundant, which could be tested by measuring how generated-material validity scales with the number of conditioning properties."],"forward_implications":["If formula and structure are nearly redundant within a dataset, then composition-based screening is a sound first step in discovery campaigns, and inverse design from formula becomes closer to designing the material itself.","Property augmentation should be guided by whether the added property supplies unique or synergetic information for the target, not by generic physical intuition about what 'should' matter.","Datasets with many polymorphs per formula are precisely where structure-based representations should retain an advantage over composition-only ones.","The framework makes forward and inverse design two instances of the same operation, mapping between low- and high-information projections, so tools developed for one can transfer to the other.","Counting the number of distinct structures per formula in a dataset gives a practical, a priori indicator of when composition-only models will be sufficient."],"supporting_citations":[{"why":"Earlier observation that composition and composition-structure models perform similarly for stable non-polymorphic materials; this is the hypothesis the paper formalizes.","marker":"[15]"},{"why":"The 126 325-material database snapshot used in all experiments; its one-to-one formula-to-material statistics supply the empirical motivation for the near-one-to-one correspondence.","marker":"[16]"},{"why":"Introduces partial information decomposition, the tool used to factor information into unique, redundant, and synergetic contributions.","marker":"[30]"},{"why":"The graph neural network architecture that was modified so the structure component can be omitted, enabling the composition versus composition-plus-structure comparison.","marker":"[40]"},{"why":"Defines the benchmark suite whose 13 tasks motivate the claim that composition-based models lead on 6 tasks.","marker":"[9]"},{"why":"The leaderboard ranking that supplies the observation of composition-based general-purpose algorithms at the top of several tasks.","marker":"[10]"}],"fun_headline_variants":["Why ML wins with just the formula","Formula and structure often say the same thing","The hidden equivalence behind simple ML models","Information theory explains simple-material ML success","Composition alone can beat structure when few polymorphs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise of the experimental section is that relative change in test mean absolute error when a feature is added faithfully reflects the mutual-information content of that feature for the target, a proxy the paper itself describes as indirect and does not prove.","fun_headline_variants_meta":{"raw":{"variants":["Why ML wins with just the formula","Formula and structure often say the same thing","The hidden equivalence behind simple ML models","Information theory explains simple-material ML success","Composition alone can beat structure when few polymorphs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1332,"prompt_tokens":850,"completion_tokens":482,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":417}},"tokens_in":466,"tokens_out":482,"duration_ms":5762,"temperature":1.0,"reasoning_tokens":417,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T00:26:03.491832+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a dataset that contains multiple structurally distinct polymorphs for the same chemical formula with differing values of a target property, and train a composition-only and a structure-based model on it. If the composition-only model's error is much larger, the near-one-to-one correspondence that the argument depends on is absent in that dataset.","supporting_citations":[{"cited_title":"What information is necessary and sufficient to predict materials properties using machine learning?, 2022","cited_arxiv_id":null,"evidence_quote":"Earlier observation that composition and composition-structure models perform similarly for stable non-polymorphic materials; this is the hypothesis the paper formalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The 126 325-material database snapshot used in all experiments; its one-to-one formula-to-material statistics supply the empirical motivation for the near-one-to-one correspondence."},{"cited_title":"Williams and Randall D","cited_arxiv_id":null,"evidence_quote":"Introduces partial information decomposition, the tool used to factor information into unique, redundant, and synergetic contributions."},{"cited_title":"Grossman","cited_arxiv_id":null,"evidence_quote":"The graph neural network architecture that was modified so the structure component can be omitted, enabling the composition versus composition-plus-structure comparison."},{"cited_title":"Benchmarking materials property prediction methods: the matbench test set and automatminer reference algorithm","cited_arxiv_id":null,"evidence_quote":"Defines the benchmark suite whose 13 tasks motivate the claim that composition-based models lead on 6 tasks."},{"cited_title":"Matbench leaderboard-property: General purpose algorithms, 2024","cited_arxiv_id":null,"evidence_quote":"The leaderboard ranking that supplies the observation of composition-based general-purpose algorithms at the top of several tasks."}],"review_version":1}