{"id":"ce0639a1-a08e-44d1-a028-6f062762c051","arxiv_id":"2607.17910","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A configurable oxidation-state filter (SMACT-based) lets users tune strictness and doubles as an RL reward, improving chemical realism of generated crystals.","lead":"This paper builds a tunable chemical filter for AI-generated inorganic crystals, using how often each oxidation state appears in real crystal data to grade plausibility from permissive to strict. It shows most popular generators produce charge-balanced formulas that rely on rare oxidation states, and the filter can also act as a reward to steer generation toward realistic chemistry.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MACE-MPA-0 energy errors may systematically inflate Ehull for rare oxidation states; the filter's low-Ehull preservation claim needs a DFT cross-check.","rationale":"The paper is a well-executed methods study with openly available code and data. The reader's weakest assumption (ICSD oxidation-state frequencies as ground truth) is a valid concern, but I find a more immediate threat to the central claim: the energy evaluation used to demonstrate the filter's thermodynamic relevance. The Ehull comparison (Figure 2, Methods §4.4) relies on MACE-MPA-0, and the reported gap is comparable to the stated MLIP error. If MLIP errors are systematically biased against rare oxidation states, the correlation between filter validity and low Ehull could be an artifact. This is a load-bearing assumption because the abstract and results emphasize that stricter filtering 'preserves low-energy compounds' as evidence of the filter's utility. The paper explicitly acknowledges MLIP errors but does not test for systematic bias across valid/invalid classes. The proposed DFT cross-check would settle whether the observed separation is robust. I partially agree with the reader: both concerns involve grounding in external data, but the MLIP bias is more directly tied to the headline quantitative result. The reader already recommended CONDITIONAL; my concern reinforces that no change in verdict is needed, but acceptance should be conditional on addressing this or softening the stability claim.","tokens_in":12257,"tokens_out":8888,"duration_ms":82805,"concrete_test":"Randomly sample 100 valid and 100 invalid generated compositions under C=3, θ=0.5. Relax each with DFT (e.g., PBEsol) and compute Ehull against the same Materials Project phase diagram. Compare the median valid–invalid gap and the fraction above 0.1 eV/atom with the MACE-MPA-0 values from Figure 2. If the gap shrinks by >20% or loses statistical significance, the filter's apparent thermodynamic selectivity is an artifact of the MLIP energy evaluator.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative evidence that the SMACT filter 'preserves low-energy compounds' is the Ehull comparison in Figure 2. However, these Ehull values are computed with MACE-MPA-0 (Methods §4.4), an MLIP with 'errors of a few tens of meV/atom' (the paper's own statement). The reported median valid–invalid gap at C=3, θ=0.5 is 0.089 eV/atom (89 meV/atom), which is the same order as the estimated error. The analysis assumes these errors are not systematically correlated with oxidation-state commonality, but MACE-MPA-0 is trained on Materials Project structures dominated by common oxidation states; rare oxidation-state compounds are out-of-distribution, and the model may systematically overestimate their energies. If so, the valid–invalid separation in Figure 2 would be inflated or even spurious, directly undermining the claim that stricter filtering retains low-Ehull compounds. The paper acknowledges MLIP errors but does not test for bias with respect to the validity grouping variable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a tunable chemical validity operator built on the SMACT oxidation-state model, with a consensus threshold C and a commonality threshold θ that control the strictness of empirical oxidation-state prior. The authors benchmark six generative models (ADiT, CDVAE, Chemeleon-DNG, DiffCSP, DiffCSP++, MatterGen) on 10,000 generated compositions each, showing that charge neutrality and the electronegativity rule are usually satisfied, while agreement with empirical oxidation-state commonality varies substantially across models. They further analyze energy above the convex hull (Ehull) using MACE-MPA-0 relaxations, reporting that stricter validity filtering correlates with lower Ehull. Finally, they demonstrate the operator as a binary or continuous RL reward for fine-tuning the Chemeleon2 latent diffusion model, claiming the continuous reward improves metastability and mSUN at the cost of uniqueness and novelty.","tokens_in":12606,"tokens_out":8903,"duration_ms":80641,"significance":"If the claims hold, this work provides a simple, interpretable, and model-agnostic chemical filter that can serve both as a post-hoc diagnostic and as an in-training reward for generative models. The threshold formulation is a pragmatic extension of SMACT, and the open-source code and benchmark data are valuable resources. The benchmarking across six state-of-the-art models is a useful contribution. However, the central Ehull-based evidence rests on MLIP energies with potentially systematic bias, and the RL results are unreplicated, so the quantitative claims require strengthening before the significance can be fully assessed.","major_comments":[{"comment":"The claim that stricter filtering preserves low-energy compounds is based on the Ehull comparison in Fig. 2, computed with MACE-MPA-0, which the paper states has errors of 'a few tens of meV/atom' (§4.4). At C=3, θ=0.5 the median valid–invalid gap is 0.089 eV/atom (89 meV/atom), the same order of magnitude as the reported MLIP error. Since MACE-MPA-0 is trained on Materials Project structures dominated by common oxidation states, energies for compounds with rare oxidation states may be systematically biased, inflating the observed separation. The paper does not test for systematic bias with respect to the validity grouping variable. I recommend a DFT cross-check on a stratified sample (e.g., 50–100 valid and invalid compounds balanced across oxidation-state commonality) or, at minimum, an analysis of residual errors against oxidation-state frequency. Without this, the 'preserves low-ener","section":"§4.4 / Fig. 2"},{"comment":"The reinforcement-learning results are presented from single training runs. Figure 6 shows one reward trajectory and one set of final metrics per reward, with no error bars or multiple seeds. The observed changes (e.g., mSUN 0.0764→0.0908, metastability 0.3279→0.4310) could be within run-to-run variance, especially with only 400 update steps and stochastic generation. The claims that the continuous SMACT reward improves metastability and mSUN, and that the binary reward does not, require reproducibility evidence. Please report means and standard deviations over at least 3–5 independent seeds, or otherwise justify the stability of the results.","section":"§2.4 / Fig. 6"}],"minor_comments":[{"comment":"The sentence 'reducing the fraction of retained compositions by a further 3–7%' is inconsistent with Table 1: from q+χ to C=3 the decreases are 1.7%, 4.2%, 3.7%, 4.0%, 3.8%, 3.9% for ADiT, CDVAE, Chemeleon-DNG, DiffCSP, DiffCSP++, and MatterGen. Similarly, the '42–55%' range for θ=0.5 omits CDVAE's 41.8% (or should be rounded consistently). Please correct the text.","section":"§2.2 / Table 1"},{"comment":"The term 'mSUN' is used in the Results but only defined in the Methods. Please define it at first use in the Results. Also, specify the number of independent runs/seeds in the figure caption.","section":"§2.4 / Fig. 6"},{"comment":"Clarify how include_zero=False interacts with the commonality calculation: is the denominator in Eq. (2) computed over the zero-inclusive or zero-excluded oxidation-state set? This could affect the filtered Ω set and the resulting retention rates.","section":"§4.1 / Eq. (2)"},{"comment":"The exclusion of Yb-containing compositions (2.5%) is stated only in the Methods. Please mention this important data exclusion in the Results text for Fig. 2, as it may affect the comparability of the Ehull distributions.","section":"§4.4 / Fig. 2"},{"comment":"The caption of Fig. 3 should state explicitly which oxidation-state list and which consensus setting (C=1 vs C=3) are used for the 'permissive' and 'strictest' panels. The current text leaves some ambiguity about the default parameters.","section":"§2.3 / Fig. 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is likely to interest the materials-informatics community. The central idea is simple and practical, and the benchmarking effort is substantial. However, the two major concerns—potential MLIP bias in the Ehull analysis and lack of replication in the RL experiments—are load-bearing and need to be addressed before publication. Both are testable and fixable within the scope of a revision. I would advise not rejecting, but requesting a revision with concrete, additional evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a solid methods paper from the SMACT group. The new thing is turning the oxidation-state validity operator into a tunable filter with two thresholds — consensus C and commonality θ — and showing it behaves sensibly across six generative models. That is genuinely useful: a cheap, model-agnostic prior that can screen outputs and steer a diffusion model through a continuous reward. The code and data are public, the benchmark is systematic, and the perovskite design-space analysis is a nice sanity check. The RL result is interesting, though it is a single run without error bars, so I would not overweight it.\n\nThe central claim — stricter filtering retains chemistry built from common oxidation states while removing rare-state compositions — is defensible on its own terms, because “valid” is explicitly defined from the ICSD oxidation-state distribution. That is a measurement convention, not a circular proof. The Ehull analysis is the piece I would want bolted down. It uses MACE-MPA-0 energies, and the paper admits errors of a few tens of meV/atom. The reported valid–invalid median gap at the strictest filter is about 89 meV/atom — same order as the error. More importantly, MACE-MPA-0 is likely trained on materials dominated by common oxidation states, so rare-state compounds may be out-of-distribution. If the error is systematically larger for exactly the compositions the filter rejects, the separation in Figure 2 could be inflated. The paper says it uses distributional comparisons, but it never tests for bias against the validity grouping. A small DFT cross-check on a random subset would settle this. I would not call it fatal, but it is the weakest link in the evidence for “preserving low-energy compounds.”\n\nThere are also two numerical sloppinesses in the text: the “further 3–7%” drop from q+χ to C=3 is actually about 1.7–4.2%, and the “42–55%” range under θ=0.5 excludes CDVAE at 41.8%. Easy fixes.\n\nThe ICSD coverage/bias limitation is acknowledged in the Discussion, and the authors are honest that the filter can penalise rare but real chemistry. That is fair.\n\nOverall: I'd send this to peer review. It is a well-scoped contribution that many groups will use, and the issues are patchable rather than structural.","headline":"Useful, well-scoped methods paper with a real soft spot in the MLIP-based stability validation; the tunable filter itself is solid and worth refereeing.","tokens_in":13013,"tokens_out":2806,"would_cite":true,"duration_ms":26581,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A tunable oxidation-state filter can screen and guide AI-generated materials, preserving low-energy crystals while rejecting chemically implausible ones.","keywords":["chemical validity","oxidation states","generative models","materials discovery","convex hull","reinforcement learning","screening","compositional filtering"],"falsifier":"Compile a list of experimentally confirmed inorganic compounds that rely on oxidation states with fewer than three database occurrences or with low relative frequency (for example, certain high-valence oxides or stabilized low-valence species). If the strict filter (C=3, θ=0.5) rejects a large share of these confirmed compounds, that is direct evidence that the database prior is too conservative; conversely, if it retains them, the prior tracks synthesizability closely.","tokens_in":12251,"feed_emoji":"🧪","tokens_out":5539,"duration_ms":52605,"temperature":0.7,"pith_summary":"The paper argues that chemical plausibility in AI-generated materials is not a binary yes/no but a matter of degree, and that it can be captured by a tunable operator built from oxidation-state statistics in experimental crystal data. The operator has two knobs: a consensus threshold, which drops oxidation states rarely seen in the database, and a commonality threshold, which keeps only states that account for a large enough share of an element's recorded occurrences. With these knobs, one can move continuously between permissive exploration and conservative down-selection. Across six generative models, the filter shows that most models get stoichiometry right but under-sample realistic oxidation-state combinations, and that stricter filtering preferentially removes high-energy compounds while retaining low-energy ones near the convex hull. The same operator can also act as a reinforcement-learning reward, steering a latent diffusion model toward more metastable, chemically grounded compositions.","feed_headline":"Chemical filter flags rare oxidation states, keeps stable crystals","feed_subtitle":"A tunable oxidation-state prior screens AI materials and doubles as a reward to steer generation toward low-energy compounds.","key_machinery":"The load-bearing component is a composition-level chemical validity operator built from empirical oxidation-state statistics. For each element, a filtered set of allowed oxidation states is constructed by keeping states that appear at least C times in a large experimental crystal-structure database (consensus) and that account for at least a fraction θ of that element's recorded occurrences (commonality). A composition is declared valid if some product assignment from these per-element sets is charge neutral and satisfies the Pauling electronegativity rule. For the continuous reward, the same machinery is turned into a score: the geometric mean, over the elements in a composition, of the obs","core_discovery":"The central claim is that oxidation-state frequencies observed in experimental crystal structures provide a model-agnostic chemical prior that can both evaluate and guide generative models. The paper formalizes this as a validity operator that accepts only compositions for which some assignment of element oxidation states is charge neutral and obeys the electronegativity rule, where the allowed states are those that survive two adjustable thresholds: consensus C (a minimum absolute number of database occurrences) and commonality θ (a minimum share of an element's occurrences). Benchmarks on six generative models show that while 93–97% of generated compositions pass charge neutrality, strict","pith_inferences":["Because the operator is tunable, it could be adapted to synthesis-specific goals: loosening θ to propose compounds that require unusual stabilisation, or tightening it near the end of a pipeline to favor well-trodden chemistry.","The correlation between validity and Ehull may reflect that common oxidation states are often those with accessible synthesis routes; if so, the filter indirectly encodes kinetic or thermodynamic accessibility beyond what the convex-hull calculation captures. This connection is not tested in the paper.","A testable extension: use the continuous reward in a multi-objective setup with an exploration bonus, which could recover some of the lost novelty while retaining the metastability gain.","As the underlying database grows and updates, the same operator can be re-derived without changing the method; quantifying how much rare-but-real chemistry is discarded at each threshold remains an open empirical question."],"forward_implications":["Generative models that score well on standard validity checks can still be far from the empirical oxidation-state distribution; the stricter commonality filter exposes this gap, with retention dropping to 42–55% at θ=0.5.","Charge neutrality and electronegativity consistency are not enough to predict stability; empirical oxidation-state commonality adds a chemical signal that correlates with low energy above the convex hull.","Stricter filtering preferentially removes high-Ehull compounds, so the operator acts as a cheap first-pass stability prior before expensive structure relaxation.","The same operator can be used as an in-training reward without retraining the generative architecture, since it is model-agnostic.","Oxidation-state-aware rewards can trade novelty and diversity for metastability; the continuous reward raised metastability while lowering uniqueness and novelty."],"fun_headline_variants":["Chemical filter steers AI to stable crystals","Oxidation-state filter blocks rare states in AI materials","Tunable prior filters and guides AI materials","AI materials get a chemical sanity filter"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire screening and reward signal is only as good as the oxidation-state frequencies in the experimental crystal-structure database, so if that database systematically under-records rare but synthesizable oxidation states, the filter will discard legitimate chemistry.","fun_headline_variants_meta":{"raw":{"variants":["Chemical filter steers AI to stable crystals","Oxidation-state filter blocks rare states in AI materials","Tunable prior filters and guides AI materials","AI materials get a chemical sanity filter"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000268,"raw_usage":{"total_tokens":1427,"prompt_tokens":688,"completion_tokens":739,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":691}},"tokens_in":432,"tokens_out":739,"duration_ms":8736,"temperature":1.0,"reasoning_tokens":691,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T16:37:11.731853+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compile a list of experimentally confirmed inorganic compounds that rely on oxidation states with fewer than three database occurrences or with low relative frequency (for example, certain high-valence oxides or stabilized low-valence species). If the strict filter (C=3, θ=0.5) rejects a large share of these confirmed compounds, that is direct evidence that the database prior is too conservative; conversely, if it retains them, the prior tracks synthesizability closely.","supporting_citations":[],"review_version":1}