Pith. sign in

REVIEW 5 minor 37 references

The Precursor Genome: A Pairwise Reaction Dataset for Solid-State Synthesis

T0 review · 0 major / 5 minor · reviewed 2026-07-14 · grok-4.5

Pith's one-line read A machine-readable map of 1,035 pairwise solid-state reactions reports full protocols, raw XRD, and expert-validated outcomes so models of inorganic synthesis can finally be trained and tested.

desk verdict Solid, high-value experimental data release: 1,035 pairwise solid-state outcomes with full provenance, raw XRD, and dual-expert scores; the short Tammann dwell is a deliberate scope choice, not a hidden flaw. read the letter →

arxiv 2607.09903 v1 pith:GG7KF5TH submitted 2026-07-10 cond-mat.mtrl-sci

classification cond-mat.mtrl-sci
keywords solid-statesynthesisprecursorgenomeself-drivinglaboratoryRietveldrefinementpowderX-raydiffractionFAIRmaterialsdatareactionoutcomedatasetinorganic
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Solid-state synthesis still dominates inorganic materials making, yet published records usually list only successes, omit negatives, and leave out the raw logs and diffraction needed for true reuse. This paper fills that gap with the Precursor Genome: 1,035 pairwise reactions among 46 common precursors spanning 39 elements, all run on one autonomous laboratory under a single standardized protocol. Every entry carries measured thermal profiles, masses, instrument settings, the raw powder XRD scans, and Rietveld phase assignments that human experts scored on a three-tier quality scale. The whole collection is released as a Pydantic-validated JSON ledger with complete provenance from precursor pair to final phase label. The authors present it as a FAIR, reusable benchmark so that first-principles, data-driven, and machine-learning models of solid-state reactivity can be trained and evaluated against consistent experimental ground truth rather than sparse, success-biased literature.

What carries the argument

The Precursor Genome itself: a hierarchical Pydantic-validated JSON ledger that joins each precursor pair to its full synthesis metadata, raw XRD scans, ranked Rietveld candidates, and human quality scores under a controlled reaction-outcome vocabulary.

What would settle it

Re-running a statistically meaningful subset of the same precursor pairs for much longer dwell times or at different temperatures and checking whether the expert-validated phase categories systematically change.

Watch

Extended reading notes

Core claim

The Precursor Genome is a complete, machine-readable dataset of 1,035 pairwise solid-state reactions that reports every experimental protocol detail, every raw XRD pattern, and every expert-validated phase assignment with unbroken provenance, thereby establishing the first large FAIR benchmark for predictive models of solid-state reactivity.

Load-bearing premise

That a single one-hour heat at the Tammann-rule temperature is long enough and representative enough to serve as a reliable ground-truth label of solid-state reactivity for machine-learning models.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The manuscript presents the Precursor Genome, a FAIR dataset of 1,035 pairwise solid-state reactions among 46 common inorganic precursors (39 elements) executed on the A-Lab autonomous platform. Every reaction is accompanied by full experimental metadata (measured thermal profiles, precursor/recovered masses, instrument configuration), raw powder XRD (1,351 scans), and automated Rietveld refinements via the Dara framework (1,950 cases) that are independently scored by human experts on a three-tier quality scale. Outcomes are labeled with a controlled vocabulary (unreacted, transformed, partially reacted, completely reacted, physical failure). The data are released as a Pydantic-validated hierarchical JSON ledger with complete provenance from precursor pair to final phase assignment, plus raw patterns, serialized refinements, and tutorial notebooks. The short 1-hour Tammann-rule dwell is intentional for capturing early intermediates and negative outcomes.

Significance. No comparably large, machine-readable solid-state synthesis dataset with consistent provenance, raw characterization, dual-expert validation, and explicit negative/partial outcomes currently exists. The resource directly addresses a well-documented bottleneck for data-driven synthesis science. Strengths include full instrument logging, Pydantic schema validation, dual (or triple) human quality scoring, explicit flagging of physical failures and bad scans, open CC-BY release on Zenodo/GitHub, and accompanying loader notebooks. These features make the dataset immediately usable as a benchmark for predictive models of solid-state reactivity and for training phase-identification algorithms.

minor comments (5)
  1. Methods §2.2 and Eq. (1): The Tammann-rule temperature (two-thirds lowest Tm/decomposition point, rounded down to nearest 100 °C, clamped 200–1100 °C) and fixed 1-hour dwell are free experimental choices. The manuscript already frames them as intentional for early intermediates; a short explicit caveat in Usage Notes that the labels are kinetic snapshots rather than equilibrium synthetic accessibility would further protect downstream users.
  2. Technical Validation / Usage Notes: ICSD/COD coverage gaps are acknowledged via the three-tier quality scores. Consider adding a one-sentence quantitative summary (e.g., fraction of quality-1 vs quality-3 refinements) so users can immediately gauge label reliability without parsing the full ledger.
  3. Fig. 1 caption and OutcomeEntry: The controlled vocabulary is clear, but a brief note on how multi-phase or ambiguous cases were assigned to “transformed” vs “partially reacted” would reduce residual ambiguity for binary-classifier users.
  4. Table 1: Several precursors list air-sensitivity >10 % (e.g., LiOH·H2O, K2CO3·1.5H2O). A short remark on whether these were handled under special conditions or simply accepted as-is would help reproducibility.
  5. Data Availability / Code Availability: Zenodo DOI and GitHub URL are given; confirming that the exact ledger version used for the manuscript figures is tagged would strengthen long-term provenance.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: experimental data release with independent observations, not a fitted or self-defined derivation.

full rationale

The Precursor Genome is a FAIR data-release resource, not a first-principles derivation or predictive model. Its load-bearing claim is the existence of 1,035 pairwise solid-state reaction outcomes with full experimental metadata, raw XRD, and dual-expert-validated Rietveld assignments under a controlled outcome vocabulary. Those outcomes are measured experimental observations (A-Lab synthesis + powder XRD + human-verified Dara refinements), not quantities defined by or fitted to the paper's own inputs. DFT reaction energies appear only as optional side fields queried from Materials Project and are not used to assign labels. Self-citations to A-Lab infrastructure and the Dara refinement framework supply methods and tooling; they do not define or force the reported phase outcomes. Tammann-rule temperature selection (Eq. 1) is an explicit experimental protocol choice, not a prediction of reactivity. No self-definitional loop, fitted-input-as-prediction, uniqueness import, or renaming of a known result is present. Circularity score is therefore zero.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

As a data-release paper the central claim rests on experimental execution and expert validation rather than on free parameters or invented physical entities. The main modeling choices that affect label interpretation are the Tammann temperature rule, the 1-hour dwell, the 1:1 cation stoichiometry target, and the three-tier human quality scale; these are domain conventions or protocol decisions, not fitted constants. No new particles, forces, or conserved quantities are postulated.

free parameters (3)
  • Tammann dwell temperature rule (Tt = 100 * floor((2/3)*min(Tm)/100), clamped 200–1100 °C)
    Protocol choice that sets every reaction temperature; not fitted to the outcome data but still a free modeling decision that determines which kinetic regime is sampled.
  • 1-hour isothermal dwell
    Fixed short-time protocol chosen to capture early intermediates and negatives; duration is not optimized against any reactivity metric.
  • 1:1 metal-cation target stoichiometry
    Uniform mixing rule applied to every pair; alternative ratios could change observed products.
assumptions (4)
  • domain assumption Tammann’s rule (solid-state reaction temperature ≈ 2/3 of the lowest melting point) is a useful heuristic for choosing dwell temperature even for precursors that decompose before melting.
    Invoked in Methods §2.2 and Eq. 1; the paper notes it is an approximation for decomposing compounds.
  • domain assumption ICSD and COD together supply a sufficiently complete set of candidate crystal structures for automated Rietveld phase identification of the product mixtures.
    Used by the Dara pipeline (Methods §2.3); residual unidentified intensity is captured by the quality-3 score rather than claimed to be zero.
  • domain assumption Two (or three) human experts scoring refinements on a 1–3 scale produce reliable ground-truth phase labels for downstream ML.
    Technical Validation section; modal score after arbitration is treated as authoritative.
  • standard math Standard crystallographic and thermodynamic tools (BGMN kernel, Materials Project DFT energies, pymatgen phase diagrams) are correct for the purposes of this dataset.
    Background infrastructure cited and used without re-derivation.
invented entities (1)
  • Precursor Genome ledger (Pydantic-validated hierarchical schema with controlled reaction-outcome vocabulary) independent evidence
    purpose: Machine-readable container that links every precursor pair to raw data, refinements, and expert labels.
    New data structure introduced by the paper; it is a software artifact, not a physical entity, and is fully specified and released.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Precursor Genome: A Pairwise Reaction Dataset for Solid-State Synthesis." pith.science (2026). https://pith.science/paper/GG7KF5TH

@misc{pith2026260709903,
  author       = {Pith},
  title        = {Pith review of: The Precursor Genome: A Pairwise Reaction Dataset for Solid-State Synthesis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GG7KF5TH}},
  note         = {Machine review of arXiv:2607.09903}
}
read the original abstract

Solid-state reactions remain the dominant route to inorganic materials, yet no large, machine-readable dataset reports their experimental protocols and outcomes with consistent provenance; this gap obstructs first-principles, data-driven, and machine-learning approaches to synthesis science. Here, we present the Precursor Genome, a dataset of 1,035 pairwise solid-state reactions generated autonomously by the A-Lab self-driving laboratory, spanning 46 precursors and 39 elements. Every reaction is reported together with its full experimental metadata, including measured thermal profiles, precursor and recovered masses, and instrument configuration. Every product mixture is identified from raw X-ray diffraction (1,351 scans) through automated Rietveld refinement with the Dara framework, yielding 1,950 refinement cases that are independently validated by human experts on a three-tier quality scale. Raw pattern files, serialized refinement objects, and reviewer annotations are distributed through a Pydantic-validated JSON ledger, preserving full traceability from each precursor pair to its final phase assignment. The Precursor Genome establishes a FAIR, reusable benchmark for training and evaluating predictive models of solid-state reactivity.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

37 extracted references · 2 canonical work pages

  1. [1]

    J.et al.An autonomous laboratory for the accelerated synthesis of novel materials.Nature624, 86–91 (2023)

    Szymanski, N. J.et al.An autonomous laboratory for the accelerated synthesis of novel materials.Nature624, 86–91 (2023)

  2. [2]

    P.et al.Self-driving laboratory for accelerated discovery of thin-film materials

    MacLeod, B. P.et al.Self-driving laboratory for accelerated discovery of thin-film materials. Sci. Adv.6(2020)

  3. [3]

    N.et al.Pascal: the perovskite automated spin coat assembly line accelerates composition screening in triple-halide perovskite alloys.Digit

    Cakan, D. N.et al.Pascal: the perovskite automated spin coat assembly line accelerates composition screening in triple-halide perovskite alloys.Digit. Disc.3, 1236–1246 (2024). 17

  4. [4]

    Burger, B.et al.A mobile robotic chemist.Nature583, 237–241 (2020)

  5. [5]

    Rev.124, 9633–9732 (2024)

    Tom, G.et al.Self-driving laboratories for chemistry and materials science.Chem. Rev.124, 9633–9732 (2024)

  6. [6]

    URL https://doi.org/10.1063/1.4812323

    Jain, A.et al.Commentary: The Materials Project: A materials genome approach to accelerating materials innovation.APL Mater.1, 011002 (2013). URL https://doi.org/10.1063/1.4812323

  7. [7]

    URL http://dx.doi.org/10.1038/s41586-023-06735-9

    Merchant, A.et al.Scaling deep learning for materials discovery.Nature624, 80–85 (2023). URL http://dx.doi.org/10.1038/s41586-023-06735-9

  8. [8]

    URL https://doi.org/10.1038/s41586-025-08628-5

    Zeni, C.et al.A generative model for inorganic materials design.Nature639, 624–632 (2025). URL https://doi.org/10.1038/s41586-025-08628-5

Show all 37 references
  1. [9]

    URL https://arxiv.org/abs/2512.09169

    Cavignac, T.et al.Ai-driven expansion and application of the alexandria database (2025). URL https://arxiv.org/abs/2512.09169

  2. [10]

    Stein, A., Keller, S. W. & Mallouk, T. E. Turning down the heat: Design and mechanism in solid-state synthesis.Science259, 1558–1564 (1993). URL http://www.jstor.org/stable/ 2880656

  3. [11]

    Synergies and competition between current approaches to materials discovery.C

    Jansen, M. Synergies and competition between current approaches to materials discovery.C. R. Chim.21, 958–968 (2018)

  4. [12]

    Mater.29, 9436–9444 (2017)

    Kim, E.et al.Materials synthesis insights from scientific literature via text extraction and machine learning.Chem. Mater.29, 9436–9444 (2017). URL https://doi.org/10.1021/acs. chemmater.7b03500

  5. [13]

    Nature533, 73–76 (2016)

    Raccuglia, P.et al.Machine-learning-assisted materials discovery using failed experiments. Nature533, 73–76 (2016). URL https://doi.org/10.1038/nature17439

  6. [14]

    URL https://doi.org/10.1038/s41586-019-1540-5

    Jia, X.et al.Anthropogenic biases in chemical reaction data hinder exploratory inorganic synthesis.Nature573, 251–255 (2019). URL https://doi.org/10.1038/s41586-019-1540-5

  7. [15]

    Discov.4, 1602–1611 (2025)

    Baibakova, V.et al.Precursor reaction pathway leading to bifeo3 formation: insights from text-mining and chemical reaction network analyses.Digit. Discov.4, 1602–1611 (2025)

  8. [16]

    URL https://arxiv.org/abs/2405.13930

    Fei, Y.et al.Alabos: A python-based reconfigurable workflow management framework for autonomous laboratories (2024). URL https://arxiv.org/abs/2405.13930. 18

  9. [17]

    URL https://books.google.com/books?id=t65OAAAAMAAJ

    Tammann, G.Lehrbuch der Metallographie: Chemie und Physik der Metalle und ihrer Legierungen(Voss, 1923). URL https://books.google.com/books?id=t65OAAAAMAAJ

  10. [18]

    & Maier, J

    Merkle, R. & Maier, J. On the tammann–rule.Z. anorg. allg. Chem.631, 1163–1166 (2005)

  11. [19]

    J., Rom, C

    Fei, Y., McDermott, M. J., Rom, C. L., Wang, S. & Ceder, G. Dara: Automated multiple- hypothesis phase identification and refinement from powder x-ray diffraction.Chemistry of Materials38, 1364–1376 (2026)

  12. [20]

    & Kleeberg, R

    Bergmann, J., Friedel, P. & Kleeberg, R. Handling of unusual instrumental profiles by the bgmn rietveld program.Mater. Sci. Forum321–324, 192–197 (2000)

  13. [21]

    & Rehme, S

    Zagorac, D., M¨ uller, H., Ruehl, S., Zagorac, J. & Rehme, S. Recent developments in the inorganic crystal structure database: theoretical crystal structure data and related features.J. Appl. Crystallogr.52, 918–925 (2019)

  14. [22]

    Jain, A.et al.Commentary: The materials project: A materials genome approach to accelerating materials innovation.APL Mater.1(2013)

  15. [23]

    P.et al.Python materials genomics (pymatgen): A robust, open-source python library for materials analysis.Comput

    Ong, S. P.et al.Python materials genomics (pymatgen): A robust, open-source python library for materials analysis.Comput. Mater. Sci.68, 314–319 (2013)

  16. [24]

    P., Wang, L., Kang, B

    Ong, S. P., Wang, L., Kang, B. & Ceder, G. Li-fe-p-o 2 phase diagram from first principles calculations.Chem. Mater.20, 1798–1807 (2008)

  17. [25]

    D., Miara, L

    Richards, W. D., Miara, L. J., Wang, Y., Kim, J. C. & Ceder, G. Interface stability in solid-state batteries.Chem. Mater.28, 266–273 (2015)

  18. [26]

    Xiao, Y.et al.Understanding interface stability in solid-state batteries.Nat. Rev. Mater.5, 105–126 (2019)

  19. [27]

    J., Dwaraknath, S

    McDermott, M. J., Dwaraknath, S. S. & Persson, K. A. A graph-based network for predicting chemical reaction pathways in solid-state materials synthesis.Nat. Commun.12(2021)

  20. [28]

    J.et al.Physical descriptor for the gibbs energy of inorganic crystalline solids and temperature-dependent materials chemistry.Nat

    Bartel, C. J.et al.Physical descriptor for the gibbs energy of inorganic crystalline solids and temperature-dependent materials chemistry.Nat. Commun.9(2018)

  21. [29]

    Cheminform.15(2023)

    Vaitkus, A.et al.A workflow for deriving chemical entities from crystallographic data and its application to the crystallography open database.J. Cheminform.15(2023). 19

  22. [30]

    Cheminform.15(2023)

    Merkys, A.et al.Graph isomorphism-based algorithm for cross-checking chemical and crystallographic descriptions.J. Cheminform.15(2023)

  23. [31]

    & Graˇ zulis, S

    Vaitkus, A., Merkys, A. & Graˇ zulis, S. Validation of the crystallography open database using the crystallographic information framework.J. Appl. Crystallogr.54, 661–672 (2021)

  24. [32]

    & Vaitkus, A

    Quir´ os, M., Graˇ zulis, S., Girdzijauskait˙ e, S., Merkys, A. & Vaitkus, A. Using smiles strings for the description of chemical connectivity in the crystallography open database.J. Cheminform. 10(2018)

  25. [33]

    Merkys, A.et al.Cod::cif::parser: an error-correcting cif parser for the perl language.J. Appl. Crystallogr.49, 292–301 (2016)

  26. [34]

    & Okuliˇ c-Kazarinas, M

    Graˇ zulis, S., Merkys, A., Vaitkus, A. & Okuliˇ c-Kazarinas, M. Computing stoichiometric molecular composition from crystal structures.J. Appl. Crystallogr.48, 85–91 (2015)

  27. [35]

    Graˇ zulis, S.et al.Crystallography open database (cod): an open-access collection of crystal structures and platform for world-wide collaboration.Nucleic Acids Res.40, D420–D427 (2011)

  28. [36]

    Graˇ zulis, S.et al.Crystallography open database – an open-access collection of crystal structures.J. Appl. Crystallogr.42, 726–729 (2009)

  29. [37]

    Downs, R. T. & Hall-Wallace, M. The american mineralogist crystal structure database.Am. Mineral.88, 247–250 (2003). 20

Pith tools

Reviewed July 14, 2026 · model on record in the stance chip above.