Pith. sign in

REVIEW 3 major objections 5 minor 83 references

Representative Random Sampling of Chemical Space

T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This paper claims that unbiased random samples of chemical space can be generated without enumerating molecules, by estimating stoichiometry counts from average graph-edit distances and sampling each stoichiometry in proportion to its estim

desk verdict A clever approximate sampler for chemical space with a promising core idea, but the 'unbiased' claim outruns the validation; useful if read as approximate. read the letter →

arxiv 2508.20609 v1 pith:WMQZ3L56 submitted 2025-08-28 physics.chem-ph

classification physics.chem-ph
keywords chemicalspacerandomsamplingsmall-worldnetworksgrapheditdistanceMarkovchainMonteCarlodatabaserepresentativenessstoichiometryweighting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Chemical space, the set of all molecules consistent with simple valence rules, is so vast that listing its members is impossible beyond a few dozen atoms. This paper claims you can still take unbiased random samples of such a space and estimate how many molecules it contains: the logarithm of the molecule count for a given stoichiometry is approximately a linear function of the average shortest path between molecular graphs in a 'universe graph' where neighboring molecules differ by a single graph edit. Because average graph-edit distances can be estimated cheaply from a small random subset of molecules, the count estimate needs no enumeration. The authors calibrate this relation on exhaustively enumerated small molecules, extrapolate it to molecules of up to about 30 atoms, and then use the resulting stoichiometry-weighted sampler to measure how biased current databases such as ANI-1, QM9, and GDB-13 are relative to the chemical spaces they nominally cover.

What carries the argument

The 'universe graph' U(d): one vertex per protomolecule with a given labeled degree sequence—a molecular graph in which element labels are replaced by valence-type labels, so all monovalent atoms count identically—and an edge between two protomolecules exactly when their molecular graphs sit at minimal graph edit distance. The small-world relation l_G ∼ log |U(d)| is the load-bearing identity: it turns a cheaply sampleable quantity (average minimal edit distance over a small set of random protomolecules) into an estimate of the number of graphs. For degree sequences with several element labels per valence, a combinatorial factor (Eq. 5) rescales the path length of the corresponding pure degr

What would settle it

Enumerate a new space not used in calibration—for example all C,H,N,O molecular graphs with 11–15 heavy atoms—with an independent isomer generator, and compare the exact per-stoichiometry counts with the paper's Eq. 7 estimates; if relative errors grow systematically with molecular size or with valence multiplicity, the extrapolation fails and the sampler's uniformity guarantee collapses.

Watch

Extended reading notes

Core claim

On the paper's own terms: the number of molecular graphs in a stoichiometry class can be estimated without enumeration because the log of that count is a linear function of the average shortest path in a 'universe graph' whose vertices are the protomolecules and whose edges link molecules at minimal graph edit distance. Calibrated on 285,656 enumerated stoichiometries (about 26 trillion graphs), the relation is log|U(d)| ≈ 1.220 l_G − 0.7295. A combinatorial factor extends pure-degree-sequence path lengths to labeled degree sequences, and an asymptotic multigraph-count formula is calibrated beyond 20 atoms. A Markov-chain sampler then generates random graphs for each degree sequence, with st

Load-bearing premise

The linear relation between average graph-edit distance and the logarithm of molecule count, fitted on molecules of 3–10 atoms, is assumed to keep holding for larger molecules and for chemical spaces outside the calibration set.

Editorial extensions

If this is right

  • Unbiased reference data for machine learning becomes available on demand for graph-representable spaces up to roughly 30 atoms, without first enumerating the space.
  • Existing databases can be scored against their nominal chemical space; the paper reports KS distances of 0.26 (ANI-1), 0.49 (QM9), and 0.44 (GDB-13), indicating substantial stoichiometric bias.
  • The sampler provides a lower-bound criterion on database size for representativeness, and randomized subsets often approximate the underlying space with fewer samples than current databases.
  • Because fresh molecules can always be generated, benchmark and test sets can be refreshed with unseen data, giving a more trustworthy estimate of generalization error in chemistry machine learning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The count-via-distance trick is not obviously limited to molecules: any discrete object space with a cheap edit distance and small-world connectivity could in principle be counted the same way, but the linear calibration would need to be re-established for each new space.
  • The paper's own 73% per-stoichiometry order-of-magnitude accuracy implies that individual formula counts carry systematic error even when the global sample is balanced; a user needing exact relative weights for one stoichiometry would need the enumerated data, not the estimate.
  • The representativeness measure could be inverted into an acquisition rule: sample preferentially from stoichiometries where the cumulative gap between database and space is largest, rather than using weights alone.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces Representative Random Sampling (RRS), a pipeline for approximately uniform sampling of molecular graphs from a user-defined chemical space without full enumeration. The method first enumerates all stoichiometries via integer partitions, groups them by degree sequence, estimates the number of protomolecules per degree sequence using the small-world relation between average graph-edit-distance path length l_G and log count, calibrated by two regressions (Eqs. 7 and 10), and then selects a stoichiometry with probability proportional to the estimated count and generates a random graph by MCMC. The authors apply RRS to score the representativeness of ANI-1, QM9 and GDB-13 against their estimated underlying chemical spaces using KS and KL statistics, and they propose a criterion for the minimum database size needed for representativeness. The central claim is that this yields unbiased representative random samples and reliable count estimates up to about 30 atoms.

Significance. If the count estimates were accurate, RRS would be a valuable, inexpensive alternative to exhaustive enumeration for exploring chemical space and auditing database diversity. The software release (nablachem.space) and precomputed databases are practical strengths. However, the paper's own data show that the central count estimator has only ~73% of stoichiometries within one order of magnitude and has systematic per-stoichiometry errors, and the extrapolation to larger molecules relies on an unvalidated asymptotic calibration. The abstract's 'unbiased' language is therefore not established. The contribution is best viewed as a useful approximate sampling heuristic whose bias needs characterization before it can support the stronger claims made.

major comments (3)
  1. [§III A, Eq. (7), Fig. 2B] The abstract claims 'unbiased representative random samples', but the method weights stoichiometries by counts estimated from Eq. (7), a linear fit to enumerated data. Fig. 2B shows only 73% of estimates within an order of magnitude, and the text states there is 'a systematic error for any single stoichiometry which does not average out' (§III A). Since a stoichiometry's sampling probability is proportional to its estimated count, a systematic over- or under-estimate directly biases the sampled distribution; the claim that deviations balance in cumulative distributions is not a proof of unbiasedness. The conclusion's own qualifier ('only approximately uniformly distributed') contradicts the abstract. Please either provide a formal error analysis showing the resulting distribution is unbiased despite these errors, or soften the claim and quantify the approximate nature.
  2. [§II B, Eqs. (8)–(10)] For larger molecules the count estimates rely on Eq. (10), obtained by regressing l_G onto the asymptotic count formula (Eq. 8) for 148,620 pure degree sequences. This calibrates a prefactor but does not validate the asymptotic count against exact counts in the 20–30 atom range. The t! correction for added monovalent atoms is introduced ad hoc ('we correct by this amount'), and no held-out test is reported for non-pure degree sequences or for custom chemical spaces outside the calibration set. The sampling weights for molecules above the enumerated domain therefore remain unsupported. A validation on a subset of exactly enumerated counts in the target range, or analytical error bounds, is needed.
  3. [Table III, Fig. 3] The KS/KL representativeness scores are computed using the estimated chemical-space distribution. Because the count estimates carry systematic per-stoichiometry errors, the reference CDF itself is uncertain; no confidence intervals or sensitivity analysis is given for the KS values. The 'necessary criterion for a lower bound of database sizes' is presented from empirical curves without a formal derivation. These application-level conclusions are therefore conditional on the unvalidated count estimates.
minor comments (5)
  1. [Abstract] Typo: 'method produce' should be 'method to produce'. More substantively, the term 'unbiased' is too strong given the approximations acknowledged later in the paper; consider replacing it with 'approximately uniform' or 'bias-characterized'.
  2. [§II B] The text says 'This is computationally feasible until about 20 atoms' and later 'This is computationally feasible until about 30 atoms.' Please clarify which data product each statement refers to (labeling, degree-sequence sampling, or asymptotic estimation).
  3. [Eq. (5)] The product formula is not self-contained; define c_i explicitly and state the index ranges for i and j. A short example of how the binomial factors arise would also help readability.
  4. [Fig. 3] The legend and line styles should be clarified, especially why the GDB-13 curves terminate at a smaller sampling fraction than the other datasets. Adding axis labels and a caption that explains the construction of the 'randomized subspace' curves would improve interpretability.
  5. [§III A] The statement 'this simple model performs remarkably well' would be more convincing with a quantitative error metric (e.g., RMSE, mean absolute error, or R²) for the comparison in Fig. 2D.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the count estimates are empirical extrapolations and the benchmark comparisons use external data, not reductions to fitted inputs by construction.

full rationale

The paper's central quantities (molecule counts per stoichiometry and database representativeness) are not equivalent to the inputs of the fits by construction. Eq. 7 is a linear fit of log-count vs average path length to enumerated 3-10 atom data; it is then extrapolated to larger molecules, not used to 'predict' the same calibration points as an independent check. Eq. 10 calibrates the prefactor of the external Greenhill–McKay asymptotic formula for >20-atom pure degree sequences, and is again used for extrapolation beyond the calibrated domain. The applications compare against external databases (ANI-1, QM9, GDB-13), so the representativeness conclusions are not manufactured from the fitted parameters. The small-world relation (Eq. 4) is explicitly tested in Fig. 2A rather than assumed from a self-citation, and Eq. 8 is an external theorem. The only self-citation relevant to data/code ([72]) is a software package and is not load-bearing for the derivation. The paper honestly acknowledges systematic per-stoichiometry error and approximate uniformity, which are accuracy limitations, not circularity. No circular step is present.

Assumptions & free parameters 5 free parameters · 4 assumptions · 2 invented entities

The method depends on two fitted linear relations (Eqs. 7 and 10) that convert average path length into log-count estimates, plus several domain assumptions about the small-world property, asymptotic enumeration, and MCMC mixing. No new physical entities are introduced; the invented constructs are mathematical tools.

free parameters (5)
  • slope in Eq. 7 (a=1.220) = 1.220
    Fitted to exact counts of 285,656 degree sequences to relate log count to average path length.
  • intercept in Eq. 7 (b=-0.7295) = -0.7295
    Same linear fit as slope, on logarithmized counts.
  • slope in Eq. 10 (c=0.7561) = 0.7561
    Fitted to 148,620 pure degree sequences with more than 20 atoms to calibrate asymptotic scaling.
  • intercept in Eq. 10 (d=-14.40) = -14.40
    Same calibration as slope, for pure degree sequences.
  • t! correction factor for added monovalent atoms = empirical (t! )
    Empirical correction to asymptotic scaling when adding hydrogens; the paper states 'we correct by this amount' without a formal derivation.
assumptions (4)
  • domain assumption Small-world network approximation: average shortest-path length l_G in the universe graph scales as log |U(d)| (Eq. 4).
    Invoked in Sec. II.B to justify estimating counts from average path length. The paper validates this on enumerated data up to 10 atoms, but assumes it holds for larger molecules.
  • standard math Asymptotic enumeration of sparse multigraphs (Greenhill-McKay, Eq. 8) applies to chemical molecular graphs with valence constraints.
    Used in Sec. II.B for large molecules; the theorem is mathematical but its applicability to chemical graphs with loop-free and connectivity constraints is assumed.
  • domain assumption Isomorphisms are rare for large molecular graphs, so the number of protomolecules for non-pure degree sequences follows combinatorial counting (Eq. 5).
    Stated in Sec. II.B as 'for large and complex molecules ... isomorphisms are rare', supported only by informal expectation, not proof.
  • domain assumption The MCMC switch chain for sampling multigraphs with a given degree sequence mixes quickly and is uniform after automorphism correction.
    Assumed in Sec. II.B; the paper notes connectivity issues can slow or bias sampling in some regions.
invented entities (2)
  • Universe graph U(d)
    purpose: Abstract graph whose vertices are all protomolecules of a given degree sequence, used to estimate molecule counts via average path length.
    A mathematical construct, not a physical entity. Its validity rests on the small-world approximation, which is empirically fitted.
  • Protomolecule families
    purpose: Equivalence classes of molecular graphs ignoring element identity but preserving valency types, used to reduce combinatorial complexity.
    A bookkeeping abstraction; not independently testable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Representative Random Sampling of Chemical Space." pith.science (2026). https://pith.science/paper/WMQZ3L56

@misc{pith2026250820609,
  author       = {Pith},
  title        = {Pith review of: Representative Random Sampling of Chemical Space},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WMQZ3L56}},
  note         = {Machine review of arXiv:2508.20609}
}
read the original abstract

The overwhelming majority of molecules remains unexplored. This is mostly due to the sheer number of them, which prohibits any enumeration of chemical space, the set of all such molecules. In practice, only subsets of chemical space are considered, but those subsets exhibit substantial bias, prohibiting data-driven characterization of chemical space itself. In this work, we provide a method produce unbiased representative random samples of chemical space without enumeration of constituent molecules and to estimate the number of molecules in any custom chemical space. The approach is applicable to molecules which can be represented as graph and runs efficiently even for molecules of 30 atoms. We use it to estimate the representativeness of current databases with respect to their underlying chemical space and to establish a necessary criterion for a lower bound of database sizes to be representative of an underlying chemical space.

Figures

Figures reproduced from arXiv: 2508.20609 by the authors.

Figure 1
Figure 1. FIG. 1. Overview of the random sampling procedure and applications. From a well defined chemical space based on valency [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. Accuracy of the correlations exploited in this work. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. Kolmogorov–Smirnov (KS) statistic as a function of the sampled fraction of the corresponding chemical space. Solid [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

83 extracted references · 75 canonical work pages

  1. [1]

    Reymond, The Chemical Space Project, Accounts of Chemical Research 48, 722 (2015)

    J.-L. Reymond, The Chemical Space Project, Accounts of Chemical Research 48, 722 (2015)

  2. [2]

    Montavon, M

    G. Montavon, M. Rupp, V. Gobre, A. Vazquez- Mayagoitia, K. Hansen, A. Tkatchenko, K.-R. M¨ uller, and O. A. von Lilienfeld, Machine learning of molecular electronic properties in chemical compound space, New Journal of Physics 15, 095003 (2013), publisher: IOP Publishing tex.timestamp: 2019-05-12

  3. [3]

    J. A. Keith, V. Vassilev-Galindo, B. Cheng, S. Chmiela, M. Gastegger, K.-R. M¨ uller, and A. Tkatchenko, Com- bining machine learning and computational chemistry for predictive modeling and design, Chemical Reviews 121, 9816 (2021), publisher: American Chemical Society

  4. [4]

    J. L. Medina-Franco, A. L. Ch´ avez-Hern´ andez, E. L´ opez- L´ opez, and F. I. Sald ´ ıvar-Gonz´ alez, Chemical Multiverse: An Expanded View of Chemical Space, Molecular Infor- matics 41, 2200116 (2022)

  5. [5]

    J. Hoja, L. Medrano Sandonas, B. G. Ernst, A. Vazquez- Mayagoitia, R. A. DiStasio, and A. Tkatchenko, QM7-X, a comprehensive dataset of quantum-mechanical prop- erties spanning the chemical space of small organic molecules, Scientific Data 8, 43 (2021)

  6. [6]

    Ganscha, O

    S. Ganscha, O. T. Unke, D. Ahlin, H. Maennel, S. Kashu- bin, and K.-R. M¨ uller, The QCML dataset, Quan- tum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations, Scientific Data12, 406 (2025)

  7. [7]

    J. S. Smith, O. Isayev, and A. E. Roitberg, ANI-1, A data set of 20 million calculated off-equilibrium conformations for organic molecules, Scientific Data 4, 170193 (2017). 9

  8. [8]

    J. S. Smith, R. Zubatyuk, B. Nebgen, N. Lubbers, K. Barros, A. E. Roitberg, O. Isayev, and S. Tretiak, The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules, Sci- entific Data 7, 10.1038/s41597-020-0473-z (2020), pub- lisher: Springer Science and Business Media LLC

Show all 83 references
  1. [9]

    L. C. Blum and J.-L. Reymond, 970 million druglike small molecules for virtual screening in the chemical uni- verse database GDB-13, Journal of the American Chemi- cal Society 131, 8732 (2009), publisher: American Chem- ical Society (ACS) tex.timestamp: 2019-03-13

  2. [10]

    Ruddigkeit, R

    L. Ruddigkeit, R. van Deursen, L. C. Blum, and J.- L. Reymond, Enumeration of 166 billion organic small molecules in the chemical universe database GDB-17, Journal of Chemical Information and Modeling 52, 2864 (2012), publisher: American Chemical Society (ACS) tex.timestamp: 2...

  3. [11]

    Ramakrishnan, P

    R. Ramakrishnan, P. O. Dral, M. Rupp, and O. A. von Lilienfeld, Quantum chemistry structures and properties of 134 kilo molecules, Scientific Data 1, 10.1038/sdata.2014.22 (2014), publisher: Springer Na- ture tex.timestamp: 2019-03-13

  4. [12]

    Zheng, Y

    J. Zheng, Y. Zhao, and D. G. Truhlar, The DBH24/08 database and its use to assess electronic structure model chemistries for chemical reaction barrier heights, Journal of Chemical Theory and Computation5, 808 (2009), pub- lisher: American Chemical Society (ACS) tex.timestamp: ...

  5. [13]

    Balcells and B

    D. Balcells and B. B. Skjelstad, tmQM Dataset—Quantum Geometries and Properties of 86k Transition Metal Complexes, Journal of Chemical Information and Modeling 60, 6135 (2020)

  6. [14]

    Nakata, T

    M. Nakata, T. Shimazaki, M. Hashimoto, and T. Maeda, PubChemQC PM6: Data Sets of 221 Million Molecules with Optimized Molecular Geometries and Electronic Properties, Journal of Chemical Information and Mod- eling 60, 5891 (2020)

  7. [15]

    Schreiner, A

    M. Schreiner, A. Bhowmik, T. Vegge, J. Busk, and O. Winther, Transition1x - a dataset for building gen- eralizable reactive machine learning potentials, Scientific Data 9, 779 (2022)

  8. [16]

    Q. Zhao, S. M. Vaddadi, M. Woulfe, L. A. Ogunfowora, S. S. Garimella, O. Isayev, and B. M. Savoie, Comprehen- sive exploration of graphically defined reaction spaces, Scientific Data 10, 145 (2023)

  9. [17]

    Isert, K

    C. Isert, K. Atz, J. Jim´ enez-Luna, and G. Schnei- der, QMugs, quantum mechanical properties of drug-like molecules, Scientific Data 9, 273 (2022)

  10. [18]

    Ehlert, J

    S. Ehlert, J. Hermann, T. Vogels, V. G. Satorras, S. La- nius, M. Segler, D. P. Kooi, K. Takeda, C.-W. Huang, G. Luise, R. v. d. Berg, P. Gori-Giorgi, and A. Karton, Accurate Chemistry Collection: Coupled cluster atom- ization energies for broad chemical space (2025), version ...

  11. [19]

    Chakraborty, I

    S. Chakraborty, I. Almog, and R. Gershoni-Poranne, COMPAS-4: A Data Set of (BN) 1 Substituted Cata - Condensed Polybenzenoid Hydrocarbons-Data Analysis and Feature Engineering, Journal of Chemical Informa- tion and Modeling 65, 5508 (2025)

  12. [20]

    Ullah, Y

    A. Ullah, Y. Chen, and P. O. Dral, Molecular quantum chemical data sets and databases for machine learning potentials, Machine Learning: Science and Technology 5, 041001 (2024)

  13. [21]

    D. S. Levine, M. Shuaibi, E. W. C. Spotte-Smith, M. G. Taylor, M. R. Hasyim, K. Michel, I. Batatia, G. Cs´ anyi, M. Dzamba, P. Eastman, N. C. Frey, X. Fu, V. Gharakhanyan, A. S. Krishnapriyan, J. A. Rack- ers, S. Raja, A. Rizvi, A. S. Rosen, Z. Ulissi, S. Var- gas, C. L. Zitni...

  14. [22]

    Hossain, P

    S. Hossain, P. Thiagarajan, S. Pathrudkar, S. Taylor, A. S. Gangan, A. S. Banerjee, and S. Ghosh, Surprisingly High Redundancy in Electronic Structure Data (2025), version Number: 1

  15. [23]

    Lopez Perez, E

    K. Lopez Perez, E. L´ opez-L´ opez, F. Soulage, E. Fe- lix, J. L. Medina-Franco, and R. A. Miranda-Quintana, Growth vs Diversity: A Time-Evolution Analysis of the Chemical Space, Journal of Chemical Information and Modeling 65, 6788 (2025)

  16. [24]

    C. Duan, J. P. Janet, F. Liu, A. Nandy, and H. J. Kulik, Learning from failure: Predicting electronic structure cal- culation outcomes with machine learning models, Journal of Chemical Theory and Computation 15, 2331 (2019), publisher: American Chemical Society (ACS)

  17. [25]

    Heinen, M

    S. Heinen, M. Schwilk, G. F. von Rudorff, and O. A. von Lilienfeld, Machine learning the computa- tional cost of quantum chemistry, Machine Learning: Science and Technology 1, 025002 (2019), arXiv: http://arxiv.org/abs/1908.06714v1 [physics.chem-ph] Publisher: IOP Publishing t...

  18. [26]

    G. S. Hammond, A Correlation of Reaction Rates, Jour- nal of the American Chemical Society 77, 334 (1955), publisher: American Chemical Society (ACS)

  19. [27]

    H¨ uckel, Quantentheoretische Beitr¨ age zum Benzol- problem: I

    E. H¨ uckel, Quantentheoretische Beitr¨ age zum Benzol- problem: I. Die Elektronenkonfiguration des Benzols und verwandter Verbindungen, Zeitschrift f¨ ur Physik70, 204 (1931), publisher: Springer Science and Business Media LLC

  20. [28]

    Heinen, G

    S. Heinen, G. F. von Rudorff, and O. A. von Lilienfeld, Toward the design of chemical reactions: Machine learn- ing barriers of competing mechanisms in reactant space, The Journal of Chemical Physics 155, 064105 (2021), publisher: AIP Publishing

  21. [29]

    G. A. Arteca and P. G. Mezey, Validity of the Hammond postulate and constraints on general one-dimensional re- action barriers, Journal of Computational Chemistry 9, 728 (1988), publisher: Wiley

  22. [30]

    N. M. Donahue, Revisiting the Hammond Postulate: The Role of Reactant and Product Ionic States in Regulat- ing Barrier Heights, Locations, and Transition State Fre- quencies, The Journal of Physical Chemistry A105, 1489 (2001), publisher: American Chemical Society (ACS)

  23. [31]

    Van Nyvel, M

    L. Van Nyvel, M. Alonso, and M. Sol` a, Effect of size, charge, and spin state on H¨ uckel and Baird aromatic- ity in [ N ]annulenes, Chemical Science 16, 5613 (2025), publisher: Royal Society of Chemistry (RSC)

  24. [32]

    Y. B. Apriliyanto, S. Battaglia, S. Evangelisti, N. Faginas-Lago, T. Leininger, and A. Lombardi, Toward a Generalized H¨ uckel Rule: The Electronic Structure of Carbon Nanocones, The Journal of Physical Chemistry A 125, 9819 (2021), publisher: American Chemical Society (ACS)

  25. [33]

    Huang, G

    B. Huang, G. F. von Rudorff, and O. A. von Lilienfeld, The central role of density functional theory in the AI age, Science 381, 170 (2023), publisher: American Asso- ciation for the Advancement of Science (AAAS)

  26. [34]

    Lavecchia, Navigating the frontier of drug-like chem- 10 ical space with cutting-edge generative AI models, Drug Discovery Today 29, 104133 (2024), publisher: Elsevier BV

    A. Lavecchia, Navigating the frontier of drug-like chem- 10 ical space with cutting-edge generative AI models, Drug Discovery Today 29, 104133 (2024), publisher: Elsevier BV

  27. [35]

    A. V. Sadybekov and V. Katritch, Computational ap- proaches streamlining drug discovery, Nature 616, 673 (2023), publisher: Springer Science and Business Media LLC

  28. [36]

    C. D. Griego, A. M. Maldonado, L. Zhao, B. Zulueta, B. M. Gentry, E. Lipsman, T. H. Choi, and J. A. Keith, Computationally guided searches for efficient catalysts through chemical/materials space: Progress and outlook, The Journal of Physical Chemistry C 125, 6495 (2021), publ...

  29. [37]

    Z. Tan, Q. Yang, and S. Luo, AI molecular cataly- sis: where are we now?, Organic Chemistry Frontiers 12, 2759 (2025), publisher: Royal Society of Chemistry (RSC)

  30. [38]

    S. Mace, Y. Xu, and B. N. Nguyen, Automated Transition Metal Catalysts Discovery and Optimisation with AI and Machine Learning, ChemCatChem 16, 10.1002/cctc.202301475 (2024), publisher: Wiley

  31. [39]

    G´ omez-Bombarelli, J

    R. G´ omez-Bombarelli, J. N. Wei, D. Duvenaud, J. M. Hern´ andez-Lobato, B. S´ anchez-Lengeling, D. Sheberla, J. Aguilera-Iparraguirre, T. D. Hirzel, R. P. Adams, and A. Aspuru-Guzik, Automatic chemical design us- ing a data-driven continuous representation of molecules, ACS C...

  32. [40]

    X. Cao, Y. Zhang, Z. Sun, H. Yin, and Y. Feng, Ma- chine learning in polymer science: A new lens for physical and chemical exploration, Progress in Materials Science , 101544 (2025)

  33. [41]

    Liu and Z

    Y. Liu and Z. Li, Predict Ionization Energy of Molecules Using Conventional and Graph-Based Machine Learning Models, Journal of Chemical Information and Modeling 63, 806 (2023)

  34. [42]

    M. Yang, G. Song, L. Cheng, and H. Ren, SMILES To- ken Additivity Model with Interpretability and General- izability for Fuel Property Predictions, Journal of Chem- ical Information and Modeling , acs.jcim.5c00986 (2025)

  35. [43]

    J. Sun, Y. Cao, H. Hu, and B. Qi, Equivariant learn- ing leveraging geometric invariances in 3D molecular conformers for accurate prediction of quantum chemi- cal properties, Scientific Reports15, 10.1038/s41598-025- 09842-x (2025), publisher: Springer Science and Business Media LLC

  36. [44]

    Suman, J

    D. Suman, J. Nigam, S. Saade, P. Pegolo, H. T¨ urk, X. Zhang, G. K.-L. Chan, and M. Ceriotti, Explor- ing the Design Space of Machine Learning Models for Quantum Chemistry with a Fully Differentiable Frame- work, Journal of Chemical Theory and Computation 10.1021/acs.jctc.5c00...

  37. [45]

    J. S. Smith, O. Isayev, and A. E. Roitberg, ANI-1: an ex- tensible neural network potential with DFT accuracy at force field computational cost, Chemical Science 8, 3192 (2017), publisher: Royal Society of Chemistry (RSC)

  38. [46]

    Stuke, M

    A. Stuke, M. Todorovi´ c, M. Rupp, C. Kunkel, K. Ghosh, L. Himanen, and P. Rinke, Chemical diversity in molecu- lar orbital energy predictions with kernel ridge regression, The Journal of Chemical Physics150, 10.1063/1.5086105 (2019), publisher: AIP Publishing

  39. [47]

    Serrano-Morr´ as, A

    A. Serrano-Morr´ as, A. Bertran-Mostazo, M. Mi˜ narro Lleonar, A. Comajuncosa-Creus, A. Cabello, C. Labranya, C. Escudero, T. V. Tian, I. Khutori- anska, D. S. Radchenko, Y. S. Moroz, L. Defelipe, D. Ruiz-Carrillo, M. Garcia-Alai, R. Schmidt, M. Rarey, P. Aloy, C. Galdeano, J....

  40. [48]

    Vogt, Exploring chemical space — Generative models and their evaluation, Artificial Intelligence in the Life Sciences 3, 100064 (2023)

    M. Vogt, Exploring chemical space — Generative models and their evaluation, Artificial Intelligence in the Life Sciences 3, 100064 (2023)

  41. [49]

    D. M. Anstine and O. Isayev, Generative Models as an Emerging Paradigm in the Chemical Sciences, Journal of the American Chemical Society 145, 8736 (2023), pub- lisher: American Chemical Society (ACS)

  42. [50]

    Langevin, C

    M. Langevin, C. Grebner, S. G¨ ussregen, S. Sauer, Y. Li, H. Matter, and M. Bianciotto, Impact of Applicabil- ity Domains to Generative Artificial Intelligence, ACS Omega 8, 23148 (2023), publisher: American Chemical Society (ACS)

  43. [51]

    Sosnin, Chemical space visual navigation in the era of deep learning and Big Data, Drug Discovery Today 30, 104392 (2025)

    S. Sosnin, Chemical space visual navigation in the era of deep learning and Big Data, Drug Discovery Today 30, 104392 (2025)

  44. [52]

    C. Duan, J. P. Janet, F. Liu, A. Nandy, and H. J. Kulik, Learning from Failure: Predicting Electronic Structure Calculation Outcomes with Machine Learning Models, Journal of Chemical Theory and Computation 15, 2331 (2019)

  45. [53]

    C. Bugg, R. Desiderato, and R. L. Sass, An X-Ray Diffraction Study of Nonplanar Carbanion Structures, Journal of the American Chemical Society 86, 3157 (1964)

  46. [54]

    Malischewski and K

    M. Malischewski and K. Seppelt, Crystal Structure De- termination of the Pentagonal-Pyramidal Hexamethyl- benzene Dication C6 (CH3 )6 2+, Angewandte Chemie In- ternational Edition 56, 368 (2017)

  47. [55]

    Khriachtchev, M

    L. Khriachtchev, M. Pettersson, N. Runeberg, J. Lundell, and M. R¨ as¨ anen, A stable argon compound, Nature406, 874 (2000)

  48. [56]

    W. Qian, A. Mardyukov, and P. R. Schreiner, Prepara- tion of a neutral nitrogen allotrope hexanitrogen C2h-N6, Nature 642, 356 (2025)

  49. [57]

    Y. Gao, P. Gupta, I. Ronˇ cevi´ c, C. Mycroft, P. J. Gates, A. W. Parker, and H. L. Anderson, Solution-phase stabi- lization of a cyclocarbon by catenane formation, Science 389, 708 (2025)

  50. [58]

    Benecke, T

    C. Benecke, T. Gr¨ uner, A. Kerber, R. Laue, and T. Wieland, MOLecular structure GENeration with MOLGEN, new features and future developments, Fre- senius Journal of Analytical Chemistry 359, 23 (1997), publisher: Springer Science and Business Media LLC

  51. [59]

    J. E. Peironcely, M. Rojas-Chert´ o, D. Fichera, T. Rei- jmers, L. Coulier, J.-L. Faulon, and T. Hankemeier, OMG: Open molecule generator, Journal of Chemin- formatics 4, 10.1186/1758-2946-4-21 (2012), publisher: Springer Science and Business Media LLC

  52. [60]

    M. M. Jaghoori, S.-S. T. Jongmans, F. De Boer, J. Peironcely, J.-L. Faulon, T. Reijmers, and T. Hanke- meier, PMG: Multi-core Metabolite Identification, Elec- tronic Notes in Theoretical Computer Science 299, 53 (2013)

  53. [61]

    B. D. McKay, M. A. Yirik, and C. Steinbeck, Surge: a fast open-source chemical graph generator, Journal of Chem- informatics 14, 10.1186/s13321-022-00604-9 (2022), pub- lisher: Springer Science and Business Media LLC. 11

  54. [62]

    M. A. Yirik, M. Sorokina, and C. Steinbeck, MAYGEN: an open-source chemical structure generator for consti- tutional isomers based on the orderly generation princi- ple, Journal of Cheminformatics 13, 10.1186/s13321-021- 00529-9 (2021), publisher: Springer Science and Business...

  55. [63]

    S. R. Rieder, M. P. Oliveira, S. Riniker, and P. H. H¨ unenberger, Development of an open-source software for isomer enumeration, Journal of Cheminformatics 15, 10 (2023)

  56. [64]

    Faulon, Stochastic Generator of Chemical Struc- ture

    J.-L. Faulon, Stochastic Generator of Chemical Struc- ture. 1. Application to the Structure Elucidation of Large Molecules, Journal of Chemical Information and Com- puter Sciences 34, 1204 (1994)

  57. [65]

    Gasevic, M

    T. Gasevic, M. M¨ uller, J. Sch¨ ops, S. Lanius, J. Hermann, S. Grimme, and A. Hansen, Chemical Space Exploration with Artificial ”Mindless” Molecules (2025)

  58. [66]

    Karandashev, J

    K. Karandashev, J. Weinreich, S. Heinen, D. J. Aris- mendi Arrieta, G. F. von Rudorff, K. Hermansson, and O. A. von Lilienfeld, Evolutionary monte carlo of QM properties in chemical space: Electrolyte design, Journal of Chemical Theory and Computation 19, 8861 (2023), publishe...

  59. [67]

    Krenn, Q

    M. Krenn, Q. Ai, S. Barthel, N. Carson, A. Frei, N. C. Frey, P. Friederich, T. Gaudin, A. A. Gayle, K. M. Jablonka, R. F. Lameiro, D. Lemm, A. Lo, S. M. Moosavi, J. M. N´ apoles-Duarte, A. Nigam, R. Pollice, K. Rajan, U. Schatzschneider, P. Schwaller, M. Skreta, B. Smit, F. St...

  60. [68]

    D. J. Watts and S. H. Strogatz, Collective dynamics of ‘small-world’ networks, Nature 393, 440 (1998)

  61. [69]

    Greenhill and B

    C. Greenhill and B. D. McKay, Asymptotic enumeration of sparse multigraphs with given degrees, SIAM Journal on Discrete Mathematics 27, 2064 (2013), publisher: So- ciety for Industrial & Applied Mathematics (SIAM)

  62. [70]

    G. A. Croes, A Method for Solving Traveling-Salesman Problems, Operations Research 6, 791 (1958), publisher: INFORMS

  63. [71]

    Abu-Aisheh, R

    Z. Abu-Aisheh, R. Raveaux, J.-Y. Ramel, and P. Mar- tineau, An Exact Graph Edit Distance Algorithm for Solving Pattern Recognition Problems:, in Proceedings of the International Conference on Pattern Recognition Applications and Methods (SCITEPRESS - Science and and Technology...

  64. [72]

    Banjafar, D

    A. Banjafar, D. J. Monterrubio-Chanca, and G. F. von Rudorff, NablaChem/nablachem: v25.8 (2025)

  65. [73]

    C. Greenhill, The switch Markov chain for sampling ir- regular graphs (Extended Abstract), inProceedings of the Twenty-Sixth Annual ACM-SIAM Symposium on Dis- crete Algorithms (Society for Industrial and Applied Mathematics, 2015) pp. 1564–1572

  66. [74]

    jamesross2, jamesross2/random graph (2023), original- date: 2020-04-05T22:22:04Z

  67. [75]

    J. S. Smith, O. Isayev, and A. E. Roitberg, ANI-1, A data set of 20 million calculated off-equilibrium conformations for organic molecules, Scientific Data 4, 170193 (2017)

  68. [76]

    L. C. Blum and J.-L. Reymond, 970 Million Druglike Small Molecules for Virtual Screening in the Chemical Universe Database GDB-13, Journal of the American Chemical Society 131, 8732 (2009)

  69. [77]

    F. J. Massey, The Kolmogorov-Smirnov Test for Good- ness of Fit, Journal of the American Statistical Associa- tion 46, 68 (1951), publisher: JSTOR

  70. [78]

    Kullback and R

    S. Kullback and R. A. Leibler, On Information and Suf- ficiency, The Annals of Mathematical Statistics 22, 79 (1951), publisher: Institute of Mathematical Statistics

  71. [79]

    D. N. Rassokhin and D. K. Agrafiotis, Kolmogorov- Smirnov statistic and its application in library design, Journal of Molecular Graphics and Modelling 18, 368 (2000)

  72. [80]

    S. Ji, Z. Zhang, S. Ying, L. Wang, X. Zhao, and Y. Gao, Kullback–Leibler Divergence Metric Learning, IEEE Transactions on Cybernetics 52, 2047 (2022)

  73. [81]

    C. L. McClendon, L. Hua, G. Barreiro, and M. P. Jacobson, Comparing Conformational Ensembles Using the Kullback–Leibler Divergence Expansion, Journal of Chemical Theory and Computation 8, 2115 (2012)

  74. [82]

    Lagarde, J.-F

    N. Lagarde, J.-F. Zagury, and M. Montes, Benchmarking Data Sets for the Evaluation of Virtual Ligand Screening Methods: Review and Perspectives, Journal of Chemical Information and Modeling 55, 1297 (2015), publisher: American Chemical Society (ACS)

  75. [83]

    E. R. Antoniuk, S. Zaman, T. Ben-Nun, P. Li, J. Diffend- erfer, B. Demirci, O. Smolenski, T. Hsu, A. M. Hiszpan- ski, K. Chiu, B. Kailkhura, and B. Van Essen, BOOM: Benchmarking Out-Of-distribution Molecular Property Predictions of Machine Learning Models (2025), version Number: 1

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.