Pith. sign in

REVIEW 4 major objections 6 minor 60 references

A unified 800K-mixture benchmark plus soft physical constraints lifts chemical property prediction accuracy and out-of-distribution reliability.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 18:28 UTC pith:7XH7HGRL

load-bearing objection Useful mixture benchmark packaging plus soft physics regularizers that help on many fluid tracks; global “SOTA by median RMSE” and fixed T-monotonicity signs are the soft spots, not the core resource. the 4 major comments →

arxiv 2607.28079 v1 pith:7XH7HGRL submitted 2026-07-30 cs.LG cs.AI

Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction

classification cs.LG cs.AI
keywords BenchmarkDatasetAI4ChemistryMolecular ScienceMaterials InformaticsPhysics-Informed Neural NetworksMixture property predictionChem World
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Chem World argues that chemical AI has been held back by fragmented mixture datasets and purely statistical models that ignore physical priors. The paper unifies 17 public sources into one benchmark of more than 800,000 samples across ten property tracks (conductivity, density, solubility, viscosity, mixing enthalpy, and others), with a common schema for components, fractions, and conditions. On that platform it introduces Mixture-PINN: pretrained molecular encoders feed an interaction-aware mixture aggregator that is regularized by composition alignment, pairwise interaction symmetry, temperature-monotonicity and smoothness where applicable, and physical output bounds. Across encoder–aggregator baselines, Mixture-PINN posts the best median errors and R², wins most property tracks, and degrades less under scaffold-disjoint, component-disjoint, and composition-extrapolation splits. The claim is that standardized multi-property evaluation plus lightweight chemical constraints is a practical path to more trustworthy mixture prediction for materials, electrolytes, fuels, and formulation work.

Core claim

On Chem World, coupling pretrained molecular representations with Mixture-PINN’s composition, interaction-symmetry, thermodynamic-monotonicity, and boundary losses yields state-of-the-art mixture property prediction versus strong encoder–aggregator baselines (median RMSE 0.743, median R² 0.918 with MolFormer), tops seven of ten tracks, and retains the best R² under three out-of-distribution splits—showing that soft physical regularizers improve accuracy, stability, and generalization without extra simulation labels.

What carries the argument

Mixture-PINN: a mixture representation that combines composition-weighted component embeddings with pairwise interaction terms, trained under a composite loss that adds composition consistency (attention weights aligned to mole fractions), interaction symmetry, optional temperature-derivative monotonicity/smoothness for conductivity and viscosity, and non-negativity-style physical bounds to the usual supervised error.

Load-bearing premise

The soft priors—especially fixed temperature-gradient signs for conductivity and viscosity, and forcing learned attention weights to match mole fractions—must be generally valid across the heterogeneous systems merged into each track, or the claimed superiority and trustworthiness shrink.

What would settle it

On held-out mixture families where conductivity falls with temperature or viscosity rises with temperature, or where component contribution is strongly non-stoichiometric, check whether Mixture-PINN’s ranking versus Set Transformer / DeepSets baselines collapses and whether boundary/monotonicity terms increase error rather than reduce it.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Mixture property models can be compared on one shared schema and ten tracks instead of incompatible single-dataset protocols.
  • Adding composition, symmetry, bound, and simple T-derivative constraints is a scalable default when full governing equations are unavailable.
  • OOD splits (unseen scaffolds, unseen components, composition shift) become a standard bar for “trustworthy” chemical AI claims.
  • Transport and mixing properties benefit most; solid-phase transitions may still favor molecule-centric tree models.
  • Open benchmark and agent tooling lower the cost of multi-property screening for electrolytes, solvents, fuels, and formulations.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If median metrics across differently scaled tracks drive the leaderboard, future work may need per-track normalized or multi-objective ranking to avoid scale dominance.
  • The same residual-plus-prior pattern could transfer to other set-valued scientific inputs (alloys, catalyst formulations, food matrices) where ideal mixing is a useful baseline.
  • Preserving lab-to-lab label noise rather than aggressive averaging may make uncertainty-aware heads a natural next model layer on this benchmark.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces Chem World, a unified benchmark that merges 17 public chemical datasets (>800K instances) into ten mixture-property prediction tracks (conductivity, density, solubility, viscosity, etc.), with standardized schemas, preprocessing, and both random and OOD splits. It further proposes Mixture-PINN, which encodes components with pretrained molecular models (e.g., MolFormer/MolT5), aggregates them with interaction-aware mixing, and trains under soft physics regularizers: composition–attention alignment (Eq. 2), pairwise interaction symmetry (Eq. 3), temperature-monotonicity/smoothness for conductivity and viscosity (Eqs. 4–6), and non-negativity bounds (Eq. 7). On the leaderboard (Tables 2–3, 8), MolFormer/MolT5 + Mixture-PINN reports the best median metrics and wins 7/10 tracks; it also leads under scaffold-, component-, and composition-OOD splits (Table 4), with ablations (Table 5) attributing gains to the constraint suite.

Significance. If the benchmark and method claims hold under fairer multi-track aggregation and better-justified priors, this is a useful contribution to AI-for-chemistry: a large, mixture-centric evaluation suite with open resources, plus evidence that lightweight chemical regularizers improve accuracy and OOD behavior over strong encoder–aggregator baselines. Strengths include the scale and diversity of Chem World, explicit OOD protocols, a full encoder×aggregator matrix (Table 8), constraint ablations, and released code/demo. The work is not definitionally circular—models are scored on held-out experimental labels. The main value is infrastructure plus a practical physics-regularized mixture architecture rather than a new physical theory.

major comments (4)
  1. [§5.4.1 Leaderboard / Table 2; cf. Table 5] Table 2 ranks methods by median raw RMSE/MAE/R² across ten heterogeneous tracks. Table 5 shows a large mean–median gap under the full model (mean RMSE 2.843 vs median 0.743), so the median can overweight mid-scale transport wins while down-weighting high-error or differently scaled tracks (e.g., olfactory, fusion/melting in Tables 3 and 8). NRMSE is defined in §5.3 but not used for the global ranking. Please add scale-normalized metrics (per-track z-scored/NRMSE targets and/or average rank over the 10 tracks) and report both mean and median; without that, the global “SOTA / trustworthy” claim in §5.4.1 is not fully supported.
  2. [§4.3 Eqs. (4)–(5); conductivity/viscosity tracks] Eqs. 4–5 impose fixed signs ∂ŷ/∂T ≥ 0 for all conductivity tasks and ≤ 0 for all viscosity tasks across merged systems (ILs, polymer electrolytes, DES, solvents, etc.). The manuscript does not validate that these signs hold on the Chem World subsets (e.g., fraction of samples with empirical positive/negative d(property)/dT, or failure cases). If non-monotonic or opposite-signed regimes are common, L_thermo is a harmful prior and attribution of OOD/reliability gains to “physics” (§5.4.2–5.4.3) is insecure. Please quantify empirical temperature trends per track/source and either restrict the constraint, make the sign data-adaptive, or show robustness when the prior is misspecified.
  3. [§4.2–4.3 vs Appendix B] The main-text Mixture-PINN (§4.2, Eq. 1) is an attention-weighted sum of component and pairwise embeddings, while Appendix B describes a different architecture: ideal-mixture prior + learned residual correction, FiLM conditioning, and optional Arrhenius/VFT heads. These are not clearly reconciled, so it is unclear which model produced Tables 2–5 and 8. Please unify the method description with the implemented model and state exactly which components were ablated in Table 5.
  4. [§5.1.3 vs Appendix A.3.1; §5.4.3] Split protocol is inconsistent: §5.1.3 states random 8:1:1 (unless official splits), while Appendix A.3.1 specifies shuffled 5-fold CV (~70/10/20) as the default leaderboard protocol. OOD constructions also differ slightly between §5.4.3 and the appendix. Reproducibility of the headline numbers requires one canonical protocol, fold-level variance, and clarification whether Table 2 medians pool folds, tracks, or both.
minor comments (6)
  1. [§4.3 Eq. (8)] Eq. (8) contains a double plus (“++λ4”); fix typography.
  2. [Abstract; §1; §3] Abstract/intro say “over 800,000 molecular samples” while the body emphasizes mixture-property records; align terminology (molecules vs mixture instances).
  3. [Table 1; §2] Table 1 comparison to CheMixHub and related mixture benchmarks is brief; expand coverage, licensing, and protocol differences so novelty of Chem World is clearer.
  4. [§4.3 Eq. (2)] L_comp (Eq. 2) forces attention toward mole fractions; discuss when mass/volume fractions or activity-based weights are more physical, and sensitivity to that choice.
  5. [Fig. 3; Fig. 4] Figure 3 and Figure 4 are hard to read in grayscale; improve legends and axis labels for property units after log transforms.
  6. [§4.3; §5] Report λ weights and whether they were tuned per track or fixed globally; free parameters listed implicitly affect reproducibility.

Circularity Check

0 steps flagged

No derivation circularity: Mixture-PINN’s test metrics are empirical supervised outcomes plus soft priors, not quantities forced by construction from their own inputs.

full rationale

The paper’s load-bearing claim is empirical: on held-out and OOD splits of Chem World (external experimental sources merged under a unified schema), MolFormer/MolT5 + Mixture-PINN beats encoder–aggregator baselines (Tables 2–4), with ablations attributing gains to soft losses L_comp, L_int, L_thermo, and L_bound (Eqs. 2–8, Table 5). Those losses regularize attention, pairwise symmetry, temperature-derivative sign/smoothness, and output positivity; they do not algebraically define ŷ or the reported RMSE/R². Training still minimizes L_data = ||y − ŷ||² on labeled mixtures; evaluation is on disjoint splits, so test numbers are not fitted inputs renamed as predictions. Chem World is author-built, which is ordinary benchmark practice and does not make the ranking tautological. There is no self-citation uniqueness theorem, no ansatz smuggled in as a forced law, and no step where a claimed first-principles result reduces to its own definition. Weaknesses (fixed ∂ŷ/∂T signs, median-raw ranking across heterogeneous scales) are correctness/evaluation-design issues, not circular derivation.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 2 invented entities

The central empirical claim rests on public experimental labels as ground truth, a unified but author-defined preprocessing/split regime, pretrained encoders as frozen/feature backbones, and four soft physics penalties with tunable weights. No new physical law is derived; trustworthiness is operationalized as lower error plus constraint satisfaction. Free parameters are the loss weights and architectural scales; axioms are standard mixture bookkeeping plus domain monotonicity/symmetry assumptions; invented entities are the benchmark product and the Mixture-PINN stack.

free parameters (4)
  • λ1, λ2, λ3, λ4 (loss weights on L_comp, L_int, L_thermo, L_bound)
    Eq. 8 balances data fit against physics penalties; values are not reported but directly affect the claimed gains and ablations.
  • λ_smooth (second-derivative temperature smoothness weight)
    Controls L_smooth inside L_thermo (§4.3); chosen by authors, not fixed by theory.
  • α residual correction scale (Appendix B) = described as a small correction scale
    Scales non-ideal residual on top of ideal mixing embedding; hand/architecture choice that changes capacity of Mixture-PINN.
  • Encoder and aggregator hyperparameters (MolFormer/MolT5/GNN depth, attention heads, training seeds)
    Leaderboard ranks depend on these implementation choices; only high-level model names are fixed in text.
axioms (6)
  • domain assumption Mixture samples can be represented as M={(m_i, x_i)} with normalized fractions and optional (T,P), and labels from heterogeneous literature are comparable after unit/log normalization.
    §3.2–3.3 and Appendix A.3 schema; underpins merging 17 datasets into shared tracks.
  • ad hoc to paper Learned component attention α should align with mole fractions x (L_comp = ||α−x||²).
    §4.3 Eq. 2; useful inductive bias but not required by thermodynamics (activity can decouple from x).
  • domain assumption Pairwise interaction embeddings are order-symmetric (e_ij ≈ e_ji).
    §4.3 Eq. 3; standard reciprocity prior for undirected interactions.
  • domain assumption For conductivity tasks ∂ŷ/∂T ≥ 0; for viscosity tasks ∂ŷ/∂T ≤ 0; second derivatives should be small.
    §4.3 Eqs. 4–6; broadly true for many liquids but not guaranteed for all chemistries/conditions in merged tracks.
  • domain assumption Target properties obey simple bound constraints (e.g., positivity) enforceable by hinge penalties.
    §4.3 Eq. 7 L_bound.
  • ad hoc to paper Median metrics across ten property tracks are a valid overall ranking of mixture intelligence methods.
    Table 2 leaderboard construction; contested when units and difficulties differ.
invented entities (2)
  • Chem World benchmark (unified 17-dataset, 10-track mixture suite) independent evidence
    purpose: Provide standardized multi-property evaluation and OOD protocols for chemical mixture models.
    Author-curated aggregation of existing public datasets with new schema/splits; value is infrastructural.
  • Mixture-PINN framework no independent evidence
    purpose: Combine molecular encoders, interaction-aware aggregation, and soft physics losses for mixture regression.
    Named architecture/objective bundle; not a new physical particle or field, but a postulated modeling stack evaluated only inside this paper’s benchmark.

pith-pipeline@v1.2.0-daily-grok45 · 27410 in / 4286 out tokens · 101405 ms · 2026-07-31T18:28:57.997004+00:00 · methodology

0 comments
read the original abstract

Chemical property prediction plays a critical role in accelerating scientific discovery in chemistry, materials science, and drug development. However, existing benchmarks often suffer from limited task diversity, fragmented datasets, and inconsistent evaluation protocols, making it challenging to systematically assess the reliability and generalization of AI models. In this work, we introduce Chem World, a comprehensive benchmark for chemical property prediction that integrates 17 diverse chemical datasets with over 800,000 molecular samples, covering various properties including density, electrical conductivity, solubility, and other molecular characteristics. Chem World provides a unified platform for evaluating AI models across multiple property prediction tasks. Furthermore, we propose Mixture-PINN, a physics-informed neural network based prediction framework that incorporates chemical prior knowledge into data-driven learning, improving the accuracy, robustness, and reliability of chemical property prediction. Extensive experiments on Chem World demonstrate the effectiveness of our approach compared with existing methods. By combining large-scale standardized evaluation with physics-informed learning, Chem World establishes a foundation for developing trustworthy AI systems for computational chemistry and advancing AI-driven scientific discovery.

Figures

Figures reproduced from arXiv: 2607.28079 by Fangyue Lin, Huan Wang, Mingchen Gao, Pinze Ren, Siming Dong, Tianyou Bai, Zhenlin Zhao.

Figure 1
Figure 1. Figure 1: Pipeline of Chem World and Mixture-PINN. Chem [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Illustration of the Chem World benchmark and Mixture-PINN. Chem World provides a standardized large-scale testbed [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Dataset statistics and structural composition of [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Full per-track performance of all tested models. [PITH_FULL_IMAGE:figures/full_fig_p016_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

60 extracted references · 7 linked inside Pith

  1. [1]

    Masoud Amiri. 2026. Physics-informed deep learning for molecular solubility prediction: integrating thermodynamic constraints with neural network archi- tectures.Scientific reports(2026)

  2. [2]

    Mohammad Reza Babaei, Ryan Stone, Thomas Allen Knotts Iv, and John Heden- gren. 2023. Physics-informed neural networks with group contribution methods. Journal of Chemical Theory and Computation19, 13 (2023), 4163–4171

  3. [3]

    Zeqing Bao, Gary Tom, Austin Cheng, Jeffrey Watchorn, Alán Aspuru-Guzik, and Christine Allen. 2024. Towards the prediction of drug solubility in binary solvent mixtures at various temperatures using machine learning.Journal of Cheminformatics16, 1 (2024), 117

  4. [4]

    Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al. 2018. Relational inductive biases, deep learning, and graph networks.arXiv preprint arXiv:1806.01261(2018)

  5. [5]

    Simon Batzner, Albert Musaelian, Lixin Sun, Mario Geiger, Jonathan P Mailoa, Mordechai Kornbluth, Nicola Molinari, Tess E Smidt, and Boris Kozinsky. 2022. E (3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials.Nature communications13, 1 (2022), 2453

  6. [6]

    Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction KDD, August 01–05, 2027, San Jose

    Gabriel Bradford, Jeffrey Lopez, Jurgis Ruza, Michael A Stolberg, Richard Os- terude, Jeremiah A Johnson, Rafael Gomez-Bombarelli, and Yang Shao-Horn. Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction KDD, August 01–05, 2027, San Jose

  7. [7]

    Tianqi Chen and Carlos Guestrin. 2016. Xgboost: A scalable tree boosting system. InProceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining. 785–794

  8. [8]

    Alex K Chew, Mohammad Atif Faiz Afzal, Zachary Kaplan, Eric M Collins, Suraj Gattani, Mayank Misra, Anand Chandrasekaran, Karl Leswing, and Mathew D Halls. 2025. Leveraging high-throughput molecular simulations and machine learning for the design of chemical mixtures.npj Computational Materials11, 1 (2025), 72

  9. [9]

    Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. 2020. ChemBERTa: large-scale self-supervised pretraining for molecular property prediction.arXiv preprint arXiv:2010.09885(2020)

  10. [10]

    Stefan Chmiela, Huziel E Sauceda, Klaus-Robert Müller, and Alexandre Tkatchenko. 2018. Towards exact molecular dynamics simulations with machine- learned force fields.Nature communications9, 1 (2018), 3887

  11. [11]

    Paolo de Blasio, Jonas Elsborg, Tejs Vegge, Eibar Flores, and Arghya Bhowmik

  12. [12]

    Carl Edwards, Tuan Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji

  13. [13]

    Adroit TN Fajar, Takafumi Hanada, Aditya D Hartono, and Masahiro Goto. 2024. Estimating the phase diagrams of deep eutectic solvents within an extensive chemical space.Communications Chemistry7, 1 (2024), 27

  14. [14]

    Johannes Gasteiger, Shankari Giri, Johannes T Margraf, and Stephan Günne- mann. 2020. Fast and uncertainty-aware directional message passing for non- equilibrium molecules.arXiv preprint arXiv:2011.14115(2020)

  15. [15]

    Johannes Gasteiger, Janek Groß, and Stephan Günnemann. 2020. Directional message passing for molecular graphs.arXiv preprint arXiv:2003.03123(2020)

  16. [16]

    Mojtaba Haghighatlari, Jie Li, Xingyi Guan, Oufan Zhang, Akshaya Das, Christo- pher J Stein, Farnaz Heidar-Zadeh, Meili Liu, Martin Head-Gordon, Luke Bertels, et al. 2022. NewtonNet: a Newtonian message passing network for deep learning of interatomic potentials and forces.Digital Discovery1, 3 (2022), 333–343

  17. [17]

    Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2020. Open graph benchmark: Datasets for machine learning on graphs.Advances in neural information processing systems 33 (2020), 22118–22133

  18. [18]

    Thelma Anizia Ihunde and Olufemi Olorode. 2022. Application of physics in- formed neural networks to compositional modeling.Journal of Petroleum Science and Engineering211 (2022), 110175

  19. [19]

    Anubhav Jain, Shyue Ping Ong, Geoffroy Hautier, Wei Chen, William Davidson Richards, Stephen Dacek, Shreyas Cholia, Dan Gunter, David Skinner, Gerbrand Ceder, et al. 2013. Commentary: The Materials Project: A materials genome approach to accelerating materials innovation.APL materials1, 1 (2013)

  20. [20]

    Soheil Kavian, Arian Zarriz, and Matthew J Powell-Palm. 2026. A large-scale dataset and physics-informed neural network for viscosity prediction in many- component aqueous and organic solutions.The Journal of Chemical Physics164, 8 (2026)

  21. [21]

    A Kazakov, Chris D Muzny, K Kroenlein, Vladimir Diky, Robert D Chirico, Joe W Magee, Ilmutdin M Abdulagatov, and M Frenkel. 2012. NIST/TRC source data archival system: The next-generation data model for storage of thermophysical properties.International Journal of Thermophysics33, 1 (2012), 22–33

  22. [22]

    Brian Kelley et al. 2024. GitHub-bp-kelley/descriptastorus: Descriptor computa- tion (chemistry) and (optional) storage for machine learning. (2024)

  23. [23]

    Lev Krasnov, Dmitry Malikov, Marina Kiseleva, Sergei Tatarin, Sergey Sosnin, and Stanislav Bezzubov. 2025. BigSolDB 2.0, dataset of solubility values for organic compounds in different solvents at various temperatures.Scientific Data12, 1 (2025), 1236

  24. [24]

    Ritesh Kumar, Minh Canh Vu, Peiyuan Ma, and Chibueze V Amanchukwu. 2025. Electrolytomics: a unified big data approach for electrolyte design and discovery. Chemistry of Materials37, 8 (2025), 2720–2734

  25. [25]

    Nursulu Kuzhagaliyeva, Samuel Horváth, John Williams, Andre Nicolle, and S Mani Sarathy. 2022. Artificial intelligence-driven design of fuel mixtures. Communications Chemistry5, 1 (2022), 111

  26. [26]

    Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. 2019. Set transformer: A framework for attention-based permutation-invariant neural networks. InInternational conference on machine learning. PMLR, 3744–3753

  27. [27]

    Dan Li, Xuena Zhang, Chunling Xin, and Meifang Liu. 2025. Thermophysical and excess properties of binary mixtures of dibutyl ether and components of biodiesel.Journal of Solution Chemistry54, 1 (2025), 125–139

  28. [28]

    Zongyi Li, Hongkai Zheng, Nikola Kovachki, David Jin, Haoxuan Chen, Burigede Liu, Kamyar Azizzadenesheli, and Anima Anandkumar. 2024. Physics-informed neural operator for learning partial differential equations.ACM/IMS Journal of Data Science1, 3 (2024), 1–27

  29. [29]

    Dmitry Malikov, Lev Krasnov, Marina Kiseleva, Elizaveta Meshcheriakova, Fedor Kuznetsov, Vladimir Elistratov, Matvei Vasiyarov, Sergei Tatarin, and Stanislav Bezzubov. 2026. Dataset of solubility values for organic compounds in binary mixtures of solvents at various temperatures.Scientific Data13, 1 (2026), 727

  30. [30]

    Valeria Odegova, Anastasia Lavrinenko, Timur Rakhmanov, George Sysuev, An- drei Dmitrenko, and Vladimir Vinogradov. 2024. DESignSolvents: an open plat- form for the search and prediction of the physicochemical properties of deep eutectic solvents.Green Chemistry26, 7 (2024), 3958–3967

  31. [31]

    U Onken, J Rarey-Nies, and J Gmehling. 1989. The Dortmund Data Bank: A computerized system for retrieval, correlation, and prediction of thermodynamic properties of mixtures.International Journal of Thermophysics10, 3 (1989), 739– 747

  32. [32]

    A Podgorsek, J Jacquemin, AAH Pádua, and MF Costa Gomes. 2016. Mixing enthalpy for binary mixtures containing ionic liquids.Chemical reviews116, 10 (2016), 6075–6106

  33. [33]

    Maziar Raissi, Paris Perdikaris, and George E Karniadakis. 2019. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations.Journal of Computa- tional physics378 (2019), 686–707

  34. [34]

    Ella Miray Rajaonson, Mahyar Rajabi Kochi, Luis Martin Mejia Mendoza, Mo- hamad Moosavi, and Benjamin Sanchez-Lengeling. 2026. CheMixHub: Datasets and benchmarks for chemical mixture property prediction.Advances in Neural Information Processing Systems38 (2026)

  35. [35]

    Raghunathan Ramakrishnan, Pavlo O Dral, Matthias Rupp, and O Anatole Von Lilienfeld. 2014. Quantum chemistry structures and properties of 134 kilo molecules.Scientific data1, 1 (2014), 1–7

  36. [36]

    Jan G Rittig, Kobi C Felton, Alexei A Lapkin, and Alexander Mitsos. 2023. Gibbs– Duhem-informed neural networks for binary activity coefficient prediction. (2023)

  37. [37]

    Kristof Schütt, Pieter-Jan Kindermans, Huziel Enoc Sauceda Felix, Stefan Chmiela, Alexandre Tkatchenko, and Klaus-Robert Müller. 2017. Schnet: A continuous- filter convolutional neural network for modeling quantum interactions.Advances in neural information processing systems30 (2017)

  38. [38]

    Kristof Schütt, Oliver Unke, and Michael Gastegger. 2021. Equivariant mes- sage passing for the prediction of tensorial properties and molecular spectra. In International conference on machine learning. PMLR, 9377–9388

  39. [39]

    Vidushi Sharma, Maxwell Giammona, Dmitry Zubarev, Andy Tek, Khanh Nu- gyuen, Linda Sundberg, Daniele Congiu, and Young-Hye La. 2023. Formulation graphs for mapping structure-composition of battery electrolytes to device per- formance.Journal of Chemical Information and Modeling63, 22 (2023), 6998–7010

  40. [40]

    Zhiwei Shi, Miao Ma, Hanyang Ning, Bo Yang, and Liping Ding. 2026. A multi- level cross-molecular interaction neural network for mixture property prediction. Physica A: Statistical Mechanics and its Applications686 (2026), 131340

  41. [41]

    Murat Cihan Sorkun, Abhishek Khetan, and Süleyman Er. 2019. AqSolDB, a curated reference set of aqueous solubility and 2D descriptors for a diverse set of compounds.Scientific data6, 1 (2019), 143

  42. [42]

    Gary Tom, Cher Tian Ser, Ella M Rajaonson, Stanley Lo, Hyun Suk Park, Brian K Lee, and Benjamin Sanchez-Lengeling. 2025. From Molecules to Mixtures: Learn- ing Representations of Olfactory Mixture Similarity using Inductive Biases.arXiv preprint arXiv:2501.16271(2025)

  43. [43]

    Juan M Uceda, Melissa Morales, Marcela Cartes, and Andrés Mejía. 2025. Ex- perimental determination and theoretical modeling of isobaric vapor–liquid equilibria, liquid mass density, surface tension and dynamic viscosity for the methyl butyrate and tert-butanol binary mixture.Fluid Phase Equilibria587 (2025), 114199

  44. [44]

    Oliver T Unke and Markus Meuwly. 2019. PhysNet: A neural network for pre- dicting energies, forces, dipole moments, and partial charges.Journal of chemical theory and computation15, 6 (2019), 3678–3693

  45. [45]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in neural information processing systems30 (2017)

  46. [46]

    Apri Wahyudi, Natthapong Sueviriyapan, and Uthaiporn Suriyapraphadilok

  47. [47]

    Renxiao Wang, Xueliang Fang, Yipin Lu, and Shaomeng Wang. 2004. The PDBbind database: collection of binding affinities for protein- ligand complexes with known three-dimensional structures.Journal of medicinal chemistry47, 12 (2004), 2977–2980

  48. [48]

    U Westhaus, T Dröge, and R Sass. 1999. DETHERM®—a thermophysical property database.Fluid phase equilibria158 (1999), 429–435

  49. [49]

    Fang Wu, Dragomir Radev, and Stan Z Li. 2021. Molformer: Motif-based trans- former on 3d heterogeneous molecular graphs.arXiv preprint arXiv:2110.01191 (2021)

  50. [50]

    Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Ge- niesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. 2018. MoleculeNet: a benchmark for molecular machine learning.Chemical science9, 2 (2018), 513–530

  51. [51]

    Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. 2021. Do transformers really perform badly for graph representation?Advances in neural information processing systems34 KDD, August 01–05, 2027, San Jose Bai et al. (2021), 28877–28888

  52. [52]

    Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola. 2017. Deep sets.Advances in neural information processing systems30 (2017)

  53. [53]

    Jiaqi Zang, Wenjie Zhai, Yuchang Wang, Bo Zhang, Xiyue Ma, Kai Ma, and Jianbin Zhang. 2025. Excess properties, intermolecular interaction, and CO2 capture performance of diethylene glycol monomethyl ether+ ethylenediamine binary mixed solutions.Journal of Molecular Liquids417 (2025), 126561

  54. [54]

    Hengrui Zhang, Jie Chen, James M Rondinelli, and Wei Chen. 2023. Molsets: Molecular graph deep sets learning for mixture property modeling.arXiv preprint arXiv:2312.16473(2023)

  55. [55]

    Hengrui Zhang, Tianxing Lai, Jie Chen, Arumugam Manthiram, James M Rondinelli, and Wei Chen. 2024. Learning molecular mixture property using chemistry-aware graph neural network.PRX Energy3, 2 (2024), 023006

  56. [56]

    Unique systems

    Shang Zhu, Bharath Ramsundar, Emil Annevelink, Hongyi Lin, Adarsh Dave, Pin- Wen Guan, Kevin Gering, and Venkatasubramanian Viswanathan. 2024. Differ- entiable modeling and optimization of non-aqueous Li-based battery electrolyte solutions using geometric deep learning.Nature Communications15, 1 (2024), 8649. A Chem World A.1 Dataset Composition Chem Worl...

  57. [2022]

    InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing

    Translation between molecules and natural language. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 375–413

  58. [2023]

    ACS Central Science9, 2 (2023), 206–216

    Chemistry-informed machine learning for polymer electrolyte discovery. ACS Central Science9, 2 (2023), 206–216

  59. [2024]

    CALiSol-23: Experimental electrolyte conductivity data for various Li-salts and solvent combinations.Scientific Data11, 1 (2024), 750

  60. [2025]

    Generalizable physics-informed graph neural network for CO2 solubility prediction across amine classes.Separation and Purification Technology(2025), 136395