REVIEW 3 major objections 5 minor 62 references
Exploring molecular assembly as a biosignature using mass spectrometry and machine learning
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims that molecular assembly—the shortest construction pathway for a molecule—can be predicted from single-stage mass spectra by a gradient-boosted model, cutting error threefold and opening a route to agnostic life detection.
desk verdict This proof-of-concept for predicting molecular assembly from single-stage mass spectra has real value, but its instrument-consistency claims rely on a train/test split that likely leaks molecules across energy levels. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the molecular assembly index (MA), the length of the shortest construction pathway for a molecule, using a shared pool of reusable fragments so that substructure reuse shortens the score. MA gives the signal its interpretability: high values are statistically hard to reach by random chemistry, so they can flag selection. The predictive machinery is XGBoost, a gradient-boosted decision tree, trained on intensity-normalized mass spectra binned by mass-to-charge ratio, with true MA scores as targets. Around this sits RecursiveMA, the algorithm that reconstructs MA from multi-stage fragmentation trees; it supplies the concept of measuring assembly from spectra and marks the limit that single-stage MS1 data cannot meet algorithmically, which is the gap the ML model fills.
What would settle it
Measure a set of known and unknown molecules on a flight-like single-stage mass spectrometer under mission parameters, compute their MA independently with multi-stage fragmentation (RecursiveMA), and compare with the XGBoost predictions; if the relative MSE exceeds roughly 0.07 or the underprediction of high-MA molecules worsens materially, the transfer claim fails. A second check is to retrain the simulated-energy models with a per-molecule split, so no molecule's spectra appear in both training and test sets, and see whether the reported energy-mismatch error doubling survives.
Extended reading notes
Core claim
The paper's central claim is that molecular assembly, although defined from a molecule's structure and measurable from multi-stage fragmentation trees, can be inferred from ordinary single-stage mass spectra well enough for biosignature screening. On a curated collection of electron-ionization spectra, the XGBoost model reduces relative MSE from 0.12 (best baseline) to 0.04; errors are systematic rather than random, with low-MA molecules overpredicted and high-MA molecules underpredicted, a conservative direction for life detection. The model generalizes to an independent spectral database at 0.07, still better than the baseline. Simulated multi-energy spectra show that training and testing under matched collision energy yields errors near 0.03, mixing energies roughly doubles error, and concatenating all three energies into one integrated representation gives the best result at 0.029. From this the paper concludes that standardized mass-spectrometry databases could make MA prediction reliable on future missions.
Load-bearing premise
The load-bearing premise is that a mapping learned from terrestrial, structurally characterized molecules transfers to unknown molecules measured by different instruments on other planetary bodies, a transfer the paper does not test with any extraterrestrial or analog sample.
Editorial extensions
If this is right
- A future spacecraft carrying only a single-stage gas-chromatography mass spectrometer could estimate MA scores in situ and use them to prioritize samples for caching or deeper analysis, without resolving unknown structures.
- Standardization of ionization and collision-energy settings across training and target instruments becomes a mission design requirement; even a shift from 40 to 10 eV roughly doubles prediction error.
- Because the model underpredicts high-MA molecules, it will not cry 'life' on the basis of an overestimated score, making false positives less likely in screening.
- Combining spectra from several collision energies into an integrated representation improves accuracy, suggesting that multi-modal or multi-energy acquisition would help future instruments.
- If the experimentally suggested MA threshold near 15 separates biological from abiotic molecules, then ML-predicted MA could serve as a screening biosignature, with borderline samples flagged for MSn follow-up.
Reading between the lines
- The paper does not test the model on any extraterrestrial or analog sample; a blind trial on tholin-like material with independently measured MA would directly test the transfer assumption.
- The simulated multi-energy experiments appear to split spectra rather than molecules, so the same molecule at different collision energies may appear in both training and test sets; a per-molecule split would show whether the energy-mismatch errors are overstated.
- A multimodal model that adds NMR or infrared data, which the authors mention as future work, is a natural extension and could reduce the underprediction of molecules that barely fragment in the mass spectrometer.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes molecular assembly (MA) as an agnostic biosignature for life detection and asks whether MA can be predicted directly from mass spectrometry data without structural elucidation. The authors first compare MA with Bertz and Böttcher complexity scores, arguing that MA is preferable because it is sublinear in molecular size, interpretable in terms of construction pathways, and experimentally related to multi-stage mass spectrometry. They then train XGBoost models on NIST SRD EI-MS1 spectra with MA scores computed from molecular structures, obtaining a relative MSE of 0.04 versus 0.12 for the best baseline, and report a relative MSE of 0.07 on a held-out MassBank EI-B MS1 set. Using CFM-ID simulated MS2 spectra at three collision energies, they report that single-energy models reach relative MSE around 0.03, that a mixed-energy model degrades performance, and that cross-energy evaluation roughly doubles the error, leading to the conclusion that instrument standardization is critical. The manuscript closes with recommendations for standardized MS databases and multimodal future work.
Significance. If the results hold, the paper provides a useful proof-of-concept that MA, a proposed agnostic biosignature, can be estimated from single-stage mass spectrometry without structural elucidation, which would be relevant to upcoming Solar System missions carrying GC-MS instruments. The work has several concrete strengths: the ML target (MA) is computed from molecular structures rather than from MS features, so the prediction task is not circular; the main NIST-to-MassBank generalization is externally benchmarked and shows improvement over a simple power-law baseline; and the authors state that code is publicly available. The comparison of MA with Bertz and Böttcher scores, including size-scaling analysis on three large databases and symmetry-breaking examples, is a useful contribution independent of the ML results. However, the headline claim about instrument inconsistencies doubling model error rests on simulated-data experiments whose train/test split is not demonstrated to be molecule-disjoint, and the reported MSE values lack any measure of uncertainty. These issues are load-bearing for the quantitative claims but appear fixable with additional analysis and reporting.
major comments (3)
- [Evaluating ML prediction of MA from MS1 data] The train/test split for the simulated CFM-ID experiments is described only as a random stratified split of 'the full data' (Fig. 6c), with no grouping by molecule identity. Because each molecule contributes three spectra at 10, 20, and 40 eV, a spectrum-level split can place the same molecule in both training and test sets at different energies. This would let the model memorize molecule-specific fragmentation signatures and would directly inflate the reported single-energy and split-energy accuracies, and it could also bias the non-diagonal cross-energy errors (e.g., 0.031 to 0.060 and 0.029 to 0.067) on which the 'instrument inconsistencies double model error' claim rests. Please repeat the simulated experiments with splits grouped by molecule, or provide code with a pinned commit hash so the existing split can be verified; also report the number of unique molecules in each fold.
- [Data availability and Methods] The central numerical claims are single relative-MSE values (0.04 on NIST, 0.07 on MassBank) with no confidence intervals, standard errors, or repeated-seed variation. The NIST test set has roughly 9,000 molecules, so the reported three-fold improvement over the baseline would be more convincing with bootstrap intervals or results across multiple random seeds. Please report uncertainty around the MSE values and, ideally, a paired comparison with the baseline on the same test folds.
- [Data availability and Methods] The main text states that molecules were removed from the NIST SRD set because they were 'unsuitable for analysis' (Supplementary Fig. 1) but does not state the exclusion criteria in either the main text or the Methods. Since the size and composition of the training set directly affect the reported MSE, the criteria (e.g., missing structures, malformed spectra, failed MA computation, hydrogen-only molecules) need to be specified, together with the number of molecules removed for each reason.
minor comments (5)
- [Data availability] The Code Availability statement gives a GitHub URL but no release tag or commit hash; a pinned version is needed for reproducibility, especially given the split question in the major comments.
- [Evaluating ML prediction of MA from MS1 data] The main text says 'structural eludication' where 'elucidation' is intended; please correct this typo.
- [Conclusions] The sentence 'we chose MA several reasons' is missing 'for'; please edit.
- [Table 4] The manuscript repeatedly refers to Table 4, but the table body is not included in the provided text; please ensure the table appears in the final version with the diagonal and off-diagonal error values clearly labeled.
- [Post-hoc analyses] The claim that underprediction of high-MA molecules is 'conservative' is reasonable for avoiding false positives, but the same behavior could also reflect systematic model bias; consider reporting a calibration plot or a bias-variance decomposition to support the interpretation.
Circularity Check
No circularity; the MA prediction is a genuinely held-out supervised learning task, and cited prior work supplies background support rather than the derivation.
full rationale
The central derivation chain is not circular. The ML target (MA) is computed independently from molecular structures via assemblyCPP/AssemblyGo, not from the mass spectral features used as inputs; the XGBoost model is trained on NIST MS1 spectra and evaluated on a held-out NIST split and on MassBank molecules excluded from training, so the reported three-fold error reduction is not forced by construction. The simulated CFM-ID energy experiments compare models across energy conditions; even if the spectrum-level split risks molecule leakage across collision energies, that would be a data-splitting flaw rather than a definitional reduction, and it does not make the diagonal/non-diagonal comparison equivalent to its inputs. Self-citations (refs 27, 30, 33) support the threshold claim, assembly-theory background, and RecursiveMA, but they are external published experimental/algorithmic results and are not the means by which the ML predictions are produced; the prediction task would remain well-defined even if those background claims were set aside. The paper is therefore self-contained as a prediction study, with no step that reduces by definition to its own inputs.
Assumptions & free parameters
assumptions (4)
- domain assumption Assembly index as computed by assemblyCPP from molecular structure is a valid operationalization of molecular assembly (MA).
- domain assumption NIST SRD EI-MS1 spectra and CFM-ID simulated ESI-MS2 spectra are representative of the mass spectrometry data that future life-detection missions will produce.
- domain assumption Molecules with MA above roughly 15 are reliable biosignatures based on prior experimental separation of biological from abiotic samples.
- domain assumption The copy-number detection threshold in mass spectrometry automatically satisfies the copy-number term in the assembly equation.
Cite this review
Pith. "Pith review of Exploring molecular assembly as a biosignature using mass spectrometry and machine learning." pith.science (2026). https://pith.science/paper/JJBA35SQ
@misc{pith2026250719057,
author = {Pith},
title = {Pith review of: Exploring molecular assembly as a biosignature using mass spectrometry and machine learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/JJBA35SQ}},
note = {Machine review of arXiv:2507.19057}
}
read the original abstract
Molecular assembly offers a promising path to detect life beyond Earth, while minimizing assumptions based on terrestrial life. As mass spectrometers will be central to upcoming Solar System missions, predicting molecular assembly from their data without needing to elucidate unknown structures will be essential for unbiased life detection. An ideal agnostic biosignature must be interpretable and experimentally measurable. Here, we show that molecular assembly, a recently developed approach to measure objects that have been produced by evolution, satisfies both criteria. First, it is interpretable for life detection, as it reflects the assembly of molecules with their bonds as building blocks, in contrast to approaches that discount construction history. Second, it can be determined without structural elucidation, as it can be physically measured by mass spectrometry, a property that distinguishes it from other approaches that use structure-based information measures for molecular complexity. Whilst molecular assembly is directly measurable using mass spectrometry data, there are limits imposed by mission constraints. To address this, we developed a machine learning model that predicts molecular assembly with high accuracy, reducing error by three-fold compared to baseline models. Simulated data shows that even small instrumental inconsistencies can double model error, emphasizing the need for standardization. These results suggest that standardized mass spectrometry databases could enable accurate molecular assembly prediction, without structural elucidation, providing a proof-of-concept for future astrobiology missions.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[34]
Gebhard, T. D. et al. Inferring molecular complexity from mass spectrometry data using machine learning. in Advances in Neural Information Processing Systems (2022)
work page 2022
-
[1]
Cleland, C. E. The Quest for a Universal Theory of Life: Searching for Life As We Don’t Know It. (Cambridge University Press, 2019). doi:10.1017/9781139046893
-
[2]
Gilbert, W. Origin of life: The RNA world. Nature 319, 618–618 (1986)
work page 1986
-
[3]
Moody, E. R. R. et al. The nature of the last universal common ancestor and its impact on the early Earth system. Nat. Ecol. Evol. 8, 1654–1666 (2024)
work page 2024
-
[4]
Petrov, A. S. et al. Evolution of the ribosome at atomic resolution. Proc. Natl. Acad. Sci. 111, 10251–10256 (2014)
work page 2014
-
[5]
Krasnopolsky, V. A., Maillard, J. P. & Owen, T. C. Detection of methane in the martian atmosphere: evidence for life? Icarus 172, 537–547 (2004)
work page 2004
-
[6]
Greaves, J. S. et al. Phosphine gas in the cloud decks of Venus. Nat. Astron. 5, 655–664 (2020)
work page 2020
-
[7]
Buan, N. R. Methanogens: pushing the boundaries of biology. Emerg. Top. Life Sci. 2, 629–646 (2018)
work page 2018
Show all 62 references
-
[8]
& Verstraete, W
Roels, J. & Verstraete, W. Biological formation of volatile phosphorus compounds. Bioresour. Technol. 79, 243–250 (2001). 30
2001
-
[9]
& Lunine, J
Truong, N. & Lunine, J. I. Volcanically extruded phosphides as an abiotic source of Venusian phosphine. Proc. Natl. Acad. Sci. 118, e2021689118 (2021)
2021
-
[10]
& Theloke, J
Etiope, G., Fridriksson, T., Italiano, F., Winiwarter, W. & Theloke, J. Natural emissions of methane from geothermal and volcanic sources in Europe. J. Volcanol. Geotherm. Res. 165, 76–86 (2007)
2007
-
[11]
Ferris, J. P. & Khwaja, H. Laboratory simulations of PH3 photolysis in the atmospheres of Jupiter and Saturn. Icarus 62, 415–424 (1985)
1985
-
[12]
Wilson, E. H. & Atreya, S. K. Sensitivity studies of methane photolysis and its impact on hydrocarbon chemistry in the atmosphere of Titan. J. Geophys. Res. Planets 105, 20263–20273 (2000)
2000
-
[13]
Knak Jensen, S. J. et al. A sink for methane on Mars? The answer is blowing in the wind. Icarus 236, 24–27 (2014)
2014
-
[14]
Bains, W. et al. Source of phosphine on Venus—An unsolved problem. Front. Astron. Space Sci. 11, 1372057 (2024)
2024
-
[15]
Bains, W. et al. Phosphine on Venus Cannot Be Explained by Conventional Processes. Astrobiology 21, 1277–1304 (2021)
2021
-
[16]
Webster, C. R. et al. Background levels of methane in Mars’ atmosphere show strong seasonal variations. Science 360, 1093–1096 (2018)
2018
-
[17]
Madhusudhan, N. et al. Carbon-bearing Molecules in a Possible Hycean Atmosphere. Astrophys. J. Lett. 956, L13 (2023)
2023
-
[18]
Madhusudhan, N., Piette, A. A. A. & Constantinou, S. Habitability and Biosignatures of Hycean Worlds. Astrophys. J. 918, 1 (2021)
2021
-
[19]
Kettle, A. J. & Andreae, M. O. Flux of dimethylsulfide from the oceans: A comparison of updated data sets and flux models. J. Geophys. Res. Atmospheres 105, 26793–26808 (2000). 31
2000
-
[20]
Hänni, N. et al. Evidence for Abiotic Dimethyl Sulfide in Cometary Matter. Astrophys. J. 976, 74 (2024)
2024
-
[21]
Reed, N. W. et al. Abiotic Production of Dimethyl Sulfide, Carbonyl Sulfide, and Other Organosulfur Gases via Photochemistry: Implications for Biosignatures and Metabolic Potential. Astrophys. J. Lett. 973, L38 (2024)
2024
-
[22]
Smith, H. B. & Mathis, C. Life detection in a universe of false positives: Can the Fatal Flaws of Exoplanet Biosignatures be Overcome Absent a Theory of Life? BioEssays 45, 2300050 (2023)
2023
-
[23]
S., McMahon, S
Cockell, C. S., McMahon, S. & Biddle, J. F. When is Life a Viable Hypothesis? The Case of Venusian Phosphine. Astrobiology 21, 261–264 (2021)
2021
-
[24]
L., Bartlett, S., Chen, S
Wong, M. L., Bartlett, S., Chen, S. & Tierney, L. Searching for Life, Mindful of Lyfe’s Possibilities. Life 12, 783 (2022)
2022
-
[25]
S., Anslyn, E
Johnson, S. S., Anslyn, E. V., Graham, H. V., Mahaffy, P. R. & Ellington, A. D. Fingerprinting Non-Terran Biosignatures. Astrobiology 18, 915–922 (2018)
2018
-
[26]
& Cleaves, H
Guttenberg, N., Chen, H., Mochizuki, T. & Cleaves, H. Classification of the Biogenicity of Complex Organic Mixtures for the Detection of Extraterrestrial Life. Life 11, 234 (2021)
2021
-
[27]
Marshall, S. M. et al. Identifying molecules as biosignatures with assembly theory and mass spectrometry. Nat. Commun. 12, 3033 (2021)
2021
-
[28]
Bertz, S. H. The first general index of molecular complexity. J. Am. Chem. Soc. 103, 3599–3601 (1981)
1981
-
[29]
An Additive Definition of Molecular Complexity
Böttcher, T. An Additive Definition of Molecular Complexity. J. Chem. Inf. Model. 56, 462–470 (2016)
2016
-
[30]
Sharma, A. et al. Assembly theory explains and quantifies selection and evolution. Nature 622, 321–328 (2023). 32
2023
- [31]
-
[32]
Chou, L. et al. Planetary Mass Spectrometry for Agnostic Life Detection in the Solar System. Front. Astron. Space Sci. 8, 755100 (2021)
2021
-
[33]
Jirasek, M. et al. Investigating and Quantifying Molecular Complexity Using Assembly Theory and Spectroscopy. ACS Cent. Sci. 10, 1054–1064 (2024)
2024
-
[35]
NIST Chemistry WebBook, NIST Standard Reference Database 69
Linstrom, P. NIST Chemistry WebBook, NIST Standard Reference Database 69. National Institute of Standards and Technology https://doi.org/10.18434/T4D303 (1997)
1997 doi
-
[36]
& Guestrin, C
Chen, T. & Guestrin, C. XGBoost: A Scalable Tree Boosting System. in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 785–794 (ACM, San Francisco California USA, 2016). doi:10.1145/2939672.2939785
2016
-
[37]
Draine, B. T. & Li, A. Infrared Emission from Interstellar Dust. IV. The Silicate‐ Graphite‐PAH Model in the Post‐ Spitzer Era. Astrophys. J. 657, 810–837 (2007)
2007
-
[38]
Allamandola, L. J. PAHs, They’re Everywhere! in The Cosmic Dust Connection (ed. Greenberg, J. M.) 81–102 (Springer Netherlands, Dordrecht, 1996). doi:10.1007/978-94- 011-5652-3_4
1996 doi
-
[39]
J., Sandford, S
Allamandola, L. J., Sandford, S. A. & Wopenka, B. Interstellar Polycyclic Aromatic Hydrocarbons and Carbon in Interplanetary Dust Particles and Meteorites. Science 237, 56–59 (1987)
1987
-
[40]
PAHs in Astronomy - A Review
Salama, F. PAHs in Astronomy - A Review. Proc. Int. Astron. Union 4, 357–366 (2008). 33
2008
-
[41]
López-Puertas, M. et al. LARGE ABUNDANCES OF POLYCYCLIC AROMATIC HYDROCARBONS IN TITAN’S UPPER ATMOSPHERE. Astrophys. J. 770, 132 (2013)
2013
-
[42]
The PAH Hypothesis after 25 Years
Peeters, E. The PAH Hypothesis after 25 Years. Proc. Int. Astron. Union 7, 149–161 (2011)
2011
-
[43]
Beattie, M. & A. H. Jones, O. Rate of Advancement of Detection Limits in Mass Spectrometry: Is there a Moore’s Law of Mass Spec? Mass Spectrom. 12, A0118–A0118 (2023)
2023
-
[44]
Chandrasekhar, V. et al. COCONUT 2.0: a comprehensive overhaul and curation of the collection of open natural products database. Nucleic Acids Res. 53, D634–D643 (2025)
2025
-
[45]
Chan, Q. H. S., Watson, J. S., Sephton, M. A., O’Brien, Á. C. & Hallis, L. J. The amino acid and polycyclic aromatic hydrocarbon compositions of the promptly recovered CM2 Winchcombe carbonaceous chondrite. Meteorit. Planet. Sci. 59, 1101–1130 (2024)
2024
-
[46]
Freissinet, C. et al. Long-chain alkanes preserved in a Martian mudstone. Proc. Natl. Acad. Sci. 122, e2420580122 (2025)
2025
-
[47]
& Lerch, Ph
Jennings, E., Montgomery, W. & Lerch, Ph. Stability of Coronene at High Temperature and Pressure. J. Phys. Chem. B 114, 15753–15758 (2010)
2010
-
[48]
Berné, O., Cox, N. L. J., Mulas, G. & Joblin, C. Detection of buckminsterfullerene emission in the diffuse interstellar medium. Astron. Astrophys. 605, L1 (2017)
2017
-
[49]
Becker, L. et al. Fullerenes in Meteorites and the Nature of Planetary Atmospheres. in Natural Fullerenes and Related Structures of Elemental Carbon vol. 6 95–121 (Springer Netherlands, Dordrecht, 2006)
2006
-
[50]
Kim, S. et al. PubChem 2025 update. Nucleic Acids Res. 53, D1516–D1525 (2025). 34
2025
-
[51]
Schmidt, G. A. & Frank, A. The Silurian hypothesis: would it be possible to detect an industrial civilization in the geological record? Int. J. Astrobiol. 18, 142–150 (2019)
2019
-
[52]
W., Abad, G
Lin, H. W., Abad, G. G. & Loeb, A. DETECTING INDUSTRIAL POLLUTION IN THE ATMOSPHERES OF EARTH-LIKE EXOPLANETS. Astrophys. J. 792, L7 (2014)
2014
-
[53]
MassBank/MassBank-data: Release version 2025.05.1
MassBank consortium and its contributors. MassBank/MassBank-data: Release version 2025.05.1. Zenodo https://doi.org/10.5281/ZENODO.3378723 (2025)
2025 doi
-
[54]
Wang, F. et al. CFM-ID 4.0: More Accurate ESI-MS/MS Spectral Prediction and Compound Identification. Anal. Chem. 93, 11692–11700 (2021)
2021
-
[55]
A., Mayer, M
Scharf, C. A., Mayer, M. H. & Boston, P. J. Using artificial intelligence to transform astrobiology. Nat. Astron. 8, 8–9 (2023)
2023
-
[56]
Warren-Rhodes, K. et al. Orbit-to-ground framework to decode and predict biosignature patterns in terrestrial analogues. Nat. Astron. 7, 406–422 (2023)
2023
-
[57]
Cleaves, H. J. et al. A robust, agnostic molecular biosignature based on machine learning. Proc. Natl. Acad. Sci. 120, e2307149120 (2023)
2023
-
[58]
Ward, J. L. et al. An inter-laboratory comparison demonstrates that [1H]-NMR metabolite fingerprinting is a robust technique for collaborative plant metabolomic data collection. Metabolomics 6, 263–273 (2010)
2010
-
[59]
Tait, K. T. et al. Preliminary Planning for Mars Sample Return (MSR) Curation Activities in a Sample Receiving Facility (SRF). Astrobiology 22, S-57-S-80 (2022)
2022
-
[60]
McKay, D. S. et al. Search for Past Life on Mars: Possible Relic Biogenic Activity in Martian Meteorite ALH84001. Science 273, 924–930 (1996)
1996
-
[61]
Letertre, M. P. M., Dervilly, G. & Giraudeau, P. Combined Nuclear Magnetic Resonance Spectroscopy and Mass Spectrometry Approaches for Metabolomics. Anal. Chem. 93, 500–518 (2021). 35
2021
-
[62]
rdkit/rdkit: 2025_03_2 (Q1 2025) Release
Greg Landrum et al. rdkit/rdkit: 2025_03_2 (Q1 2025) Release. Zenodo https://doi.org/10.5281/ZENODO.591637 (2025)
2025 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.