Pith. sign in

REVIEW 3 major objections 5 minor 65 references

A polymer language model generates millions of property-targeted polymer candidates, one of which was synthesized and matched predictions.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

A T5-based polymer language model pre-trained on 100 million hypothetical polymers predicts thermal, electronic, and solubility properties and generates polymers conditioned on a target glass-transition temperature, with one experimental validation.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection Solid incremental polymer T5 model; the generative-design claims rest too heavily on the model scoring its own outputs, but the property prediction and pre-training ablation give it real value. the 3 major comments →

arxiv 2510.18860 v1 pith:ULBT44M6 submitted 2025-10-21 cond-mat.mtrl-sci cond-mat.soft

An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design

classification cond-mat.mtrl-sci cond-mat.soft
keywords polymer informaticsgenerative designT5 transformerSELFIESglass transition temperaturedielectric polymersproperty predictionfoundation model
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces polyT5, a domain-specific encoder–decoder language model pre-trained on over 100 million SELFIES-encoded polymer structures. The authors claim that polyT5 has learned the intrinsic relationship between polymer chemical structure and glass transition temperature (Tg), enabling both accurate property prediction and conditional generation of chemically valid, novel polymers tailored to a specified Tg. Applied to dielectric polymer design, the model generated 6.17 million hypothetical candidates, screened them to 21,457 that meet dielectric, thermal, and solubility criteria, and one top candidate was experimentally synthesized—its measured Tg (472 K) and bandgap (4.53 eV) aligning closely with predictions (483 K and 4.45 eV). If correct, this establishes an end-to-end, informatics-driven workflow that moves from a property target to a validated polymer without exhaustive enumeration.

Core claim

The central claim is that polyT5 has effectively learned the intrinsic relationship between polymer chemical structure and Tg, enabling the generation of candidates tailored to specified target temperatures. Evidence includes the predicted Tg distribution of 6 million generated polymers centering near the 500 K target, high chemical diversity relative to the training set (Tanimoto similarity analysis), and a five-fold improvement in SELFIES reproducibility from pre-training. The screening pipeline, built on polyT5's own fine-tuned predictors, identified 21,457 candidates satisfying dielectric constant > 3, bandgap > 4 eV, Tg > 400 K, melt processability, and solubility in water or ethanol. T

What carries the argument

The key machinery is the PSELFIES representation, a polymer-specific adaptation of SELFIES where the two polymer termini are replaced by Astatine placeholder atoms, making polymers representable as valid molecular strings. The model uses a T5 encoder–decoder with bidirectional attention and span-masking pre-training on over 100 million hypothetical polymers generated via known polymerization reactions (polycondensation, click chemistry, ROMP). Fine-tuning inverts the direction: Tg values are fed as input and PSELFIES strings are generated as output, enabling property-conditional generation. Candidate validity is enforced through sequential filters—SMILES validity, training-set deduplication,

Load-bearing premise

The generative design results are certified by polyT5's own fine-tuned property predictors, so if those predictions are systematically biased for novel chemistries—unlikely to be revealed by a single validation polymer chosen for ease of synthesis—the quality of the generated candidates and the screening conclusions would be compromised.

What would settle it

Synthesize a random sample of 10–20 polymers from the 21,457 screened candidates (not just the easiest-to-synthesize one) and compare their measured Tg and bandgap with polyT5 predictions; if the median absolute error significantly exceeds the reported RMSE of 40.8 K for Tg, the screening pipeline's reliability claim fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • The same conditional-generation framework can be extended to other polymer properties (e.g., Td, Tm, Eg, dielectric constant) provided sufficiently large and diverse training datasets exist.
  • Pre-training on 100 million structures dramatically improves downstream performance: Tg RMSE drops from 89.35 K to 40.82 K and R2 rises from 0.31 to 0.86 compared to training from scratch.
  • The workflow compresses the design cycle from target property to a wet-lab-validated polymer in one experimental iteration, a proof of concept for generative polymer discovery.
  • The agentic AI wrapper, coupling polyT5 with a general-purpose LLM, makes property prediction and generative design accessible to non-experts through natural-language queries.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If polyT5's property predictors carry a systematic bias on novel chemistries, the 21,457 'promising' screened candidates could be over-optimistic; the single validation polymer, chosen for ease of synthesis, cannot detect such a bias.
  • The generated candidate set shows a heavy skew toward allyl (84.3%), ether (66.1%), and amide (56.0%) groups, suggesting the model is biased toward chemistries overrepresented in the training/pre-training data—broader reaction-template diversity could expand scaffold coverage.
  • A stronger test of the generator's learned Tg–structure relationship would be to generate polymers at extreme Tg targets (e.g., 700 K or 150 K) and check whether predicted distributions still center on the target.
  • The screening success rate of 21,457 out of 6.17 million candidates (0.35%) is implicitly presented as an enrichment over random sampling, but the paper does not quantify the baseline hit rate, which would be needed to claim true acceleration.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript presents polyT5, an encoder-decoder T5 model pre-trained on over 100 million polymer structures in a SELFIES-based representation, and fine-tuned for property prediction (Tg, Td, Tm, Eg, dielectric constant, solubility) and for conditional generation of polymer structures targeting a desired Tg. The authors report RMSE values (e.g., Tg RMSE 40.8 K, Eg RMSE 0.60 eV) across five random splits, an ablation study showing pre-training helps, and a dielectric polymer design pipeline that generated over 6 million candidates, screened them with polyT5 property predictors, and selected 21,457 promising candidates. One candidate was synthesized and its measured Tg, Tm, Td, and DFT Eg agree with predictions within error bars. An agentic AI framework integrating polyT5 with a general-purpose LLM is also demonstrated. The central claim is that polyT5 captures polymer structure–property relationships well enough to enable reliable generative design and accelerated discovery.

Significance. If the claims are supported, polyT5 would be a useful open-domain generative and predictive tool for polymer informatics, combining a T5 encoder-decoder architecture with polymer-specific SELFIES representations and large-scale pre-training. The property prediction results are presented with learning curves and five-split statistics, and the ablation study indicates that pre-training contributes substantially to performance, especially for solubility. The experimental validation of one synthesized polymer is a positive step. However, the generative design claim and the screening results rest on self-prediction with limited external verification, which currently prevents the paper from substantiating the 'accelerated generative discovery' framing. The strength of the property-prediction component is notable, but the design pipeline requires additional validation or more cautious interpretation.

major comments (3)
  1. [Downstream Task 2: Generative design; Figure S10] The generative targeting evaluation uses the fine-tuned polyT5 Tg predictor to score generated candidates: the TP metric in Figure S10 counts candidates with polyT5-predicted Tg within 500±50 K, and Figure 4D plots the same predictor's distribution. This is circular as a validation of the generator: both generator and predictor derive from the same fine-tuned model, so a generator that exploits the predictor's systematic biases would still score well. The text's claim that 'polyT5 has effectively learned the intrinsic relationship between polymer chemical structure and Tg' (Downstream Task 2) is not supported by this self-consistent loop. Please add an external check—e.g., evaluate a random sample of generated candidates with an independent model (such as polyBERT), group-contribution methods, or DFT—or temper the claim to say that the generator is consistent with the polyT5 Tg model.
  2. [Dielectric polymer design; Figure 5A; Table 1] The screening funnel in Figure 5A uses only polyT5-predicted Tg, Td, Tm, Eg, epsilon, and solubility to select the 21,457 finalists. With Eg RMSE ≈0.60 eV and a >4 eV cutoff, and Tm/Td RMSEs of about 67/79 K against a 100 K margin, the true number of candidates satisfying the property constraints may be substantially lower. The solubility model's insoluble accuracy of 0.917 also implies an ~8% false-positive rate on the insoluble class; over millions of candidates this could introduce tens of thousands of false positives. The single experimental validation in Table 1 concerns one polymer chosen for ease of synthesis and cannot validate the other 21,456 candidates. Please provide a conservative estimate that propagates prediction uncertainty through the cascade, or independently verify a random subset of screened candidates (e.g., DFT for Eg/epsilon, external ML predictors, or additional
  3. [PolyT5: Foundation model for polymers; Materials and Methods] The manuscript does not state whether the 12,473 experimentally known polymers used in pre-training were excluded from the fine-tuning datasets for Tg, Td, Tm, Eg, epsilon, and solubility. If any of these known polymers appear in the 100M pre-training corpus, the reported property-prediction metrics could be inflated by pre-training leakage, because the model has already seen the exact polymer strings. Please clarify the deduplication procedure between the pre-training corpus and each downstream fine-tuning dataset, and report the overlap if any. This is load-bearing for the property-prediction claim.
minor comments (5)
  1. [Discussion, first paragraph] The sentence reports RMSEs of 40.82, 67.07, and 78.59 K for Tg, Tm, and Tm, respectively; the last property should be Td.
  2. [Figure S2 caption] The caption says training loss is shown 'across two epochs', but the Materials and Methods state pre-training was performed for up to 5 epochs. Please align these numbers.
  3. [Figure 4A caption] The nested relationship is written as 'PV⊂DD⊂TSD⊂SV', which is backwards relative to the text. It should be SV⊃TSD⊃DD⊃PV.
  4. [Impact of pre-training] The SELFIES reproducibility (SR) metric is first used in the ablation study but is not defined until later in the same section. Define it clearly at first use.
  5. [Table S5] The non-pre-trained solubility model achieves soluble accuracy 0.98 but insoluble accuracy 0.04 and overall accuracy 0.66, which is essentially the trivial all-soluble classifier. The text should explicitly note this to clarify that the pre-trained model is meaningful and not just better on a trivial baseline.

Circularity Check

0 steps flagged

No significant circularity: generative-design evaluation is self-scored but explicitly predictive and anchored by one experimental synthesis.

full rationale

The paper's core derivation chain — pre-training a polymer T5 model on 100M PSELFIES strings, fine-tuning for property prediction and Tg-conditioned generation, screening generated candidates with property predictors, and validating one candidate experimentally — does not reduce to its own inputs by construction. Property prediction is benchmarked on held-out test splits (RMSEs of 40.8 K for Tg, 0.596 eV for Eg, etc., Figure 3), so the screening models are not fitted to the final generated candidates. The generative design claim is supported by Figure 4D and S10, where generated polymers are scored by polyT5's own predicted Tg; this is a self-consistency check rather than an external validation, and it is a genuine limitation for certifying 21,457 candidates. However, the paper explicitly says these are 'predicted' values, and it provides one experimental synthesis with measured Tg = 472 K vs predicted 483 K, Eg = 4.53 eV vs 4.45 eV. No equation-level identity or fitted-parameter-as-prediction reduction is present. Self-citations (refs 12, 13, 32, 45) supply datasets, the pseudo-SELFIES conversion tool, and prior modeling strategies; they are implementation and data sources, not a load-bearing uniqueness theorem or an ansatz smuggled in to force the conclusion. The single-synthesis anchor is thin for a 21,457-candidate screening claim, but that is a correctness/validation risk, not circularity.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The paper introduces no new physical entities or forces. The free parameters are screening thresholds and hyperparameters chosen by hand, plus the data-generation assumptions. The key hidden assumption is the absence of leakage between pre-training and fine-tuning datasets, which is not explicitly verified. The largest circularity burden comes from using the same model family for both generation and evaluation.

free parameters (3)
  • Screening thresholds for dielectric design = Tg>400 K, Eg>4 eV, epsilon>3, Tm-Tg>100 K, Td-Tg>100 K, soluble in water or ethanol
    These hand-chosen cutoffs determine the 21,457 finalists and are not derived from any principle or fitted to data; they are domain assumptions that directly shape the claimed number of promising candidates.
  • Generation inference hyperparameters for final model = Medium model: 6 fine-tuning epochs, T5 temperature 1.1, top_p 0.75
    The authors chose the combination that maximizes the number of candidates passing the PV filter (Figure 4A). This is a parameter selection made to optimize an internal filter, and the generated set depends on it.
  • Span masking configuration in pre-training = Up to 15% masked tokens, up to 8 spans of up to 3 tokens each
    Chosen from T5 conventions rather than from polymer-specific evidence; the central claim of capturing multi-token dependencies rests on this configuration.
axioms (4)
  • domain assumption Rule-based polymerization of commercially available molecules yields chemically valid, synthesizable polymer structures.
    The 100M pre-training corpus is generated by polycondensation, click chemistry, and ROMP reactions. If these reactions produce structures that are not actually synthesizable, the foundation-model training data is of questionable quality. Invoked in Materials and Methods, 'Generation of reliable hypothetical polymers for pre-training'.
  • domain assumption PSELFIES representation with two Astatine atoms as termini is a faithful encoding of a polymer that preserves chemical identity after conversion from PSMILES.
    The whole model operates on PSELFIES strings. The conversion through a cyclic intermediate and At-atom attachment is assumed to be bijective and chemically meaningful (Figure 2A).
  • ad hoc to paper The fine-tuned polyT5 property-prediction models are accurate enough to screen generated candidates without external verification.
    The dielectric screening uses polyT5-predicted Tg, Td, Tm, Eg, epsilon, and solubility to select the finalists. This assumes the model generalizes to the novel chemistries it generates. This is the load-bearing premise of the discovery workflow.
  • ad hoc to paper There is no leakage between the 100M hypothetical pre-training corpus and the property fine-tuning datasets.
    The paper does not state that the fine-tuning polymers were removed from the pre-training corpus. Since the hypothetical polymers are generated by reactions of commercial molecules, some real polymers from property datasets could appear in pre-training, inflating prediction accuracy. The authors only deduplicate generated candidates against the training set (TSD), not the reverse.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design." pith.science (2026). https://pith.science/paper/ULBT44M6

@misc{pith2026251018860,
  author       = {Pith},
  title        = {Pith review of: An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ULBT44M6}},
  note         = {Machine review of arXiv:2510.18860}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Traditional machine learning has advanced polymer discovery, yet direct generation of chemically valid and synthesizable polymers without exhaustive enumeration remains a challenge. Here we present polyT5, an encoder-decoder chemical language model based on the T5 architecture, trained to understand and generate polymer structures. polyT5 enables both property prediction and the targeted generation of polymers conditioned on desired property values. We demonstrate its utility for dielectric polymer design, seeking candidates with dielectric constant >3, bandgap >4 eV, and glass transition temperature >400 K, alongside melt-processability and solubility requirements. From over 20,000 generated promising candidates, one was experimentally synthesized and validated, showing strong agreement with predictions. To further enhance usability, we integrated polyT5 within an agentic AI framework that couples it with a general-purpose LLM, allowing natural language interaction for property prediction and generative design. Together, these advances establish a versatile and accessible framework for accelerated polymer discovery.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

65 extracted references · 18 canonical work pages

  1. [1]

    Zhang,et al., Exploring the role of large language models in the scientific method: from hypothesis to discovery.npj Artif

    Y. Zhang,et al., Exploring the role of large language models in the scientific method: from hypothesis to discovery.npj Artif. Intell.1, 14 (2025)

  2. [2]

    Peivaste,et al., Artificial intelligence in materials science and engineering: Current landscape, key challenges, and future trajectories.Compos

    I. Peivaste,et al., Artificial intelligence in materials science and engineering: Current landscape, key challenges, and future trajectories.Compos. Struct.372, 119419 (2025)

  3. [3]

    M. C. Ramos, C. J. Collison, A. D. White, A review of large language models and autonomous agents in chemistry.Chem. Sci.16, 2514–2572 (2025)

  4. [4]

    E. O. Pyzer-Knapp,et al., Foundation models for materials discovery –current state and future directions.npj Comput. Mater.11, 61 (2025), doi:10.1038/s41524-025-01538-0

  5. [5]

    M.-H. Van, P. Verma, C. Zhao, X. Wu, A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools (2025),https://arxiv.org/abs/2506.20743

  6. [6]

    D. M. Anstine, O. Isayev, Generative Models as an Emerging Paradigm in the Chemical Sciences.J. Am. Chem. Soc.145(16), 8736–8750 (2023), doi:10.1021/jacs.2c13467

  7. [7]

    Zaki, Jayadeva, Mausam, N

    M. Zaki, Jayadeva, Mausam, N. M. A. Krishnan, MaScQA: investigating materials science knowledge of large language models.Digital Discovery3, 313–327 (2024)

  8. [8]

    Zhang, X

    J. Zhang, X. Chen, X. Ye, Y. Yang, B. Ai, Large Language Model in Materials Science: Roles, Challenges, and Strategic Outlook.Advanced Intelligent Discoveryp. 202500085 (2025)

  9. [9]

    Liu,et al., MolXPT: Wrapping Molecules with Text for Generative Pre-training (2023), https://arxiv.org/abs/2305.10688

    Z. Liu,et al., MolXPT: Wrapping Molecules with Text for Generative Pre-training (2023), https://arxiv.org/abs/2305.10688

  10. [10]

    Chen,et al., An overview of domain-specific foundation model: key technologies, applica- tions and challenges (2025),https://arxiv.org/abs/2409.04267

    H. Chen,et al., An overview of domain-specific foundation model: key technologies, applica- tions and challenges (2025),https://arxiv.org/abs/2409.04267

  11. [11]

    D. M. Anisuzzaman, J. G. Malins, P. A. Friedman, Z. I. Attia, Fine-Tuning Large Language Models for Specialized Use Cases.Mayo Clinic Proceedings: Digital Health3(1), 100184 (2025), doi:https://doi.org/10.1016/j.mcpdig.2024.11.005. 21

  12. [12]

    Gupta, A

    S. Gupta, A. Mahmood, S. Shukla, R. Ramprasad, Benchmarking Large Language Models for Polymer Property Predictions (2025),https://arxiv.org/abs/2506.02129

  13. [13]

    Agarwal, A

    S. Agarwal, A. Mahmood, R. Ramprasad, Polymer Solubility Prediction Using Large Language Models.ACS Mater. Lett.7, 2017–2023 (2025)

  14. [14]

    G. Chen,et al., Machine-Learning-Assisted De Novo Design of Organic Molecules and Poly- mers: Opportunities and Challenges.Polymers12(1) (2020), doi:10.3390/polym12010163, https://www.mdpi.com/2073-4360/12/1/163

  15. [15]

    Reymond, The Chemical Space Project.Acc

    J.-L. Reymond, The Chemical Space Project.Acc. Chem. Res.48, 722–730 (2015), doi:10. 1021/ar500432k

  16. [16]

    Sahu,et al., Designing promising molecules for organic solar cells via machine learning as- sisted virtual screening.J

    H. Sahu,et al., Designing promising molecules for organic solar cells via machine learning as- sisted virtual screening.J. Mater. Chem. A7, 17480–17488 (2019), doi:10.1039/C9TA04097H

  17. [17]

    H. Park, Z. Li, A. Walsh, Has generative artificial intelligence solved inverse materials design? Matter7, 2355–2367 (2024), doi:10.1016/j.matt.2024.05.017

  18. [18]

    Dollar, N

    O. Dollar, N. Joshi, D. A. C. Beck, J. Pfaendtner, Attention-based generative models for de novo molecular design.Chem. Sci.12, 8362–8372 (2021), doi:10.1039/D1SC01050F

  19. [19]

    He,et al., Molecular optimization by capturing chemist’s intuition using deep neural net- works.J

    J. He,et al., Molecular optimization by capturing chemist’s intuition using deep neural net- works.J. Cheminform.13, 26 (2021), doi:10.1186/s13321-021-00497-0

  20. [20]

    Irwin, S

    R. Irwin, S. Dimitriadis, J. He, E. J. Bjerrum, Chemformer: a pre-trained transformer for computational chemistry.Mach. learn.: sci. technol.3, 015022 (2022)

  21. [21]

    Weininger, SMILES, a chemical language and information system

    D. Weininger, SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules.J. Chem. Inf. Comput. Sci.28, 31–36 (1988)

  22. [22]

    Weininger, A

    D. Weininger, A. Weininger, J. L. Weininger, SMILES. 2. Algorithm for generation of unique SMILES notation.J. Chem. Inf. Comput. Sci.29, 97–101 (1989), doi:10.1021/ci00062a008

  23. [23]

    Sattari, Y

    K. Sattari, Y. Xie, J. Lin, Data-driven algorithms for inverse design of polymers.Soft Matter 17, 7607–7622 (2021), doi:10.1039/D1SM00725D. 22

  24. [24]

    Batra,et al., Polymers for Extreme Conditions Designed Using Syntax-Directed Variational Autoencoders.Chem

    R. Batra,et al., Polymers for Extreme Conditions Designed Using Syntax-Directed Variational Autoencoders.Chem. Mat.32, 10489–10500 (2020), doi:10.1021/acs.chemmater.0c03332

  25. [25]

    M. A. Skinnider, Invalid SMILES are beneficial rather than detrimental to chemical language models.Nat. Mach. Intell.6, 437–448 (2024), doi:10.1038/s42256-024-00821-x

  26. [26]

    Krenn, F

    M. Krenn, F. H ¨ase, A. K. Nigam, P. Friederich, A. Aspuru-Guzik, Self-referencing embedded strings (SELFIES): A 100representation.Mach. Learn.: Sci. Technol.1, 045024 (2020), doi: 10.1088/2632-2153/aba947

  27. [27]

    Krenn,et al., SELFIES and the future of molecular string representations.Patterns3, 100588 (2022), doi:10.1016/j.patter.2022.100588

    M. Krenn,et al., SELFIES and the future of molecular string representations.Patterns3, 100588 (2022), doi:10.1016/j.patter.2022.100588

  28. [28]

    T. Xu, N. Velzeboer, Y. Maruyama, Chemist-Computer Interaction: Representation Learning for Chemical Design via Refinement of SELFIES V AE, inHCI International 2023 – Late Breaking Posters, C. Stephanidis, M. Antona, S. Ntoa, G. Salvendy, Eds. (Springer Nature Switzerland, Cham) (2024), pp. 353–361

  29. [29]

    S. Piao, J. Choi, S. Seo, S. Park, SELF-EdiT: Structure-constrained molecular optimi- sation using SELFIES editing transformer.Appl. Intell.53, 25868–25880 (2023), doi: 10.1007/s10489-023-04915-8

  30. [30]

    M. T. Albrijawi, R. Alhajj, LSTM-driven drug design using SELFIES for target-focused de novo generation of HIV-1 protease inhibitor candidates for AIDS treatment.PLOS ONE19, 1–30 (2024), doi:10.1371/journal.pone.0303597

  31. [31]

    IBM Research,https://huggingface.co/ibm-research/materials.selfies-ted

  32. [32]

    Savit, H

    A. Savit, H. Sahu, S. Shukla, W. Xiong, R. Ramprasad, polyBART: A Chemical Linguist for Polymer Property Prediction and Generative Design (2025),https://arxiv.org/abs/ 2506.04233

  33. [33]

    Raffel,et al., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Trans- former (2023),https://arxiv.org/abs/1910.10683

    C. Raffel,et al., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Trans- former (2023),https://arxiv.org/abs/1910.10683. 23

  34. [34]

    Edwards,et al., Translation between Molecules and Natural Language (2022),https: //arxiv.org/abs/2204.11817

    C. Edwards,et al., Translation between Molecules and Natural Language (2022),https: //arxiv.org/abs/2204.11817

  35. [35]

    Rothchild, A

    D. Rothchild, A. Tamkin, J. Yu, U. Misra, J. Gonzalez, C5T5: Controllable Generation of Organic Molecules with Transformers (2021),https://arxiv.org/abs/2108.10307

  36. [36]

    J. Lu, Y. Zhang, Unified Deep Learning Model for Multitask Reaction Predictions with Expla- nation.J. Chem. Inf. Model.62(6), 1376–1387 (2022), doi:10.1021/acs.jcim.1c01467

  37. [37]

    Gurnani,et al., AI-assisted discovery of high-temperature dielectrics for energy storage

    R. Gurnani,et al., AI-assisted discovery of high-temperature dielectrics for energy storage. Nat. Commun.15, 6107 (2024), doi:10.1038/s41467-024-50413-x

  38. [38]

    S. Kim, C. M. Schroeder, N. E. Jackson, Open Macromolecular Genome: Generative Design of Synthetically Accessible Polymers.ACS Polymers Au3, 318–330 (2023), doi:10.1021/ acspolymersau.3c00003

  39. [39]

    M. Ohno, Y. Hayashi, Q. Zhang, Y. Kaneko, R. Yoshida, SMiPoly: Generation of a Synthe- sizable Polymer Virtual Library Using Rule-Based Polymerization Reactions.J. Chem. Inf. Model.63, 5539–5548 (2023), doi:10.1021/acs.jcim.3c00329

  40. [40]

    T. Yue, J. He, Y. Li, Polyuniverse: generation of a large-scale polymer library using rule-based polymerization reactions for polymer informatics.Digital Discovery3, 2465–2478 (2024), doi:10.1039/D4DD00196F

  41. [41]

    H. C. Kolb, M. G. Finn, K. B. Sharpless, Click Chemistry: Diverse Chemical Function from a Few Good Reactions.Angew. Chem. Int. Ed.40, 2004–2021 (2001)

  42. [42]

    Z. Geng, J. J. Shin, Y. Xi, C. J. Hawker, Click chemistry strategies for the accelerated synthesis of functional macromolecules.J. Polym. Sci.59, 963–1042 (2021), doi:https://doi.org/10.1002/ pol.20210126

  43. [43]

    Odian,Ring-Opening Polymerization(John Wiley & Sons, Ltd), chap

    G. Odian,Ring-Opening Polymerization(John Wiley & Sons, Ltd), chap. 7, pp. 544–618 (2004), doi:https://doi.org/10.1002/047147875X.ch7

  44. [44]

    Materials and methods are available as supplementary material. 24

  45. [45]

    Kuenneth, R

    C. Kuenneth, R. Ramprasad, polyBERT: a chemical language model to enable fully machine-driven ultrafast polymer informatics.Nat. Commun.14, 4099 (2023), doi:10.1038/ s41467-023-39868-6

  46. [46]

    Liang,et al., All organic polymer dielectrics for high-temperature energy storage from the classification of heat-resistant insulation grades.J

    Y. Liang,et al., All organic polymer dielectrics for high-temperature energy storage from the classification of heat-resistant insulation grades.J. Polym. Sci.61(22), 2777–2795 (2023), doi:https://doi.org/10.1002/pol.20230334

  47. [47]

    Landrum, RDKit: Open-source cheminformatics (2016),https://www.rdkit.org

    G. Landrum, RDKit: Open-source cheminformatics (2016),https://www.rdkit.org

  48. [48]

    C. M. Alder,et al., Updating and further expanding GSK’s solvent sustainability guide.Green Chem.18, 3879–3890 (2016), doi:10.1039/C6GC00611F

  49. [49]

    OpenAI, GPT-5-Nano: API Documentation,https://platform.openai.com/docs/ models/gpt-5-nano(2025)

  50. [50]

    Odian,Step Polymerization(John Wiley & Sons, Ltd), chap

    G. Odian,Step Polymerization(John Wiley & Sons, Ltd), chap. 2, pp. 39–197 (2004), doi: https://doi.org/10.1002/047147875X.ch2

  51. [51]

    Sterling, J

    T. Sterling, J. J. Irwin, ZINC 15 – Ligand Discovery for Everyone.J. Chem. Inf. Model.55(11), 2324–2337 (2015), doi:10.1021/acs.jcim.5b00559

  52. [52]

    Gaulton,et al., The ChEMBL database in 2017.Nucleic Acids Res.45, D945–D954 (2016), doi:10.1093/nar/gkw1074

    A. Gaulton,et al., The ChEMBL database in 2017.Nucleic Acids Res.45, D945–D954 (2016), doi:10.1093/nar/gkw1074

  53. [53]

    emolecules.com/

    eMolecules Inc., eMolecules database (n.d.), retrieved May 31, 2024, fromhttps://www. emolecules.com/

  54. [54]

    C. E. Hoyle, C. N. Bowman, Thiol–Ene Click Chemistry.Angew. Chem. Int. Ed.49(9), 1540–1573 (2010), doi:https://doi.org/10.1002/anie.200903924

  55. [55]

    A. B. Lowe, C. E. Hoyle, C. N. Bowman, Thiol-yne click chemistry: A powerful and versatile methodology for materials synthesis.J. Mater. Chem.20, 4745–4750 (2010), doi:10.1039/ B917102A. 25

  56. [56]

    Zhang,et al., Thiol–bromo click polymerization for multifunctional polymers: synthesis, light refraction, aggregation-induced emission and explosive detection.Polym

    Y. Zhang,et al., Thiol–bromo click polymerization for multifunctional polymers: synthesis, light refraction, aggregation-induced emission and explosive detection.Polym. Chem.6, 97– 105 (2015), doi:10.1039/C4PY01164C

  57. [57]

    Gandini, The furan/maleimide Diels–Alder reaction: A versatile click–unclick tool in macromolecular synthesis.Prog

    A. Gandini, The furan/maleimide Diels–Alder reaction: A versatile click–unclick tool in macromolecular synthesis.Prog. Polym. Sci.38, 1–29 (2013), doi:https://doi.org/10.1016/j. progpolymsci.2012.04.002

  58. [58]

    Wang,et al., SuFEx-Based Polysulfonate Formation from Ethenesulfonyl Fluoride–Amine Adducts.Angew

    H. Wang,et al., SuFEx-Based Polysulfonate Formation from Ethenesulfonyl Fluoride–Amine Adducts.Angew. Chem. Int. Ed.56, 11203–11208 (2017), doi:https://doi.org/10.1002/anie. 201701160

  59. [59]

    Collins, Z

    J. Collins, Z. Xiao, A. Espinosa-Gomez, B. P. Fors, L. A. Connal, Extremely rapid and versatile synthesis of high molecular weight step growth polymers via oxime click chemistry.Polym. Chem.7, 2581–2588 (2016), doi:10.1039/C6PY00372A

  60. [60]

    Priyadarsini,et al., SELF-BART : A Transformer-based Molecular Representation Model using SELFIES (2024),https://arxiv.org/abs/2410.12348

    I. Priyadarsini,et al., SELF-BART : A Transformer-based Molecular Representation Model using SELFIES (2024),https://arxiv.org/abs/2410.12348

  61. [61]

    T. Kudo, J. Richardson, SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing (2018),https://arxiv.org/abs/1808.06226

  62. [62]

    Sahu, K.-H

    H. Sahu, K.-H. Shen, J. H. Montoya, H. Tran, R. Ramprasad, Polymer Structure Predictor (PSP): A Python Toolkit for Predicting Atomic-Level Structural Models for a Range of Polymer Geometries.J. Chem. Theory Comput.18, 2737–2748 (2022)

  63. [63]

    Kresse, J

    G. Kresse, J. Furthm¨ uller, Efficiency of ab-initio total energy calculations for metals and semiconductors using a plane-wave basis set.Comput. Mater. Sci.6, 15–50 (1996), doi:https: //doi.org/10.1016/0927-0256(96)00008-0

  64. [64]

    J. P. Perdew, K. Burke, M. Ernzerhof, Generalized Gradient Approximation Made Simple. Phys. Rev. Lett.77, 3865–3868 (1996)

  65. [65]

    There are no competing interests to declare

    J. Heyd, G. E. Scuseria, M. Ernzerhof, Hybrid functionals based on a screened Coulomb potential.J. Chem. Phys.118, 8207–8215 (2003). 26 Acknowledgments Funding:The authors acknowledge financial support by the Office of Naval Research through grants N00014-19-1-2103 and N00014-20-1-2175. Author contributions:H. S. was the primary architect of the project, ...

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.