REVIEW 2 major objections 2 minor 45 references
No Gaussian release of neural hidden states sits in the moderate-utility, moderate-privacy middle against an adaptive attacker.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 16:11 UTC pith:UZCKXFF2
load-bearing objection We only have the abstract for Hidden-State Privacy; the cached full text is a different paper, so the empty-middle claim cannot be checked. the 2 major comments →
Hidden-State Privacy Has an Empty Middle
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Of 1,536 Gaussian release covariances tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate privacy against an adaptive retrieval attacker. Every full-rank Gaussian release at O(1) Fisher utility admits a direction whose Mahalanobis signal grows linearly in hidden width, ruling out uniform Gaussian safety and matching the empirical empty middle. The diagonal inverse-Fisher release is the unique minimax-optimal diagonal mechanism at a fixed first-order KL budget and the only release with worst-attacker top-1 at most 0.001 across a 32-point model-layer grid, yet it lives on an edge rather than in a usable interior. Architecture co-design (a split-memory
What carries the argument
The Fisher-ball lower bound: any full-rank Gaussian release whose Fisher utility stays O(1) necessarily leaves a coordinate direction whose Mahalanobis signal scales linearly with hidden width, so no uniform safety exists inside the Gaussian class. The companion object is the diagonal inverse-Fisher release Σ★_diag(K) = (2K/d) diag(1/F_ii), unique minimax-optimal among diagonal mechanisms at first-order KL budget K.
Load-bearing premise
The argument treats an adaptive Mahalanobis (or full-trajectory) retrieval attacker that knows the release covariance and Fisher structure as the right threat, and treats the chosen moderate utility and privacy cutoffs as the right operational thresholds.
What would settle it
Find a full-rank Gaussian release covariance that, at O(1) Fisher utility on a wide hidden layer, keeps adaptive Mahalanobis top-1 retrieval below the paper’s moderate-privacy threshold on the same 32-point model-layer grid, or show a direction-free bound that does not grow with width.
If this is right
- Mechanism design that stays inside isotropic or full-rank Gaussian noise cannot fill a usable privacy–utility interior for hidden-state release.
- Evaluations that only use Euclidean retrieval will overstate privacy; adaptive Mahalanobis (and sequence) attackers must be the default test.
- The inverse-Fisher diagonal is the default safe diagonal release at a fixed KL budget, but operators should expect an edge tradeoff rather than a comfortable middle.
- Gains large enough to matter will come from architecture or release co-design (e.g., split memory), not from retuning Gaussian covariances alone.
- Pretrained GPT-style models are far weaker on the paper’s Mahalanobis privacy gain metric than models trained from scratch with split memory under the same budget.
Where Pith is reading between the lines
- If the empty-middle result holds for multi-layer or residual streams as well, API providers that currently expose intermediate activations would need structural isolation, not just more noise.
- The linear-in-width Mahalanobis growth suggests width scaling laws and privacy may be in tension unless the release interface itself is redesigned.
- A natural next stress test is whether non-Gaussian or quantized releases can occupy the middle that Gaussians cannot, or whether the same Fisher geometry reappears.
- Split-memory training may trade ordinary language-modeling quality or transfer for privacy gain; measuring that transfer cost would decide whether the co-design path is deployable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. From the abstract alone, the paper claims that of 1,536 tested full-rank Gaussian release covariances for single-layer hidden-state privacy, none simultaneously achieve moderate utility and moderate privacy against an adaptive retrieval attacker. It asserts a complementary Fisher-ball lower bound: every full-rank Gaussian release at O(1) Fisher utility admits a direction whose Mahalanobis signal grows linearly in hidden width, ruling out uniform Gaussian safety in that class. The diagonal inverse-Fisher mechanism Σ★_diag(K)=(2K/d) diag(1/F_ii) is claimed unique minimax-optimal among diagonal releases at first-order KL budget K and the only release with worst-attacker top-1 ≤ 0.001 on a 32-point model-layer grid, yet to sit on a privacy/utility edge rather than in the middle. Supporting results include collapse of a generalized-eigen mechanism under Mahalanobis attack, a sequence inverter recovering 94% of clean GPT-2 prefixes (0% under Σ_diag), and a split-memory transformer with G_Mah ∈ [20,33] at 90M and a 6–24× advantage over same-budget GPT baselines. The abstract concludes that hidden-state release should be reframed as architecture or release co-design rather than Gaussian mechanism design.
Significance. If the empty-middle empirical claim and the matching Fisher-ball lower bound hold under clearly stated threat and utility definitions, the result would be a substantial negative result for the Gaussian release class in hidden-state privacy, with clear implications for private inference, KV-cache sharing, and intermediate-representation APIs. Naming a unique minimax-optimal diagonal mechanism and showing that Euclidean-Pareto gains can vanish under adaptive Mahalanobis attack would be useful guidance. The architectural co-design direction (split-memory transformer) is a constructive follow-on rather than pure impossibility. These strengths cannot be credited as verified here: the supplied full-text body is a different manuscript (LLM-AutoSciLab / arXiv:2605.24043), so proofs, experimental protocol, and architecture results are not checkable from the materials provided for 2605.24042.
major comments (2)
- Manuscript identity mismatch: the abstract, title, and arXiv id concern Hidden-State Privacy / Gaussian hidden-state release, but the full manuscript body supplied for review is LLM-AutoSciLab (closed-loop scientific discovery; ActiveSciBench; NewtonBench). No section, equation, table, or figure of the claimed paper is present. The Fisher-ball lower bound, the 1,536-covariance protocol, the definition of moderate utility/privacy, the adaptive Mahalanobis attacker, Σ★_diag optimality, the 32-layer grid, the sequence-inverter numbers, and the split-memory transformer results therefore cannot be verified. This is load-bearing: the central empty-middle claim is uncheckable on the materials given.
- Even restricting attention to the abstract, the empty-middle conclusion depends on operational cutoffs for “moderate” utility and privacy and on a threat model in which the attacker has adaptive Mahalanobis (or full-trajectory) access with knowledge of the release covariance and Fisher structure. Without the missing sections that define Fisher utility, the KL budget K, the top-1 thresholds, and the attacker information set, one cannot assess whether the middle is empty under weaker or differently informed attackers, or whether the thresholds are set so stringently that the conclusion is an artifact of the cutoffs. Those definitions are load-bearing for the main claim and are not available in the provided body.
minor comments (2)
- Abstract notation is dense (Σ★_diag(K), G_Mah, first-order KL budget K) without a one-line definition of Fisher utility or the adaptive attacker; once the correct manuscript is supplied, a short notation paragraph early in the introduction would help non-specialists.
- The abstract’s “13× Pareto reduction under Euclidean retrieval” that “collapses to 100% top-1 under the adaptive Mahalanobis attacker” is a strong contrast; the correct paper should make the Euclidean vs. Mahalanobis attacker comparison explicit in a single table or figure so the collapse is not only narrative.
Circularity Check
No circularity can be established: the provided full text is a different paper, and the target abstract shows no by-construction reductions.
full rationale
The claimed paper is Hidden-State Privacy Has an Empty Middle (arXiv:2605.24042). The CACHEABLE full manuscript is instead LLM-AutoSciLab (arXiv:2605.24043), so the Fisher-ball lower bound, the 1,536-covariance grid, the minimax argument for Σ★_diag, the adaptive Mahalanobis attacker, and the split-memory transformer results are not present to walk. On the only available target text—the abstract—none of the load-bearing claims reduce by construction to their inputs: the empty-middle statement is an empirical count over tested Gaussian releases; the complementary lower bound is stated as a proved existence of a linear-in-width Mahalanobis direction under full-rank O(1) Fisher utility, not as a redefinition of that utility; the diagonal inverse-Fisher release is claimed unique minimax-optimal at first-order KL budget K among diagonal mechanisms, which is a standard mechanism-design claim rather than a fitted parameter renamed as a prediction; and the architectural G_Mah numbers are reported experimental outcomes against baselines. No self-citation chain, uniqueness import, or ansatz smuggling appears in the abstract. Without the actual derivation sections, circularity cannot be exhibited by quote-and-reduction, so the honest score is 0 with empty steps. Residual concerns about threshold choice and attacker class are assumption/scope issues, not circularity.
Axiom & Free-Parameter Ledger
free parameters (2)
- first-order KL budget K
- moderate utility / moderate privacy thresholds
axioms (3)
- domain assumption Releases are full-rank Gaussian with covariance chosen by the defender; utility is measured via Fisher information (O(1) Fisher utility) and privacy via adaptive Mahalanobis / retrieval top-1.
- domain assumption First-order KL budget K is the right resource constraint for comparing Gaussian releases.
- ad hoc to paper Single-layer hidden-state release is the right primary setting; multi-layer or full-trajectory behavior is secondary (sequence inverter is reported as supporting evidence).
invented entities (2)
-
split-memory transformer
no independent evidence
-
diagonal inverse-Fisher release Σ★_diag(K)
no independent evidence
read the original abstract
Of $1{,}536$ Gaussian release covariances we tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate privacy against an adaptive retrieval attacker. We prove a complementary Fisher-ball lower bound: every full-rank Gaussian release at $O(1)$ Fisher utility admits a direction whose Mahalanobis signal grows linearly in hidden width, ruling out uniform Gaussian safety in the class and matching the empirical empty middle. The diagonal inverse-Fisher release $\Sigma^\star_{\mathrm{diag}}(\mathcal{K}) = (2\mathcal{K}/d)\,\mathrm{diag}(1/F_{ii})$ is the unique minimax-optimal diagonal mechanism at first-order KL budget $\mathcal{K}$ and the only release with worst-attacker top-1 $\le 0.001$ at every point of a 32 model-layer grid, but it sits on a privacy/utility edge rather than filling the middle. A generalized-eigen mechanism reaching $13\times$ Pareto reduction under Euclidean retrieval collapses to $100\%$ top-1 under the adaptive Mahalanobis attacker, and a full-trajectory sequence inverter recovers $94\%$ of clean GPT-2 prefixes but $0\%$ under $\Sigma_{\mathrm{diag}}$. A split-memory transformer trained from scratch reaches $G_{\mathrm{Mah}} \in [20, 33]$ at 90M and maintains a $6$--$24\times$ advantage over same-budget GPT baselines from 30M to 1B at a fixed-token language-modeling loss penalty; pretrained models top out at 9.3. These results reframe hidden-state release from mechanism-design within the Gaussian class to architecture or release co-design.
Figures
Reference graph
Works this paper leans on
-
[1]
Nikhil Abhyankar, Sanchit Kabra, Saaketh Desai, and Chandan K. Reddy. LLEMA: Evolution- ary search with LLMs for multi-objective materials discovery. InThe Fourteenth International Conference on Learning Representations, 2026
2026
-
[2]
The rise of self-driving labs in chemical and materials sciences.Nature Synthesis, 2:483 – 492, 2023
Milad Abolhasani and Eugenia Kumacheva. The rise of self-driving labs in chemical and materials sciences.Nature Synthesis, 2:483 – 492, 2023
2023
-
[3]
Autodiscovery: Open-ended scientific discovery via bayesian surprise
Dhruv Agarwal, Bodhisattwa Prasad Majumder, Reece Adamson, Megha Chakravorty, Satvika Reddy Gavireddy, Aditya Parashar, Harshit Surana, Bhavana Dalvi Mishra, Andrew McCallum, Ashish Sabharwal, and Peter Clark. Autodiscovery: Open-ended scientific discovery via bayesian surprise. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026
2026
-
[4]
Microsoft Research AI4Science and Microsoft Azure Quantum. The impact of large lan- guage models on scientific discovery: a preliminary study using gpt-4.arXiv preprint arXiv:2311.07361, 2023
Pith/arXiv arXiv 2023
-
[5]
Deep batch active learning for drug discovery
Michael Bailey, Saeed Moayedpour, Ruijiang Li, Alejandro Corrochano-Navarro, Alexander Kötter, Lorenzo Kogler-Anele, Saleh Riahi, Christoph Grebner, Gerhard Hessler, Hans Matter, Marc Bianciotto, Pablo Mas, Ziv Bar-Joseph, and Sven Jager. Deep batch active learning for drug discovery. January 2024
2024
-
[6]
Pouya Behzadifar, Parshin Shojaee, Sanchit Kabra, Kazem Meidani, and Chandan K. Reddy. Decompose, adapt, and evolve: Towards efficient scientific equation discovery with large language models. InNeurIPS 2025 AI for Science Workshop, 2025
2025
-
[7]
Discrimination among mechanistic models.Technomet- rics, 9(1):57–71, 1967
George EP Box and WILLIAM J Hill. Discrimination among mechanistic models.Technomet- rics, 9(1):57–71, 1967
1967
-
[8]
Qiguang Chen, Mingda Yang, Libo Qin, Jinhao Liu, Zheng Yan, Jiannan Guan, Dengyun Peng, Yiyan Ji, Hanjing Li, Mengkang Hu, et al. Ai4research: A survey of artificial intelligence for scientific research.arXiv preprint arXiv:2507.01903, 2025
Pith/arXiv arXiv 2025
-
[9]
Tingting Chen, Beibei Lin, Zifeng Yuan, Qiran Zou, Hongyu He, Anirudh Goyal, Yew-Soon Ong, and Dianbo Liu. Hypospace: Evaluating llm creativity as set-valued hypothesis generators under underdetermination.arXiv preprint arXiv:2510.15614, 2025
Pith/arXiv arXiv 2025
-
[10]
A large-scale benchmark for network inference from single-cell perturbation data.Communica- tions Biology, 8(1):412, 2025
Mathieu Chevalley, Yusuf H Roohani, Arash Mehrjou, Jure Leskovec, and Patrick Schwab. A large-scale benchmark for network inference from single-cell perturbation data.Communica- tions Biology, 8(1):412, 2025
2025
-
[11]
Interpretable machine learning for science with pysr and symbolicregression
Miles Cranmer. Interpretable machine learning for science with pysr and symbolicregression. jl. arXiv preprint arXiv:2305.01582, 2023
Pith/arXiv arXiv 2023
-
[12]
ODEFormer: Symbolic regression of dynamical systems with transformers
Stéphane d’Ascoli, Sören Becker, Philippe Schwaller, Alexander Mathis, and Niki Kilbertus. ODEFormer: Symbolic regression of dynamical systems with transformers. InThe Twelfth International Conference on Learning Representations, 2024
2024
-
[13]
Autoscilab: A self-driving laboratory for interpretable scientific discovery
Saaketh Desai, Sadhvikas Addamane, Jeffrey Y Tsao, Igal Brener, Laura P Swiler, Remi Dingreville, and Prasad P Iyer. Autoscilab: A self-driving laboratory for interpretable scientific discovery. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 146–154, 2025
2025
-
[14]
Jingru Gan, Peichen Zhong, Yuanqi Du, Yanqiao Zhu, Chenru Duan, Haorui Wang, Daniel Schwalbe-Koda, Carla P Gomes, Kristin A Persson, and Wei Wang. Matllmsearch: Crystal struc- ture discovery with evolution-guided large language models.arXiv preprint arXiv:2502.20933, 2025
arXiv 2025
-
[15]
Symbolic regression with a learned concept library.Advances in Neural Information Processing Systems, 37:44678–44709, 2024
Arya Grayeli, Atharva Sehgal, Omar Costilla-Reyes, Miles Cranmer, and Swarat Chaudhuri. Symbolic regression with a learned concept library.Advances in Neural Information Processing Systems, 37:44678–44709, 2024. 10
2024
-
[16]
Olympus: a benchmarking framework for noisy optimization and experiment planning.Machine Learning: Science and Technology, 2(3):035021, 2021
Florian Häse, Matteo Aldeghi, Riley J Hickman, Loïc M Roch, Melodie Christensen, Elena Liles, Jason E Hein, and Alán Aspuru-Guzik. Olympus: a benchmarking framework for noisy optimization and experiment planning.Machine Learning: Science and Technology, 2(3):035021, 2021
2021
-
[17]
Characterization and greedy learning of interventional markov equivalence classes of directed acyclic graphs.The Journal of Machine Learning Research, 13(1):2409–2464, 2012
Alain Hauser and Peter Bühlmann. Characterization and greedy learning of interventional markov equivalence classes of directed acyclic graphs.The Journal of Machine Learning Research, 13(1):2409–2464, 2012
2012
-
[18]
Sequential optimal experimental design of perturbation screens guided by multi-modal priors.bioRxiv, 2023
Kexin Huang, Romain Lopez, Jan-Christian Hütter, Takamasa Kudo, Antonio Rios, and Aviv Regev. Sequential optimal experimental design of perturbation screens guided by multi-modal priors.bioRxiv, 2023
2023
-
[19]
Inferring regulatory networks from expression data using tree-based methods.PLoS ONE, 5, 2010
Vân Anh Huynh-Thu, Alexandre Irrthum, Louis Wehenkel, and Pierre Geurts. Inferring regulatory networks from expression data using tree-based methods.PLoS ONE, 5, 2010
2010
-
[20]
Generating literature-driven scientific theories at scale.arXiv preprint arXiv:2601.16282, 2026
Peter Jansen, Peter Clark, Doug Downey, and Daniel S Weld. Generating literature-driven scientific theories at scale.arXiv preprint arXiv:2601.16282, 2026
Pith/arXiv arXiv 2026
-
[21]
Active symbolic discovery of ordinary differential equations via phase portrait sketching
Nan Jiang, Md Nasim, and Yexiang Xue. Active symbolic discovery of ordinary differential equations via phase portrait sketching. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 17626–17634, 2025
2025
-
[22]
Sanchit Kabra, Shobhnik Kriplani, Parshin Shojaee, and Chandan K. Reddy. SURFACEBENCH: A geometry-aware benchmark for symbolic surface discovery.Transactions on Machine Learning Research, 2026
2026
-
[23]
On-the- fly closed-loop materials discovery via bayesian active learning.Nature communications, 11(1):5966, 2020
A Gilad Kusne, Heshan Yu, Changming Wu, Huairuo Zhang, Jason Hattrick-Simpers, Brian DeCost, Suchismita Sarker, Corey Oses, Cormac Toher, Stefano Curtarolo, et al. On-the- fly closed-loop materials discovery via bayesian active learning.Nature communications, 11(1):5966, 2020
2020
-
[24]
Kyro, Anton Morgunov, Rafael I
Gregory W. Kyro, Anton Morgunov, Rafael I. Brent, and Victor S. Batista. Chemspaceal: An efficient active learning methodology applied to protein-specific molecular generation.Journal of Chemical Information and Modeling, 64(3):653–665, January 2024
2024
-
[25]
Integrated systems for computational scientific discovery.Proceedings of the AAAI Conference on Artificial Intelligence, 38(20):22598–22606, Mar
Pat Langley. Integrated systems for computational scientific discovery.Proceedings of the AAAI Conference on Artificial Intelligence, 38(20):22598–22606, Mar. 2024
2024
-
[26]
Julia Ling, Maxwell Hutchinson, Erin Antono, Sean Paradiso, and Bryce Meredig. High- dimensional materials and process optimization using data-driven experimental design with well-calibrated uncertainty estimates.Integrating Materials and Manufacturing Innovation, 6(3):207–217, 2017
2017
-
[27]
B. P. MacLeod, F. G. L. Parlane, T. D. Morrissey, F. Häse, L. M. Roch, K. E. Dettelbach, R. Moreira, L. P. E. Yunker, M. B. Rooney, J. R. Deeth, V . Lai, G. J. Ng, H. Situ, R. H. Zhang, M. S. Elliott, T. H. Haley, D. J. Dvorak, A. Aspuru-Guzik, J. E. Hein, and C. P. Berlinguette. Self-driving laboratory for accelerated discovery of thin-film materials.Sci...
2020
-
[28]
Data-driven discovery with large generative models.arXiv preprint arXiv:2402.13610, 2024
Bodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal, Sanchaita Hazra, Ashish Sabharwal, and Peter Clark. Data-driven discovery with large generative models.arXiv preprint arXiv:2402.13610, 2024
Pith/arXiv arXiv 2024
-
[29]
Melnikov, Hendrik Poulsen Nautrup, Mario Krenn, Vedran Dunjko, Markus Tiersch, Anton Zeilinger, and Hans J
Alexey A. Melnikov, Hendrik Poulsen Nautrup, Mario Krenn, Vedran Dunjko, Markus Tiersch, Anton Zeilinger, and Hans J. Briegel. Active learning machine learns to create new quantum experiments.Proceedings of the National Academy of Sciences, 115(6):1221–1226, 2018
2018
-
[30]
Long Ouyang, Michael Henry Tessler, Daniel Ly, and Noah Goodman. Practical optimal experiment design with probabilistic programs.arXiv preprint arXiv:1608.05046, 2016
Pith/arXiv arXiv 2016
-
[31]
Mundhenk, Claudio Prata Santiago, Soo Kyung Kim, and Joanne Taery Kim
Brenden K Petersen, Mikel Landajuela Larma, Terrell N. Mundhenk, Claudio Prata Santiago, Soo Kyung Kim, and Joanne Taery Kim. Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradients. InInternational Conference on Learning Representations, 2021. 11
2021
-
[32]
Jalihal, Jeffrey N
Aditya Pratapa, Amogh P. Jalihal, Jeffrey N. Law, Aditya Bharadwaj, and T. M. Murali. Benchmarking algorithms for gene regulatory network inference from single-cell transcriptomic data.bioRxiv, 2019
2019
-
[33]
Active learning for efficient discovery of optimal gene combinations in the combinatorial perturbation space
Jason Qin, Hans-Hermann Wessels, Carlos Fernandez-Granda, and Yuhan Hao. Active learning for efficient discovery of optimal gene combinations in the combinatorial perturbation space. In NeurIPS 2024 Workshop on AI for New Drug Modalities, 2024
2024
-
[34]
Towards scientific discovery with generative ai: Progress, opportunities, and challenges
Chandan K Reddy and Parshin Shojaee. Towards scientific discovery with generative ai: Progress, opportunities, and challenges. InProceedings of the AAAI conference on artificial intelligence, volume 39, pages 28601–28609, 2025
2025
-
[35]
Mathematical discoveries from program search with large language models
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. Mathematical discoveries from program search with large language models. Nature, 625(7995):468–475, 2024
2024
-
[36]
Genenetweaver: in silico bench- mark generation and performance profiling of network inference methods.Bioinformatics, 27(16):2263–2270, 08 2011
Thomas Schaffter, Daniel Marbach, and Dario Floreano. Genenetweaver: in silico bench- mark generation and performance profiling of network inference methods.Bioinformatics, 27(16):2263–2270, 08 2011
2011
-
[37]
Parshin Shojaee, Kazem Meidani, Shashank Gupta, Amir Barati Farimani, and Chandan K. Reddy. LLM-SR: Scientific equation discovery via programming with large language models. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[38]
Parshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani, Khoa D Doan, and Chandan K. Reddy. LLM-SRBench: A new benchmark for scientific equation discovery with large language models. InForty-second International Conference on Machine Learning, 2025
2025
-
[39]
Pdebench: An extensive benchmark for scientific machine learning.Advances in neural information processing systems, 35:1596–1611, 2022
Makoto Takamoto, Timothy Praditia, Raphael Leiteritz, Daniel MacKinlay, Francesco Alesiani, Dirk Pflüger, and Mathias Niepert. Pdebench: An extensive benchmark for scientific machine learning.Advances in neural information processing systems, 35:1596–1611, 2022
2022
-
[40]
Ai feynman: A physics-inspired method for symbolic regression.Science advances, 6(16):eaay2631, 2020
Silviu-Marian Udrescu and Max Tegmark. Ai feynman: A physics-inspired method for symbolic regression.Science advances, 6(16):eaay2631, 2020
2020
-
[41]
Scientific discovery in the age of artificial intelligence.Nature, 620(7972):47–60, 2023
Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, et al. Scientific discovery in the age of artificial intelligence.Nature, 620(7972):47–60, 2023
2023
-
[42]
Efficient evolutionary search over chemical space with large language models
Haorui Wang, Marta Skreta, Cher Tian Ser, Wenhao Gao, Lingkai Kong, Felix Strieth-Kalthoff, Chenru Duan, Yuchen Zhuang, Yue Yu, Yanqiao Zhu, Yuanqi Du, Alan Aspuru-Guzik, Kirill Neklyudov, and Chao Zhang. Efficient evolutionary search over chemical space with large language models. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[43]
Newtonbench: Benchmarking generalizable scientific law discovery in LLM agents
Tianshi Zheng, Kelvin Kiu Wai Tam, Newt Nguyen Kim Hue Nam, Baixuan Xu, Zhaowei Wang, Cheng Jiayang, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Tianqing Fang, Yangqiu Song, Ginny Wong, and Simon See. Newtonbench: Benchmarking generalizable scientific law discovery in LLM agents. InThe Fourteenth International Conference on Learning Representations, 2026
2026
-
[44]
Dags with no tears: Continuous optimization for structure learning.Advances in neural information processing systems, 31, 2018
Xun Zheng, Bryon Aragam, Pradeep K Ravikumar, and Eric P Xing. Dags with no tears: Continuous optimization for structure learning.Advances in neural information processing systems, 31, 2018
2018
-
[45]
bounds": {
Yangqiaoyu Zhou, Haokun Liu, Tejes Srivastava, Hongyuan Mei, and Chenhao Tan. Hypothesis generation with large language models. InProceedings of the 1st Workshop on NLP for Science (NLP4Science), pages 117–139, 2024. 12 Reproducibility Statement To ensure reproducibility, we provide the relevant implementation and experimental details throughout the paper...
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.