Pith. sign in

REVIEW 2 major objections 2 minor 45 references

No Gaussian release of neural hidden states sits in the moderate-utility, moderate-privacy middle against an adaptive attacker.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 16:11 UTC pith:UZCKXFF2

load-bearing objection We only have the abstract for Hidden-State Privacy; the cached full text is a different paper, so the empty-middle claim cannot be checked. the 2 major comments →

arxiv 2605.24042 v3 pith:UZCKXFF2 submitted 2026-05-21 cs.LG cs.AI

Hidden-State Privacy Has an Empty Middle

classification cs.LG cs.AI
keywords hidden-state privacyGaussian release mechanismsadaptive Mahalanobis attackerFisher informationinverse-Fisher diagonalprivacy-utility tradeoffsplit-memory transformerretrieval attack
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that releasing single-layer neural hidden states with Gaussian noise has an empty middle: across 1,536 tested release covariances, none deliver both moderate utility and moderate privacy once an adaptive retrieval attacker is allowed. A matching Fisher-ball lower bound shows why: any full-rank Gaussian release that keeps Fisher utility at constant order leaves a direction whose Mahalanobis signal grows linearly with hidden width, so uniform safety inside the Gaussian class is impossible. The unique best diagonal mechanism (inverse-Fisher) sits on a privacy/utility edge rather than filling that middle, and mechanisms that look strong under Euclidean retrieval collapse under the adaptive Mahalanobis attacker. A full-trajectory inverter recovers most clean GPT-2 prefixes under ordinary releases but none under the inverse-Fisher diagonal. The authors therefore reframe the problem away from designing better Gaussians and toward co-designing architecture or release structure, showing a split-memory transformer trained from scratch that substantially beats same-budget GPT baselines on a Mahalanobis privacy gain metric.

Core claim

Of 1,536 Gaussian release covariances tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate privacy against an adaptive retrieval attacker. Every full-rank Gaussian release at O(1) Fisher utility admits a direction whose Mahalanobis signal grows linearly in hidden width, ruling out uniform Gaussian safety and matching the empirical empty middle. The diagonal inverse-Fisher release is the unique minimax-optimal diagonal mechanism at a fixed first-order KL budget and the only release with worst-attacker top-1 at most 0.001 across a 32-point model-layer grid, yet it lives on an edge rather than in a usable interior. Architecture co-design (a split-memory

What carries the argument

The Fisher-ball lower bound: any full-rank Gaussian release whose Fisher utility stays O(1) necessarily leaves a coordinate direction whose Mahalanobis signal scales linearly with hidden width, so no uniform safety exists inside the Gaussian class. The companion object is the diagonal inverse-Fisher release Σ★_diag(K) = (2K/d) diag(1/F_ii), unique minimax-optimal among diagonal mechanisms at first-order KL budget K.

Load-bearing premise

The argument treats an adaptive Mahalanobis (or full-trajectory) retrieval attacker that knows the release covariance and Fisher structure as the right threat, and treats the chosen moderate utility and privacy cutoffs as the right operational thresholds.

What would settle it

Find a full-rank Gaussian release covariance that, at O(1) Fisher utility on a wide hidden layer, keeps adaptive Mahalanobis top-1 retrieval below the paper’s moderate-privacy threshold on the same 32-point model-layer grid, or show a direction-free bound that does not grow with width.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Mechanism design that stays inside isotropic or full-rank Gaussian noise cannot fill a usable privacy–utility interior for hidden-state release.
  • Evaluations that only use Euclidean retrieval will overstate privacy; adaptive Mahalanobis (and sequence) attackers must be the default test.
  • The inverse-Fisher diagonal is the default safe diagonal release at a fixed KL budget, but operators should expect an edge tradeoff rather than a comfortable middle.
  • Gains large enough to matter will come from architecture or release co-design (e.g., split memory), not from retuning Gaussian covariances alone.
  • Pretrained GPT-style models are far weaker on the paper’s Mahalanobis privacy gain metric than models trained from scratch with split memory under the same budget.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the empty-middle result holds for multi-layer or residual streams as well, API providers that currently expose intermediate activations would need structural isolation, not just more noise.
  • The linear-in-width Mahalanobis growth suggests width scaling laws and privacy may be in tension unless the release interface itself is redesigned.
  • A natural next stress test is whether non-Gaussian or quantized releases can occupy the middle that Gaussians cannot, or whether the same Fisher geometry reappears.
  • Split-memory training may trade ordinary language-modeling quality or transfer for privacy gain; measuring that transfer cost would decide whether the co-design path is deployable.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. From the abstract alone, the paper claims that of 1,536 tested full-rank Gaussian release covariances for single-layer hidden-state privacy, none simultaneously achieve moderate utility and moderate privacy against an adaptive retrieval attacker. It asserts a complementary Fisher-ball lower bound: every full-rank Gaussian release at O(1) Fisher utility admits a direction whose Mahalanobis signal grows linearly in hidden width, ruling out uniform Gaussian safety in that class. The diagonal inverse-Fisher mechanism Σ★_diag(K)=(2K/d) diag(1/F_ii) is claimed unique minimax-optimal among diagonal releases at first-order KL budget K and the only release with worst-attacker top-1 ≤ 0.001 on a 32-point model-layer grid, yet to sit on a privacy/utility edge rather than in the middle. Supporting results include collapse of a generalized-eigen mechanism under Mahalanobis attack, a sequence inverter recovering 94% of clean GPT-2 prefixes (0% under Σ_diag), and a split-memory transformer with G_Mah ∈ [20,33] at 90M and a 6–24× advantage over same-budget GPT baselines. The abstract concludes that hidden-state release should be reframed as architecture or release co-design rather than Gaussian mechanism design.

Significance. If the empty-middle empirical claim and the matching Fisher-ball lower bound hold under clearly stated threat and utility definitions, the result would be a substantial negative result for the Gaussian release class in hidden-state privacy, with clear implications for private inference, KV-cache sharing, and intermediate-representation APIs. Naming a unique minimax-optimal diagonal mechanism and showing that Euclidean-Pareto gains can vanish under adaptive Mahalanobis attack would be useful guidance. The architectural co-design direction (split-memory transformer) is a constructive follow-on rather than pure impossibility. These strengths cannot be credited as verified here: the supplied full-text body is a different manuscript (LLM-AutoSciLab / arXiv:2605.24043), so proofs, experimental protocol, and architecture results are not checkable from the materials provided for 2605.24042.

major comments (2)
  1. Manuscript identity mismatch: the abstract, title, and arXiv id concern Hidden-State Privacy / Gaussian hidden-state release, but the full manuscript body supplied for review is LLM-AutoSciLab (closed-loop scientific discovery; ActiveSciBench; NewtonBench). No section, equation, table, or figure of the claimed paper is present. The Fisher-ball lower bound, the 1,536-covariance protocol, the definition of moderate utility/privacy, the adaptive Mahalanobis attacker, Σ★_diag optimality, the 32-layer grid, the sequence-inverter numbers, and the split-memory transformer results therefore cannot be verified. This is load-bearing: the central empty-middle claim is uncheckable on the materials given.
  2. Even restricting attention to the abstract, the empty-middle conclusion depends on operational cutoffs for “moderate” utility and privacy and on a threat model in which the attacker has adaptive Mahalanobis (or full-trajectory) access with knowledge of the release covariance and Fisher structure. Without the missing sections that define Fisher utility, the KL budget K, the top-1 thresholds, and the attacker information set, one cannot assess whether the middle is empty under weaker or differently informed attackers, or whether the thresholds are set so stringently that the conclusion is an artifact of the cutoffs. Those definitions are load-bearing for the main claim and are not available in the provided body.
minor comments (2)
  1. Abstract notation is dense (Σ★_diag(K), G_Mah, first-order KL budget K) without a one-line definition of Fisher utility or the adaptive attacker; once the correct manuscript is supplied, a short notation paragraph early in the introduction would help non-specialists.
  2. The abstract’s “13× Pareto reduction under Euclidean retrieval” that “collapses to 100% top-1 under the adaptive Mahalanobis attacker” is a strong contrast; the correct paper should make the Euclidean vs. Mahalanobis attacker comparison explicit in a single table or figure so the collapse is not only narrative.

Circularity Check

0 steps flagged

No circularity can be established: the provided full text is a different paper, and the target abstract shows no by-construction reductions.

full rationale

The claimed paper is Hidden-State Privacy Has an Empty Middle (arXiv:2605.24042). The CACHEABLE full manuscript is instead LLM-AutoSciLab (arXiv:2605.24043), so the Fisher-ball lower bound, the 1,536-covariance grid, the minimax argument for Σ★_diag, the adaptive Mahalanobis attacker, and the split-memory transformer results are not present to walk. On the only available target text—the abstract—none of the load-bearing claims reduce by construction to their inputs: the empty-middle statement is an empirical count over tested Gaussian releases; the complementary lower bound is stated as a proved existence of a linear-in-width Mahalanobis direction under full-rank O(1) Fisher utility, not as a redefinition of that utility; the diagonal inverse-Fisher release is claimed unique minimax-optimal at first-order KL budget K among diagonal mechanisms, which is a standard mechanism-design claim rather than a fitted parameter renamed as a prediction; and the architectural G_Mah numbers are reported experimental outcomes against baselines. No self-citation chain, uniqueness import, or ansatz smuggling appears in the abstract. Without the actual derivation sections, circularity cannot be exhibited by quote-and-reduction, so the honest score is 0 with empty steps. Residual concerns about threshold choice and attacker class are assumption/scope issues, not circularity.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 2 invented entities

Abstract-only review. Load-bearing modeling choices visible from the abstract: Gaussian release class, Fisher utility and first-order KL budget K, Mahalanobis/adaptive retrieval attacker, single-layer hidden-state focus, and the operational 'moderate' utility/privacy cutoffs that define the empty middle. No free parameters or invented physical entities are named with fitted values in the abstract; the diagonal mechanism formula is derived from Fisher diagonals and budget K rather than fit to attack success.

free parameters (2)
  • first-order KL budget K
    Mechanism family is parameterized by K; the diagonal inverse-Fisher release is defined at budget K. Abstract does not report a data-fit for K itself, but all privacy/utility statements are relative to chosen K.
  • moderate utility / moderate privacy thresholds
    The empty-middle claim depends on operational cutoffs for 'moderate' on both axes; abstract does not give numerical thresholds beyond top-1 ≤ 0.001 for the diagonal edge case.
axioms (3)
  • domain assumption Releases are full-rank Gaussian with covariance chosen by the defender; utility is measured via Fisher information (O(1) Fisher utility) and privacy via adaptive Mahalanobis / retrieval top-1.
    Defines the mechanism class and attacker against which the empty middle and lower bound are stated.
  • domain assumption First-order KL budget K is the right resource constraint for comparing Gaussian releases.
    Minimax optimality of Σ★_diag is claimed at this budget.
  • ad hoc to paper Single-layer hidden-state release is the right primary setting; multi-layer or full-trajectory behavior is secondary (sequence inverter is reported as supporting evidence).
    Abstract focuses the empty-middle claim on single-layer releases.
invented entities (2)
  • split-memory transformer no independent evidence
    purpose: Architectural co-design that separates memory streams to achieve high G_Mah while keeping language-modeling loss competitive.
    Positive constructive result offered after Gaussian mechanisms fail; independent evidence would require released models/code and external attack evaluations.
  • diagonal inverse-Fisher release Σ★_diag(K) no independent evidence
    purpose: Unique minimax-optimal diagonal Gaussian mechanism at KL budget K; only release with worst-attacker top-1 ≤ 0.001 on the reported grid.
    Named closed-form mechanism; not a new physical entity but a paper-specific optimal object within the diagonal subclass.

pith-pipeline@v1.1.0-grok45 · 16098 in / 2950 out tokens · 30892 ms · 2026-07-12T16:11:31.236334+00:00 · methodology

0 comments
read the original abstract

Of $1{,}536$ Gaussian release covariances we tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate privacy against an adaptive retrieval attacker. We prove a complementary Fisher-ball lower bound: every full-rank Gaussian release at $O(1)$ Fisher utility admits a direction whose Mahalanobis signal grows linearly in hidden width, ruling out uniform Gaussian safety in the class and matching the empirical empty middle. The diagonal inverse-Fisher release $\Sigma^\star_{\mathrm{diag}}(\mathcal{K}) = (2\mathcal{K}/d)\,\mathrm{diag}(1/F_{ii})$ is the unique minimax-optimal diagonal mechanism at first-order KL budget $\mathcal{K}$ and the only release with worst-attacker top-1 $\le 0.001$ at every point of a 32 model-layer grid, but it sits on a privacy/utility edge rather than filling the middle. A generalized-eigen mechanism reaching $13\times$ Pareto reduction under Euclidean retrieval collapses to $100\%$ top-1 under the adaptive Mahalanobis attacker, and a full-trajectory sequence inverter recovers $94\%$ of clean GPT-2 prefixes but $0\%$ under $\Sigma_{\mathrm{diag}}$. A split-memory transformer trained from scratch reaches $G_{\mathrm{Mah}} \in [20, 33]$ at 90M and maintains a $6$--$24\times$ advantage over same-budget GPT baselines from 30M to 1B at a fixed-token language-modeling loss penalty; pretrained models top out at 9.3. These results reframe hidden-state release from mechanism-design within the Gaussian class to architecture or release co-design.

Figures

Figures reproduced from arXiv: 2605.24042 by Alexander Okezue Bell.

Figure 1
Figure 1. Figure 1: Attack success vs. KL cost for four defenses at [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The empty middle. Every Gaussian release cell across the 5-model 32-layer sweep, plotted [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: SMT architecture. Each layer’s hidden state is factored into a narrow predictive trunk [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Perturbation analysis by category. (A) The scaffold fraction of feature flips (green) is [PITH_FULL_IMAGE:figures/full_fig_p013_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: KL divergence under PB versus PI projection at k = 128 across five models. The gap between the two curves widens from 0.5× on GPT-2 Small (identity more destructive) to 153× on DeepSeek-R1-14B (behavior more destructive), matching the direction-flip reported in [PITH_FULL_IMAGE:figures/full_fig_p017_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Utility-margin Pareto frontier on GPT-2 Small at [PITH_FULL_IMAGE:figures/full_fig_p018_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Gradient-covariance eigenspectrum on GPT-2 Medium at layer 12. The spectrum spans [PITH_FULL_IMAGE:figures/full_fig_p019_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Cumulative gradient-covariance energy versus subspace dimension across six models. The [PITH_FULL_IMAGE:figures/full_fig_p019_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: (A) Cumulative Fisher energy Ek vs. k/d for the ten models in [PITH_FULL_IMAGE:figures/full_fig_p020_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Behavior-margin fraction mB/mfull vs. p k/d at k = 128, ℓ = L/2. Dashed line is y = p k/d from Proposition 1; linear fit gives R2 = 0.93. geometry because the inter-prefix difference distribution is close to isotropic at the relevant scale, so any rank-k projector, Fisher-aligned or not, captures a fraction of ℓ2 mass close to k/d. The GPT-2 direction flip is absent from the margin table, and the Mistral-… view at source ↗
Figure 11
Figure 11. Figure 11: (A) Per-prefix anti-correlation between behavior-projection and identity-projection frac [PITH_FULL_IMAGE:figures/full_fig_p024_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Gradient-energy concentration for reasoning (DeepSeek-R1-Distill-Qwen-14B) versus [PITH_FULL_IMAGE:figures/full_fig_p025_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: CLT feature analysis across 600 sampled features (50 per layer). (A) Behavioral score [PITH_FULL_IMAGE:figures/full_fig_p032_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Circuit graph for “The capital of France is.” Token nodes (blue, bottom) connect to [PITH_FULL_IMAGE:figures/full_fig_p034_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: CLT mechanistic summary on GPT-2 Small. (A) Reconstruction error under five ablation [PITH_FULL_IMAGE:figures/full_fig_p034_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Original utility-margin sweep on GPT-2 Small. (A) Utility cost [PITH_FULL_IMAGE:figures/full_fig_p035_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: (A) Empirical worst-case Mahalanobis max∆∈Aemp ∆⊤Σ −1 α ∆, normalized to α = 0, averaged over layers within each model. Bands are ±1 standard deviation across layers and seeds. (B) Realised retrieval top-1 attack success at matched utility budget. Dotted line marks α = 1.0, the theoretical Fisher-ball minimax point. Three findings. First, the qualitative direction predicted by Theorem 15 holds robustly: t… view at source ↗
Figure 18
Figure 18. Figure 18: GEuc vs GMah across the five models. The Euclidean predictor spans three orders of magnitude and massively overstates the practical privacy gap. The Mahalanobis predictor ranges from 1.71 to 8.30 and is the correct scalar against an adaptive attacker. Mechanism σ KL ℓ2-full PI -restricted M isotropic 1.0 7.12 0.936 0.904 0.936 isotropic 2.0 8.04 0.192 0.162 0.192 complement 1.0 7.11 0.948 0.904 0.948 comp… view at source ↗
Figure 19
Figure 19. Figure 19: Best-attack top-1 at σ = 5 vs relative layer depth across 5 models. A: isotropic. B: diagonal-minimax Σ ⋆ diag(K) = (2K/d)D−1 (implementation parameterization Σdiag = σ 2diag(1/Fii), σ 2 = 2K/d). Isotropic fails on every modern model except mid-Mistral; Σdiag is at ≤ 0.001 on every one of the 32 points tested. The two compounding effects are (i) isotropic’s utility cost 1 2 tr(F Σiso) = 1 2 σ 2 tr(F) shri… view at source ↗
Figure 20
Figure 20. Figure 20: Leading-order mean KL vs empirical ε(δ = 10−6 ) on GPT-2 Small for isotropic Gaussian and Σ ⋆ Mah releases. Σ ⋆ Mah achieves 20–30% lower ε than iso on GPT-2 at every σ, but the ordering reverses on Mistral and Phi-2 (iso is tighter) and reverts to Σ ⋆ Mah again on Qwen3-14B / DeepSeek (see [PITH_FULL_IMAGE:figures/full_fig_p058_20.png] view at source ↗
Figure 21
Figure 21. Figure 21: Ablation KL for each attention head at five layers of GPT-2 Small. The maximum per-head [PITH_FULL_IMAGE:figures/full_fig_p062_21.png] view at source ↗
Figure 22
Figure 22. Figure 22: Predictive quotient release. The frozen language model is split at layer [PITH_FULL_IMAGE:figures/full_fig_p063_22.png] view at source ↗
Figure 23
Figure 23. Figure 23: (A) GMah vs. parameters. SMT (orange circles) maintains a 6–24× advantage from 30M to 1B. Matched-d baselines (blue squares) remain flat at ≈ 1.3. Faint markers at 90M are individual seeds. (B) E128 (top-128 Fisher mass) decays in SMT as the trunk dimension r grows. The dashed reference is min(1, 128/r), the value expected if Fisher were uniform on the trunk. (C) Margin mass qB in the top-128 Fisher subsp… view at source ↗
Figure 24
Figure 24. Figure 24: Training-time objectives on GPT-2 Small. (A) Trajectory in margin-fraction space. [PITH_FULL_IMAGE:figures/full_fig_p070_24.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

45 extracted references · 7 linked inside Pith

  1. [1]

    Nikhil Abhyankar, Sanchit Kabra, Saaketh Desai, and Chandan K. Reddy. LLEMA: Evolution- ary search with LLMs for multi-objective materials discovery. InThe Fourteenth International Conference on Learning Representations, 2026

  2. [2]

    The rise of self-driving labs in chemical and materials sciences.Nature Synthesis, 2:483 – 492, 2023

    Milad Abolhasani and Eugenia Kumacheva. The rise of self-driving labs in chemical and materials sciences.Nature Synthesis, 2:483 – 492, 2023

  3. [3]

    Autodiscovery: Open-ended scientific discovery via bayesian surprise

    Dhruv Agarwal, Bodhisattwa Prasad Majumder, Reece Adamson, Megha Chakravorty, Satvika Reddy Gavireddy, Aditya Parashar, Harshit Surana, Bhavana Dalvi Mishra, Andrew McCallum, Ashish Sabharwal, and Peter Clark. Autodiscovery: Open-ended scientific discovery via bayesian surprise. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026

  4. [4]

    The impact of large lan- guage models on scientific discovery: a preliminary study using gpt-4.arXiv preprint arXiv:2311.07361, 2023

    Microsoft Research AI4Science and Microsoft Azure Quantum. The impact of large lan- guage models on scientific discovery: a preliminary study using gpt-4.arXiv preprint arXiv:2311.07361, 2023

  5. [5]

    Deep batch active learning for drug discovery

    Michael Bailey, Saeed Moayedpour, Ruijiang Li, Alejandro Corrochano-Navarro, Alexander Kötter, Lorenzo Kogler-Anele, Saleh Riahi, Christoph Grebner, Gerhard Hessler, Hans Matter, Marc Bianciotto, Pablo Mas, Ziv Bar-Joseph, and Sven Jager. Deep batch active learning for drug discovery. January 2024

  6. [6]

    Pouya Behzadifar, Parshin Shojaee, Sanchit Kabra, Kazem Meidani, and Chandan K. Reddy. Decompose, adapt, and evolve: Towards efficient scientific equation discovery with large language models. InNeurIPS 2025 AI for Science Workshop, 2025

  7. [7]

    Discrimination among mechanistic models.Technomet- rics, 9(1):57–71, 1967

    George EP Box and WILLIAM J Hill. Discrimination among mechanistic models.Technomet- rics, 9(1):57–71, 1967

  8. [8]

    Ai4research: A survey of artificial intelligence for scientific research.arXiv preprint arXiv:2507.01903, 2025

    Qiguang Chen, Mingda Yang, Libo Qin, Jinhao Liu, Zheng Yan, Jiannan Guan, Dengyun Peng, Yiyan Ji, Hanjing Li, Mengkang Hu, et al. Ai4research: A survey of artificial intelligence for scientific research.arXiv preprint arXiv:2507.01903, 2025

  9. [9]

    Hypospace: Evaluating llm creativity as set-valued hypothesis generators under underdetermination.arXiv preprint arXiv:2510.15614, 2025

    Tingting Chen, Beibei Lin, Zifeng Yuan, Qiran Zou, Hongyu He, Anirudh Goyal, Yew-Soon Ong, and Dianbo Liu. Hypospace: Evaluating llm creativity as set-valued hypothesis generators under underdetermination.arXiv preprint arXiv:2510.15614, 2025

  10. [10]

    A large-scale benchmark for network inference from single-cell perturbation data.Communica- tions Biology, 8(1):412, 2025

    Mathieu Chevalley, Yusuf H Roohani, Arash Mehrjou, Jure Leskovec, and Patrick Schwab. A large-scale benchmark for network inference from single-cell perturbation data.Communica- tions Biology, 8(1):412, 2025

  11. [11]

    Interpretable machine learning for science with pysr and symbolicregression

    Miles Cranmer. Interpretable machine learning for science with pysr and symbolicregression. jl. arXiv preprint arXiv:2305.01582, 2023

  12. [12]

    ODEFormer: Symbolic regression of dynamical systems with transformers

    Stéphane d’Ascoli, Sören Becker, Philippe Schwaller, Alexander Mathis, and Niki Kilbertus. ODEFormer: Symbolic regression of dynamical systems with transformers. InThe Twelfth International Conference on Learning Representations, 2024

  13. [13]

    Autoscilab: A self-driving laboratory for interpretable scientific discovery

    Saaketh Desai, Sadhvikas Addamane, Jeffrey Y Tsao, Igal Brener, Laura P Swiler, Remi Dingreville, and Prasad P Iyer. Autoscilab: A self-driving laboratory for interpretable scientific discovery. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 146–154, 2025

  14. [14]

    Matllmsearch: Crystal struc- ture discovery with evolution-guided large language models.arXiv preprint arXiv:2502.20933, 2025

    Jingru Gan, Peichen Zhong, Yuanqi Du, Yanqiao Zhu, Chenru Duan, Haorui Wang, Daniel Schwalbe-Koda, Carla P Gomes, Kristin A Persson, and Wei Wang. Matllmsearch: Crystal struc- ture discovery with evolution-guided large language models.arXiv preprint arXiv:2502.20933, 2025

  15. [15]

    Symbolic regression with a learned concept library.Advances in Neural Information Processing Systems, 37:44678–44709, 2024

    Arya Grayeli, Atharva Sehgal, Omar Costilla-Reyes, Miles Cranmer, and Swarat Chaudhuri. Symbolic regression with a learned concept library.Advances in Neural Information Processing Systems, 37:44678–44709, 2024. 10

  16. [16]

    Olympus: a benchmarking framework for noisy optimization and experiment planning.Machine Learning: Science and Technology, 2(3):035021, 2021

    Florian Häse, Matteo Aldeghi, Riley J Hickman, Loïc M Roch, Melodie Christensen, Elena Liles, Jason E Hein, and Alán Aspuru-Guzik. Olympus: a benchmarking framework for noisy optimization and experiment planning.Machine Learning: Science and Technology, 2(3):035021, 2021

  17. [17]

    Characterization and greedy learning of interventional markov equivalence classes of directed acyclic graphs.The Journal of Machine Learning Research, 13(1):2409–2464, 2012

    Alain Hauser and Peter Bühlmann. Characterization and greedy learning of interventional markov equivalence classes of directed acyclic graphs.The Journal of Machine Learning Research, 13(1):2409–2464, 2012

  18. [18]

    Sequential optimal experimental design of perturbation screens guided by multi-modal priors.bioRxiv, 2023

    Kexin Huang, Romain Lopez, Jan-Christian Hütter, Takamasa Kudo, Antonio Rios, and Aviv Regev. Sequential optimal experimental design of perturbation screens guided by multi-modal priors.bioRxiv, 2023

  19. [19]

    Inferring regulatory networks from expression data using tree-based methods.PLoS ONE, 5, 2010

    Vân Anh Huynh-Thu, Alexandre Irrthum, Louis Wehenkel, and Pierre Geurts. Inferring regulatory networks from expression data using tree-based methods.PLoS ONE, 5, 2010

  20. [20]

    Generating literature-driven scientific theories at scale.arXiv preprint arXiv:2601.16282, 2026

    Peter Jansen, Peter Clark, Doug Downey, and Daniel S Weld. Generating literature-driven scientific theories at scale.arXiv preprint arXiv:2601.16282, 2026

  21. [21]

    Active symbolic discovery of ordinary differential equations via phase portrait sketching

    Nan Jiang, Md Nasim, and Yexiang Xue. Active symbolic discovery of ordinary differential equations via phase portrait sketching. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 17626–17634, 2025

  22. [22]

    Sanchit Kabra, Shobhnik Kriplani, Parshin Shojaee, and Chandan K. Reddy. SURFACEBENCH: A geometry-aware benchmark for symbolic surface discovery.Transactions on Machine Learning Research, 2026

  23. [23]

    On-the- fly closed-loop materials discovery via bayesian active learning.Nature communications, 11(1):5966, 2020

    A Gilad Kusne, Heshan Yu, Changming Wu, Huairuo Zhang, Jason Hattrick-Simpers, Brian DeCost, Suchismita Sarker, Corey Oses, Cormac Toher, Stefano Curtarolo, et al. On-the- fly closed-loop materials discovery via bayesian active learning.Nature communications, 11(1):5966, 2020

  24. [24]

    Kyro, Anton Morgunov, Rafael I

    Gregory W. Kyro, Anton Morgunov, Rafael I. Brent, and Victor S. Batista. Chemspaceal: An efficient active learning methodology applied to protein-specific molecular generation.Journal of Chemical Information and Modeling, 64(3):653–665, January 2024

  25. [25]

    Integrated systems for computational scientific discovery.Proceedings of the AAAI Conference on Artificial Intelligence, 38(20):22598–22606, Mar

    Pat Langley. Integrated systems for computational scientific discovery.Proceedings of the AAAI Conference on Artificial Intelligence, 38(20):22598–22606, Mar. 2024

  26. [26]

    Julia Ling, Maxwell Hutchinson, Erin Antono, Sean Paradiso, and Bryce Meredig. High- dimensional materials and process optimization using data-driven experimental design with well-calibrated uncertainty estimates.Integrating Materials and Manufacturing Innovation, 6(3):207–217, 2017

  27. [27]

    B. P. MacLeod, F. G. L. Parlane, T. D. Morrissey, F. Häse, L. M. Roch, K. E. Dettelbach, R. Moreira, L. P. E. Yunker, M. B. Rooney, J. R. Deeth, V . Lai, G. J. Ng, H. Situ, R. H. Zhang, M. S. Elliott, T. H. Haley, D. J. Dvorak, A. Aspuru-Guzik, J. E. Hein, and C. P. Berlinguette. Self-driving laboratory for accelerated discovery of thin-film materials.Sci...

  28. [28]

    Data-driven discovery with large generative models.arXiv preprint arXiv:2402.13610, 2024

    Bodhisattwa Prasad Majumder, Harshit Surana, Dhruv Agarwal, Sanchaita Hazra, Ashish Sabharwal, and Peter Clark. Data-driven discovery with large generative models.arXiv preprint arXiv:2402.13610, 2024

  29. [29]

    Melnikov, Hendrik Poulsen Nautrup, Mario Krenn, Vedran Dunjko, Markus Tiersch, Anton Zeilinger, and Hans J

    Alexey A. Melnikov, Hendrik Poulsen Nautrup, Mario Krenn, Vedran Dunjko, Markus Tiersch, Anton Zeilinger, and Hans J. Briegel. Active learning machine learns to create new quantum experiments.Proceedings of the National Academy of Sciences, 115(6):1221–1226, 2018

  30. [30]

    Practical optimal experiment design with probabilistic programs.arXiv preprint arXiv:1608.05046, 2016

    Long Ouyang, Michael Henry Tessler, Daniel Ly, and Noah Goodman. Practical optimal experiment design with probabilistic programs.arXiv preprint arXiv:1608.05046, 2016

  31. [31]

    Mundhenk, Claudio Prata Santiago, Soo Kyung Kim, and Joanne Taery Kim

    Brenden K Petersen, Mikel Landajuela Larma, Terrell N. Mundhenk, Claudio Prata Santiago, Soo Kyung Kim, and Joanne Taery Kim. Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradients. InInternational Conference on Learning Representations, 2021. 11

  32. [32]

    Jalihal, Jeffrey N

    Aditya Pratapa, Amogh P. Jalihal, Jeffrey N. Law, Aditya Bharadwaj, and T. M. Murali. Benchmarking algorithms for gene regulatory network inference from single-cell transcriptomic data.bioRxiv, 2019

  33. [33]

    Active learning for efficient discovery of optimal gene combinations in the combinatorial perturbation space

    Jason Qin, Hans-Hermann Wessels, Carlos Fernandez-Granda, and Yuhan Hao. Active learning for efficient discovery of optimal gene combinations in the combinatorial perturbation space. In NeurIPS 2024 Workshop on AI for New Drug Modalities, 2024

  34. [34]

    Towards scientific discovery with generative ai: Progress, opportunities, and challenges

    Chandan K Reddy and Parshin Shojaee. Towards scientific discovery with generative ai: Progress, opportunities, and challenges. InProceedings of the AAAI conference on artificial intelligence, volume 39, pages 28601–28609, 2025

  35. [35]

    Mathematical discoveries from program search with large language models

    Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. Mathematical discoveries from program search with large language models. Nature, 625(7995):468–475, 2024

  36. [36]

    Genenetweaver: in silico bench- mark generation and performance profiling of network inference methods.Bioinformatics, 27(16):2263–2270, 08 2011

    Thomas Schaffter, Daniel Marbach, and Dario Floreano. Genenetweaver: in silico bench- mark generation and performance profiling of network inference methods.Bioinformatics, 27(16):2263–2270, 08 2011

  37. [37]

    Parshin Shojaee, Kazem Meidani, Shashank Gupta, Amir Barati Farimani, and Chandan K. Reddy. LLM-SR: Scientific equation discovery via programming with large language models. InThe Thirteenth International Conference on Learning Representations, 2025

  38. [38]

    Parshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani, Khoa D Doan, and Chandan K. Reddy. LLM-SRBench: A new benchmark for scientific equation discovery with large language models. InForty-second International Conference on Machine Learning, 2025

  39. [39]

    Pdebench: An extensive benchmark for scientific machine learning.Advances in neural information processing systems, 35:1596–1611, 2022

    Makoto Takamoto, Timothy Praditia, Raphael Leiteritz, Daniel MacKinlay, Francesco Alesiani, Dirk Pflüger, and Mathias Niepert. Pdebench: An extensive benchmark for scientific machine learning.Advances in neural information processing systems, 35:1596–1611, 2022

  40. [40]

    Ai feynman: A physics-inspired method for symbolic regression.Science advances, 6(16):eaay2631, 2020

    Silviu-Marian Udrescu and Max Tegmark. Ai feynman: A physics-inspired method for symbolic regression.Science advances, 6(16):eaay2631, 2020

  41. [41]

    Scientific discovery in the age of artificial intelligence.Nature, 620(7972):47–60, 2023

    Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, et al. Scientific discovery in the age of artificial intelligence.Nature, 620(7972):47–60, 2023

  42. [42]

    Efficient evolutionary search over chemical space with large language models

    Haorui Wang, Marta Skreta, Cher Tian Ser, Wenhao Gao, Lingkai Kong, Felix Strieth-Kalthoff, Chenru Duan, Yuchen Zhuang, Yue Yu, Yanqiao Zhu, Yuanqi Du, Alan Aspuru-Guzik, Kirill Neklyudov, and Chao Zhang. Efficient evolutionary search over chemical space with large language models. InThe Thirteenth International Conference on Learning Representations, 2025

  43. [43]

    Newtonbench: Benchmarking generalizable scientific law discovery in LLM agents

    Tianshi Zheng, Kelvin Kiu Wai Tam, Newt Nguyen Kim Hue Nam, Baixuan Xu, Zhaowei Wang, Cheng Jiayang, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Tianqing Fang, Yangqiu Song, Ginny Wong, and Simon See. Newtonbench: Benchmarking generalizable scientific law discovery in LLM agents. InThe Fourteenth International Conference on Learning Representations, 2026

  44. [44]

    Dags with no tears: Continuous optimization for structure learning.Advances in neural information processing systems, 31, 2018

    Xun Zheng, Bryon Aragam, Pradeep K Ravikumar, and Eric P Xing. Dags with no tears: Continuous optimization for structure learning.Advances in neural information processing systems, 31, 2018

  45. [45]

    bounds": {

    Yangqiaoyu Zhou, Haokun Liu, Tejes Srivastava, Hongyuan Mei, and Chenhao Tan. Hypothesis generation with large language models. InProceedings of the 1st Workshop on NLP for Science (NLP4Science), pages 117–139, 2024. 12 Reproducibility Statement To ensure reproducibility, we provide the relevant implementation and experimental details throughout the paper...