Pith. sign in

REVIEW 2 major objections 6 minor 64 references

Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable

T0 review · 2 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A two-sided audit with a decidable negative side shows that one fallible RNA oracle can inflate a 43/60 capability claim to 1/60 under a three-predictor panel.

desk verdict The decidable negative side and the 43-to-1 collapse are genuinely useful, but the empirical side is model-relative and the paper knows it; a non-thermodynamic judge is owed. read the letter →

arxiv 2608.00981 v1 pith:NGDTH7RX submitted 2026-08-02 cs.AI

classification cs.AI
keywords agenticsciencecapabilityauditingRNApseudoknotdesignoracleover-optimizationChomskyhierarchyrealizable-targetlanguagecross-predictoradjudicationself-improvingagents
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to give capability claims by self-improving AI-for-science systems a two-sided audit: one side states exactly what the prior verifier could not certify, and the other measures what a new procedure actually does under adjudicators it never optimized. On the negative side the audit is decidable: a pseudoknot-free folding oracle's output range lies entirely in the nested (context-free) language, so no policy using it can ever be certified to realize a crossing pseudoknot target—a theorem discharged before any run. On the positive side the paper shows how much a single fallible oracle can inflate a claim: an invented, solver-free operator solves $43/60$ crossing RNA targets under the predictor it optimizes, but only $1/60$ under a three-predictor panel, and on the same 43 targets a predictor it never saw confirms $2$ of its designs against $26$ for a minimum-free-energy solver ($p=8\times10^{-7}$). The paper also reports that two agent-written operators, frozen after invention, carry over to the held-out predictor about three times as often as the hand-built operator while spending a fraction of the oracle calls, though no mechanism for that difference is identified. If the audit is right, benchmark deltas, compression gates, and p-values are insufficient evidence for "new capability"; claims need an outside adjudicator, a compute control, and a decidable statement about the prior verifier's reach.

What carries the argument

The central object is the realizable-target language, the set of target structures a policy realizes through a given oracle. Lemma 1 caps it by the oracle's image, and Proposition 1 computes that image for the pseudoknot-free verifier, making the negative side decidable by a crossing test quadratic in the number of pairs. On the positive side, the carrying instruments are the three-predictor panel (pKiss, ProbKnot, ShapeKnots), the sealed-commitment protocol that freezes candidates before any held-out adjudication, and the conditional carry-over statistic: the probability that a predictor never in an arm's loop confirms a design, given that the arm's own success predicate is satisfied. This

What would settle it

Re-run the paired 43-target adjudication with a pseudoknot predictor from a different model class (for example, a learned predictor trained without nearest-neighbour thermodynamic features) in place of ShapeKnots. If the pKiss-optimized operator's designs are confirmed at a rate comparable to the minimum-free-energy solver's rather than 2 versus 26, the collapse and the carry-over gap are artifacts of shared thermodynamic model family, not evidence of oracle-gaming or agent capability.

Watch

Extended reading notes

Core claim

Capability claims need two facts: what the prior verifier could not certify, and what a new procedure realizes under an adjudicator outside its loop. The negative side here is a theorem: a pseudoknot-free oracle's output is always nested, so no policy can realize a crossing target through it. The positive side is empirical: an invented solver-free operator scores $43/60$ under pKiss but $1/60$ under a three-predictor panel; on the same 43 targets, a predictor it never saw confirms $2$ designs vs $26$ for a minimum-free-energy solver ($p=8\times10^{-7}$). Frozen agent-written operators carry over at $0.293$ vs $0.095$ ($p=5\times10^{-5}$) at a fraction of oracle calls; no mechanism is identif

Load-bearing premise

The load-bearing premise is that the three-predictor panel (pKiss, ProbKnot, ShapeKnots) is a meaningful arbiter of real pseudoknot formation; all three share nearest-neighbour thermodynamic parameters, two agree strongly on natives, and the conjunction recovers only about 1% of native pseudoknots, so if the panel is systematically biased or low-recall the measured 1/60 collapse and carry-over gap describe agreement with a model family rather than a real capability.

Editorial extensions

If this is right

  • A single-oracle benchmark delta or compression gate would have accepted the 43/60 result as a capability gain; the 43-to-1 collapse is invisible to any statistic computed from the system and its own oracle, so one-sided evidence is not enough.
  • The negative side is a theorem for any verifier with a characterized output range: once an oracle's image is known to exclude a target class, no search budget lets any policy be certified on that class through it.
  • Re-optimizing for cross-oracle agreement recovers only a small set of panel-unanimous designs (3–8/60 for the hand-built operator, up to 17/60 for agent-invented ones), so even with the audit in place, crossing design under a stringent panel is rare.
  • Frozen agent-written operators carry over to the held-out predictor at 0.293 versus 0.095 on 951 paired, target-clustered units while spending 4.6–10x fewer oracle calls; this establishes difference (D1) and not-bought-with-compute (D2), but mechanism transfer (D3) remains open.
  • The structure-aware clusters of Pseudobase++ limit natural-corpus transfer tests: 251 targets collapse to 60 clusters, only 24 clean, so direction can replicate but magnitudes cannot; the unbounded H-type family supplies the only estimable magnitude (+0.316 on a 100-target grid).

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the two-stage attribution gate (screen candidates by difference and correlation, then intervene on each survivor) should be portable to other learned-system comparisons; on the paper's data a purely observational screen would have produced two mutually independent false mechanisms with clean p-values, so pre-committing to intervention is a testable methodological norm.
  • Editorial extension: with pKiss and ShapeKnots forming a correlated block (κ=0.673) and ProbKnot's independence poorly determined at n=80, a unanimous three-predictor verdict is worth less than three votes; the natural next audit should either weight the panel by measured agreement or add a non-thermodynamic predictor before trusting absolute confirmation rates.
  • Editorial extension: the positive control's advantage survives compute, composition, and transplanted mechanisms, which suggests the difference may live in the search trajectory itself rather than in any design-level feature; the unbounded H-type grid is the place to test trajectory-level statistics such as acceptance timing and repair-step distributions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper proposes a two-sided audit for capability-acquisition claims in agentic science. The negative side is a decidable containment statement: for a pseudoknot-free oracle, the realizable-target language of any policy is contained in the nested language (Prop. 1), so no crossing target can be certified through that verifier. The positive side is a model-relative, panel-adjudicated empirical measurement. On a 60-target RNA pseudoknot design pool, an invented, solver-free operator reaches 43/60 under the pKiss oracle it optimizes, but only 1/60 under a three-predictor panel; a paired comparison on the same 43 targets gives 2/43 of its designs confirmed by ProbKnot versus 26/43 for a minimum-free-energy solver (exact McNemar p ≈ 8e-7). A positive control using frozen LLM-written operators achieves higher carry-over to a held-out predictor at lower oracle-call counts (0.293 vs 0.095 on 951 paired units, target-clustered p = 5e-5). The paper reports extensive caveats, including correlated predictors, low native recall, its own failure to separate from undirected search, and the absence of a non-thermodynamic adjudicator.

Significance. If accepted, the paper provides a transferable audit protocol that separates exact verifier-range confinement from measured empirical gains and detects single-oracle over-optimization. The formal side is correct and cleanly discharged offline, though it is a near-tautological corollary of the oracle's definition. The empirical side is unusually careful: exact paired statistics, target-clustered intervals, deterministic bit-reproducible re-runs, released scripts, frozen committed predictions, and seven interventional attribution tests, four of which expose previously plausible false mechanisms. The paper also reports its own failures in detail, which strengthens confidence in the instrument. The central limitation is that all positive claims are relative to a panel of three thermodynamically correlated predictors; the paper explicitly acknowledges this but the headline numbers should not be read as biophysical ground truth. The main contribution is a rigorous, honestly bounded demonstration of oracle-gaming and a practical template for auditing such claims.

major comments (2)
  1. [§5.1, Tables 4–5, §3.1] The headline '43/60 under one oracle, 1/60 under three' and the paired '2 vs 26' contrast are load-bearing for the paper's central claim that single-oracle capability claims can be massively inflated. The manuscript discloses, but mainly in later sections, that the three adjudicators are not independent voices: all share nearest-neighbour thermodynamic parameters, pKiss and ShapeKnots agree at κ=0.673, the three-way conjunction has ~1% native recall, and ShapeKnots is run in unvalidated de-novo mode. ProbKnot, used for the 2-vs-26 comparison, shares the thermodynamic parameter family with DesiRNA's MFE objective. As a result, the measured collapse demonstrates overfitting to a single thermodynamic predictor family, not necessarily the absence of any real pseudoknot-design capability. Because the paper's conceptual claim is about fallible oracles in general, and because the paper itself s
  2. [§4, Prop. 1] The formal negative side is correct, but as stated it is an immediate corollary of Lemma 1 plus the definition of Ω_cf (im Ω_cf ⊆ L_nest). The paper's contribution list presents this as a formalization of 'the object a pumping lemma can constrain' and as a decidable negative side, which overstates the novelty. The paper itself calls the exclusion 'shallow' in the introduction, but the contribution list and title emphasis on 'negative side decidable' invite a stronger reading. I recommend explicitly describing Prop. 1 as a framing/accounting result—useful because it fixes the right object and discharges the negative side offline—rather than as a new theorem. The substantive novelty lies in the audit protocol and the empirical demonstration, not in the one-line containment proof.
minor comments (6)
  1. [Abstract and title] Several numbers run into adjacent words without spaces in the rendered text, e.g. 'credits43designs' in the title and 'puts84%' in the abstract. Please fix the typography.
  2. [§5.3 and Table 2] The hand-built operator's carry-over is quoted as 0.091 (48/530) in §5.3 and as 0.095 in the D1 row of Table 2. These come from different denominators (all cells vs. the paired n=951 units), but the presentation is confusing; clarify the denominator each time.
  3. [References] Reference [4] (Baker and Pixley) appears in the bibliography but is not cited in the body text. Either add a citation or remove the entry.
  4. [Figure 5] The caption lists raw counts (57/222, 23/209, etc.) while the bars are labeled with rates. Add the rates to the bar labels or the axis for readability, since the two are not immediately reconcilable.
  5. [Table 7] The term 'development/evaluation pool' is potentially misleading because the same pool is used to fix all hyperparameters. Consider renaming it 'development pool' or explicitly stating that all rates on it are in-development performance, so readers do not infer a held-out test set.
  6. [§5.4] The median Δbp for the development/evaluation pool is reported as 0.059 in Table 7, while the text says the median of the 43 pKiss-solves is Δbp=0. Clarify that these are medians over different subsets (all targets vs. solved targets).

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: Prop. 1 is a declared-verifier consequence, and the positive claims are adjudicated by a predictor outside every optimization loop.

full rationale

The paper's formal negative side is self-contained and not circular. Lemma 1 (R_Omega(pi) subseteq im Omega) is definitional, and Proposition 1 is the scoped consequence of declaring Omega_cf to be a pseudoknot-free folder whose output alphabet is {(,),.}; the paper itself calls the exclusion 'shallow' and does not upgrade the empirical 0/60 floor into a theorem. No fitted parameter is renamed as a prediction. On the empirical side, the headline 43/60 -> 1/60 collapse is a measured overfitting phenomenon: the operator optimizes pKiss, and the panel adds predictors outside that loop, with ShapeKnots the only member in no arm's objective. The D1 positive control is paired, target-clustered, and scored by the same held-out predictor, so it does not reduce to fitting the winning arm's own objective. The one self-citation, companion report [8] with overlapping first authorship, supplies the verifier-ladder vocabulary and some panel calibration, but the paper re-measures the load-bearing numbers itself (Cohen's kappa on 80 natives, ShapeKnots' 11.3% native ceiling), so [8] is not the sole support for any central claim. The strongest legitimate concern, that the three panel predictors share nearest-neighbour thermodynamic parameters and pKiss/ShapeKnots agree at kappa=0.673, is a correctness or external-validity limitation, not a circularity; the paper explicitly concedes that carry-over follows shared model family rather than correctness. The derivation chain does not reduce to its inputs, and the verdict is a normal, mostly honest non-finding.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The audit's exact side rests on a software-output fact about ViennaRNA; the empirical side rests on the panel of three thermodynamic predictors, whose limits the paper itself measures. The listed free parameters are design budgets and thresholds; none is fitted to the headline effect, but the 15% clustering threshold and the 60-draw probe budget shape the reported intervals and the 84% attribution.

free parameters (4)
  • Single-linkage clustering threshold = 0.15
    Defines structural clusters for the disjoint pool and for cluster-resampled confidence intervals; changing it changes the 24-cluster/27-target result and interval widths (Table 7, App. D).
  • Random compatible sequences per target in search-free probe = 60
    Budget for the Delta_bp margin probe; sampling more can only lower Delta_bp, so the 84% share is a lower bound and the residual +0.016 an upper bound, but it is still a hand-chosen budget (Eq. 8, Section 5.4).
  • Equal-allowance pKiss call cap = 150
    Chosen above the roughly 142.7 calls realized under the wall-clock cap; sets the equal-allowance ablation design (App. B, Table 6).
  • ProbKnot acceptance gate = pk_bp <= 3
    Operator hyperparameter (App. F) used in the repair loop; hand-chosen and affects which designs are accepted before adjudication.
assumptions (5)
  • standard math Ogden lemma / pumping lemma for context-free languages
    Used to argue that the H-type family L_H is not context-free (App. A, [18, 36]); invoked, not proved.
  • standard math H-type pseudoknot family is non-context-free and in MCFG dimension 2
    Invoked from Nebel and Weinberg [33] and Kato et al. [22]; the paper verifies only the interleaving witness for its own L_H family (App. A).
  • domain assumption ViennaRNA RNA.fold outputs only nested structures, so im Omega_cf is a subset of L_nest
    A property of the pinned software version (Turner 2004, pseudoknot-free); the audit's exact negative side depends on this (Section 2, Prop. 1, Table 8).
  • domain assumption Panel unanimity among pKiss, ProbKnot, and ShapeKnots approximates real pseudoknot formation
    Load-bearing for all positive rates; the paper measures the panel's limits (shared nearest-neighbour parameters, kappa = 0.673 for pKiss/ShapeKnots, about 1% native recall for the three-way conjunction) and explicitly labels the positive side model-relative (Section 3.1, Section 6.2).
  • domain assumption ShapeKnots run de novo without SHAPE restraints is an acceptable held-out adjudicator
    The paper runs ShapeKnots outside its calibrated regime and flags it as unvalidated (App. C.2); all D1/D2 numbers depend on this adjudicator.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable." pith.science (2026). https://pith.science/paper/NGDTH7RX

@misc{pith2026260800981,
  author       = {Pith},
  title        = {Pith review of: Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NGDTH7RX}},
  note         = {Machine review of arXiv:2608.00981}
}
read the original abstract

When a self-improving AI-for-science system claims a new capability, the evidence is usually a benchmark delta, a description-length gate, or a p-value. None separates a real gain from extra search, from a changed verifier, or from adaptation to a fallible oracle. We build a two-sided audit whose negative side is a formal fact: a pseudoknot-free oracle provably cannot represent a crossing base pair, so the prior verifier's range is bounded exactly, offline, before any run. "New" is relative to the agent's prior self, never to the base model. First, how far a single fallible oracle can inflate a capability claim. An invented, solver-free operator solves 43/60 crossing RNA targets under the predictor it optimizes, above a context-free floor of 0/60; under three predictors, 1/60 survives. Paired on the same 43 targets, a predictor the operator never saw confirms 2 of its designs against 26 for a minimum-free-energy solver (p = 8e-7). No statistic computed from the system and its own oracle sees that gap. Second, agent-written procedures can beat a human-written one under a judge no objective can flatter, at a fraction of the compute. Of six frontier models, the two whose operators ran without timeouts carry over at 0.293 against our 0.095 (n = 951 paired units, target-clustered [+0.108, +0.297], p = 5e-5) while spending 4.6-10x fewer oracle calls. Three rungs: difference under an outside adjudicator (reached), not bought with compute (reached, both directions), mechanism identified and transferable (not reached; seven candidates tested, none moves the statistic). The ceiling is the panel itself: its three predictors share nearest-neighbour thermodynamic parameters, two agreeing at kappa = 0.673. The audit is as unsparing about our own system: matched undirected search is an exact zero, and a search-free probe puts 84% of our headline effect on targets a random sequence already solves.

Figures

Figures reproduced from arXiv: 2608.00981 by the authors.

Figure 1
Figure 1. End-to-end architecture. A target is graded by family type and crossing-stem count and assigned to a pool of [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. The realizable-target language, confinement, and the ascent. [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. The panel exposes oracle-gaming (measured, [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: The statement about the headline effect that uses no interval at all. [PITH_FULL_IMAGE:figures/full_fig_p022_4.png]
Figure 5
Figure 5. Figure 5: Conditional carry-over C (Eq. (7)): the metric-neutral comparison. Each arm is scored on the fraction of its own successes that a predictor never in its loop confirms, so the denominators differ by design and that is what makes the statistic neutral. Bars are exact bin…
Figure 6
Figure 6. Figure 6: Nine interventions on the hand-built operator, against the arm they are trying to [PITH_FULL_IMAGE:figures/full_fig_p026_6.png]
Figure 7
Figure 7. Figure 7: The largest caveat this paper raises against its own positive result. [PITH_FULL_IMAGE:figures/full_fig_p028_7.png]
Figure 8
Figure 8. Figure 8: Why the rising confirmed rate on LH does not mean what it looks like. The panel￾unanimous rate rmix (grey) rises with the family parameter, which reads as a fixed-description operator holding up as the family grows. It is a product, rmix = rin × carry-over, and only th…
Figure 9
Figure 9. Figure 9: Why Table 10 is alphabetical [PITH_FULL_IMAGE:figures/full_fig_p044_9.png]
Figure 10
Figure 10. Figure 10: One rung above panel-unanimity, as a distribution rather than a median. [PITH_FULL_IMAGE:figures/full_fig_p045_10.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

64 extracted references · 38 canonical work pages

  1. [1]

    Buldú, Michael Stich, and Susanna C

    Jacobo Aguirre, Javier M. Buldú, Michael Stich, and Susanna C. Manrubia. Topological struc- ture of the space of phenotypes: The case of RNA neutral networks.PLoS ONE, 6(10):e26324,

  2. [2]

    Concrete problems in AI safety.arXiv preprint arXiv:1606.06565, 2016

    Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. Concrete problems in AI safety.arXiv preprint arXiv:1606.06565, 2016

  3. [3]

    Pellock, Tamuka M

    Ivan Anishchenko, Samuel J. Pellock, Tamuka M. Chidyausiku, et al. De novo protein design by deep network hallucination.Nature, 600:547–552, 2021

  4. [4]

    Baker and Alden F

    Kirby A. Baker and Alden F. Pixley. Polynomial interpolation and the Chinese remainder theorem for algebraic systems.Mathematische Zeitschrift, 143(2):165–174, 1975. doi: 10.1007/ BF01187059

  5. [5]

    Stanislav Bellaousov and David H. Mathews. ProbKnot: Fast prediction of RNA secondary structure including pseudoknots.RNA, 16(10):1870–1880, 2010. doi: 10.1261/rna.2125310

  6. [6]

    Designing RNA secondary structures is hard.Journal of Computational Biology, 27(3):302–316, 2020

    Édouard Bonnet, Paweł Rz¸ ażewski, and Florian Sikora. Designing RNA secondary structures is hard.Journal of Computational Biology, 27(3):302–316, 2020. doi: 10.1089/cmb.2019.0420. prelim. RECOMB 2018; arXiv:1710.11513

  7. [7]

    INFO-RNA—a fast approach to inverse RNA folding.Bioin- formatics, 22(15):1823–1831, 2006

    Anke Busch and Rolf Backofen. INFO-RNA—a fast approach to inverse RNA folding.Bioin- formatics, 22(15):1823–1831, 2006. doi: 10.1093/bioinformatics/btl194

  8. [8]

    “solved” is a choice of verifier, not a fact: Deterministic cost–robustness frontiers of verifier choice in RNA and SAT

    Wenhui Chen, Qingqing Mao, and Chenghua Wang. “solved” is a choice of verifier, not a fact: Deterministic cost–robustness frontiers of verifier choice in RNA and SAT. Technical report, InceptLabs-Shanghai, 2026. companion technical report

Show all 64 references
  1. [9]

    Three models for the description of language.IRE Transactions on Informa- tion Theory, 2(3):113–124, 1956

    Noam Chomsky. Three models for the description of language.IRE Transactions on Informa- tion Theory, 2(3):113–124, 1956. doi: 10.1109/TIT.1956.1056813

  2. [10]

    Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, and Pedro A. Ortega. Neural networks and the Chomsky hierarchy. InICLR, 2023. arXiv:2207.02098

  3. [11]

    Dowell and Sean R

    Robin D. Dowell and Sean R. Eddy. Evaluation of several lightweight stochastic context-free grammars for RNA secondary structure prediction.BMC Bioinformatics, 5:71, 2004. doi: 10.1186/1471-2105-5-71

  4. [12]

    Eddy, Anders Krogh, and Graeme Mitchison.Biological Sequence Analysis

    Richard Durbin, Sean R. Eddy, Anders Krogh, and Graeme Mitchison.Biological Sequence Analysis. Cambridge University Press, 1998

  5. [13]

    Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mo- hammadamin Barekatain, Alexander Novikov, Francisco J. R. Ruiz, Julian Schrittwieser, Grze- gorz Swirszcz, David Silver, Demis Hassabis, and Pushmeet Kohli. Discovering faster matrix multipl...

  6. [14]

    Scaling laws for reward model overoptimization

    Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. InICML, PMLR 202, pages 10835–10866, 2023. arXiv:2210.10760. 47

  7. [15]

    Towards an AI co-scientist, 2025

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, et al. Towards an AI co-scientist, 2025. arXiv:2502.18864

  8. [16]

    Grünwald.The Minimum Description Length Principle

    Peter D. Grünwald.The Minimum Description Length Principle. MIT Press, 2007

  9. [17]

    Hajdin, Stanislav Bellaousov, Wayne Huggins, Christopher W

    Christine E. Hajdin, Stanislav Bellaousov, Wayne Huggins, Christopher W. Leonard, David H. Mathews, and Kevin M. Weeks. Accurate SHAPE-directed RNA secondary structure modeling, including pseudoknots.PNAS, 110(14):5498–5503, 2013. doi: 10.1073/pnas.1219988110

  10. [18]

    Hopcroft and Jeffrey D

    John E. Hopcroft and Jeffrey D. Ullman.Introduction to Automata Theory, Languages, and Computation. Addison-Wesley, 1979

  11. [19]

    Li, Emmanuel Candès, and Jure Leskovec

    Kexin Huang, Ying Jin, Ryan Li, Michael Y. Li, Emmanuel Candès, and Jure Leskovec. Automated hypothesis validation with agentic sequential falsifications.arXiv preprint arXiv:2502.09858, 2025. ICML 2025

  12. [20]

    The RNA shapes studio.Bioinformatics, 31(3):423–425,

    Stefan Janssen and Robert Giegerich. The RNA shapes studio.Bioinformatics, 31(3):423–425,

  13. [21]

    Aravind K. Joshi. Tree adjoining grammars: How much context-sensitivity is required to pro- vide reasonable structural descriptions? In David R. Dowty, Lauri Karttunen, and Arnold M. Zwicky, editors,Natural Language Parsing, pages 206–250. Cambridge University Press, 1985

  14. [22]

    RNA pseudoknotted structure prediction using stochastic multiple context-free grammar.IPSJ Digital Courier, 2:655–664, 2006

    Yuki Kato, Hiroyuki Seki, and Tadao Kasami. RNA pseudoknotted structure prediction using stochastic multiple context-free grammar.IPSJ Digital Courier, 2:655–664, 2006. doi: 10. 2197/ipsjdc.2.655

  15. [23]

    RNA secondary structure prediction using stochastic context- free grammars and evolutionary history.Bioinformatics, 15(6):446–454, 1999

    Bjarne Knudsen and Jotun Hein. RNA secondary structure prediction using stochastic context- free grammars and evolutionary history.Bioinformatics, 15(6):446–454, 1999. doi: 10.1093/ bioinformatics/15.6.446

  16. [24]

    Koodli, Boris Rudolfs, Hannah K

    Rohan V. Koodli, Boris Rudolfs, Hannah K. Wayment-Steele, and Rhiju Das. Redesigning the Eterna100 for the Vienna 2 folding engine.bioRxiv, 2021. doi: 10.1101/2021.08.26.457839. preprint

  17. [25]

    Are agents just automata? on the formal equivalence between agentic AI and the Chomsky hierarchy.arXiv preprint arXiv:2510.23487, 2025

    Roham Koohestani, Ziyou Li, Anton Podkopaev, and Maliheh Izadi. Are agents just automata? on the formal equivalence between agentic AI and the Chomsky hierarchy.arXiv preprint arXiv:2510.23487, 2025

  18. [26]

    Cambridge University Press, 1978

    Imre Lakatos.The Methodology of Scientific Research Programmes. Cambridge University Press, 1978

  19. [27]

    Let’s verify step by step.arXiv preprint arXiv:2305.20050, 2023

    Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step.arXiv preprint arXiv:2305.20050, 2023. ICLR 2024

  20. [28]

    AIGS: Generating science from AI-powered automated falsification.arXiv preprint arXiv:2411.11910, 2024

    Zijun Liu, Kaiming Liu, Yiqi Zhu, Xuanyu Lei, Zonghan Yang, Zhenhe Zhang, Peng Li, and Yang Liu. AIGS: Generating science from AI-powered automated falsification.arXiv preprint arXiv:2411.11910, 2024

  21. [29]

    Bernhart, Christian Höner zu Siederdissen, Hakim Tafer, Christoph Flamm, Peter F

    Ronny Lorenz, Stephan H. Bernhart, Christian Höner zu Siederdissen, Hakim Tafer, Christoph Flamm, Peter F. Stadler, and Ivo L. Hofacker. ViennaRNA package 2.0.Algorithms for Molecular Biology, 6(1):26, 2011. doi: 10.1186/1748-7188-6-26. 48

  22. [30]

    Lyngsø and Christian N

    Rune B. Lyngsø and Christian N. S. Pedersen. RNA pseudoknot prediction in energy- based models.Journal of Computational Biology, 7(3-4):409–427, 2000. doi: 10.1089/ 106652700750050862

  23. [31]

    Mankowitz, Andrea Michi, Anton Zhernov, Marco Gelmi, Marco Selvi, Cosmin Padu- raru, et al

    Daniel J. Mankowitz, Andrea Michi, Anton Zhernov, Marco Gelmi, Marco Selvi, Cosmin Padu- raru, et al. Faster sorting algorithms discovered using deep reinforcement learning.Nature, 618 (7964):257–263, 2023. doi: 10.1038/s41586-023-06004-9

  24. [32]

    McCaskill

    John S. McCaskill. The equilibrium partition function and base pair binding probabilities for RNA secondary structure.Biopolymers, 29(6-7):1105–1119, 1990. doi: 10.1002/bip.360290621

  25. [33]

    Nebel and Frank Weinberg

    Markus E. Nebel and Frank Weinberg. Algebraic and combinatorial properties of common RNA pseudoknot classes.Journal of Computational Biology, 19(10):1134–1150, 2012. PMC3469209

  26. [34]

    AlphaEvolve: A coding agent for scientific and algorithmic discovery.arXiv preprint arXiv:2506.13131, 2025

    Alexander Novikov, Ngân V˜ u, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, et al. AlphaEvolve: A coding agent for scientific and algorithmic discovery.arXiv preprint arXiv:2506.13131, 2025

  27. [35]

    Jacobson

    Ruth Nussinov and Ann B. Jacobson. Fast algorithm for predicting the secondary structure of single-stranded RNA.PNAS, 77(11):6309–6313, 1980. doi: 10.1073/pnas.77.11.6309

  28. [36]

    A helpful result for proving inherent ambiguity.Mathematical Systems Theory, 2(3):191–194, 1968

    William Ogden. A helpful result for proving inherent ambiguity.Mathematical Systems Theory, 2(3):191–194, 1968. doi: 10.1007/BF01694004

  29. [37]

    Popper.The Logic of Scientific Discovery

    Karl R. Popper.The Logic of Scientific Discovery. Hutchinson, London, 1959. transl. of Logik der Forschung, 1934

  30. [38]

    Design, implementation and evaluation of a practical pseudoknot folding algorithm based on thermodynamics.BMC Bioinformatics, 5:104, 2004

    Jens Reeder and Robert Giegerich. Design, implementation and evaluation of a practical pseudoknot folding algorithm based on thermodynamics.BMC Bioinformatics, 5:104, 2004. doi: 10.1186/1471-2105-5-104

  31. [39]

    Elena Rivas and Sean R. Eddy. A dynamic programming algorithm for RNA structure pre- diction including pseudoknots.Journal of Molecular Biology, 285(5):2053–2068, 1999. doi: 10.1006/jmbi.1998.2436

  32. [40]

    Elena Rivas and Sean R. Eddy. The language of RNA: A formal grammar that includes pseudoknots.Bioinformatics, 16(4):334–340, 2000. doi: 10.1093/bioinformatics/16.4.334

  33. [41]

    RNAInvBench and Pseudobase++: Pseudoknotted RNA inverse-design tar- gets, 2024

    RNAInvBench. RNAInvBench and Pseudobase++: Pseudoknotted RNA inverse-design tar- gets, 2024. inverse-design benchmark suite

  34. [42]

    Pawan Kumar, Emilien Dupont, Francisco J

    Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, Pushmeet Kohli, and Alhussein Fawzi. Mathematical discoveries from program search with larg...

  35. [43]

    Learning to design RNA

    Frederic Runge, Danny Stoll, Stefan Falkner, and Frank Hutter. Learning to design RNA. In ICLR, 2019. arXiv:1812.11951

  36. [44]

    Saira Mian, Kimmen Sjölander, Rebecca C

    Yasubumi Sakakibara, Michael Brown, Richard Hughey, I. Saira Mian, Kimmen Sjölander, Rebecca C. Underwood, and David Haussler. Stochastic context-free grammars for tRNA modeling.Nucleic Acids Research, 22(23):5112–5120, 1994. doi: 10.1093/nar/22.23.5112. 49

  37. [45]

    Schaefer

    Thomas J. Schaefer. The complexity of satisfiability problems. InSTOC, pages 216–226, 1978. doi: 10.1145/800133.804350

  38. [46]

    Stadler, and Ivo L

    Peter Schuster, Walter Fontana, Peter F. Stadler, and Ivo L. Hofacker. From sequences to shapes and back: A case study in RNA secondary structures.Proceedings of the Royal Society of London B, 255(1344):279–284, 1994. doi: 10.1098/rspb.1994.0040

  39. [47]

    David B. Searls. The linguistics of DNA.American Scientist, 80(6):579–591, 1992

  40. [48]

    David B. Searls. The language of genes.Nature, 420(6912):211–217, 2002. doi: 10.1038/ nature01255

  41. [49]

    On multiple context- freegrammars.Theoretical Computer Science, 88(2):191–229, 1991

    Hiroyuki Seki, Takashi Matsumura, Mamoru Fujii, and Tadao Kasami. On multiple context- freegrammars.Theoretical Computer Science, 88(2):191–229, 1991. doi: 10.1016/0304-3975(91) 90374-B

  42. [50]

    Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. Defining and characterizing reward hacking. InNeurIPS, volume 35, 2022. arXiv:2209.13085

  43. [51]

    Vijay-Shanker and David J

    K. Vijay-Shanker and David J. Weir. The equivalence of four extensions of context-free gram- mars.Mathematical Systems Theory, 27(6):511–546, 1994. doi: 10.1007/BF01191624

  44. [52]

    Self-revisingdiscoverysystemsforscience: Acategorical framework for agentic AI.arXiv preprint arXiv:2606.01444, 2026

    FionaY.WangandMarkusJ.Buehler. Self-revisingdiscoverysystemsforscience: Acategorical framework for agentic AI.arXiv preprint arXiv:2606.01444, 2026

  45. [53]

    Voyager: An open-ended embodied agent with large language models, 2023

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, et al. Voyager: An open-ended embodied agent with large language models, 2023. arXiv:2305.16291

  46. [54]

    Watson, David Juergens, Nathaniel R

    Joseph L. Watson, David Juergens, Nathaniel R. Bennett, et al. De novo design of protein structure and function with RFdiffusion.Nature, 620:1089–1100, 2023

  47. [55]

    From AI for science to agentic science: A survey on autonomous scientific discovery, 2025

    Jiaqi Wei, Yuejin Yang, Xiang Zhang, et al. From AI for science to agentic science: A survey on autonomous scientific discovery, 2025. arXiv:2508.14111

  48. [56]

    Wirecki, Grzegorz Lach, Nagendar Goud Badepally, S

    Tomasz K. Wirecki, Grzegorz Lach, Nagendar Goud Badepally, S. Naeim Moafinejad, Farhang Jaryani, Gaja Klaudel, Kalina Nec, Eugene F. Baulin, and Janusz M. Bujnicki. DesiRNA: Structure-based design of RNA sequences with a replica exchange monte carlo approach.Nu- cleic Acids Re...

  49. [57]

    Position: Falsify, don’t just discover—AI-generated discoveries are not born scientific

    Ce Xu et al. Position: Falsify, don’t just discover—AI-generated discoveries are not born scientific. InICML, Position Paper Track, 2025. OpenReview SlgXCLZFj3

  50. [58]

    Zadeh, Conrad D

    Joseph N. Zadeh, Conrad D. Steenberg, Justin S. Bois, Brian R. Wolfe, Marshall B. Pierce, Asif R. Khan, Robert M. Dirks, and Niles A. Pierce. NUPACK: Analysis and design of nucleic acid systems.Journal of Computational Chemistry, 32(1):170–173, 2011. doi: 10.1002/jcc. 21596

  51. [59]

    Mathews, and Liang Huang

    Tianshuo Zhou, Ning Dai, Sizhen Li, Max Ward, David H. Mathews, and Liang Huang. RNA design via structure-aware multifrontier ensemble optimization.Bioinformatics, 39(Supplement 1):i563–i571, 2023. doi: 10.1093/bioinformatics/btad252

  52. [60]

    Mathews, and Liang Huang

    Tianshuo Zhou, Wei Yu Tang, David H. Mathews, and Liang Huang. Undesignable RNA struc- ture identification via rival structure generation and structure decomposition. InRECOMB, LNCS 14758, pages 270–287, 2024. doi: 10.1007/978-1-0716-3989-4_17. arXiv:2311.08339. 50

  53. [61]

    Mathews, and Liang Huang

    Tianshuo Zhou, Apoorv Malik, Wei Yu Tang, David H. Mathews, and Liang Huang. Scalable and interpretable identification of minimal undesignable RNA structure motifs with rotational invariance. InRECOMB, 2025. arXiv:2402.17206

  54. [62]

    Optimal computer folding of large RNA sequences using thermodynamics and auxiliary information.Nucleic Acids Research, 9(1):133–148, 1981

    Michael Zuker and Patrick Stiegler. Optimal computer folding of large RNA sequences using thermodynamics and auxiliary information.Nucleic Acids Research, 9(1):133–148, 1981. doi: 10.1093/nar/9.1.133. 51

  55. [2011]

    doi: 10.1371/journal.pone.0026324

  56. [2015]

    doi: 10.1093/bioinformatics/btu649

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.