REVIEW 5 major objections 5 minor 98 references
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
T0 review · 5 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Move AI evaluation from answers to verifiable investigations
desk verdict A serious evaluation-framework paper whose discovery claims outrun its verifiers; the TRACES/HDS6 machinery deserves referee time, but the AAV SOTA number needs external validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the environment–task–episode abstraction with its fixed episode interface and two verification channels. The episode interface is the single contract through which every solver and baseline passes: it fixes inputs, tools, budgets, and submission format on one side, and the hidden verifier, hard anti-gaming gates, and outcome metrics on the other, which is what allows observed differences to be attributed to a specific solver component. Overlaid on this is the HDS6 process verifier, which scores the recorded trajectory—blind to outcome and to solver identity—along six capabilities that spell TRACES (Tools, Repair, Alternatives, Coherence, Evidence, Scope), using per-task subrubrics plus a load-bearing-step analysis. Verification output doubles as a repair note returned to the solver, converting evaluation from a verdict into a loop. The paper treats the construction of faithful environments with trustworthy hidden verifiers as a contribution on its own, explicitly noting that building such an environment can be as hard as solving the task.
What would settle it
Run the four AAV capsid tasks scoring the original published methods' own released outputs or official implementations, rather than the paper's reproductions, against the same hidden verifiers; if apodex-1.1's per-task scores (0.904 viability, 0.635 tropism, 0.649 structure, 0.180 design) stop exceeding the published states of the art (0.878, 0.622, 0.605, 0.116), the headline 7% claim fails.
Extended reading notes
Core claim
The paper's central claim is that open-ended real-world problems can be converted into tractable, verifiable, repairable benchmark environments, and that evaluating the complete solver system rather than the model alone is what enables progress toward genuine discovery. The transformation has four load-bearing pieces: a problem manifest that fixes the question, success criterion, task decomposition, tools, budgets, and hidden verifier; a reality-based environment that returns fresh observations as the solver acts; outcome verification against hidden ground truth or measurable proxies; and blind process verification through HDS6, whose repair notes close a measure–diagnose–improve loop without leaking ground truth. The paper further claims this machinery works in practice: in adeno-associated virus capsid design, the framework exceeded the published state of the art at every one of the four pipeline stages, and in drug repurposing and reformulation, adding the task-specific environment improved mean normalized prediction scores of two GPT backbones by 2.5 and 7.6 points over the same closed-book backbone. It also claims that the blind HDS6 process score correlates positively with hidden outcome scores, with a pooled Spearman coefficient of 0.51 over 409 trajectories, supporting the use of process metrics when definitive ground truth is delayed.
Load-bearing premise
The whole framework assumes that each hidden verifier's measurable proxy—a held-out experimental fitness, a clinical-trial stage, a predicted selectivity rank—really tracks the true real-world objective, and that the reproduced baselines used as published state of the art faithfully represent the original methods; if either assumption gives way, a solver can score high without achieving a genuine discovery.
Editorial extensions
If this is right
- Benchmarks can be built around complete solver systems rather than isolated prompts, with final-answer scoring and process scoring reported side by side.
- Problems whose ground truth is delayed, incomplete, or non-existent become evaluable: outcome verifiers handle measurable proxies, HDS6 supplies an immediate process score, and the repair loop turns diagnosis into a measurable improvement.
- Across the four AAV capsid tasks, a domain-specific environment plus a capable solver surpasses the previous per-task published state of the art, from viability prediction to generative design.
- The controlled episode interface makes component attribution possible, as shown by ablations where skill guidance and terminal verifier guidance improve outcome scores on LLM-engineering tasks.
- A blind process score that correlates positively with hidden outcomes can act as a provisional quality signal while definitive outcomes are still pending.
Reading between the lines
- The framework's registry of 423 problems suggests that almost any domain with a monitorable success criterion could be turned into an adversarial evaluation environment, making benchmark construction more of an industrial process than a hand-crafted exercise.
- Because the HDS6-outcome correlation is positive but moderate, process scores are most defensible as early indicators for curation and repair, not as replacements for outcome verification when high-stakes decisions ride on the result.
- A natural next experiment is to ablate only the biomedical evidence tools in drug repurposing while holding the episode interface fixed; the paper's attribution logic predicts the reported gains should be traceable to specific tool categories.
- If hidden verifiers based on proxies become standard, the field will need periodic calibration checks that compare proxy scores against long-delayed true outcomes, otherwise solvers will optimize to the proxy instead of the real objective.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Apodex Discovery, a framework for building and evaluating 'discoverative AI' through heavy-duty solvers, together with TRACES, a benchmark of executable environments, tasks, and episodes with hidden verifiers. It describes a problem-scouting process that produced a 423-problem registry, an environment-task-episode abstraction with isolation and anti-gaming gates, a blind process metric (HDS6) scoring Tools, Repair, Alternatives, Coherence, Evidence, and Scope, and verification-driven repair loops. The empirical sections report that the Apodex solver surpassed published state of the art on four AAV capsid design tasks, that a task-specific biomedical environment improved normalized prediction scores of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points on drug repurposing, and that controlled ablations attribute performance differences to solver components. The authors release TRACES with 17 environments and 218 episodes and position the work as moving AI evaluation beyond predefined benchmarks toward verifiable investigations.
Significance. If the framework and results hold, this is a substantial step toward evaluating AI on open-ended, consequential problems. The release of executable environments with trajectory recording, hidden verifiers, anti-gaming gates, and a blind process metric is valuable infrastructure, and the internal validation of HDS6 against outcome scores (pooled Spearman 0.51 over 409 trajectories) is a meaningful first check. The controlled ablations and the verification-repair case studies are also useful contributions. However, the headline empirical claims currently depend on author-reproduced baselines, surrogate oracles, and a high-anchor normalization, so the significance of the specific quantitative results is conditional on additional validation.
major comments (5)
- [4.1.3 and Table 19] The claim that Apodex 'surpassed the published state of the art by 7%' is not defined precisely, and the generative-design comparison is against the authors' own reproductions, not published numbers. Table 19 states 'All three reference methods are our own reproductions under the identical episode budget,' yet AAVDiff is labeled 'published SOTA.' The 7% figure appears to depend on including Task 4, where the baseline is an in-house reimplementation scored by the same surrogate oracle. Please state the exact formula for the 7% aggregate, report the original published scores alongside the reproductions, and provide code and hyperparameters for all baseline reproductions so the comparison is verifiable.
- [4.1.5, Task 4 (Generative design)] The design-task oracle is a trained classifier (receptor instances) or a predicted-rank quantile (tissue instances), and the paper provides no validation that these surrogates track true experimental fitness in the out-of-distribution regime of multi-substitution novel sequences. Since the benchmark intentionally scores sequences far from the seed distribution, a solver can score well by exploiting surrogate blind spots without producing genuinely on-target capsids. Please add an oracle-calibration analysis against held-out experimental measurements, report oracle error as a function of mutation distance from the training set, or explicitly reframe the design result as an in-silico benchmark rather than evidence of genuine discovery.
- [4.2.2, Eq. (4), and Table 7] The random anchor S_R=0.86925 is very close to the oracle anchor S_O=1, so the normalized score in Eq. (4) compresses raw score differences by a factor of about 7.65 (1/0.13075). The reported gains of +2.52 and +7.60 normalized points correspond to raw gains of +0.00330 and +0.00993, and with three runs the standard deviations overlap substantially (GPT-5.5: 53.78±3.22 vs 56.30±1.56; GPT-5.6-sol: 54.21±2.20 vs 61.81±3.92). The abstract and conclusions present these as clean gains; please report raw scores with confidence intervals, add significance tests or explicit non-significance statements, and describe the results as preliminary.
- [3.1.2 and Table 3] HDS6 is introduced as a measure of process quality, but its validation currently consists of correlation with the authors' own outcome verifiers (pooled Spearman 0.51). No inter-annotator agreement is reported for the judge/reviewer/arbiter procedure, and no comparison with human expert process ratings is provided. Since HDS6 is a central contribution, please report reliability statistics (e.g., judge-reviewer agreement, human-HDS6 agreement) or temper the claim that HDS6 'evaluates' process quality rather than reflecting a proposed operationalization.
- [4.4 and Table 13] The verification-driven repair section reports a mean outcome improvement of +0.155 over 434 deficient trajectories, but the design has no matched control group: the same deficient trajectories are not re-run without the repair note. The text acknowledges this limitation, yet the section title and the framing of the aggregate result present the improvement as an effect of repair. Please add a matched control (identical trajectories re-run without the repair note) or present the table strictly as a pilot demonstration, with regression-to-mean and run-to-run variation discussed as alternative explanations.
minor comments (5)
- [Abstract] The abstract omits the caveats stated in Section 4.2.3 ('these gains should be interpreted descriptively') and in Table 19 (baselines are author reproductions); please add brief qualifiers so the claims match the body's caution.
- [Table 19] The label 'published SOTA' on the AAVDiff row conflicts with the table note that all three reference methods are reproductions; use 'reproduced baseline' throughout or clearly separate published numbers from reimplementations.
- [4.1.3, last paragraph] The conclusion states that Apodex 'exceeded prior human baselines,' but Table 5 shows that apodex-1.1 is not the strongest solver in the authors' own set (e.g., kimi-k3 scores higher on all four tasks); please specify that the comparison is to published methods, not human experts or the strongest frontier models.
- [3.1.2, Evidence definition] There is a typo: 'disocoverative AI' should be 'discoverative AI'; the same section also uses 'Table 2' and 'Table 3' before their first mention, which is fine but should be cross-checked.
- [4.4.1] The 'virtual clinician' is introduced without specification of its interface or whether it is a simulated tool or a human in the loop; please clarify its role in the environment and whether it is available to all solvers.
Circularity Check
One disclosed, non-central procedure-design result reduces to retrieval from the reference source; the main AAV, drug-repurposing, and HDS6 claims are self-contained empirical comparisons rather than circular derivations.
-
fitted input called prediction
[Section 4.2.5 (Repurposing and Reformulation Procedure Design)]
"Because the retrieved eligibility record is also the source from which the population reference is derived, the population comparison measures evidence retrieval and synthesis rather than fully prospective procedure generation."
In the LM + context condition, the model receives the structured eligibility context retrieved from registered clinical-trial records, and the reference population value is derived from those same records. The reported +0.89 population gain is therefore largely an artifact of the answer being present in the input, not an independent prospective prediction. The paper explicitly disclaims the comparison, so this is a disclosed, non-load-bearing limitation rather than a hidden circularity. It affects only the secondary procedure-design evaluation; the main outcome-prediction results (Table 7) and AAV comparisons use held-out labels or experimental proxies and do not share this construction.
full rationale
The central derivation chain is not circular. The AAV tasks are scored against held-out experimental data (Bryant et al. deep mutagenesis, Ogden et al., Fit4Function, LY6A/LY6C1 binding) or against post-cutoff deposited structures; the generative-design oracle is a surrogate trained on real fitness data, so a high score requires generalization to novel sequences rather than reproduction of the training labels. The drug-repurposing outcome metric uses held-out clinical and regulatory outcomes with temporal filtering and decontamination, and the normalization anchor SR is fixed from development labels and applied identically to both conditions, so the reported gains are not artifacts of the anchor. HDS6 is validated by correlating a blind process judge against a separate hidden outcome verifier; both channels are author-built, which is a validity limitation, but the correlation is an empirical measurement rather than a definitional equivalence. The only identifiable by-construction step is the disclosed procedure-design population comparison, where the context contains the reference source. No load-bearing self-citation, imported uniqueness theorem, or ansatz-smuggling chain is present; the AAV SOTA baselines are reproductions of published methods, which is a reproducibility concern only if the reproductions are unfaithful, not a circularity.
Assumptions & free parameters
free parameters (4)
- DRR tier scale values =
0.00, 0.15, 0.45, 0.65, 1.00
- Random-anchor score S_R =
0.86925
- AAV baseline reproductions =
AAVGen 0.109, AAVDiff 0.116, ALICE 0.110
- HDS6 subrubric weights =
e.g., 2:2:2:1 for Evidence subrubrics
assumptions (4)
- domain assumption Measurable proxies (e.g., packaging fitness, tissue enrichment, surrogate classifiers) are valid stand-ins for the true real-world objective (therapeutic benefit).
- domain assumption The hidden ground truth and scoring rubrics are not reachable by the solver despite tool access and environment isolation.
- domain assumption LLM-based process judges provide valid, bias-free scores after stripping chain-of-thought and identity, with re-grounding against the log.
- domain assumption The clinical outcome tier scale and its historical distribution are representative of future drug repurposing outcomes.
invented entities (1)
-
Virtual clinician
Cite this review
Pith. "Pith review of Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence." pith.science (2026). https://pith.science/paper/OEVTJQBT
@misc{pith2026260811341,
author = {Pith},
title = {Pith review of: Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence},
year = {2026},
howpublished = {\url{https://pith.science/paper/OEVTJQBT}},
note = {Machine review of arXiv:2608.11341}
}
read the original abstract
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form. We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success. In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.
Reference graph
Works this paper leans on
-
[1]
International Conference on Learning Representations (ICLR) , year =
Measuring Massive Multitask Language Understanding , author =. International Conference on Learning Representations (ICLR) , year =. 2009.03300 , archivePrefix =
arXiv 2009
-
[2]
, booktitle =
Rein, David and Hou, Betty Li and Stickland, Asa Cooper and Petty, Jackson and Pang, Richard Yuanzhe and Dirani, Julien and Michael, Julian and Bowman, Samuel R. , booktitle =. 2024 , eprint =
2024
-
[3]
Measuring Mathematical Problem Solving With the
Hendrycks, Dan and Burns, Collin and Kadavath, Saurav and Arora, Akul and Basart, Steven and Tang, Eric and Song, Dawn and Steinhardt, Jacob , booktitle =. Measuring Mathematical Problem Solving With the. 2021 , eprint =
2021
-
[4]
Glazer, Elliot and Erdil, Ege and Besiroglu, Tamay and Chicharro, Diego and Chen, Evan and Gunning, Alex and Falkman Olsson, Caroline and Denain, Jean-Stanislas and Ho, Anson and Santos, Emily de Oliveira and others , year =. 2411.04872 , archivePrefix =
-
[5]
Humanity's Last Exam , author =. 2025 , howpublished =. 2501.14249 , archivePrefix =
arXiv 2025
-
[6]
arXiv preprint arXiv:2107.03374 , year =
Evaluating Large Language Models Trained on Code , author =. arXiv preprint arXiv:2107.03374 , year =
-
[7]
and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =
Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , booktitle =. 2024 , eprint =
2024
-
[8]
2024 , howpublished =
Introducing. 2024 , howpublished =
2024
Show all 98 references
-
[9]
2025 , howpublished =
2025
-
[10]
Zhou, Shuyan and Xu, Frank F. and Zhu, Hao and Zhou, Xuhui and Lo, Robert and Sridhar, Abishek and Cheng, Xianyi and Ou, Tianyue and Bisk, Yonatan and Fried, Daniel and Alon, Uri and Neubig, Graham , booktitle =. 2024 , eprint =
2024
-
[11]
2024 , eprint =
Xie, Tianbao and Zhang, Danyang and Chen, Jixuan and Li, Xiaochuan and Zhao, Siheng and Cao, Ruisheng and Hua, Toh Jing and Cheng, Zhoujun and Shin, Dongchan and Lei, Fangyu and others , booktitle =. 2024 , eprint =
2024
-
[12]
Yao, Shunyu and Shinn, Noah and Razavi, Pedram and Narasimhan, Karthik , journal =
-
[13]
arXiv preprint arXiv:2110.14168 , year =
Training Verifiers to Solve Math Word Problems , author =. arXiv preprint arXiv:2110.14168 , year =
-
[14]
Transactions on Machine Learning Research (TMLR) , year =
Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models , author =. Transactions on Machine Learning Research (TMLR) , year =
-
[15]
Think You Have Solved Question Answering? Try
Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind , journal =. Think You Have Solved Question Answering? Try
-
[16]
2024 , eprint =
Liu, Xiao and Yu, Hao and Zhang, Hanchen and others , booktitle =. 2024 , eprint =
2024
-
[17]
International Conference on Learning Representations (ICLR) , year =
Mialon, Gr. International Conference on Learning Representations (ICLR) , year =. 2311.12983 , archivePrefix =
-
[18]
2024 , eprint =
Koh, Jing Yu and Lo, Robert and Jang, Lawrence and Duvvur, Vikram and Lim, Ming Chong and Huang, Po-Yu and Neubig, Graham and Zhou, Shuyan and Salakhutdinov, Ruslan and Fried, Daniel , booktitle =. 2024 , eprint =
2024
-
[19]
2024 , eprint =
Trivedi, Harsh and Khot, Tushar and Hartmann, Mareike and Manku, Ruskin and Dong, Vinty and Li, Edward and Gupta, Shashank and Sabharwal, Ashish and Balasubramanian, Niranjan , booktitle =. 2024 , eprint =
2024
-
[20]
and others , journal =
Zhang, Andy K. and others , journal =
-
[21]
Chan, Jun Shern and others , journal =
-
[22]
Wijk, Hjalmar and others , journal =
-
[23]
Starace, Giulio and others , journal =
-
[24]
Chen, Ziru and others , journal =
-
[25]
Majumder, Bodhisattwa Prasad and others , journal =
-
[26]
and others , journal =
Laurent, Jon M. and others , journal =
-
[27]
Training Software Engineering Agents and Verifiers with
Pan, Jiayi and others , journal =. Training Software Engineering Agents and Verifiers with
-
[28]
2023 , eprint =
Yao, Shunyu and Zhao, Jeffrey and Yu, Dian and Du, Nan and Shafran, Izhak and Narasimhan, Karthik and Cao, Yuan , booktitle =. 2023 , eprint =
2023
-
[29]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Tree of Thoughts: Deliberate Problem Solving with Large Language Models , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =. 2305.10601 , archivePrefix =
-
[30]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Reflexion: Language Agents with Verbal Reinforcement Learning , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =. 2303.11366 , archivePrefix =
-
[31]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Self-Refine: Iterative Refinement with Self-Feedback , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =. 2303.17651 , archivePrefix =
-
[32]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Toolformer: Language Models Can Teach Themselves to Use Tools , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =. 2302.04761 , archivePrefix =
-
[33]
and Burger, Doug and Wang, Chi , journal =
Wu, Qingyun and Bansal, Gagan and Zhang, Jieyu and Wu, Yiran and Li, Beibin and Zhu, Erkang and Jiang, Li and Zhang, Xiaoyun and Zhang, Shaokun and Liu, Jiale and Awadallah, Ahmed Hassan and White, Ryen W. and Burger, Doug and Wang, Chi , journal =
-
[34]
arXiv preprint arXiv:2305.16291 , year =
Voyager: An Open-Ended Embodied Agent with Large Language Models , author =. arXiv preprint arXiv:2305.16291 , year =
-
[35]
and Wettig, Alexander and Lieret, Kilian and Yao, Shunyu and Narasimhan, Karthik and Press, Ofir , booktitle =
Yang, John and Jimenez, Carlos E. and Wettig, Alexander and Lieret, Kilian and Yao, Shunyu and Narasimhan, Karthik and Press, Ofir , booktitle =. 2024 , eprint =
2024
-
[36]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , booktitle =. Judging. 2023 , eprint =
2023
-
[37]
2023 , eprint =
Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang , booktitle =. 2023 , eprint =
2023
-
[38]
Zhu, Lianghui and Wang, Xinggang and Wang, Xinlong , journal =
-
[39]
Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models , author =. Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =. 2405.01535 , archivePrefix =
-
[40]
2024 , eprint =
Ye, Seonghyeon and Kim, Doyoung and Kim, Sungdong and Hwang, Hyeonbin and Kim, Seungone and Jo, Yongrae and Thorne, James and Kim, Juho and Seo, Minjoon , booktitle =. 2024 , eprint =
2024
-
[41]
and others , journal =
Arora, Rahul K. and others , journal =
-
[42]
arXiv preprint arXiv:2211.14275 , year =
Solving Math Word Problems with Process- and Outcome-based Feedback , author =. arXiv preprint arXiv:2211.14275 , year =
-
[43]
International Conference on Learning Representations (ICLR) , year =
Let's Verify Step by Step , author =. International Conference on Learning Representations (ICLR) , year =. 2305.20050 , archivePrefix =
-
[44]
2024 , eprint =
Wang, Peiyi and Li, Lei and Shao, Zhihong and Xu, Runxin and Dai, Damai and Li, Yifei and Chen, Deli and Wu, Yu and Sui, Zhifang , booktitle =. 2024 , eprint =
2024
-
[45]
arXiv preprint arXiv:2408.15240 , year =
Generative Verifiers: Reward Modeling as Next-Token Prediction , author =. arXiv preprint arXiv:2408.15240 , year =
-
[46]
arXiv preprint arXiv:2501.12948 , year =
-
[47]
Lambert, Nathan and others , journal =
-
[48]
Concrete Problems in
Amodei, Dario and Olah, Chris and Steinhardt, Jacob and Christiano, Paul and Schulman, John and Man. Concrete Problems in. arXiv preprint arXiv:1606.06565 , year =
-
[49]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Defining and Characterizing Reward Hacking , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =. 2209.13085 , archivePrefix =
-
[50]
Findings of the Association for Computational Linguistics: EMNLP , year =
Sainz, Oscar and Campos, Jon Ander and Garc. Findings of the Association for Computational Linguistics: EMNLP , year =. 2310.18018 , archivePrefix =
-
[51]
Time Travel in
Golchin, Shahriar and Surdeanu, Mihai , booktitle =. Time Travel in. 2024 , eprint =
2024
-
[52]
International Conference on Learning Representations (ICLR) , year =
Shridhar, Mohit and Yuan, Xingdi and C. International Conference on Learning Representations (ICLR) , year =. 2010.03768 , archivePrefix =
2010 arXiv
-
[53]
2022 , eprint =
Yao, Shunyu and Chen, Howard and Yang, John and Narasimhan, Karthik , booktitle =. 2022 , eprint =
2022
-
[54]
2023 , eprint =
Deng, Xiang and Gu, Yu and Zheng, Boyuan and Chen, Shijie and Stevens, Samuel and Wang, Boshi and Sun, Huan and Su, Yu , booktitle =. 2023 , eprint =
2023
-
[55]
2023 , eprint =
Yang, John and Prabhakar, Akshara and Narasimhan, Karthik and Yao, Shunyu , booktitle =. 2023 , eprint =
2023
-
[56]
2024 , eprint =
Qin, Yujia and Liang, Shihao and Ye, Yining and Zhu, Kunlun and Yan, Lan and Lu, Yaxi and Lin, Yankai and Cong, Xin and Tang, Xiangru and Qian, Bill and others , booktitle =. 2024 , eprint =
2024
-
[57]
2024 , eprint =
Huang, Qian and Vora, Jian and Liang, Percy and Leskovec, Jure , booktitle =. 2024 , eprint =
2024
-
[58]
and Bashir, Ali and Sinai, Sam and Jain, Nina K
Bryant, Drew H. and Bashir, Ali and Sinai, Sam and Jain, Nina K. and Ogden, Pierce J. and Riley, Patrick F. and Church, George M. and Colwell, Lucy J. and Kelsic, Eric D. , journal =. Deep Diversification of an
-
[59]
Lu, Chris and Lu, Cong and Lange, Robert Tjarko and Foerster, Jakob and Clune, Jeff and Ha, David , journal =. The
-
[60]
Science (New York, N.Y.) , year=
Comprehensive AAV capsid fitness landscape reveals a viral gene and enables machine-guided design , author=. Science (New York, N.Y.) , year=
-
[61]
Nature Communications , volume=
Systematic multi-trait AAV capsid engineering for efficient gene delivery , author=. Nature Communications , volume=. 2024 , publisher=
2024
-
[62]
PLOS Biology , year=
Targeting AAV vectors to the central nervous system by engineering capsid–receptor interactions that enable crossing of the blood–brain barrier , author=. PLOS Biology , year=
-
[63]
Nucleic acids research , year=
The Protein Data Bank , author=. Nucleic acids research , year=
-
[64]
Human Gene Therapy , year=
Prediction of Adeno-Associated Virus Fitness with a Protein Language-Based Machine Learning Model , author=. Human Gene Therapy , year=
-
[65]
PLoS ONE , year=
DockQ: A Quality Measure for Protein-Protein Docking Models , author=. PLoS ONE , year=
-
[66]
ArXiv , year=
Structured Denoising Diffusion Models in Discrete State-Spaces , author=. ArXiv , year=
-
[67]
bioRxiv , year=
Mapping AAV capsid sequences to functions through function-guided in silico evolution , author=. bioRxiv , year=
-
[68]
ArXiv , year=
AAVGen: Precision Engineering of Adeno-associated Viral Capsids for Renal Selective Targeting , author=. ArXiv , year=
-
[69]
Nature , volume=
Accurate structure prediction of biomolecular interactions with AlphaFold 3 , author=. Nature , volume=. 2024 , publisher=
2024
-
[70]
Nature , year=
Highly accurate protein structure prediction with AlphaFold , author=. Nature , year=
-
[71]
Cell Reports Medicine , year=
Complete neutralizing antibody evasion by serodivergent non-mammalian AAVs enables gene therapy redosing , author=. Cell Reports Medicine , year=
-
[72]
arXiv preprint arXiv:2602.00359 , year=
Position: Agentic Evolution is the Path to Evolving LLMs , author=. arXiv preprint arXiv:2602.00359 , year=
-
[73]
International Conference on Learning Representations (ICLR) , volume=
Openhands: An open platform for ai software developers as generalist agents , author=. International Conference on Learning Representations (ICLR) , volume=
-
[74]
arXiv preprint arXiv:2404.10573 , year=
AAVDiff: Experimental Validation of Enhanced Viability and Diversity in Recombinant Adeno-Associated Virus (AAV) Capsids through Diffusion Generation , author=. arXiv preprint arXiv:2404.10573 , year=
-
[75]
arXiv preprint arXiv:2510.04374 , year =
Patwardhan, Tejal and Dias, Rachel and Proehl, Elizabeth and Kim, Grace and Wang, Michele and Watkins, Olivia and Fishman, Sim. arXiv preprint arXiv:2510.04374 , year =
-
[76]
Measuring
Kwa, Thomas and West, Ben and Becker, Joel and Deng, Amy and Garcia, Katharyn and Hasin, Max and Jawhar, Sami and Kinniment, Megan and others , booktitle =. Measuring
-
[77]
arXiv preprint arXiv:2407.00215 , year =
McAleese, Nat and Pokorny, Rai Michael and Cer. arXiv preprint arXiv:2407.00215 , year =
-
[78]
Zheng, Chujie and Zhang, Zhenru and Zhang, Beichen and Lin, Runji and Lu, Keming and Yu, Bowen and Zhou, Jingren and Lin, Junyang , booktitle =
-
[79]
International Conference on Learning Representations (ICLR) , year =
Large Language Models Cannot Self-Correct Reasoning Yet , author =. International Conference on Learning Representations (ICLR) , year =
-
[80]
Nature , volume =
Autonomous chemical research with large language models , author =. Nature , volume =
-
[81]
Towards an
Gottweis, Juraj and Weng, Wei-Hung and Daryin, Alexander and Tu, Tao and Palepu, Anil and Sirkovic, Petar and others , journal =. Towards an
-
[82]
Nature Machine Intelligence , volume =
Augmenting large language models with chemistry tools , author =. Nature Machine Intelligence , volume =
-
[83]
Empowering biomedical discovery with
Gao, Shanghua and Fang, Ada and Huang, Yepeng and Giunchiglia, Valentina and Noori, Ayush and Schwarz, Jonathan Richard and Ektefaie, Yasha and Kondic, Jovana and Zitnik, Marinka , journal =. Empowering biomedical discovery with
-
[84]
Gou, Zhibin and Shao, Zhihong and Gong, Yeyun and Shen, Yelong and Yang, Yujiu and Duan, Nan and Chen, Weizhu , booktitle =
-
[85]
Miserendino, Samuel and Wang, Michele and Patwardhan, Tejal and Heidecke, Johannes , journal =
-
[86]
arXiv preprint arXiv:2606.05405 , year =
Agents' Last Exam , author =. arXiv preprint arXiv:2606.05405 , year =
-
[87]
2026 , howpublished =
Discovery Loop , author =. 2026 , howpublished =
2026
-
[88]
Nature , volume =
Scientific discovery in the age of artificial intelligence , author =. Nature , volume =
-
[89]
Novikov, Alexander and others , journal =
-
[90]
Nature , volume =
Scaling deep learning for materials discovery , author =. Nature , volume =
-
[91]
Nature , volume =
An autonomous laboratory for the accelerated synthesis of novel materials , author =. Nature , volume =
-
[92]
Brockman, Greg and Cheung, Vicki and Pettersson, Ludwig and Schneider, Jonas and Schulman, John and Tang, Jie and Zaremba, Wojciech , journal =
-
[93]
Journal of Artificial Intelligence Research , volume =
The Arcade Learning Environment: An Evaluation Platform for General Agents , author =. Journal of Artificial Intelligence Research , volume =
-
[94]
Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =
Wang, Ruoyao and Jansen, Peter and C. Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =
-
[95]
Nature biotechnology , volume=
Learning protein fitness models from evolutionary and assay-labeled data , author=. Nature biotechnology , volume=. 2022 , publisher=
2022
-
[96]
Cell Press Blue , volume=
Mapping AAV capsid sequences to functions through function-guided in silico evolution , author=. Cell Press Blue , volume=. 2026 , publisher=
2026
-
[97]
BioRxiv , pages=
Boltz-1 democratizing biomolecular interaction modeling , author=. BioRxiv , pages=
-
[98]
Computers in biology and medicine , volume=
ComDock: A novel approach for protein-protein docking with an efficient fusing strategy , author=. Computers in biology and medicine , volume=. 2023 , publisher=
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.