REVIEW 4 major objections 4 minor 197 references
Training AI Scientists to Replicate Research
T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper reports that Faraday, a 27B agent post-trained on Replica, a space of 310 figure-replication tasks, scores above Claude Opus 4.8 and GPT-5.5 on most tasks according to an auto-generated rubric judge, and that this advantage…
desk verdict Solid, honest, partly circular: the same rubric judge supplies reward and headline metric, human validation is weak (p=0.109), but the task space and CAT training recipe are real contributions worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the per-task rubric judge: an auto-generated scoring guide covering five dimensions, executed by a coding agent that inspects the rollout's container, code, git history, and the gold figure, and sampled three times to reduce noise. This judge supplies the GRPO reward, so its validity determines everything downstream. Around it sits the coding-agent-as-tool (CAT) setup, in which the 27B Faraday model plans experiments and delegates implementation to a frontier coding agent through a shell tool, plus a training recipe with turn-level credit assignment weights from the judge, normalized per token, that stabilizes long-horizon, non-verifiable reinforcement learning.
What would settle it
Take a random sample of Replica test tasks (not selected for large judge margins), have expert humans blindly rank rollouts from Faraday, Claude, and Codex, and compare human rankings with the rubric judge's scores; if human-judge agreement on the full sample is no better than the baseline judge's agreement, or if humans prefer the frontier baselines on average, then the rubric judge is not capturing human replication taste and Faraday's headline advantage is an artifact of the reward loop.
Extended reading notes
Core claim
The paper's central claim is that long-horizon reinforcement learning against a rubric-based judge turns a 27B base language model into an agent that replicates research figures more faithfully than frontier coding agents. Replica tasks are constructed by redacting one results plot from a paper and requiring the agent to produce the plot by running real experiments within an hour on a single GPU slice, scaling down when necessary. The judge, a coding agent scored against a per-task rubric auto-generated by Claude Opus 4.7, grades five dimensions: visual fidelity, claim reproduction, implementation fidelity, experimental depth, and scientific integrity. Faraday, post-trained with a turn-level credit variant of GRPO using three judge samples per rollout, outperforms Claude Opus 4.8 and GPT-5.5 under the same budget; qualitative inspection of rollouts shows Faraday implementing the mechanism behind the claim rather than hard-coding outputs, and it transfers to held-out AI-for-science papers and to variants with changed claims or datasets (19 of 20 'imagined' tasks). The paper is careful to note that the headline comparison is made by the rubric judge itself, with only partial human validation.
Load-bearing premise
The entire result rests on the rubric judge measuring true replication quality: the paper's own human-agreement check is weak (participants side with the judge on 63% of disputed pairs, a result not statistically significant), so if the judge rewards judge-pleasing behavior instead of faithful reproduction, Faraday's edge could vanish.
Editorial extensions
If this is right
- If the rubric judge tracks human taste, replication skill can be trained at scale without hand-verified rewards; the paper estimates that top ML conferences alone could yield about 36,000 new tasks per year.
- A small open-weights model can direct a frontier coding agent and beat the frontier agent running alone, so post-training a 'researcher layer' is a viable alternative to engineering specialized harnesses.
- The learned behavior transfers to papers outside the training distribution (AI-for-science) and to altered claims or datasets, indicating generalization rather than memorization.
- Swapping the inner coding agent for a stronger one at evaluation time improves performance, so the trained scientist layer can ride improvements in frontier coding models without retraining.
- The training recipe with three judge samples and turn-level credit stabilizes GRPO in a long-horizon, non-verifiable setting; without it, the paper's ablation shows training collapsing after about 50 steps.
Reading between the lines
- A skeptical reading: because the same rubric judge supplies the reward and the headline evaluation, some of Faraday's edge may be judge-pleasing behavior; a decisive test would rerun the main comparison on a randomly selected sample of tasks with human expert ranking as ground truth, rather than only on the strong-margin subset.
- If the judge's human-agreement evidence is as weak as it appears (participants side with the rubric judge on 63% of disputed pairs, p = 0.109), averaging three judge samples could inflate Faraday's apparent advantage by rewarding consistency that humans do not value.
- A natural extension is to scale Replica toward full-paper replication and multi-figure consistency, where the judge would have to check cross-figure coherence; if Faraday's skill transfers there, the paper's thesis that replication is a curriculum step toward innovation would gain support.
- The CAT architecture suggests a cost-benefit question the paper leaves open: how to split size and capability between an outer judgment model and an inner execution model, and whether the outer model can remain small, open, and inspectable as coding tools improve.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Replica, a task space of 310 figure-replication tasks derived from 100 ML and AI-for-science papers, where each task asks an agent to reproduce a redacted results figure from a paper within a one-hour budget. To supply training signal, the authors build an auto-generated, per-task rubric-based judge implemented with Codex GPT-5.5, which scores rollouts along five dimensions and produces turn-level credit weights. They post-train Faraday, a 27B-parameter model, with a modified GRPO recipe using this judge as reward, and report that Faraday outperforms Claude Opus 4.8 and GPT-5.5 on 73% of in-distribution tasks and 60% of held-out AI-for-science tasks according to the same rubric judge, with average gains of 6% and 8% respectively. The paper also includes human studies aimed at validating the judge, a prompt-optimization control, qualitative case analyses, and generalization experiments to larger compute budgets and stronger coding tools.
Significance. If the central claim holds, the paper would be a meaningful step toward training small 'AI scientist' agents that can direct frontier coding agents on long-horizon, underspecified scientific tasks. The strengths of the paper are real: the Replica task space is scalable and automatically generated, the training recipe is described in unusual detail with stage-by-stage hyperparameters, the ablations (turn-level credit assignment, coding-agent tool, judge noise) are informative, and the authors include machine-checkable experiments and a prompt-optimized baseline that rules out one class of judge-pleasing. However, the significance is contingent on whether the rubric judge measures human-valued replication quality. The current evidence for that validity is thin, and because the same judge is used for both the training reward and the headline evaluation metric, the reported advantage is vulnerable to learned judge-pleasing that would not transfer to human judgment.
major comments (4)
- [§3.2, §3.5, §4.3] The rubric judge is both the GRPO training reward (Section 3.2 and 3.5) and the sole evaluation metric for the headline comparison (Section 4.3, Figure 2). Consequently, any policy-level reward hacking or judge-pleasing behavior learned during RL is inherited directly by the test evaluation, because the same judge scores the held-out rollouts. The prompt-optimization control in Figure 5 (left) rules out prompt-level judge-pleasing, but it does not rule out policy-level learned behaviors, such as producing certain procedural artifacts or stylistic traces that the judge rewards. The paper's central claim of superiority over Claude and Codex therefore rests entirely on the untested assumption that the judge is not exploiting these artifacts. To support the claim, the authors need either an independent evaluation metric (e.g., a different judge model or a non-LLM objective measure) on the test split, or a human-validated judge on the test split itself.
- [§4.1, Appendix F.1] The human-validation evidence for the rubric judge is not statistically significant and is restricted to the training split. Participants side with the rubric judge over the baseline judge on only 63% of disputed pairs, with p = 0.109 from a binomial mixed-effects model, and the rubric-human Kendall tau is 0.19 versus 0.15 for the baseline judge, a difference of 0.04 with no reported confidence interval. The ten tasks used in this study are all drawn from the train split and were selected specifically for judge disagreement, so the study does not establish judge validity on the AI-for-science test split where the out-of-distribution claim (60% task superiority) is made. The abstract's statement that the judge 'agrees with human assessment of replication quality' is stronger than this evidence supports.
- [§4.4, Appendix F.2] The agent-comparison human study is conditional on the rubric judge's high-margin selection: only rollouts where Faraday scores at least 0.2 above both Claude and Codex on the judge's scale are shown to humans. The paper explicitly and correctly notes that this design cannot support conclusions about average human preference. As external grounding for the judge, this study is therefore a sanity check that the judge's most confident calls align with humans, but it does not validate the judge for the typical cases that drive the aggregate results in Figure 2. The 71% preference rate among the 41 selected rollouts is compatible with a judge that is only directionally correct on extreme margins while being wrong or noisy on the majority of comparisons.
- [§3.2, Appendix D.1] The discussion of pre-training contamination in Appendix D.1 addresses the agent's exposure to the papers but does not address the judge's exposure. The judge is a Codex model that has likely seen many of the Replica papers, including their figures, in its pre-training data. Because the judge is given the gold plot and the redacted paper, it could reward rollouts that happen to match its memorized representation of the original figure, rather than rollouts that faithfully replicate the underlying experiment. Since the judge is also the training reward, this could bias Faraday toward mimicking memorized outputs rather than learning robust replication skill. The authors should at least analyze whether judge scores correlate with the judge model's (or the rubric generator's) familiarity with the paper, e.g., by publication year or by held-out citation counts, or by comparing judge scores on papers published after the judge's knowledge cutoff.
minor comments (4)
- [Figure 3 (left)] The reported Kendall tau values of 0.19 (rubric vs humans) and 0.15 (baseline vs humans) are shown as single numbers with no confidence intervals or significance tests; adding these would help readers assess the practical difference.
- [Table 1] The qualitative case analyses appear to be based on the authors' own inspection of rollouts; please state whether the inspection was performed blind to the model identity, or at least acknowledge the potential for confirmation bias in selecting and describing these examples.
- [Appendix G.1] The Faraday system prompt contains the literal template variable '{coding_agent_budget} tokens', which suggests an unresolved placeholder; this should be rendered as a concrete value in the published appendix.
- [Abstract and §4.3] The abstract claims the judge 'agrees with human assessment of replication quality' and Section 4.3 describes a 'comprehensive uplift' in performance; both statements should be tempered to reflect the non-significant human-agreement result and the conditional nature of the human preference study.
Circularity Check
Headline advantage is measured by the same rubric judge that supplied the training reward; human validation of that judge is non-significant (p=0.109) and conditional.
-
fitted input called prediction
[Section 3.2 (Reward function) and Figure 1 caption; Section 4.3 (Faraday replicates better than Claude and Codex), Figure 2; Appendix F.1]
"The judge provides an overall reward and per-turn credit assignment weights, which are used to train the Faraday agent using a modified version of GRPO. ... We compare Faraday against baselines across the entire Replica task distribution (Figure 2). ... Faraday outperforms both Claude and Codex on 73% of tasks. ... Participants side with the rubric judge on 63% of disputed pairs, higher than chance but not significantly so (p = 0.109)."
The rubric-based judge is both the GRPO training reward and the headline evaluation metric. Faraday's policy weights were post-trained to maximize this judge's score, so the reported in-distribution advantage over Claude and Codex is, on the train split, the agent beating the very function it was optimized against; the 'prediction' is the fitted objective. The held-out AI-for-science split is scored by the same judge that shaped Faraday's behavior, and the paper's own external grounding of that judge is not statistically significant (p=0.109). The average 6-8% gain therefore reduces substantially to 'the policy optimized against J scores higher on J', unless judge validity is independently established, which the paper's data do not do.
full rationale
The central quantitative claim is evaluatively self-referential: the same auto-generated rubric judge supplies the GRPO reward (Section 3.2, Figure 1) and the headline test metric (Section 4.3, Figure 2). On the training distribution, Faraday's advantage over non-optimized baselines is the expected consequence of maximizing that score; on the held-out AI-for-science split, the same judge is used and its human-validation evidence is weak. The paper does make genuine attempts at external checks — prompt-optimization control, a conditional human-preference study, qualitative rollout analysis, and original-paper-author feedback — so the circularity is partial, not total: the test-split generalization and qualitative differences are independent content. I found no load-bearing self-citation chain or imported uniqueness theorem; the issue is that the headline metric is the training objective, with no significant external anchor. Score 6 reflects one central 'prediction' that reduces by construction to the fitted reward.
Assumptions & free parameters
free parameters (4)
- GRPO/LoRA training hyperparameters =
LoRA rank 128, alpha 128, lr 6e-6, KL beta 3e-3, eps_low 0.15, eps_high 0.35
- Judge sample count and rollout count =
3 judge samples per rollout during training; 8 evaluation rollouts per task
- Rubric dimension weights =
Equal weights, 1/5 per dimension
- Task time and compute budget =
60 minutes; one-seventh H200 MIG slice
assumptions (4)
- domain assumption The five rubric dimensions and their equal weighting define replication quality.
- domain assumption A scaled-down version of an experiment can preserve the paper's core claim.
- domain assumption A 10-minute LLM judge inspecting code, git history, container, and gold plot can reliably score replication.
- domain assumption Pre-training exposure to the original papers and figures is not a validity threat.
invented entities (3)
-
Replica task space
-
Rubric-based judge
independent evidence
-
Faraday
Cite this review
Pith. "Pith review of Training AI Scientists to Replicate Research." pith.science (2026). https://pith.science/paper/GCIZLPZC
@misc{pith2026260813331,
author = {Pith},
title = {Pith review of: Training AI Scientists to Replicate Research},
year = {2026},
howpublished = {\url{https://pith.science/paper/GCIZLPZC}},
note = {Machine review of arXiv:2608.13331}
}
read the original abstract
The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research. In this work, we develop Replica, a scalable task space for paper replication. To provide reward signal, we introduce an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality. We post-train Faraday, a 27B-parameter "AI Scientist" agent that leverages coding agents as tools, surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Qualitative analysis of individual rollouts reveals that Faraday adopts a more scientifically-principled approach. We believe that our results provide a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2503.11926 , year=
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation , author=. arXiv preprint arXiv:2503.11926 , year=
-
[2]
2026 , eprint=
Measuring Reward-Seeking via Contrastive Belief Updates , author=. 2026 , eprint=
2026
-
[3]
2026 , howpublished=
Hugging Face model evaluation security incident , author=. 2026 , howpublished=
2026
-
[4]
The Annals of Statistics , volume=
Additive logistic regression: a statistical view of boosting , author=. The Annals of Statistics , volume=
-
[5]
How do language models learn facts?
Zucchet, Nicolas and Bornschein, J. How do language models learn facts?. arXiv preprint arXiv:2503.21676 , year=
-
[6]
Learning precise timing with
Gers, Felix A and Schraudolph, Nicol N and Schmidhuber, J. Learning precise timing with. Journal of Machine Learning Research , volume=
-
[7]
arXiv preprint arXiv:1412.6980 , year=
Adam: A method for stochastic optimization , author=. arXiv preprint arXiv:1412.6980 , year=
-
[8]
2006 , howpublished=
Artificial scientists , author=. 2006 , howpublished=
2006
Show all 197 references
-
[9]
3d Conference on Artificial General Intelligence (AGI-2010) , pages=
Artificial scientists & artists based on the formal theory of creativity , author=. 3d Conference on Artificial General Intelligence (AGI-2010) , pages=. 2010 , organization=
2010
-
[10]
Stanford Humanities Review , volume=
Creativity and unpredictability , author=. Stanford Humanities Review , volume=
-
[11]
ECAI 2012 , pages=
Computational Creativity: The Final Frontier? , author=. ECAI 2012 , pages=. 2012 , publisher=
2012
-
[12]
2011 , publisher =
David Deutsch , title =. 2011 , publisher =
2011
-
[13]
Leakage and the Reproducibility Crisis in
Kapoor, Sayash and Narayanan, Arvind , journal=. Leakage and the Reproducibility Crisis in. 2022 , url=
2022
-
[14]
AI Magazine , volume =
Reproducibility in Machine-Learning-Based Research: Overview, Barriers, and Drivers , author =. AI Magazine , volume =. 2025 , doi =
2025
-
[15]
arXiv preprint arXiv:2603.19461 , year=
Hyperagents , author=. arXiv preprint arXiv:2603.19461 , year=
-
[16]
2026 , howpublished=
2026
-
[17]
Incompressible knowledge probes: Estimating black-box
Li, Bojie , journal=. Incompressible knowledge probes: Estimating black-box. 2026 , url=
2026
-
[18]
International Conference on Learning Representations , year=
Learning to orchestrate agents in natural language with the conductor , author=. International Conference on Learning Representations , year=
-
[19]
2025 , url=
Su, Hongjin and Diao, Shizhe and Lu, Ximing and Liu, Mingjie and Xu, Jiacheng and Dong, Xin and Fu, Yonggan and Belcak, Peter and Ye, Hanrong and Yin, Hongxu and others , journal=. 2025 , url=
2025
-
[20]
Xie, Yutao and Thomas, Nathaniel and Hansen, Nick and Fu, Yang and Li, Li and Wang, Xiaolong , booktitle=
-
[21]
Physical Review Letters , volume=
Fast and accurate modeling of molecular atomization energies with machine learning , author=. Physical Review Letters , volume=. 2012 , publisher=
2012
-
[22]
arXiv preprint arXiv:2205.06175 , year=
A generalist agent , author=. arXiv preprint arXiv:2205.06175 , year=
-
[23]
2015 , publisher =
Why Greatness Cannot Be Planned: The Myth of the Objective , author =. 2015 , publisher =. doi:10.1007/978-3-319-15524-1 , isbn =
2015 doi
-
[24]
2018 , publisher=
Cognitive gadgets: The cultural evolution of thinking , author=. 2018 , publisher=
2018
-
[25]
Organization Science , volume=
More versus better: Artificial intelligence, incentives, and the emerging crisis in peer review , author=. Organization Science , volume=. 2026 , publisher=
2026
-
[26]
Advances in Neural Information Processing Systems , year=
Thinking fast and slow with deep learning and tree search , author=. Advances in Neural Information Processing Systems , year=
-
[27]
Mastering the game of
Silver, David and Huang, Aja and Maddison, Chris J and Guez, Arthur and Sifre, Laurent and Van Den Driessche, George and Schrittwieser, Julian and Antonoglou, Ioannis and Panneershelvam, Veda and Lanctot, Marc and others , journal=. Mastering the game of. 2016 , publisher=
2016
-
[28]
On scalable oversight with weak
Kenton, Zachary and Siegel, Noah Y and Kram. On scalable oversight with weak. Advances in Neural Information Processing Systems , year=
-
[29]
arXiv preprint arXiv:2211.03540 , year=
Measuring progress on scalable oversight for large language models , author=. arXiv preprint arXiv:2211.03540 , year=
-
[30]
Concrete problems in
Amodei, Dario and Olah, Chris and Steinhardt, Jacob and Christiano, Paul and Schulman, John and Man. Concrete problems in. arXiv preprint arXiv:1606.06565 , year=
-
[31]
Proceedings of the SIGCHI Conference on Human Factors in Computing Systems , year=
Picbreeder: evolving pictures collaboratively online , author=. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems , year=
-
[32]
Philosophical Transactions of the Royal Society B: Biological Sciences , volume=
Innovation in the collective brain , author=. Philosophical Transactions of the Royal Society B: Biological Sciences , volume=
-
[33]
Nature Communications , volume=
Learning few-shot imitation as cultural transmission , author=. Nature Communications , volume=. 2023 , publisher=
2023
-
[34]
2025 , url=
Shao, Rulin and Asai, Akari and Shen, Shannon Zejiang and Ivison, Hamish and Kishore, Varsha and Zhuo, Jingming and Zhao, Xinran and Park, Molly and Finlayson, Samuel G and Sontag, David and others , journal=. 2025 , url=
2025
-
[35]
2026 , url=
Li, Gaotang and Mishra, Bhavana Dalvi and Wang, Zifeng and Yan, Jun and Chen, Yanfei and Li, Chun-Liang and Le, Long T and Han, Rujun and Lee, George and Tong, Hanghang and others , journal=. 2026 , url=
2026
-
[36]
Peter Kirgis and Sayash Kapoor and Andrew Schwartz and Stephan Rabanser and David Africa and Konstantinos Voudouris and Viet Nguyen and Toby Pilditch and Magda Dubois and Harry Coppock and Cozmin Ududec and Nitya Nadgir and Matilda Orona and Tilman Bayer and Derrick Chan-Sew a...
2026
-
[37]
Lu, Chris and Lu, Cong and Lange, Robert Tjarko and Foerster, Jakob and Clune, Jeff and Ha, David , journal=. The. 2024 , url=
2024
-
[38]
2026 , publisher =
David Epstein , title =. 2026 , publisher =
2026
-
[39]
, title=
Sutton, Richard S. , title=. 2019 , month=
2019
-
[40]
arXiv preprint arXiv:2510.21614 , year=
Wang, Wenyi and Pi. arXiv preprint arXiv:2510.21614 , year=
-
[41]
Advances in Neural Information Processing Systems , year=
Checklists are better than reward models for aligning language models , author=. Advances in Neural Information Processing Systems , year=
-
[42]
arXiv preprint arXiv:2507.17746 , year=
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains , author=. arXiv preprint arXiv:2507.17746 , year=
-
[43]
Nature , pages=
A multi-agent system for automating scientific discovery , author=. Nature , pages=. 2026 , publisher=
2026
-
[44]
2025 , publisher=
Ghafarollahi, Alireza and Buehler, Markus J , journal=. 2025 , publisher=
2025
-
[45]
Towards an
Gottweis, Juraj and Weng, Wei-Hung and Daryin, Alexander and Tu, Tao and Palepu, Anil and Sirkovic, Petar and Myaskovsky, Artiom and Weissenberger, Felix and Rong, Keran and Tanno, Ryutaro and others , journal=. Towards an. 2025 , url=
2025
-
[46]
arXiv preprint arXiv:2107.12808 , year=
Open-ended learning leads to generally capable agents , author=. arXiv preprint arXiv:2107.12808 , year=
-
[47]
Rethinking Rubric Generation for Improving
Shen, William F and Qiu, Xinchi and Whitehouse, Chenxi and Alazraki, Lisa and Goel, Shashwat and Barbieri, Francesco and Willi, Timon and Mathur, Akhil and Leontiadis, Ilias , journal=. Rethinking Rubric Generation for Improving. 2026 , url=
2026
-
[48]
Ticking all the boxes: Generated checklists improve
Cook, Jonathan and Rockt. Ticking all the boxes: Generated checklists improve. arXiv preprint arXiv:2410.03608 , year=
-
[49]
Hong, Hanhua and Li, Yizhi and Chen, Jiaoyan and Huy, Luu Gia and Ananiadou, Sophia and Kim, Jung-jae and Lin, Chenghua , journal=. Can. 2026 , url=
2026
-
[50]
Training
Goel, Shashwat and Hazra, Rishi and Jayalath, Dulhan and Willi, Timon and Jain, Parag and Shen, William F and Leontiadis, Ilias and Barbieri, Francesco and Bachrach, Yoram and Geiping, Jonas and others , journal=. Training. 2025 , url=
2025
-
[51]
Agent laboratory: Using
Schmidgall, Samuel and Su, Yusheng and Wang, Ze and Sun, Ximeng and Wu, Jialian and Yu, Xiaodong and Liu, Jiang and Moor, Michael and Liu, Zicheng and Barsoum, Emad , journal=. Agent laboratory: Using. 2025 , publisher=
2025
-
[52]
Automating
Kirsch, Louis , school =. Automating. 2025 , month = jun, type =
2025
-
[53]
Paleobiology , volume=
Exaptation—a missing term in the science of form , author=. Paleobiology , volume=. 1982 , publisher=
1982
-
[54]
arXiv preprint arXiv:2603.17863 , year=
Procedural Generation of Algorithm Discovery Tasks in Machine Learning , author=. arXiv preprint arXiv:2603.17863 , year=
-
[55]
Advances in Neural Information Processing Systems , year=
Qiang, Rushi and Zhuang, Yuchen and Li, Yinghao and Sagar V K, Dingu and Zhang, Rongzhi and Li, ChangHao and Wong, Ian and Yang, Sherry and Liang, Percy and Zhang, Chao and Dai, Bo , title=. Advances in Neural Information Processing Systems , year=
-
[56]
2026 , url=
Karpathy, Andrej , title=. 2026 , url=
2026
-
[57]
arXiv preprint arXiv:2506.13131 , year=
Novikov, Alexander and V. arXiv preprint arXiv:2506.13131 , year=
-
[58]
2026 , url=
Alisia Lupidi and Bhavul Gauri and Thomas Simon Foster and Bassel Al Omari and Despoina Magka and Alberto Pepe and Alexis Audran-Reiss and Muna Aghamelu and Nicolas Baldwin and Lucia Cipolina-Kun and Jean-Christophe Gagnon-Audet and Chee Hau Leow and Sandra Lefdal and Hossam M...
2026
-
[59]
2024 , url=
Bodhisattwa Prasad Majumder and Harshit Surana and Dhruv Agarwal and Bhavana Dalvi Mishra and Abhijeetsingh Meena and Aryan Prakhar and Tirth Vora and Tushar Khot and Ashish Sabharwal and Peter Clark , journal=. 2024 , url=
2024
-
[60]
2406.06769 , archivePrefix=
Peter Jansen and Marc-Alexandre Côté and Tushar Khot and Erin Bransom and Bhavana Dalvi Mishra and Bodhisattwa Prasad Majumder and Oyvind Tafjord and Peter Clark , year=. 2406.06769 , archivePrefix=
-
[61]
2024 , url=
Qian Huang and Jian Vora and Percy Liang and Jure Leskovec , journal=. 2024 , url=
2024
-
[62]
2025 , url=
Jun Shern Chan and Neil Chowdhury and Oliver Jaffe and James Aung and Dane Sherburn and Evan Mays and Giulio Starace and Kevin Liu and Leon Maksin and Tejal Patwardhan and Lilian Weng and Aleksander Mądry , journal=. 2025 , url=
2025
-
[63]
2025 , url=
Deepak Nathani and Lovish Madaan and Nicholas Roberts and Nikolay Bashlykov and Ajay Menon and Vincent Moens and Amar Budhiraja and Despoina Magka and Vladislav Vorotilov and Gaurav Chaurasia and Dieuwke Hupkes and Ricardo Silveira Cabral and Tatiana Shavrina and Jakob Foerste...
2025
-
[64]
2026 , url=
Ben Rank and Hardik Bhatnagar and Ameya Prabhu and Shira Eisenberg and Karina Nguyen and Matthias Bethge and Maksym Andriushchenko , journal=. 2026 , url=
2026
-
[65]
2025 , url=
Hjalmar Wijk and Tao Lin and Joel Becker and Sami Jawhar and Neev Parikh and Thomas Broadley and Lawrence Chan and Michael Chen and Josh Clymer and Jai Dhyani and Elena Ericheva and Katharyn Garcia and Brian Goodrich and Nikola Jurkovic and Holden Karnofsky and Megan Kinniment...
2025
-
[66]
Baker and Benjamin Burns and Daniel Adu-Ampratwum and Xuhui Huang and Xia Ning and Song Gao and Yu Su and Huan Sun , journal=
Ziru Chen and Shijie Chen and Yuting Ning and Qianheng Zhang and Boshi Wang and Botao Yu and Yifei Li and Zeyi Liao and Chen Wei and Zitong Lu and Vishal Dey and Mingyi Xue and Frazier N. Baker and Benjamin Burns and Daniel Adu-Ampratwum and Xuhui Huang and Xia Ning and Song G...
2025
-
[67]
2409.07440 , archivePrefix=
Ben Bogin and Kejuan Yang and Shashank Gupta and Kyle Richardson and Erin Bransom and Peter Clark and Ashish Sabharwal and Tushar Khot , year=. 2409.07440 , archivePrefix=
-
[68]
Miller and Abhishek Charnalia and Derek Dunfield and Carole-Jean Wu and Pontus Stenetorp and Nicola Cancedda and Jakob Nicolaus Foerster and Yoram Bachrach , journal=
Edan Toledo and Karen Hambardzumyan and Martin Josifoski and Rishi Hazra and Nicolas Baldwin and Alexis Audran-Reiss and Michael Kuchnik and Despoina Magka and Minqi Jiang and Alisia Maria Lupidi and Andrei Lupu and Roberta Raileanu and Kelvin Niu and Tatiana Shavrina and Jean...
2025
-
[69]
2025 , url=
Zhengyao Jiang and Dominik Schmidt and Dhruv Srikanth and Dixing Xu and Ian Kaplan and Deniss Jacenko and Yuxiang Wu , journal=. 2025 , url=
2025
-
[70]
2026 , url=
Karen Hambardzumyan and Nicolas Baldwin and Edan Toledo and Rishi Hazra and Michael Kuchnik and Bassel Al Omari and Thomas Simon Foster and Anton Protopopov and Jean-Christophe Gagnon-Audet and Ishita Mediratta and Kelvin Niu and Michael Shvartsman and Alisia Lupidi and Alexis...
2026
-
[71]
2511.08522 , archivePrefix=
Zhaojian Yu and Kaiyue Feng and Yilun Zhao and Shilin He and Xiao-Ping Zhang and Arman Cohan , year=. 2511.08522 , archivePrefix=
-
[72]
Nature , volume=
Mathematical discoveries from program search with large language models , author=. Nature , volume=. 2024 , publisher=
2024
-
[73]
2025 , eprint=
Mathematical exploration and discovery at scale , author=. 2025 , eprint=
2025
-
[74]
2025 , url=
Robert Tjarko Lange and Yuki Imajuku and Edoardo Cetin , journal=. 2025 , url=
2025
-
[75]
Wider or Deeper?
Yuichi Inoue and Kou Misaki and Yuki Imajuku and So Kuroki and Taishi Nakamura and Takuya Akiba , journal=. Wider or Deeper?. 2025 , url=
2025
-
[76]
Algorithm Discovery With
Anja Surina and Amin Mansouri and Lars Quaedvlieg and Amal Seddas and Maryna Viazovska and Emmanuel Abbe and Caglar Gulcehre , journal=. Algorithm Discovery With. 2025 , url=
2025
-
[77]
Nature , volume=
Discovering faster matrix multiplication algorithms with reinforcement learning , author=. Nature , volume=. 2022 , publisher=
2022
-
[78]
arXiv preprint arXiv:2301.07608 , year=
Human-Timescale Adaptation in an Open-Ended Task Space , author=. arXiv preprint arXiv:2301.07608 , year=
-
[79]
International Conference on Learning Representations , volume=
Reinforcement learning for machine learning engineering agents , author=. International Conference on Learning Representations , volume=
-
[80]
Nature , volume=
Olympiad-level formal mathematical reasoning with reinforcement learning , author=. Nature , volume=. 2026 , publisher=
2026
-
[81]
2026 , url=
Jenny Zhang and Shengran Hu and Cong Lu and Robert Lange and Jeff Clune , journal=. 2026 , url=
2026
-
[82]
2024 , eprint=
Open-Endedness is Essential for Artificial Superhuman Intelligence , author=. 2024 , eprint=
2024
-
[83]
Gottweis, Juraj and Weng, Wei-Hung and Daryin, Alexander and Tu, Tao and Sirkovic, Petar and Myaskovsky, Artiom and Glowaty, Grzegorz and Weissenberger, Felix and Orlandi, Alessio and Popovici, Dan and Palepu, Anil and Rong, Keran and Tanno, Ryutaro and Saab, Khaled and Zhang,...
-
[84]
2503.18102 , archivePrefix=
Samuel Schmidgall and Michael Moor , year=. 2503.18102 , archivePrefix=
-
[85]
Chenglei Si and Diyi Yang and Tatsunori Hashimoto , year=. Can. 2409.04109 , archivePrefix=
-
[86]
2025 , url=
Yixuan Weng and Minjun Zhu and Guangsheng Bao and Hongbo Zhang and Jindong Wang and Yue Zhang and Linyi Yang , journal=. 2025 , url=
2025
-
[87]
2408.14033 , archivePrefix=
Ruochen Li and Teerth Patel and Qingyun Wang and Xinya Du , year=. 2408.14033 , archivePrefix=
-
[88]
2404.07738 , archivePrefix=
Jinheon Baek and Sujay Kumar Jauhar and Silviu Cucerzan and Sung Ju Hwang , year=. 2404.07738 , archivePrefix=
-
[89]
2026 , url=
Meysam Alizadeh and Mohsen Mosleh and Fabrizio Gilardi and Atoosa Kasirzadeh and Joshua Tucker , journal=. 2026 , url=
2026
-
[90]
2025 , url=
Jiabin Tang and Lianghao Xia and Zhonghang Li and Chao Huang , journal=. 2025 , url=
2025
-
[91]
arXiv preprint arXiv:2605.00803 , year=
Can Coding Agents Reproduce Findings in Computational Materials Science? , author=. arXiv preprint arXiv:2605.00803 , year=
-
[92]
2026 , eprint=
Coding-agents can replicate scientific machine learning papers , author=. 2026 , eprint=
2026
-
[93]
Siegel and Sayash Kapoor and Nitya Nadgir and Benedikt Stroebl and Arvind Narayanan , journal=
Zachary S. Siegel and Sayash Kapoor and Nitya Nadgir and Benedikt Stroebl and Arvind Narayanan , journal=. 2026 , url=
2026
-
[94]
2025 , url=
Patrick Tser Jern Kon and Jiachen Liu and Xinyi Zhu and Qiuyi Ding and Jingjia Peng and Jiarong Xing and Yibo Huang and Yiming Qiu and Jayanth Srinivasa and Myungjin Lee and Mosharaf Chowdhury and Matei Zaharia and Ang Chen , journal=. 2025 , url=
2025
-
[95]
arXiv preprint arXiv:2506.19724 , year=
From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking , author=. arXiv preprint arXiv:2506.19724 , year=
-
[96]
2025 , url=
Shuo Yan and Ruochen Li and Ziming Luo and Zimu Wang and Daoyang Li and Liqiang Jing and Kaiyu He and Peilin Wu and George Michalopoulos and Yue Zhang and Ziyang Zhang and Mian Zhang and Zhiyu Chen and Xinya Du , journal=. 2025 , url=
2025
-
[97]
2505.19955 , archivePrefix=
Hui Chen and Miao Xiong and Yujie Lu and Wei Han and Ailin Deng and Yufei He and Jiaying Wu and Yibo Li and Yue Liu and Bryan Hooi , year=. 2505.19955 , archivePrefix=
-
[98]
2026 , url=
Sasi Kiran Gaddipati and Diyana Muhammed and Farhana Keya and Gollam Rabby and Sören Auer , journal=. 2026 , url=
2026
-
[99]
2025 , url=
Starace, Giulio and Jaffe, Oliver and Sherburn, Dane and Aung, James and Chan, Jun Shern and Maksin, Leon and Dias, Rachel and Mays, Evan and Kinsella, Benjamin and Thompson, Wyatt and Heidecke, Johannes and Glaese, Amelia and Patwardhan, Tejal , journal=. 2025 , url=
2025
-
[100]
2026 , url=
Shi Qiu and Junyi Deng and Yiwei Deng and Haoran Dong and Jieyu Fu and Mao Li and Zeyu Li and Zhaolong Zhang and Huiwen Zheng and Leidong Bao and Anqi Lv and Zihan Mo and Yadi Niu and Yiyang Peng and Yu Tian and Yili Wang and Ziyu Wang and Zi-Yu Wang and Jiashen Wei and Liuhen...
2026
-
[101]
2025 , url=
Chuxuan Hu and Liyun Zhang and Yeji Lim and Aum Wadhwani and Austin Peters and Daniel Kang , journal=. 2025 , url=
2025
-
[102]
Truong and Weixin Liang and Fan-Yun Sun and Nick Haber , journal=
Tianyu Hua and Harper Hua and Violet Xiang and Benjamin Klieger and Sang T. Truong and Weixin Liang and Fan-Yun Sun and Nick Haber , journal=. 2025 , url=
2025
-
[103]
2025 , url=
Yanzheng Xiang and Hanqi Yan and Shuyin Ouyang and Lin Gui and Yulan He , journal=. 2025 , url=
2025
-
[104]
Miller and Oisin Mac Aodha and Jakob Foerster and Yoram Bachrach , journal=
Bingchen Zhao and Despoina Magka and Minqi Jiang and Xian Li and Roberta Raileanu and Tatiana Shavrina and Jean-Christophe Gagnon-Audet and Kelvin Niu and Shagun Sodhani and Michael Shvartsman and Andrei Lupu and Alisia Lupidi and Edan Toledo and Karen Hambardzumyan and Martin...
2025
-
[105]
2026 , url=
Xuanle Zhao and Zilin Sang and Yuxuan Li and Qi Shi and Weilun Zhao and Shuo Wang and Duzhen Zhang and Xu Han and Zhiyuan Liu and Maosong Sun , journal=. 2026 , url=
2026
-
[106]
2026 , url=
Haokun Liu and Filbert Aurelian Tjiaranata and Chenhao Tan , journal=. 2026 , url=
2026
-
[107]
2026 , url=
Minju Seo and Jinheon Baek and Seongyun Lee and Sung Ju Hwang , journal=. 2026 , url=
2026
-
[108]
2506.11763 , archivePrefix=
Mingxuan Du and Benfeng Xu and Chiwei Zhu and Xiaorui Wang and Zhendong Mao , year=. 2506.11763 , archivePrefix=
-
[109]
Rahul K. Arora and Jason Wei and Rebecca Soskin Hicks and Preston Bowman and Joaquin Quiñonero-Candela and Foivos Tsimpourlas and Michael Sharman and Meghan Shah and Andrea Vallone and Alex Beutel and Johannes Heidecke and Karan Singhal , year=. 2505.08775 , archivePrefix=
-
[110]
Liu, Amelia and Ho, Andrew and Droste, Anne Marie and Martin, David and Wong, Edmund and Zhou, Edward and Zhou, Isabelle and Park, Joshua and Jiao, Joy and Skelly, Katie-Rose and others , journal=
-
[111]
2403.07974 , archivePrefix=
Naman Jain and King Han and Alex Gu and Wen-Ding Li and Fanjia Yan and Tianjun Zhang and Sida Wang and Armando Solar-Lezama and Koushik Sen and Ion Stoica , year=. 2403.07974 , archivePrefix=
-
[112]
Zhihong Shao and Peiyi Wang and Qihao Zhu and Runxin Xu and Junxiao Song and Xiao Bi and Haowei Zhang and Mingchuan Zhang and Y. K. Li and Y. Wu and Daya Guo , journal=. 2024 , url=
2024
-
[113]
Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen-Zhu and Yuanzhi Li and Shean Wang and Lu Wang and Weizhu Chen , booktitle=
Edward J. Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen-Zhu and Yuanzhi Li and Shean Wang and Lu Wang and Weizhu Chen , booktitle=
-
[114]
Back to Basics: Revisiting
Arash Ahmadian and Chris Cremer and Matthias Gall. Back to Basics: Revisiting. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , year=
-
[115]
2025 , url=
Qiying Yu and Zheng Zhang and Ruofei Zhu and Yufeng Yuan and Xiaochen Zuo and Yu Yue and Tiantian Fan and Gaohong Liu and Lingjun Liu and Xin Liu and Haibin Lin and Zhiqi Lin and Bole Ma and Guangming Sheng and Yuxuan Tong and Chi Zhang and Mofan Zhang and Wang Zhang and Hang ...
2025
-
[116]
Advances in Neural Information Processing Systems , year=
Meta learning backpropagation and improving it , author=. Advances in Neural Information Processing Systems , year=
-
[117]
hook , author=
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-... hook , author=. 1987 , school=
1987
-
[118]
International Conference on Learning Representations , year=
Improving Generalization in Meta Reinforcement Learning using Learned Objectives , author=. International Conference on Learning Representations , year=
-
[119]
Self-taught optimizer (
Zelikman, Eric and Lorch, Eliana and Mackey, Lester and Kalai, Adam Tauman , journal=. Self-taught optimizer (. 2023 , url=
2023
-
[120]
arXiv preprint arXiv:2212.14392 , year=
Eliminating Meta Optimization Through Self-Referential Meta Learning , author=. arXiv preprint arXiv:2212.14392 , year=
-
[121]
Shifting inductive bias with success-story algorithm, adaptive
Schmidhuber, J. Shifting inductive bias with success-story algorithm, adaptive. Machine Learning , volume=. 1997 , publisher=
1997
-
[122]
Frontiers in psychology , volume=
Schmidhuber, J. Frontiers in psychology , volume=. 2013 , publisher=
2013
-
[123]
International Conference on Artificial Neural Networks , year=
A ‘self-referential’weight matrix , author=. International Conference on Artificial Neural Networks , year=
-
[124]
A possibility for implementing curiosity and boredom in model-building neural controllers , author=. Proc. of the international conference on simulation of adaptive behavior: From animals to animats , year=
-
[125]
Claude Models Overview , year=
-
[126]
arXiv preprint arXiv:2510.18855 , year=
Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model , author=. arXiv preprint arXiv:2510.18855 , year=
-
[127]
arXiv preprint arXiv:2605.02572 , year=
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length , author=. arXiv preprint arXiv:2605.02572 , year=
-
[128]
and Sun, Renliang and Zhu, Yanqiao and Cong, Jason and Sun, Yizhou and Wang, Wei , journal=
Wang, Xiaoxuan and Zhang, Han and Wang, Haixin and Shi, Yidan and Li, Ruoyan and Han, Kaiqiao and Tong, Chenyi and Deng, Haoran and Taylor, Alexander K. and Sun, Renliang and Zhu, Yanqiao and Cong, Jason and Sun, Yizhou and Wang, Wei , journal=. 2026 , url=
2026
-
[129]
arXiv preprint arXiv:2509.25598 , year=
Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks , author=. arXiv preprint arXiv:2509.25598 , year=
-
[130]
Proceedings of the AAAI Conference on Artificial Intelligence , year =
Deep Reinforcement Learning that Matters , author =. Proceedings of the AAAI Conference on Artificial Intelligence , year =
-
[131]
Journal of Machine Learning Research , volume =
Variance Reduction Techniques for Gradient Estimates in Reinforcement Learning , author =. Journal of Machine Learning Research , volume =
-
[132]
arXiv preprint arXiv:2210.10760 , year =
Scaling Laws for Reward Model Overoptimization , author =. arXiv preprint arXiv:2210.10760 , year =
-
[133]
2022 , url=
Kueue: Kubernetes-native Job Queueing , author=. 2022 , url=
2022
-
[134]
and Stoica, Ion , booktitle=
Moritz, Philipp and Nishihara, Robert and Wang, Stephanie and Tumanov, Alexey and Liaw, Richard and Liang, Eric and Elibol, Melih and Yang, Zongheng and Paul, William and Jordan, Michael I. and Stoica, Ion , booktitle=. Ray: A Distributed Framework for Emerging
-
[135]
Megatron-
Shoeybi, Mohammad and Patwary, Mostofa and Puri, Raul and LeGresley, Patrick and Casper, Jared and Catanzaro, Bryan , journal=. Megatron-. 2019 , url=
2019
-
[136]
and Zhang, Hao and Stoica, Ion , booktitle =
Kwon, Woosuk and Li, Zhuohan and Zhuang, Siyuan and Sheng, Ying and Zheng, Lianmin and Yu, Cody Hao and Gonzalez, Joseph E. and Zhang, Hao and Stoica, Ion , booktitle =. Efficient Memory Management for Large Language Model Serving with
-
[137]
Gated Delta Networks: Improving
Yang, Songlin and Kautz, Jan and Hatamizadeh, Ali , booktitle=. Gated Delta Networks: Improving
-
[138]
Handbook of Evolutionary Machine Learning , pages=
Evolution through large models , author=. Handbook of Evolutionary Machine Learning , pages=. 2023 , publisher=
2023
-
[139]
General-Purpose
Kirsch, Louis and Harrison, James and Sohl-Dickstein, Jascha and Metz, Luke , journal=. General-Purpose. 2022 , note=
2022
-
[140]
Zhuge, Mingchen and Zhao, Changsheng and Ashley, Dylan R. and Wang, Wenyi and Khizbullin, Dmitrii and Xiong, Yunyang and Liu, Zechun and Chang, Ernie and Krishnamoorthi, Raghuraman and Tian, Yuandong and Shi, Yangyang and Chandra, Vikas and Schmidhuber, J. Proceedings of the 4...
-
[141]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , booktitle =. Judging. 2023 , note =
2023
-
[142]
Forty-first International Conference on Machine Learning , year=
Mingchen Zhuge and Wenyi Wang and Louis Kirsch and Francesco Faccio and Dmitrii Khizbullin and J. Forty-first International Conference on Machine Learning , year=
-
[143]
arXiv preprint arXiv:2408.08435 , year=
Automated design of agentic systems , author=. arXiv preprint arXiv:2408.08435 , year=
-
[144]
International Conference on Learning Representations , year=
Large language models as optimizers , author=. International Conference on Learning Representations , year=
-
[145]
Rubric-based
Fang, Junfeng and Hong, Zhepei and Zheng, Mao and Song, Mingyang and Li, Gengsheng and Jiang, Houcheng and Zhang, Dan and Guo, Haiyun and Wang, Xiang and Chua, Tat-Seng , journal =. Rubric-based
-
[146]
arXiv preprint arXiv:2507.06261 , year=
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities , author=. arXiv preprint arXiv:2507.06261 , year=
-
[147]
Advances in Neural Information Processing Systems , year=
Measuring what matters: Construct validity in large language model benchmarks , author=. Advances in Neural Information Processing Systems , year=
-
[148]
, author=
Construct validity in psychological tests. , author=. Psychological Bulletin , volume=. 1955 , publisher=
1955
-
[149]
2020 , organization=
Real, Esteban and Liang, Chen and So, David and Le, Quoc , booktitle=. 2020 , organization=
2020
-
[150]
Advances in Neural Information Processing Systems , year=
Discovering reinforcement learning algorithms , author=. Advances in Neural Information Processing Systems , year=
-
[151]
Preprints Conf
On the optimization of a synaptic learning rule , author=. Preprints Conf. Optimality in Artificial and Biological Neural Networks , year=
-
[152]
Transactions on Machine Learning Research , issn=
Voyager: An Open-Ended Embodied Agent with Large Language Models , author=. Transactions on Machine Learning Research , issn=. 2024 , url=
2024
-
[153]
ACS Central Science , volume =
Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules , author =. ACS Central Science , volume =. 2018 , doi =
2018
-
[154]
Nature , volume =
Scaling Deep Learning for Materials Discovery , author =. Nature , volume =. 2023 , doi =
2023
-
[155]
2025 , eprint=
Reinforcing General Reasoning without Verifiers , author=. 2025 , eprint=
2025
-
[156]
Language Gamification-NeurIPS 2024 Workshop , year=
On Reward Functions For Self-Improving Chain-of-Thought Reasoning Without Supervised Datasets (Abridged Version) , author=. Language Gamification-NeurIPS 2024 Workshop , year=
2024
-
[157]
2022 , eprint=
Learning to summarize from human feedback , author=. 2022 , eprint=
2022
-
[158]
2022 , eprint=
Training language models to follow instructions with human feedback , author=. 2022 , eprint=
2022
-
[159]
2026 , eprint=
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision , author=. 2026 , eprint=
2026
-
[160]
Why Is That Relevant?
McDonnell, Tyler and Lease, Matthew and Kutlu, Mucahid and Elsayed, Tamer , booktitle=. Why Is That Relevant?
-
[161]
arXiv preprint arXiv:1710.10324 , year=
Crystal Graph Convolutional Neural Networks for an Accurate and Interpretable Prediction of Material Properties , author=. arXiv preprint arXiv:1710.10324 , year=
-
[162]
2020 , url=
Chithrananda, Seyone and Grand, Gabriel and Ramsundar, Bharath , journal=. 2020 , url=
2020
-
[163]
arXiv preprint arXiv:1502.02072 , year=
Massively Multitask Networks for Drug Discovery , author=. arXiv preprint arXiv:1502.02072 , year=
-
[164]
2025 , url=
Xu, Changwen and Zhu, Shang and Viswanathan, Venkatasubramanian , journal=. 2025 , url=
2025
-
[165]
2026 , url=
Bhattacharjee, Satadeep , journal=. 2026 , url=
2026
-
[166]
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in
Gu, Hengrui and Han, Xiaotian and Bian, Yujing and Wang, Feiyi and Zhou, Kaixiong , journal=. Asymmetric Advantage Modulation Calibrates Entropy Dynamics in. 2026 , url=
2026
-
[167]
Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of
Lin, Jianzhe , journal=. Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of. 2026 , url=
2026
-
[168]
Understanding Diversity Collapse in
Yuan, Suqin and Chen, Jinkun and Zheng, Jiyang and Li, Muyang and Feng, Lei and Wang, Dadong and Xiang, Tao and Liu, Tongliang and An, Bo , journal=. Understanding Diversity Collapse in. 2026 , url=
2026
-
[169]
Science , volume=
Reducing the dimensionality of data with neural networks , author=. Science , volume=
-
[170]
A foundation model for the
Bodnar, Cristian and Bruinsma, Wessel P and Lucic, Ana and Stanley, Megan and Allen, Anna and Brandstetter, Johannes and Garvan, Patrick and Riechert, Maik and Weyn, Jonathan A and Dong, Haiyu and Gupta, Jayesh K and Thambiratnam, Kit and Archibald, Alexander T and Wu, Chun-Ch...
-
[171]
BMC Bioinformatics , volume=
Efficient discovery of responses of proteins to compounds using active learning , author=. BMC Bioinformatics , volume=
-
[172]
Physics informed deep learning (
Raissi, Maziar and Perdikaris, Paris and Karniadakis, George Em , journal=. Physics informed deep learning (. 2017 , url=
2017
-
[173]
International Conference on Machine Learning , year=
Social influence as intrinsic motivation for multi-agent deep reinforcement learning , author=. International Conference on Machine Learning , year=
-
[174]
Towards execution-grounded automated
Si, Chenglei and Yang, Zitong and Choi, Yejin and Cand. Towards execution-grounded automated. arXiv preprint arXiv:2601.14525 , year=
-
[175]
International Conference on Machine Learning , year=
Asynchronous methods for deep reinforcement learning , author=. International Conference on Machine Learning , year=
-
[176]
Mnih, Volodymyr and Kavukcuoglu, Koray and Silver, David and Graves, Alex and Antonoglou, Ioannis and Wierstra, Daan and Riedmiller, Martin , journal=. Playing
-
[177]
Auto-encoding variational
Kingma, Diederik P and Welling, Max , journal=. Auto-encoding variational. 2013 , url=
2013
-
[178]
Journal of Machine Learning Research , volume=
Exploring strategies for training deep neural networks , author=. Journal of Machine Learning Research , volume=
-
[179]
Proceedings of the IEEE , volume=
Gradient-based learning applied to document recognition , author=. Proceedings of the IEEE , volume=
-
[180]
Chawla, Nitesh V and Bowyer, Kevin W and Hall, Lawrence O and Kegelmeyer, W Philip , journal=
-
[181]
Scaling in-context online learning capability of
Lin, Xiaofeng and Zhu, Sirou and Chen, Yilei and Chen, Mingyu and Sang, Hejian and Paschalidis, Ioannis and Wang, Zhipeng and Pacchiano, Aldo and Zhang, Xuezhou , journal=. Scaling in-context online learning capability of. 2026 , url=
2026
-
[182]
The Annals of Statistics , volume=
Greedy function approximation: A gradient boosting machine , author=. The Annals of Statistics , volume=
-
[183]
Journal of Machine Learning Research , volume=
Manifold regularization: A geometric framework for learning from labeled and unlabeled examples , author=. Journal of Machine Learning Research , volume=
-
[184]
Journal of Machine Learning Research , volume=
Dropout: A simple way to prevent neural networks from overfitting , author=. Journal of Machine Learning Research , volume=
-
[185]
Krizhevsky, Alex and Sutskever, Ilya and Hinton, Geoffrey E , booktitle=
-
[186]
Advances in Neural Information Processing Systems , year=
Recht, Benjamin and R. Advances in Neural Information Processing Systems , year=
-
[187]
International Conference on Learning Representations , year=
Diversity is all you need: Learning skills without a reward function , author=. International Conference on Learning Representations , year=
-
[188]
Advances in Neural Information Processing Systems , year=
Toolformer: Language models can teach themselves to use tools , author=. Advances in Neural Information Processing Systems , year=
-
[189]
Advances in Neural Information Processing Systems , year=
Convolutional networks on graphs for learning molecular fingerprints , author=. Advances in Neural Information Processing Systems , year=
-
[190]
Advances in Neural Information Processing Systems , year=
Batatia, Ilyes and Kov. Advances in Neural Information Processing Systems , year=
-
[191]
Molecular
Schwaller, Philippe and Laino, Teodoro and Gaudin, Th. Molecular. ACS Central Science , volume=
-
[192]
Proceedings of the National Academy of Sciences , volume=
Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences , author=. Proceedings of the National Academy of Sciences , volume=
-
[193]
Accurate structure prediction of biomolecular interactions with
Abramson, Josh and Adler, Jonas and Dunger, Jack and Evans, Richard and Green, Tim and Pritzel, Alexander and Ronneberger, Olaf and Willmore, Lindsay and Ballard, Andrew J and Bambrick, Joshua and others , journal=. Accurate structure prediction of biomolecular interactions with
-
[194]
Nature , volume=
A generative model for inorganic materials design , author=. Nature , volume=
-
[195]
Science , volume=
Learning skillful medium-range global weather forecasting , author=. Science , volume=
-
[196]
2022 , url=
Pathak, Jaideep and Subramanian, Shashank and Harrington, Peter and Raja, Sanjeev and Chattopadhyay, Ashesh and Mardani, Morteza and Kurth, Thorsten and Hall, David and Li, Zongyi and Azizzadenesheli, Kamyar and Hassanzadeh, Pedram and Kashinath, Karthik and Anandkumar, Animas...
2022
-
[197]
arXiv preprint arXiv:2412.09560 , year=
Foundational large language models for materials research , author=. arXiv preprint arXiv:2412.09560 , year=
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.