Pith. sign in

REVIEW 2 major objections 1 minor 221 references

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments

T0 review · 2 major / 1 minor · reviewed 2026-05-15 · grok-4.3

Pith's one-line read Reinforcement learning environments are splitting into LLM-driven semantic systems and domain-specific physical generalization systems.

desk verdict The paper scales literature analysis to map RL environments into a taxonomy and claims a split between LLM-driven semantic and domain-specific generalization ecosystems, but the shift rests on an unvalidated automated pipeline. read the letter →

arxiv 2603.23964 v2 submitted 2026-03-25 cs.AI

classification cs.AI
keywords reinforcementlearningenvironmentstaxonomylargelanguagemodelsparadigmshiftempiricalstudycognitivecapabilitiesgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper conducts a large-scale quantitative analysis of more than 2,000 RL publications to map how training environments have evolved from isolated physical simulations toward generalist language-driven agents. It introduces a multi-dimensional taxonomy that classifies environments according to their application domains and the specific cognitive capabilities they test. Automated semantic and statistical processing of the literature identifies a data-supported bifurcation of the field into a Semantic Prior ecosystem centered on large language models and a Domain-Specific Generalization ecosystem. The study further extracts cognitive fingerprints for each branch to explain patterns of skill transfer, interference, and zero-shot generalization. These results supply a concrete roadmap for building Embodied Semantic Simulators that link continuous physical control with high-level logical reasoning.

What carries the argument

A novel multi-dimensional taxonomy that classifies RL benchmarks by application domains and required cognitive capabilities, derived through automated processing of publication data.

What would settle it

A manual audit of several hundred recent papers that places the majority outside both the Semantic Prior and Domain-Specific Generalization categories or shows no statistical evidence of bifurcation.

Watch

Extended reading notes

Core claim

Automated semantic and statistical analysis of a corpus of over 2,000 RL publications reveals a paradigm shift in which the field bifurcates into a Semantic Prior ecosystem dominated by Large Language Models and a Domain-Specific Generalization ecosystem; each ecosystem carries distinct cognitive fingerprints that govern cross-task synergy, multi-domain interference, and zero-shot generalization.

Load-bearing premise

That programmatically analyzing a large corpus of publications produces an unbiased map of RL environment trends without meaningful selection or interpretation bias in the pipeline.

Editorial extensions

If this is right

  • Designers of new agents can target the two ecosystems separately before combining their strengths.
  • Cognitive fingerprints offer a practical way to forecast and improve zero-shot transfer between tasks.
  • Embodied Semantic Simulators can be built by deliberately bridging pixel-level control with language-level reasoning.
  • Environment selection for training can be guided by the quantitative trends rather than qualitative judgment alone.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Separate training regimes for semantic and physical skills may become standard before integration into single agents.
  • Applying the same automated taxonomy method to other AI domains could expose parallel splits in research focus.
  • New environments released after the study can be classified under the taxonomy to test whether the bifurcation continues.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript presents a large-scale empirical study analyzing over 2,000 publications on reinforcement learning environments. It proposes a multi-dimensional taxonomy to map the evolution from physical simulations to language-driven agents and claims a paradigm shift bifurcating the field into a 'Semantic Prior' ecosystem dominated by LLMs and a 'Domain-Specific Generalization' ecosystem, while characterizing 'cognitive fingerprints' for cross-task synergy and zero-shot generalization.

Significance. If the automated analysis is shown to be robust, the quantitative taxonomy could provide a valuable roadmap for designing next-generation embodied semantic simulators that bridge continuous physical control and high-level reasoning, moving the field beyond purely qualitative reviews.

major comments (2)
  1. [Methodology] The automated semantic and statistical analysis section provides no description of the embedding model, dimensionality reduction technique, clustering algorithm, or any other implementation details used to derive the multi-dimensional taxonomy and the bifurcation into Semantic Prior and Domain-Specific Generalization ecosystems. This prevents assessment of whether the reported paradigm shift is an artifact of the pipeline choices.
  2. [Results] No human-annotated ground-truth dataset, precision/recall metrics, or external benchmark comparisons are reported to validate the automated labels or the claimed cross-task synergy and zero-shot generalization patterns. The central claim of a 'data-verified paradigm shift' therefore rests on an unvalidated process.
minor comments (1)
  1. [Abstract] The abstract introduces 'cognitive fingerprints' without a concise operational definition, which should be clarified early to aid reader comprehension.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback on our empirical analysis of reinforcement learning environments. The comments highlight important areas for improving methodological transparency and validation, which we address below.

read point-by-point responses
  1. Referee: [Methodology] The automated semantic and statistical analysis section provides no description of the embedding model, dimensionality reduction technique, clustering algorithm, or any other implementation details used to derive the multi-dimensional taxonomy and the bifurcation into Semantic Prior and Domain-Specific Generalization ecosystems. This prevents assessment of whether the reported paradigm shift is an artifact of the pipeline choices.

    Authors: We agree that the original manuscript omits key implementation details of the automated pipeline. In the revised version, we will insert a dedicated 'Implementation Details' subsection describing the embedding model (all-MiniLM-L6-v2 via sentence-transformers), dimensionality reduction (UMAP with n_neighbors=15 and min_dist=0.1), clustering (HDBSCAN with min_cluster_size=5), and the statistical procedures used to detect the bifurcation and cognitive fingerprints. These additions will support reproducibility and allow readers to evaluate whether the observed paradigm shift depends on specific pipeline choices. revision: yes

  2. Referee: [Results] No human-annotated ground-truth dataset, precision/recall metrics, or external benchmark comparisons are reported to validate the automated labels or the claimed cross-task synergy and zero-shot generalization patterns. The central claim of a 'data-verified paradigm shift' therefore rests on an unvalidated process.

    Authors: We acknowledge the absence of quantitative validation in the submitted manuscript. The revised version will add a validation subsection that reports a human annotation study on a random subset of 150 papers, yielding precision/recall figures and inter-annotator agreement against the automated labels. We will also include direct comparisons with prior qualitative taxonomies from the RL literature. While exhaustive ground-truth labeling of the full corpus remains resource-intensive, these targeted validations will provide concrete support for the bifurcation, cross-task synergy, and zero-shot patterns. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical taxonomy derived from corpus analysis

full rationale

The paper conducts a data-driven literature review by programmatically processing >2000 publications to propose a novel multi-dimensional taxonomy and then applies automated semantic/statistical analysis to identify a bifurcation into Semantic Prior and Domain-Specific Generalization ecosystems. This bifurcation is reported as an output of the analysis on the processed corpus rather than a definitional premise or fitted parameter renamed as a prediction. No equations, self-citations, uniqueness theorems, or ansatzes are invoked in the provided sections that would reduce the central claims to their own inputs by construction. The methodology remains self-contained as an empirical mapping exercise.

Assumptions & free parameters 0 free parameters · 1 assumptions · 3 invented entities

Central claims rest on the representativeness of the selected literature corpus and the validity of automated semantic extraction; new categories are introduced without external falsifiable tests.

assumptions (1)
  • domain assumption The corpus of over 2,000 core publications accurately represents the evolution of RL environments
    All quantitative findings and the paradigm-shift claim are derived from processing this corpus.
invented entities (3)
  • Semantic Prior ecosystem
    purpose: Label environments dominated by LLMs for semantic understanding
    Introduced as one branch of the bifurcation without independent validation data.
  • Domain-Specific Generalization ecosystem
    purpose: Label environments focused on task-specific generalization
    Introduced as the contrasting branch of the paradigm shift.
  • cognitive fingerprints
    purpose: Describe domain-specific mechanisms of synergy and interference
    New term coined to characterize analysis outputs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments." pith.science (2026). https://pith.science/paper/2603.23964

@misc{pith2026260323964,
  author       = {Pith},
  title        = {Pith review of: From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2603.23964}},
  note         = {Machine review of arXiv:2603.23964}
}
read the original abstract

The remarkable progress of reinforcement learning (RL) is intrinsically tied to the environments used to train and evaluate artificial agents. Moving beyond traditional qualitative reviews, this work presents a large-scale, data-driven empirical investigation into the evolution of RL environments. By programmatically processing a massive corpus of academic literature and rigorously distilling over 2,000 core publications, we propose a quantitative methodology to map the transition from isolated physical simulations to generalist, language-driven foundation agents. Implementing a novel, multi-dimensional taxonomy, we systematically analyze benchmarks against diverse application domains and requisite cognitive capabilities. Our automated semantic and statistical analysis reveals a profound, data-verified paradigm shift: the bifurcation of the field into a "Semantic Prior" ecosystem dominated by Large Language Models (LLMs) and a "Domain-Specific Generalization" ecosystem. Furthermore, we characterize the "cognitive fingerprints" of these distinct domains to uncover the underlying mechanisms of cross-task synergy, multi-domain interference, and zero-shot generalization. Ultimately, this study offers a rigorous, quantitative roadmap for designing the next generation of Embodied Semantic Simulators, bridging the gap between continuous physical control and high-level logical reasoning.

Figures

Figures reproduced from arXiv: 2603.23964 by the authors.

Figure 1
Figure 1. The Evolution of Reinforcement Learning Environments: A chronological visual timeline illustrating [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The Evolutionary Tree of Reinforcement Learning Environments: The Ascent of Cognitive Abstraction. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. The taxonomy of multi-dimensional spectrum of reinforcement learning task types [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (14 more)
Figure 4
Figure 4. Figure 4: WebArena: The Frontier of Vision-Language￾Action (VLA) Fusion. Representing the modern multi￾modal landscape, this environment requires agents to ground open-ended natural language instructions into dense visual interfaces. It forces a complex synthesis of image-based …
Figure 5
Figure 5. Figure 5: The multi-dimensional landscape of requisite agent capabilities. The diagram illustrates the diverse skill set [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: A taxonomic overview of diverse reinforcement learning application domains. The progression illustrates [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: The CARLA Autonomous Driving Simulator. Illustrating the pinnacle of the Autonomous Systems & Nav￾igation domain, CARLA forces agents to process multi￾modal, heterogeneous sensor streams (including RGB-D, LiDAR point clouds, and GPS). Operating under severe partial obs…
Figure 10
Figure 10. Figure 10 [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]
Figure 9
Figure 9. Figure 9: AlphaStar Mastering StarCraft II. A landmark achievement in Real-Time Strategy, AlphaStar masters the immense complexity of StarCraft II. By integrating deep neural networks with a multi-agent reinforcement learning league, it overcomes imperfect information and a mass…
Figure 11
Figure 11. Figure 11: Insilico Medicine’s Fully Automated Robotics Laboratory. Representing the frontier of AI￾driven drug discovery, this system automates complex wet-lab processes. It integrates reinforcement learning to autonomously optimize experimental strategies and pro￾cess control …
Figure 12
Figure 12. Figure 12: A retrospective analysis of the temporal distribution of RL environments by primary application domain. The [PITH_FULL_IMAGE:figures/full_fig_p024_12.png]
Figure 13
Figure 13. Figure 13: The paradigm shift in modality distribution. The figure illustrates the breakdown of single-modality [PITH_FULL_IMAGE:figures/full_fig_p024_13.png]
Figure 14
Figure 14. Figure 14: Temporal Evolution of Capability Requirements in LLM Environments. This alluvial plot tracks the shifting cognitive demands of environments utilizing LLMs as active agents, based on their inception dates. The data illustrates a rapid escalation from foundational langu…
Figure 15
Figure 15. Figure 15: The evolution of primary application domains in a broader field. The trajectory illustrates a shift from [PITH_FULL_IMAGE:figures/full_fig_p028_15.png]
Figure 16
Figure 16. Figure 16: Evolutionary Trajectory of Agent Capabilities Across Four RL Eras. This citation-weighted alluvial plot illustrates the longitudinal shift in cognitive and physical requirements of RL environments from 2013 to the present. The temporal axis spans four major algorithmi…
Figure 17
Figure 17. Figure 17: Cognitive Fingerprints of LLMs engaged RL Application Subdomains. This row-normalized clustered heatmap illustrates the proportional distribution of required agent capabilities across various domains. Rows (application subdomains) are ordered via hierarchical clusteri…
Figure 18
Figure 18. Figure 18: Evolutionary Trajectory of Agent Capabilities. This citation-weighted alluvial plot illustrates the shifting cognitive and physical requirements of RL environments across four major algorithmic epochs: Classic DRL & Physics, Scalable Games & MARL, Offline & Pre-traini…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

221 extracted references · 221 canonical work pages

  1. [1]

    MIT press, 2018

    Richard S Sutton and Andrew G Barto.Reinforce- ment learning: An introduction. MIT press, 2018

  2. [2]

    Mastering the game of go with deep neural networks and tree search.Nature, 529(7587):484– 489, 2016

    David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Pan- neershelvam, Marc Lanctot, Sander Dieleman, Do- minik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Ko- ray Kavukcuoglu, Thore Graepel, and Demis Has- sabis. Masterin...

  3. [3]

    Grandmaster level in star- craft ii using multi-agent reinforcement learning

    Oriol Vinyals, Igor Babuschkin, Wojciech M Czar- necki, Michaël Mathieu, Andrew Dudzik, Juny- oung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Hor- gan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P Agapiou, Max Jaderberg, Alexander S Vezhnevets, Rémi Leblond, Tobias Pohlen, Valentin Dalib...

  4. [4]

    End-to-end training of deep visuomo- tor policies.Journal of Machine Learning Research, 17(39):1–40, 2016

    Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. End-to-end training of deep visuomo- tor policies.Journal of Machine Learning Research, 17(39):1–40, 2016

  5. [5]

    Rusu, Joel Veness, Marc G

    V olodymyr Mnih, Koray Kavukcuoglu, David Sil- ver, Andrei A. Rusu, Joel Veness, Marc G. Belle- mare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. Human-level con- trol through deep reinforc...

  6. [6]

    Deep reinforcement learning that matters

    Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. Deep reinforcement learning that matters. InPro- ceedings of the AAAI Conference on Artificial Intel- ligence, volume 32, 2018

  7. [7]

    The arcade learning environment: An evaluation platform for general agents.Jour- nal of Artificial Intelligence Research, 47:253–279, 2013

    Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. The arcade learning environment: An evaluation platform for general agents.Jour- nal of Artificial Intelligence Research, 47:253–279, 2013

  8. [8]

    Mu- JoCo: A physics engine for model-based control

    Emanuel Todorov, Tom Erez, and Yuval Tassa. Mu- JoCo: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelli- gent Robots and Systems, pages 5026–5033. IEEE, 2012

Show all 221 references
  1. [9]

    Unity: A general platform for intelligent agents.arXiv:1809.02627, 2018

    Arthur Juliani, Vincent-Pierre Berges, Ervin Teng, Andrew Cohen, Jonathan Harper, Chris Elion, Christopher Goy, Yuan Gao, Hunter Henry, Mar- wan Mattar, and Danny Lange. Unity: A general platform for intelligent agents.arXiv:1809.02627, 2018

  2. [10]

    Language models are few-shot learners.Advances in Neural Information Process- ing Systems, 33:1877–1901, 2020

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, et al. Language models are few-shot learners.Advances in Neural Information Process- ing Systems, 33:1877–1901, 2020

  3. [11]

    Gpt-4 technical report.arXiv:2303.08774, 2023

    OpenAI. Gpt-4 technical report.arXiv:2303.08774, 2023

  4. [12]

    Neuronlike adaptive elements that can solve difficult learning control problems.IEEE Transactions on Systems, Man, and Cybernetics, (5):834–846, 1983

    Andrew G Barto, Richard S Sutton, and Charles W Anderson. Neuronlike adaptive elements that can solve difficult learning control problems.IEEE Transactions on Systems, Man, and Cybernetics, (5):834–846, 1983

  5. [13]

    Domain randomization for transferring deep neural networks from simulation to the real world

    Josh Tobin, Rachel Fong, Alex Ray, Jonas Schnei- der, Wojciech Zaremba, and Pieter Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 23–30, 2017

  6. [14]

    Leveraging procedural generation 34 An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments to benchmark reinforcement learning

    Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman. Leveraging procedural generation 34 An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments to benchmark reinforcement learning. InInter- national Conference on Machine L...

  7. [15]

    A markovian decision process

    Richard Bellman. A markovian decision process. Journal of Mathematics and Mechanics, pages 679– 684, 1957

  8. [16]

    Textworld: A learning environment for text-based games

    Marc-Alexandre Côté, Akos Kádár, Xingdi Yuan, Ben Kybartas, Tavian Barnes, Emery Fine, James Moore, Matthew Hausknecht, Layla cross El Asri, Mahmoud Adada, Wendy Tay, and Adam Trischler. Textworld: A learning environment for text-based games. InWorkshop on Computer Games, page...

  9. [17]

    Re- act: Synergizing reasoning and acting in language models, 2023

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. Re- act: Synergizing reasoning and acting in language models, 2023

  10. [18]

    Transfer learning for reinforcement learning domains: A survey.Journal of Machine Learning Research, 10(Jul):1633–1685, 2009

    Matthew E Taylor and Peter Stone. Transfer learning for reinforcement learning domains: A survey.Journal of Machine Learning Research, 10(Jul):1633–1685, 2009

  11. [19]

    Model-agnostic meta-learning for fast adaptation of deep networks

    Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. InInternational Conference on Machine Learning, pages 1126–1135, 2017

  12. [20]

    The episodic buffer: a new com- ponent of working memory?Trends in Cognitive Sciences, 4(11):417–423, 2000

    Alan D Baddeley. The episodic buffer: a new com- ponent of working memory?Trends in Cognitive Sciences, 4(11):417–423, 2000

  13. [21]

    Darwin’s mistake: Explaining the discon- tinuity between human and nonhuman minds.Be- havioral and Brain Sciences, 31(2):109–130, 2008

    Derek C Penn, Keith J Holyoak, and Daniel J Povinelli. Darwin’s mistake: Explaining the discon- tinuity between human and nonhuman minds.Be- havioral and Brain Sciences, 31(2):109–130, 2008

  14. [22]

    How to grow a mind: Statistics, structure, and abstraction.Science, 331(6022):1279–1285, 2011

    Joshua B Tenenbaum, Charles Kemp, Thomas L Griffiths, and Noah D Goodman. How to grow a mind: Statistics, structure, and abstraction.Science, 331(6022):1279–1285, 2011

  15. [23]

    Does the chim- panzee have a theory of mind?Behavioral and Brain Sciences, 1(4):515–526, 1978

    David Premack and Guy Woodruff. Does the chim- panzee have a theory of mind?Behavioral and Brain Sciences, 1(4):515–526, 1978

  16. [24]

    Prospective memory: Theoretical considerations and opera- tional definitions.The Cognitive Neuroscience of Memory, pages 112–128, 2007

    Sam J Gilbert and Paul W Burgess. Prospective memory: Theoretical considerations and opera- tional definitions.The Cognitive Neuroscience of Memory, pages 112–128, 2007

  17. [25]

    Abstract representations of numbers in the animal and human brain.Trends in Neurosciences, 21(8):355–361, 1998

    Stanislas Dehaene, Ghislaine Dehaene-Lambertz, and Laurent Cohen. Abstract representations of numbers in the animal and human brain.Trends in Neurosciences, 21(8):355–361, 1998

  18. [26]

    The sensorimotor foundations of higher cognition

    Daniel M Wolpert, Zoubin Ghahramani, and J Ran- dall Flanagan. The sensorimotor foundations of higher cognition. InCommon Minds: Themes from the Philosophy of Philip Pettit. Oxford University Press, 2003

  19. [27]

    Three models for the description of language.IRE Transactions on Information Theory, 2(3):113–124, 1956

    Noam Chomsky. Three models for the description of language.IRE Transactions on Information Theory, 2(3):113–124, 1956

  20. [28]

    Planning and acting in partially observable stochastic domains.Artificial Intelli- gence, 101(1-2):99–134, 1998

    Leslie Pack Kaelbling, Michael L Littman, and An- thony R Cassandra. Planning and acting in partially observable stochastic domains.Artificial Intelli- gence, 101(1-2):99–134, 1998

  21. [29]

    Superhuman AI for multiplayer poker.Science, 365(6456):885– 890, 2019

    Noam Brown and Tuomas Sandholm. Superhuman AI for multiplayer poker.Science, 365(6456):885– 890, 2019

  22. [30]

    Continuous con- trol with deep reinforcement learning

    Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous con- trol with deep reinforcement learning. InInter- national Conference on Learning Representations (ICLR), 2016

  23. [31]

    Deep recur- rent q-learning for partially observable mdps

    Matthew Hausknecht and Peter Stone. Deep recur- rent q-learning for partially observable mdps. In 2015 AAAI Fall Symposium Series, 2015

  24. [32]

    Policy invariance under reward transforma- tions: Theory and application to reward shaping

    Andrew Y Ng, Daishi Harada, and Stuart Rus- sell. Policy invariance under reward transforma- tions: Theory and application to reward shaping. InInternational Conference on Machine Learning (ICML), pages 278–287, 1999

  25. [33]

    Exploration by random network dis- tillation

    Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov. Exploration by random network dis- tillation. InInternational Conference on Learning Representations (ICLR), 2019

  26. [34]

    A survey of multi- objective sequential decision-making.Journal of Artificial Intelligence Research, 48:67–113, 2013

    Diederik M Roijers, Peter Vamplew, Shimon White- son, and Richard Dazeley. A survey of multi- objective sequential decision-making.Journal of Artificial Intelligence Research, 48:67–113, 2013

  27. [35]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kel- ton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, a...

  28. [36]

    Openai gym.arXiv:1606.01540, 2016

    Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wo- jciech Zaremba. Openai gym.arXiv:1606.01540, 2016

  29. [37]

    SWE-bench: Can language models resolve real-world GitHub issues? InThe Twelfth International Conference on Learning Rep- resentations, 2024

    Carlos E Jimenez, John Yang Murphy, Alexander Shirinov, Kweon Chen, Austin McMillan, Guil- laume Lample, et al. SWE-bench: Can language models resolve real-world GitHub issues? InThe Twelfth International Conference on Learning Rep- resentations, 2024

  30. [38]

    We- bArena: A realistic web environment for building autonomous agents

    Shuyan Zhou, Frank F Hou, Yikang Cheng, Keisuke Hong, Graham Neubig, and Pengcheng Yin. We- bArena: A realistic web environment for building autonomous agents. InThe Twelfth International Conference on Learning Representations, 2024

  31. [39]

    A com- prehensive survey on safe reinforcement learning

    Javier García and Fernando Fernández. A com- prehensive survey on safe reinforcement learning. 35 An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments Journal of Machine Learning Research, 16(1):1437– 1480, 2015

  32. [40]

    On the measure of intelligence

    François Chollet. On the measure of intelligence. arXiv:1911.01547, 2019

  33. [41]

    Processbench: Identifying process errors in mathematical reasoning, 2025

    Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin. Processbench: Identifying process errors in mathematical reasoning, 2025

  34. [42]

    Reward is enough.Artificial Intelligence, 299:103535, 2021

    David Silver, Satinder Singh, Doina Precup, and Richard S Sutton. Reward is enough.Artificial Intelligence, 299:103535, 2021

  35. [43]

    Vizdoom: A doom-based ai research platform for visual re- inforcement learning

    Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Ja´skowski. Vizdoom: A doom-based ai research platform for visual re- inforcement learning. In2016 IEEE Conference on Computational Intelligence and Games (CIG), pages 1–8, 2016

  36. [44]

    Charles Beattie, Joel Z. Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, Julian Schrittwieser, Keith Anderson, Sarah York, Max Cant, Adam Cain, Adrian Bolton, Stephen Gaffney, Helen King, Demis Ha...

  37. [45]

    Minimalistic gridworld environment for OpenAI Gym

    Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal. Minimalistic gridworld environment for OpenAI Gym. GitHub repository, 2018

  38. [46]

    Habitat: A platform for embodied AI research

    Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Ma- lik, et al. Habitat: A platform for embodied AI research. InProceedings of the IEEE/CVF Interna- tional Conference on Computer Visio...

  39. [47]

    Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

    Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on Robot Learning, pages 1094–1100, 2020

  40. [48]

    ALFWorld: Aligning text and embod- ied environments for interactive learning

    Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. ALFWorld: Aligning text and embod- ied environments for interactive learning. InInter- national Conference on Learning Representations, 2021

  41. [49]

    Brax–a differentiable physics engine for large scale rigid body simulation

    C Daniel Freeman, Erik Frey, Anton Raichuk, Ser- tan Girgin, Igor Mordatch, and Olivier Bachem. Brax–a differentiable physics engine for large scale rigid body simulation. InThirty-Fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2021

  42. [50]

    Minedojo: Building open-ended embodied agents with internet-scale knowledge, 2022

    Linxi Fan, Guanzhi Wang, Yunfan Jiang, Ajay Man- dlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, and Anima Anandkumar. Minedojo: Building open-ended embodied agents with internet-scale knowledge, 2022

  43. [51]

    WebShop: Towards scalable real- world web interaction with grounded language agents

    Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. WebShop: Towards scalable real- world web interaction with grounded language agents. InAdvances in Neural Information Pro- cessing Systems, volume 35, pages 20744–20757, 2022

  44. [52]

    Training software engineering agents and verifiers with swe-gym, 2025

    Jiayi Pan, Xingyao Wang, Graham Neubig, Navdeep Jaitly, Heng Ji, Alane Suhr, and Yizhe Zhang. Training software engineering agents and verifiers with swe-gym, 2025

  45. [53]

    Sparks of artificial general intelli- gence: Early experiments with gpt-4, 2023

    Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. Sparks of artificial general intelli- gence: Early experiments with...

  46. [54]

    OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xi- aochuan Li, Ruiyuan Zhao, Ruisheng Cao, et al. OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. arXiv preprint arXiv:2404.07972, 2024

  47. [55]

    Android- world: A dynamic benchmarking environment for autonomous agents.arXiv:2405.14573, 2024

    Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy Lillicrap. Android- world: A dynamic benchmarking environment for autonomous agents.arXiv:2405.14573, 2024

  48. [56]

    Mle-bench: Evaluating machine learning agents on machine learning engineering.arXiv:2410.07095, 2024

    Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, Lilian Weng, and Aleksander M ˛ adry. Mle-bench: Evaluating machine learning agents on machine learning engineering.arXiv:2410.07095, 2024

  49. [57]

    Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021

  50. [58]

    Measuring mathematical problem solving with the math dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. InProceed- ings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2021

  51. [59]

    Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021

  52. [60]

    The lean 4 theorem prover and programming language

    Leonardo de Moura and Sebastian Ullrich. The lean 4 theorem prover and programming language. In Automated Deduction–CADE 28, pages 625–635, 2021. 36 An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments

  53. [61]

    miniF2F: A cross-system benchmark for for- mal olympiad-level mathematics

    Kunhao Zheng, Jesse Michael Han, and Stanislas Polu. miniF2F: A cross-system benchmark for for- mal olympiad-level mathematics. InInternational Conference on Learning Representations, 2022

  54. [62]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark, 2023

  55. [63]

    Live- codebench: Holistic and contamination free eval- uation of large language models for code, 2024

    Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. Live- codebench: Holistic and contamination free eval- uation of large language models for code, 2024. arXiv:2403.07974

  56. [64]

    Olympiad- bench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scien- tific problems

    Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, and Maosong Sun. Olympiad- bench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scien...

  57. [65]

    Tree of thoughts: Deliberate problem solving with large language models

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. InAdvances in Neural Information Processing Systems, volume 36, 2023

  58. [66]

    Deepseek-r1: Incentivizing reason- ing capability in llms via reinforcement learning

    DeepSeek-AI. Deepseek-r1: Incentivizing reason- ing capability in llms via reinforcement learning. arXiv:2501.12948, 2025

  59. [67]

    From local to global: A graph rag approach to query-focused summarization, 2025

    Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. From local to global: A graph rag approach to query-focused summarization, 2025

  60. [68]

    Can llm al- ready serve as a database interface? a big bench for large-scale database grounded text-to-sqls

    Jinyang Li, Binyuan Hui, Chengwei Qu, Binhua Li, Ruiying Geng, Bowen Li, Bailin Wang, Bowen Qin, Ruiyao Dong, Chenhao Zhang, et al. Can llm al- ready serve as a database interface? a big bench for large-scale database grounded text-to-sqls. InAd- vances in Neural Information P...

  61. [69]

    Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals

    Yujia Li, David Choi, Junyoung Chung, Nate Kush- man, Julian Schrittwieser, Rémi Leblond, Tom Ec- cles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Mas- son d’Autume, Igor Babuschkin, Xinyun Chen, Po- Sen Huang, Johannes Welbl, Sven Gow...

  62. [70]

    Visualwebarena: Evaluating multi- modal agents on realistic visual web tasks, 2024

    Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried. Visualwebarena: Evaluating multi- modal agents on realistic visual web tasks, 2024

  63. [71]

    Solving sokoban using hierarchi- cal reinforcement learning with landmarks, 2025

    Sergey Pastukhov. Solving sokoban using hierarchi- cal reinforcement learning with landmarks, 2025

  64. [72]

    Reasoning gym: Reason- ing environments for reinforcement learning with verifiable rewards, 2025

    Zafir Stojanovski, Oliver Stanley, Joe Sharratt, Richard Jones, Abdulhakeem Adefioye, Jean Kad- dour, and Andreas Köpf. Reasoning gym: Reason- ing environments for reinforcement learning with verifiable rewards, 2025

  65. [73]

    Athena scientific, 2012

    Dimitri P Bertsekas.Dynamic programming and optimal control: Vol I. Athena scientific, 2012

  66. [74]

    Foerster, Yannis M

    Jakob N. Foerster, Yannis M. Assael, Nando de Fre- itas, and Shimon Whiteson. Learning to communi- cate with deep multi-agent reinforcement learning, 2016

  67. [75]

    A general reinforcement learn- ing algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140–1144, 2018

    David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis. A general reinforcement learn- ing algorithm that masters chess...

  68. [76]

    Multi-agent actor-critic for mixed cooperative-competitive environments, 2020

    Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. Multi-agent actor-critic for mixed cooperative-competitive environments, 2020

  69. [77]

    Magent: A many- agent reinforcement learning platform for artificial collective intelligence, 2017

    Lianmin Zheng, Jiacheng Yang, Han Cai, Weinan Zhang, Jun Wang, and Yong Yu. Magent: A many- agent reinforcement learning platform for artificial collective intelligence, 2017

  70. [78]

    Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H

    Nolan Bard, Jakob N. Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H. Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, Iain Dunning, Shibl Mourad, Hugo Larochelle, Marc G. Bellemare, and Michael Bowl- ing. The hanabi challenge: A new fro...

  71. [79]

    Application of self-play rein- forcement learning to a four-player game of imper- fect information, 2018

    Henry Charlesworth. Application of self-play rein- forcement learning to a four-player game of imper- fect information, 2018

  72. [80]

    Google re- search football: A novel reinforcement learning en- vironment

    Karol Kurach, Anton Raichuk, Piotr Sta ´nczyk, Michał Zaj ˛ ac, Olivier Bachem, Lasse Espeholt, Car- los Riquelme, Damien Vincent, Marcin Michalski, Olivier Bousquet, and Sylvain Gelly. Google re- search football: A novel reinforcement learning en- vironment. InProceedings of ...

  73. [81]

    37 An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments On the utility of learning about humans for human- ai coordination

    Micah Carroll, Rohin Shah, Mark K Ho, Tom Grif- fiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan. 37 An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments On the utility of learning about humans for human- ai coordination. InAdv...

  74. [82]

    Albrecht

    Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V . Albrecht. Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks, 2021

  75. [83]

    Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks.arXiv:2006.07869, 2020

    Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V Albrecht. Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks.arXiv:2006.07869, 2020

  76. [84]

    Turner, Oleksandr Maksymets, Zsolt Kira, Mrinal Kalakr- ishnan, Jitendra Malik, Devendra Singh Chaplot, Unnat Jain, Dhruv Batra, Akshara Rai, and Roozbeh Mottaghi

    Xavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote, Tsung-Yen Yang, Ruslan Part- sey, Ruta Desai, Alexander William Clegg, Michal Hlavac, So Yeon Min, Vladimír V ondruš, Theophile Gervet, Vincent-Pierre Berges, John M. Turner, Oleksandr Maksymets, Zsolt Kira, Mri...

  77. [85]

    Vmas: A vectorized multi-agent simulator for col- lective robotics.IEEE Robotics and Automation Letters, 7(2):5323–5330, 2022

    Matteo Bettini, Ryan Corsi, and Amanda Prorok. Vmas: A vectorized multi-agent simulator for col- lective robotics.IEEE Robotics and Automation Letters, 7(2):5323–5330, 2022

  78. [86]

    Isaac gym: High performance gpu-based physics simulation for robot learning.arXiv:2108.10470, 2021

    Viktor Makoviychuk, Lukasz Wawrzyniak, Yuriy Gavrilov, Anton Korobov, Ray Lu, Sho Tsushima, Gavriel State, Miles Macklin, and Ankur Handa. Isaac gym: High performance gpu-based physics simulation for robot learning.arXiv:2108.10470, 2021

  79. [87]

    J. K. Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar, Ananth Hari, Ryan Sullivan, Luis Santos, Rodrigo Perez, Caroline Horsch, Clemens Dieffendahl, Niall L. Williams, Yashas Lokesh, and Praveen Ravi. Pettingzoo: Gym for multi-agent reinforcement learning, 2021

  80. [88]

    Jaxmarl: Multi-agent rl environments and algorithms in jax, 2024

    Alexander Rutherford, Benjamin Ellis, Matteo Gal- lici, Jonathan Cook, Andrei Lupu, Gardar Ingvars- son, Timon Willi, Ravi Hammond, Akbir Khan, Christian Schroeder de Witt, Alexandra Souly, Saptarashmi Bandyopadhyay, Mikayel Samvelyan, Minqi Jiang, Robert Tjarko Lange, Shimon ...

  81. [89]

    O’Brien, Carrie J

    Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior, 2023

  82. [90]

    Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors, 2023

    Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu, Maosong Sun, and Jie Zhou. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors, 2023

  83. [91]

    Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xi- awu Zheng, Yuheng Cheng, Ceyao Zhang, Jin- lin Wang, Zili Wang, Steven S. K. Yau, Zijian Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. Metagpt: Meta programming for a multi-agent collaborative f...

  84. [92]

    Mindagent: Emergent gaming inter- action, 2023

    Ran Gong, Qiuyuan Huang, Xiaojian Ma, Hoi V o, Zane Durante, Yusuke Noda, Zilong Zheng, Song- Chun Zhu, Demetri Terzopoulos, Li Fei-Fei, and Jianfeng Gao. Mindagent: Emergent gaming inter- action, 2023

  85. [93]

    A novel approach to chest x-ray lung segmentation using u-net and modified convolu- tional block attention module, 2024

    Mohammad Ali Labbaf Khaniki and Mohammad Manthouri. A novel approach to chest x-ray lung segmentation using u-net and modified convolu- tional block attention module, 2024

  86. [94]

    Tau-bench: A benchmark for tool-agent-user interaction in real-world domains

    Shun Fang, Zhenyu Zhang, Zhaowei Hou, Lingyao Li, Ning Bian, Jinjie Ni, Xiaoliang Wang, and Renrui Zhang. Tau-bench: A benchmark for tool-agent-user interaction in real-world domains. arXiv:2406.12045, 2024

  87. [95]

    Sweet-rl: Training multi-turn llm agents on collaborative reasoning tasks, 2025

    Yifei Zhou, Song Jiang, Yuandong Tian, Jason We- ston, Sergey Levine, Sainbayar Sukhbaatar, and Xian Li. Sweet-rl: Training multi-turn llm agents on collaborative reasoning tasks, 2025

  88. [96]

    AMEX: Android multi- annotation expo dataset for mobile GUI agents

    Yuxiang Chai, Siyuan Huang, Yazhe Niu, Han Xiao, Liang Liu, Guozhi Wang, Dingyu Zhang, Shuai Ren, and Hongsheng Li. AMEX: Android multi- annotation expo dataset for mobile GUI agents. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Findi...

  89. [97]

    Association for Computational Linguistics

  90. [98]

    Reso: A reward-driven self-organizing llm-based multi-agent system for reasoning tasks, 2025

    Heng Zhou, Hejia Geng, Xiangyuan Xue, Li Kang, Yiran Qin, Zhiyong Wang, Zhenfei Yin, and Lei Bai. Reso: A reward-driven self-organizing llm-based multi-agent system for reasoning tasks, 2025

  91. [99]

    Hima-ecom: Enabling joint training of hierarchical multi-agent e-commerce as- sistants, 2026

    Junxing Hu, Ai Han, Haolan Zhan, Pu Wei, Zhiqian Zhang, Yuhang Guo, Jiawei Lu, Zhen Chen, Haoran Li, and Zicheng Zhang. Hima-ecom: Enabling joint training of hierarchical multi-agent e-commerce as- sistants, 2026

  92. [100]

    Textarena, 2025

    Leon Guertler, Bobby Cheng, Simon Yu, Bo Liu, Leshem Choshen, and Cheston Tan. Textarena, 2025

  93. [101]

    Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi- turn reinforcement learning, 2026

    Bo Liu, Leon Guertler, Simon Yu, Zichen Liu, Penghui Qi, Daniel Balcells, Mickel Liu, Cheston Tan, Weiyan Shi, Min Lin, Wee Sun Lee, and Natasha Jaques. Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi- turn reinforcement learning, 2026. 38 An E...

  94. [102]

    aha mo- ments

    Xiaoqing Zhang, Huabin Zheng, Ang Lv, Yuhan Liu, Zirui Song, Xiuying Chen, Rui Yan, and Flood Sung. Divide-fuse-conquer: Eliciting "aha mo- ments" in multi-scenario games, 2025

  95. [103]

    Synlogic: Synthesizing verifiable reasoning data at scale for learning logical reasoning and beyond, 2025

    Junteng Liu, Yuanxiang Fan, Zhuo Jiang, Han Ding, Yongyi Hu, Chi Zhang, Yiqi Shi, Shitong Weng, Aili Chen, Shiqi Chen, Yunan Huang, Mozhi Zhang, Pengyu Zhao, Junjie Yan, and Junxian He. Synlogic: Synthesizing verifiable reasoning data at scale for learning logical reasoning an...

  96. [104]

    Measuring general intelligence with generated games, 2025

    Vivek Verma, David Huang, William Chen, Dan Klein, and Nicholas Tomlin. Measuring general intelligence with generated games, 2025

  97. [105]

    Marft: Multi-agent reinforcement fine- tuning, 2025

    Junwei Liao, Muning Wen, Jun Wang, and Weinan Zhang. Marft: Multi-agent reinforcement fine- tuning, 2025

  98. [106]

    Stochastic games.Proceedings of the national academy of sciences, 39(10):1095– 1100, 1953

    Lloyd S Shapley. Stochastic games.Proceedings of the national academy of sciences, 39(10):1095– 1100, 1953

  99. [107]

    Markov games as a framework for multi-agent reinforcement learning

    Michael L Littman. Markov games as a framework for multi-agent reinforcement learning. InMachine Learning Proceedings 1994, pages 157–163, 1994

  100. [108]

    Mikayel Samvelyan, Tabish Rashid, Chris- tian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philip H. S. Torr, Jakob Foerster, and Shimon Whiteson. The starcraft multi-agent challenge. In Proceedings of the 18th International Conference...

  101. [109]

    Qmix: Monotonic value func- tion factorisation for deep multi-agent reinforce- ment learning

    Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. Qmix: Monotonic value func- tion factorisation for deep multi-agent reinforce- ment learning. InInternational Conference on Ma- chine Learning, pages 4295–4304. PMLR, 2018

  102. [110]

    Improving fac- tuality and reasoning in language models through multiagent debate.arXiv:2305.14325, 2023

    Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch. Improving fac- tuality and reasoning in language models through multiagent debate.arXiv:2305.14325, 2023

  103. [111]

    Llm collaboration with multi-agent reinforcement learning, 2025

    Shuo Liu, Tianle Chen, Zeyu Liang, Xueguang Lyu, and Christopher Amato. Llm collaboration with multi-agent reinforcement learning, 2025

  104. [112]

    Alfred: A bench- mark for interpreting grounded instructions for ev- eryday tasks

    Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. Alfred: A bench- mark for interpreting grounded instructions for ev- eryday tasks. InProceedings of the IEEE/CVF con- ference on computer vision and pat...

  105. [113]

    Attention, learn to solve routing problems! InInter- national Conference on Learning Representations, 2019

    Wouter Kool, Herke van Hoof, and Max Welling. Attention, learn to solve routing problems! InInter- national Conference on Learning Representations, 2019

  106. [114]

    Deep- mind control suite, 2018

    Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy Lillicrap, and Martin Riedmiller. Deep- mind control suite, 2018

  107. [115]

    robosuite: A modular simulation framework and benchmark for robot learning.arXiv:2009.12293, 2020

    Yuke Zhu, Josiah Wong, Ajay Mandlekar, Roberto Martín-Martín, Silvio Savarese, and Li Fei-Fei. robosuite: A modular simulation framework and benchmark for robot learning.arXiv:2009.12293, 2020

  108. [116]

    Quantifying generaliza- tion in reinforcement learning

    Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman. Quantifying generaliza- tion in reinforcement learning. In Kamalika Chaud- huri and Ruslan Salakhutdinov, editors,Proceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings...

  109. [117]

    The animal-ai environment: Train- ing and testing animal-like artificial cognition, 2019

    Benjamin Beyret, José Hernández-Orallo, Lucy Cheke, Marta Halina, Murray Shanahan, and Matthew Crosby. The animal-ai environment: Train- ing and testing animal-like artificial cognition, 2019

  110. [118]

    Model-based reinforcement learning for atari, 2024

    Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Pi- otr Kozakowski, Sergey Levine, Afroz Mohiud- din, Ryan Sepassi, George Tucker, and Henryk Michalewski. Model-based reinforcement learning for a...

  111. [119]

    The nethack learning environment.Advances in Neural Information Pro- cessing Systems, 33:7671–7684, 2020

    Heinrich Küttler, Nantas Nardelli, Alexander Miller, Roberta Raileanu, Marco Selig, Edward Grefen- stette, and Tim Rocktäschel. The nethack learning environment.Advances in Neural Information Pro- cessing Systems, 33:7671–7684, 2020

  112. [120]

    Benchmarking the spectrum of agent capabilities, 2022

    Danijar Hafner. Benchmarking the spectrum of agent capabilities, 2022

  113. [121]

    Babyai: A platform to study the sample efficiency of grounded language learning, 2019

    Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio. Babyai: A platform to study the sample efficiency of grounded language learning, 2019

  114. [122]

    Interactive fiction games: A colossal adventure, 2020

    Matthew Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, and Xingdi Yuan. Interactive fiction games: A colossal adventure, 2020

  115. [123]

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan H...

  116. [124]

    Toolllm: Fa- cilitating large language models to master 16000+ real-world apis

    Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Rian Lauren, Ruobing Tian, Ruoko Xie, Jie Zhou, Mark Gerstein, Dahua Li, Zhiyuan Liu, and Maosong Sun. Toolllm: Fa- cilitating large language models to ...

  117. [125]

    Mmlu-pro: A more robust and challenging multi-task language under- standing benchmark, 2024

    Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. Mmlu-pro: A more robust and challenging multi-task language under-...

  118. [126]

    Supergpqa: Scaling llm evaluation across 285 graduate disci- plines, 2025

    P Team, Xinrun Du, Yifan Yao, Kaijing Ma, Bingli Wang, Tianyu Zheng, King Zhu, Minghao Liu, Yim- ing Liang, Xiaolong Jin, Zhenlin Wei, Chujie Zheng, Kaixin Deng, Shawn Gavin, Shian Jia, Sichao Jiang, Yiyan Liao, Rui Li, Qinrui Li, Sirun Li, Yizhi Li, Yunwen Li, David Ma, Yuans...

  119. [127]

    Deepscaler: Surpassing o1- preview with a 1.5b model by scaling rl, 2025

    Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Tang, Manan Roongta, Colin Cai, Jef- frey Luo, Tianjun Zhang, Erran Li, Raluca Ada Popa, and Ion Stoica. Deepscaler: Surpassing o1- preview with a 1.5b model by scaling rl, 2025. No- tion Blog

  120. [128]

    Deepmind control suite.arXiv:1801.00690, 2018

    Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, et al. Deepmind control suite.arXiv:1801.00690, 2018

  121. [129]

    Soundspaces: Audio-visual navigation in 3d environments

    Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Santhosh Ramakrishnan, and Kristen Grauman. Soundspaces: Audio-visual navigation in 3d environments. In European Conference on Computer Vision, pages 17–36, 2020

  122. [130]

    Finrl: A deep reinforcement learning library for automated stock trading in quantitative finance

    Xiao-Yang Liu, Hongyang Yang, Qian Chen, Runjia Zhang, Linyun Yang, Bowen Xiao, and William Wang. Finrl: A deep reinforcement learning library for automated stock trading in quantitative finance. arXiv:2011.09607, 2020

  123. [131]

    Deep reinforcement learning for sepsis treatment

    Aniruddh Raghu, Matthieu Komorowski, Leo A Celi, Peter Szolovits, and Marzyeh Ghassemi. Deep reinforcement learning for sepsis treatment. InMa- chine Learning for Healthcare Conference, pages 174–182, 2017

  124. [132]

    Chartqa: A benchmark for question answering about charts with visual and logical reasoning, 2022

    Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. Chartqa: A benchmark for question answering about charts with visual and logical reasoning, 2022

  125. [133]

    Mind2web: Towards a generalist agent for the web

    Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su. Mind2web: Towards a generalist agent for the web. InAdvances in Neural Information Processing Sys- tems, volume 36, 2023

  126. [134]

    Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

    Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

  127. [135]

    Ferret: Refer and ground anything anywhere at any granularity

    Haoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du, Bowen Zhang, Zirui Wang, Liang Cao, Shih- Fu Chang, and Yinfei Yang. Ferret: Refer and ground anything anywhere at any granularity. In International Conference on Learning Representa- tions, 2024

  128. [136]

    Mmmu: A massive multi-discipline multimodal un- derstanding and reasoning benchmark for expert agi, 2024

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wen- hao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A...

  129. [137]

    V-irl: Grounding virtual in- telligence in real life

    Jihan Yang, Runyu Ding, Ellis Brown, Xiaojuan Qi, and Saining Xie. V-irl: Grounding virtual in- telligence in real life. InEuropean Conference on Computer Vision (ECCV), 2024

  130. [138]

    Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis, 2025

    Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Caifeng Shan, Ran He, and Xing Sun. Video-mme: The first-ev...

  131. [139]

    Carla: An open urban driving simulator

    Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. Carla: An open urban driving simulator. InConference on Robot Learning, pages 1–16, 2017

  132. [140]

    Ai2- thor: An interactive 3d environment for visual ai

    Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi. Ai2- thor: An interactive 3d environment for visual ai. arXiv:1712.05474, 2017

  133. [141]

    Virtualhome: Simulating household activities via programs, 2018

    Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba. Virtualhome: Simulating household activities via programs, 2018

  134. [142]

    igibson 1.0: a simulation environment for interactive tasks in large realistic scenes

    Bokui Shen, Fei Xia, Chengshu Li, Roberto Martín- Martín, Linxi Fan, Guanzhi Wang, Shyamal Buch, Claudia D’Arpino, Sanjana Srivastava, Lyne P Tchapmi, Micael E Tchapmi, Kent Vainio, Li Fei- Fei, and Silvio Savarese. igibson 1.0: a simulation environment for interactive tasks i...

  135. [143]

    Man- iskill2: A unified benchmark for generalizable ma- nipulation skills

    Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, Xiaodi Yuan, Pengwei Xie, Zhiao Huang, Rui Chen, and Hao Su. Man- iskill2: A unified benchmark for generalizable ma- nipulation skills. InInternational Conf...

  136. [144]

    Libero: Bench- marking knowledge transfer for lifelong robot learn- ing

    Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Biao, Qiang Zhang, Zhenjia Xu, Jianwei Zhang, Aliang Guo, Yuke Zhu, and Peter Stone. Libero: Bench- marking knowledge transfer for lifelong robot learn- ing. InAdvances in Neural Information Processing Systems, volume 36, 2023

  137. [145]

    Zhao, and Chelsea Finn

    Zipeng Fu, Tony Z. Zhao, and Chelsea Finn. Mobile aloha: Learning bimanual mobile manip- ulation with low-cost whole-body teleoperation. arXiv:2401.02117, 2024

  138. [146]

    Loos, Markus N

    Kshitij Bansal, Sarah M. Loos, Markus N. Rabe, Christian Szegedy, and Stewart Wilcox. Holist: An environment for machine learning of higher-order theorem proving, 2019

  139. [147]

    Compilergym: robust, performant compiler optimization environ- ments for ai research

    Chris Cummins, Bram Wasti, Jiadong Guo, Bran- don Cui, Jason Ansel, Sahir Gomez, Shobha Murali, Hugh Leather, and Yuandong Tian. Compilergym: robust, performant compiler optimization environ- ments for ai research. InProceedings of the 2022 IEEE/ACM International Symposium on ...

  140. [148]

    Language models are greedy reasoners: A systematic formal analysis of chain-of-thought, 2023

    Abulhair Saparov and He He. Language models are greedy reasoners: A systematic formal analysis of chain-of-thought, 2023

  141. [149]

    Intercode: Standardizing and benchmarking interactive coding with execution feedback, 2023

    John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao. Intercode: Standardizing and benchmarking interactive coding with execution feedback, 2023

  142. [150]

    Verilogeval: Evaluating large language models for verilog code generation

    Mingjie Liu, Nathaniel Pinckney, Brucek Khailany, and Haoxing Ren. Verilogeval: Evaluating large language models for verilog code generation. arXiv:2309.07554, 2023

  143. [151]

    Opencodeinterpreter: Integrat- ing code generation with execution and refinement

    Tianyu Zheng, Ge Zhang, Tianqi Shen, Xuel- ing Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue. Opencodeinterpreter: Integrat- ing code generation with execution and refinement. arXiv:2402.14658, 2024

  144. [152]

    Optimization of molecules via deep reinforcement learning.Scientific reports, 9(1):10752, 2019

    Zhenpeng Zhou, Steven Kearnes, Li Li, Richard N Zare, and Patrick Riley. Optimization of molecules via deep reinforcement learning.Scientific reports, 9(1):10752, 2019

  145. [153]

    Citylearn: Standardizing research in multi-agent reinforcement learning for demand response and urban energy management, 2020

    Jose R Vazquez-Canteli, Sourav Dey, Gregor Henze, and Zoltan Nagy. Citylearn: Standardizing research in multi-agent reinforcement learning for demand response and urban energy management, 2020

  146. [154]

    Trans- fer learning enables predictions in network biology

    Christina V Theodoris, Ling Xiao, Anant Chopra, Mark D Chaffin, Zeina R Al Sayed, Matthew C Hill, Helene Mantineo, Elizabeth M Brydon, Zexian Zeng, X Shirley Liu, and Patrick T Ellinor. Trans- fer learning enables predictions in network biology. Nature, 618(7965):616–624, 2023

  147. [155]

    Large lan- guage models encode clinical knowledge.Nature, 620(7972):172–180, 2023

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Kannan, Philip Mansfield, Michal Lukasik, Bradley Green, Ewa Dominowska, Blaise Aguera y Arcas, Joelle Barral, Dale Schuurmans, Ivan Kashnitsky, and Vivek Natarajan. Large l...

  148. [156]

    Solving olympiad geometry with- out human demonstrations.Nature, 625(7995):476– 482, 2024

    Trieu H Trinh, Yuhuai Wu, Quoc V Le, He He, and Thang Luong. Solving olympiad geometry with- out human demonstrations.Nature, 625(7995):476– 482, 2024

  149. [157]

    Pokorny, Xiao Huang, and Xinrun Wang

    Chang Yang, Ruiyu Wang, Junzhe Jiang, Qi Jiang, Qinggang Zhang, Yanchen Deng, Shuxin Li, Shuyue Hu, Bo Li, Florian T. Pokorny, Xiao Huang, and Xinrun Wang. Nondeterministic polynomial- time problem challenge: An ever-scaling reasoning benchmark for llms.arxiv:2504.11239, 2025

  150. [158]

    Le, James Laudon, Richard Ho, Roger Carpenter, and Jeff Dean

    Azalia Mirhoseini, Anna Goldie, Mustafa Yazgan, Joe Wenjie Jiang, Ebrahim Songhori, Shen Wang, Young-Joon Lee, Eric Johnson, Omkar Pathak, Azade Nazi, Jiwoo Pak, Andy Tong, Kavya Srini- vasa, William Hang, Emre Tuncer, Quoc V . Le, James Laudon, Richard Ho, Roger Carpenter, an...

  151. [159]

    Dream to control: Learning behaviors by latent imagination

    Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination. InInternational Conference on Learning Representations, 2019

  152. [160]

    An investigation of model-free plan- ning

    Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sébastien Racanière, Theophane Weber, David Raposo, Adam Santoro, Oriol Vinyals, and 41 An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments David Silver. An investigation of m...

  153. [161]

    A survey of monte carlo tree search methods.IEEE Trans- actions on Computational Intelligence and AI in games, 4(1):1–43, 2012

    Cameron B Browne, Edward Powley, Daniel White- house, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyri- don Samothrakis, and Simon Colton. A survey of monte carlo tree search methods.IEEE Trans- actions on Computational Intelligence and ...

  154. [162]

    Lean- dojo: Theorem proving with retrieval-augmented language models

    Kaiyu Yang, Aidan M Swope, Alex Gu, Rahul Cha- lamala, Peiyang Song, Shixing Yu, Saad Godil, Ryan Prenger, and Anima Anandkumar. Lean- dojo: Theorem proving with retrieval-augmented language models. InAdvances in Neural Informa- tion Processing Systems, volume 36, 2023

  155. [163]

    Chain-of-thought prompting elicits reasoning in large language mod- els

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language mod- els. InAdvances in Neural Information Processing Systems, volume 35, pages 24824–24837, 2022

  156. [164]

    Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645:633–638, 2025

    Daya Guo et al. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645:633–638, 2025

  157. [165]

    Wang, Michael King, Nicolas Porcel, Zeb Kurth-Nelson, Tina Zhu, Charlie Deck, Pe- ter Choy, Mary Cassin, Malcolm Reynolds, Fran- cis Song, Gavin Buttimore, David P

    Jane X. Wang, Michael King, Nicolas Porcel, Zeb Kurth-Nelson, Tina Zhu, Charlie Deck, Pe- ter Choy, Mary Cassin, Malcolm Reynolds, Fran- cis Song, Gavin Buttimore, David P. Reichert, Neil Rabinowitz, Loic Matthey, Demis Hass- abis, Alexander Lerchner, and Matthew Botvinick. Al...

  158. [166]

    Evaluating long-term memory in 3d mazes, 2022

    Jurgis Pasukonis, Timothy Lillicrap, and Danijar Hafner. Evaluating long-term memory in 3d mazes, 2022

  159. [167]

    Benchmarking Safe Exploration in Deep Reinforce- ment Learning

    Alex Ray, Joshua Achiam, and Dario Amodei. Benchmarking Safe Exploration in Deep Reinforce- ment Learning. 2019

  160. [168]

    D4rl: Datasets for deep data-driven reinforcement learning, 2021

    Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine. D4rl: Datasets for deep data-driven reinforcement learning, 2021

  161. [169]

    Solving the rubik’s cube with deep reinforcement learning and search

    Forest Agostinelli, Stephen McAleer, Alexander Shmakov, and Pierre Baldi. Solving the rubik’s cube with deep reinforcement learning and search. Nature Machine Intelligence, 1(8):356–363, 2019

  162. [170]

    Emergent tool use from multi-agent autocur- ricula

    Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mor- datch. Emergent tool use from multi-agent autocur- ricula. InInternational Conference on Learning Representations, 2020

  163. [171]

    Superhuman ai for heads-up no-limit poker: Libratus beats top professionals.Science, 359(6374):418–424, 2018

    Noam Brown and Tuomas Sandholm. Superhuman ai for heads-up no-limit poker: Libratus beats top professionals.Science, 359(6374):418–424, 2018

  164. [172]

    Miller, Sasha Mitts, Adithya Rendleman, Stephen Roller, Dirk Rowe, Jared Salter, Kurt Shus- ter, Michael Spisak, Alexander Wei, David Wu, Hugh Zhang, and Markus Zijlstra

    Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexan- der H. Miller, Sasha Mitts, Adithya Rendleman, Stephen...

  165. [173]

    Ex- ploring large language models for communica- tion games: An empirical study on werewolf

    Yuzhuang Xu, Shuo Wang, Peng Li, Fuwen Luo, Xiaolong Wang, Weidong Liu, and Yang Liu. Ex- ploring large language models for communica- tion games: An empirical study on werewolf. arXiv:2309.04658, 2023

  166. [174]

    Preparing for the unknown: Learning a universal policy with online system identification

    Wenhao Yu, Jie Tan, C Karen Liu, and Greg Turk. Preparing for the unknown: Learning a universal policy with online system identification. arXiv:1702.02453, 2017

  167. [175]

    Metadrive: Composing diverse driving scenarios for gener- alizable reinforcement learning.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 45(3):3461–3475, 2022

    Quanyi Li, Zhenghao Peng, Lan Feng, Qihang Zhang, Zhenghai Xue, and Bolei Zhou. Metadrive: Composing diverse driving scenarios for gener- alizable reinforcement learning.IEEE Transac- tions on Pattern Analysis and Machine Intelligence, 45(3):3461–3475, 2022

  168. [176]

    Roboballet: Planning for multirobot reaching with graph neural networks and reinforcement learning

    Matthew Lai, Keegan Go, Zhibin Li, Torsten Kröger, Stefan Schaal, Kelsey Allen, and Jonathan Scholz. Roboballet: Planning for multirobot reaching with graph neural networks and reinforcement learning. Science Robotics, 10(106), September 2025

  169. [177]

    Magnetic control of tokamak plasmas through deep reinforce- ment learning.Nature, 602(7897):414–419, 2022

    Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de las Casas, et al. Magnetic control of tokamak plasmas through deep reinforce- ment learning.Nature, 602(7897):414–419, 2022

  170. [178]

    Scicode: A research coding benchmark curated by scientists, 2024

    Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, Shengyan Liu, Di Luo, Yutao Ma, Hao Tong, Kha Trinh, Chenyu Tian, Zihan Wang, Bohao Wu, Yanyu Xiong, Shengzhu Yin, Minhui Zhu, Kilian Lieret, Yanx...

  171. [179]

    The artificial intelligence clinician learns optimal treatment strate- gies for sepsis in intensive care.Nature Medicine, 24(11):1716–1720, 2018

    Matthieu Komorowski, Leo A Celi, Omar Badawi, Anthony C Gordon, and A Aldo Faisal. The artificial intelligence clinician learns optimal treatment strate- gies for sepsis in intensive care.Nature Medicine, 24(11):1716–1720, 2018

  172. [180]

    Kegg: Kyoto encyclopedia of genes and genomes.Nucleic Acids Research, 28(1):27–30, 2000

    Minoru Kanehisa and Susumu Goto. Kegg: Kyoto encyclopedia of genes and genomes.Nucleic Acids Research, 28(1):27–30, 2000. 42 An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments

  173. [181]

    What dis- ease does this patient have? a large-scale open do- main question answering dataset from medical ex- ams.Applied Sciences, 11(14):6421, 2021

    Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What dis- ease does this patient have? a large-scale open do- main question answering dataset from medical ex- ams.Applied Sciences, 11(14):6421, 2021

  174. [182]

    Medxpertqa: Benchmarking expert-level medical reasoning and understanding,

    Yuxin Zuo, Shang Qu, Yifei Li, Zhangren Chen, Xuekai Zhu, Ermo Hua, Kaiyan Zhang, Ning Ding, and Bowen Zhou. Medxpertqa: Benchmarking expert-level medical reasoning and understanding,

  175. [183]

    Train- ing llms for ehr-based reasoning tasks via reinforce- ment learning, 2025

    Jiacheng Lin, Zhenbang Wu, and Jimeng Sun. Train- ing llms for ehr-based reasoning tasks via reinforce- ment learning, 2025. arXiv:2505.24105

  176. [184]

    Let’s verify step by step.arXiv:2305.20050, 2023

    Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step.arXiv:2305.20050, 2023

  177. [185]

    Deepseek- math: Pushing the limits of mathematical reasoning in open language models.arXiv:2402.03300, 2024

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Jiashuo Wang, Mingchuan Zhang, Yuqi Qiao, Zhengxu Qiao, Jianbing Dong, Fengji Zhang, Zhenda Xie, Bowen Obeng, Huanyu Liu, Wenhai Wang, Xingjian Shi, Renrui Zhang, et al. Deepseek- math: Pushing the limits of mathematical reasonin...

  178. [186]

    Gym-anytrading: Anytrading is a col- lection of openai gym environments for reinforce- ment learning-based trading algorithms

    Amin Alaee. Gym-anytrading: Anytrading is a col- lection of openai gym environments for reinforce- ment learning-based trading algorithms. https:// github.com/AminHp/gym-anytrading, 2018

  179. [187]

    Qlib: An ai-oriented quantitative investment platform.arXiv:2009.11189, 2020

    Xiao Yang, Weiqing Liu, Dong Zhou, Jiang Bian, and Tie-Yan Liu. Qlib: An ai-oriented quantitative investment platform.arXiv:2009.11189, 2020

  180. [188]

    Exploring the limitations of behavior cloning for autonomous driving.Inter- national Conference on Computer Vision(ICCV), 2019

    Felipe Codevilla, Eder Santana, Antonio M López, and Adrien Gaidon. Exploring the limitations of behavior cloning for autonomous driving.Inter- national Conference on Computer Vision(ICCV), 2019

  181. [189]

    Chang, Leonidas J

    Fanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia, Hao Zhu, Fangchen Liu, Minghua Liu, Hanxiao Jiang, Yifu Yuan, He Wang, Li Yi, Angel X. Chang, Leonidas J. Guibas, and Hao Su. SAPIEN: A sim- ulated part-based interactive environment. InThe IEEE Conference on Computer Vision and ...

  182. [190]

    Finqa: A dataset of numerical reasoning over financial data, 2022

    Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. Finqa: A dataset of numerical reasoning over financial data, 2022

  183. [191]

    Sara Mahdavi, Joelle Bar- ral, Dale Webster, Greg S

    Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Aguera y ...

  184. [192]

    Proximal policy optimization algorithms.arXiv:1707.06347, 2017

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms.arXiv:1707.06347, 2017

  185. [193]

    Soft actor-critic: Off-policy max- imum entropy deep reinforcement learning with a stochastic actor

    Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy max- imum entropy deep reinforcement learning with a stochastic actor. InInternational Conference on Machine Learning, pages 1861–1870, 2018

  186. [194]

    Openai o1 system card, 2024

    OpenAI. Openai o1 system card, 2024. https://openai.com/index/openai-o1-system-card/

  187. [195]

    Seephys: Does seeing help thinking? – benchmarking vision- based physics reasoning, 2025

    Kun Xiang, Heng Li, Terry Jingchen Zhang, Yinya Huang, Zirong Liu, Peixin Qu, Jixi He, Jiaqi Chen, Yu-Jie Yuan, Jianhua Han, Hang Xu, Hanhui Li, Mrinmaya Sachan, and Xiaodan Liang. Seephys: Does seeing help thinking? – benchmarking vision- based physics reasoning, 2025

  188. [196]

    Mastering diverse domains through world models.arXiv:2301.04104, 2023

    Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models.arXiv:2301.04104, 2023

  189. [197]

    A generalist agent, 2022

    Scott Reed, Konrad Zolna, Emilio Parisotto, Ser- gio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol V...

  190. [198]

    Human-timescale adap- tation in an open-ended task space

    Jakob Bauer, Kate Baumli, Satinder Baveja, Fer- yal Behbahani, Avishkar Bhoopchand, Nathalie Bradley-Schmieg, Michael Chang, Natalie Clay, Adrian Collister, Vibhavari Dasagi, Lucy Gon- zalez, Karol Gregor, Edward Hughes, Sheleem Kashem, Maria Loks-Thompson, Hannah Open- shaw, ...

  191. [199]

    Building a subspace of policies for scal- able continual learning.arXiv:2211.10445, 2022

    Jean-Baptiste Gaya, Thang Doan, Lucas Caccia, Laure Soulier, Ludovic Denoyer, and Roberta Raileanu. Building a subspace of policies for scal- able continual learning.arXiv:2211.10445, 2022

  192. [200]

    Agentbench: Evaluating llms as agents, 2023

    Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xu- anyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kai- wen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Min- lie Huang, Yuxiao Dong, and Jie Tang. Agent...

  193. [201]

    Tool-star: Empowering llm- brained multi-tool reasoner via reinforcement learn- ing.arXiv:2505.16410, 2025

    Guanting Dong, Yifei Chen, Zhicheng Dou, Yux- uan Wang, et al. Tool-star: Empowering llm- brained multi-tool reasoner via reinforcement learn- ing.arXiv:2505.16410, 2025

  194. [202]

    Sws: Self-aware weakness- driven problem synthesis in reinforcement learning for llm reasoning.arXiv:2506.08989, 2025

    Xiao Liang, Zhong-Zhi Li, Yeyun Gong, Yang Wang, Hengyuan Zhang, Yelong Shen, Ying Nian Wu, and Weizhu Chen. Sws: Self-aware weakness- driven problem synthesis in reinforcement learning for llm reasoning.arXiv:2506.08989, 2025

  195. [203]

    Llada 1.5: Variance-reduced preference optimization for large language diffusion models.arXiv:2505.19223, 2025

    Fengqi Zhu, Rongzhen Wang, Shen Nie, Xiaolu Zhang, Chunwei Wu, Jun Hu, Jun Zhou, Jianfei Chen, Yankai Lin, Ji-Rong Wen, et al. Llada 1.5: Variance-reduced preference optimization for large language diffusion models.arXiv:2505.19223, 2025

  196. [204]

    Killian, Mikhail Yurochkin, Zhengzhong Liu, Eric P

    Zhoujun Cheng, Shibo Hao, Tianyang Liu, Fan Zhou, Yutao Xie, Feng Yao, Yuexin Bian, Yong- hao Zhuang, Nilabjo Dey, Yuheng Zha, Yi Gu, Kun Zhou, Yuqi Wang, Yuan Li, Richard Fan, Jianshu She, Chengqian Gao, Abulhair Saparov, Haonan Li, Taylor W. Killian, Mikhail Yurochkin, Zheng...

  197. [205]

    Can one domain help others? a data-centric study on multi-domain reason- ing via reinforcement learning.arXiv:2507.17512, 2025

    Yu Li, Zhuoshi Pan, Honglin Lin, Mengyuan Sun, Conghui He, and Lijun Wu. Can one domain help others? a data-centric study on multi-domain reason- ing via reinforcement learning.arXiv:2507.17512, 2025

  198. [206]

    American invitational mathematics examination (AIME)

    Mathematical Association of America. American invitational mathematics examination (AIME). https://maa.org/math-competitions/aime, 2024

  199. [207]

    Can one domain help others? a data-centric study on multi-domain rea- soning via reinforcement learning, 2025

    Yu Li, Zhuoshi Pan, Honglin Lin, Mengyuan Sun, Conghui He, and Lijun Wu. Can one domain help others? a data-centric study on multi-domain rea- soning via reinforcement learning, 2025

  200. [208]

    Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning.arXiv:2411.02337, 2024

    Zehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Xinyue Yang, Jiadai Sun, Yu Yang, Shuntian Yao, Tianjie Zhang, Wei Xu, Jie Tang, and Yuxiao Dong. Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning.arXiv:2411.02337, 2024

  201. [209]

    Huatuogpt-o1, towards medical com- plex reasoning with llms.arXiv:2412.18925, 2024

    Junying Chen, Zhenyang Cai, Ke Ji, Xidong Wang, Wanlong Liu, Rongsheng Wang, Jianye Hou, and Benyou Wang. Huatuogpt-o1, towards medical com- plex reasoning with llms.arXiv:2412.18925, 2024

  202. [210]

    Medagent- gym: A scalable agentic training environment for code-centric reasoning in biomedical data science

    Ran Xu, Yuchen Zhuang, Yishan Zhong, Yue Yu, Xiangru Tang, Hang Wu, May D Wang, Peifeng Ruan, Donghan Yang, Tao Wang, Guanghua Xiao, Carl Yang, Yang Xie, and Wenqi Shi. Medagent- gym: A scalable agentic training environment for code-centric reasoning in biomedical data science...

  203. [211]

    Ui-tars: Pioneering automated gui interaction with native agents.arXiv:2501.12326, 2025

    Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, et al. Ui-tars: Pioneering automated gui interaction with native agents.arXiv:2501.12326, 2025

  204. [212]

    Ui-r1: En- hancing efficient action prediction of gui agents by reinforcement learning.arXiv:2503.21620, 2025

    Z Lu, Y Chai, Y Guo, X Yin, L Liu, H Wang, H Xiao, S Ren, G Xiong, and H Li. Ui-r1: En- hancing efficient action prediction of gui agents by reinforcement learning.arXiv:2503.21620, 2025

  205. [213]

    Gui-r1: A generalist r1-style vision-language action model for gui agents

    Z Luo et al. Gui-r1: A generalist r1-style vision-language action model for gui agents. arXiv:2504.10458, 2025

  206. [214]

    Polite Pool,

    DeepSeek-AI. Deepseek-v3.2: Pushing the frontier of open large language models.arXiv:2512.02556, 2025. 44 An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments A Data Collection Strategy & Data Preprocessing To ensure a comprehensiv...

  207. [215]

    The initial API query filtered works where the title or abstract contained foundational reinforcement learning terminology (e.g.,reinforcement learning, MARL, DRL, RLHF , offline RL). Dynamic Citation ThresholdingTo objectively identify milestone environments without succumbin...

  208. [216]

    Lexical BlacklistingWe established an absolute blacklist to filter out papers focused exclusively on al- gorithmic convergence or methodology. Unless overridden by a strong environment signal, papers containing key- words such asalgorithm, policy optimization, q-learning, acto...

  209. [217]

    Strong Semantic AnchoringA paper was imme- diately classified as an environment milestone if its title contained unambiguous benchmark indicators. This in- cluded exact matches for terminology (e.g.,benchmark, simulator, testbed, arena) or regular expression matches for establ...

  210. [218]

    release pattern

    Syntactic Action Parsing (Abstract Inverted Index) For papers exhibiting weak or ambiguous signals in the title (e.g., containing general terms likeframework, plat- form, orproblem), we executed a deep syntactic parse utilizing the OpenAlex abstract inverted index. A paper was...

  211. [219]

    dataset,

    Modality-Specific Nuance: The "Dataset" Excep- tionIn classic RL, static datasets do not constitute en- vironments. However, in the era of Offline RL and Large Language Model (LLM) agents, static datasets are fre- quently wrapped into interactive cognitive environments. To acc...

  212. [220]

    The Golden Pathway (Explicit Recognition):Papers whose titles explicitly matched a curated registry of widely recognized LLM environments and foundation benchmarks (e.g.,WebArena, SWE-bench, ToolLLM, GSM8K, ALFWorld) were automatically preserved to guarantee the inclusion of i...

  213. [221]

    Cognitive Fingerprints

    The Semantic Release Pathway:For novel or lesser- known environments, the abstract was required to sat- isfy a strict tripartite syntactic condition. It must simul- taneously contain (a) an LLM domain identifier (e.g., commonsense reasoning, math word problem), (b) an active r...

Pith tools

Reviewed May 15, 2026 · model on record in the stance chip above.