REVIEW 3 major objections 6 minor 178 references
A Research Agenda for Usability and Generalisation in Reinforcement Learning
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper argues that reinforcement learning environments should be described in user-friendly domain-specific or natural languages, and that complete descriptions, supplied to agents as context, are the route to zero-shot generalization…
desk verdict A clear, honest position paper that usefully spells out a research agenda for DSL/natural-language environment descriptions in RL—but its load-bearing assumption about user-friendliness remains quantitatively untested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is the environment description as a dual-use object: a user-friendly formulation in a DSL or natural language that a compiler or language model translates into a runnable simulator, and that is simultaneously provided to the agent as context conditioning its policy or value function. Formalized as a contextual (PO)MDP or Markov game, this context must be complete enough to disambiguate between environments; incompleteness turns the collection of possible environments into a partially observable problem. The shared vocabulary of the description language is what makes the loop work in both directions: environments can be generated from contexts and contexts from environments, enabling procedural generation of training tasks and, in principle, zero-shot transfer to unseen descriptions.
What would settle it
Run a controlled usability study in which people with no programming background describe the same dozen tasks (board games, simple control problems) in a user-friendly DSL, in natural language, and in a general-purpose language, then measure whether the DSL and natural-language descriptions are more accurate, complete, and faster to produce. If novices produce unusable or incomplete descriptions at comparable rates, the usability and generalisation arguments for the agenda collapse.
Extended reading notes
Core claim
The paper's central claim is that the customary workflow—an engineer implementing each environment directly in a general-purpose programming language or a hardware-acceleration framework—is itself an obstacle to RL adoption and to generalization. It proposes that environments should be described in a shared, user-friendly language, with complete descriptions that can be compiled to executable simulators; those same descriptions, supplied to an agent as context, are argued to be a prerequisite for unrestricted zero-shot generalisation across every task expressible in the language. The paper supports this with a concrete example of a game description language that lets a user write tic-tac-toe as a short high-level script, and notes a library of over 1400 game descriptions contributed in part by non-programmers. It also identifies a practical rupture: succinct description languages often make it impossible to infer the full action space in advance, which violates a common assumption in deep RL APIs. The conclusion is a position statement: the RL community should place greater focus on benchmarks with environments defined in user-friendly DSLs or natural languages.
Load-bearing premise
The entire agenda rests on the assumption that non-engineers can write complete, unambiguous environment descriptions in a DSL or natural language more easily than in code; the paper itself admits it has no quantitative evidence for this.
Editorial extensions
If this is right
- Benchmark suites should be built around description languages rather than hand-written simulators, and evaluation should test agents on unseen descriptions in the same language.
- Non-engineers—private individuals, small organisations, and domain experts—could specify their own tasks and receive a policy without writing code, provided the compiler and a sufficiently general agent exist.
- Complete, compileable descriptions are claimed as a prerequisite for unrestricted zero-shot generalisation; partial contexts such as numeric goal coordinates or short instructions cannot do the same job.
- Existing assumptions like knowing the full action space in advance may fail for user-friendly DSLs, so deep RL methods will need to handle action aliasing or variable action spaces.
- Procedural generation of new descriptions in the same language can supply a curriculum, letting agents learn the semantics of the language and generalise across the whole describable space.
Reading between the lines
- The agenda implicitly predicts a convergence between RL environment design and language-model-driven program synthesis: if natural-language descriptions become the interface, the reliability of translating language to simulators becomes a core RL benchmark question rather than a side concern.
- A testable extension is to measure how much zero-shot transfer performance scales with the number and diversity of descriptions seen during pretraining; the paper does not make a scaling-law claim, but its argument suggests such a relationship.
- The complete-description requirement may be too strong for physical-world tasks, where dynamics are not fully describable in language; the paper allows reward-only descriptions for such cases, which leaves a gap between virtual and physical generalization.
- If the position is adopted, evaluation methodology must separate what an agent learned about the description language from what it learned about general RL competence, since performance on unseen descriptions could come from either.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that the customary practice of implementing RL environments in general-purpose programming languages imposes a usability barrier on non-engineer users and also obstructs progress on generalization, because there is no shared formalism in which different problems are represented. The authors advocate for a research agenda centered on describing environments in user-friendly domain-specific languages (DSLs) or natural language, such that (i) users with little programming expertise can formally describe their problems and (ii) algorithms can use the resulting complete descriptions as context to generalize zero-shot across all tasks describable in the chosen language. The paper states two explicit assumptions (Section 3.4), discusses potential issues such as action-space inference and simulation speed, notes other barriers to RL adoption (Section 5), and lays out desiderata and research directions (Section 6). It is an extended version of an earlier position paper, with additions covering other barriers and a more detailed agenda.
Significance. If the central position is correct, it would motivate a substantial shift in benchmark practice away from bespoke simulator code toward description-language-native environments, potentially democratizing RL for non-engineers and opening a new axis for zero-shot generalization research. The paper is honest and internally consistent: the two core assumptions are stated explicitly, the generalization claims in Section 4.3 are hedged, and Section 6 openly concedes the lack of quantitative evidence for the user-friendliness trade-off. It also credibly connects the agenda to concrete existing artifacts (Ludii, GAVEL, Ludax), gives a concrete desiderata list, and distinguishes its proposal from related DSL/context work. Its main weakness is that the load-bearing Assumption 1 rests on informal evidence and anecdotal experience rather than a user study; nevertheless the agenda is constructed so that this assumption could, in principle, be tested, which is a strength worth acknowledging.
major comments (3)
- [Section 3.4 and Section 6] Assumption 1 (that DSLs or natural language are more user-friendly than general-purpose programming languages for defining environments) is load-bearing for both halves of the agenda: the usability motivation in Section 3 and the generalization thesis in Section 4.3 both require non-engineers to be able to author complete, unambiguous descriptions. The only support offered is Ludii's design goal, the count of over 1400 game descriptions, and third-party forum contributions, while Section 6 admits "we have no quantitative evidence at this point." Because the entire agenda collapses if this assumption fails for the intended end users, the manuscript should either include or cite a user study that measures whether non-programmers can successfully write and validate complete environment descriptions, or explicitly elevate this to the first, falsifiable item of the research agenda with pre-registered success criteria.
- [Section 3.3] The proposed natural-language workflow requires an LLM to translate the description into a DSL, after which "a user can inspect the generated description and make corrections if necessary before it is compiled into a simulator." This verification-and-correction step still demands the ability to read and edit DSL code, which is precisely the expertise the agenda aims to remove. The paper should address how much DSL/verification competence is assumed of the end user and whether the verification step is realistically feasible for the target population; otherwise the usability advantage of natural-language descriptions is substantially weakened.
- [Section 4.3] The statement that complete environment descriptions "are likely to be a prerequisite for unrestricted, zero-shot generalisation in RL" is supported only by an analogy to humans learning new board games from rules, and the paper itself gives a video-game fire example where humans generalize without a complete description of the environment. This is not an internally inconsistent claim, but it is underspecified: the term "unrestricted" is never defined, and no concrete evidence or formal argument is given for why completeness is necessary rather than merely helpful. The claim should be reframed as a falsifiable hypothesis with a precise scope (e.g., across the set of tasks describable in a given DSL), and the authors should specify what experimental comparisons (such as context completeness versus zero-shot transfer on a DSL benchmark suite) would support or refute it.
minor comments (6)
- [Section 3.4] The word "exectuable" should be "executable."
- [Section 2.1] The phrase "when action according to a policy" should be "when acting according to a policy."
- [Section 3.2] The Ludii example may confuse readers unfamiliar with the language because the comment says some rules are omitted as defaults without explaining what those defaults are; a short note on the default turn-taking and draw conditions would improve readability.
- [Section 6] The t-SNE figure (Fig. 2) is described only as "reduced from a larger feature space" with a citation; the caption could usefully state which features from [116] were used and how the embedding was computed.
- [References] Reference [76] contains "hum4n l4ngu4ge" which appears to be either a deliberate obfuscation or a transcription error; if deliberate, the authors should add a note, as it may confuse readers.
- [Section 4.1] The term "unrestricted, zero-shot generalisation" is used in Section 4.3 before being defined; a definition or at least a clarifying sentence in Section 4.1 would help.
Circularity Check
No circular derivation in this position paper; the only borderline issue is self-referential evidence for Assumption 1, which is not a load-bearing circular step.
full rationale
This is a position paper, not a derivation chain. It contains no fitted parameters, no equations whose outputs are also inputs, and no 'prediction' that is statistically forced. The central Position (Section 3.4) is a normative call for more DSL/natural-language benchmarks, supported by two explicitly stated assumptions. Section 4.3 argues that complete environment descriptions are 'likely to be a prerequisite' for unrestricted zero-shot generalisation, citing external theoretical work ([56]) and using an analogy to human board-game learning; it is not presented as a theorem and does not reduce to its inputs. The main self-reference appears in Assumption 1: the paper supports the claim that DSLs can be user-friendly by citing the authors' own Ludii design goal [115] and the >1400-game library on the authors' platform, and Section 6 concedes 'we have no quantitative evidence at this point' for the user-friendliness/generality trade-off. This is self-referential evidence rather than circular derivation: the conclusion is not equivalent to the citation by construction, and the library statistics are publicly checkable. Under the hard rule requiring a quoted reduction, no actual circular step can be exhibited, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Assumption 1: Defining environments in DSLs or natural languages can be more user-friendly than defining them in general-purpose programming languages.
- domain assumption Assumption 2: Enabling environments to be defined in more user-friendly ways is desirable.
- domain assumption Complete, compilable environment descriptions are a prerequisite for unrestricted zero-shot generalisation in RL.
- domain assumption The full action space cannot be reliably inferred from user-friendly DSLs such as Ludii.
Cite this review
Pith. "Pith review of A Research Agenda for Usability and Generalisation in Reinforcement Learning." pith.science (2026). https://pith.science/paper/3L2STC2L
@misc{pith2026241216970,
author = {Pith},
title = {Pith review of: A Research Agenda for Usability and Generalisation in Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/3L2STC2L}},
note = {Machine review of arXiv:2412.16970}
}
read the original abstract
It is common practice in reinforcement learning (RL) research to train and deploy agents in bespoke simulators, typically implemented by engineers directly in general-purpose programming languages or hardware acceleration frameworks such as CUDA or JAX. This means that programming and engineering expertise is not only required to develop RL algorithms, but is also required to use already developed algorithms for novel problems. The latter poses a problem in terms of the usability of RL, in particular for private individuals and small organisations without substantial engineering expertise. We also perceive this as a challenge for effective generalisation in RL, in the sense that is no standard, shared formalism in which different problems are represented. As we typically have no consistent representation through which to provide information about any novel problem to an agent, our agents also cannot instantly or rapidly generalise to novel problems. In this position paper, we advocate for a research agenda centred around the use of user-friendly description languages for describing problems, such that (i) users with little to no engineering expertise can formally describe the problems they would like to be tackled by RL algorithms, and (ii) algorithms can leverage problem descriptions to effectively generalise among all problems describable in the language of choice.
Figures
Reference graph
Works this paper leans on
-
[1]
In: 2019 ICML Workshop on Human in the Loop Learning (2019)
Abid, A., Abdalla, A., Abid, A., Khan, D., Alfozan, A., Zou, J.: Gradio: Hassle- free sharing and testing of ML models in the wild. In: 2019 ICML Workshop on Human in the Loop Learning (2019)
2019
-
[2]
In: AAAI 2024 Workshop on Synergy of Reinforcement Learning and Large Language Models (2024)
Afshar, A., Li, W.: DeLF: Designing learning environments with foundation mod- els. In: AAAI 2024 Workshop on Synergy of Reinforcement Learning and Large Language Models (2024)
2024
-
[3]
In: Ranzato, M., Beygelzimer,A.,Dauphin,Y.,Liang,P.,Vaughan, J.W.(eds.)Advances inNeural InformationProcessingSystems.vol.34,pp.29304–29320.CurranAssociates,Inc
Agarwal, R., Schwarzer, M., Castro, P.S., Courville, A., Bellemare, M.G.: Deep reinforcement learning at the edge of the statistical precipice. In: Ranzato, M., Beygelzimer,A.,Dauphin,Y.,Liang,P.,Vaughan, J.W.(eds.)Advances inNeural InformationProcessingSystems.vol.34,pp.29304–29320.CurranAssociates,Inc. (2021)
2021
-
[4]
In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A
Agarwal, R., Schwarzer, M., Castro, P.S., Courville, A.C., Bellemare, M.: Reincar- nating reinforcement learning: Reusing prior computation to accelerate progress. In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A. (eds.) Advances in Neural Information Processing Systems. vol. 35, pp. 28955–28971. Curran Associates, Inc. (2022)
2022
-
[5]
In: IEEE ICRA 2024 Workshop on Vision-Language Models for Navigation and Ma- nipulation (2024) 20 D.J.N.J
Ahn, M., Dwibedi, D., Finn, C., Gonzalez Arenas, M., Gopalakrishnan, K., Haus- man, K., Ichter, B., Irpan, A., Joshi, N., Julian, R., Kirmani, S., Leal, I., Lee, E., Levine, S., Lu, Y., Maddineni, S., Rao, K., Sadigh, D., Sanketi, P., Sermanet, P., Vuong, Q., Welker, S., Xia, F., Xiao, T., Xu, P., Xu, S., Xu, Z.: AutoRT: Embodied foundation models for lar...
2024
-
[6]
In: Proceedings of the 34th International Conference on Machine Learning
Andreas, J., Klein, D., Levine, S.: Modular multitask reinforcement learning with policy sketches. In: Proceedings of the 34th International Conference on Machine Learning. vol. 70, pp. 166–175. PMLR (2017)
2017
-
[7]
In: 2021 International Conference on Learning Representations (2021)
Andrychowicz, M., Raichuk, A., Stań’czyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., Gelly, S., Bachem, O.: What matters for on-policy deep actor-critic methods? a large-scale study. In: 2021 International Conference on Learning Representations (2021)
2021
-
[8]
Journal of Internet Services and Applications6(13) (2015)
Aram, M., Neumann, G.: Multilayered analysis of co-development of business information systems. Journal of Internet Services and Applications6(13) (2015)
2015
Show all 178 references
-
[9]
In: AAAI-21 Workshop on Reinforcement Learning in Games (2021)
Bamford, C., Huang, S., Lucas, S.: Griddly: A platform for AI research in games. In: AAAI-21 Workshop on Reinforcement Learning in Games (2021)
2021
-
[10]
In: The 20th International Joint Conference on Artificial Intelligence
Banerjee, B., Stone, P.: General game learning using knowledge transfer. In: The 20th International Joint Conference on Artificial Intelligence. pp. 672–677 (2007)
2007
-
[11]
https://arxiv.org/abs/2301.08028 (2023)
Beck, J., Vuorio, R., Liu, E.Z., Xiong, Z., Zintgraf, L., Finn, C., Whiteson, S.: A survey of meta-reinforcement learning. https://arxiv.org/abs/2301.08028 (2023)
2023 arXiv
-
[12]
Nature588, 77–82 (2020)
Bellemare, M.G., Candido, S., Castro, P.S., Gong, J., Machado, M.C., Moitra, S., Ponda, S.S., Wang, Z.: Autonomous navigation of stratospheric balloons using reinforcement learning. Nature588, 77–82 (2020)
2020
-
[13]
Journal of Artificial Intelli- gence Research 47(1), 253–279 (2013)
Bellemare, M.G., Naddaf, Y., Veness, J., Bowling, M.: The arcade learning envi- ronment: An evaluation platform for general agents. Journal of Artificial Intelli- gence Research 47(1), 253–279 (2013)
2013
-
[14]
Transactions on Machine Learning Research (2023)
Benjamins, C., Eimer, T., Schubert, F., Mohan, A., Döhler, S., Biedenkapp, A., Rosenhahn, B., Hutter, F., Lindauer, M.: Contextualize me – the case for context in reinforcement learning. Transactions on Machine Learning Research (2023)
2023
-
[15]
In: Proceedings of the 16th International Symposium on Distributed Autonomous Robotic Systems
Bettini, M., Kortvelesy, R., Blumenkamp, J., Prorok, A.: VMAS: A vectorized multi-agent simulator for collective robot learning. In: Proceedings of the 16th International Symposium on Distributed Autonomous Robotic Systems. DARS ’22, Springer (2022)
2022
-
[16]
Voleti, Z.E., Letts, A., Jampani, V., Rombach, R.: Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., anda V. Voleti, Z.E., Letts, A., Jampani, V., Rombach, R.: Stable video diffusion: Scaling latent video diffusion models to large datasets. https: //arxiv.org/abs/2311.15127 (2023)
2023 arXiv
-
[17]
In: Proceedings of the 42nd International Conference on Machine Learning (2025), to appear
Blili-Hamelin, B., Graziul, C., Hancox-Li, L., Hazan, H., El-Mhamdi, E.M., Ghosh, A., Heller, K., Metcalf, J., Murai, F., Salvaggio, E., Smart, A., Snider, T., Tighanimine, M., Ringer, T., Mitchell, M., Dori-Hacohen, S.: Position: Stop treating ‘AGI’ as the north-star goal of ...
2025
-
[18]
In: Proceedings of the International Conference on Learning Represen- tations (2024)
Bonnet, C., Luo, D., Byrne, D., Surana, S., Abramowitz, S., Duckworth, P., Coyette, V., Midgley, L.I., Tegegn, E., Kalloniatis, T., Mahjoub, O., Macfarlane, M., Smit, A.P., Grinsztajn, N., Bolge, R., Waters, C.N., Mimouni, M.A., Sob, U.A.M., de Kock, R., Singh, S., Furelos-Bla...
2024
-
[19]
In: Xing, E.P., Jebara, T
Bou Ammar, H., Eaton, E., Ruvolo, P., Taylor, M.E.: Online multi-task learn- ing for policy gradient methods. In: Xing, E.P., Jebara, T. (eds.) Proceedings of the 31st International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 32, pp. 1206–1214 (2014)
2014
-
[20]
com/google/jax
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M.J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., Zhang, Q.: JAX: A Research Agenda for Usability and Generalisation in RL 21 composable transformations of Python+NumPy programs (2018),...
2018
-
[21]
Journal of Artificial Intelligence Research43, 661–704 (2012)
Branavan, S.R.K., Silver, D., Barzilay, R.: Learning to win by reading manuals in a Monte-Carlo framework. Journal of Artificial Intelligence Research43, 661–704 (2012)
2012
-
[22]
https://arxiv.org/abs/1606.01540 (2016)
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., Zaremba, W.: OpenAI gym. https://arxiv.org/abs/1606.01540 (2016)
2016 arXiv
-
[23]
ludii.games/downloads/LudiiLanguageReference.pdf (2020)
Browne, C., Soemers, D.J.N.J., Piette, É., Stephenson, M., Crist, W.: Ludii lan- guage reference. ludii.games/downloads/LudiiLanguageReference.pdf (2020)
2020
-
[24]
Phd thesis, Faculty of Information Technology, Queensland University of Tech- nology, Queensland, Australia (2009)
Browne, C.B.: Automatic Generation and Evaluation of Recombination Games. Phd thesis, Faculty of Information Technology, Queensland University of Tech- nology, Queensland, Australia (2009)
2009
-
[25]
In: Daumé III, H., Singh, A
Cobbe, K., Hesse, C., Hilton, J., Schulman, J.: Leveraging procedural generation to benchmark reinforcement learning. In: Daumé III, H., Singh, A. (eds.) Pro- ceedings of the 37th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 119,...
2020
-
[26]
In: Chaudhuri, K., Salakhutdinov, R
Cobbe, K., Klimov, O., Hesse, C., Kim, T., Schulman, J.: Quantifying gener- alization in reinforcement learning. In: Chaudhuri, K., Salakhutdinov, R. (eds.) Proceedings of the 36th International Conference on Machine Learning. Proceed- ings of Machine Learning Research, vol. 9...
2019
-
[27]
https://arxiv.org/abs/2109
Cummins, C., Wasti, B., Guo, J., Cui, B., Ansel, J., Gomez, S., Jain, S., Liu, J., Teytaud, O., Steiner, B., Tian, Y., Leather, H.: Compilergym: Robust, performant compiler optimization environments for ai research. https://arxiv.org/abs/2109. 08267 (2021)
2021
-
[28]
In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H
Dalton, S., Frosio, I.: Accelerating reinforcement learning through GPU Atari emulation. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H. (eds.) Advances in Neural Information Processing Systems. vol. 33, pp. 19773–19782. Curran Associates, Inc. (2020)
2020
-
[29]
In: Finding the Frame workshop @ Reinforcement Learning Con- ference (2024)
Davidson, G., Gureckis, T.M.: Toward complex and structured goals in reinforce- ment learning. In: Finding the Frame workshop @ Reinforcement Learning Con- ference (2024)
2024
-
[30]
https://arxiv.org/abs/2405.13242 (2024)
Davidson,G.,Todd,G.,Togelius,J.,Gureckis,T.M.,Lake,B.M.:Goalsasreward- producing programs. https://arxiv.org/abs/2405.13242 (2024)
2024 arXiv
-
[31]
Nature602, 414–419 (2022)
Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de las Casas, D., Donner, C., Fritz, L., Galperti, C., Huber, A., Keeling, J., Tsimpoukelli, M., Kay, J., Merle, A., Moret, J.M., Noury, S., Pesamosca, F., Pfa...
2022
-
[32]
In: Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA)
Deisenroth, M.P., Englert, P., Peters, J., Fox, D.: Multi-task policy search for robotics. In: Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA). pp. 3876–3881 (2014)
2014
-
[33]
In: Advances in Neural Information Processing Systems
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A., Russell, S., Critch, A., Levine, S.: Emergent complexity and zero-shot transfer via unsupervised environment design. In: Advances in Neural Information Processing Systems. vol. 33, pp. 13049–13061 (2020)
2020
-
[34]
In: Proceedings of the 38th International Conference on Software Engineering
Desai, A., Gulwani, S., Hingorani, V., Jain, N., Karkare, A., Marron, M., R, S., Roy, S.: Program synthesis using natural language. In: Proceedings of the 38th International Conference on Software Engineering. p. 345–356. Association for Computing Machinery (2016) 22 D.J.N.J. ...
2016
-
[35]
In: Pro- ceedings of the Thirtieth International Joint Conference on Artificial Intelligence
Eimer, T., Biedenkapp, A., Reimer, M., Adriaensen, S., Hutter, F., Lindauer, M.: DACBench: A benchmark library for dynamic algorithm configuration. In: Pro- ceedings of the Thirtieth International Joint Conference on Artificial Intelligence. pp. 1668–1674 (2021)
2021
-
[36]
In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J
Eimer, T., Lindauer, M., Raileanu, R.: Hyperparameters in reinforcement learning and how to tune them. In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J. (eds.) Proceedings of the 40th International Conference on Machine Learning. Proceedings of M...
2023
-
[37]
In: Advances in Neural Information Processing Systems (2023), accepted
Ellis, B., Cook, J., Moalla, S., Samvelyan, M., Sun, M., Mahajan, A., Foerster, J.N.,Whiteson,S.:SMACv2:Animprovedbenchmarkforcooperativemulti-agent reinforcement learning. In: Advances in Neural Information Processing Systems (2023), accepted
2023
-
[38]
https://arxiv.org/abs/1810.00123 (2018)
Farebrother, J., Machado, M.C., Bowling, M.: Generalization and regularization in DQN. https://arxiv.org/abs/1810.00123 (2018)
2018 arXiv
-
[39]
Progress in AI2(1), 13–27 (2013)
Fernandéz, F., Veloso, M.: Learning domain structure through probabilistic policy reuse in reinforcement learning. Progress in AI2(1), 13–27 (2013)
2013
-
[40]
In: Precup, D., Teh, Y.W
Finn, C., Abeel, P., Levine, S.: Model-agnostic meta-learning for fast adaptation of deep networks. In: Precup, D., Teh, Y.W. (eds.) Proceedings of the 34th In- ternational Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 70, pp. 1126–1135 (2017)
2017
-
[41]
Freeman, C.D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., Bachem, O.: Brax - a differentiable physics engine for large scale rigid body simulation (2021), http: //github.com/google/brax
2021
-
[42]
In: ICAIF ’23: Proceedings of the Fourth ACM International Conference on AI in Finance
Frey, S., Li, K., Nagy, P., Sapora, S., Lu, C., Zohren, S., Foerster, J., Calinescu, A.: JAX-LOB: A GPU-accelerated limit order book simulator to unlock large scale reinforcement learning for trading. In: ICAIF ’23: Proceedings of the Fourth ACM International Conference on AI ...
2023
-
[43]
In: Proceedings of the 36th International Conference on Machine Learning
Gamrian, S., Goldberg, Y.: Transfer learning for related reinforcement learning tasks via image-to-image translation. In: Proceedings of the 36th International Conference on Machine Learning. pp. 2063–2072 (2019)
2019
-
[44]
Synthesis Lectures on Ar- tificial Intelligence and Machine Learning, Morgan & Claypool Publishers (2014)
Genesereth, M., Thielscher, M.: General Game Playing. Synthesis Lectures on Ar- tificial Intelligence and Machine Learning, Morgan & Claypool Publishers (2014)
2014
-
[45]
In: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W
Ghosh, D., Rahme, J., Kumar, A., Zhang, A., Adams, R.P., Levine, S.: Why gen- eralization in rl is difficult: Epistemic pomdps and implicit partial observability. In: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W. (eds.) Advances in Neural Information Proc...
2021
-
[46]
In: Brazilian Conference on Intelligent Systems (BRACIS)
Glatt, R., da Silva, F.L., Costa, A.H.R.: Towards knowledge transfer in deep re- inforcement learning. In: Brazilian Conference on Intelligent Systems (BRACIS). pp. 91–96. IEEE (2016)
2016
-
[47]
Ex- pert Systems with Applications156 (2020)
Glatt, R., Silva, F.L.D., da Costa Bianchi, R.A., Costa, A.H.R.: DECAF: Deep case-based policy inference for knowledge transfer in reinforcement learning. Ex- pert Systems with Applications156 (2020)
2020
-
[48]
Goldie, A.D., Lu, C., Jackson, M.T., Whiteson, S., Foerster, J.N.: Can learned optimization make reinforcement learning less difficult? In: AutoRL Workshop @ ICML 2024 (2024)
2024
-
[49]
In: The Thirty-Fourth AAAI Conference on Artificial Intelligence
Goldwaser, A., Thielscher, M.: Deep reinforcement learning for general game play- ing. In: The Thirty-Fourth AAAI Conference on Artificial Intelligence. pp. 1701–
-
[50]
In: Proceedings of the 2020 Conference on Robot Learning
Goyal, P., Niekum, S., Mooney, R.J.: PixL2R: Guiding reinforcement learning using natural language by mapping pixels to rewards. In: Proceedings of the 2020 Conference on Robot Learning. PMLR, vol. 155, pp. 485–497 (2021)
2021
-
[51]
In: Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (2023)
Gulino, C., Fu, J., Luo, W., Tucker, G., Bronstein, E., Lu, Y., Harb, J., Pan, X., Wang, Y., Chen, X., Co-Reyes, J.D., Agarwal, R., Roelofs, R., Lu, Y., Montali, N., Mougin, P., Yang, Z., B, W., Faust, A., McAllister, R., Anguelov, D., Sapp, B.: Waymax: An accelerated, data-dr...
2023
-
[52]
In: Proceedings of the 42nd International Conference on Machine Learning (2025), to appear
Hartman, S., Ong, C.S., Powles, J., Kuhnert, P.: Position: We need responsible, application-driven (RAD) AI research. In: Proceedings of the 42nd International Conference on Machine Learning (2025), to appear
2025
-
[53]
In: Proceedings of the 32nd AAAI Confer- ence on Artificial Intelligence
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., Meger, D.: Deep reinforcement learning that matters. In: Proceedings of the 32nd AAAI Confer- ence on Artificial Intelligence. pp. 3207–3214. AAAI (2018)
2018
-
[54]
Communications of the ACM64(12), 58–65 (2021)
Hooker, S.: The hardware lottery. Communications of the ACM64(12), 58–65 (2021)
2021
-
[55]
The MIT Press, Cambridge, Massachusetts (1960)
Howard, R.A.: Dynamic Programming and Markov Processes. The MIT Press, Cambridge, Massachusetts (1960)
1960
-
[56]
In: ICML 2019 Workshop on Understanding and Improving Generalization in Deep Learning (2019)
Irpan, A., Song, X.: The principle of unchanged optimality in reinforcement learn- ing generalization. In: ICML 2019 Workshop on Understanding and Improving Generalization in Deep Learning (2019)
2019
-
[57]
In: Advances in Neural Information Processing Systems
Jackson, M.T., Jiang, M., Parker-Holder, J., Vuorio, R., Lu, C., Farquhar, G., Whiteson, S., Foerster, J.N.: Discovering general reinforcement learning algo- rithms with adversarial environment design. In: Advances in Neural Information Processing Systems. vol. 36, pp. 79980–7...
2023
-
[58]
it’s unwieldy and it takes a lot of time
Jacob, M., Devlin, S., Hofmann, K.: “it’s unwieldy and it takes a lot of time” — challengesandopportunitiesforcreatingagentsincommercialgames.In:Proceed- ings of the Sixteenth AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment. pp. 88–94 (2020)
2020
-
[59]
In: Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F
Jordan, S.M., White, A., da Silva, B.C., White, M., Thomas, P.S.: Position: Benchmarking is limited in reinforcement learning research. In: Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F. (eds.) Proceedings of the 41st Internatio...
2024
-
[60]
https://nips.cc/ virtual/2022/63891 (2022), opinion talk contributed to the Deep Reinforcement Learning Workshop at NeurIPS 2022
Jordan, S.M.: Scientific experiments in reinforcement learning. https://nips.cc/ virtual/2022/63891 (2022), opinion talk contributed to the Deep Reinforcement Learning Workshop at NeurIPS 2022
2022
-
[61]
In: Advances in Neural Information Processing Systems
Jothimurugan, K., Alur, R., Bastani, O.: A composable specification language for reinforcement learning tasks. In: Advances in Neural Information Processing Systems. vol. 32, pp. 13041–13051 (2019)
2019
-
[62]
In: NeurIPS 2018 Workshop on Deep Reinforcement Learning (2018)
Justesen, N., Torrado, R.R., Bontrager, P., Khalifa, A., Togelius, J., Risi, S.: Illuminating generalization in deep reinforcement learning through procedural level generation. In: NeurIPS 2018 Workshop on Deep Reinforcement Learning (2018)
2018
-
[63]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Kaiser, Ł., Stafiniak, Ł.: First-order logic with counting for general game playing. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 25, pp. 791–796 (2011) 24 D.J.N.J. Soemers et al
2011
-
[64]
Transactions on Machine Learning Research (2025)
Kaufmann, T., Weng, P., Bengs, V., Hüllermeier, E.: A survey of reinforce- ment learning from human feedback. Transactions on Machine Learning Research (2025)
2025
-
[65]
https: //arxiv.org/abs/2401.02991 (2024)
Kharyal, C., Krishna Gottipati, S., Kumar Sinha, T., Das, S., Taylor, M.E.: GLIDE-RL: Grounded language instruction through DEmonstration in RL. https: //arxiv.org/abs/2401.02991 (2024)
2024 arXiv
-
[66]
Journal of Artificial Intelligence Research 76, 201–264 (2023)
Kirk, R., Zhang, A., Grefenstette, E., Rocktäschel, T.: A survey of zero-shot generalisation in deep reinforcement learning. Journal of Artificial Intelligence Research 76, 201–264 (2023)
2023
-
[67]
In: Proceedings of the 2020 IEEE Conference on Games
Kowalksi, J., Miernik, R., Mika, M., Pawlik, W., Sutowicz, J., Szykuła, M., Tkaczyk, A.: Efficient reasoning in regular boardgames. In: Proceedings of the 2020 IEEE Conference on Games. pp. 455–462. IEEE (2020)
2020
-
[68]
In: Proceedings of the 33rd AAAI Conference on Artificial Intelligence
Kowalski, J., Maksymilian, M., Sutowicz, J., Szykuła, M.: Regular boardgames. In: Proceedings of the 33rd AAAI Conference on Artificial Intelligence. vol. 33, pp. 1699–1706. AAAI Press (2019)
2019
-
[69]
In: Advances in Neural Information Processing Systems (2023)
Koyamada, S., Okano, S., Nishimori, S., Murata, Y., Habara, K., Kita, H., Ishii, S.:Pgx:Hardware-acceleratedparallelgamesimulatorsforreinforcementlearning. In: Advances in Neural Information Processing Systems (2023)
2023
-
[70]
In: Kok, J., Koronacki, J., Mantaras, R., Matwin, S., Mladenič, D., Skowron, A
Kuhlmann, G., Stone, P.: Graph-based domain mapping for transfer learning in general games. In: Kok, J., Koronacki, J., Mantaras, R., Matwin, S., Mladenič, D., Skowron, A. (eds.) Machine Learning: ECML 2007. Lecture Notes in Computer Science, vol. 4071, pp. 188–200. Springer, ...
2007
-
[71]
Lange, R.T.: gymnax: A JAX-based reinforcement learning environment library (2022), http://github.com/RobertTLange/gymnax
2022
-
[72]
In: Neural Networks
Lange, S., Riedmiller, M.: Deep auto-encoder neural networks in reinforcement learning. In: Neural Networks. International Joint Conference. 2010. (IJCNN 2010). pp. 1623–1630. IEEE (2010)
2010
-
[73]
In: Wiering, M., van Otterlo, M
Lazaric, A.: Transfer in reinforcement learning: a framework and a survey. In: Wiering, M., van Otterlo, M. (eds.) Reinforcement Learning. Adaptation, Learn- ing, and Optimization, vol. 12, pp. 143–173. Springer, Berlin, Heidelberg (2012)
2012
-
[74]
Nature521(7553), 436–444 (2015)
LeCun, Y., Bengio, Y., Hinton, G.: Deep learning. Nature521(7553), 436–444 (2015)
2015
-
[75]
https:// arxiv.org/abs/2306.14892 (2023)
Lee, J.N., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., Brunskill, E.: Supervised pretraining can learn in-context reinforcement learning. https:// arxiv.org/abs/2306.14892 (2023)
2023 arXiv
-
[76]
Leivada, E., Marcus, G., Günther, F., Murphy, E.: A sentence is worth a thousand pictures: Can large language models understand hum4n l4ngu4ge and the w0rld behind w0rds? https://arxiv.org/abs/2308.00109 (2024)
2024 arXiv
-
[77]
Journal of Machine Learning Research10(40), 1131–1186 (2009)
Li, H., Liao, X., Carin, L.: Multi-task reinforcement learning in partially ob- servable stochastic environments. Journal of Machine Learning Research10(40), 1131–1186 (2009)
2009
-
[78]
In: Agmon, N., Taylor, M.E., Veloso, E.E.M
Li, X., Zhang, J., Bian, J., Tong, Y., Liu, T.Y.: A cooperative multi-agent rein- forcement learning framework for resource balancing in complex logistics network. In: Agmon, N., Taylor, M.E., Veloso, E.E.M. (eds.) Proceedings of the 18th Inter- national Conference on Autonomo...
2019
-
[79]
https://arxiv.org/abs/2306.00937 (2023)
Lifschitz, S., Paster, K., Chan, H., Ba, J., McIlraith, S.: Steve-1: A generative model for text-to-behavior in Minecraft. https://arxiv.org/abs/2306.00937 (2023)
2023 arXiv
-
[80]
In: Proceedings of the International Conference on Automated Planning and Scheduling
Lin, S., Bercher, P.: On the expressive power of planning formalisms in conjunc- tion with LTL. In: Proceedings of the International Conference on Automated Planning and Scheduling. vol. 32, pp. 231–240 (2022) A Research Agenda for Usability and Generalisation in RL 25
2022
-
[81]
In: Proceedings of the 41st International Conference on Machine Learning
Lindauer, M., Karl, F., Klier, A., Moosbauer, J., Tornede, A., Mueller, A., Hutter, F., Feurer, M., Bischl, B.: Position: A call to action for a human-centered Au- toML paradigm. In: Proceedings of the 41st International Conference on Machine Learning. PMLR, vol. 235, pp. 3056...
2024
-
[82]
In: Proceedings of the Eleventh International Conference on Machine Learn- ing
Littman, M.L.: Markov games as a framework for multi-agent reinforcement learn- ing. In: Proceedings of the Eleventh International Conference on Machine Learn- ing. pp. 157–163 (1994)
1994
-
[83]
Love, N., Hinrichs, T., Haley, D., Schkufza, E., Genesereth, M.: General game playing: Game description language specification. Tech. Rep. LG-2006-01, Stan- ford Logic Group (2008)
2008
-
[84]
In: Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19
Luketina, J., Nardelli, N., Farquhar, G., Foerster, J., Andreas, J., Grefenstette, E., Whiteson, S., Rocktä"schel, T.: A survey of reinforcement learning informed by natural language. In: Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligenc...
2019
-
[85]
In: Towards Generalist Robots: Learning Paradigms for Scalable Skill Acquisition @ CoRL2023 (2023)
Luo, J., Hu, Z., Xu, C., Gadipudi, S., Sharma, A., Ahmad, R., Schaal, S., Finn, C., Gupta, A., Levine, S.: SERL: A software suite for sample-efficient robotic reinforcement learning. In: Towards Generalist Robots: Learning Paradigms for Scalable Skill Acquisition @ CoRL2023 (2023)
2023
-
[86]
Journal of Machine Learning Research 9(86), 2579–2605 (2008)
van der Maaten, L., Hinton, G.: Visualizing data using t-sne. Journal of Machine Learning Research 9(86), 2579–2605 (2008)
2008
-
[87]
Journal of Artificial Intelligence Research61, 523–562 (2018)
Machado, M.C., Bellemare, M.G., Talvitie, E., Veness, J., Hausknecht, M., Bowl- ing, M.: Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents. Journal of Artificial Intelligence Research61, 523–562 (2018)
2018
-
[88]
Machine Learning 22, 251–281 (1996)
Maclin, R., Shavlik, J.W.: Creating advice-taking reinforcement learners. Machine Learning 22, 251–281 (1996)
1996
-
[89]
In: Vanschoren, J., Yeung, S
Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., State, G.: Isaac gym: High per- formance GPU based physics simulation for robot learning. In: Vanschoren, J., Yeung, S. (eds.) Proceedings of the Neural ...
2021
-
[90]
Malik, D., Li, Y., Ravikumar, P.: When is generalizable reinforcement learning tractable? In: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W.(eds.)AdvancesinNeuralInformationProcessingSystems.vol.34,pp.8032–
-
[91]
https://arxiv.org/abs/2301.01320 (2023)
Mannor, S., Tamar, A.: Towards deployable RL – what’s broken with RL research and a potential fix. https://arxiv.org/abs/2301.01320 (2023)
2023 arXiv
-
[92]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Maras, M., Kępa, M., Kowalski, J., Szykuła, M.: Fast and knowledge-free deep learning for general game playing (student abstract). In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 38, pp. 23576–23578 (2024)
2024
-
[93]
McDermott, D., Ghallab, M., Howe, A., Knoblock, C., Ram, A., Veloso, M., Weld, D., Wilkins, D.: PDDL—the planning domain definition language. Tech. Rep. CVC TR98003/DCS TR1165, New Haven, CT: Yale Center for Computational Vision and Control (1998)
1998
-
[94]
In: 2024 International Conference on Learning Represen- tations (2024)
Mediratta, I., You, Q., Jiang, M., Raileanu, R.: A study of generalization in offline reinforcement learning. In: 2024 International Conference on Learning Represen- tations (2024)
2024
-
[95]
ACM Computing Surveys37(4), 316–344 (2005) 26 D.J.N.J
Mernik, M., Heering, J., Sloane, A.M.: When and how to develop domain-specific languages. ACM Computing Surveys37(4), 316–344 (2005) 26 D.J.N.J. Soemers et al
2005
-
[96]
Nature594, 207–212 (2021)
Mirhoseini, A., Goldie, A., Yazgan, M., Jiang, J.W., Songhori, E., Wang, S., Lee, Y.J., Johnson, E., Pathak, O., Nazi, A., Pak, J., Tong, A., Srinivasa, K., Hang, W., Tuncer, E., Le, Q.V., Laudon, J., Ho, R., Carpenter, R., Dean, J.: A graph placement methodology for fast chip...
2021
-
[97]
In: Palmer, M., Hwa, R., Riedel, S
Misra, D., Langford, J., Artzi, Y.: Mapping instructions and visual observations to actions with reinforcement learning. In: Palmer, M., Hwa, R., Riedel, S. (eds.) Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. pp. 1004–1015 (2017)
2017
-
[98]
In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
Mittel, A., Munukutla, P.S.: Visual transfer between Atari games using competi- tive reinforcement learning. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). pp. 499–501 (2019)
2019
-
[99]
https://arxiv
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., Riedmiller, M.: Playing Atari with deep reinforcement learning. https://arxiv. org/abs/1312.5602 (2013)
2013 arXiv
-
[100]
In: Faust, A., Garnett, R., White, C., Hutter, F., Gardner, J.R
Mohan, A., Benjamins, C., Wienecke, K., Dockhorn, A., Lindauer, M.: AutoRL hyperparameter landscapes. In: Faust, A., Garnett, R., White, C., Hutter, F., Gardner, J.R. (eds.) International Conference on Automated Machine Learning. Proceedings of Machine Learning Research, vol. ...
2023
-
[101]
Journal of Artificial Intelligence Research79, 1167– 1236 (2024)
Mohan, A., Zhang, A., Lindauer, M.: Structure in deep reinforcement learning: A survey and open problems. Journal of Artificial Intelligence Research79, 1167– 1236 (2024)
2024
-
[102]
Müller-Brockhausen, M., Preuss, M., Plaat, A.: Procedural content generation: Betterbenchmarksfortransferreinforcementlearning.In:Proceedingsofthe2021 IEEE Conference on Games. pp. 924–931 (2021)
2021
-
[103]
In: Proceedings of the 2019 International Conference on Learning Representations (2019)
Nagabandi, A., Clavera, I., Liu, S., Fearing, R.S., Abbeel, P., Levine, S., Finn, C.: Learning to adapt in dynamic, real-world environments through meta- reinforcement learning. In: Proceedings of the 2019 International Conference on Learning Representations (2019)
2019
-
[104]
https://arxiv.org/abs/1804.03720 (2018)
Nichol, A., Pfau, V., Hesse, C., Klimov, O., Schulman, J.: Gotta learn fast: A new benchmark for generalization in RL. https://arxiv.org/abs/1804.03720 (2018)
2018 arXiv
-
[105]
In: Reinforcement Learning Conference (2024), accepted
Obando-Ceron, J., Araú’jo, J.G.M., Courville, A., Castro, P.S.: On the consis- tency of hyper-parameter selection in value-based deep reinforcement learning. In: Reinforcement Learning Conference (2024), accepted
2024
-
[106]
In: Meila, M., Zhang, T
Obando-Ceron, J.S., Castro, P.S.: Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research. In: Meila, M., Zhang, T. (eds.) Proceedings of the 38th International Conference on Machine Learning. pp. 1373–
-
[107]
In: Proceedings of the 34th International Conference on Machine Learning
Oh, J., Singh, S., Lee, H., Kohli, P.: Zero-shot task generalization with multi-task deep reinforcement learning. In: Proceedings of the 34th International Conference on Machine Learning. pp. 2661–2670. PMLR (2017)
2017
-
[108]
https://openai.com/blog/chatgpt (2022), ac- cessed: 2024-01-02
OpenAI: Introducing ChatGPT. https://openai.com/blog/chatgpt (2022), ac- cessed: 2024-01-02
2022
-
[109]
In: Proceedings of the International Con- ference on Automated Planning and Scheduling
Oswald, J., Srinivas, K., Kokel, H., Lee, J., Katz, M., Sohrabi, S.: Large language models as planning domain generators. In: Proceedings of the International Con- ference on Automated Planning and Scheduling. vol. 34, pp. 423–431 (2024)
2024
-
[110]
In: Koyejo, S., A Research Agenda for Usability and Generalisation in RL 27 Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., Lowe, R.: Train- ing language models to follo...
2022
-
[111]
In: Proceedings of the 39th International Conference on Machine Learning
Parker-Holder, J., Jiang, M., Dennis, M., Samvelyan, M., Foerster, J., Grefen- stette, E., Rocktäschel, T.: Evolving curricula with regret-based environment de- sign. In: Proceedings of the 39th International Conference on Machine Learning. PMLR, vol. 162, pp. 17473–17498 (2022)
2022
-
[112]
Journal of Artificial Intelligence Research74, 517–568 (2022)
Parker-Holder, J., Rajan, R., Song, X., Biedenkapp, A., Miao, Y., Eimer, T., Zhang, B., Nguyen, V., Calandra, R., Faust, A., Hutter, F., Lindauer, M.: Auto- mated reinforcement learning (autoRL): A survey and open problems. Journal of Artificial Intelligence Research74, 517–568 (2022)
2022
-
[113]
https://arxiv.org/abs/2304.01315 (2023)
Patterson, A., Neumann, S., White, M., White, A.: Empirical design in reinforce- ment learning. https://arxiv.org/abs/2304.01315 (2023)
2023 arXiv
-
[114]
Stratega
Perez-Liebana, D., Dockhorn, A., Grueso, J.H., Jeurissen, D.: The design of "Stratega": A general strategy games framework. In: Osborn, J.C. (ed.) Joint Pro- ceedings of the AIIDE 2020 Workshops co-located with 16th AAAI Conference on Artificial Intelligence and Interactive Di...
2020
-
[115]
In: Giacomo, G.D., Catala, A., Dilkina, B., Milano, M., Barro, S., Bugarín, A., Lang, J
Piette, É., Soemers, D.J.N.J., Stephenson, M., Sironi, C.F., Winands, M.H.M., Browne, C.: Ludii – the ludemic general game system. In: Giacomo, G.D., Catala, A., Dilkina, B., Milano, M., Barro, S., Bugarín, A., Lang, J. (eds.) Proceedings of the 24th European Conference on Art...
2020
-
[116]
In: Proceedings of the 2021 IEEE Conference on Games (CoG)
Piette, É., Stephenson, M., Soemers, D.J.N.J., Browne, C.: General board game concepts. In: Proceedings of the 2021 IEEE Conference on Games (CoG). pp. 932–939. IEEE (2021)
2021
-
[117]
https://arxiv.org/abs/2307.01952 (2023)
Podell, D., English, Z., Lacey, K.,Blattmann, A., Dockhorn,T., Müller, J., Penna, J., Rombach, R.: SDXL: Improving latent diffusion models for high-resolution image synthesis. https://arxiv.org/abs/2307.01952 (2023)
2023 arXiv
-
[118]
In: Reinforcement Learning Conference (2025), to appear
Ponse, K., Kleuker, J.F., Moerland, T.M., Plaat, A.: Chargax: A JAX accelerated EV charging simulator. In: Reinforcement Learning Conference (2025), to appear
2025
-
[119]
In: Chaudhuri, K., Salakhutdinov, R
Rakelly, K., Zhou, A., Quillen, D., Finn, C., Levine, S.: Efficient off-policy meta- reinforcement learning via probabilistic context variables. In: Chaudhuri, K., Salakhutdinov, R. (eds.) Proceedings of the 36th International Conference on Machine Learning. Proceedings of Mac...
2019
-
[120]
https://arxiv.org/abs/2204
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., Chen, M.: Hierarchical text- conditional image generation with CLIP latents. https://arxiv.org/abs/2204. 06125 (2022)
2022
-
[121]
rewards: A comparative study of ob- jective specification mechanisms
Rani, S., Booth, S., Sreedharan, S.: Goals vs. rewards: A comparative study of ob- jective specification mechanisms. In: Reinforcement Learning Conference (2025), to appear
2025
-
[122]
https://arxiv
Raparthy, S.C., Hambro, E., Kirk, R., Henaff, M., Raileanu, R.: Generalization to new sequential decision making tasks with in-context learning. https://arxiv. org/abs/2312.03801 (2023)
2023 arXiv
-
[123]
Transactions on Machine Learning Research (2023) 28 D.J.N.J
Reed, S., Żołna, K., Parisotto, E., Colmenarejo, S.G., Novikov, A., Barth-Maron, G., Giménez, M., Sulsky, Y., Kay, J., Springenberg, J.T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., de Freitas, N.: A generalist a...
2023
-
[124]
In: Machine Learning: ECML 2005
Riedmiller, M.: Neural fitted Q iteration - first experiences with a data efficient neural reinforcement learning method. In: Machine Learning: ECML 2005. Lec- ture Notes in Computer Science, vol. 3720, pp. 317–328. Springer (2005)
2005
-
[125]
In: International Conference on Learning Representations (2024)
Rigter, M., Jiang, M., Posner, I.: Reward-free curricula for training robust world models. In: International Conference on Learning Representations (2024)
2024
-
[126]
In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J
Rodriguez-Sanchez, R., Spiegel, B.A., Wang, J., Patel, R., Tellex, S., Konidaris, G.: RLang: A declarative language for describing partial world knowledge to rein- forcement learning agents. In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J. (eds....
2023
-
[127]
In: Proceedings of the 41st International Conference on Machine Learning
Rolnick, D., Aspuru-Guzik, A., Beery, S., Dilkina, B., Donti, P.L., Ghassemi, M., Kerner, H., Monteleoni, C., Rolf, E., Tambe, M., White, A.: Position: Application- driven innovation in machine learning. In: Proceedings of the 41st International Conference on Machine Learning....
2024
-
[128]
Journal of Artificial Intelligence Research 67, 673–703 (2020)
Rostami, M., Isele, D., Eaton, E.: Using task descriptions in lifelong machine learning for improved performance and zero-shot transfer. Journal of Artificial Intelligence Research 67, 673–703 (2020)
2020
-
[129]
https: //arxiv.org/abs/1606.04671 (2016)
Rusu, A.A., Rabinowitz, N.C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., Hadsell, R.: Progressive neural networks. https: //arxiv.org/abs/1606.04671 (2016)
2016 arXiv
-
[130]
https://arxiv.org/abs/2311.10090 (2023)
Rutherford, A., Ellis, B., Gallici, M., Cook, J., Lupu, A., Ingvarsson, G., Willi, T., Khan, A., de Witt, C.S., Souly, A., Bandyopadhyay, S., Samvelyan, M., Jiang, M., Lange, R.T., Whiteson, S., Lacerda, B., Hawes, N., Rocktäschel, T., Lu, C., Foerster, J.N.: JaxMARL: Multi-ag...
2023 arXiv
-
[131]
In: Proceedings of IEEE International Conference on Robotics and Automation (ICRA) (2025), to appear
Sakçak, B., Shell, D.A., O’Kane, J.M.: Limits of specifiability for sensor-based robotic planning tasks. In: Proceedings of IEEE International Conference on Robotics and Automation (ICRA) (2025), to appear
2025
-
[132]
In: International Conference on Learning Representations (2023)
Samvelyan, M., Khan, A., Dennis, M., Jiang, M., Parker-Holder, J., Foerster, J., Raileanu, R., Rocktäschel, T.: MAESTRO: Open-ended environment design for multi-agent reinforcement learning. In: International Conference on Learning Representations (2023)
2023
-
[133]
In: Advances in Neural Information Processing Systems (2021)
Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Küttler, H., Grefenstette, E., Rocktäschel, T.: Minihack the planet: A sandbox for open-ended reinforcement learning research. In: Advances in Neural Information Processing Systems (2021)
2021
-
[134]
In: Proceedings of the IEEE Conference on Computational Intelligence in Games
Schaul, T.: A video game description language for model-based or interactive learning. In: Proceedings of the IEEE Conference on Computational Intelligence in Games. pp. 193–200. IEEE (2013)
2013
-
[135]
In: Proceedings of the 32nd International Conference on Machine Learning
Schaul,T.,Horgan,D.,Gregor,K.,Silver,D.:Universalvaluefunctionapproxima- tors. In: Proceedings of the 32nd International Conference on Machine Learning. JLMR: W&CP, vol. 37, pp. 1312–1320 (2015)
2015
-
[136]
IEEE Transac- tions on Computational Intelligence and AI in Games6(4), 325–331 (Dec 2014)
Schaul, T.: An extensible description language for video games. IEEE Transac- tions on Computational Intelligence and AI in Games6(4), 325–331 (Dec 2014). https://doi.org/10.1109/TCIAIG.2014.2352795
2014
-
[137]
Schmidhuber, J.: On learning how to learn learning strategies. Tech. Rep. FKI- 198-94, Institut für Informatik, Technische Universität München (1994)
1994
-
[138]
In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society
Seger, E., Ovadya, A., Siddarth, D., Garfinkel, B., Dafoe, A.: Democratising AI: Multiple meanings, goals, and methods. In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society. pp. 715–722 (2023) A Research Agenda for Usability and Generalisation in RL 29
2023
-
[139]
In: International Conference on Learning Representations (2018)
Shu, T., Xiong, C., Socher, R.: Hierarchical and interpretable skill acquisition in multi-task reinforcement learning. In: International Conference on Learning Representations (2018)
2018
-
[140]
In: ICAPS Workshop on Bridging the Gap Between AI Planning and Reinforcement Learning (PRL) (2020)
Silver, T., Chitnis, R.: PDDLGym: Gym environments from PDDL problems. In: ICAPS Workshop on Bridging the Gap Between AI Planning and Reinforcement Learning (PRL) (2020)
2020
-
[141]
it is there, and you need it, so why do you not use it?
Simkute, A., Luger, E., Evans, M., Jones, R.: “it is there, and you need it, so why do you not use it?” achieving better adoption of AI systems by domain experts, in the case study of natural science research. https://arxiv.org/abs/2403.16895 (2024)
2024 arXiv
-
[142]
https://arxiv.org/abs/1807.11074 (2018)
Sobol, D., Wolf, L., Taigman, Y.: Visual analogies between atari games for study- ing transfer learning in rl. https://arxiv.org/abs/1807.11074 (2018)
2018 arXiv
-
[143]
ICGA Journal43(3), 146–161 (2022)
Soemers, D.J.N.J., Mella, V., Browne, C., Teytaud, O.: Deep learning for general game playing with Ludii and Polygames. ICGA Journal43(3), 146–161 (2022)
2022
-
[144]
Transactions on Machine Learning Research (2023)
Soemers, D.J.N.J., Mella, V., Piette, É., Stephenson, M., Browne, C., Teytaud, O.: Towards a general transfer approach for policy-value networks. Transactions on Machine Learning Research (2023)
2023
-
[145]
In: Proceedings of the 2024 IEEE Conference on Games
Soemers, D.J.N.J., Piette, É., Stephenson, M., Browne, C.: The Ludii game de- scription language is universal. In: Proceedings of the 2024 IEEE Conference on Games. pp. 1–8 (2024)
2024
-
[146]
In: Rocha, A.P., Steels, L., van den Herik, H.J
Soemers, D.J.N.J., Samothrakis, S., Driessens, K., Winands, M.H.M.: Environ- ment descriptions for usability and generalisation in reinforcement learning. In: Rocha, A.P., Steels, L., van den Herik, H.J. (eds.) Proceedings of the 17th In- ternational Conference on Agents and A...
2025
-
[147]
https://stability.ai/research/stable-audio-efficient-timing-latent-diffusion (2023), accessed: 2024-1-4
Stability AI: Stable audio: Fast timing-conditioned latent audio diffusion. https://stability.ai/research/stable-audio-efficient-timing-latent-diffusion (2023), accessed: 2024-1-4
2023
-
[148]
In: Browne, C., Kishimoto, A., Schaeffer, J
Stephenson, M., Soemers, D.J.N.J., Piette, É., Browne, C.: Measuring board game distance. In: Browne, C., Kishimoto, A., Schaeffer, J. (eds.) Computers and Games. CG 2022. Lecture Notes in Computer Science, vol. 13865, pp. 121–130. Springer, Cham (2023)
2023
-
[149]
https: //arxiv.org/abs/2101.02722 (2021)
Stone, A., Ramirez, O., Konolige, K., Jonschkowski, R.: The distracting control suite – a challenging benchmark for reinforcement learning from pixels. https: //arxiv.org/abs/2101.02722 (2021)
2021 arXiv
-
[150]
In: International Confer- ence on Learning Representations (2020)
Sun, S.H., Wu, T.L., Lim, J.J.: Program guided agent. In: International Confer- ence on Learning Representations (2020)
2020
-
[151]
MIT Press, Cambridge, MA, 2 edn
Sutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction. MIT Press, Cambridge, MA, 2 edn. (2018)
2018
-
[152]
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T., Riedmiller, M.: DeepMind control suite (2018)
2018
-
[153]
In: Mahadevan, S
Taylor, M.E., Stone, P.: Transfer learning for reinforcement learning domains: A survey. In: Mahadevan, S. (ed.) Journal of Machine Learning Research. vol. 10, pp. 1633–1685 (2009)
2009
-
[154]
In: Advances in Neural Information Processing Systems
Terry, J.K., Black, B., Grammel, N., Jayakumar, M., Hari, A., Sullivan, R., San- tos,L.,Perez,R.,Horsch,C.,Dieffendahl,C.,Williams,N.L.,Lokesh,Y.,Ravi,P.: Pettingzoo: A standard API for multi-agent reinforcement learning. In: Advances in Neural Information Processing Systems. ...
2021
-
[155]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D., Mannor, S.: A deep hierar- chical approach to lifelong learning in minecraft. In: Proceedings of the AAAI Conference on Artificial Intelligence. pp. 1553–1561. AAAI (2017)
2017
-
[156]
In: Proceedings of the Twenty-second International Joint Conference on Artificial Intelligence, IJCAI-11
Thielscher, M.: The general game playing description language is universal. In: Proceedings of the Twenty-second International Joint Conference on Artificial Intelligence, IJCAI-11. pp. 1107–1112 (2011)
2011
-
[157]
In: Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C
Todd, G., Padula, A., Stephenson, M., Piette, É., Soemers, D.J.N.J., Togelius, J.: GAVEL: Generating games via evolution and language models. In: Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C. (eds.) Advances in Neural Information Processi...
2024
-
[158]
https://arxiv.org/abs/ 2506.22609 (2025)
Todd, G., Padula, A.G., Soemers, D.J.N.J., Togelius, J.: Ludax: A GPU- accelerated domain specific language for board games. https://arxiv.org/abs/ 2506.22609 (2025)
2025
-
[159]
https://arxiv.org/ abs/2101.04808 (2021)
Trofin, M., Qian, Y., Brevdo, E., Lin, Z., Choromanski, K., Li, D.: MLGO: a machine learning guided compiler optimizations framework. https://arxiv.org/ abs/2101.04808 (2021)
2021 arXiv
-
[160]
In: Finding the Frame workshop @ Reinforcement Learning Conference (2024)
Voelcker, C., Hussing, M., Eaton, E.: Can we hop in general? A discussion of benchmark selection and design using the Hopper environment. In: Finding the Frame workshop @ Reinforcement Learning Conference (2024)
2024
-
[161]
van der Wal, J.: Stochastic Dynamic Programming. No. 139 in Mathematical Centre tracts, Morgan Kaufmann, Amsterdam (1981)
1981
-
[162]
In: 2011 IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning (ADPRL)
Whiteson, S., Tanner, B., Taylor, M.E., Stone, P.: Protecting against evalua- tion overfitting in empirical reinforcement learning. In: 2011 IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning (ADPRL). pp. 120–127 (2011)
2011
-
[163]
In: 2018 IEEE International Conference on Robotics and Automation
Williams, E.C., Gopalan, N., Rhee, M., Tellex, S.: Learning to parse natural language to gronuded reward functions with weak supervision. In: 2018 IEEE International Conference on Robotics and Automation. pp. 4430–4436 (2018)
2018
-
[164]
In: Proceedings of the 24th International Confer- ence on Machine Learning
Wilson, A., Fern, A., Ray, S., Tadepalli, P.: Multi-task reinforcement learning: A hierarchical Bayesian approach. In: Proceedings of the 24th International Confer- ence on Machine Learning. pp. 1015–1022 (2007)
2007
-
[165]
https:// arxiv.org/abs/2101.02230 (2021)
Yang, K.: Learn dynamic-aware state embedding for transfer learning. https:// arxiv.org/abs/2101.02230 (2021)
2021 arXiv
-
[166]
In: International Confer- ence on Learning Representations (2024)
Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Kaelbling, L., Schuurmans, D., Abbeel, P.: Learning interactive real-world simulators. In: International Confer- ence on Learning Representations (2024)
2024
-
[167]
In: Proceedings of the International Conference on Learning Representations (2021)
Yoon, D., Hong, S., Lee, B.J., Kim, K.E.: Winning the L2RPN challenge: Power grid management via semi-Markov afterstate actor critic. In: Proceedings of the International Conference on Learning Representations (2021)
2021
-
[168]
https://arxiv.org/abs/1806.07937 (2020)
Zhang, A., Ballas, N., Pineau, J.: A dissection of overfitting and generalization in continuous reinforcement learning. https://arxiv.org/abs/1806.07937 (2020)
2020 arXiv
-
[169]
https://arxiv.org/abs/1804.06893 (2018)
Zhang, C., Vinyals, O., Munos, R., Bengio, S.: A study on overfitting in deep reinforcement learning. https://arxiv.org/abs/1804.06893 (2018)
2018 arXiv
-
[170]
In: 2020 IEEE Symposium Series on Com- putational Intelligence (SSCI)
Zhao, W., Queralta, J.P., Westerlund, T.: Sim-to-real transfer in deep reinforce- ment learning for robotics: A survey. In: 2020 IEEE Symposium Series on Com- putational Intelligence (SSCI). pp. 737–744 (2020)
2020
-
[171]
https://arxiv.org/abs/2303.18223 (2023), accessed: 25-01-2024 A Research Agenda for Usability and Generalisation in RL 31
Zhao, W.X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.Y., Wen, J.R.: A survey of large language models. https://arxiv.org/abs/2303...
2023 arXiv
-
[172]
In: International Conference on Learning Repre- sentations (2020)
Zhong, V., Rocktäschel, T., Grefenstette, E.: RTFM: Generalising to novel envi- ronment dynamics via reading. In: International Conference on Learning Repre- sentations (2020)
2020
-
[173]
https://arxiv
Zhu, Z., de Salvo Braz, R., Bhandari, J., Jiang, D., Wan, Y., Efroni, Y., Wang, L., Xu, R., Guo, H., Nikulkov, A., Korenkevych, D., Dogan, U., Cheng, F., Wu, Z., Xu, W.: Pearl: A production-ready reinforcement learning agent. https://arxiv. org/abs/2312.03814 (2023)
2023 arXiv
-
[174]
IEEE Transactions on Pattern Analysis and Machine Intelli- gence 45(11), 13344–13362 (2023)
Zhu, Z., Lin, K., Jain, A.K., Zhou, J.: Transfer learning in deep reinforcement learning: A survey. IEEE Transactions on Pattern Analysis and Machine Intelli- gence 45(11), 13344–13362 (2023)
2023
-
[175]
Zintgraf, L.: Fast Adaptation via Meta Reinforcement Learning. Ph.D. thesis, University of Oxford, Oxford, United Kingdom (2022)
2022
-
[176]
https://arxiv
Zuo, M., Velez, F.P., Li, X., Littman, M.L., Bach, S.H.: Planetarium: A rigorous benchmark for translating text to structured planning languages. https://arxiv. org/abs/2407.03321 (2024)
2024
-
[1708]
AAAI Press (2020) A Research Agenda for Usability and Generalisation in RL 23
2020
-
[8045]
Curran Associates, Inc. (2021)
2021
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.