Pith. sign in

REVIEW 4 major objections 5 minor 299 references

Playstyle and Artificial Intelligence: An Initial Blueprint Through the Lens of Video Games

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This dissertation argues that playstyle — an agent's consistent pattern of decision-making — can be measured by one label-free metric: the divergence of two agents' action distributions on the discrete states both have visited, and that thi

desk verdict The dissertation's playstyle metric is a real but not-new contribution; the 'general' claim outruns the evidence, yet the framework and ablations deserve refereeing. read the letter →

arxiv 2508.19152 v1 pith:NRSSFURU submitted 2025-08-26 cs.AI cs.LGcs.MAcs.SC

classification cs.AIcs.LGcs.MAcs.SC
keywords playstyledecision-makingstyleHierarchicalStateDiscretizationDistancemeasurementreinforcementlearningstrategicdiversityagentbehavior
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This dissertation tries to establish that 'playstyle' — the consistent way an agent makes decisions when several viable choices exist — is a measurable, transferable property of intelligent agents, not an aesthetic label. The core proposal is a general metric: a learned encoder maps observations into discrete symbolic states, the comparison keeps only the states two agents have both visited, and playstyle distance is the divergence between their action distributions conditioned on those shared states. The dissertation reports that this metric separates 25 rule-based racing-agent configurations, human-defined style dimensions in a second driving game, and four deep reinforcement learning algorithms across seven Atari games, with no task-specific labels or handcrafted features. On the same foundation it builds quantitative notions of strategic diversity and competitive balance, and argues that style — as an expression of belief and value — deserves a place beside rationality in how we model, train, and evaluate intelligent agents.

What carries the argument

The load-bearing object is the Hierarchical State Discretization (HSD) encoder, a vector-quantized variational autoencoder whose discrete codes are organized into a multiscale hierarchy and trained with two joint objectives: an observation reconstruction loss that preserves perceptual capacity, and a policy prediction loss that forces codes to retain decision-relevant semantics. On top of it, Playstyle Distance compares only the intersection of symbolic states two agents have both visited, aggregating local 2-Wasserstein distances between their action distributions with expected averaging. A multiscale composition of three granularities (coarse singleton, intermediate 2^20 codes, fine hierar

What would settle it

A dual test would settle it. First, self-distance: split one agent's trajectory in half and compute the metric between halves with the same encoder — it should be near zero and far below the distance between different policies; an equal or larger self-distance means the metric tracks sampling noise, not style. Second, oracle consistency: build three rule-based TORCS agents that human observers unanimously rate as A ≈ B in style and C clearly different; the paper's stated consistency criterion requires d(A,C) < d(B,C) to be reproduced, so a single inversion would falsify the semantic-alignment

Watch

Extended reading notes

Core claim

Playstyle is defined as the decision-making style of an agent expressed through interaction with a responsive environment, where multiple viable choices exist and external consequences give actions meaning. The central operational claim is that this latent construct can be measured by Playstyle Distance: project both agents' observations into a discrete symbolic state space via Hierarchical State Discretization (HSD), intersect the states both agents visited, and take the expected 2-Wasserstein divergence between their empirical action distributions on that intersection. Because style manifests in how agents choose among comparable situations, restricting comparison to genuinely shared conte

Load-bearing premise

The load-bearing assumption is that the discrete symbolic states learned from human gameplay mean the same decision context for every agent, so that comparing two agents' action choices only on the states both happened to visit is a fair and sufficient basis for measuring their stylistic difference.

Editorial extensions

If this is right

  • Playstyle measurement becomes label-free and domain-agnostic: the same encoder-plus-intersection procedure works for rule-based bots, human players, and learned agents, across racing games and Atari.
  • Strategic diversity and competitive balance become quantifiable on the same distance: the dissertation defines diverse trajectory counts, Top-D diversity, and Top-B balance directly from the playstyle metric.
  • Style becomes a trainable target: the capacity and popularity dimensions can be folded into loss functions, letting reinforcement and imitation learning produce agents with specified stylistic tendencies.
  • Imitation fidelity can be scored as playstyle distance between an expert and a learner, a process-level complement to outcome-based measures of success.
  • If style is a core dimension of intelligent behavior, modeling playstyle becomes a candidate ingredient for agents built to express recognizable values, not just optimize rewards.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The intersection strategy is a generic answer to the state-alignment problem in behavior comparison; it should transfer to any sequential decision domain (robotics, driving, human-computer interaction) where a shared discrete state space can be learned, though the dissertation trains a fresh encoder per environment, leaving single-encoder cross-environment transfer as an open, testable question.
  • The exponential similarity kernel predicts a specific psychophysical signature: human similarity judgments between playstyle traces should track e^{-d} rather than raw distance; a paired user study could confirm or reject that cognitive grounding.
  • The capacity-popularity axes suggest an operational definition of 'meta shifts' in live games: tracking how a population's styles move through this two-dimensional space over time would turn a qualitative community term into a measurable quantity.
  • Because Playstyle Distance is a metric, it induces a geometry on the space of agents; interpolating between policies would yield 'midpoint styles,' a natural and untested route to style blending and novel style generation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This dissertation-style manuscript proposes 'playstyle' as a decision-making style exhibited by agents in interactive environments, and develops a conceptual blueprint spanning philosophy, measurement, expression, and applications. The central technical contribution is a label-free Playstyle Distance metric based on Hierarchical State Discretization (HSD): observations are mapped to discrete symbolic states, and two agents are compared by the divergence of their empirical action distributions over the intersection of visited discrete states. Experiments are reported across three domains: TORCS (rule-based agents), RGSK (six human participants), and seven Atari games (RL agents trained with DQN, C51, Rainbow, IQN). The metric is extended to multiscale states, a probabilistic similarity kernel, continuous playstyle spectra, and downstream diversity/balance indicators. The manuscript also discusses imitation learning and human-likeness, and concludes with speculative connections to AGI and 'soul' digitization.

Significance. If the proposed Playstyle Distance is valid, it would constitute a notable step toward a general, label-free, transferable measurement of decision-making style, with applications in game AI, player modeling, and agent diversity. The paper's strengths include a concrete algorithmic proposal (HSD with dual reconstruction/policy decoders), a relevant ablation showing that policy supervision is necessary for usable intersections (Table 5), and experiments spanning rule-based, human, and learning-based agents. However, the generality claim is considerably broader than the evidence: the human sample is very small, semantic alignment of discrete states across agents is assumed rather than tested, and several reported accuracy numbers are inconsistent across tables. The core construct-validity question is therefore left open.

major comments (4)
  1. [§6.5.1–6.5.2, Eqs. defining S_φ(A,B) and d_φ(A,B|s)] The Playstyle Distance is defined as the divergence of action distributions on the intersection of HSD discrete states. This only measures playstyle if the same symbolic state in two agents' datasets refers to the same decision context. The evidence offered—CIFAR-10 classification (Table 2), state utilization (Table 3), noise robustness (Table 4), and the policy-decoder ablation (Table 5)—shows that HSD codes preserve object-class information and are influenced by policy supervision, but it does not demonstrate cross-agent semantic alignment. Intersection sizes are reported (e.g., 187.53 with the policy decoder in Table 5), but no coverage statistics, per-state sample counts, or tests of state-conditional action correspondence are given. Since the metric's validity depends entirely on this alignment, the 'general metric' claim is under-supported. Please add direct evidence: distribution
  2. [§6.5.2, 'Evaluation on RGSK' and Table 8] The human-player evaluation uses only six participants, each instructed to maintain distinct playstyles. With such a small, instruction-conditioned sample, near-perfect accuracies (93–99%) largely reflect the experimental setup rather than generalizable human playstyle discrimination. The manuscript does not report per-participant variance, leave-one-participant-out results, or any statistical significance. This is load-bearing because the abstract claims a 'general' metric across human, rule-based, and RL agents, and the human evidence is a cornerstone of that claim. I recommend additional participants, cross-participant validation, and explicit reporting of variability.
  3. [Table 13 vs Tables 7–8] There is a substantial unexplained discrepancy in the central accuracy numbers. Table 13 reports single-scale 220 (t=2) accuracy of 73.3±8.2% for TORCS and 79.2±7.9% for RGSK, whereas Tables 7 and 8 report 91.20% (5 Speed) and 93.11% (Nitro) for what appears to be the same configuration. If the tasks differ (e.g., 25-class aggregate vs 5-class per-dimension), this must be stated explicitly. As written, a reader cannot determine which numbers support the main claim. Given that the accuracy tables are the primary experimental evidence, this inconsistency must be resolved.
  4. [§6.5.4, Eq. (6.3) and surrounding text] The exponential perceptual kernel P(d)=1/e^d is asserted on the basis of a proof 'provided in the Appendix' of a prior paper, but no derivation appears in this manuscript. The similarity measure P S∩ is a core extension, and the kernel's functional form and the normalizer D_M^Φ are not derived or validated here. Please include the derivation in full, or explicitly state the kernel as a modeling assumption and justify it with data. As written, this is a load-bearing gap in the similarity extension.
minor comments (5)
  1. [§11.2] The heading 'Video Game Industry' appears twice consecutively, disrupting the section numbering and flow.
  2. [§6.5.4] The notation P S∩ and D_M^Φ is visually cluttered and not clearly defined in the main text. Please clarify the indexing and the exact set over which the sum is taken.
  3. [§6.5.2, Table 12 and surrounding text] MKL is called 'approx. JS' in the table but the text states it is not equivalent to Jensen-Shannon divergence. Reconcile this label or rename the metric.
  4. [Throughout] There are minor typographical issues, such as 'V AEs' in §2.4 and irregular spacing around mathematical symbols. A copyedit pass would improve readability.
  5. [Reproducibility] No code, data, or trained encoders are made available. Given the centrality of the HSD training procedure and the transfer claims, a public release of the evaluation code and datasets would substantially strengthen the contribution.

Circularity Check

1 steps flagged · score 2.0 of 10

Core playstyle-distance derivation is self-contained; minor load-bearing self-citation for the perceptual similarity kernel.

  1. self citation load bearing [Section 6.5.4, 'Perception of Similarity' (around Eq. 6.3)]
    "This perceptual relation is the only relation under our assumptions from human cognition and probability. The original paper by Lin et al. [181] (2024) provides a proof using differential equations in the Appendix."

    The exponential similarity kernel P(d)=1/e^d is asserted to be the unique consequence of the stated assumptions, but the proof is not reproduced in this dissertation; it is deferred to the authors' own prior work (Lin et al., 2024). The uniqueness claim is therefore load-bearing on a self-citation rather than on a self-contained derivation. This is a minor circularity concern because the main Playstyle Distance accuracy results and the HSD-based comparisons do not depend on this perceptual kernel; the kernel primarily underlies the Playstyle Similarity extension.

full rationale

The central measurement claim—Playstyle Distance as the divergence between action distributions conditioned on HSD-discretized intersection states—is not circular. The HSD encoder is trained on human TORCS gameplay without style labels; the style dimensions and agent identities are used only as held-out evaluation targets. Accuracy is reported on rule-based TORCS agents, human RGSK players, and Atari RL agents, so the main results are not fitted to the target labels. The state-space size and intersection threshold are selected on TORCS and then transferred to RGSK and Atari; this is hyperparameter transfer rather than a fitted-input-called-prediction reduction. The only notable issue is the perceptual similarity kernel in Section 6.5.4, whose uniqueness proof is deferred to a self-citation; this is ancillary to the core Playstyle Distance validation. Overall, the paper shows no significant circularity in its central derivation.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The dissertation introduces no new physical entities such as particles, forces, or dimensions. Its conceptual constructs, including 'style boundary', 'dual loops', and 'self as belief system', are analytical devices rather than entities with independent falsifiable handles. The free parameters listed above are the most consequential data-dependent choices in the measurement pipeline.

free parameters (4)
  • HSD state space size (default) = 2^20 = 1,048,576 states
    Chosen after TORCS experiments as the default symbolic state space, balancing capacity and intersection stability.
  • Intersection threshold t = 2
    Selected on TORCS to filter rare or unstable states, then fixed as the default for RGSK and Atari.
  • Playstyle similarity normalizer D_M^Phi = Computed as the average of all observed pairwise distances in the comparison set
    Used in Equation 6.3 to rescale distances so the expected distance converges to 1; fitted to the same data being compared.
  • Multiscale mixture weights = Equal weights over encoders {1, 2^20 HSD, 256-res HSD}
    The multiscale distance in Equation 6.1 averages encoders with equal weights; this is a design choice rather than a tuned parameter.
assumptions (5)
  • ad hoc to paper Playstyle Consistency: preference comparisons are anchored to a current belief state b via PF(b', b), avoiding global transitivity or ordering.
    Introduced in Section 5.4.1 to make the agent's internal decision loop self-consistent; it is an assumption proposed by the paper, not an established result.
  • domain assumption Beliefs are foundational, unprovable premises that initiate action, and the self is an internal system of beliefs.
    Chapter 1.2; a philosophical premise used to ground the playstyle framework, not empirically verified.
  • domain assumption Human perception of similarity follows a logarithmic or exponential Weber-Fechner relationship, leading to P(d) = exp(-d).
    Section 6.5.4; the proof of uniqueness is deferred to an appendix of the original paper that is not included in this arXiv version.
  • domain assumption The environment is a 'perfect generator': states produced by environment interaction are inherently valid, so GAN-style fake/real evaluation is inapplicable.
    Section 5.5.1; this framing motivates the choice of distance-based playstyle metrics over generative-style fidelity metrics.
  • ad hoc to paper The dual-loop model (external interaction loop plus internal deliberation loop) underlies all agent decision-making and style formation.
    Chapters 1.3 and 5.4.3; a structural assumption proposed by the dissertation to organize playstyle analysis.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Playstyle and Artificial Intelligence: An Initial Blueprint Through the Lens of Video Games." pith.science (2026). https://pith.science/paper/NRSSFURU

@misc{pith2026250819152,
  author       = {Pith},
  title        = {Pith review of: Playstyle and Artificial Intelligence: An Initial Blueprint Through the Lens of Video Games},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NRSSFURU}},
  note         = {Machine review of arXiv:2508.19152}
}
read the original abstract

Contemporary artificial intelligence (AI) development largely centers on rational decision-making, valued for its measurability and suitability for objective evaluation. Yet in real-world contexts, an intelligent agent's decisions are shaped not only by logic but also by deeper influences such as beliefs, values, and preferences. The diversity of human decision-making styles emerges from these differences, highlighting that "style" is an essential but often overlooked dimension of intelligence. This dissertation introduces playstyle as an alternative lens for observing and analyzing the decision-making behavior of intelligent agents, and examines its foundational meaning and historical context from a philosophical perspective. By analyzing how beliefs and values drive intentions and actions, we construct a two-tier framework for style formation: the external interaction loop with the environment and the internal cognitive loop of deliberation. On this basis, we formalize style-related characteristics and propose measurable indicators such as style capacity, style popularity, and evolutionary dynamics. The study focuses on three core research directions: (1) Defining and measuring playstyle, proposing a general playstyle metric based on discretized state spaces, and extending it to quantify strategic diversity and competitive balance; (2) Expressing and generating playstyle, exploring how reinforcement learning and imitation learning can be used to train agents exhibiting specific stylistic tendencies, and introducing a novel approach for human-like style learning and modeling; and (3) Practical applications, analyzing the potential of these techniques in domains such as game design and interactive entertainment. Finally, the dissertation outlines future extensions, including the role of style as a core element in building artificial general intelligence (AGI).

Figures

Figures reproduced from arXiv: 2508.19152 by the authors.

Figure 1
Figure 1. This diagram illustrates the internal structure and outward expression of an intelligent [PITH_FULL_IMAGE:figures/full_fig_p019_1.png] view at source ↗
Figure 2
Figure 2. A conceptual style map based on Capacity and Popularity. Each vector denotes a [PITH_FULL_IMAGE:figures/full_fig_p038_2.png] view at source ↗
Figure 3
Figure 3. The interplay between the three key components of playstyle: Measurement, Expres [PITH_FULL_IMAGE:figures/full_fig_p052_3.png] view at source ↗
Figures from the paper (32 more)
Figure 4
Figure 4. Figure 4: Illustration of the Hierarchical State Discretization (HSD) framework. Each hierarchy [PITH_FULL_IMAGE:figures/full_fig_p092_4.png]
Figure 5
Figure 5. Figure 5: Examples of CIFAR-10 images with increasing levels of Gaussian noise (Std = 0, 2, [PITH_FULL_IMAGE:figures/full_fig_p094_5.png]
Figure 6
Figure 6. Figure 6: Degree of Similarity: This demonstrates how multiple candidate points C can share [PITH_FULL_IMAGE:figures/full_fig_p106_6.png]
Figure 7
Figure 7. Figure 7: Comparison of Efficacy: Probabilistic vs. Distance Approaches. The plot illustrates [PITH_FULL_IMAGE:figures/full_fig_p111_7.png]
Figure 8
Figure 8. Figure 8: Transforming Playstyle Distance into Playstyle Similarity. The pipeline begins by [PITH_FULL_IMAGE:figures/full_fig_p114_8.png]
Figure 9
Figure 9. Figure 9: Evaluation of different playstyle metrics in TORCS, RGSK, and Atari. Accuracy [PITH_FULL_IMAGE:figures/full_fig_p117_9.png]
Figure 10
Figure 10. Figure 10: Heatmap of Playstyle Similarity (mix) using Speed60N0 as the reference. Rows [PITH_FULL_IMAGE:figures/full_fig_p118_10.png]
Figure 11
Figure 11. Figure 11: Accuracy comparison of potential unsupervised similarity measures on TORCS [PITH_FULL_IMAGE:figures/full_fig_p120_11.png]
Figure 12
Figure 12. Figure 12: Accuracy comparison of potential unsupervised similarity measures on Atari games. [PITH_FULL_IMAGE:figures/full_fig_p121_12.png]
Figure 13
Figure 13. Figure 13: Game environments used in this study. From left to right: TORCS, RGSK, Atari, [PITH_FULL_IMAGE:figures/full_fig_p122_13.png]
Figure 14
Figure 14. Figure 14: An illustration of counter relationships in games. Player A may have a higher Elo [PITH_FULL_IMAGE:figures/full_fig_p142_14.png]
Figure 15
Figure 15. Figure 15: Neural Rating Table (NRT) predicts win probabilities between two teams using [PITH_FULL_IMAGE:figures/full_fig_p145_15.png]
Figure 16
Figure 16. Figure 16: Effect of VQ Mean Loss on codebook utilization. Left: Without VQ Mean Loss, [PITH_FULL_IMAGE:figures/full_fig_p148_16.png]
Figure 17
Figure 17. Figure 17: Architecture of the Neural Counter Table [PITH_FULL_IMAGE:figures/full_fig_p149_17.png]
Figure 18
Figure 18. Figure 18: Illustration of the learning pipeline for the Neural Rating Table ( [PITH_FULL_IMAGE:figures/full_fig_p150_18.png]
Figure 19
Figure 19. Figure 19: Scalar rating (NRT) for 45 civilizations in Age of Empires II. [PITH_FULL_IMAGE:figures/full_fig_p164_19.png]
Figure 20
Figure 20. Figure 20: Learned 9 × 9 counter table for Age of Empires II. Each cell estimates the counter residual between categories. 152 [PITH_FULL_IMAGE:figures/full_fig_p165_20.png]
Figure 21
Figure 21. Figure 21: Rating table for Hearthstone decks, showcasing the relative strength of top [PITH_FULL_IMAGE:figures/full_fig_p168_21.png]
Figure 22
Figure 22. Figure 22: The 9 × 9 counter table for Hearthstone, visualizing inter-category counter relation￾ships. Green indicates advantage, red indicates disadvantage. Control decks (Category 1) lose to OTK finishers (Category 8) but handle board-centric decks (Category 4) well. Aggro (Ca…
Figure 23
Figure 23. Figure 23: An example of extending the classical Rock-Paper-Scissors to more complex cases. [PITH_FULL_IMAGE:figures/full_fig_p173_23.png]
Figure 24
Figure 24. Figure 24: Illustration of the EM algorithm. Residual win values for all categories are learned [PITH_FULL_IMAGE:figures/full_fig_p176_24.png]
Figure 25
Figure 25. Figure 25: A taxonomy of four canonical decision problem types, defined by whether actions [PITH_FULL_IMAGE:figures/full_fig_p189_25.png]
Figure 26
Figure 26. Figure 26: Seven environmental dimensions influencing decision complexity and playstyle di [PITH_FULL_IMAGE:figures/full_fig_p190_26.png]
Figure 27
Figure 27. Figure 27: A visual summary of major developments in machine learning, deep learning, and [PITH_FULL_IMAGE:figures/full_fig_p193_27.png]
Figure 28
Figure 28. Figure 28: Mean Human Normalized Scores (HNS) show some agents, like [PITH_FULL_IMAGE:figures/full_fig_p207_28.png]
Figure 29
Figure 29. Figure 29: Median HNS provides a more robust view by downweighting score inflation from [PITH_FULL_IMAGE:figures/full_fig_p208_29.png]
Figure 30
Figure 30. Figure 30: SABER Mean scores cap extreme outliers. LBC (2023) achieves 76.67%, with Agent57 close behind at 76.26%. wide range of DRL algorithms—spanning from classic baselines to state-of-the-art architec￾tures—have been evaluated on this platform. Among them, some were specifi…
Figure 31
Figure 31. Figure 31: SABER Median reveals stronger generality. LBC remains state-of-the-art at 64.93%, [PITH_FULL_IMAGE:figures/full_fig_p210_31.png]
Figure 32
Figure 32. Figure 32: A conceptual taxonomy of imitation learning paradigms, categorized by the agent’s [PITH_FULL_IMAGE:figures/full_fig_p217_32.png]
Figure 33
Figure 33. Figure 33: Comparison of agent behavior trained with different reward functions. [PITH_FULL_IMAGE:figures/full_fig_p221_33.png]
Figure 34
Figure 34. Figure 34: Fragmentation of human-likeness in behavior space. Each colored circle represents [PITH_FULL_IMAGE:figures/full_fig_p229_34.png]
Figure 35
Figure 35. Figure 35: From Foundational Gears to the Next Stage: A Dual-Tier Gear System of Playstyle [PITH_FULL_IMAGE:figures/full_fig_p262_35.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

299 extracted references · 65 canonical work pages

  1. [1]

    Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng

    Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. Tensorflow: A system fo...

  2. [2]

    Pieter Abbeel and Andrew Y . Ng. Apprenticeship learning via inverse reinforcement learning. In International Conference on Machine Learning (ICML), 2004

  3. [3]

    Machado, Pablo Samuel Castro, and Marc G

    Rishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro, and Marc G. Bellemare. Contrastive behavioral similarity embeddings for generalization in reinforcement learn- ing. In International Conference on Learning Representations (ICLR), 2021

  4. [4]

    Christiano, John Schulman, and Dan Mané

    Dario Amodei, Chris Olah, Jacob Steinhardt, Paul F. Christiano, John Schulman, and Dan Mané. Concrete problems in ai safety. CoRR, abs/1606.06565, 2016

  5. [5]

    Summa Theologica , volume 5

    Thomas Aquinas. Summa Theologica , volume 5. Guerin, 1274. Originally written 1265–1274. Printed edition used: Guerin, 1869

  6. [6]

    Metaphysics

    Aristotle. Metaphysics. Clarendon Press, 350 BCE. Originally written circa 350 BCE; this edition translated by Christopher Kirwan and published in 1993

  7. [7]

    Wasserstein gan

    Martín Arjovsky, Soumith Chintala, and Léon Bottou. Wasserstein gan. CoRR, 2017

  8. [8]

    Finite-time analysis of the multi- armed bandit problem

    Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer. Finite-time analysis of the multi- armed bandit problem. Machine Learning, 47:235–256, 2002. 258

Show all 299 references
  1. [9]

    Bernard J. Baars. A Cognitive Theory of Consciousness . Cambridge University Press, 1988

  2. [10]

    Agent57: Outperforming the atari human benchmark

    Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, and Charles Blundell. Agent57: Outperforming the atari human benchmark. In International Conference on Machine Learning (ICML) , 2020

  3. [11]

    Never give up: Learning directed exploration strate- gies

    Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andrew Bolt, and Charles Blundell. Never give up: Learning directed exploration strate- gies. In International...

  4. [12]

    Markov, Yi Wu, Glenn Powell, Bob Mc- Grew, and Igor Mordatch

    Bowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu, Glenn Powell, Bob Mc- Grew, and Igor Mordatch. Emergent tool use from multi-agent autocurricula. In Inter- national Conference on Learning Representations (ICLR), 2020

  5. [13]

    Video pretraining (vpt): Learning to act by watching unlabeled online videos

    Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, and Jeff Clune. Video pretraining (vpt): Learning to act by watching unlabeled online videos. In Advances in Neural Information Processing Systems (NeurIPS), 2022

  6. [14]

    Re-evaluating evaluation

    David Balduzzi, Karl Tuyls, Julien Pérolat, and Thore Graepel. Re-evaluating evaluation. In Advances in Neural Information Processing Systems (NeurIPS), 2018

  7. [15]

    Open-ended learning in symmetric zero-sum games

    David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Pérolat, Max Jaderberg, and Thore Graepel. Open-ended learning in symmetric zero-sum games. In International Conference on Machine Learning (ICML), 2019

  8. [16]

    Hunt, Tom Schaul, Hado P

    André Barreto, Will Dabney, Rémi Munos, Jonathan J. Hunt, Tom Schaul, Hado P. Van Hasselt, and David Silver. Successor features for transfer in reinforcement learn- ing. In Advances in Neural Information Processing Systems (NeurIPS), 2017

  9. [17]

    Hearts, clubs, diamonds, spades: Players who suit muds

    Richard Bartle. Hearts, clubs, diamonds, spades: Players who suit muds. Journal of MUD Research, 1996

  10. [18]

    What is game balancing? - an examination of concepts

    Alexander Becker and Daniel Görlich. What is game balancing? - an examination of concepts. ParadigmPlus, 2020. 259

  11. [19]

    M. G. Bellemare, W. Dabney, and R. Munos. A distributional perspective on reinforce- ment learning. In International Conference on Machine Learning (ICML), 2017

  12. [20]

    Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling

    Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. The arcade learn- ing environment: An evaluation platform for general agents. Journal of Artificial Intelli- gence Research, 2013

  13. [21]

    Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos

    Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos. Unifying count-based exploration and intrinsic motivation. In Advances in Neural Information Processing Systems (NeurIPS), 2016

  14. [22]

    Representation learning: A review and new perspectives

    Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8):1798–1828, 2013

  15. [23]

    Culture, Class, Distinction

    Tony Bennett, Mike Savage, Elizabeth Bortolaia Silva, Alan Warde, Modesto Gayo-Cal, and David Wright. Culture, Class, Distinction. Routledge, London, 2009

  16. [24]

    Theano: A cpu and gpu math expression compiler

    James Bergstra, Olivier Breuleux, Frédéric Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and Yoshua Bengio. Theano: A cpu and gpu math expression compiler. In Proceedings of the Python for Scientific Computing Conference (Sci...

  17. [25]

    A Treatise Concerning the Principles of Human Knowledge

    George Berkeley. A Treatise Concerning the Principles of Human Knowledge. Clarendon Press, 1710. Reprinted in many editions; standard version from the 1710 first edition

  18. [26]

    Daniel E. Berlyne. Curiosity and exploration: Animals spend much of their time seeking stimuli whose significance raises problems for psychology. Science, 153(3731):25–33, 1966

  19. [27]

    System neural diversity: Measuring behavioral heterogeneity in multi-agent learning

    Matteo Bettini, Ajay Shankar, and Amanda Prorok. System neural diversity: Measuring behavioral heterogeneity in multi-agent learning. CoRR, abs/2305.02128, 2023

  20. [28]

    On a measure of divergence between two multinomial populations

    Anil Bhattacharyya. On a measure of divergence between two multinomial populations. Sankhy¯a: The Indian Journal of Statistics, 1946

  21. [29]

    The Holy Bible

    Bible. The Holy Bible. Church of England, 600 BCE. Compilation of texts from circa 600 BCE to 100 CE; this edition is the King James Version (KJV), first published in 1611. 260

  22. [30]

    Christopher M. Bishop. Pattern Recognition and Machine Learning. Springer, 2006

  23. [31]

    Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba

    Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba. End to end learning for self-driving cars. CoRR, abs/1604.07316, 2016

  24. [32]

    Are you living in a computer simulation? The Philosophical Quarterly, 53(211):243–255, 2003

    Nick Bostrom. Are you living in a computer simulation? The Philosophical Quarterly, 53(211):243–255, 2003

  25. [33]

    Distinction: A Social Critique of the Judgement of Taste

    Pierre Bourdieu. Distinction: A Social Critique of the Judgement of Taste . Harvard University Press, Cambridge, MA, 1984. Originally published in French in 1979 as La distinction

  26. [34]

    Heads-up limit hold’em poker is solved

    Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin. Heads-up limit hold’em poker is solved. Science, 347(6218):145–149, 2015

  27. [35]

    Ralph Allan Bradley and Milton E. Terry. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 1952

  28. [36]

    Disagreement-regularized imitation learn- ing

    Kianté Brantley, Wen Sun, and Mikael Henaff. Disagreement-regularized imitation learn- ing. In International Conference on Learning Representations (ICLR), 2020

  29. [37]

    Random forests

    Leo Breiman. Random forests. Machine Learning, 45(1):5–32, 2001

  30. [38]

    Openai gym

    Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. Openai gym. CoRR, abs/1606.01540, 2016

  31. [39]

    Driving event detection and driving style classification using artificial neural networks

    Patrick Brombacher, Johannes Masino, Michael Frey, and Frank Gauterin. Driving event detection and driving style classification using artificial neural networks. In IEEE Inter- national Conference on Industrial Technology (ICIT), 2017

  32. [40]

    Bentz, Léon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Eduard Säckinger, and Roopak Shah

    Jane Bromley, James W. Bentz, Léon Bottou, Isabelle Guyon, Yann LeCun, Cliff Moore, Eduard Säckinger, and Roopak Shah. Signature verification using a "siamese" time delay neural network. International Journal of Pattern Recognition and Artificial Intelligence (IJPRAI), 7, 1993

  33. [41]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini 261 Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, ...

  34. [42]

    Jerome S. Bruner. Acts of Meaning. Harvard University Press, 1990

  35. [43]

    Perception and the Representative Design of Psychological Experi- ments

    Egon Brunswik. Perception and the Representative Design of Psychological Experi- ments. University of California Press, 1956

  36. [44]

    Human Motivation and Emotion

    Ross Buck. Human Motivation and Emotion. John Wiley & Sons, 1988

  37. [45]

    Storkey, and Oleg Klimov

    Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov. Exploration by random network distillation. In International Conference on Learning Representations (ICLR), 2019

  38. [46]

    Peter Burkholder, Donald Jay Grout, and Claude V

    J. Peter Burkholder, Donald Jay Grout, and Claude V . Palisca. A History of Western Music: Tenth International Student Edition. W. W. Norton & Company, New York, 10th, international student edition edition, 2019

  39. [47]

    Joseph Jr

    Murray Campbell, A. Joseph Jr. Hoane, and Feng-Hsiung Hsu. Deep blue. Artificial Intelligence, 134(1–2):57–83, 2002

  40. [48]

    William E. Caplin. Classical Form: A Theory of Formal Functions for the Instrumental Music of Haydn, Mozart, and Beethoven. Oxford University Press, New York, 1998

  41. [49]

    Bellemare

    Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G. Bellemare. Dopamine: A research framework for deep reinforcement learning. CoRR, abs/1812.06110, 2018

  42. [50]

    scary robots

    Stephen Cave, Kate Coughlan, and Kanta Dihal. "scary robots": Examining public re- sponses to AI. In AAAI/ACM Conference on AI, Ethics, and Society (AIES), 2019

  43. [51]

    Vision-language models as a source of rewards

    Harris Chan, V olodymyr Mnih, Feryal Behbahani, Michael Laskin, Luyu Wang, Fabio Pardo, Maxime Gazeau, Himanshu Sahni, Dan Horgan, Kate Baumli, Yannick 262 Schroecker, Stephen Spencer, Richie Steigerwald, John Quan, Gheorghe Comanici, Se- bastian Flennerhag, Alexander Neitz, L...

  44. [52]

    Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, and Wen Sun

    Jonathan D. Chang, Masatoshi Uehara, Dhruv Sreenivas, Rahul Kidambi, and Wen Sun. Mitigating covariate shift in imitation learning via offline data with partial coverage. In Advances in Neural Information Processing Systems (NeurIPS), 2021

  45. [53]

    Similarity estimation techniques from rounding algorithms

    Moses Charikar. Similarity estimation techniques from rounding algorithms. In ACM Symposium on Theory of Computing (STOC), 2002

  46. [54]

    Detecting individual decision-making style in the game of go

    Chun-Jung Chen. Detecting individual decision-making style in the game of go. Master’s thesis, National Yang Ming Chiao Tung University, 2023

  47. [55]

    Learning to evaluate the artness of ai-generated images

    Junyu Chen, Jie An, Hanjia Lyu, Christopher Kanan, and Jiebo Luo. Learning to evaluate the artness of ai-generated images. IEEE Transactions on Multimedia, 26, 2024

  48. [56]

    Diffusion model-augmented behavioral cloning

    Shang-Fu Chen, Hsiang-Chun Wang, Ming-Hao Hsu, Chun-Mao Lai, and Shao-Hua Sun. Diffusion model-augmented behavioral cloning. InInternational Conference on Machine Learning (ICML), 2024

  49. [57]

    Modeling intransitivity in matchup and comparison data

    Shuo Chen and Thorsten Joachims. Modeling intransitivity in matchup and comparison data. In International Conference on Web Search and Data Mining (WSDM), 2016

  50. [58]

    Infogan: Interpretable representation learning by information maximizing generative ad- versarial nets

    Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel. Infogan: Interpretable representation learning by information maximizing generative ad- versarial nets. In Advances in Neural Information Processing Systems (NeurIPS), 2016

  51. [59]

    Open Innovation: The New Imperative for Creating and Profiting from Technology

    Henry William Chesbrough. Open Innovation: The New Imperative for Creating and Profiting from Technology. Harvard Business Press, 2003

  52. [60]

    Learning phrase representations us- ing rnn encoder–decoder for statistical machine translation

    Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. Learning phrase representations us- ing rnn encoder–decoder for statistical machine translation. In Conference on Empirical Methods in Natural Language Pr...

  53. [61]

    Christensen

    Clayton M. Christensen. The Innovator’s Dilemma: When New Technologies Cause Great Firms to Fail. Harvard Business Review Press, 1997

  54. [62]

    Maximizing cosine similarity between spatial features for unsupervised domain adaptation in semantic segmentation

    Inseop Chung, Daesik Kim, and Nojun Kwak. Maximizing cosine similarity between spatial features for unsupervised domain adaptation in semantic segmentation. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2022

  55. [63]

    Cuisine and Culture: A History of Food and People

    Linda Civitello. Cuisine and Culture: A History of Food and People. John Wiley & Sons, Hoboken, NJ, 2011

  56. [64]

    Surfing Uncertainty: Prediction, Action, and the Embodied Mind

    Andy Clark. Surfing Uncertainty: Prediction, Action, and the Embodied Mind . Oxford University Press, 2015

  57. [65]

    de Condorcet

    M. de Condorcet. Essai sur l’application de l’analyse à la probabilité des décisions rendues à la pluralité des voix. Imprimerie Royale, Paris, 1785

  58. [66]

    The Analects

    Confucius. The Analects. Ballantine Books, 475 BCE. Originally compiled in the 5th century BCE; this edition published in 2003

  59. [67]

    Support-vector networks

    Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine Learning, 20 (3):273–297, 1995

  60. [68]

    Costa, Paul T

    Jr. Costa, Paul T. and Robert R. McCrae. Domains and facets: Hierarchical personality assessment using the revised neo personality inventory. Journal of Personality Assess- ment, 1995

  61. [69]

    Efficient selectivity and backup operators in monte-carlo tree search

    Rémi Coulom. Efficient selectivity and backup operators in monte-carlo tree search. In International Conference on Computer and Games (CG), pp. 72–83, 2006

  62. [70]

    elo ratings

    Rémi Coulom. Computing “elo ratings” of move patterns in the game of go. Journal of the International Computer Games Association (ICGA Journal), 30(4):198–206, 2007

  63. [71]

    Whole-history rating: A bayesian rating system for players of time- varying strength

    Rémi Coulom. Whole-history rating: A bayesian rating system for players of time- varying strength. In International Conference on Computers and Games (CG), 2008

  64. [72]

    Flow: The Psychology of Optimal Experience

    Mihály Csíkszentmihályi. Flow: The Psychology of Optimal Experience. Harper & Row, 1990. 264

  65. [73]

    Robots that can adapt like animals

    Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret. Robots that can adapt like animals. Nature, 521(7553):503–507, 2015

  66. [74]

    Implicit quantile net- works for distributional reinforcement learning

    Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos. Implicit quantile net- works for distributional reinforcement learning. InInternational Conference on Machine Learning (ICML), 2018

  67. [75]

    Primal wasser- stein imitation learning

    Robert Dadashi, Léonard Hussenot, Matthieu Geist, and Olivier Pietquin. Primal wasser- stein imitation learning. In International Conference on Learning Representations (ICLR), 2021

  68. [76]

    Rae, Andreas Glaese, Yujia Rist, Tom Goodside, Catherine Olsson, Jarrid Rae, Thomas Van Overveldt, Daniel Levy, Xinyun Li, Anton Bakhtin, Bryan McCann, Panos Mandilaras, Samuel R

    Sumanth Dathathri, Abigail See, Sumedh Ghaisas, Po-Sen Huang, Rob McAdam, Jo- hannes Welbl, Vandana Bachani, Alex Kaskasoli, Robert Stanforth, Tatiana Matejovi- cova, James Sutherland, Chaitanya Malaviya, Nicholas Nunez, Luke Belrose, Samuel Humeau, Abhinav Sridhar, Tom Jones,...

  69. [77]

    Actions, reasons, and causes

    Donald Davidson. Actions, reasons, and causes. The Journal of Philosophy , 60(23): 685–700, 1963

  70. [78]

    Gureckis, and Brenden M

    Guy Davidson, Graham Todd, Julian Togelius, Todd M. Gureckis, and Brenden M. Lake. Goals as reward-producing programs. Nature Machine Intelligence, 2025

  71. [79]

    Dempster, Nan M

    Arthur P. Dempster, Nan M. Laird, and Donald B. Rubin. Maximum likelihood from incomplete data via the em algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 39(1):1–22, 1977

  72. [80]

    Daniel C. Dennett. The Intentional Stance. MIT Press, 1987

  73. [81]

    Meditations on First Philosophy

    René Descartes. Meditations on First Philosophy . Hackett Publishing, Indianapolis, 3rd edition, 1641. Originally published in Latin in 1641. This edition: 3rd ed., 1993, translated by Donald A. Cress. 265

  74. [82]

    BERT: Pre- training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre- training of deep bidirectional transformers for language understanding. North American Chapter of the Association for Computational Linguistics (NAACL), 2019

  75. [83]

    Jacob D. Dodson. The relation of strength of stimulus to rapidity of habit-formation in the kitten. Journal of Animal Behavior, 1915

  76. [84]

    Flownet: Learning optical flow with convolutional networks

    Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Häusser, Caner Hazirbas, Vladimir Golkov, Patrick van der Smagt, Daniel Cremers, and Thomas Brox. Flownet: Learning optical flow with convolutional networks. In IEEE International Conference on Com- puter Vision (ICCV), 2015

  77. [85]

    Yannakakis

    Anders Drachen, Alessandro Canossa, and Georgios N. Yannakakis. Player modeling us- ing self-organization in tomb raider: Underworld. InIEEE Conference on Computational Intelligence and Games (CIG), 2009

  78. [86]

    Temporal cycle-consistency learning

    Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet, and Andrew Zis- serman. Temporal cycle-consistency learning. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  79. [87]

    Elo-mmr: A rating system for massive multiplayer compe- titions

    Aram Ebtekar and Paul Liu. Elo-mmr: A rating system for massive multiplayer compe- titions. In The Web Conference (WWW), 2021

  80. [88]

    Reason for Being: A Meditation on Ecclesiastes

    Jacques Ellul. Reason for Being: A Meditation on Ecclesiastes . Wm. B. Eerdmans Publishing, Grand Rapids, MI, 1990

  81. [89]

    Arpad E. Elo. The USCF Rating System: Its Development, Theory, and Applications . United States Chess Federation, 1966

  82. [90]

    Impala: Scalable distributed deep-rl with importance weighted actor- learner architectures

    Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, V olodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu. Impala: Scalable distributed deep-rl with importance weighted actor- learner architectures. In Internatio...

  83. [91]

    A density-based al- gorithm for discovering clusters in large spatial databases with noise

    Martin Ester, Hans-Peter Kriegel, Jörg Sander, and Xiaowei Xu. A density-based al- gorithm for discovering clusters in large spatial databases with noise. In International Conference on Knowledge Discovery and Data Mining (KDD), 1996. 266

  84. [92]

    Diversity is all you need: Learning skills without a reward function

    Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine. Diversity is all you need: Learning skills without a reward function. In International Conference on Learning Representations (ICLR), 2019

  85. [93]

    Generalized data distribution iteration

    Jiajun Fan and Changnan Xiao. Generalized data distribution iteration. In International Conference on Machine Learning (ICML), 2022

  86. [94]

    Learnable behavior control: Breaking atari human world records via sample-efficient behavior selection

    Jiajun Fan, Yuzheng Zhuang, Yuecheng Liu, Jianye Hao, Bin Wang, Jiangcheng Zhu, Hao Wang, and Shu-Tao Xia. Learnable behavior control: Breaking atari human world records via sample-efficient behavior selection. InInternational Conference on Learning Representations (ICLR), 2023

  87. [95]

    G. T. Fechner. Elements of Psychophysics. Holt, Rinehart and Winston, 1966. Originally published in 1860

  88. [96]

    Testing the manifold hy- pothesis

    Charles Fefferman, Sanjoy Mitter, and Hariharan Narayanan. Testing the manifold hy- pothesis. Journal of the American Mathematical Society, 29(4):983–1049, 2016

  89. [97]

    Guided cost learning: Deep inverse op- timal control via policy optimization

    Chelsea Finn, Sergey Levine, and Pieter Abbeel. Guided cost learning: Deep inverse op- timal control via policy optimization. In International Conference on Machine Learning (ICML), 2016

  90. [98]

    Ramirez, Ayzaan Wahid, Laura Downs, Adrian Wong, Johnny Lee, Igor Mordatch, and Jonathan Tompson

    Pete Florence, Corey Lynch, Andy Zeng, Oscar A. Ramirez, Ayzaan Wahid, Laura Downs, Adrian Wong, Johnny Lee, Igor Mordatch, and Jonathan Tompson. Implicit behavioral cloning. In Conference on Robot Learning (CoRL), 2021

  91. [99]

    Noisy networks for exploration

    Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hes- sel, Ian Osband, Alex Graves, V olodymyr Mnih, Rémi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg. Noisy networks for exploration. In Inter- national Conference on Lear...

  92. [100]

    Fox go, 2024

    Fox Go. Fox go, 2024. URL https://www.foxwq.com/

  93. [101]

    Friedman

    Jerome H. Friedman. Greedy function approximation: A gradient boosting machine. Annals of Statistics, 29(5):1189–1232, 2001

  94. [102]

    Learning robust rewards with adversarial in- verse reinforcement learning

    Justin Fu, Katie Luo, and Sergey Levine. Learning robust rewards with adversarial in- verse reinforcement learning. In International Conference on Learning Representations (ICLR), 2018. 267

  95. [103]

    Evalu- ating human-like behaviors of video-game agents autonomously acquired with biological constraints

    Nobuto Fujii, Yuichi Sato, Hironori Wakama, Koji Kazai, and Haruhiro Katayose. Evalu- ating human-like behaviors of video-game agents autonomously acquired with biological constraints. In Advances in Computer Entertainment Technology (ACE), 2013

  96. [104]

    Extreme q-learning: Maxent rl without entropy

    Divyansh Garg, Joey Hejna, Matthieu Geist, and Stefano Ermon. Extreme q-learning: Maxent rl without entropy. In International Conference on Learning Representations (ICLR), 2023

  97. [105]

    Gatys, Alexander S

    Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge. Image style transfer using convolutional neural networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016

  98. [106]

    Combining online and offline knowledge in uct

    Sylvain Gelly and David Silver. Combining online and offline knowledge in uct. In International Conference on Machine Learning (ICML), pp. 273–280, 2007

  99. [107]

    Gut Feelings: The Intelligence of the Unconscious

    Gerd Gigerenzer. Gut Feelings: The Intelligence of the Unconscious. Penguin, 2007

  100. [108]

    Rich feature hierar- chies for accurate object detection and semantic segmentation

    Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierar- chies for accurate object detection and semantic segmentation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014

  101. [109]

    Glickman

    Mark E. Glickman. Parameter estimation in large dynamic paired comparison experi- ments. Journal of the Royal Statistical Society Series C: Applied Statistics, 1999

  102. [110]

    Clément Godard, Oisin Mac Aodha, and Gabriel J. Brostow. Unsupervised monocular depth estimation with left-right consistency. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017

  103. [111]

    Über formal unentscheidbare sätze der principia mathematica und ver- wandter systeme i

    Kurt Gödel. Über formal unentscheidbare sätze der principia mathematica und ver- wandter systeme i. Monatshefte für Mathematik und Physik, 38:173–198, 1931

  104. [112]

    Gonzalez and Richard E

    Rafael C. Gonzalez and Richard E. Woods. Digital Image Processing . Pearson, 4th edition, 2017

  105. [113]

    Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C

    Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems (NeurIPS), 2014

  106. [114]

    Ways of Worldmaking

    Nelson Goodman. Ways of Worldmaking. Hackett Publishing, Indianapolis, 1978. 268

  107. [115]

    Yannakakis

    Daniele Gravina, Antonios Liapis, and Georgios N. Yannakakis. Quality diversity through surprise. IEEE Transactions on Evolutionary Computation , 23(4):603–616, 2019

  108. [116]

    Shuyue Guan and Murray H. Loew. A novel measure to evaluate generative adversarial networks based on direct analysis of generated images. Neural Computing and Applica- tions, 2021

  109. [117]

    Reinforcement learn- ing with deep energy-based policies

    Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine. Reinforcement learn- ing with deep energy-based policies. In International Conference on Machine Learning (ICML), 2017

  110. [118]

    Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

    Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning (ICML), 2018

  111. [119]

    The Code of Hammurabi

    Hammurabi. The Code of Hammurabi. University of Chicago Press, 1754 BCE. Orig- inally inscribed circa 1754 BCE; this edition translated from the original Akkadian and published in 1904

  112. [120]

    The Elements of Statistical Learning

    Trevor Hastie, Robert Tibshirani, and Jerome Friedman. The Elements of Statistical Learning. Springer, 2009

  113. [121]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016

  114. [122]

    Mask r-cnn

    Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask r-cnn. IEEE International Conference on Computer Vision (ICCV), 2017

  115. [123]

    Artistic styles: Revisiting the analysis of modern artists’ careers

    Christiane Hellmanzik. Artistic styles: Revisiting the analysis of modern artists’ careers. Journal of Cultural Economics, 33:201–232, 2009

  116. [124]

    Trueskill™: A bayesian skill rating system

    Ralf Herbrich, Tom Minka, and Thore Graepel. Trueskill™: A bayesian skill rating system. In Advances in Neural Information Processing Systems (NIPS), 2006

  117. [125]

    Clipscore: A reference-free evaluation metric for image captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning. In Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021. 269

  118. [126]

    Hessel, J

    M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Gheshlaghi Azar, and D. Silver. Rainbow: Combining improvements in deep reinforcement learning. In AAAI Conference on Artificial Intelligence (AAAI), 2018

  119. [127]

    Hester, M

    T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, G. Dulac-Arnold, J. P. Agapiou, J. Z. Leibo, and A. Gruslys. Deep q-learning from demonstrations. In AAAI Conference on Artificial Intelligence (AAAI), 2018

  120. [128]

    Agapiou, Joel Z

    Todd Hester, Matej Vecerík, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, Gabriel Dulac-Arnold, John P. Agapiou, Joel Z. Leibo, and Audrunas Gruslys. Deep q-learning from demonstrations. In AAAI Conference on Arti...

  121. [129]

    Gans trained by a two time-scale update rule converge to a local nash equi- librium

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equi- librium. In Advances in Neural Information Processing Systems (NeurIPS), 2017

  122. [130]

    Generative adversarial imitation learning

    Jonathan Ho and Stefano Ermon. Generative adversarial imitation learning. In Advances in Neural Information Processing Systems (NeurIPS), 2016

  123. [131]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (NeurIPS), 2020

  124. [132]

    Towards human-like rl: Taming non-naturalistic behavior in deep rl via adap- tive behavioral costs in 3d games

    Kuo-Hao Ho, Ping-Chun Hsieh, Chiu-Chou Lin, You-Ren Luo, Feng-Jian Wang, and I- Chen Wu. Towards human-like rl: Taming non-naturalistic behavior in deep rl via adap- tive behavioral costs in 3d games. In Asian Conference on Machine Learning (ACML), 2023

  125. [133]

    Leviathan

    Thomas Hobbes. Leviathan. Penguin Classics, London, 1651. This edition: Penguin Classics, 1985. Originally published in 1651

  126. [134]

    Long short-term memory

    Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computa- tion, 9:1735–1780, 1997

  127. [135]

    Distributed prioritized experience replay

    Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado van Hasselt, and David Silver. Distributed prioritized experience replay. In International Conference on Learning Representations (ICLR), 2018. 270

  128. [136]

    Multilayer feedforward net- works are universal approximators

    Kurt Hornik, Maxwell Stinchcombe, and Halbert White. Multilayer feedforward net- works are universal approximators. Neural Networks, 2(5):359–366, 1989

  129. [137]

    Image quality metrics: Psnr vs

    Alain Horé and Djemel Ziou. Image quality metrics: Psnr vs. ssim. In International Conference on Pattern Recognition (ICPR), 2010

  130. [138]

    Measuring policy distance for multi-agent reinforcement learning

    Tianyi Hu, Zhiqiang Pu, Xiaolin Ai, Tenghai Qiu, and Jianqiang Yi. Measuring policy distance for multi-agent reinforcement learning. In International Conference on Au- tonomous Agents and Multiagent Systems (AAMAS), 2024

  131. [139]

    Efficient action-constrained reinforce- ment learning via acceptance-rejection method and augmented mdps

    Wei Hung, Shao-Hua Sun, and Ping-Chun Hsieh. Efficient action-constrained reinforce- ment learning via acceptance-rejection method and augmented mdps. In International Conference on Learning Representations (ICLR), 2025

  132. [140]

    The case for dynamic difficulty adjustment in games

    Robin Hunicke. The case for dynamic difficulty adjustment in games. In ACM SIGCHI International Conference on Advances in Computer Entertainment Technology (ACE) , 2005

  133. [141]

    Ideas Pertaining to a Pure Phenomenology and to a Phenomenologi- cal Philosophy: First Book: General Introduction to a Pure Phenomenology

    Edmund Husserl. Ideas Pertaining to a Pure Phenomenology and to a Phenomenologi- cal Philosophy: First Book: General Introduction to a Pure Phenomenology . Springer Science & Business Media, Dordrecht, 1913. Originally published in German in 1913. This edition: English transla...

  134. [142]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift

    Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning (ICML), 2015

  135. [143]

    The Art of Color: The Subjective Experience and Objective Rationale of Color

    Johannes Itten. The Art of Color: The Subjective Experience and Objective Rationale of Color. Van Nostrand Reinhold Company, New York, 1973. Translated from German by Ernst van Haagen. Originally published as Kunst der Farbe

  136. [144]

    Étude comparative de la distribution florale dans une portion des alpes et des jura

    Paul Jaccard. Étude comparative de la distribution florale dans une portion des alpes et des jura. Bulletin de la Société Vaudoise des Sciences Naturelles, 37:547–579, 1901

  137. [145]

    Czarnecki, Jeff Don- ahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, Chrisantha 271 Fernando, and Koray Kavukcuoglu

    Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M. Czarnecki, Jeff Don- ahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, Chrisantha 271 Fernando, and Koray Kavukcuoglu. Population based training of neural networks.CoRR, abs/1711.09846, 2017

  138. [146]

    The Principles of Psychology, volume 1–2

    William James. The Principles of Psychology, volume 1–2. Henry Holt and Company, New York, 1890. Originally published in two volumes

  139. [147]

    Lawrence Zitnick, and Ross Girshick

    Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Judy Hoffman, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. Inferring and executing programs for visual reasoning. In IEEE International Conference on Computer Vision (ICCV), 2017

  140. [148]

    Jones and Benjamin K

    Cameron R. Jones and Benjamin K. Bergen. Large language models pass the turing test. CoRR, abs/2503.23674, 2025

  141. [149]

    Unity: A general platform for intelligent agents

    Arthur Juliani, Vincent-Pierre Berges, Esh Vckay, Yuan Gao, Hunter Henry, Marwan Mattar, and Danny Lange. Unity: A general platform for intelligent agents. CoRR, abs/1809.02627, 2018

  142. [150]

    John M. Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A A Kohl, Andy Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav...

  143. [151]

    Daniel Jurafsky and James H. Martin. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition . Prentice Hall, 2 edition, 2008

  144. [152]

    Littman, and Anthony R

    Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra. Planning and acting in partially observable stochastic domains. Artificial Intelligence, 101(1–2):99– 134, 1998

  145. [153]

    Kahneman and A

    D. Kahneman and A. Tversky. Prospect theory: An analysis of decision under risk. Econometrica, 47(2):263–291, 1979

  146. [154]

    Thinking, Fast and Slow

    Daniel Kahneman. Thinking, Fast and Slow. Farrar, Straus and Giroux, New York, 2011. 272

  147. [155]

    Benchmarking end-to-end behavioural cloning on video games

    Anssi Kanervisto, Joonas Pussinen, and Ville Hautamäki. Benchmarking end-to-end behavioural cloning on video games. In IEEE Conference on Games (CoG), 2020

  148. [156]

    Critique of Pure Reason

    Immanuel Kant. Critique of Pure Reason . Cambridge University Press, Cambridge,

  149. [157]

    Re- current experience replay in distributed reinforcement learning

    Steven Kapturowski, Georg Ostrovski, John Quan, Rémi Munos, and Will Dabney. Re- current experience replay in distributed reinforcement learning. In International Confer- ence on Learning Representations (ICLR), 2019

  150. [158]

    A style-based generator architecture for generative adversarial networks

    Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  151. [159]

    The road to artificial superintelligence: A comprehen- sive survey of superalignment

    HyunJin Kim, Xiaoyuan Yi, Jing Yao, Jianxun Lian, Muhua Huang, Shitong Duan, JinYeong Bak, and Xing Xie. The road to artificial superintelligence: A comprehen- sive survey of superalignment. CoRR, abs/2412.16468, 2024

  152. [160]

    Kingma and Max Welling

    Diederik P. Kingma and Max Welling. Auto-encoding variational bayes. InInternational Conference on Learning Representations (ICLR), 2014

  153. [161]

    The soviet scientific programme on ai: if a machine cannot ‘think’, can it ‘control’? BJHS Themes, 8:111–125, 2023

    Olessia Kirtchik. The soviet scientific programme on ai: if a machine cannot ‘think’, can it ‘control’? BJHS Themes, 8:111–125, 2023

  154. [162]

    Gaming industry report 2025: Market size & trends,

    Andrea Knezovic. Gaming industry report 2025: Market size & trends,

  155. [163]

    Bandit based monte-carlo planning

    Levente Kocsis and Csaba Szepesvári. Bandit based monte-carlo planning. In European Conference on Machine Learning (ECML), pp. 282–293, 2006

  156. [164]

    Deep neural decision forests

    Peter Kontschieder, Madalina Fiterau, Antonio Criminisi, and Samuel Rota Bulò. Deep neural decision forests. In IEEE International Conference on Computer Vision (ICCV), 2015

  157. [165]

    Offline reinforcement learning with implicit q-learning

    Ilya Kostrikov, Ashvin Nair, and Sergey Levine. Offline reinforcement learning with implicit q-learning. In International Conference on Learning Representations (ICLR) , 2022. 273

  158. [166]

    Learning multiple layers of features from tiny images

    Alex Krizhevsky and Geoffrey Hinton. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009

  159. [167]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Sys- tems (NeurIPS), 2012

  160. [168]

    Conservative q-learning for offline reinforcement learning

    Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. In Advances in Neural Information Processing Sys- tems (NeurIPS), 2020

  161. [169]

    Diffusion-reward adversarial imitation learning

    Chun-Mao Lai, Hsiang-Chun Wang, Ping-Chun Hsieh, Frank Wang, Min-Hung Chen, and Shao-Hua Sun. Diffusion-reward adversarial imitation learning. In Advances in Neural Information Processing Systems (NIPS), 2024

  162. [170]

    Can agents run relay race with strangers? generalization of rl to out-of-distribution trajectories

    Li-Cheng Lan, Huan Zhang, and Cho-Jui Hsieh. Can agents run relay race with strangers? generalization of rl to out-of-distribution trajectories. In International Con- ference on Learning Representations (ICLR), 2023

  163. [171]

    A unified game-theoretic ap- proach to multiagent reinforcement learning

    Marc Lanctot, Vinícius Flores Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolet, David Silver, and Thore Graepel. A unified game-theoretic ap- proach to multiagent reinforcement learning. In Advances in Neural Information Pro- cessing Systems (NeurIPS), 2017

  164. [172]

    Deep learning

    Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. Nature, 521(7553): 436–444, 2015

  165. [173]

    Gradient- based regularization for action smoothness in robotic control with reinforcement learn- ing

    I Lee, Hoang-Giang Cao, Cong-Tinh Dao, Yu-Cheng Chen, and I-Chen Wu. Gradient- based regularization for action smoothness in robotic control with reinforcement learn- ing. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2024

  166. [174]

    G. W. Leibniz. The monadology. In L. E. Loemker (ed.), Philosophical Papers and Letters, pp. 643–653. Springer, Dordrecht, 1714. This edition from: *Philosophical Papers and Letters*, 1989. Originally written in 1714. 274

  167. [175]

    Nonlinear inverse reinforcement learning with gaussian processes

    Sergey Levine, Zoran Popovic, and Vladlen Koltun. Nonlinear inverse reinforcement learning with gaussian processes. InAdvances in Neural Information Processing Systems (NIPS), 2011

  168. [176]

    Neural image beauty predictor based on bradley- terry model

    Shiyu Li, Hao Ma, and Xiangyu Hu. Neural image beauty predictor based on bradley- terry model. CoRR, abs/2111.10127, 2021

  169. [177]

    Infogail: Interpretable imitation learning from visual demonstrations

    Yunzhu Li, Jiaming Song, and Stefano Ermon. Infogail: Interpretable imitation learning from visual demonstrations. In Advances in Neural Information Processing Systems (NeurIPS), 2017

  170. [178]

    Online learning of counter categories and ratings in pvp games

    Chiu-Chou Lin and I-Chen Wu. Online learning of counter categories and ratings in pvp games. Proceedings of the Annual Conference of JSAI , JSAI2025:3K5IS2b05– 3K5IS2b05, 2025

  171. [179]

    An unsupervised video game playstyle metric via state discretization

    Chiu-Chou Lin, Wei-Chen Chiu, and I-Chen Wu. An unsupervised video game playstyle metric via state discretization. In Conference on Uncertainty in Artificial Intelligence (UAI), 2021

  172. [180]

    Method for training ai bot in computer game, February 2022

    Chiu-Chou Lin, Ying-Hau Wu, Kuan-Ming Lin, Pei-Wen Huang, I-Chen Wu, and Cheng- Lun Tsai. Method for training ai bot in computer game, February 2022. United States patent

  173. [181]

    Perceptual similarity for measuring decision-making style and policy diversity in games

    Chiu-Chou Lin, Wei-Chen Chiu, and I-Chen Wu. Perceptual similarity for measuring decision-making style and policy diversity in games. Transactions on Machine Learning Research, 2024

  174. [182]

    Identifying and clustering counter relationships of team com- positions in pvp games for efficient balance analysis

    Chiu-Chou Lin, Yu-Wei Shih, Kuei-Ting Kuo, Yu-Cheng Chen, Chien-Hua Chen, Wei- Chen Chiu, and I-Chen Wu. Identifying and clustering counter relationships of team com- positions in pvp games for efficient balance analysis. Transactions on Machine Learning Research, 2024

  175. [183]

    Method for training ai bot in computer game, October 2024

    Chiu-Chou Lin, Ying-Hau Wu, Kuan-Ming Lin, Pei-Wen Huang, I-Chen Wu, and Cheng- Lun Tsai. Method for training ai bot in computer game, October 2024. Taiwan patent

  176. [184]

    Method for training ai bot in computer game, February 2025

    Chiu-Chou Lin, I-Chen Wu, Jung-Chang Kuo, Ying-Hau Wu, An-Lun Teng, and Pei- Wen Huang. Method for training ai bot in computer game, February 2025. United States patent. 275

  177. [185]

    Breuel, and Jan Kautz

    Ming-Yu Liu, Thomas M. Breuel, and Jan Kautz. Unsupervised image-to-image transla- tion networks. In Advances in Neural Information Processing Systems (NeurIPS), 2017

  178. [186]

    Re-evaluating open-ended evaluation of large language models

    Siqi Liu, Ian Gemp, Luke Marris, Georgios Piliouras, Nicolas Heess, and Marc Lanc- tot. Re-evaluating open-ended evaluation of large language models. In International Conference on Learning Representations (ICLR), 2025

  179. [187]

    Towards unifying behavioral and response diversity for open-ended learning in zero-sum games

    Xiangyu Liu, Hangtian Jia, Ying Wen, Yujing Hu, Yingfeng Chen, Changjie Fan, Zhipeng Hu, and Yaodong Yang. Towards unifying behavioral and response diversity for open-ended learning in zero-sum games. In Advances in Neural Information Pro- cessing Systems (NeurIPS), 2021

  180. [188]

    Imitation from obser- vation: Learning to imitate behaviors from raw video via context translation

    Yuxuan Liu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine. Imitation from obser- vation: Learning to imitate behaviors from raw video via context translation. In IEEE International Conference on Robotics and Automation (ICRA), 2018

  181. [189]

    A unified diversity measure for multiagent reinforcement learning

    Zongkai Liu, Chao Yu, Yaodong Yang, Peng Sun, Zifan Wu, and Yuan Li. A unified diversity measure for multiagent reinforcement learning. In Advances in Neural Infor- mation Processing Systems (NeurIPS), 2022

  182. [190]

    An Essay Concerning Human Understanding

    John Locke. An Essay Concerning Human Understanding. Kay & Troutman, Philadel- phia, 1689. This edition: Kay & Troutman, 1847 (Philadelphia). Public domain edition scanned by Google Books. Originally published in 1689

  183. [191]

    The psychology of curiosity: A review and reinterpretation

    George Loewenstein. The psychology of curiosity: A review and reinterpretation. Psy- chological Bulletin, 116(1):75–98, 1994

  184. [192]

    A study of human-like deep reinforcement learning agents

    You-Ren Luo. A study of human-like deep reinforcement learning agents. Master’s thesis, National Chiao Tung University, 2019

  185. [193]

    Foerster

    Andrei Lupu, Brandon Cui, Hengyuan Hu, and Jakob N. Foerster. Trajectory diversity for zero-shot coordination. In International Conference on Machine Learning (ICML) , 2021

  186. [194]

    Sugarman, and Sarah Hickinbottom

    Jack Martin, Jeff H. Sugarman, and Sarah Hickinbottom. Persons: Understanding Psy- chological Selfhood and Agency. Springer, 2009. 276

  187. [195]

    Abraham H. Maslow. A theory of human motivation. Psychological Review, 50(4): 370–396, 1943

  188. [196]

    McCulloch and Walter Pitts

    Warren S. McCulloch and Walter Pitts. A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics, 5(4):115–133, 1943

  189. [197]

    Kleinberg, and Ashton Ander- son

    Reid McIlroy-Young, Yu Wang, Siddhartha Sen, Jon M. Kleinberg, and Ashton Ander- son. Detecting individual decision-making style: Exploring behavioral stylometry in chess. In Advances in Neural Information Processing Systems (NeurIPS), 2021

  190. [198]

    Charles W. Millard. Fauvism. The Hudson Review, 29(4):576–580, 1976

  191. [199]

    The Rolling Stone Illustrated History of Rock and Roll

    Jim Miller (ed.). The Rolling Stone Illustrated History of Rock and Roll. Random House, New York, 2nd edition, 1981. Originally published in 1976

  192. [200]

    Perceptrons: An Introduction to Computational Geometry

    Marvin Minsky and Seymour Papert. Perceptrons: An Introduction to Computational Geometry. MIT Press, 1969

  193. [201]

    Conditional generative adversarial nets

    Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets. CoRR, 2014

  194. [202]

    Riedmiller

    V olodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller. Playing atari with deep reinforcement learning. CoRR, abs/1312.5602, 2013

  195. [203]

    Rusu, Joel Veness, Marc G

    V olodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Os- trovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstr...

  196. [204]

    Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu

    V olodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning (ICML), 2016

  197. [205]

    Giovanni Molinaro and Andrew G. E. Collins. A goal-centric outlook on learning.Trends in Cognitive Sciences, 27(12):1150–1164, 2023. 277

  198. [206]

    Position: Levels of agi for operationalizing progress on the path to agi

    Meredith Ringel Morris, Jascha Sohl-Dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clément Farabet, and Shane Legg. Position: Levels of agi for operationalizing progress on the path to agi. In International Conference on Machine Learning (ICML), 2024

  199. [207]

    McCaulley, Naomi L

    Isabel Briggs Myers, Mary H. McCaulley, Naomi L. Quenk, and Allen L. Hammer.MBTI Manual: A Guide to the Development and Use of the Myers-Briggs Type Indicator. Con- sulting Psychologists Press, 1998

  200. [208]

    Regularizing action policies for smooth control with reinforcement learning

    Siddharth Mysore, Bassel Mabsout, Renato Mancuso, and Kate Saenko. Regularizing action policies for smooth control with reinforcement learning. In IEEE International Conference on Robotics and Automation (ICRA), 2021

  201. [209]

    The View from Nowhere

    Thomas Nagel. The View from Nowhere. Oxford University Press, 1986

  202. [210]

    Discrete, compositional, and symbolic representations through attractor dynamics

    Andrew Nam, Eric Elmoznino, Nikolay Malkin, James McClelland, Yoshua Bengio, and Guillaume Lajoie. Discrete, compositional, and symbolic representations through attractor dynamics. CoRR, abs/2310.01807, 2023

  203. [211]

    Improving image generation with better captions

    Charlie Nash, William Chan, Yilun Du, Alexander Kirillov, Ross Girshick, Hanzi Liu, Aditya Ramesh, and Barret Zoph. Improving image generation with better captions. CoRR, 2023

  204. [212]

    Ng and Stuart Russell

    Andrew Y . Ng and Stuart Russell. Algorithms for inverse reinforcement learning. In International Conference on Machine Learning (ICML), 2000

  205. [213]

    Ng, Daishi Harada, and Stuart Russell

    Andrew Y . Ng, Daishi Harada, and Stuart Russell. Policy invariance under reward trans- formations: Theory and application to reward shaping. In International Conference on Machine Learning (ICML), 1999

  206. [214]

    Information-directed exploration for deep reinforcement learning

    Nikolay Nikolov, Johannes Kirschner, Felix Berkenkamp, and Andreas Krause. Information-directed exploration for deep reinforcement learning. In International Con- ference on Learning Representations (ICLR), 2019

  207. [215]

    Donald A. Norman. The Design of Everyday Things: Revised and Expanded Edition. Ba- sic Books, New York, revised and expanded edition edition, 2013. Originally published in 1988. 278

  208. [216]

    Game Development Essentials

    Jeannie Novak, Meaghan O’Brien, and Jim Gish. Game Development Essentials. Delmar Cengage Learning, 3 edition, 2012

  209. [217]

    Cuda programming guide

    NVIDIA. Cuda programming guide. https://docs.nvidia.com/cuda/, 2008. Ac- cessed: 2025-06-18

  210. [218]

    Basketball on Paper: Rules and Tools for Performance Analysis

    Dean Oliver. Basketball on Paper: Rules and Tools for Performance Analysis. University of Nebraska Press, Lincoln, NE, 2004

  211. [219]

    Czarnecki, Marc Lanctot, Julien Perolat, and Remi Munos

    Shayegan Omidshafiei, Christos Papadimitriou, Georgios Piliouras, Karl Tuyls, Mark Rowland, Jean-Baptiste Lespiau, Wojciech M. Czarnecki, Marc Lanctot, Julien Perolat, and Remi Munos. Alpha-rank: Multi-agent evaluation by evolution. Scientific Reports, 9(1):9937, 2019

  212. [220]

    OpenAI Five

    OpenAI. OpenAI Five. https://blog.openai.com/openai-five/, 2018

  213. [221]

    Introducing gpt-5

    OpenAI. Introducing gpt-5. https://openai.com/zh-Hant/index/ introducing-gpt-5/, 2025. Accessed: 2025-08-15

  214. [222]

    Deep explo- ration via bootstrapped dqn

    Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. Deep explo- ration via bootstrapped dqn. In Advances in Neural Information Processing Systems (NeurIPS), 2016

  215. [223]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welin- der, Paul F. Christiano, Jan Le...

  216. [224]

    Bleu: A method for automatic evaluation of machine translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: A method for automatic evaluation of machine translation. In Annual Meeting of the Association for Computational Linguistics (ACL), 2002

  217. [225]

    O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S

    Joon Sung Park, Joseph C. O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In ACM Symposium on User Interface Software and Technology (UIST), 2023. 279

  218. [226]

    Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmai- son, Andreas Köpf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner...

  219. [227]

    Efros, and Trevor Darrell

    Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. Curiosity-driven ex- ploration by self-supervised prediction. In International Conference on Machine Learn- ing (ICML), 2017

  220. [228]

    Ivan P. Pavlov. Conditioned Reflexes: An Investigation of the Physiological Activity of the Cerebral Cortex. Oxford University Press, 1927. Translated by G. V . Anrep

  221. [229]

    Arithmetices Principia: Nova Methodo Exposita

    Giuseppe Peano. Arithmetices Principia: Nova Methodo Exposita. Fratres Bocca, Turin,

  222. [230]

    Bilevel entropy based mech- anism design for balancing meta in video games

    Sumedh Pendurkar, Chris Chow, Luo Jie, and Guni Sharon. Bilevel entropy based mech- anism design for balancing meta in video games. In International Conference on Au- tonomous Agents and Multiagent Systems (AAMAS), 2023

  223. [231]

    Human- ity’s last exam

    Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Summer Yue, Alexa...

  224. [232]

    Randy C. Ploetz. Panama disease: An old nemesis rears its ugly head: Part 1. the begin- nings of the banana export trades. Plant Health Progress, 6(1):18, 2005

  225. [234]

    Observe and look further: Achieving consistent performance on atari

    Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado Van Hasselt, John Quan, Mel Vecerík, Mat- teo Hessel, Rémi Munos, and Olivier Pietquin. Observe and look further: Achieving consistent performance on atari. Co...

  226. [235]

    Pomerleau

    Dean A. Pomerleau. Alvinn: An autonomous land vehicle in a neural network. In Advances in Neural Information Processing Systems (NeurIPS), 1988

  227. [236]

    Matchmaking problems in moba games

    Muhammad Farrel Pramono, Kevin Renalda, and Harco Leslie Hendric Spits Warnars. Matchmaking problems in moba games. Indonesian Journal of Electrical Engineering and Computer Science, 2018

  228. [237]

    Pugh, Lawrence B

    Justin K. Pugh, Lawrence B. Soros, and Kenneth O. Stanley. Quality diversity: A new frontier for evolutionary computation. Frontiers in Robotics and AI, 3:40, 2016

  229. [238]

    Connor, Neil Burch, Thomas W

    Julien Pérolat, Bart De Vylder, Daniel Hennes, Eugene Tarassov, Florian Strub, Vin- cent de Boer, Paul Muller, Jerome T. Connor, Neil Burch, Thomas W. Anthony, Stephen McAleer, Romuald Elie, Sarah H. Cen, Zhe Wang, Audrunas Gruslys, Aleksandra Maly- sheva, Mina Khan, Sherjil O...

  230. [239]

    Ross Quinlan

    J. Ross Quinlan. Induction of decision trees. Machine Learning, 1:81–106, 1986

  231. [240]

    Ross Quinlan

    J. Ross Quinlan. C4.5: Programs for Machine Learning. Morgan Kaufmann, 1993

  232. [241]

    Improving lan- guage understanding by generative pre-training, 2018

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving lan- guage understanding by generative pre-training, 2018. Technical report

  233. [242]

    Language models are unsupervised multitask learners

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. Technical report, OpenAI, 2019. Technical report

  234. [243]

    Language models are unsupervised mul- titask learners

    Alec Radford, Jeffrey Wu, Rewon Child, et al. Language models are unsupervised mul- titask learners. OpenAI Blog, 1(8), 2019. 281

  235. [244]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervi- sion. In International C...

  236. [245]

    Bayesian inverse reinforcement learning

    Deepak Ramachandran and Eyal Amir. Bayesian inverse reinforcement learning. In International Joint Conference on Artificial Intelligence (IJCAI), 2007

  237. [246]

    Zero-shot text-to-image generation

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning (ICML), 2021

  238. [247]

    Scott E. Reed, Konrad Zolna, Emilio Parisotto, Sergio Gómez Colmenarejo, Alexan- der Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Orio...

  239. [248]

    Faster r-cnn: Towards real- time object detection with region proposal networks

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real- time object detection with region proposal networks. In Advances in Neural Information Processing Systems (NeurIPS), 2015

  240. [249]

    Intrinsic motivation and flow

    Falko Rheinberg. Intrinsic motivation and flow. Motivation Science, 2020

  241. [250]

    The roles of inducer size and distance in the ebbinghaus illusion (titchener circles)

    Brian Roberts, Mike G Harris, and Tim A Yates. The roles of inducer size and distance in the ebbinghaus illusion (titchener circles). Perception, 34(7):847–856, 2005

  242. [251]

    Vision-language models are zero-shot reward models for reinforcement learning

    Juan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez, and David Lindner. Vision-language models are zero-shot reward models for reinforcement learning. In In- ternational Conference on Learning Representations (ICLR), 2024

  243. [252]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

  244. [253]

    The perceptron: A probabilistic model for information storage and organization in the brain

    Frank Rosenblatt. The perceptron: A probabilistic model for information storage and organization in the brain. Psychological Review, 65(6):386–408, 1958. 282

  245. [254]

    Gordon, and Drew Bagnell

    Stéphane Ross, Geoffrey J. Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2011

  246. [255]

    Rumelhart, Geoffrey E

    David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. Learning representa- tions by back-propagating errors. Nature, 323:533–536, 1986

  247. [256]

    Human-compatible artificial intelligence

    Stuart Russell. Human-compatible artificial intelligence. In Human-Like Machine Intel- ligence, pp. 3–23. Oxford University Press, 2022

  248. [257]

    Artificial Intelligence: A Modern Approach (4th Edi- tion)

    Stuart Russell and Peter Norvig. Artificial Intelligence: A Modern Approach (4th Edi- tion). Pearson, 2020

  249. [258]

    Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text- to-image diffusion models with d...

  250. [259]

    Mehdi S. M. Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly. Assessing generative models via precision and recall. InAdvances in Neural Information Processing Systems (NeurIPS), 2018

  251. [260]

    Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen

    Tim Salimans, Ian J. Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In Advances in Neural Information Processing Systems (NIPS), 2016

  252. [261]

    Learning to fly

    Claude Sammut, Scott Hurst, Dana Kedzier, and Donald Michie. Learning to fly. In International Workshop on Machine Learning (ML), 1992

  253. [262]

    The Art of Game Design: A Book of Lenses

    Jesse Schell. The Art of Game Design: A Book of Lenses. CRC Press, 2008

  254. [263]

    Mastering atari, go, chess and shogi by planning with a learned model

    Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver. Mastering atari, go, chess and shogi by planning with a learned model. Natur...

  255. [264]

    Prox- imal policy optimization algorithms

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Prox- imal policy optimization algorithms. CoRR, abs/1707.06347, 2017

  256. [265]

    Schumpeter

    Joseph A. Schumpeter. The Theory of Economic Development. Harvard University Press, 1934

  257. [266]

    Time-contrastive net- works: Self-supervised learning from multi-view observation

    Pierre Sermanet, Corey Lynch, Jasmine Hsu, and Sergey Levine. Time-contrastive net- works: Self-supervised learning from multi-view observation. In IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPR Workshops), 2017

  258. [267]

    Neurosymbolic artificial intelligence (why, what, and how)

    Amit Sheth, Kaushik Roy, and Manas Gaur. Neurosymbolic artificial intelligence (why, what, and how). IEEE Intelligent Systems, 38(3):56–62, 2023

  259. [268]

    Exploration into translation-equivariant image quantization

    Woncheol Shin, Gyubok Lee, Jiyoung Lee, Eunyi Lyou, Joonseok Lee, and Edward Choi. Exploration into translation-equivariant image quantization. In International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023

  260. [269]

    Bridging the gap between ethics and practice: Guidelines for reliable, safe, and trustworthy human-centered AI systems

    Ben Shneiderman. Bridging the gap between ethics and practice: Guidelines for reliable, safe, and trustworthy human-centered AI systems. ACM Transactions on Interactive Intelligent Systems, 10, 2020

  261. [270]

    David Silver and Richard S. Sutton. Welcome to the era of experience. https://storage.googleapis.com/deepmind-media/Era-of-Experience% 20/The%20Era%20of%20Experience%20Paper.pdf, 2025

  262. [271]

    David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap,...

  263. [272]

    Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis

    David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy P. Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Maste...

  264. [273]

    A general reinforcement learning algorithm that masters chess, shogi, and go through self-play

    David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, and Demis Hassabis. A general reinforcement learning algorithm that masters chess, shogi, and go through self-play. Sc...

  265. [274]

    David Silver, Satinder Singh, Doina Precup, and Richard S. Sutton. Reward is enough. Artificial Intelligence, 299:103535, 2021

  266. [275]

    Georg Simmel. Fashion. American Journal of Sociology, 62(6):541–558, 1957

  267. [276]

    Herbert A. Simon. Models of Man: Social and Rational. Wiley, 1957

  268. [277]

    Very deep convolutional networks for large- scale image recognition

    Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large- scale image recognition. International Conference on Learning Representations (ICLR), 2015

  269. [278]

    Denoising diffusion implicit models

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations (ICLR), 2021

  270. [279]

    Baruch Spinoza. Ethics. Anonymous, 1677. English translation included in *A Spinoza Reader: The Ethics and Other Works*, translated and edited by Edwin Curley, Princeton University Press, 1994

  271. [280]

    Dropout: A simple way to prevent neural networks from overfitting

    Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15:1929–1958, 2014

  272. [281]

    The Berg Companion to Fashion

    Valerie Steele (ed.). The Berg Companion to Fashion. Bloomsbury Publishing, London, 2015

  273. [282]

    Fusarial Wilt (Panama Disease) of Bananas and Other Musa Species

    Robert Harry Stover. Fusarial Wilt (Panama Disease) of Bananas and Other Musa Species. Commonwealth Mycological Institute, 1962

  274. [283]

    Pilgrim in the Microworld

    David Sudnow. Pilgrim in the Microworld. Warner Books, New York, 1983

  275. [284]

    Richard S. Sutton. The reward hypothesis. http://incompleteideas.net/rlai.cs. ualberta.ca/RLAI/rewardhypothesis.html, 2004. 285

  276. [285]

    Sutton and Andrew G

    Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, second edition, 2018

  277. [286]

    Going deeper with convolutions

    Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015

  278. [287]

    Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich

    Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015

  279. [288]

    #exploration: A study of count-based ex- ploration for deep reinforcement learning

    Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel. #exploration: A study of count-based ex- ploration for deep reinforcement learning. InAdvances in Neural Information Processing Systems (NeurIPS), 2017

  280. [289]

    Thorndike

    Edward L. Thorndike. Animal intelligence: An experimental study of the associative processes in animals. The Psychological Review: Monograph Supplements , 2(4):i–109,

  281. [290]

    On aims and methods of ethology

    Niko Tinbergen. On aims and methods of ethology. Zeitschrift für Tierpsychologie, 20 (4):410–433, 1963

  282. [291]

    Champandard, Pier Luca Lanzi, Michael Mateas, Ana Paiva, Mike Preuss, and Kenneth O

    Julian Togelius, Alex J. Champandard, Pier Luca Lanzi, Michael Mateas, Ana Paiva, Mike Preuss, and Kenneth O. Stanley. Procedural content generation: Goals, challenges and actionable steps. In Artificial and Computational Intelligence in Games . Schloss Dagstuhl - Leibniz-Zent...

  283. [292]

    Behavioral cloning from observation

    Faraz Torabi, Garrett Warnell, and Peter Stone. Behavioral cloning from observation. In International Joint Conference on Artificial Intelligence (IJCAI), 2018

  284. [293]

    Generative adversarial imitation from observation

    Faraz Torabi, Garrett Warnell, and Peter Stone. Generative adversarial imitation from observation. In ICML Workshop on Imitation, Intent, and Interaction (I3), 2019. 286

  285. [294]

    Klassen, Richard Anthony Valenzano, and Sheila A

    Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Anthony Valenzano, and Sheila A. McIl- raith. Reward machines: Exploiting reward function structure in reinforcement learning. Journal of Artificial Intelligence Research, 73:173–208, 2022

  286. [295]

    Is deep reinforcement learning really superhuman on atari? leveling the playing field

    Marin Toromanoff, Emilie Wirbel, and Fabien Moutarde. Is deep reinforcement learning really superhuman on atari? leveling the playing field. CoRR, abs/1908.04683, 2019

  287. [296]

    Pure end-to-end training car racing game ai bot using deep reinforce- ment learning

    Cheng-Lun Tsai. Pure end-to-end training car racing game ai bot using deep reinforce- ment learning. Master’s thesis, National Chiao Tung University, 2018

  288. [1781]

    Translation based on the 1781 (A) and 1787 (B) editions

    This edition published in 1998. Translation based on the 1781 (A) and 1787 (B) editions

  289. [1889]

    Title translation: The Principles of Arithmetic, Presented by a New Method

    In Latin. Title translation: The Principles of Arithmetic, Presented by a New Method

  290. [1898]

    Doctoral dissertation, Columbia University

  291. [2025]

    URL https://www.blog.udonis.co/mobile-marketing/mobile-games/ gaming-industry. Udonis

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.