Pith. sign in

REVIEW 4 major objections 4 minor 7 cited by

A theory of appropriateness with applications to generative artificial intelligence

T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper argues that human decision making is pattern completion, not reward optimization, and that AI alignment should be reframed as context-specific appropriateness.

desk verdict A genuinely synthetic theory paper, well-written and honest about its limits, but the central equivalence claim hinges on an underspecified 'sufficiently powerful' assumption and needs empirical teeth. read the letter →

arxiv 2412.19010 v1 pith:BCOXV4KG submitted 2024-12-26 cs.AI

classification cs.AI
keywords appropriatenesspredictivepatterncompletionlargelanguagemodelsreward-freedecisionmakingnormsconventionssanctioningAIalignment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Appropriateness—the sense that some actions fit a situation and others do not—is the paper's central topic. The paper argues that human decision making is best explained by predictive pattern completion, an autoregressive next-symbol operation like the one inside a large language model, rather than by reward or utility maximization. It claims that every phenomenon traditionally explained with scalar rewards or utilities can be equivalently explained without them, and that the reward-free version is more parsimonious. On this basis it redefines norms as generically conventional patterns of sanctioning and proposes that AI alignment be replaced by appropriateness tailored to specific communities and contexts. If correct, the theory would reorient both cognitive science and the governance of generative AI.

What carries the argument

The central object is predictive pattern completion: an autoregressive next-symbol prediction model, implemented like a large language model, that samples continuations conditioned on a global workspace containing memory assemblies and current perceptions. It carries the argument by replacing reward maximization with pattern completion as the fundamental operation; goals become text in the conditioning sequence, so maximizing a stated reward is just one possible completion. The other key machinery is the formal account of conventions and norms: convention-sensitive actors reproduce patterns because of the weight of precedent, and normative behavior is defined as behavior suggested by generically conventional patterns of sanctioning. These definitions are what let the paper connect individual decision making to social stability, norm change, and sanctioning.

What would settle it

A decisive observation would be a human choice pattern, driven by acute pain, hunger, or pharmacological reward, that a simple reward-based model predicts substantially better than an LLM-based pattern-completion actor given only textual observations and no reward labels.

Watch

Extended reading notes

Core claim

The paper's central claim is that pattern completion is the fundamental operation of human decision making, not reward optimization. It models an individual as a generative actor whose choices are autoregressive samples from a next-symbol predictor conditioned on a global workspace that holds recalled memories and current observations; the decision procedure is the 'logic of appropriateness' rendered as a chain of thought: what kind of situation is this, what kind of person am I, what does a person like me do. Because any goal can appear as a text fragment in the conditioning sequence, a sufficiently powerful predictor can simulate reward maximization, so the pattern-completion theory is claimed to have explanatory power equal to reward/utility theories. The paper further claims it is more parsimonious, since it never requires the modeler to specify where rewards come from, and it handles endogenous preference formation more naturally. Normative appropriateness is then defined as behavior suggested by generically conventional patterns of sanctioning, with conventions understood as patterns reproduced because of the weight of precedent, yielding the stylized facts of context dependence, arbitrariness, automaticity, dynamism, and sanctioning.

Load-bearing premise

The argument stands on treating a human as a text-pretrained next-symbol predictor with no reward signal; if human decisions are causally shaped by non-linguistic, embodied reward processing that such a predictor cannot reproduce, the equivalence claim collapses.

Editorial extensions

If this is right

  • If the reward-free equivalence holds, modelers of human choice will no longer need to write down a reward function; preferences can be treated as patterns in memory and culture.
  • AI alignment would be redefined as appropriateness for a given community and context, with sanctioning as the natural interface for steering systems.
  • Norms and conventions become formal, computable objects: patterns reproduced because of precedent, making norm change and polarization open to computational analysis.
  • Generative agents built on LLMs become scientific models of human social behavior, testable in linguistic multi-actor environments.
  • The theory predicts that failures of AI systems such as jailbreaks and sycophancy are better described as norm violations than as goal misgeneralization.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extension: if pattern completion is sufficient, then the burden of proof shifts to reward-based theories to exhibit a behavior that a text-pretrained predictor cannot reproduce; no such behavior is identified in the paper.
  • Extension: non-linguistic embodied drivers such as pain, hunger, and drug reward would need to be re-described as inputs to the predictor rather than rewards; testing that re-description is a natural next step.
  • Extension: the view of norms as conflict-resolution technology suggests that AI governance should be evaluated by whether it keeps losing groups engaged, not by whether it finds consensus values.
  • Extension: one could measure the paper's parsimony claim by comparing how many free parameters an LLM actor needs versus a reward model to fit the same corpus of human choice data.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a unified theory of appropriateness, defined as socially constructed standards that guide human and AI behavior, and argues that human decision making should be modeled as predictive pattern completion by an LLM-like next-symbol predictor rather than as reward or utility maximization. It introduces a formal actor model with memory, global workspace, and summary functions; defines conventions as patterns reproduced by weight of precedent and norms as behaviors supported by generically conventional patterns of sanctioning; and uses this machinery to explain stylized facts about context dependence, arbitrariness, automaticity, dynamism, and sanctioning. The final parts apply the framework to generative AI, proposing decentralized post-training, sanctioning as a universal interface, and norm-sensitive AI as an alternative to alignment-based approaches.

Significance. If the central equivalence claim were established, the paper would offer a substantial reframing of both cognitive science and AI alignment: instead of optimizing a reward, agents can be understood as completing patterns from culturally shared contexts, and alignment becomes the deployment of appropriateness in specific communities. The manuscript is genuinely ambitious and interdisciplinary, and it has several concrete strengths: a formal actor model in Section 5.2.6 with explicit equations; definitions of convention-sensitivity and sanctioning in Section 6.2-6.3 that are precise enough to simulate; a concrete mechanism for polarization in Section 5.6.4; and an extensive engagement with philosophical and social-science literature. However, the paper's load-bearing claim—that pattern completion has explanatory power equal to reward/utility theories—is asserted rather than demonstrated, and the parsimony argument is not carried out quantitatively. The manuscript therefore currently reads as a rich theoretical proposal whose central thesis needs substantially more support.

major comments (4)
  1. [§5.5, with §5.2.3] The central equivalence claim is not established. The construction in Section 5.5 conditions a pattern completion network on 'The goal of game G is to maximize r' and the current state, and asserts that a 'sufficiently powerful' p will emit the maximizing action. But Section 5.2.3 models p as an LLM pretrained on text, and the paper gives no argument or evidence that a text-pretrained next-symbol predictor solves arbitrary maximization problems, especially for reward functions or domains poorly represented in its pretraining distribution. If 'sufficiently powerful' means an arbitrary computational function, the equivalence is a Turing-completeness remark with no bearing on the proposed human decision model; if it means the LLM of Section 5.2.3, the claim is an unverified empirical hypothesis. In addition, the construction embeds a quantified scalar reward r directly in the conditioning text. Even if no scalar reward appears in the actor's loss function, the decision-relevant context still contains a reward specification, which weakens the paper's stated claim that the theory has 'no role for quantified scalar representations.'
  2. [§5.5 (second parsimony argument)] The parsimony comparison is not actually performed. The paper argues that the pattern completion theory is simpler because it avoids positing a reward function, but the proposed actor model includes many additional free components: the similarity threshold epsilon, the counter-factual weight of precedent f, the rate of reproduction r, the number of global workspace assemblies K, and the choice of summary functions q_k and framing functions phi_k. A reward-based model may require specifying a reward, but the pattern completion model requires specifying all of these functions and hyperparameters. The paper does not compare the total specification complexity of the two frameworks, so the conclusion that pattern completion is 'the more parsimonious' is unsupported. This is load-bearing because the authors explicitly state in Section 5.5 that their argument is one of parsimony, not proof of superior fit.
  3. [§6.1, §6.2.1, §6.3] The account of norm emergence is close to tautological. Definitions 3-6 define actors as convention-sensitive, and Section 6.1 lists as hypotheses that convention-sensitive and sanction-sensitive actors 'can develop' conventional sanctioning patterns that create norms. Section 6.3 then defines norms as behaviors induced by generically conventional sanctioning. Given these definitions, a population of actors that are defined to follow precedent and respond to sanctions will, by construction, exhibit behavior that the paper calls normative. The paper should specify an independent behavioral prediction—for example, a condition under which a population of such actors would fail to develop a particular norm—and state what empirical observations would falsify the account. Without such content, the explanation of norms is a definitional consequence rather than a scientific explanation.
  4. [§5.2, §5.3.2] The foundational assumption that human decision making can be modeled by an LLM pretrained on text is stated but not defended against non-linguistic, embodied accounts of decision making. Section 5.2 says 'We model p as an LLM,' and Section 5.3.2 assumes 'Within a culture, all adult individuals share approximately the same pattern completion network p.' The paper's equivalence argument in Section 5.5 depends on this assumption, because it claims that all behavior explicable by reward maximization can be reproduced by pattern completion over a linguistic context. If human choices are causally shaped by non-linguistic reward processing, as is common in affective neuroscience and embodied cognition, the pattern completion model may be an incomplete description. The paper acknowledges that the language domain is 'fully general' only in the sense that any action can be described in text, but description is not the same as causal mechanism. This point needs explicit discussion and, ideally, supporting empirical evidence.
minor comments (4)
  1. [§5.2.6, Eq. (1)] The symbol 'B' used in Equations (1) and (3) is not defined; it appears to mean 'is defined as' or 'is implemented by,' but it should be stated explicitly.
  2. [§6.2.1, Eqs. (10)-(13)] The displayed equations contain unbalanced parentheses and inconsistent typesetting, for example the extra closing parenthesis in Eq. (10) and the misplaced double parenthesis in Eq. (13). These should be corrected so the definitions are unambiguous.
  3. [§5.5] The phrase 'all the data traditionally explained using reward-based and utility-based theories' is too vague to be tested. The paper should either list specific datasets or experimental paradigms, or clarify that the claim is about a class of models rather than empirical data.
  4. [§4] The stylized facts are clearly presented, but the connection between each stylized fact and a specific theorem or derivation in Sections 5-6 could be made more explicit; currently the explanatory link is often asserted in prose rather than shown from the formal definitions.

Circularity Check

2 steps flagged · score 6.0 of 10

Norm emergence and the reward-free equivalence reduce to their own definitions; the 'pattern completion is all you need' argument assumes the maximization capability it claims to explain.

  1. self definitional [Glossary (p. 3); Section 6.3.2 heading]
    "Norm: A behavior is normative when it is encouraged (or its complement discouraged) by a generically scoped conventional pattern of sanctioning. "Norms are induced by generically conventional patterns of sanctioning""

    The Glossary defines 'norm' in terms of generically scoped conventional sanctioning, and Section 6.3.2 announces that norms are induced by exactly that pattern. The explanation's explanans is the definiens of its explanandum: any behavior produced by a generically conventional sanctioning pattern is normative by definition. The paper's 'hypothesis' that convention- and sanction-sensitive actors develop such patterns is an additional empirical assumption, not a derivation. Thus the claimed account of norm emergence restates the definition rather than explaining how norms arise.

  2. self definitional [Section 5.5, 'Pattern completion is all you need']
    "Consider a sequence of perceived symbols beginning with the words 'The goal of game G is to maximize r', then later in the sequence the words 'the current state of G is s, what does actor A choose to do next?'. As long as the pattern completion network p is sufficiently powerful then it will emit the action most likely to maximize r, perhaps after first emitting some number of other symbols corresponding to reasoning steps."

    The advertised equivalence is conditional on p being powerful enough to solve the maximization problem posed in the prompt—the very capability the reward/utility theory is supposed to explain. If 'sufficiently powerful' means 'emits the action most likely to maximize r', the simulation is true by construction and has no independent content; if it means 'LLM pretrained on text', it is an unsupported empirical claim. The scalar r is also placed in the conditioning sequence, so the theory's claim of 'no role for quantified scalar representations of rewards' is not discharged: the reward signal enters as an input symbol.

full rationale

This paper is a conceptual framework rather than an empirical study. The genuinely constructive content—the LMAE formalism, the memory/global-workspace actor model, summary functions, endogenous preference formation, and the AI governance proposals—is self-contained and does not reduce to its inputs. However, two load-bearing moves are circular as stated. First, 'norm' is defined as behavior encouraged by a generically scoped conventional pattern of sanctioning, and the paper then presents 'Norms are induced by generically conventional patterns of sanctioning' as a result. The definition already contains the mechanism that is supposed to explain norms; the only non-tautological part, that convention- and sanction-sensitive actors will actually develop such patterns, is assumed as a hypothesis rather than derived. Second, the central equivalence argument in Section 5.5 makes the conclusion true by assumption: a sufficiently powerful pattern-completion network is one that will emit the action most likely to maximize r when the prompt contains the reward specification. This is not a derivation from the pattern-completion model; it is a restatement of the capability to be explained. It also reintroduces the scalar reward r into the conditioning context, undercutting the paper's claim that its theory has no role for quantified scalar reward representations. The paper's reliance on Vezhnevets et al. (2023), a co-authored prior work, for the two assumptions underlying the model is worth noting, but the assumptions are stated openly rather than concealed behind the citation, so I do not treat the self-citation as an additional circular step. The overall score reflects partial circularity in the two central claims, not a wholesale absence of independent content.

Assumptions & free parameters 5 free parameters · 6 assumptions · 1 invented entities

The paper's central model rests on describing human decision-making as next-symbol prediction in an LLM-like network. The formal definitions introduce several uncalibrated parameters (epsilon, f, r, K) without empirical grounding. The theory additionally assumes language fully captures action and that all actors in a culture share the same pattern-completion network.

free parameters (5)
  • epsilon (meaning similarity threshold)
    Used in Definition 1 to define when two utterances have epsilon-similar meaning; no value specified or estimated.
  • f (counter-factual weight of precedent)
    Defined in Definition 4 as the fraction of memories edited in counterfactual operations; used to define convention sensitivity but not calibrated.
  • r (rate of reproduction)
    Used in Definition 6 to define reproduction rate; no empirical value given.
  • K (number of global workspace assemblies)
    Introduced in Section 5.2.6 as a formal parameter for the number of summary-function-generated assemblies; no method for choosing it is provided.
  • Choice of summary functions {q_k}
    The paper says 'a broad range of interdependent summary functions can be used' (Section 5.2.5) but does not specify which ones are needed for the theory's predictions.
assumptions (6)
  • domain assumption Human decision-making can be modeled as an autoregressive next-symbol predictor (LLM) trained on large text corpora.
    Stated in Section 5.2.3 and Section 5.5; this is the core computational assumption and is not empirically validated.
  • domain assumption All situations and actions can be fully described in natural language text.
    Section 5.1 asserts the LMAE framework is generic because language is generic.
  • ad hoc to paper Within a culture, all adult individuals share approximately the same pattern completion network p.
    Section 6.2 makes this simplifying assumption to derive collective properties; the paper acknowledges it is an idealization.
  • ad hoc to paper A sufficiently powerful pattern completion network can simulate any reward maximization behavior.
    Section 5.5 asserts this without proof; it is the basis for the claimed equivalence to reward-based theories.
  • domain assumption Appropriateness functions as a culturally evolved conflict resolution technology.
    The paper's thesis in Sections 1 and 6; it is a hypothesis rather than an established fact.
  • standard math Standard probability theory and Arrow's theorem apply to the formal model.
    Used in definitions and the Arrow's theorem discussion in Section 5.6.3.
invented entities (1)
  • Linguistic Multi-Actor Environment (LMAE)
    purpose: A formal text-based multi-agent environment used to ground the computational model of human decision making.
    Introduced in Section 5.1 as a generic modeling framework; no direct empirical referent.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A theory of appropriateness with applications to generative artificial intelligence." pith.science (2026). https://pith.science/paper/BCOXV4KG

@misc{pith2026241219010,
  author       = {Pith},
  title        = {Pith review of: A theory of appropriateness with applications to generative artificial intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BCOXV4KG}},
  note         = {Machine review of arXiv:2412.19010}
}
read the original abstract

What is appropriateness? Humans navigate a multi-scale mosaic of interlocking notions of what is appropriate for different situations. We act one way with our friends, another with our family, and yet another in the office. Likewise for AI, appropriate behavior for a comedy-writing assistant is not the same as appropriate behavior for a customer-service representative. What determines which actions are appropriate in which contexts? And what causes these standards to change over time? Since all judgments of AI appropriateness are ultimately made by humans, we need to understand how appropriateness guides human decision making in order to properly evaluate AI decision making and improve it. This paper presents a theory of appropriateness: how it functions in human society, how it may be implemented in the brain, and what it means for responsible deployment of generative AI technology.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power

    cs.AI 2025-07 conditional novelty 7.0 of 10

    A new AI objective, ICCEA power, aggregates humans' ability to reach many possible goals with inequality and risk aversion, and a soft-maximizing agent learns cooperative behavior without knowing human goals.

  2. Structural transparency of societal AI alignment through Institutional Logics

    cs.CY 2026-02 conditional novelty 6.0 of 10

    Introduces a five-component analytical framework, grounded in Institutional Logics, for making visible the organizational and institutional decisions that shape AI alignment.

  3. Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions

    cs.MA 2025-07 conditional novelty 6.0 of 10

    Language model agents with personality and theory-of-mind prompts reproduce human third-party punishment and gossip effects, and predict lower anonymous punishment and higher cooperation after group discussion.

  4. When One LLM Drools, Multi-LLM Collaboration Rules

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A position paper that introduces a four-level taxonomy of multi-LLM collaboration (API, text, logit, weight) and argues it is essential for reliability, pluralism, and democratization.

  5. Virtual Agent Economies

    cs.AI 2025-09 conditional novelty 5.0 of 10

    Proposes a two-axis framework (emergent versus intentional, permeable versus impermeable) for the coming AI agent economy and argues for proactively designing steerable agent markets.

  6. Multi-Actor Generative Artificial Intelligence as a Game Engine

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Generative multi-actor AI platforms can be built on the Entity-Component pattern, treating the environment (Game Master) as a composable entity, so that one library serves simulation, storytelling, and evaluation goals.

  7. Jackpot! Alignment as a Maximal Lottery

    cs.AI 2025-01 conditional novelty 5.0 of 10

    Nash Learning from Human Feedback is shown to approximate the maximal lottery voting rule, which the authors argue better reflects majority preferences than Borda-like RLHF.

Reference graph

Works this paper leans on

71 extracted references · 29 canonical work pages · cited by 7 Pith papers

  1. [1]

    Abdulhai, G

    M. Abdulhai, G. Serapio-Garcia, C. Crepy, D. Valter, J. Canny, and N. Jaques. Moral foundations of large language models.arXiv preprint arXiv:2310.15337,

  2. [5]

    Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McK- innon, et al. Constitutional AI: Harmlessness from AI feedback.arXiv preprint arXiv:2212.08073,

  3. [7]

    Bourdieu

    P. Bourdieu. Distinction a social critique of the judgement of taste. InInequality, pages 287–318. Routledge, 1984/2018. S. Bowles. Endogenous preferences: The cultural consequences of markets and other economic institutions. Journal of economic literature, 36(1):75–111,

  4. [13]

    Dafoe, E

    A. Dafoe, E. Hughes, Y. Bachrach, T. Collins, K. R. McKee, J. Z. Leibo, K. Larson, and T. Graepel. Open problems in cooperative AI.arXiv preprint arXiv:2012.08630,

  5. [19]

    Hadfield-Menell and G

    D. Hadfield-Menell and G. K. Hadfield. Incomplete contracting and ai alignment. InProceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, pages 417–422,

  6. [23]

    J. J. Horton. Large language models as simulated economic agents: What can we learn from homo silicus? arXiv preprint arXiv:2301.07543,

  7. [24]

    Huang, F

    W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, etal. Innermonologue: Embodiedreasoningthroughplanningwithlanguagemodels. arXivpreprint arXiv:2207.05608,

  8. [25]

    Jaques, J

    N. Jaques, J. H. Shen, A. Ghandeharioun, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard. Human-centric dialog training via offline reinforcement learning. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3985–4003,

Show all 71 references
  1. [27]

    Kaplan, S

    J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361,

  2. [28]

    Köster, K

    104 A theory of appropriateness with applications to generative artificial intelligence R. Köster, K. R. McKee, R. Everett, L. Weidinger, W. S. Isaac, E. Hughes, E. A. Duéñez-Guzmán, T. Graepel, M. Botvinick, and J. Z. Leibo. Model-free conventions in multi-agent reinforcement...

  3. [29]

    A. K. Lampinen, I. Dasgupta, S. C. Chan, H. R. Sheahan, A. Creswell, D. Kumaran, J. L. McClelland, and F. Hill. Language models show human-like content effects on reasoning tasks.arXiv preprint arXiv:2207.07051,

  4. [30]

    S. Lazar. Governing the algorithmic city.arXiv preprint arXiv:2410.20720,

  5. [32]

    common is moral

    C. Li, Z. Gan, Z. Yang, J. Yang, L. Li, L. Wang, and J. Gao. Multimodal foundation models: From specialists to general-purpose assistants.arXiv preprint arXiv:2309.10020, 10, 2023a. G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem. CAMEL: Communicative agents for “min...

  6. [33]

    Machery and S

    E. Machery and S. Stich. The Moral/Conventional Distinction. In E. N. Zalta, editor,The Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University, Summer 2022 edition,

  7. [36]

    Accessed: 2024-11-21. S. Mathew and R. Boyd. Punishment sustains large-scale cooperation in prestate warfare.Proceedings of the National Academy of Sciences, 108(28):11375–11380,

  8. [37]

    Mills and H

    S. Mills and H. S. Sætra. Algorithms in the room: AI, representation, and decisions about sustainable futures. Representation, and Decisions about Sustainable Futures (September 10, 2024),

  9. [38]

    Mukobi, H

    G. Mukobi, H. Erlebach, N. Lauffer, L. Hammond, A. Chan, and J. Clifton. Welfare diplomacy: Benchmarking language model cooperation.arXiv preprint arXiv:2310.08901,

  10. [40]

    was it “stated

    R. Patel and E. Pavlick. “was it “stated” or was it “claimed”?: How linguistic bias affects generative language models. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10080–10095,

  11. [41]

    Perez, S

    E. Perez, S. Ringer, K. Lukoši¯ut˙e, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath, et al. Discovering language model behaviors with model-written evaluations.arXiv preprint arXiv:2212.09251,

  12. [42]

    Perez, C

    J. Perez, C. Léger, M. Ovando-Tellez, C. Foulon, J. Dussauld, P.-Y. Oudeyer, and C. Moulin-Frier. Cultural evolution in populations of large language models.arXiv preprint arXiv:2403.08882,

  13. [48]

    Safdari, G

    M. Safdari, G. Serapio-García, C. Crepy, S. Fitz, P. Romero, L. Sun, M. Abdulhai, A. Faust, and M. Matarić. Personality traits in large language models.arXiv preprint arXiv:2307.00184,

  14. [49]

    D. Sally. Conversation and cooperation in social dilemmas: A meta-analysis of experiments from 1958 to 1992.Rationality and society, 7(1):58–92,

  15. [50]

    Schrimpf, I

    M. Schrimpf, I. Blank, G. Tuckute, C. Kauf, E. A. Hosseini, N. Kanwisher, J. Tenenbaum, and E. Fe- dorenko. Artificial neural networks accurately predict language processing in the brain.BioRxiv, pages 2020–06,

  16. [52]

    Shanahan

    M. Shanahan. The Frame Problem. In E. N. Zalta, editor,The Stanford Encyclopedia of Philosophy. Metaphysics Research Lab, Stanford University, Spring 2016 edition,

  17. [53]

    Sharma, M

    M. Sharma, M. Tong, T. Korbak, D. Duvenaud, A. Askell, S. R. Bowman, N. Cheng, E. Durmus, Z. Hatfield-Dodds, S. R. Johnston, et al. Towards understanding sycophancy in language models. arXiv preprint arXiv:2310.13548,

  18. [54]

    Sorensen, J

    T. Sorensen, J. Moore, J. Fisher, M. Gordon, N. Mireshghallah, C. M. Rytting, A. Ye, L. Jiang, X. Lu, N. Dziri, et al. A roadmap to pluralistic alignment.arXiv preprint arXiv:2402.05070,

  19. [56]

    Taylor, M

    R. Taylor, M. Kardas, G. Cucurull, T. Scialom, A. Hartshorn, E. Saravia, A. Poulton, V. Kerkez, and R. Stojnic. Galactica: A large language model for science.arXiv preprint arXiv:2211.09085,

  20. [57]

    M. H. Tessler, J. Madeano, P. A. Tsividis, B. Harper, N. D. Goodman, and J. B. Tenenbaum. Learn- ing to solve complex tasks by growing knowledge culturally across generations.arXiv preprint arXiv:2107.13377,

  21. [59]

    A. S. Vezhnevets, J. P. Agapiou, A. Aharon, R. Ziv, J. Matyas, E. A. Duéñez-Guzmán, W. A. Cunningham, S. Osindero, D. Karmon, and J. Z. Leibo. Generative agent-based modeling with actions grounded in physical, social, or digital space using concordia.arXiv preprint arXiv:2312.03664,

  22. [61]

    Weidinger, J

    L. Weidinger, J. Mellor, B. G. Pegueroles, N. Marchal, R. Kumar, K. Lum, C. Akbulut, M. Diaz, S. Bergman, M. Rodriguez, et al. STAR: Sociotechnical approach to red teaming language models. arXiv preprint arXiv:2406.11757,

  23. [64]

    114 A theory of appropriateness with applications to generative artificial intelligence Q. Wu, G. Bansal, J. Zhang, Y. Wu, S. Zhang, E. Zhu, B. Li, L. Jiang, X. Zhang, and C. Wang. Autogen: Enabling next-gen llm applications via multi-agent conversation framework. arXiv prepri...

  24. [65]

    A. Xu, E. Pathak, E. Wallace, S. Gururangan, M. Sap, and D. Klein. Detoxifying language models risks marginalizing minority voices. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,...

  25. [66]

    Yaman, J

    A. Yaman, J. Z. Leibo, G. Iacca, and S. W. Lee. The emergence of division of labour through decentralized social sanctioning.Proceedings of the Royal Society B: Biological Sciences, 290(2009): 20231716,

  26. [67]

    J. C. Yang, M. Korecki, D. Dailisan, C. I. Hausladen, and D. Helbing. LLM voting: Human choices and AI collective decision making.arXiv preprint arXiv:2402.01766, 2024a. L. Yang, S. Constantino, B. Grenfell, E. Weber, S. Levin, and V. Vasconcelos. Sociocultural determi- nants ...

  27. [68]

    S. Yang, J. Walker, J. Parker-Holder, Y. Du, J. Bruce, A. Barreto, P. Abbeel, and D. Schuurmans. Video as the new language for real-world decision making.arXiv preprint arXiv:2402.17139, 2024b. S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan. Tree o...

  28. [70]

    C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y. Mao, X. Ma, A. Efrat, P. Yu, L. Yu, et al. LIMA: Less is more for alignment.arXiv preprint arXiv:2305.11206,

  29. [71]

    D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving. Fine-tuning language models from human preferences.arXiv preprint arXiv:1909.08593,

  30. [1956]

    Trivedi, N

    R. Trivedi, N. Chandak, A. I. Muresanu, S. Zhu, A. Sarkar, J. Z. Leibo, D. Hadfield-Menell, and G. K. Hadfield. Altared environments: Normative institutions support AI alignment for cooperation in multi-agent systems. InAgentic Markets Workshop at ICML 2024,

  31. [1957]

    Z. Zhao, W. S. Lee, and D. Hsu. Large language models as commonsense knowledge for large-scale task planning.arXiv preprint arXiv:2305.14078,

  32. [1977]

    105 A theory of appropriateness with applications to generative artificial intelligence Y. Mao, M. G. Reinecke, M. Kunesch, E. A. Duéñez-Guzmán, R. Comanescu, J. Haas, and J. Z. Leibo. Doing the right thing for the right reason: Evaluating artificial moral cognition by probing...

  33. [1980]

    C. S. Crandall, J. M. Miller, and M. H. White. Changing norms following the 2016 US presidential election: The Trump effect on prejudice.Social Psychological and Personality Science, 9(2):186–192,

  34. [1988]

    J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein. Generative agents: Interactive simulacra of human behavior.arXiv preprint arXiv:2304.03442,

  35. [1990]

    Plappert, R

    M. Plappert, R. Sampedro, T. Xu, I. Akkaya, V. Kosaraju, P. Welinder, R. D’Sa, A. Petron, H. P. d. O. Pinto, A. Paino, H. Noh, L. Weng, Q. Yuan, C. Chu, and W. Zaremba. Asymmetric self-play for automatic goal discovery in robotic manipulation.Preprint arXiv:2101.04882,

  36. [1991]

    Carroll, D

    M. Carroll, D. Foote, A. Siththaranjan, S. Russell, and A. Dragan. AI alignment with changing and influenceable reward functions.arXiv preprint arXiv:2405.17713,

  37. [1994]

    Rorty.Philosophy and the Mirror of Nature

    109 A theory of appropriateness with applications to generative artificial intelligence R. Rorty.Philosophy and the Mirror of Nature. Princeton university press, 1978/2009. R. Rorty.Pragmatism as Anti-authoritarianism. Harvard University Press,

  38. [1995]

    S. Hong, X. Zheng, J. Chen, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu, and J. Schmidhuber. Metagpt: Meta programming for multi-agent collaborative framework. arXiv preprint arXiv:2308.00352,

  39. [1996]

    M. B. Johanson, E. Hughes, F. Timbers, and J. Z. Leibo. Emergent bartering behaviour in multi-agent reinforcement learning.arXiv preprint arXiv:2205.06760,

  40. [1997]

    J. X. Wang, E. Hughes, C. Fernando, W. M. Czarnecki, E. A. Duéñez-Guzmán, and J. Z. Leibo. Evolving intrinsic motivations for altruistic behavior. InProceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems, pages 683–692, 2019a. R. Wang, J. ...

  41. [1998]

    Firebaugh and K

    G. Firebaugh and K. E. Davis. Trends in antiblack prejudice, 1972-1984: Region and cohort effects. American Journal of Sociology, 94(2):251–272,

  42. [1999]

    doi: 10.1162/003355399556151. B. Felbo, A. Mislove, A. Søgaard, I. Rahwan, and S. Lehmann. Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm. InProceed- ings of the 2017 Conference on Empirical Methods in Natur...

  43. [2000]

    Haidt et al

    J. Haidt et al. The moral emotions.Handbook of affective sciences, 11(2003):852–870,

  44. [2001]

    Glaese, N

    A. Glaese, N. McAleese, M. Trębacz, J. Aslanides, V. Firoiu, T. Ewalds, M. Rauh, L. Weidinger, M. Chadwick, P. Thacker, L. Campbell-Gillingham, J. Uesato, P.-S. Huang, R. Comanescu, F. Yang, A. See, S. Dathathri, R. Greig, C. Chen, D. Fritz, J. Sanchez Elias, R. Green, S. Mokr...

  45. [2003]

    J. P. Agapiou, A. S. Vezhnevets, E. A. Duéñez-Guzmán, J. Matyas, Y. Mao, P. Sunehag, R. Köster, U. Madhushani, K. Kopparapu, R. Comanescu, D. Strouse, M. B. Johanson, S. Singh, J. Haas, I. Mordatch, D. Mobbs, and J. Z. Leibo. Melting pot 2.0.arXiv preprint arXiv:2211.13746,

  46. [2004]

    Goalmisgeneralization: why correct specifications aren’t enough for correct goals.arXiv preprint arXiv:2210.01790,

    R.Shah,V.Varma,R.Kumar,M.Phoung,V.Krakovna,J.Uesato,andZ.Kenton. Goalmisgeneralization: why correct specifications aren’t enough for correct goals.arXiv preprint arXiv:2210.01790,

  47. [2006]

    Stastny, M

    111 A theory of appropriateness with applications to generative artificial intelligence J. Stastny, M. Riché, A. Lyzhov, J. Treutlein, A. Dafoe, and J. Clifton. Normative disagreement as a challenge for cooperative ai.arXiv preprint arXiv:2111.13872,

  48. [2007]

    J. Z. Leibo, E. Hughes, M. Lanctot, and T. Graepel. Autocurricula and the emergence of innova- tion from social interaction: A manifesto for multi-agent intelligence research.arXiv preprint arXiv:1903.00742, 2019a. J. Z. Leibo, J. Perolat, E. Hughes, S. Wheelwright, A. H. Marb...

  49. [2008]

    Rafailov, A

    R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn. Direct preference optimization: Your language model is secretly a reward model.arXiv preprint arXiv:2305.18290,

  50. [2009]

    Phelps and Y

    S. Phelps and Y. I. Russell. Investigating emergent goal-like behaviour in large language models using experimental economics.arXiv preprint arXiv:2305.07970,

  51. [2010]

    G. H. Resende, L. F. Nery, F. Benevenuto, S. Zannettou, and F. Figueiredo. A comprehensive view of the biases of toxicity and sentiment analysis methods towards utterances with african american english expressions.arXiv preprint arXiv:2401.12720,

  52. [2011]

    doi: 10.1093/oxfordhb/9780199604456.013.0024. A. Marmor. Social conventions: from language to law. InSocial Conventions. Princeton University Press,

  53. [2013]

    Chiappa and W

    S. Chiappa and W. S. Isaac. A causal bayesian networks viewpoint on fairness.Privacy and Identity Management. Fairness, Accountability, and Transparency in the Age of Big Data: 13th IFIP WG 9.2, 9.6/11.7, 11.6/SIG 9.2. 2 International Summer School, Vienna, Austria, August 20-...

  54. [2014]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Che...

  55. [2016]

    Welbl, A

    J. Welbl, A. Glaese, J. Uesato, S. Dathathri, J. Mellor, L. A. Hendricks, K. Anderson, P. Kohli, B. Coppin, and P.-S. Huang. Challenges in detoxifying language models. InFindings of the Association for Computational Linguistics: EMNLP 2021, pages 2447–2469,

  56. [2017]

    Y. Wolf, N. Wies, D. Shteyman, B. Rothberg, Y. Levine, and A. Shashua. Tradeoffs between alignment and helpfulness in language models.arXiv preprint arXiv:2401.16332,

  57. [2018]

    Framework-basedqualitativeanaly- sis of free responses of large language models: Algorithmic fidelity.arXiv preprint arXiv:2309.06364,

    A.Amirova, T.Fteropoulli, N.Ahmed, M.R.Cowie, andJ.Z.Leibo. Framework-basedqualitativeanaly- sis of free responses of large language models: Algorithmic fidelity.arXiv preprint arXiv:2309.06364,

  58. [2019]

    E. A. Duéñez-Guzmán, K. R. McKee, Y. Mao, B. Coppin, S. Chiappa, A. S. Vezhnevets, M. A. Bakker, Y. Bachrach, S. Sadedin, W. Isaac, et al. Statistical discrimination in learning agents.arXiv preprint arXiv:2110.11404,

  59. [2020]

    Bubeck, V

    S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4.arXiv preprint arXiv:2303.12712,

  60. [2021]

    H. L. A. Hart.The concept of law. Oxford University Press, 1961/2012. J. Haugeland. Heidegger on being a person.Nous, pages 15–26,

  61. [2022]

    Blundell, B

    C. Blundell, B. Uria, A. Pritzel, Y. Li, A. Ruderman, J. Z. Leibo, J. Rae, D. Wierstra, and D. Hassabis. Model-free episodic control.arXiv preprint arXiv:1606.04460,

  62. [2023]

    Amodei, C

    D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané. Concrete problems in AI safety. arXiv preprint arXiv:1606.06565,

  63. [2024]

    Q. Dong, L. Li, D. Dai, C. Zheng, Z. Wu, B. Chang, X. Sun, J. Xu, and Z. Sui. A survey for in-context learning. arXiv preprint arXiv:2301.00234,

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.