Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Scopes of Alignment

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Alignment is not one workflow: the paper proposes three scopes — competence, transience, and audience — along which AI alignment should be designed.

desk verdict A clear, readable position paper that gives alignment a useful three-scope vocabulary, but the axes are asserted rather than derived and the categories are under-specified. read the letter →

arxiv 2501.12405 v1 pith:4SSOSS54 submitted 2025-01-15 cs.CY cs.AIcs.CL

classification cs.CYcs.AIcs.CL
keywords AIalignmentlargelanguagemodelsscopedcompetencetransienceaudiencepluralisticvalues
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that AI alignment should not be a single one-size-fits-all process of instilling helpful, harmless, and honest values into a large language model. Instead, alignment work should be chosen according to three scopes: competence (knowledge, skills, or behaviors), transience (episodic, tied to a time and place, or semantic, general across contexts), and audience (mass, public, small-group, or dyadic). The authors' point is that the standard practice — mass semantic alignment for behaviors — is only one cell in a much larger design space. A sympathetic reader should care because different scopes imply different data, optimization, and inference-time methods, and because properly scoping alignment may avoid the value conflicts that pluralistic alignment tries to mediate. The paper offers the framework as a prerequisite for more pluralistic alignment work.

What carries the argument

The organizing device is the three-axis scope framework: competence by knowledge, skills, or behaviors; transience by episodic or semantic; and audience by mass, public, small-group, or dyadic. The framework works as a classification space for alignment goals and methods, letting the authors position existing techniques as particular cells — for example, reward-based preference learning as suitable only for local uses, low-rank adapters as a path to episodic alignment, and mutual theory of mind as a mechanism for dyadic alignment. It carries the argument by turning the vague term 'alignment' into a design space in which each axis is supposed to point to different technical choices.

What would settle it

Find two deployment contexts that differ on exactly one scope — say audience, dyadic versus small-group, with competence and transience held fixed — and measure whether the optimal alignment method or outcome differs meaningfully; if no systematic difference appears across several such pairs, the scopes do not carve alignment at its joints. A complementary test would look for a context that differs on all three scopes yet still requires the same alignment workflow, showing the axes are not the load-bearing choices.

Watch

Extended reading notes

Core claim

The paper's central claim is that 'alignment should not be just one activity, technical approach, or workflow,' but should be 'more specific and scoped for the different needs of different groups over time.' It proposes three dimensions along which an alignment effort can be placed: competence, which covers the knowledge, skills, or behaviors the model must have; transience, which distinguishes episodic alignment bound to a particular time, place, or context from semantic alignment about general facts of the world; and audience, which runs from a single user in a dyadic relationship, through small groups and publics, to mass alignment with all of humanity. The claim is that most current research and frontier-model practice occupies one corner of this space — mass semantic alignment for behaviors such as helpfulness, harmlessness, and honesty — and that this narrowness hides the technical requirements of other scopes. If the claim is right, alignment is a family of workflows with different data sources, learning algorithms, and inference-time mechanisms, not one method applied to one target.

Load-bearing premise

The three scopes are assumed to vary independently and to cover the meaningful decisions in alignment, but the paper's own example changes all three at once when the country changes, so if they co-vary, the proposed design space claims more degrees of freedom than exist.

Editorial extensions

If this is right

  • Researchers would stop treating mass semantic behavioral alignment as the default and would report which scope their method targets.
  • Alignment technologies would diverge by competence: knowledge alignment favors question-answer generation from content, skills favor repeated examples, and behaviors favor preference data with scenarios.
  • Episodic alignment would be done at inference time, for example by swapping low-rank adapters on the fly, rather than by permanently fine-tuning the model.
  • Dyadic and small-group alignment would be bidirectional, with both human and model adapting, unlike the unidirectional alignment used for mass audiences.
  • Properly scoped alignment could resolve many value conflicts before pluralistic mediation is needed, because narrowing the scope can remove conflicting values.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The three axes are proposed as independent but the paper's motivating example varies all three together when moving from one country to another; an editorial reading is that the axes may co-vary, so the true design space may be smaller than the 3-by-2-by-4 grid suggests.
  • A natural extension the paper leaves implicit is an empirical test: hold two scopes constant, vary the third, and compare which data, optimization, or inference-time method succeeds, turning the taxonomy into a predictive theory.
  • If scoping succeeds, evaluation and red-teaming may also need to be scoped, since a model aligned for a dyadic therapy relationship should not be judged by a mass-semantic benchmark.
  • Another extension is to treat scope selection itself as a governed, human-centered design decision, because the paper notes scopes should be found through reflective equilibrium but does not prescribe a process for doing so.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper is a position statement arguing that AI alignment should not be reduced to a single generic activity. It proposes three "scopes" of alignment: competence (knowledge, skills, or behaviors), transience (episodic versus semantic), and audience (mass, public, small-group, or dyadic). The authors claim that prevailing practice, mass semantic alignment for behaviors, is only one point in this larger space, and they use a hypothetical mental-health counseling example in the U.S. versus China to illustrate how the scopes could differ by context. Section 5 argues that the scopes have technical implications, such as full fine-tuning versus inference-time adapters and unidirectional versus bidirectional alignment procedures, and suggests that reflective equilibrium could be used to choose a scope and thereby dissolve value conflicts. The paper is explicitly non-empirical and positions itself as a precursor to pluralistic alignment.

Significance. If the proposed taxonomy is accepted, it offers a useful descriptive vocabulary for discussing alignment targets and draws attention to alignment settings beyond the dominant one. The paper is clear in its motivation and points to a non-empty space, citing existing work on personalized alignment, contextual alignment, episodic memory, and knowledge-oriented alignment. However, its central contribution is a conceptual scheme whose utility depends on the categories being clearly applicable and reasonably complete. The paper does not provide a formal derivation, but for a position paper that is not itself disqualifying; the honesty of the writing and the modest scope of the claims are strengths. The three-axis taxonomy could serve as a starting point for more operational work, but as it stands the framework's applicability is not yet demonstrated.

major comments (4)
  1. [Section 3, competence] The three competence categories (knowledge, skills, behaviors) are not defined with mutually exclusive boundaries. The paper calls 'converting texts into hip-hop raps' a skill and 'politeness' and 'verbosity' behaviors, but both are action-level properties of model output. If the distinction between skill and behavior is not operational, then Section 5's claim that different competences require different data and optimization strategies cannot be grounded. A usable taxonomy needs either an explicit assignment rule or a statement that the categories are overlapping perspectives rather than disjoint dimensions.
  2. [Section 3, transience] The episodic/semantic distinction is introduced with examples but no operational criterion for deciding whether a given knowledge, skill, or behavior is 'bound to a time, place, or other context' or 'general about the world.' For instance, 'obtaining manager pre-approval before booking air tickets' is described as episodic, yet in another organization it could be a standing rule. Without a test or at least a clear decision procedure, two annotators could plausibly classify the same alignment target into different transience categories, which undermines the claimed design space.
  3. [Section 4 and Section 3] The motivating example changes all three axes simultaneously when moving from the U.S. to China: competence shifts from individualized to collectivist, transience from episodic to semantic, and audience from individual to community. This co-variation suggests that the three axes may not be independent as the design-space framing implies, and it does not provide evidence for a 3-by-2-by-4 space. The paper should either give an example where one axis varies while the other two are fixed, or soften the claim that the three scopes are independent dimensions.
  4. [Section 5, reflective equilibrium] The claim that 'we will also implicitly deal with the problem of conflicting values because the reflective equilibrium will have resolved conflicts by narrowing the scope until there are few remaining' is asserted without a method or justification. Narrowing the scope can avoid a conflict, but that is different from resolving it, and no procedure is given for finding a scope with minimal internal conflict. The paper itself acknowledges that 'such a process is clearly necessary,' so this section currently promises more than it delivers.
minor comments (5)
  1. [Section 5] The reference style is inconsistent: the text says 'As Tan Zhi-Xuan et al. submit,' but the bibliography entry is 'Zhi-Xuan et al. 2024.' Please standardize.
  2. [Section 3] The transliteration 's¯adh¯aran.a-dharma and vi´ses.a-dharma' contains formatting artifacts; it should read 'sādhāraṇa dharma' and 'viśeṣa dharma'.
  3. [Section 5] The sentence beginning 'It is important to differentiate the transience of alignment because while semantic alignment implies ' is missing a comma after 'because' and is long enough to obscure the conditional structure; consider splitting it.
  4. [Section 4] The mental-health example is labeled hypothetical, but the paper does not qualify that the characterization of U.S. and Chinese cultural norms is a simplification; adding a caveat would prevent the example from being read as an empirical claim.
  5. [Figure 1] The figure is not described in the text beyond a parenthetical about the dotted line; the lifecycle stages in the figure should be enumerated in the prose to make the figure self-contained.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the three-scope taxonomy is a conceptual proposal, not derived from or reducible to its illustrative self-citations.

full rationale

This paper is a conceptual position paper rather than a derivation or fitting exercise. It proposes three scopes of alignment (competence, transience, audience), grounding each in external literatures: knowledge/skills/behaviors from the Army KSB framework, episodic/semantic from neuroscience, and audience categories from communication theory. The central claim that alignment should be scoped differently for different groups and contexts is argued from moral psychology and dyadic morality theory, not derived from the scopes themselves. Section 5's technical implications (knowledge via question-answer data, skills via examples, behaviors via preferences; semantic alignment implying full fine-tuning, episodic alignment implying inference-time adapters; audience determining bidirectional versus unidirectional procedures) are asserted as practical heuristics, not equations fitted to data. The Related Work section cites several works with overlapping authorship, including Alignment Studio, Value Alignment from Unstructured Text, Larimar, Contextual Value Alignment, and Decolonial AI Alignment, to illustrate that alignment beyond mass semantic behavioral alignment is not empty space. These citations are illustrative rather than load-bearing: none is invoked as a uniqueness theorem, as a derivation of the taxonomy, or as a fitted parameter. Even if these self-citations were removed, the taxonomy and its implications would stand or fall on the paper's own conceptual arguments. The paper explicitly acknowledges that it 'does not detail choosing the right scope of alignment, but such a process is clearly necessary,' which is an honest limitation about assignment rules and usability, not a circular step. No equation reduces to its own input, no fitted parameter is renamed as a prediction, and no self-citation is used to forbid alternative frameworks. Consequently, there is no significant circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

This paper introduces a conceptual taxonomy, so the ledger contains no fitted numbers or invented entities. The relevant assumptions are the choice of the three axes, the borrowed neuroscience and communication-theory dichotomies, and the moral-psychology framing that motivates patient-centered scopes.

assumptions (5)
  • domain assumption Alignment is an 'empty signifier' whose meaning is determined by practice rather than fixed content.
    Invoked in Section 1 via Kirk et al. 2023 to justify redefining alignment in terms of scopes.
  • ad hoc to paper The three proposed scopes are a useful and sufficiently complete decomposition of alignment.
    Section 3 introduces competence, transience, and audience without proof of completeness, mutual exclusivity, or orthogonality.
  • ad hoc to paper The episodic or semantic distinction from neuroscience applies to AI knowledge, skills, and behaviors.
    Section 3 borrows the neuroscience terminology and assumes it maps directly to alignment transience.
  • ad hoc to paper Audience categories from communication theory, dyadic, small-group, public, and mass, correspond to distinct alignment mechanisms.
    Section 3 and Figure 2 draw the analogy to communication theory; no empirical support is provided for this mapping.
  • domain assumption Dyadic morality theory's agent-patient harm model is a valid lens for understanding moral differences relevant to alignment.
    Section 2 uses Schein and Gray 2018 to motivate patient-centered scoped alignment.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Scopes of Alignment." pith.science (2026). https://pith.science/paper/4SSOSS54

@misc{pith2026250112405,
  author       = {Pith},
  title        = {Pith review of: Scopes of Alignment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4SSOSS54}},
  note         = {Machine review of arXiv:2501.12405}
}
read the original abstract

Much of the research focus on AI alignment seeks to align large language models and other foundation models to the context-less and generic values of helpfulness, harmlessness, and honesty. Frontier model providers also strive to align their models with these values. In this paper, we motivate why we need to move beyond such a limited conception and propose three dimensions for doing so. The first scope of alignment is competence: knowledge, skills, or behaviors the model must possess to be useful for its intended purpose. The second scope of alignment is transience: either semantic or episodic depending on the context of use. The third scope of alignment is audience: either mass, public, small-group, or dyadic. At the end of the paper, we use the proposed framework to position some technologies and workflows that go beyond prevailing notions of alignment.

Figures

Figures reproduced from arXiv: 2501.12405 by the authors.

Figure 1
Figure 1. Typical application development lifecycle for LLMs considered by frontier model providers. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Different audiences of alignment. This example illustrates why a flexible, multi-scope ap￾proach is essential for AI alignment, as country-specific social and cultural norms shape how AI should interact, evolve, and serve its intended audience. 5 Implications What is the use of defining these three scopes of alignment? The first use is simply to be able to more precisely discern that alignment is not a singular proc… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions

    cs.CY 2026-02 unverdicted novelty 4.0 of 10

    Current XAI methods for DNNs and LLMs rest on paradoxes and false assumptions that demand a paradigm shift to verification protocols, scientific foundations, context-aware design, and faithful model analysis rather th...

Reference graph

Works this paper leans on

37 extracted references · 23 canonical work pages · cited by 1 Pith paper

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence 'output.state := if if FUNCTION not #0 #1 if FUNCTION and 'skip pop #0 if FUNCTIO...

  2. [2]

    E.; Bhattacharjee, B.; Bouneffouf, D.; Chaudhury, S.; Chen, P.-Y.; Chiazor, L.; Daly, E

    Achintalwar, S.; Alvarado Garcia, A.; Anaby-Tavor, A.; Baldini, I.; Berger, S. E.; Bhattacharjee, B.; Bouneffouf, D.; Chaudhury, S.; Chen, P.-Y.; Chiazor, L.; Daly, E. M.; de Paula, R. A.; Dognin, P.; Farchi, E.; Ghosh, S.; Hind, M.; Horesh, R.; Kour, G.; Lee, J. Y.; Miehling, E.; Murugesan, K.; Nagireddy, M.; Padhi, I.; Piorkowski, D.; Rawat, A.; Raz, O....

  3. [3]

    A.; and Varshney, K

    Achintalwar, S.; Baldini, I.; Bouneffouf, D.; Byamugisha, J.; Chang, M.; Dognin, P.; Farchi, E.; Makondo, N.; Mojsilovi \'c , A.; Nagireddy, M.; Natesan Ramamurthy, K.; Padhi, I.; Raz, O.; Rios, J.; Sattigeri, P.; Singh, M.; Thwala, S.; Uceda-Sosa, R. A.; and Varshney, K. R. 2024b. Alignment studio: Aligning large language models to particular contextual ...

  4. [4]

    Anthropic. 2024. System prompts. https://docs.anthropic.com/en/release-notes/system-prompts\#july-12th-2024

  5. [5]

    R.; Hatfield-Dodds, Z.; Mann, B.; Amodei, D.; Joseph, N.; McCandlish, S.; Brown, T.; and Kaplan, J

    Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; Chen, C.; Olsson, C.; Olah, C.; Hernandez, D.; Drain, D.; Ganguli, D.; Li, D.; Tran-Johnson, E.; Perez, E.; Kerr, J.; Mueller, J.; Ladish, J.; Landau, J.; Ndousse, K.; Lukosuite, K.; Lovitt, L.; Sellitto, M.; Elhage, N.; Schiefer, N.; ...

  6. [6]

    A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M

    Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M. S.; Bohg, J.; Bosselut, A.; Brunskill, E.; Brynjolfsson, E.; Buch, S.; Card, D.; Castellon, R.; Chatterji, N.; Chen, A.; Creel, K.; Davis, J. Q.; Demszky, D.; Donahue, C.; Doumbouya, M.; Durmus, E.; Ermon, S.; Etchemendy, J.; Ethayarajh, K.; Li, F.-F.; Finn, C.; Gal...

  7. [7]

    Das, P.; Chaudhury, S.; Nelson, E.; Melnyk, I.; Swaminathan, S.; Dai, S.; Lozano, A.; Kollias, G.; Chenthamarakshan, V.; Dan, S.; and Chen, P.-Y. 2024. Larimar: Large language models with episodic memory control. In Proceedings of the International Conference on Machine Learning

  8. [8]

    D.; Nagireddy, M.; Varshney, K

    Dognin, P.; Rios, J.; Luss, R.; Sattigeri, P.; Liu, M.; Padhi, I.; Riemer, M. D.; Nagireddy, M.; Varshney, K. R.; and Bouneffouf, D. 2025. Contextual value alignment. In IEEE International Conference on Acoustics, Speech, and Signal Processing

Show all 37 references
  1. [9]

    Durmus, E.; Nyugen, K.; Liao, T. I.; Schiefer, N.; Askell, A.; Bakhtin, A.; Chen, C.; Hatfield-Dodds, Z.; Hernandez, D.; Joseph, N.; Lovitt, L.; McCandlish, S.; Sikder, O.; Tamkin, A.; Thamkul, J.; Kaplan, J.; Clark, J.; and Ganguli, D. 2023. Towards measuring the representati...

  2. [10]

    Y.; Liu, Y.; and Tsvetkov, Y

    Feng, S.; Park, C. Y.; Liu, Y.; and Tsvetkov, Y. 2023. From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models. arXiv:2305.08283

  3. [11]

    Gilbert, E. 2024. HCC is all you need: Alignment---the sensible kind anyway---is just human-centered computing. arXiv:2405.03699

  4. [12]

    P.; and Ditto, P

    Graham, J.; Haidt, J.; Koleva, S.; Motyl, M.; Iyer, R.; Wojcik, S. P.; and Ditto, P. H. 2013. Moral foundations theory: The pragmatic validity of moral pluralism. Advances in Experimental Social Psychology 47:55--130

  5. [13]

    Haemmerl, K.; Deiseroth, B.; Schramowski, P.; Libovick \'y , J.; Rothkopf, C.; Fraser, A.; and Kersting, K. 2023. Speaking multiple languages affects the moral bias of language models. In Findings of the Association for Computational Linguistics , 2137--2156

  6. [14]

    A., and Fleming, A

    Homer, B. A., and Fleming, A. C. 2023. Soldier perspectives on the knowledge, skills, and behaviors of a good teammate. Technical Report Research Note 2023-07, Army Research Institute

  7. [15]

    Hu, E.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022. LoRA : Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations

  8. [16]

    Y.; Dai, J.; Pan, X.; O'Gara, A.; Lei, Y.; Xu, H.; Tse, B.; Fu, J.; McAleer, S.; Yang, Y.; Wang, Y.; Zhu, S.-C.; Guo, Y.; and Gao, W

    Ji, J.; Qiu, T.; Chen, B.; Zhang, B.; Lou, H.; Wang, K.; Duan, Y.; He, Z.; Zhou, J.; Zhang, Z.; Zeng, F.; Ng, K. Y.; Dai, J.; Pan, X.; O'Gara, A.; Lei, Y.; Xu, H.; Tse, B.; Fu, J.; McAleer, S.; Yang, Y.; Wang, Y.; Zhu, S.-C.; Guo, Y.; and Gao, W. 2023. AI alignment: A comprehe...

  9. [17]

    R.; Vidgen, B.; R \"o ttger, P.; and Hale, S

    Kirk, H. R.; Vidgen, B.; R \"o ttger, P.; and Hale, S. A. 2023. The empty signifier problem: Towards clearer paradigms for operationalising ``alignment''' in large language models. arXiv:2310.02457

  10. [18]

    R.; Vidgen, B.; R \"o ttger, P.; and Hale, S

    Kirk, H. R.; Vidgen, B.; R \"o ttger, P.; and Hale, S. A. 2024. The benefits, risks and bounds of personalizing the alignment of large language models to individuals. Nature Machine Intelligence 6:83--392

  11. [19]

    Li, J.; Consul, S.; Zhou, E.; Wong, J.; Farooqui, N.; Ye, Y.; Manohar, N.; Wei, Z.; Wu, T.; Echols, B.; Zhou, S.; and Diamos, G. 2024. Banishing LLM hallucinations requires rethinking generalization. arXiv:2406.17642

  12. [20]

    Lipton, Z. 2024. Alignment is now defined so broadly. https://x.com/zacharylipton/status/1771177444088685045

  13. [21]

    M \"o ller, N. 2016. Value uncertainty. In Hansson, S. O., and Hadorn, G. H., eds., The Argumentative Turn in Policy Analysis: Reasoning About Uncertainty . Switzerland: Springer. 105--133

  14. [22]

    Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022. Training language model...

  15. [23]

    Padhi, I.; Natesan Ramamurthy, K.; Sattigeri, P.; Nagireddy, M.; Dognin, P.; and Varshney, K. R. 2024. Value alignment from unstructured text. In Proceedings of the Conference on Empirical Methods in Natural Language Processing , 1083--1095

  16. [24]

    Qi, X.; Zeng, Y.; Xie, T.; Chen, P.-Y.; Jia, R.; Mittal, P.; and Henderson, P. 2023. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv:2310.03693

  17. [25]

    Schein, C., and Gray, K. 2018. The theory of dyadic morality: Reinventing moral judgment by redefining harm. Personality and Social Psychology Review 22(1):32--70

  18. [26]

    Scherrer, N.; Shi, C.; Feder, A.; and Blei, D. 2024. Evaluating the moral beliefs encoded in LLMs . Advances in Neural Information Processing Systems 36

  19. [27]

    Shelby, R.; Rismani, S.; Henne, K.; Moon, A.; Rostamzadeh, N.; Nicholas, P.; Yilla, N.; Gallegos, J.; Smart, A.; Garcia, E.; and Virk, G. 2023. Sociotechnical harms: Scoping a taxonomy for harm reduction. In Proceedings of the AAAI/ACM Conference on Artificial Intelligence, Et...

  20. [28]

    Shen, T.; Jin, R.; Huang, Y.; Liu, C.; Dong, W.; Guo, Z.; Wu, X.; Liu, Y.; and Xiong, D. 2023. Large language model alignment: A survey. arXiv:2309.15025

  21. [29]

    Shneiderman, B., and Muller, M. 2023. On AI anthropomorphism. https://medium.com/human-centered-ai/on-ai-anthropomorphism-abff4cecc5ae

  22. [30]

    M.; Ye, A.; Jiang, L.; Lu, X.; Dziri, N.; Althoff, T.; and Choi, Y

    Sorensen, T.; Moore, J.; Fisher, J.; Gordon, M.; Mireshghallah, N.; Rytting, C. M.; Ye, A.; Jiang, L.; Lu, X.; Dziri, N.; Althoff, T.; and Choi, Y. 2024. A roadmap to pluralistic alignment. In Proceedings of the International Conference on Machine Learning

  23. [31]

    D.; and Srivastava, A

    Sudalairaj, S.; Bhandwaldar, A.; Pareja, A.; Xu, K.; Cox, D. D.; and Srivastava, A. 2024. LAB : Large-scale alignment for chatbots. arXiv:2403.01081

  24. [32]

    Z.; Yu, P

    Sun, L.; Huang, Y.; Wang, H.; Wu, S.; Zhang, Q.; Gao, C.; Huang, Y.; Lyu, W.; Zhang, Y.; Li, X.; Liu, Z.; Liu, Y.; Wang, Y.; Zhang, Z.; Vidgen, B.; Kailkhura, B.; Xiong, C.; Xiao, C.; Li, C.; Xing, E.; Huang, F.; Liu, H.; Ji, H.; Wang, H.; Zhang, H.; Yao, H.; Kellis, M.; Zitni...

  25. [33]

    Varshney, K. R. 2024. Decolonial AI alignment: Openness, vi\' s e s a-dharma, and including excluded knowledges. In Proceedings of the AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society , 1467--1481

  26. [34]

    Wang, Q., and Goel, A. K. 2022. Mutual theory of mind for human- AI communication. arXiv:2210.03842

  27. [35]

    D.; and Goel, A

    Wang, Q.; Walsh, S.; Si, M.; Kephart, J.; Weisz, J. D.; and Goel, A. K. 2024. Theory of mind in human- AI interaction. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems , 493

  28. [36]

    A.; Rimell, L.; Isaac, W.; Haas, J.; Legassick, S.; Irving, G.; and Gabriel, I

    Weidinger, L.; Uesato, J.; Rauh, M.; Griffin, C.; Huang, P.-S.; Mellor, J.; Glaese, A.; Cheng, M.; Balle, B.; Kasirzadeh, A.; Biles, C.; Brown, S.; Kenton, Z.; Hawkins, W.; Stepleton, T.; Birhane, A.; Hendricks, L. A.; Rimell, L.; Isaac, W.; Haas, J.; Legassick, S.; Irving, G....

  29. [37]

    Zhi-Xuan, T.; Carroll, M.; Franklin, M.; and Ashton, H. 2024. Beyond preferences in AI alignment. Philosophical Studies

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.