REVIEW 4 major objections 5 minor 1 cited by
Scopes of Alignment
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Alignment is not one workflow: the paper proposes three scopes — competence, transience, and audience — along which AI alignment should be designed.
desk verdict A clear, readable position paper that gives alignment a useful three-scope vocabulary, but the axes are asserted rather than derived and the categories are under-specified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is the three-axis scope framework: competence by knowledge, skills, or behaviors; transience by episodic or semantic; and audience by mass, public, small-group, or dyadic. The framework works as a classification space for alignment goals and methods, letting the authors position existing techniques as particular cells — for example, reward-based preference learning as suitable only for local uses, low-rank adapters as a path to episodic alignment, and mutual theory of mind as a mechanism for dyadic alignment. It carries the argument by turning the vague term 'alignment' into a design space in which each axis is supposed to point to different technical choices.
What would settle it
Find two deployment contexts that differ on exactly one scope — say audience, dyadic versus small-group, with competence and transience held fixed — and measure whether the optimal alignment method or outcome differs meaningfully; if no systematic difference appears across several such pairs, the scopes do not carve alignment at its joints. A complementary test would look for a context that differs on all three scopes yet still requires the same alignment workflow, showing the axes are not the load-bearing choices.
Extended reading notes
Core claim
The paper's central claim is that 'alignment should not be just one activity, technical approach, or workflow,' but should be 'more specific and scoped for the different needs of different groups over time.' It proposes three dimensions along which an alignment effort can be placed: competence, which covers the knowledge, skills, or behaviors the model must have; transience, which distinguishes episodic alignment bound to a particular time, place, or context from semantic alignment about general facts of the world; and audience, which runs from a single user in a dyadic relationship, through small groups and publics, to mass alignment with all of humanity. The claim is that most current research and frontier-model practice occupies one corner of this space — mass semantic alignment for behaviors such as helpfulness, harmlessness, and honesty — and that this narrowness hides the technical requirements of other scopes. If the claim is right, alignment is a family of workflows with different data sources, learning algorithms, and inference-time mechanisms, not one method applied to one target.
Load-bearing premise
The three scopes are assumed to vary independently and to cover the meaningful decisions in alignment, but the paper's own example changes all three at once when the country changes, so if they co-vary, the proposed design space claims more degrees of freedom than exist.
Editorial extensions
If this is right
- Researchers would stop treating mass semantic behavioral alignment as the default and would report which scope their method targets.
- Alignment technologies would diverge by competence: knowledge alignment favors question-answer generation from content, skills favor repeated examples, and behaviors favor preference data with scenarios.
- Episodic alignment would be done at inference time, for example by swapping low-rank adapters on the fly, rather than by permanently fine-tuning the model.
- Dyadic and small-group alignment would be bidirectional, with both human and model adapting, unlike the unidirectional alignment used for mass audiences.
- Properly scoped alignment could resolve many value conflicts before pluralistic mediation is needed, because narrowing the scope can remove conflicting values.
Reading between the lines
- The three axes are proposed as independent but the paper's motivating example varies all three together when moving from one country to another; an editorial reading is that the axes may co-vary, so the true design space may be smaller than the 3-by-2-by-4 grid suggests.
- A natural extension the paper leaves implicit is an empirical test: hold two scopes constant, vary the third, and compare which data, optimization, or inference-time method succeeds, turning the taxonomy into a predictive theory.
- If scoping succeeds, evaluation and red-teaming may also need to be scoped, since a model aligned for a dyadic therapy relationship should not be judged by a mass-semantic benchmark.
- Another extension is to treat scope selection itself as a governed, human-centered design decision, because the paper notes scopes should be found through reflective equilibrium but does not prescribe a process for doing so.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a position statement arguing that AI alignment should not be reduced to a single generic activity. It proposes three "scopes" of alignment: competence (knowledge, skills, or behaviors), transience (episodic versus semantic), and audience (mass, public, small-group, or dyadic). The authors claim that prevailing practice, mass semantic alignment for behaviors, is only one point in this larger space, and they use a hypothetical mental-health counseling example in the U.S. versus China to illustrate how the scopes could differ by context. Section 5 argues that the scopes have technical implications, such as full fine-tuning versus inference-time adapters and unidirectional versus bidirectional alignment procedures, and suggests that reflective equilibrium could be used to choose a scope and thereby dissolve value conflicts. The paper is explicitly non-empirical and positions itself as a precursor to pluralistic alignment.
Significance. If the proposed taxonomy is accepted, it offers a useful descriptive vocabulary for discussing alignment targets and draws attention to alignment settings beyond the dominant one. The paper is clear in its motivation and points to a non-empty space, citing existing work on personalized alignment, contextual alignment, episodic memory, and knowledge-oriented alignment. However, its central contribution is a conceptual scheme whose utility depends on the categories being clearly applicable and reasonably complete. The paper does not provide a formal derivation, but for a position paper that is not itself disqualifying; the honesty of the writing and the modest scope of the claims are strengths. The three-axis taxonomy could serve as a starting point for more operational work, but as it stands the framework's applicability is not yet demonstrated.
major comments (4)
- [Section 3, competence] The three competence categories (knowledge, skills, behaviors) are not defined with mutually exclusive boundaries. The paper calls 'converting texts into hip-hop raps' a skill and 'politeness' and 'verbosity' behaviors, but both are action-level properties of model output. If the distinction between skill and behavior is not operational, then Section 5's claim that different competences require different data and optimization strategies cannot be grounded. A usable taxonomy needs either an explicit assignment rule or a statement that the categories are overlapping perspectives rather than disjoint dimensions.
- [Section 3, transience] The episodic/semantic distinction is introduced with examples but no operational criterion for deciding whether a given knowledge, skill, or behavior is 'bound to a time, place, or other context' or 'general about the world.' For instance, 'obtaining manager pre-approval before booking air tickets' is described as episodic, yet in another organization it could be a standing rule. Without a test or at least a clear decision procedure, two annotators could plausibly classify the same alignment target into different transience categories, which undermines the claimed design space.
- [Section 4 and Section 3] The motivating example changes all three axes simultaneously when moving from the U.S. to China: competence shifts from individualized to collectivist, transience from episodic to semantic, and audience from individual to community. This co-variation suggests that the three axes may not be independent as the design-space framing implies, and it does not provide evidence for a 3-by-2-by-4 space. The paper should either give an example where one axis varies while the other two are fixed, or soften the claim that the three scopes are independent dimensions.
- [Section 5, reflective equilibrium] The claim that 'we will also implicitly deal with the problem of conflicting values because the reflective equilibrium will have resolved conflicts by narrowing the scope until there are few remaining' is asserted without a method or justification. Narrowing the scope can avoid a conflict, but that is different from resolving it, and no procedure is given for finding a scope with minimal internal conflict. The paper itself acknowledges that 'such a process is clearly necessary,' so this section currently promises more than it delivers.
minor comments (5)
- [Section 5] The reference style is inconsistent: the text says 'As Tan Zhi-Xuan et al. submit,' but the bibliography entry is 'Zhi-Xuan et al. 2024.' Please standardize.
- [Section 3] The transliteration 's¯adh¯aran.a-dharma and vi´ses.a-dharma' contains formatting artifacts; it should read 'sādhāraṇa dharma' and 'viśeṣa dharma'.
- [Section 5] The sentence beginning 'It is important to differentiate the transience of alignment because while semantic alignment implies ' is missing a comma after 'because' and is long enough to obscure the conditional structure; consider splitting it.
- [Section 4] The mental-health example is labeled hypothetical, but the paper does not qualify that the characterization of U.S. and Chinese cultural norms is a simplification; adding a caveat would prevent the example from being read as an empirical claim.
- [Figure 1] The figure is not described in the text beyond a parenthetical about the dotted line; the lifecycle stages in the figure should be enumerated in the prose to make the figure self-contained.
Circularity Check
No significant circularity: the three-scope taxonomy is a conceptual proposal, not derived from or reducible to its illustrative self-citations.
full rationale
This paper is a conceptual position paper rather than a derivation or fitting exercise. It proposes three scopes of alignment (competence, transience, audience), grounding each in external literatures: knowledge/skills/behaviors from the Army KSB framework, episodic/semantic from neuroscience, and audience categories from communication theory. The central claim that alignment should be scoped differently for different groups and contexts is argued from moral psychology and dyadic morality theory, not derived from the scopes themselves. Section 5's technical implications (knowledge via question-answer data, skills via examples, behaviors via preferences; semantic alignment implying full fine-tuning, episodic alignment implying inference-time adapters; audience determining bidirectional versus unidirectional procedures) are asserted as practical heuristics, not equations fitted to data. The Related Work section cites several works with overlapping authorship, including Alignment Studio, Value Alignment from Unstructured Text, Larimar, Contextual Value Alignment, and Decolonial AI Alignment, to illustrate that alignment beyond mass semantic behavioral alignment is not empty space. These citations are illustrative rather than load-bearing: none is invoked as a uniqueness theorem, as a derivation of the taxonomy, or as a fitted parameter. Even if these self-citations were removed, the taxonomy and its implications would stand or fall on the paper's own conceptual arguments. The paper explicitly acknowledges that it 'does not detail choosing the right scope of alignment, but such a process is clearly necessary,' which is an honest limitation about assignment rules and usability, not a circular step. No equation reduces to its own input, no fitted parameter is renamed as a prediction, and no self-citation is used to forbid alternative frameworks. Consequently, there is no significant circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption Alignment is an 'empty signifier' whose meaning is determined by practice rather than fixed content.
- ad hoc to paper The three proposed scopes are a useful and sufficiently complete decomposition of alignment.
- ad hoc to paper The episodic or semantic distinction from neuroscience applies to AI knowledge, skills, and behaviors.
- ad hoc to paper Audience categories from communication theory, dyadic, small-group, public, and mass, correspond to distinct alignment mechanisms.
- domain assumption Dyadic morality theory's agent-patient harm model is a valid lens for understanding moral differences relevant to alignment.
Cite this review
Pith. "Pith review of Scopes of Alignment." pith.science (2026). https://pith.science/paper/4SSOSS54
@misc{pith2026250112405,
author = {Pith},
title = {Pith review of: Scopes of Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/4SSOSS54}},
note = {Machine review of arXiv:2501.12405}
}
read the original abstract
Much of the research focus on AI alignment seeks to align large language models and other foundation models to the context-less and generic values of helpfulness, harmlessness, and honesty. Frontier model providers also strive to align their models with these values. In this paper, we motivate why we need to move beyond such a limited conception and propose three dimensions for doing so. The first scope of alignment is competence: knowledge, skills, or behaviors the model must possess to be useful for its intended purpose. The second scope of alignment is transience: either semantic or episodic depending on the context of use. The third scope of alignment is audience: either mass, public, small-group, or dyadic. At the end of the paper, we use the proposed framework to position some technologies and workflows that go beyond prevailing notions of alignment.
Figures
Forward citations
Cited by 1 Pith paper
-
Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions
Current XAI methods for DNNs and LLMs rest on paradoxes and false assumptions that demand a paradigm shift to verification protocols, scientific foundations, context-aware design, and faithful model analysis rather th...
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence 'output.state := if if FUNCTION not #0 #1 if FUNCTION and 'skip pop #0 if FUNCTIO...
-
[2]
E.; Bhattacharjee, B.; Bouneffouf, D.; Chaudhury, S.; Chen, P.-Y.; Chiazor, L.; Daly, E
Achintalwar, S.; Alvarado Garcia, A.; Anaby-Tavor, A.; Baldini, I.; Berger, S. E.; Bhattacharjee, B.; Bouneffouf, D.; Chaudhury, S.; Chen, P.-Y.; Chiazor, L.; Daly, E. M.; de Paula, R. A.; Dognin, P.; Farchi, E.; Ghosh, S.; Hind, M.; Horesh, R.; Kour, G.; Lee, J. Y.; Miehling, E.; Murugesan, K.; Nagireddy, M.; Padhi, I.; Piorkowski, D.; Rawat, A.; Raz, O....
-
[3]
Achintalwar, S.; Baldini, I.; Bouneffouf, D.; Byamugisha, J.; Chang, M.; Dognin, P.; Farchi, E.; Makondo, N.; Mojsilovi \'c , A.; Nagireddy, M.; Natesan Ramamurthy, K.; Padhi, I.; Raz, O.; Rios, J.; Sattigeri, P.; Singh, M.; Thwala, S.; Uceda-Sosa, R. A.; and Varshney, K. R. 2024b. Alignment studio: Aligning large language models to particular contextual ...
-
[4]
Anthropic. 2024. System prompts. https://docs.anthropic.com/en/release-notes/system-prompts\#july-12th-2024
work page 2024
-
[5]
R.; Hatfield-Dodds, Z.; Mann, B.; Amodei, D.; Joseph, N.; McCandlish, S.; Brown, T.; and Kaplan, J
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; Chen, C.; Olsson, C.; Olah, C.; Hernandez, D.; Drain, D.; Ganguli, D.; Li, D.; Tran-Johnson, E.; Perez, E.; Kerr, J.; Mueller, J.; Ladish, J.; Landau, J.; Ndousse, K.; Lukosuite, K.; Lovitt, L.; Sellitto, M.; Elhage, N.; Schiefer, N.; ...
arXiv 2022
-
[6]
A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M
Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M. S.; Bohg, J.; Bosselut, A.; Brunskill, E.; Brynjolfsson, E.; Buch, S.; Card, D.; Castellon, R.; Chatterji, N.; Chen, A.; Creel, K.; Davis, J. Q.; Demszky, D.; Donahue, C.; Doumbouya, M.; Durmus, E.; Ermon, S.; Etchemendy, J.; Ethayarajh, K.; Li, F.-F.; Finn, C.; Gal...
arXiv 2021
-
[7]
Das, P.; Chaudhury, S.; Nelson, E.; Melnyk, I.; Swaminathan, S.; Dai, S.; Lozano, A.; Kollias, G.; Chenthamarakshan, V.; Dan, S.; and Chen, P.-Y. 2024. Larimar: Large language models with episodic memory control. In Proceedings of the International Conference on Machine Learning
work page 2024
-
[8]
D.; Nagireddy, M.; Varshney, K
Dognin, P.; Rios, J.; Luss, R.; Sattigeri, P.; Liu, M.; Padhi, I.; Riemer, M. D.; Nagireddy, M.; Varshney, K. R.; and Bouneffouf, D. 2025. Contextual value alignment. In IEEE International Conference on Acoustics, Speech, and Signal Processing
work page 2025
Show all 37 references
-
[9]
Durmus, E.; Nyugen, K.; Liao, T. I.; Schiefer, N.; Askell, A.; Bakhtin, A.; Chen, C.; Hatfield-Dodds, Z.; Hernandez, D.; Joseph, N.; Lovitt, L.; McCandlish, S.; Sikder, O.; Tamkin, A.; Thamkul, J.; Kaplan, J.; Clark, J.; and Ganguli, D. 2023. Towards measuring the representati...
2023 arXiv
-
[10]
Y.; Liu, Y.; and Tsvetkov, Y
Feng, S.; Park, C. Y.; Liu, Y.; and Tsvetkov, Y. 2023. From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair NLP models. arXiv:2305.08283
2023 arXiv
-
[11]
Gilbert, E. 2024. HCC is all you need: Alignment---the sensible kind anyway---is just human-centered computing. arXiv:2405.03699
2024 arXiv
-
[12]
P.; and Ditto, P
Graham, J.; Haidt, J.; Koleva, S.; Motyl, M.; Iyer, R.; Wojcik, S. P.; and Ditto, P. H. 2013. Moral foundations theory: The pragmatic validity of moral pluralism. Advances in Experimental Social Psychology 47:55--130
2013
-
[13]
Haemmerl, K.; Deiseroth, B.; Schramowski, P.; Libovick \'y , J.; Rothkopf, C.; Fraser, A.; and Kersting, K. 2023. Speaking multiple languages affects the moral bias of language models. In Findings of the Association for Computational Linguistics , 2137--2156
2023
-
[14]
A., and Fleming, A
Homer, B. A., and Fleming, A. C. 2023. Soldier perspectives on the knowledge, skills, and behaviors of a good teammate. Technical Report Research Note 2023-07, Army Research Institute
2023
-
[15]
Hu, E.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022. LoRA : Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations
2022
-
[16]
Y.; Dai, J.; Pan, X.; O'Gara, A.; Lei, Y.; Xu, H.; Tse, B.; Fu, J.; McAleer, S.; Yang, Y.; Wang, Y.; Zhu, S.-C.; Guo, Y.; and Gao, W
Ji, J.; Qiu, T.; Chen, B.; Zhang, B.; Lou, H.; Wang, K.; Duan, Y.; He, Z.; Zhou, J.; Zhang, Z.; Zeng, F.; Ng, K. Y.; Dai, J.; Pan, X.; O'Gara, A.; Lei, Y.; Xu, H.; Tse, B.; Fu, J.; McAleer, S.; Yang, Y.; Wang, Y.; Zhu, S.-C.; Guo, Y.; and Gao, W. 2023. AI alignment: A comprehe...
2023 arXiv
-
[17]
R.; Vidgen, B.; R \"o ttger, P.; and Hale, S
Kirk, H. R.; Vidgen, B.; R \"o ttger, P.; and Hale, S. A. 2023. The empty signifier problem: Towards clearer paradigms for operationalising ``alignment''' in large language models. arXiv:2310.02457
2023 arXiv
-
[18]
R.; Vidgen, B.; R \"o ttger, P.; and Hale, S
Kirk, H. R.; Vidgen, B.; R \"o ttger, P.; and Hale, S. A. 2024. The benefits, risks and bounds of personalizing the alignment of large language models to individuals. Nature Machine Intelligence 6:83--392
2024
-
[19]
Li, J.; Consul, S.; Zhou, E.; Wong, J.; Farooqui, N.; Ye, Y.; Manohar, N.; Wei, Z.; Wu, T.; Echols, B.; Zhou, S.; and Diamos, G. 2024. Banishing LLM hallucinations requires rethinking generalization. arXiv:2406.17642
2024 arXiv
-
[20]
Lipton, Z. 2024. Alignment is now defined so broadly. https://x.com/zacharylipton/status/1771177444088685045
2024
-
[21]
M \"o ller, N. 2016. Value uncertainty. In Hansson, S. O., and Hadorn, G. H., eds., The Argumentative Turn in Policy Analysis: Reasoning About Uncertainty . Switzerland: Springer. 105--133
2016
-
[22]
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022. Training language model...
2022
-
[23]
Padhi, I.; Natesan Ramamurthy, K.; Sattigeri, P.; Nagireddy, M.; Dognin, P.; and Varshney, K. R. 2024. Value alignment from unstructured text. In Proceedings of the Conference on Empirical Methods in Natural Language Processing , 1083--1095
2024
-
[24]
Qi, X.; Zeng, Y.; Xie, T.; Chen, P.-Y.; Jia, R.; Mittal, P.; and Henderson, P. 2023. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv:2310.03693
2023 arXiv
-
[25]
Schein, C., and Gray, K. 2018. The theory of dyadic morality: Reinventing moral judgment by redefining harm. Personality and Social Psychology Review 22(1):32--70
2018
-
[26]
Scherrer, N.; Shi, C.; Feder, A.; and Blei, D. 2024. Evaluating the moral beliefs encoded in LLMs . Advances in Neural Information Processing Systems 36
2024
-
[27]
Shelby, R.; Rismani, S.; Henne, K.; Moon, A.; Rostamzadeh, N.; Nicholas, P.; Yilla, N.; Gallegos, J.; Smart, A.; Garcia, E.; and Virk, G. 2023. Sociotechnical harms: Scoping a taxonomy for harm reduction. In Proceedings of the AAAI/ACM Conference on Artificial Intelligence, Et...
2023
-
[28]
Shen, T.; Jin, R.; Huang, Y.; Liu, C.; Dong, W.; Guo, Z.; Wu, X.; Liu, Y.; and Xiong, D. 2023. Large language model alignment: A survey. arXiv:2309.15025
2023 arXiv
-
[29]
Shneiderman, B., and Muller, M. 2023. On AI anthropomorphism. https://medium.com/human-centered-ai/on-ai-anthropomorphism-abff4cecc5ae
2023
-
[30]
M.; Ye, A.; Jiang, L.; Lu, X.; Dziri, N.; Althoff, T.; and Choi, Y
Sorensen, T.; Moore, J.; Fisher, J.; Gordon, M.; Mireshghallah, N.; Rytting, C. M.; Ye, A.; Jiang, L.; Lu, X.; Dziri, N.; Althoff, T.; and Choi, Y. 2024. A roadmap to pluralistic alignment. In Proceedings of the International Conference on Machine Learning
2024
-
[31]
D.; and Srivastava, A
Sudalairaj, S.; Bhandwaldar, A.; Pareja, A.; Xu, K.; Cox, D. D.; and Srivastava, A. 2024. LAB : Large-scale alignment for chatbots. arXiv:2403.01081
2024 arXiv
-
[32]
Z.; Yu, P
Sun, L.; Huang, Y.; Wang, H.; Wu, S.; Zhang, Q.; Gao, C.; Huang, Y.; Lyu, W.; Zhang, Y.; Li, X.; Liu, Z.; Liu, Y.; Wang, Y.; Zhang, Z.; Vidgen, B.; Kailkhura, B.; Xiong, C.; Xiao, C.; Li, C.; Xing, E.; Huang, F.; Liu, H.; Ji, H.; Wang, H.; Zhang, H.; Yao, H.; Kellis, M.; Zitni...
2024 arXiv
-
[33]
Varshney, K. R. 2024. Decolonial AI alignment: Openness, vi\' s e s a-dharma, and including excluded knowledges. In Proceedings of the AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society , 1467--1481
2024
-
[34]
Wang, Q., and Goel, A. K. 2022. Mutual theory of mind for human- AI communication. arXiv:2210.03842
2022 arXiv
-
[35]
D.; and Goel, A
Wang, Q.; Walsh, S.; Si, M.; Kephart, J.; Weisz, J. D.; and Goel, A. K. 2024. Theory of mind in human- AI interaction. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems , 493
2024
-
[36]
A.; Rimell, L.; Isaac, W.; Haas, J.; Legassick, S.; Irving, G.; and Gabriel, I
Weidinger, L.; Uesato, J.; Rauh, M.; Griffin, C.; Huang, P.-S.; Mellor, J.; Glaese, A.; Cheng, M.; Balle, B.; Kasirzadeh, A.; Biles, C.; Brown, S.; Kenton, Z.; Hawkins, W.; Stepleton, T.; Birhane, A.; Hendricks, L. A.; Rimell, L.; Isaac, W.; Haas, J.; Legassick, S.; Irving, G....
2022
-
[37]
Zhi-Xuan, T.; Carroll, M.; Franklin, M.; and Ashton, H. 2024. Beyond preferences in AI alignment. Philosophical Studies
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.