REVIEW 3 major objections 3 minor 277 references
This paper argues that pluralistic alignment research has not yet shaped deployed AI systems and lays out three research directions to make adoption happen.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 05:09 UTC pith:TF2D7JNH
load-bearing objection Useful audit and a sensible agenda, but the abstract's 'no public evidence it shaped training' overstates what the paper itself shows. the 3 major comments →
A Roadmap to Impactful Pluralistic Alignment Research
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is an absence: as of mid-2026, there is no public evidence that pluralistic value alignment shapes any deployed frontier model. Four of five audited labs prescribe some multi-perspective behavior in a constitution, model spec, or system prompt, but no lab's model or system card evaluates whether the model follows those provisions; the only dedicated pluralism evaluation found measures individual preference prediction rather than value pluralism and was dropped in the next release. The authors argue this gap is not accidental but structural: the field's normative justifications are untested, the goal itself is underspecified, and existing benchmarks are not designed for
What carries the argument
The audit's central object is the pair of public documents labs release: behavior documents that state intended behavior (constitutions, model specs, system prompts) and evaluations that report actual behavior (model or system cards, technical reports). The argument runs on the gap between them: documents prescribe multi-perspective defaults in several labs, while evaluations measure political bias, hedging, or first-person fairness—constructs distinct from pluralism. The paper's proposed replacement machinery is a 'hill-climbable' evaluation: a stated behavioral target in a behavior document, a scoring rubric tied to it, paired-prompt designs, and co-reported measures of accuracy, helpfulne
Load-bearing premise
The audit treats the absence of pluralism from public documents as proof that pluralism is absent from deployed models, even though labs could be pursuing pluralism internally without publishing it.
What would settle it
A single frontier lab publishing a model card that names pluralism as a goal, reports a dedicated pluralism evaluation tied to specific behavior-document provisions, and documents training toward that behavior would contradict the adoption problem; so would an internal document or independent audit showing production models are trained and evaluated for pluralism despite no public reporting.
If this is right
- If the audit is right, users and regulators currently have no way to know whether a deployed model represents diverse perspectives, and no way to hold developers accountable for it.
- The field's implicit theory of change—research produces methods and benchmarks, labs adopt them—has not operated yet, so continued research-scale progress alone will not achieve the field's stated goals.
- Labs already commit to multi-perspective defaults in writing, so a release-ready evaluation tied to those provisions is an achievable near-term deliverable rather than a distant ideal.
- Political bias evaluations, the closest current proxy, measure a different construct and may even trade off against pluralism, so bias scores cannot stand in for pluralism scores.
- Open-weight model families show the same absence of pluralism, and because they lack in-house behavior-specification processes, they may be the path of least resistance for adoption of community-built evaluations.
Where Pith is reading between the lines
- The audit's evidence is about public documentation; if frontier labs are pursuing pluralism internally without publishing, the adoption problem is overstated even though the transparency problem remains.
- A quick empirical test of the paper's second barrier: measure whether people agree more about when a pluralistic answer is warranted than about the answer itself; if they do, there is a learnable signal for pluralistic defaults.
- The paper's own logic implies regulation would be a downstream beneficiary rather than an alternative: no mandate can be written without the empirical evidence and concrete goal the paper calls for.
- A testable extension of the 'hill-climbable' criterion: benchmark whether any existing pluralism score is improved by prompting tricks rather than by genuine training, which would reveal which evaluations can guide production optimization.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that pluralistic value alignment research has not yet been adopted by frontier AI labs, and that the field should redirect its efforts toward empirical, operational, and evaluation work that can be integrated into deployed systems. The authors audit public behavior documents (constitutions, model specs, system prompts) and release evaluations from OpenAI, Anthropic, Google, xAI, Meta, and major open-weight model families, finding that no lab names pluralism as an explicit goal, no system card evaluates pluralistic behavior as defined in the paper, and no externally developed pluralism benchmark appears in release materials. They attribute this gap to three causes: unmeasured benefits of pluralism, an unsettled normative account of when pluralism is warranted, and evaluations/methods that are not suitable for production. The paper proposes three corresponding research directions: empirical studies of pluralism's benefits and costs, an account of when and how pluralism should be operationalized, and release-ready, trade-off-aware evaluations and methods.
Significance. If the audit is read with its stated scope, the paper makes a useful and checkable empirical contribution: it documents, with citations, that frontier labs' public release materials do not currently report dedicated pluralism evaluations, and it clearly separates pluralism from political neutrality and personalization. The proposed research agenda is reasonable and actionable, and the authors provide reproducible code/data for their literature-trend analysis. The paper is less convincing as a demonstration of an 'adoption problem' in the strong sense, because its central empirical claim is overstated in the abstract and partially contradicted by the paper's own appendix. With careful reframing of the claim from 'no adoption' to 'no public evidence of adoption and no public accountability', the paper would be a valuable field-level critique.
major comments (3)
- [Abstract & §A.1] The central claim 'no public evidence that it has shaped the training or evaluation of the AI systems people actually use' is contradicted by the paper's own audit. §A.1 identifies OpenAI's Model Spec and Anthropic's Constitution as 'the most directly integrated into model training' (A.1 intro), and quotes the Model Spec requiring the assistant to 'fairly describe significant views' and 'present perspectives from any point of an opinion spectrum' (A.1.2), and the Constitution requiring Claude to 'represent multiple perspectives in cases where there is a lack of empirical or moral consensus' (A.1.1). Under the paper's §2.1 definition ('represent and serve diverse human values and perspectives'), these are pluralistic prescriptions, and the public documents say they shape training. The finding is only salvageable if narrowed to 'no lab names pluralism as a goal or reports a dedicated plura
- [§3] The premise 'unpublished evaluations ... give users and other stakeholders no transparency ... Therefore, we treat them as equivalent to no evaluation for the purposes of this audit' is defensible as a transparency norm, but it smuggles a normative judgment into an empirical audit. The paper's conclusion (end of §1) says 'pluralistic alignment research is not reaching deployed frontier models,' a factual claim about adoption. Absence of published evaluations cannot establish that claim if labs may be training/evaluating internally; at most it establishes an absence of public evidence and accountability. Please separate the two claims and rephrase the conclusion as 'there is no public evidence of adoption and no public accountability for it,' rather than 'there has been no adoption.'
- [fn. 7, §A.2.2] OpenAI's Model Spec Evals, a public evaluation that scores the Model Spec's multiple-perspectives provisions and reports shortfalls, is excluded because it is not release documentation. The same applies to OpenAI's standalone political-bias evaluation. This scoping choice is reasonable for studying release accountability, but the abstract presents the claim as 'no public evidence' without this restriction. A reader of the abstract alone will find the Model Spec Evals evidence contradicting it. Please state the release-documentation scope in the abstract and in §1, or include standalone public evaluations as evidence and adjust the conclusion.
minor comments (3)
- [§2.4] Typo: 'the sate of the user' should be 'the state of the user'.
- [§4.5] 'conservation-level' should be 'conversation-level' (cf. §2.3 'conversation-level signals').
- [§2.3, §4.4] Benchmarks from Poole-Dayan et al. (2026) are used to assert that 'current models are not pluralistic' and that neutrality ratings correlate negatively with pluralism. These are the authors' own benchmarks; given the paper's audit-like framing, note this dependence and, if possible, cite independent replications or acknowledge the limitation explicitly.
Circularity Check
Audit's adoption criterion (must name 'pluralism') makes the headline 'no public evidence it shaped training' true by construction, while §A.1's training-to-Model-Spec evidence undercuts the behavioral reading.
specific steps
-
self definitional
[Abstract; §2.1; §3; §A.1 intro and §A.1.2]
"Pluralistic alignment is the goal of building AI systems that represent and serve diverse human values and perspectives. ... None of the labs explicitly name pluralism as a goal, but four of the five labs prescribe some form of multi-perspective behavior. ... The two most substantial examples are OpenAI's Model Spec and Anthropic's constitution, which are the most comprehensive and the most directly integrated into model training."
The audit operationalizes 'adoption' as explicitly naming pluralism or reporting a dedicated pluralism evaluation (§1, §3), rather than using the paper's own §2.1 behavioral definition. Under that behavioral definition, the Model Spec and constitution clauses are pluralistic, and §A.1.2 states OpenAI trains models 'to align to the principles in the Model Spec,' including 'fairly describe significant views' and 'present perspectives from any point of an opinion spectrum'; similarly, Anthropic's constitution directs 'represent multiple perspectives.' Thus the abstract's claim that there is 'no public evidence that it has shaped the training or evaluation' is true only if 'it' means the label or the community's artifacts, not the §2.1 construct. The headline finding is therefore equivalent to
full rationale
Most of the audit's evidentiary base is independent: the paper reads primary lab documents, system cards, and open-weight technical reports, and the self-citations (Poole-Dayan et al. 2026; Sorensen et al. 2024) supply benchmarks and definitions that are externally checkable, not fitted to the conclusion. The stated premise that unpublished evaluations count as no evaluation is a transparency assumption, not a circular reduction. The definitional problem is the audit's adoption criterion: the paper defines pluralistic alignment behaviorally, then requires the word 'pluralism' or a dedicated pluralism evaluation to count as adoption. This guarantees the finding that no lab names pluralism, and it conflicts with the paper's own §A.1 evidence that pluralism-adjacent Model Spec and constitution provisions are integrated into training. The evaluation-gap finding and the open-weight audit retain independent content, so the circularity is partial rather than total; score 6 reflects that the central 'shaped training' claim reduces to the chosen criterion.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Absence of pluralism from public behavior documents and evaluations is evidence that labs are not adopting pluralistic alignment; unpublished evaluations count as no evaluation.
- domain assumption Impact of pluralistic alignment should be measured by adoption in deployed frontier models and the most-used open-weight families, because usage is concentrated.
- domain assumption Evaluations drive progress and are the main way AI researchers measure and report progress, so absence of evaluations implies absence of adoption pressure.
read the original abstract
Pluralistic value alignment---the goal of building AI systems that represent and serve diverse human values and perspectives---has emerged as an active research agenda. Yet, there's no public evidence that it has shaped the training or evaluation of the AI systems people actually use. We audit the public behavior documents and evaluations of frontier labs, finding none name pluralism as a goal, and as of this writing, no clear indication that production models are explicitly trained or tested for it. This goes against the primary motivations and goals of pluralistic alignment, which revolve around making a positive difference in the models serving billions of users worldwide. We argue that the pluralistic alignment research community should focus on supporting impact and adoption in deployed, widely-used AI systems. We provide evidence for the adoption problem, present three main reasons behind it, and discuss three corresponding areas for future research to address it: 1. The primary justifications for pluralistic alignment so far have been normative or speculative. We need studies showing empirically how pluralistic AI benefits users or society. 2. The pluralistic alignment research community has not settled when pluralistic behavior is warranted or what pluralism ideally looks like in practice. We need to establish a concrete goal for developers to operationalize. 3. Current methods trade off against other desiderata of LLMs in ways that are largely unmeasured, and existing metrics are not "hill-climbable." We need trade-off-aware evaluations and methods that meet the requirements of production systems. This paper serves as a collective call to action for the pluralistic alignment researchers: progress requires moving beyond normative justification toward empirical foundations, a concrete account of ideal pluralistic behavior, and practical methods and evaluations built for adoption.
Figures
Reference graph
Works this paper leans on
-
[1]
Position:
Kazeev, Nikita and Huyen, Phan Bui Nhat , year =. Position:. Pluralistic
-
[2]
and Balicer, Ran and Brendel, Rebecca Weintraub and Dagan, Noa and Kohane, Isaac S
Chandak, Payal and Alkin, Victoria and Wu, David and Dagan, Maya and Roy, Taposh Dutta and Menezes, Maria Clara Saad and Noori, Ayush and Somia, Nirali and Brownstein, John S. and Balicer, Ran and Brendel, Rebecca Weintraub and Dagan, Noa and Kohane, Isaac S. and Brat, Gabriel A. , year =. What. Pluralistic
-
[3]
On the conversational persuasiveness of
Salvi, Francesco and Horta Ribeiro, Manoel and Gallotti, Riccardo and West, Robert , month = may, year =. On the conversational persuasiveness of. Nature Human Behaviour , publisher =. doi:10.1038/s41562-025-02194-6 , number =
-
[4]
Tsirtsis, Stratis and Rawal, Kai and Russell, Chris and Mittelstadt, Brent and Wachter, Sandra , year =
-
[5]
and Jacobs, Bob M
Conitzer, Vincent and Freedman, Rachel and Heitzig, Jobst and Holliday, Wesley H. and Jacobs, Bob M. and Lambert, Nathan and Mossé, Milan and Pacuit, Eric and Russell, Stuart and Schoelkopf, Hailey and Tewolde, Emanuel and Zwicker, William S. , year =. Position: social choice should guide. Proceedings of the 41st
-
[6]
Tomašev, Nenad and Franklin, Matija and Osindero, Simon , year =
-
[8]
and Nickel, Maximilian , year =
Milli, Smitha and Berker, Ratip Emin and Kraiczy, Sonja and Shi, Claudia and Kussman, Jack and Bose, Avinandan and Elkind, Edith and Bhattacharjee, Himaghna and Procaccia, Ariel D. and Nickel, Maximilian , year =. For. Pluralistic
-
[9]
, year =
Yu, Fangyi and Seedat, Nabeel and Schwarz, Jonathan Richard and Bean, Andrew M. , year =. To. Pluralistic
-
[10]
Introducing
Guo, Alan and Wolfe, Jason , month = mar, year =. Introducing
-
[11]
Building
Davani, Aida and Prabhakaran, Vinodkumar , month = feb, year =. Building
-
[12]
Lambert, Nathan and Brand, Florian , year =. The
-
[13]
State of
Aubakirova, Malika and Atallah, Alex and Clark, Chris and Summerville, Justin and Midha, Anjney , month = dec, year =. State of
-
[14]
Open models lag state-of-the-art closed models by 4 months , url =
Edwards, Jack and Emberson, Luke , year =. Open models lag state-of-the-art closed models by 4 months , url =
-
[15]
Evaluating
Stelling, Lily and Murray, Malcolm and Galizzi, Bruno and Schaffelder, Max and Campos, Siméon and Papadatos, Henry , year =. Evaluating
-
[16]
Trans. Mach. Learn. Res. , author =
-
[18]
Plurality of value pluralism and
Kasirzadeh, Atoosa , year =. Plurality of value pluralism and. Pluralistic
-
[20]
Transactions on Machine Learning Research , author =
Open. Transactions on Machine Learning Research , author =
-
[21]
and Durmus, Esin and Hatfield-Dodds, Zac and Johnston, Scott R
Sharma, Mrinank and Tong, Meg and Korbak, Tomasz and Duvenaud, David and Askell, Amanda and Bowman, Samuel R. and Durmus, Esin and Hatfield-Dodds, Zac and Johnston, Scott R. and Kravec, Shauna M. and Maxwell, Timothy and McCandlish, Sam and Ndousse, Kamal and Rausch, Oliver and Schiefer, Nicholas and Yan, Da and Zhang, Miranda and Perez, Ethan , month = o...
-
[22]
2025 , note =
Qwen3. 2025 , note =
2025
-
[23]
2025 , note =
Executive. 2025 , note =
2025
-
[24]
Model evaluation for extreme risks , url =
Shevlane, Toby and Farquhar, Sebastian and Garfinkel, Ben and Phuong, Mary and Whittlestone, Jess and Leung, Jade and Kokotajlo, Daniel and Marchal, Nahema and Anderljung, Markus and Kolt, Noam and Ho, Lewis and Siddarth, Divya and Avin, Shahar and Hawkins, Will and Kim, Been and Gabriel, Iason and Bolina, Vijay and Clark, Jack and Bengio, Yoshua and Chri...
-
[26]
and Zhao, Zhe and Houlsby, Neil and Diaz, Fernando and Metzler, Donald and Vinyals, Oriol , year =
Dehghani, Mostafa and Tay, Yi and Gritsenko, Alexey A. and Zhao, Zhe and Houlsby, Neil and Diaz, Fernando and Metzler, Donald and Vinyals, Oriol , year =. The
-
[27]
and Hanna, Alex and Paullada, Amandalynne , editor =
Raji, Inioluwa Deborah and Denton, Emily and Bender, Emily M. and Hanna, Alex and Paullada, Amandalynne , editor =. Proceedings of the
-
[28]
Google, Gemini Team , year =. What is
-
[29]
Our approach to the
Google, Gemini Team , year =. Our approach to the
-
[30]
and Sap, Maarten , year =
Ghate, Kshitish and Liu, Andy and Jain, Devansh and Sorensen, Taylor and Kasirzadeh, Atoosa and Caliskan, Aylin and Diab, Mona T. and Sap, Maarten , year =
-
[31]
Askell, Amanda and Bai, Yuntao and Chen, Anna and Drain, Dawn and Ganguli, Deep and Henighan, Tom and Jones, Andy and Joseph, Nicholas and Mann, Ben and DasSarma, Nova and Elhage, Nelson and Hatfield-Dodds, Zac and Hernandez, Danny and Kernion, Jackson and Ndousse, Kamal and Olsson, Catherine and Amodei, Dario and Brown, Tom and Clark, Jack and McCandlish...
-
[32]
and Shaw, Alexander Glenn and Carlini, Nicholas and Li, Boxuan and Raj, Harsh and Bercovich, Ivan and Shi, Lin and Shin, Jeong Yeon and Walshe, Thomas and Buchanan, E
Merrill, Mike A. and Shaw, Alexander Glenn and Carlini, Nicholas and Li, Boxuan and Raj, Harsh and Bercovich, Ivan and Shi, Lin and Shin, Jeong Yeon and Walshe, Thomas and Buchanan, E. Kelly and Shen, Junhong and Ye, Guanghao and Lin, Haowei and Poulos, Jason and Wang, Maoyu and Nezhurina, Marianna and Lu, Di and Mastromichalakis, Orfeas Menis and Xu, Zhi...
-
[33]
Chollet, Francois and Knoop, Mike and Kamradt, Gregory and Landers, Bryan and Pinkard, Henry , year =
-
[34]
A benchmark of expert-level academic questions to assess
Phan, Long and Gatti, Alice and Li, Nathaniel and Hu, Josephina and Zhang, Hugh and Zhang, Chen Bo Calvin and Shaaban, Mohamed and Ling, John and Shi, Sean and Choi, Michael and Agrawal, Anish and Chopra, Arnav and Khoja, Adam and Kim, Ryan and Ren, Richard and Hausenloy, Jason and Zhang, Oliver and Mazeika, Mantas and Yue, Summer and Wang, Alexandr and H...
-
[35]
and Ermon, Stefano and Finn, Chelsea , editor =
Rafailov, Rafael and Sharma, Archit and Mitchell, Eric and Manning, Christopher D. and Ermon, Stefano and Finn, Chelsea , editor =. Direct. Advances in
-
[37]
and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik R
Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik R. , year =. The
-
[38]
, year =
Rein, David and Hou, Betty Li and Stickland, Asa Cooper and Petty, Jackson and Pang, Richard Yuanzhe and Dirani, Julien and Michael, Julian and Bowman, Samuel R. , year =. First
-
[39]
Measuring
Hendrycks, Dan and Burns, Collin and Basart, Steven and Zou, Andy and Mazeika, Mantas and Song, Dawn and Steinhardt, Jacob , year =. Measuring. 9th
-
[40]
Shah, Rohin and Irpan, Alex and Turner, Alexander Matt and Wang, Anna and Conmy, Arthur and Lindner, David and Brown-Cohen, Jonah and Ho, Lewis and Nanda, Neel and Popa, Raluca Ada and Jain, Rishub and Greig, Rory and Albanie, Samuel and Emmons, Scott and Farquhar, Sebastian and Krier, Sébastien and Rajamanoharan, Senthooran and Bridgers, Sophie and Ijito...
-
[41]
and Gu, Keren and Brakman, Anna-Luisa and Mishkin, Pamela and Shah, Meghan and Heidecke, Johannes and Weng, Lilian and Kalai, Adam Tauman , year =
Eloundou, Tyna and Beutel, Alex and Robinson, David G. and Gu, Keren and Brakman, Anna-Luisa and Mishkin, Pamela and Shah, Meghan and Heidecke, Johannes and Weng, Lilian and Kalai, Adam Tauman , year =. First-. The
-
[42]
Menghini, Cristina and Ney, Peter and Kwisaba, Hamza and. Muse. 2026 , note =
2026
-
[43]
Mehta, Ivan , month = jun, year =
-
[44]
and Hitzig, Zoe and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , year =
Chatterji, Aaron and Cunningham, Thomas and Deming, David J. and Hitzig, Zoe and Ong, Christopher and Shan, Carl Yan and Wadman, Kevin , year =. How
-
[45]
Americans and
Gottfried, Jeffrey and Bishop, William and Anderson, Monica and Faverio, Michelle and Park, Eugenie and McClain, Colleen , month = jun, year =. Americans and
-
[46]
Gemini 3
Google, Gemini Team , month = nov, year =. Gemini 3
-
[47]
Gemini 3.1
Google, Gemini Team , month = feb, year =. Gemini 3.1
-
[48]
Google, Gemini Team , year =. Gemini 2.5:. doi:10.48550/arXiv.2507.06261 , abstract =
-
[49]
Zhang, Lily H. and Milli, Smitha and Jusko, Karen Long and Smith, Jonathan and Amos, Brandon and Bouaziz, Wassim and Revel, Manon and Kussman, Jack and Sheynin, Yasha and Titus, Lisa and Radharapu, Bhaktipriya and Yu, Jane and Sarma, Vidya and Rose, Kristopher and Nickel, Maximilian , year =. Cultivating. The
-
[50]
OpenAI , month = apr, year =
-
[51]
OpenAI , month = dec, year =
-
[52]
2026 , note =
Claude's. 2026 , note =
2026
-
[53]
grok-prompts , url =
-
[54]
2025 , note =
Measuring political bias in. 2025 , note =
2025
-
[55]
2026 , note =
System. 2026 , note =
2026
-
[56]
Qiu, Tianyi and He, Zhonghao and Chugh, Tejasveer and Kleiman-Weiner, Max , month = jun, year =. The. Forty-second
-
[57]
Nie, Shangrui and Omoomi, Kian and Flek, Lucie and Zhao, Zhixue and Welch, Charles , month = oct, year =. The
-
[58]
, year =
Poole-Dayan, Elinor and Wu, Jiayi and Sorensen, Taylor and Pei, Jiaxin and Bakker, Michiel A. , year =. Benchmarking. The
-
[59]
Ali, Dalia and Zhao, Dora and Koenecke, Allison and Papakyriakopoulos, Orestis , year =. Operationalizing pluralistic values in large language model alignment reveals trade-offs in safety, inclusivity, and model behavior , isbn =. Proceedings of the. doi:10.1609/aaai.v40i44.41053 , abstract =
-
[60]
and Ren, Xiang and Sap, Maarten , editor =
Zhou, Kaitlyn and Hwang, Jena D. and Ren, Xiang and Sap, Maarten , editor =. Relying on the. Proceedings of the 62nd. 2024 , pages =. doi:10.18653/v1/2024.acl-long.198 , abstract =
-
[61]
Hosking, Tom and Blunsom, Phil and Bartolo, Max , year =. Human. The
-
[63]
Argumentative
Shi, Li and Liu, Houjiang and Wong, Yian and Mujumdar, Utkarsh and Zhang, Dan and Gwizdka, Jacek and Lease, Matthew , year =. Argumentative
-
[64]
Nature , author =. 2025 , pages =. doi:10.1038/s41586-025-09422-z , abstract =
-
[65]
Consequences of
Zhuang, Simon and Hadfield-Menell, Dylan , editor =. Consequences of. Advances in
-
[68]
Understanding the
Kirk, Robert and Mediratta, Ishita and Nalmpantis, Christoforos and Luketina, Jelena and Hambro, Eric and Grefenstette, Edward and Raileanu, Roberta , year =. Understanding the. The
-
[69]
Vishwarupe, Varad and Shadbolt, Nigel and Jirotka, Marina , year =. From
-
[70]
Gabriel, Iason and Manzini, Arianna and Keeling, Geoff and Hendricks, Lisa Anne and Rieser, Verena and Iqbal, Hasan and Tomašev, Nenad and Ktena, Ira and Kenton, Zachary and Rodriguez, Mikel and El-Sayed, Seliem and Brown, Sasha and Akbulut, Canfer and Trask, Andrew and Hughes, Edward and Bergman, A. Stevie and Shelby, Renee and Marchal, Nahema and Griffi...
-
[71]
Distributional
Siththaranjan, Anand and Laidlaw, Cassidy and Hadfield-Menell, Dylan , year =. Distributional. The
-
[72]
Controllable
Zhang, Jingyu and Elgohary, Ahmed and Magooda, Ahmed and Khashabi, Daniel and Durme, Benjamin Van , year =. Controllable. The
-
[73]
Berlin, Isaiah , year =. Four
-
[74]
Political
Rawls, John , year =. Political
-
[75]
Röttger, Paul and Kirk, Hannah and Vidgen, Bertie and Attanasio, Giuseppe and Bianchi, Federico and Hovy, Dirk , editor =. Proceedings of the 2024. 2024 , pages =. doi:10.18653/v1/2024.naacl-long.301 , abstract =
-
[76]
Advances in
Aroyo, Lora and Taylor, Alex and Díaz, Mark and Homan, Christopher and Parrish, Alicia and Serapio-García, Gregory and Prabhakaran, Vinodkumar and Wang, Ding , editor =. Advances in. 2023 , pages =
2023
-
[77]
Pluralistic
Chen, Daiwei and Chen, Yi and Rege, Aniket and Vinayak, Ramya Korlakai , year =. Pluralistic
-
[78]
and Yu, Tong and Kumar, Sachin and Majumder, Bodhisattwa Prasad and Shang, Jingbo and Ammanabrolu, Prithviraj and McAuley, Julian , year =
Xie, Zhouhang and Wu, Junda and Shen, Yiran and Jain, Raghav and Xia, Yu and Li, Xintong and Chang, Aaron and Rossi, Ryan A. and Yu, Tong and Kumar, Sachin and Majumder, Bodhisattwa Prasad and Shang, Jingbo and Ammanabrolu, Prithviraj and McAuley, Julian , year =. A. Second
-
[79]
The benefits, risks and bounds of personalizing the alignment of large language models to individuals , volume =
Kirk, HR and Vidgen, B and Röttger, P and Hale, SA , year =. The benefits, risks and bounds of personalizing the alignment of large language models to individuals , volume =. Nature Machine Intelligence , publisher =
-
[80]
Bao, Han and Huang, Yue and Wang, Xiaoda and Zhang, Zheyuan and Zhou, Yujun and Yang, Carl and Zhang, Xiangliang and Ye, Yanfang , year =
-
[81]
Edelman, Joe and Zhi-Xuan, Tan and Lowe, Ryan and Klingefjord, Oliver and Wang-Mascianica, Vincent and Franklin, Matija and Kearns, Ryan Othniel and Hain, Ellie and Sarkar, Atrisha and Bakker, Michiel and Barez, Fazl and Duvenaud, David and Foerster, Jakob and Gabriel, Iason and Gubbels, Joseph and Goodman, Bryce and Haupt, Andreas and Heitzig, Jobst and ...
-
[83]
Pluralistic
Alavi, Khashayar and Flek, Lucie and Mai, Florian , year =. Pluralistic. Second
-
[84]
Political
Stray, Jonathan and Yang, David Zhai and Luo, Steven and Takagi, Miu Nicole and Chang, Serina , year =. Political
-
[85]
Philosophical
Lazar, Seth , month = dec, year =. Philosophical
-
[86]
Gabriel, Iason and Ghazavi, Vafa , year =. The
-
[88]
Lee, Sunbowen and Zhou, Junting and Ao, Chang and Li, Kaige and Du, Xeron and He, Sirui and Wu, Haihong and Liu, Tianci and Liu, Jiaheng and Alinejad-Rokny, Hamid and Yang, Min and Liang, Yitao and Wen, Zhoufutu and Ni, Shiwen , editor =. Quantification of. 2025 , pages =. doi:10.18653/V1/2025.ACL-LONG.248 , booktitle =
-
[89]
Xu, Xiaohan and Li, Ming and Tao, Chongyang and Shen, Tao and Cheng, Reynold and Li, Jinyang and Xu, Can and Tao, Dacheng and Zhou, Tianyi , year =. A
-
[90]
Freedman, Rachel , year =. Adaptive. doi:https://doi.org/10.48550/arXiv.2605.01642 , booktitle =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.