Pith. sign in

REVIEW 3 major objections 6 minor 97 references

Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper argues that ordinary preference-tuning data contain the reasons behind user choices, and that those reasons can be recovered as personas to make LLMs personalize their responses.

desk verdict Solid, honest paper on persona-augmented preference tuning; the central claim is plausible but rests heavily on an LLM judge with modest human agreement. read the letter →

arxiv 2501.11549 v2 pith:GNHPBQM5 submitted 2025-01-20 cs.CL

classification cs.CL
keywords personalizationpreferencetuningpersonainferenceabductivereasoningdirectoptimizationLLMalignmentdatarejectedresponses
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish that ordinary preference-tuning data contain the reasons behind user preferences, and that those reasons can be recovered and used to make language models personalize their answers. It proposes two steps: Persona Inference, where an LLM abductively infers a short persona describing the kind of user who would prefer the chosen response and the kind who would prefer the rejected one, and Persona Tailoring, where models are trained to condition their outputs on such personas. The main claim is that training on persona-augmented preference data with direct preference optimization boosts judged personalization substantially over standard DPO, with the largest gains (a 66% average improvement) on uncommon but valid needs drawn from rejected responses. The paper also claims this transfers to personas written by real users, so the improved personalization is not an artifact of the inferred training personas.

What carries the argument

The device that carries the argument is a two-step pipeline. Persona Inference (PI) prompts an LLM with a prompt and a chosen/rejected response pair and asks it to abductively produce a one-sentence persona for each response: "The user is [attribute] and prefers [explanation]." Persona Tailoring (PT) then feeds the prompt together with the inferred persona to a model trained by few-shot prompting, supervised fine-tuning, or direct preference optimization, so the model learns to condition its answer on the inferred need. The load-bearing identity is that a pairwise preference judgment can be read as evidence about the type of user who prefers each answer, turning one comparison into a training signal for personalization.

What would settle it

Have independent annotators, who cannot see which model produced each response, compare the persona-tailored DPO outputs with standard DPO outputs on rejected-response personas, rating only whether the response truly fits the user's stated need. If human preference does not reproduce the reported 66% average improvement (or at least a large positive edge) over DPO, the paper's central personalization claim would be falsified.

Watch

Extended reading notes

Core claim

The central discovery is that the rejected response in a preference pair is not noise: it encodes a valid but less common user need that can be surfaced by abduction. Across question answering, dialogue, and education datasets, LLMs infer personas with high accuracy (LLaMA-405B at 91% by a GPT-4o judge that itself agrees with humans 90% of the time), and chosen and rejected personas are rated as comparable in quality though rejected ones apply to fewer users. When a small model is trained on chosen-response personas via DPO, it outperforms a DPO baseline trained with no personas on both personalization and response quality, and the margin is larger on rejected-response personas than on chosen ones. An eight-user study with 144 user-written personas confirms that users judge the persona-tailored model as more personalized than the baseline without sacrificing answerability.

Load-bearing premise

The load-bearing premise is that the Prometheus-7B judge's personalization scores reflect genuine tailoring rather than superficial persona-matching, but the judge agrees with human annotators only 62% of the time on a 50-item personalization sample, and the human studies are small.

Editorial extensions

If this is right

  • Existing preference datasets can be upgraded into personalized training corpora without any new user-data collection: run PI once to annotate the data, then train with PT.
  • Rejected-response personas provide a harder and more meaningful evaluation of personalization, so future alignment work can use them to test whether a model serves long-tail user needs.
  • Persona-tailored DPO generalizes to personas that users write themselves, meaning the method is not confined to the LLM-inferred personas used during training.
  • Persona inference doubles as a content-analysis tool, revealing hidden biases in preference datasets (for example, a systematic preference for verbose over concise answers in existing safety data).

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because PI needs only pairwise comparisons, it could in principle be applied to any existing preference corpus, including domain-specific or multilingual data, without additional human annotation; that is a direct extension the paper does not run.
  • The strong results on rejected personas suggest a broader route to pluralistic alignment: deliberately mining disagreements in preference data, rather than discarding losing responses, could help models serve minority needs in other settings.
  • A natural extension would use multiple inferred personas per preference pair, or personas inferred from several related pairs, to capture multi-faceted user needs; the paper itself notes that its one-example PI is a limitation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper proposes to recover the latent reasons behind pairwise preference judgments and to use them for personalization. In the Persona Inference (PI) phase, a teacher LLM abductively infers personas—short descriptions of needs and interests of users who would prefer either the chosen or the rejected response—for entries of four preference datasets (BeaverTails, SHP, Anthropic HHH, Mnemonic). In the Persona Tailoring (PT) phase, a LLaMA-8B student is trained on persona-augmented data via three strategies (few-shot prompting, SFT, DPO), with the persona prepended to the prompt, and compared against the corresponding no-persona baseline using the Prometheus-7B judge on personalization and response quality, summarized by a ΔPQ metric. The paper reports that PT improves personalization across all three generation strategies (Table 2), that PT DPO is the strongest variant (Table 3), and that PT DPO outperforms a standard DPO model especially when the input persona is derived from the rejected response (Table 4; the abstract claims 66% average improvement on rejected personas). A user study with eight users who wrote 144 personas (§5.4) supports the personalization gain on BeaverTails and shows smaller gains on Anthropic HHH, and the paper releases code and persona-augmented datasets.

Significance. If the headline result holds, the contribution is substantial: it converts existing pairwise preference data into a training signal for personalization without collecting new per-user data, and it demonstrates transfer to user-written personas in a small user study. The execution has real strengths that should be credited: comparisons are made within generation strategy; verbosity is controlled in the §5.2 ablations; several human checks are included (90% agreement for the persona-accuracy judge, 62% for the personalization judge, human ratings of persona plausibility/applicability/harmfulness, and the eight-user study); code and persona-augmented datasets are released; and the Limitations and Ethical Considerations sections are candid. The main risk is concentrated on the empirical anchor of the central claim: the headline rejected-persona numbers are produced by a single open-source judge whose calibration was measured on only 50 comparisons by two of the authors, and no human judgment was collected on that specific subset.

major comments (3)
  1. [Abstract; §5.3; Table 4; Appendix A.4] The abstract's lead quantitative claim—'PT DPO is judged as much stronger than DPO on ... personas from rejected responses (66% average improvement in personalization)'—is a Prometheus-7B-only statistic whose computation is never stated; it matches the average of p_win = win/(win+loss) over the five rejected-persona rows of Table 4 (e.g., 45.1/(45.1+23.2) ≈ 0.660 for BeaverTails Reject+Pretr, averaging to ≈65.9%). Appendix A.4 reports only 62% raw agreement with two author annotators on 50 personalization comparisons and gives no confidence interval, and the §5.4 user study covers only user-written personas rather than the rejected personas used for this headline. The paper should provide either human annotations on a sample of the Table 4 rejected-persona comparisons (with per-subset agreement), or an explicit downgrade of the claim to a judge-relative statement, and it should state in the main text how the 66% figure is computed.
  2. [§4.3; Appendix A.4; Tables 2–4] The calibration of the personalization judge is too thin to support the fine-grained win rates in Tables 2–4: it rests on 50 pairwise comparisons judged by two of the authors, who agree only at Kendall's tau = 0.63, with the 62% judge-annotator agreement reported as a single point estimate and no chance-adjusted measure; with three categories the implied chance level is about 33%, and a Wilson interval on 62% of 50 spans roughly 48–75%. Tables 2 and 4 then report win-rate differences as small as 2–3 points (e.g., HHH Chosen quality 35.0 vs 37.0 in Table 4) without any uncertainty estimate. I also ask for a test that the judge, and the PT DPO outputs, are not simply mirroring persona wording (for example, correlating personalization wins with persona-output lexical overlap), since the §3.3 overfitting check applies to the personas themselves rather than to the evaluated outputs or the judge.
  3. [§3.3; §5.3; §5.4] The interpretation of the Table 4 rejected-persona rows as 'uncommon but still valid needs' rests on human evidence collected only for BeaverTails personas (80 personas from GPT-4o/L-405B in §3.3), while the headline generalization in contribution 3 covers HHH and Mnemonic as well; §3.2's quality-tie results for those datasets come from the same GPT-4o judge that is itself one of the persona-producing models, and the §5.4 user study does not elicit personas of the rejected-response type. The manuscript should either extend the human applicability/plausibility ratings to HHH and Mnemonic personas, or scope the 'valid needs' wording to the datasets where human evidence exists.
minor comments (6)
  1. [Abstract; §5.3] Define the '66% average improvement' in the text (the average of p_win over the rejected-persona rows of Table 4), and likewise define the '91% accuracy' attributed to LLaMA-405B in the Introduction, since Figure 3 aggregates many model/dataset cells.
  2. [Table 2] In the SFT block, the labels 'PT FT+Pretr' and 'PT SFT+Pgold' are inconsistent; use one naming convention (PT SFT+Pretr and PT SFT+Pgold) throughout the table and text.
  3. [Appendix A.7; §5.2; §5.3] The caption of Table 9 refers to '§3.1' when the verbosity-controlled comparison is discussed in §5.2/§5.3; correct the cross-reference, and consider moving the length-discrepancy justification (Table 7) into the main text of §5.3 because it directly affects how the headline comparison is read.
  4. [§3.1; Appendix A.4] GPT-4o is both one of the nine PI models whose personas are judged and the judge of §3.1 accuracy; the 90% human-agreement figure is aggregated over the full set of judgments, and the reported Fleiss' kappa of 0.59 is moderate—report the agreement per dataset and state whether the human-check sample includes the harder Mnemonic condition.
  5. [Figure 3] The y-axis values in Figure 3 are not legible at print size; providing numerical values or a companion table would let readers verify the '0.06 gap for L-405B' claim and the Mnemonic floor effect.
  6. [§4.1; Table 5] The statement 'We sample 2449, 1059, and 328 training entries' is confusing because Table 5 splits these totals into SFT and DPO folds plus validation sets; clarify that each number is the union of the SFT and DPO non-test splits.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the persona-inference and persona-tailoring pipeline is an empirical training/evaluation chain with separate data, separate judges, and a human validation component.

full rationale

The paper's claimed derivation chain is empirical rather than definitional. Personas are inferred by teacher LLMs from preference pairs (PI); the student is trained on persona-augmented chosen responses (PT); and the resulting outputs are judged by a different model (Prometheus-7B) plus a small human study. No parameter is fitted to the quantity that is later 'predicted': the headline rejected-persona result compares PT-DPO, trained only on chosen personas and responses, against DPO given the same rejected persona as an inference-time input, so the improvement is measured rather than forced by construction. The persona-accuracy check uses GPT-4o as a judge with 90% human agreement; that is an evaluation reliability issue, not a circularity step. The reported 62% Prometheus-human agreement on personalization is a substantive external-validity concern, but it does not make any equation or construction reduce to its own input. The paper's self-citations, such as the Mnemonic dataset and the abduction metric reference, are independent dataset or method references rather than load-bearing uniqueness claims used to forbid alternatives. No step was found where an output variable is defined in terms of the target claim, a fitted parameter is renamed as a prediction, or a cited prior result by the same authors supplies the force of the derivation. Therefore the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper rests on several domain assumptions about the validity of LLM-based evaluation and the representativeness of LLM-inferred personas. There are no newly invented physical or formal entities. The central claim is empirical, and its soundness hinges on these assumptions.

assumptions (4)
  • domain assumption LLM judges (GPT-4o, Prometheus-7B) provide valid measurements of persona accuracy and personalization.
    The core evaluations in §3 and §5 rely on LLM judgments; the paper provides partial human agreement but assumes the automated judge is a reliable proxy.
  • domain assumption Human preference labels in BeaverTails, SHP, HHH, and Mnemonic reflect genuine user preferences.
    The preference datasets are taken as ground truth for chosen/rejected responses; no external verification of the labels is provided.
  • domain assumption Personas inferred from one pairwise comparison are a sufficient representation of user needs for training a personalized model.
    The method uses a single example per persona (PI), and the authors acknowledge in §8 that multi-example personas could be more nuanced.
  • domain assumption Training on chosen-response personas (PC) transfers to user-written personas at inference.
    PT models are trained only on PC, yet the claim includes generalization to user-written personas (§5.4); the user study is small.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas." pith.science (2026). https://pith.science/paper/GNHPBQM5

@misc{pith2026250111549,
  author       = {Pith},
  title        = {Pith review of: Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GNHPBQM5}},
  note         = {Machine review of arXiv:2501.11549}
}
read the original abstract

LLMs are aligned to follow input instructions by learning which of two responses users prefer for a prompt. However, such preference data do not convey why users prefer responses that are chosen or rejected, so LLMs trained on these datasets cannot tailor responses to varied user needs. To surface these parameters of personalization, we apply abductive reasoning to preference data, inferring needs and interests of users, i.e., personas, that may prefer either response. We test this idea in two steps: Persona Inference (PI), abductively inferring personas of users who prefer chosen or rejected outputs, and Persona Tailoring (PT), training models to tailor outputs to personas from PI. We show: 1) LLMs infer personas accurately explaining why different users may prefer both chosen or rejected outputs; 2) Training on preference data augmented with PI personas via PT boosts personalization and generalizes to supporting user-written personas; and 3) Rejected response personas form harder personalization evaluations, showing PT better aids users with uncommon preferences versus typical alignment methods. We argue for an abductive view of preferences for personalization, asking not only which response is better but when, why, and for whom.

Figures

Figures reproduced from arXiv: 2501.11549 by the authors.

Figure 1
Figure 1. Training methods with typical preference datasets like DPO cannot fully cater to a user’s specified personas. To overcome this, we train models on preference data augmented with LLM-inferred personas, which we call persona tailoring. While preference datasets are valuable, they as￾sume chosen responses are universally better, fail￾ing to consider why users prefer responses (Joshi et al., 2025). In reality, some user… view at source ↗
Figure 2
Figure 2. Overview of this paper. Preference data has a prompt, chosen response, and rejected response (left). Most users prefer the chosen response, but there are valid reasons and personas of users that may prefer either response, which we uncover via abductive reasoning in PERSONA INFERENCE (PI, middle). We then study PERSONA TAILORING (PT) to use the personas for data augmentation in few-shot prompts, fine-tuning, and dir… view at source ↗
Figure 3
Figure 3. GPT-4o judgments on if LLM personas accurately infer users who prefer chosen/rejected responses. Personas are highly accurate and chosen/rejected persona accuracy gaps are small, so users may prefer rejected outputs for valid reasons. L-8B L-70B L-405B Haiku Sonnet Opus GPT-3.5 GPT-4 GPT-4o 0.0 0.5 1.0 BeaverTails L-8B L-70B L-405B Haiku Sonnet Opus GPT-3.5 GPT-4 GPT-4o 0.0 0.5 1.0 Stanford Human Preferences L-8B L-… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Persona quality comparison. Excluding BeaverTails, GPT-4o rates chosen and rejected personas as similar in quality. r1 may be favored over r2 (§2). GPT-4o—a judge with 90% agreement with three Ph.D. students (Ap￾pendix A.4)—evaluates this, judging if a user de￾scribed …
Figure 5
Figure 5. Figure 5: Qualitative sanity check of personas. Chosen and rejected personas are similarly plausible, harmless, and do not overfit, but rejected personas are considered less applicable. ference of 0.1), further suggesting users can prefer rejected outputs for reasons as valid as…
Figure 6
Figure 6. Figure 6: Testing how well models aid user-specified needs. On BeaverTails, PT largely boosts personalization without losing answerability (Dror et al., 2018, 95% bootstrapped CIs). On BeaverTails, both models have high answer￾ability but PTDPO is significantly more personal￾ize…
Figure 8
Figure 8. Figure 8: Rejected response personas are as valid as chosen ones. Prompting LLMs with personas often switches [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: Instructions given to annotators when writing personas for input prompts. [PITH_FULL_IMAGE:figures/full_fig_p020_9.png]
Figure 10
Figure 10. Figure 10: Instructions given to annotators when evaluating responses for prompts and personas on personalization [PITH_FULL_IMAGE:figures/full_fig_p020_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

97 extracted references · 29 canonical work pages

  1. [1]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  2. [2]

    Stevie Bergman, Jennifer Chien, Mark D\' az, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R

    William Agnew, A. Stevie Bergman, Jennifer Chien, Mark D\' az, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R. McKee. 2024. https://doi.org/10.1145/3613904.3642703 The illusion of artificial inclusion . In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI '24, New York, NY, USA. Association for Computing Machinery

  3. [3]

    Anthropic. 2023. Meet claude. https://www.anthropic.com/product. Accessed: 2024-09-10

  4. [4]

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 a . Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862

  5. [5]

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022 b . Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073

  6. [6]

    Michiel Bakker, Martin Chadwick, Hannah Sheahan, Michael Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matt Botvinick, et al. 2022. Fine-tuning language models to find agreement among humans with diverse preferences. Advances in Neural Information Processing Systems, 35:38176--38189

  7. [7]

    Nishant Balepur, Feng Gu, Abhilasha Ravichander, Shi Feng, Jordan Lee Boyd-Graber, and Rachel Rudinger. 2025 a . https://aclanthology.org/2025.naacl-short.5/ Reverse question answering: Can an LLM write a question so hard (or bad) that it can`t answer? In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Comp...

  8. [8]

    Nishant Balepur, Shramay Palta, and Rachel Rudinger. 2024 a . It’s not easy being wrong: Large language models struggle with process of elimination reasoning. In Findings of the Association for Computational Linguistics ACL 2024, pages 10143--10166

Show all 97 references
  1. [9]

    Nishant Balepur, Abhilasha Ravichander, and Rachel Rudinger. 2024 b . https://aclanthology.org/2024.acl-long.555 Artifacts or abduction: How do LLM s answer multiple-choice questions without the question? In Proceedings of the 62nd Annual Meeting of the Association for Computa...

  2. [10]

    Nishant Balepur, Matthew Shu, Alexander Hoyle, Alison Robey, Shi Feng, Seraphina Goldfarb-Tarrant, and Jordan Lee Boyd-Graber. 2024 c . https://doi.org/10.18653/v1/2024.emnlp-main.786 A SMART mnemonic sounds like `` glue tonic '' : Mixing LLM s with student feedback to make mn...

  3. [11]

    Nishant Balepur, Alexa Siu, Nedim Lipka, Franck Dernoncourt, Tong Sun, Jordan Lee Boyd-Graber, and Puneet Mathur. 2025 b . https://aclanthology.org/2025.naacl-long.20/ M o DS : Moderating a mixture of document speakers to summarize debatable queries in document collections . I...

  4. [12]

    Matthew L Bernacki, Meghan J Greene, and Nikki G Lobczowski. 2021. A systematic review of research on personalized learning: Personalized by whom, to what, how, and for what purpose (s)? Educational Psychology Review, 33(4):1675--1715

  5. [13]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  6. [14]

    Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Dinesh Manocha, Furong Huang, Amrit Bedi, and Mengdi Wang. 2024. https://proceedings.mlr.press/v235/chakraborty24b.html M ax M in- RLHF : Alignment with diverse human preferences . In Proceedings of the 41st Internationa...

  7. [15]

    Jiangjie Chen, Xintao Wang, Rui Xu, Siyu Yuan, Yikai Zhang, Wei Shi, Jian Xie, Shuang Li, Ruihan Yang, Tinghui Zhu, Aili Chen, Nianqi Li, Lida Chen, Caiyu Hu, Siye Wu, Scott Ren, Ziquan Fu, and Yanghua Xiao. 2024 a . https://openreview.net/forum?id=xrO70E8UIZ From persona to p...

  8. [16]

    Jin Chen, Zheng Liu, Xu Huang, Chenwang Wu, Qi Liu, Gangwei Jiang, Yuanhao Pu, Yuxuan Lei, Xiaolong Chen, Xingmei Wang, Kai Zheng, Defu Lian, and Enhong Chen. 2024 b . https://doi.org/10.1007/s11280-024-01276-1 When large language models meet personalization: perspectives of c...

  9. [17]

    Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinso...

  10. [18]

    Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H Holliday, Bob M Jacobs, Nathan Lambert, Milan Mosse, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, et al. 2024. Social choice for ai alignment: Dealing with diverse human feedback. In International Joint Conference o...

  11. [19]

    Budhaditya Deb, Ahmed Hassan Awadallah, and Guoqing Zheng. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.456 Boosting natural language generation from instructions with meta-learning . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processin...

  12. [20]

    Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.88 Toxicity in chatgpt: Analyzing persona-assigned language models . In Findings of the Association for Computational Linguistics:...

  13. [21]

    Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018. https://doi.org/10.18653/v1/P18-1128 The hitchhiker ' s guide to testing statistical significance in natural language processing . In Proceedings of the 56th Annual Meeting of the Association for Computational Lin...

  14. [22]

    Li Du, Xiao Ding, Ting Liu, and Bing Qin. 2021. https://doi.org/10.18653/v1/2021.acl-long.403 Learning event graph knowledge for abductive reasoning . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Co...

  15. [23]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  16. [24]

    Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. 2022. Understanding dataset difficulty with V -usable information. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 5988--6008. PMLR

  17. [25]

    Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. https://doi.org/10.18653/v1/p19-1346 ELI5: long form question answering . In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, ...

  18. [26]

    Shangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher, Chan Young Park, Yejin Choi, and Yulia Tsvetkov. 2024. https://aclanthology.org/2024.emnlp-main.240 Modular pluralism: Pluralistic alignment via multi- LLM collaboration . In Proceedings of the 2024 Conference on Empir...

  19. [27]

    Joseph L Fleiss, Bruce Levin, Myunghee Cho Paik, et al. 1981. The measurement of interrater agreement. Statistical methods for rates and proportions, 2(212-236):22--23

  20. [28]

    Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao. 2021. Knowledge distillation: A survey. International Journal of Computer Vision, 129(6):1789--1819

  21. [29]

    Kunal Handa, Yarin Gal, Ellie Pavlick, Noah Goodman, Jacob Andreas, Alex Tamkin, and Belinda Z Li. 2024. Bayesian preference elicitation with language models. arXiv preprint arXiv:2403.05534

  22. [30]

    Miguel A Hern \'a n, Sonia Hern \'a ndez-D \' az, and James M Robins. 2004. A structural approach to selection bias. Epidemiology, 15(5):615--625

  23. [31]

    Yu Hou, Hal Daum \'e Iii, and Rachel Rudinger. 2025. https://aclanthology.org/2025.naacl-long.611/ Language models predict empathy gaps between social in-groups and out-groups . In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for...

  24. [32]

    Alexander Hoyle, Rupak Sarkar, Pranav Goel, and Philip Resnik. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.815 Natural language decompositions of implicit content enable better text representations . In Proceedings of the 2023 Conference on Empirical Methods in Natural L...

  25. [33]

    Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. https://openreview.net/forum?id=nZeVKeeFYf9 Lo RA : Low-rank adaptation of large language models . In International Conference on Learning Representations

  26. [34]

    EunJeong Hwang, Bodhisattwa Majumder, and Niket Tandon. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.393 Aligning language models to user opinions . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 5906--5919, Singapore. Association for ...

  27. [35]

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. 2023. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674

  28. [36]

    Pegah Jandaghi, Xianghai Sheng, Xinyi Bai, Jay Pujara, and Hakim Sidahmed. 2024. https://aclanthology.org/2024.nlp4convai-1.8 Faithful persona-based conversational dataset generation with large language models . In Proceedings of the 6th Workshop on NLP for Conversational AI (...

  29. [37]

    Joel Jang, Seungone Kim, Bill Yuchen Lin, Yizhong Wang, Jack Hessel, Luke Zettlemoyer, Hannaneh Hajishirzi, Yejin Choi, and Prithviraj Ammanabrolu. 2024. https://openreview.net/forum?id=EMrnoPRvxe Personalized soups: Personalized large language model alignment via post-hoc par...

  30. [38]

    Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. 2024. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. Advances in Neural Information Processing Systems, 36

  31. [39]

    Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al. 2023. Ai alignment: A comprehensive survey. arXiv preprint arXiv:2310.19852

  32. [40]

    Zhijing Jin, Nils Heil, Jiarui Liu, Shehzaad Dhuliawala, Yahang Qi, Bernhard Sch \"o lkopf, Rada Mihalcea, and Mrinmaya Sachan. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.717 Implicit personalization in language models: A systematic study . In Findings of the Associ...

  33. [41]

    Zhijing Jin, Julius von K \"u gelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, and Bernhard Schoelkopf. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.748 Causal direction of data collection matters: Implications of causal and anticausal learning for NLP ....

  34. [42]

    Brihi Joshi, Xiang Ren, Swabha Swayamdipta, Rik Koncel-Kedziorski, and Tim Paek. 2025. Improving llm personas via rationalization with psychological scaffolds. arXiv preprint arXiv:2504.17993

  35. [43]

    Anjali Kantharuban, Jeremiah Milbauer, Emma Strubell, and Graham Neubig. 2024. Stereotype or personalization? user identity biases chatbot recommendations. arXiv preprint arXiv:2410.05613

  36. [44]

    Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024. https://aclanthology.org/2024.emnlp-main.248 Prometheus 2: An open source language model specialized in evaluating other langu...

  37. [45]

    Hannah Rose Kirk, Alexander Whitefield, Paul R \"o ttger, Andrew Michael Bean, Katerina Margatina, Rafael Mosquera, Juan Manuel Ciro, Max Bartolo, Adina Williams, He He, Bertie Vidgen, and Scott A. Hale. 2024. https://openreview.net/forum?id=DFr5hteojx The PRISM alignment data...

  38. [46]

    Joshua Klayman. 1995. Varieties of confirmation bias. Psychology of learning and motivation, 32:385--418

  39. [47]

    o pf, Yannic Kilcher, Dimitri von R\

    Andreas K\" o pf, Yannic Kilcher, Dimitri von R\" u tte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Rich\' a rd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu N...

  40. [48]

    Seongyun Lee, Sue Hyun Park, Seungone Kim, and Minjoon Seo. 2024. https://openreview.net/forum?id=recsheQ7e8 Aligning to thousands of preferences via system message generalization . In The Thirty-eighth Annual Conference on Neural Information Processing Systems

  41. [49]

    Li, Alex Tamkin, Noah Goodman, and Jacob Andreas

    Belinda Z. Li, Alex Tamkin, Noah Goodman, and Jacob Andreas. 2025. https://openreview.net/forum?id=LvDwwAgMEW Eliciting human preferences with language models . In The Thirteenth International Conference on Learning Representations

  42. [50]

    Andy Liu, Mona Diab, and Daniel Fried. 2024 a . https://doi.org/10.18653/v1/2024.findings-acl.586 Evaluating large language model biases in persona-steered generation . In Findings of the Association for Computational Linguistics: ACL 2024, pages 9832--9850, Bangkok, Thailand....

  43. [51]

    Tianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J Liu, and Jialu Liu. 2024 b . https://openreview.net/forum?id=xbjSwwrQOe Statistical rejection sampling improves preference optimization . In The Twelfth International Conference on Learning Representations

  44. [52]

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.153 G -eval: NLG evaluation using gpt-4 with better human alignment . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language ...

  45. [53]

    Yantao Liu, Zhao Zhang, Zijun Yao, Shulin Cao, Lei Hou, and Juanzi Li. 2024 c . Aligning teacher with student preferences for tailored training data generation. arXiv preprint arXiv:2406.19227

  46. [54]

    Chaitanya Malaviya, Joseph Chee Chang, Dan Roth, Mohit Iyyer, Mark Yatskar, and Kyle Lo. 2024. Contextualized evaluations: Taking the guesswork out of language model evaluations. arXiv preprint arXiv:2411.07237

  47. [55]

    Nicholas Meade, Elinor Poole-Dayan, and Siva Reddy. 2022. https://doi.org/10.18653/v1/2022.acl-long.132 An empirical survey of the effectiveness of debiasing techniques for pre-trained language models . In Proceedings of the 60th Annual Meeting of the Association for Computati...

  48. [56]

    Hussein Mozannar, Valerie Chen, Mohammed Alsobay, Subhro Das, Sebastian Zhao, Dennis Wei, Manish Nagireddy, Prasanna Sattigeri, Ameet Talwalkar, and David Sontag. 2025. https://openreview.net/forum?id=hGaWq5Buj7 The realhumaneval: Evaluating large language models abilities to ...

  49. [57]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 3...

  50. [58]

    Vishakh Padmakumar, Chuanyang Jin, Hannah Rose Kirk, and He He. 2024. https://openreview.net/forum?id=M2Yqg68jVW Beyond the binary: Capturing diverse preferences with reward regularization . In Workshop on Socially Responsible Language Modelling Research

  51. [59]

    Susanna Pardini, Silvia Gabrielli, Marco Dianti, Caterina Novara, Gesualdo M Zucco, Ornella Mich, and Stefano Forti. 2022. The role of personalization in the user experience, preferences and engagement with virtual reality environments for relaxation. International Journal of ...

  52. [60]

    Charles Sanders Peirce. 1974. Collected papers of charles sanders peirce, volume 5. Harvard University Press

  53. [61]

    Dan Peng, Zhihui Fu, and Jun Wang. 2024. https://aclanthology.org/2024.privatenlp-1.10/ P ocket LLM : Enabling on-device fine-tuning for personalized LLM s . In Proceedings of the Fifth Workshop on Privacy in Natural Language Processing, pages 91--96, Bangkok, Thailand. Associ...

  54. [62]

    Silviu Pitis, Ziang Xiao, Nicolas Le Roux, and Alessandro Sordoni. 2024. https://openreview.net/forum?id=52r4XJYzjg Improving context-aware preference modeling for language models . In The Thirty-eighth Annual Conference on Neural Information Processing Systems

  55. [63]

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36

  56. [64]

    Kavel Rao, Liwei Jiang, Valentina Pyatkin, Yuling Gu, Niket Tandon, Nouha Dziri, Faeze Brahman, and Yejin Choi. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.812 What makes it ok to set a fire? iterative self-distillation of contexts and rationales for disambiguating d...

  57. [65]

    Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. 2024. https://doi.org/10.18653/v1/2024.acl-long.399 L a MP : When large language models meet personalization . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volu...

  58. [66]

    Chowdhury, and Bernard J

    Joni Salminen, Kathleen Guan, Soon-Gyo Jung, Shammur A. Chowdhury, and Bernard J. Jansen. 2020. https://doi.org/10.1145/3313831.3376502 A literature review of quantitative persona creation . In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, CHI '...

  59. [67]

    Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, and Matei Zaharia. 2022. https://doi.org/10.18653/v1/2022.naacl-main.272 C ol BERT v2: Effective and efficient retrieval via lightweight late interaction . In Proceedings of the 2022 Conference of the North Am...

  60. [68]

    Sander Schulhoff, Michael Ilie, Nishant Balepur, Konstantine Kahadze, Amanda Liu, Chenglei Si, Yinheng Li, Aayush Gupta, HyoJung Han, Sevien Schulhoff, et al. 2024. The prompt report: A systematic survey of prompting techniques. arXiv preprint arXiv:2406.06608

  61. [69]

    Catharina M Serino, Christopher P Furner, and Cindi Smatt. 2005. Making it personal: How personalization affects trust over time. In Proceedings of the 38th annual Hawaii international conference on system sciences, pages 170a--170a. IEEE

  62. [70]

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Esin DURMUS, Zac Hatfield-Dodds, Scott R Johnston, Shauna M Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Pere...

  63. [71]

    Matthew Shu, Nishant Balepur, Shi Feng, and Jordan Lee Boyd-Graber. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.784 KARL : Knowledge-aware retrieval and representations aid retention and learning in students . In Proceedings of the 2024 Conference on Empirical Methods in...

  64. [72]

    Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi

    Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell L. Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. 2024. https://openreview.net/forum?id=gQpBnRHwxM Position: A roadmap to pluralisti...

  65. [73]

    Moritz Stephan, Alexander Khazatsky, Eric Mitchell, Annie S Chen, Sheryl Hsu, Archit Sharma, and Chelsea Finn. 2024. Rlvf: learning from verbal feedback without overgeneralization. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org

  66. [74]

    Robert S Taylor. 1962. The process of asking questions. American documentation, 13(4):391--396

  67. [75]

    Yu-Min Tseng, Yu-Chao Huang, Teng-Yun Hsiao, Wei-Lin Chen, Chao-Wei Huang, Yu Meng, and Yun-Nung Chen. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.969 Two tales of persona in LLM s: A survey of role-playing and personalization . In Findings of the Association for Com...

  68. [76]

    Michael V \"o lske, Martin Potthast, Shahbaz Syed, and Benno Stein. 2017. https://doi.org/10.18653/v1/W17-4508 TL ; DR : Mining R eddit to learn automatic summarization . In Proceedings of the Workshop on New Frontiers in Summarization, pages 59--63, Copenhagen, Denmark. Assoc...

  69. [77]

    Andrew Wang, Cristina Aggazzotti, Rebecca Kotula, Rafael Rivera Soto, Marcus Bishop, and Nicholas Andrews. 2023. Can authorship representation learning capture stylistic features? Transactions of the Association for Computational Linguistics, 11:1416--1431

  70. [78]

    Haoxiang Wang, Yong Lin, Wei Xiong, Rui Yang, Shizhe Diao, Shuang Qiu, Han Zhao, and Tong Zhang. 2024 a . https://doi.org/10.18653/v1/2024.acl-long.468 Arithmetic control of LLM s for diverse user preferences: Directional preference alignment with multi-objective rewards . In ...

  71. [79]

    Rose E Wang, Ana T Ribeiro, Carly D Robinson, Susanna Loeb, and Dora Demszky. 2024 b . Tutor copilot: A human-ai approach for scaling real-time expertise. arXiv preprint arXiv:2410.03017

  72. [80]

    Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, and Michael Zeng. 2022. https://doi.org/10.18653/v1/2022.acl-long.226 Training data is more valuable than you think: A simple and effective method by retrieving from training data . In Procee...

  73. [81]

    Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin. 2024 c . Do-not-answer: Evaluating safeguards in llms. In Findings of the Association for Computational Linguistics: EACL 2024, pages 896--911

  74. [82]

    Fung, Cheng Qian, Jeonghwan Kim, Dilek Hakkani-Tur, and Heng Ji

    Shujin Wu, Yi R. Fung, Cheng Qian, Jeonghwan Kim, Dilek Hakkani-Tur, and Heng Ji. 2025. https://aclanthology.org/2025.coling-main.511/ Aligning LLM s with individual preferences via interaction . In Proceedings of the 31st International Conference on Computational Linguistics,...

  75. [83]

    Xiaohan Xu, Chongyang Tao, Tao Shen, Can Xu, Hongbo Xu, Guodong Long, Jian-Guang Lou, and Shuai Ma. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.871 Re-reading improves reasoning in large language models . In Proceedings of the 2024 Conference on Empirical Methods in Natu...

  76. [84]

    Yang, Maxime Robeyns, Thomas Coste, Jun Wang, Haitham Bou Ammar, and Laurence Aitchison

    Adam X. Yang, Maxime Robeyns, Thomas Coste, Jun Wang, Haitham Bou Ammar, and Laurence Aitchison. 2024. https://openreview.net/forum?id=asgCeFRVjt Bayesian reward models for LLM alignment . In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models

  77. [85]

    Hongzhi Yin, Bin Cui, Jing Li, Junjie Yao, and Chen Chen. 2012. https://doi.org/10.14778/2311906.2311916 Challenging the long tail recommendation . Proc. VLDB Endow., 5(9):896–907

  78. [86]

    Dun Zeng, Yong Dai, Pengyu Cheng, Longyue Wang, Tianhao Hu, Wanshun Chen, Nan Du, and Zenglin Xu. 2024. https://aclanthology.org/2024.findings-emnlp.538 On diversified preferences of large language model alignment . In Findings of the Association for Computational Linguistics:...

  79. [87]

    Lechen Zhang, Tolga Ergen, Lajanugen Logeswaran, Moontae Lee, and David Jurgens. 2024 a . Sprig: Improving large language model performance by system prompt optimization. arXiv preprint arXiv:2410.14826

  80. [88]

    Yusen Zhang, Nan Zhang, Yixin Liu, Alexander Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao, Dragomir Radev, Kathleen McKeown, and Rui Zhang. 2024 b . https://doi.org/10.18653/v1/2024.naacl-long.187 Fair abstractive summarization of diverse perspectives . ...

  81. [89]

    Zhehao Zhang, Ryan A Rossi, Branislav Kveton, Yijia Shao, Diyi Yang, Hamed Zamani, Franck Dernoncourt, Joe Barrow, Tong Yu, Sungchul Kim, et al. 2024 c . Personalization of large language models: A survey. arXiv preprint arXiv:2411.00027

  82. [90]

    Wenting Zhao, Justin Chiu, Claire Cardie, and Alexander Rush. 2023. https://doi.org/10.18653/v1/2023.acl-long.831 Abductive commonsense reasoning exploiting mutually exclusive explanations . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguis...

  83. [91]

    Wenting Zhao, Justin Chiu, Jena Hwang, Faeze Brahman, Jack Hessel, Sanjiban Choudhury, Yejin Choi, Xiang Li, and Alane Suhr. 2024. https://doi.org/10.18653/v1/2024.naacl-long.469 UN commonsense reasoning: Abductive reasoning about uncommon situations . In Proceedings of the 20...

  84. [92]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024 a . Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36

  85. [93]

    Mingqian Zheng, Jiaxin Pei, Lajanugen Logeswaran, Moontae Lee, and David Jurgens. 2024 b . https://doi.org/10.18653/v1/2024.findings-emnlp.888 When a helpful assistant is not really helpful: Personas in system prompts do not improve performances of large language models . In F...

  86. [94]

    Ruiqi Zhong, Peter Zhang, Steve Li, Jinwoo Ahn, Dan Klein, and Jacob Steinhardt. 2023. Goal driven discovery of distributional differences via language descriptions. Advances in Neural Information Processing Systems, 36:40204--40237

  87. [95]

    Yanmengqian Zhou and Lijiang Shen. 2022. Confirmation bias and the persistence of misinformation on climate change. Communication Research, 49(4):500--523

  88. [96]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...

  89. [97]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.