REVIEW 3 major objections 6 minor 57 references
Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Giving an LLM relevance judge a different persona acts as a controlled diagnostic probe: verdicts shift in localized ways while global system rankings hold for capable models, and sensitivity concentrates on particular system types.
desk verdict Solid, well-designed empirical study of persona-conditioned LLM judging with a real but addressable external validity gap around summary-based inputs; worth a careful review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is persona conditioning itself: prefixing the fixed UMBRELA relevance-judging prompt with "You are acting as {persona}." Five assessor roles perturb only the assessor perspective: Query-Aligned (query-specific intent interpretation), Domain-Expert (a nine-domain taxonomy with per-query domain assignment at inter-annotator agreement $\kappa = 0.86$), Orthogonal (a deliberately dissimilar interpretation), Evidence-Verification (factual correctness and source credibility), and the Global Assessor Persona (a professional search-quality rater). Each document is summarized once into an approximately 80-token GPT-4o summary and reused across all conditions, so observed differences are attributed to the persona rather than to document context. Sensitivity is read out at three levels: quadratic-weighted Cohen's kappa for judgment agreement, Kendall's tau and Rank-Biased Overlap ($\phi = 0.9$) for system-ranking agreement with human-derived rankings, and mean absolute rank displacement $\text{Sensitivity}(s) = \frac{1}{|P|}\sum_{p\in P}|\Delta r(s,p)|$ relative to UMBRELA for localized system movement.
What would settle it
Re-run the persona protocol on DL20 with full documents instead of the reused 80-token summaries: if persona-conditioned kappa spreads and mean rank displacements shrink or vanish, the reported sensitivity is an artifact of summary judging. A complementary check: hand the same five persona instructions to human assessors; if the structured sensitivity pattern disappears, it is a property of LLM judges rather than of the assessment task itself.
Extended reading notes
Core claim
Persona conditioning produces structured, model-dependent sensitivity rather than uniform evaluation instability. Across six backbones (GPT-4o, GPT-4o-mini, LLaMA-3.1-70B, LLaMA-3.1-8B, Qwen-2.5-72B, Qwen-2.5-7B) and two datasets (TREC DL20, RAG24), persona-conditioned labels generally remain close to the UMBRELA baseline, with differences appearing as localized shifts in assessment strictness, evidential threshold, or interpretation emphasis instead of widespread relevance inversions. At the system level, Kendall's tau between persona-derived and human-derived rankings stays high for high-capacity models, while smaller models produce substantially larger rank displacement (up to a mean of 5.19 on DL20 for LLaMA-3.1-8B). Local rank-displacement analysis shows that sensitivity is concentrated on particular systems and system types: transformer-based neural ranking and reranking runs on DL20 and retrieval-augmented and generation-oriented pipelines on RAG24. Persona source, abstract PersonaHub profiles versus skill-grounded USPersona profiles, has a secondary effect relative to assessor role and model capacity; the contrastive USPersona Orthogonal role induces the largest mean rank displacement on both datasets, positioning persona-conditioned judging as a stress test rather than an alternative labeling strategy.
Load-bearing premise
The paper assumes that judging from a single reused 80-token summary of each document, rather than the full text, preserves how much a judge's persona changes verdicts and rankings; that premise is carried over from prior summarization work and is not revalidated under persona conditioning.
Editorial extensions
If this is right
- Evaluation pipelines can run persona-conditioned judging as a stress test: systems whose ranks move sharply under contrasting assessor roles are flagged as evaluator-sensitive, while a stable global ranking indicates the evaluation is insensitive to framing.
- Smaller LLM judges should not be used for persona-conditioned evaluation: LLaMA-3.1-8B and Qwen-2.5-7B convert persona instructions into broad judgment instability rather than controlled perspective shifts.
- The contrastive, skill-grounded Orthogonal persona is the most effective probe, producing the largest mean rank displacement on both DL20 (2.31) and RAG24 (2.66), while Domain perspectives consistently induce among the smaller shifts.
- Evaluation conclusions about transformer-based neural ranking and reranking systems on DL20 and about RAG-oriented pipelines on RAG24 carry extra uncertainty, because these are the system types where assessor framing moves ranks most.
- Persona source is a second-order factor: abstract PersonaHub and skill-grounded USPersona profiles produce broadly similar sensitivity patterns, with differences confined to specific role-model combinations.
Reading between the lines
- The same probe plausibly transfers to other LLM-as-a-judge settings, such as summarization, dialogue, or code generation, where a parallel claim would be that persona sensitivity concentrates on particular output types and architectures.
- If the findings hold, evaluation practice would likely move toward reporting persona-robustness alongside agreement, for example publishing rank-displacement intervals or a small persona battery next to each NDCG table so consumers can see which system comparisons depend on framing.
- The concentration of sensitivity in neural rerankers and RAG pipelines hints at a mechanism the paper does not test: these systems may produce outputs that share surface characteristics with LLM-preferred text, so certain personas change how much an LLM judge rewards those cues, a hypothesis a paraphrase or cue-injection experiment could isolate.
- Because persona source mattered little while role mattered a lot, the effective variable may be the semantic position of the role instruction in prompt space rather than the persona's descriptive content, testable by measuring sensitivity against instruction embeddings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes using persona conditioning as a diagnostic probe for LLM-based IR evaluation. It instantiates five assessor roles (Query-Aligned, Domain-Expert, Orthogonal, Evidence-Verification, and GAP) drawn from two persona sources (PersonaHub and Nemotron-Personas-USA), compares them against a standard UMBRELA baseline, and measures judgment-level agreement, system-ranking stability, and local rank displacement across six LLM backbones on TREC DL20 and RAG24. The main empirical claims are that persona conditioning produces structured, model-dependent sensitivity rather than uniform instability: judgments usually remain close to the UMBRELA baseline, global system rankings stay stable for high-capacity models, and sensitivity concentrates on particular retrieval system types. The paper concludes that persona-conditioned judging can act as a controlled stress test for LLM-based evaluation pipelines.
Significance. If the empirical findings hold, the paper provides a practical and interpretable method for exposing assessor framing effects in LLM-based IR evaluation, which is a timely topic given the rapid adoption of LLM judges. The study has notable methodological strengths: it ships public code, uses temperature 0, fixes the UMBRELA prompt, reuses fixed summaries to isolate persona variation, and computes 95% bootstrap confidence intervals for model-level sensitivity estimates. It also makes falsifiable predictions about the role of model capacity and assessor role, and it tests two persona sources. These strengths make the paper useful as a systematic sensitivity analysis. However, the central claims depend on an unvalidated summary-based judging setup, and several key sensitivity claims lack statistical support (persona-level displacement in Table 5, system-type concentration in Table 7). These issues do not invalidate the design internally but require additional work to make the diagnostic claims convincing.
major comments (3)
- [Section 3.4] The summary-based judging design reuses a single approximately 80-token GPT-4o summary per document across all persona conditions, with the validity of this choice referenced to the authors' prior work [37]. That prior work establishes that concise summaries preserve judgment behavior and system-level stability for a standard UMBRELA-style configuration, but it does not establish that 80-token abstractive summaries preserve the cues that the Evidence-Verification, Orthogonal, and Domain-Expert personas are explicitly designed to respond to (e.g., source credibility, hedging, domain-specific terminology, and alternative framings, as defined in Section 3.2.1). Because the paper's central claim is that persona conditioning reveals structured sensitivity in LLM-based evaluation, the fidelity of the probe for persona-conditioned judging is an unvalidated input. I request a validation subset (for example, one dataset judged on full documents or substantially longer passages under at least a few persona/model combinations) or a clear delimitation of all conclusions to summary-based judging.
- [Section 4.4.2, Table 7] The conclusion that sensitivity concentrates on particular retrieval system types (neural ranking/reranking systems on DL20, retrieval-augmented/generation-oriented pipelines on RAG24) rests on a small set of 'representative systems' with no statistical comparison to the distribution of rank displacement across all systems. The reported Mean|Δr| ranges (3.69–4.19 on DL20 and 4.31–8.75 on RAG24) have no confidence intervals and are not tested against a null distribution. Under the bootstrap procedure already used for Table 6, the authors should report whether these systems' displacements are significantly larger than the system-level average, or temper the system-type conclusion to 'illustrative examples' rather than a claim of concentration.
- [Section 4.4.1, Table 5] Persona-level mean absolute rank displacement is reported without confidence intervals, despite the same bootstrap resampling approach being available that was used for model-level sensitivity in Table 6. The claim that 'USPersona Orthogonal yields the highest mean absolute rank displacement on both datasets' may be within sampling error; for example, on DL20 USPersona Orthogonal (2.31) is close to Evidence (2.21) and GAP (2.11). Please add confidence intervals or a significance test for all persona-level values in Table 5, or state explicitly that these differences are descriptive and not tested for statistical significance.
minor comments (6)
- [Section 3.2.2] For Orthogonal persona retrieval, the paper says 'top three candidate personas' are retrieved for inspection, but it does not specify whether this means the three with lowest cosine similarity, how ties are broken, or whether a dissimilarity threshold is enforced. Please clarify the retrieval procedure for the Orthogonal role.
- [Section 3.4] The summary length is described as 'approximately 80-token'; please state whether the prompt enforces a hard token limit and report the actual distribution of summary lengths across documents.
- [Table 4] Several RBO values are notably higher than the corresponding UMBRELA baseline (e.g., RAG24 LLaMA-3.1-70B PersonaHub Query RBO 0.991 vs. UMBRELA 0.639). Please explain whether these values reflect genuine top-rank agreement or artifacts such as tied NDCG scores or small numbers of systems.
- [Section 4.4.2] The sentence 'only 7 out of 472 system–persona pairs show consistent directional behavior' should define the denominator and the consistency criterion (e.g., all six models moving in the same direction with a nonzero rank shift).
- [Section 3.1] The paper says 'we use instruction-tuned conversational variants of the open-weight models [39]' but reference [39] is the InstructGPT paper; please specify the exact model identifiers (e.g., meta-llama/Meta-Llama-3.1-8B-Instruct) in the experimental setup.
- [Section 3.2.1] The Domain-Expert role is described as emphasizing 'domain-specific relevance criteria,' but the taxonomy in Table 1 includes a 'General Knowledge and Reasoning' domain. Please clarify how domain-specific criteria are operationalized for such a broad domain.
Circularity Check
No significant circularity; the central claims are empirical measurements compared against independent human judgments and a fixed UMBRELA baseline.
full rationale
The paper's conclusions about persona-conditioned assessor sensitivity are not derived from its assumptions but are direct measurements of LLM outputs. For each query–document pair, persona-conditioned judgments are generated independently under a fixed UMBRELA judging template and compared against both UMBRELA and human labels; system-level stability is then computed from these labels using standard rank-correlation metrics. No parameter is fitted to the target outcome, and no result is defined into existence by the experimental setup. The single notable self-citation is [37], used to justify reusing one 80-token GPT-4o summary per document across persona conditions; even if that prior work were unverified, the persona-sensitivity findings would remain conditional empirical observations rather than consequences of the citation. The summary-reuse issue is an external-validity threat about whether summaries preserve persona-relevant cues, not a circularity in the paper's reasoning. Accordingly, no circular step meeting the evidence standard is present.
Assumptions & free parameters
free parameters (2)
- Document summary length =
~80 tokens
- Top-k persona retrieval count =
3
assumptions (5)
- domain assumption GPT-4o-generated 80-token document summaries preserve the relevance judgment and ranking behavior needed for persona sensitivity analysis.
- domain assumption TREC human relevance labels for DL20 and RAG24 are a valid gold standard for human agreement and human-derived system rankings.
- domain assumption Prepending 'You are acting as {persona}' to the UMBRELA prompt is sufficient to instantiate the intended assessor perspective.
- ad hoc to paper Semantic similarity in all-MiniLM-L6-v2 space selects personas that represent desired assessor perspectives, including dissimilarity for Orthogonal.
- standard math Quadratic-weighted Cohen's kappa, NDCG@10, Kendall's tau, and RBO are appropriate measures for the comparisons made.
Cite this review
Pith. "Pith review of Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation." pith.science (2026). https://pith.science/paper/UTHIBBFQ
@misc{pith2026260810385,
author = {Pith},
title = {Pith review of: Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/UTHIBBFQ}},
note = {Machine review of arXiv:2608.10385}
}
read the original abstract
Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We study persona conditioning as a diagnostic mechanism for exposing LLM assessor sensitivity. Using task-oriented personas drawn from two complementary sources (PersonaHub and NVIDIA Nemotron-Personas-USA), we instantiate five assessor roles emphasizing intent interpretation, domain expertise, contrastive judgment, evidence verification, and global search-quality assessment, compared with a standard UMBRELA baseline. Across six LLM backbones on TREC DL20 and RAG24, our analyses reveal structured rather than uniform assessor sensitivity. Judgments usually remain close to the baseline while shifting assessment strictness, evidential threshold, or interpretation emphasis rather than producing widespread relevance reversals. At the system level, high-capacity models preserve system-ranking agreement, while smaller models amplify persona-induced instability. Local rank-displacement analysis shows sensitivity concentrates on particular retrieval systems and system types, especially neural ranking/reranking systems on DL20 and RAG-oriented pipelines on RAG24. Persona source matters less than assessor role and model capacity. These findings position persona-conditioned judging as a controlled sensitivity probe for stress-testing LLM-based IR evaluation pipelines and identifying systems whose evaluation outcomes are sensitive to assessor framing.
Figures
Reference graph
Works this paper leans on
-
[37]
Samaneh Mohtadi, Kevin Roitero, Stefano Mizzaro, and Gianluca Demartini. 2026. The Effect of Document Summarization on LLM-Based Relevance Judgments. In Proceedings of the 48th European Conference on Information Retrieval (ECIR 2026). arXiv:2512.05334 [cs.IR] https://arxiv.org/abs/2512.05334
-
[2]
Marwah Alaofi, Paul Thomas, Falk Scholer, and Mark Sanderson. 2026. On the Use of LLMs for Relevance Labelling.ACM Transactions on Information Systems (2026). doi:10.1145/3788872
doi:10.1145/3788872 2026
-
[4]
Peter Bailey, Nick Craswell, Ian Soboroff, Paul Thomas, Arjen P. de Vries, and Emine Yilmaz. 2008. Relevance Assessment: Are Judges Exchangeable and Does It Matter?. InProceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 667–674. doi:10. 1145/1390334.1390447
arXiv 2008
-
[6]
Pietro Bernardelle, Stefano Civelli, Leon Fröhling, Riccardo Lunardi, Kevin Roi- tero, and Gianluca Demartini. 2025. Political ideology shifts in large language models.arXiv preprintarXiv:2508.16013 (2025). https://arxiv.org/abs/2508.16013
arXiv 2025
-
[7]
Pietro Bernardelle, Leon Froehling, Stefano Civelli, and Gianluca Demartini. 2026. SubData: Bridging Heterogeneous Datasets to Enable Theory-Driven Evaluation of Political and Demographic Perspectives in LLMs. InProceedings of the the fifth edition of NLPerspectives, Shiran Dudy, Gavin Abercrombie, Valerio Basile, Elisa Leonardelli, and Simona Frenda (Eds...
-
[8]
Pietro Bernardelle, Leon Fröhling, Stefano Civelli, Riccardo Lunardi, Kevin Roi- tero, and Gianluca Demartini. 2025. Mapping and influencing the political ideol- ogy of large language models using synthetic personas. InCompanion Proceedings of the ACM on Web Conference 2025. 864–867. doi:10.1145/3701716.3715578
arXiv 2025
-
[9]
Ben Carterette and Ian Soboroff. 2010. The Effect of Assessor Error on IR System Evaluation. InProceedings of the 33rd International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 539–546. doi:10.1145/ 1835449.1835540
arXiv 2010
- [10]
Show all 57 references
-
[11]
Nuo Chen, Jiqun Liu, Xiaoyu Dong, Qijiong Liu, Tetsuya Sakai, and Xiao-Ming Wu. 2024. AI Can Be Cognitively Biased: An Exploratory Study on Threshold Priming in LLM-Based Batch Relevance Assessment. InProceedings of the 2024 Annual International ACM SIGIR Conference on Researc...
2024
-
[12]
Stefano Civelli, Pietro Bernardelle, and Gianluca Demartini. 2025. The Impact of Persona-based Political Perspectives on Hateful Content Detection. InCompanion Proceedings of the ACM on Web Conference 2025. 1963–1968. doi:10.1145/3701716. 3718383
2025 doi
-
[13]
Pratama, and Gianluca Demartini
Stefano Civelli, Pietro Bernardelle, Nardiena A. Pratama, and Gianluca Demartini
-
[14]
Charles L. A. Clarke and Laura Dietz. 2025. LLM-based Relevance Assessment Still Can’t Replace Human Relevance Assessment. InProceedings of the 11th International Workshop on Evaluating Information Access (EVIA 2025). National Institute of Informatics, Tokyo, Japan. doi:10.207...
2025 doi
-
[15]
Jacob Cohen. 1968. Weighted Kappa: Nominal Scale Agreement Provision for Scaled Disagreement or Partial Credit.Psychological Bulletin70, 4 (1968), 213–220. doi:10.1037/h0026256
1968 doi
-
[16]
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, and Daniel Campos. 2020. Overview of the TREC 2020 Deep Learning Track. InProceedings of the Twenty-Ninth Text REtrieval Conference (TREC 2020) (NIST Special Publication, Vol. 1266). National Institute of Standards and Technology (NI...
2020
-
[18]
Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth
Guglielmo Faggioli, Laura Dietz, Charles L.A. Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, and Henning Wachsmuth. 2023. Perspectives on Large Language Models for Relevance Judgment. InProceedings of ...
2023
-
[19]
Hanpei Fang, Sijie Tao, Nuo Chen, Kai-Xin Chang, and Tetsuya Sakai. 2025. Do Large Language Models Favor Recent Content? A Study on Recency Bias in LLM-Based Reranking. InProceedings of the 2025 Annual International ACM SIGIR Conference on Research and Development in Informati...
2025
-
[21]
Naghmeh Farzi and Laura Dietz. 2025. Does UMBRELA Work on Other LLMs?. InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’25). ACM, 3214–3222. doi:10.1145/ 3726302.3730317
2025
-
[22]
Ryan Lin Feng, Keyu Tian, Hanming Zheng, Congjing Zhang, Li Zeng, and Shuai Huang. 2025. CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models.arXiv preprint arXiv:2512.07890(2025). https://arxiv. org/abs/2512.07890
2025
-
[23]
Leon Fröhling, Gianluca Demartini, and Dennis Assenmacher. 2025. Personas with Attitudes: Controlling LLMs for Diverse Data Annotation. InProceedings of the 9th Workshop on Online Abuse and Harms (WOAH). Association for Computa- tional Linguistics, Vienna, Austria, 468–481. ht...
2025
-
[24]
Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. 2024. Scaling Synthetic Data Creation with 1,000,000,000 Personas.arXiv preprint arXiv:2406.20094 (2024). https://arxiv.org/abs/2406.20094
2024 arXiv
-
[25]
Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, and Asaf Yehudai. 2025. JuStRank: Benchmarking LLM Judges for System Ranking. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for...
2025 doi
-
[26]
Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated Gain-Based Evaluation of IR Techniques.ACM Transactions on Information Systems20, 4 (2002), 422–446. doi:10.1145/582415.582418
2002
-
[27]
Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang, and Yixin Zhu. 2023. Evaluating and Inducing Personality in Pre-trained Language Models. InProceedings of the 37th International Conference on Neural Infor- mation Processing Systems (NeurIPS ’23). Curran Assoc...
2023
-
[28]
Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara
-
[29]
Jüri Keller, Maik Fröbe, Björn Engelmann, Fabian Haak, Timo Breuer, Birger Larsen, and Philipp Schaer. 2026. Formalized Information Needs Improve Large- Language-Model Relevance Judgments. InProceedings of the 49th International ACM SIGIR Conference on Research and Development...
2026 arXiv
-
[30]
Maurice G. Kendall. 1938. A New Measure of Rank Correlation.Biometrika30, 1-2 (1938), 81–93. doi:10.1093/biomet/30.1-2.81
1938 doi
-
[31]
David La Barbera, Riccardo Lunardi, Mengdie Zhuang, and Kevin Roitero. 2025. Impersonating the Crowd: Evaluating LLMs’ Ability to Replicate Human Judg- ment in Misinformation Assessment. InProceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Th...
2025
-
[32]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguisti...
2023 doi
-
[33]
Shane Culpepper, Alistair Moffat, Sachin Pathiyan Cherumanal, Falk Scholer, and Johanne Trippas
Angel Felipe Magnossão de Paula, J. Shane Culpepper, Alistair Moffat, Sachin Pathiyan Cherumanal, Falk Scholer, and Johanne Trippas. 2025. The Effects of Demographic Instructions on LLM Personas. InProceedings of the 48th International ACM SIGIR Conference on Research and Deve...
2025
-
[34]
Yev Meyer and Dane Corneil. 2025. Nemotron-Personas-USA: Synthetic Personas Aligned to Real-World Distributions. https://huggingface.co/datasets/nvidia/ Nemotron-Personas-USA Accessed: 2025-09-27
2025
-
[35]
Stefano Mizzaro. 1997. Relevance: The Whole History.Journal of the American Society for Information Science48, 9 (1997), 810–832. doi:10.1002/(SICI)1097- 4571(199709)48:9<810::AID-ASI6>3.0.CO;2-U
1997 doi
-
[36]
Samaneh Mohtadi and Gianluca Demartini. 2026. Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis. InProceedings of the 48th European Conference on Information Retrieval (ECIR). arXiv:2601.01751 [cs.IR] https://arxiv.org/abs/2601.01751
2026
-
[38]
NIST TREC Organizers. 2024. TREC 2024 Retrieval-Augmented Generation (RAG) Track Guidelines. https://trec-rag.github.io/annoucements/2024-track- guidelines/ Accessed: 2025-09-27
2024
-
[39]
Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training Language Models to Follow Instructions with Human Feedback. In Proceedings of the 36th International Conference ...
2022 doi
-
[40]
Ronak Pradeep, Nandan Thakur, Sahel Sharifymoghaddam, Eric Zhang, Ryan Nguyen, Daniel Campos, Nick Craswell, and Jimmy Lin. 2025. Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track. InProceedings of the 47th European Conference ...
2025 doi
-
[41]
Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles L
Hossein A. Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra, Paul Thomas, and Emine Yilmaz. 2025. Judging the Judges: A Collection of LLM-Generated Relevance Judgements.arXiv preprint arXiv:2502.13908(2025). ...
2025 arXiv
-
[42]
Tefko Saracevic. 1996. Relevance Reconsidered. InInformation Science: Integration in Perspectives (Proceedings of the Second Conference on Conceptions of Library and Information Science (CoLIS 2)), Peter Ingwersen and Niels Ole Pors (Eds.). Copenhagen, Denmark, 201–218
1996
- [43]
-
[44]
Ian Soboroff. 2024. Don’t Use LLMs to Make Relevance Judgments.Information Retrieval Research Journal(2024). doi:10.54195/irrj.19625
2024 doi
-
[45]
Eero Sormunen. 2002. Liberal Relevance Criteria of TREC: Counting on Negligible Documents?. InProceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 324–330. doi:10. 1145/564376.564433
2002
-
[46]
Yamshchikov
Aleksandra Sorokovikova, Sharwin Rezagholi, Natalia Fedorova, and Ivan P. Yamshchikov. 2024. LLMs Simulate Big5 Personality Traits: Further Evidence. InProceedings of the 1st Workshop on Personalization of Generative AI Systems (PERSONALIZE 2024). Association for Computational...
2024
-
[47]
Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024. Large Language Models Can Accurately Predict Searcher Preferences. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24). 1930–1940. doi:...
2024
- [48]
-
[49]
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur, Daniel Campos, Nick Craswell, Ian Soboroff, and Jimmy Lin. 2025. A Large-Scale Study of Relevance Assessments with Large Language Models Using UMBRELA. InProceedings of the 2025 International ACM SIGIR Conference on Innovative Co...
2025
- [50]
-
[52]
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024. Large Language Models are not Fair Evaluators. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...
2024 doi
-
[53]
Yumeng Wang, Jirui Qi, Catherine Chen, Panagiotis Eustratiadis, and Suzan Verberne. 2026. How Role-Play Shapes Relevance Judgment in Zero-Shot LLM Rankers. InAdvances in Information Retrieval (ECIR ’26). Springer, 228–242. doi:10. 1007/978-3-032-21289-4_15
2026
-
[54]
William Webber, Alistair Moffat, and Justin Zobel. 2010. A Similarity Measure for Indefinite Rankings.ACM Transactions on Information Systems28, 4 (2010), 20:1–20:38. doi:10.1145/1852102.1852106
2010
-
[55]
Chawla, and Xiangliang Zhang
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V. Chawla, and Xiangliang Zhang. 2025. Justice or Prejudice? Quantifying Biases in LLM-as-a- Judge. InProceedings of the International Conference on...
2025
-
[56]
Chuting Yu, Hang Li, Guido Zuccon, Joel Mackenzie, and Teerapong Leelanupab
-
[57]
Pengwei Zhan, Zhen Xu, Qian Tan, Jie Song, and Ru Xie. 2024. Unveiling the Lex- ical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Comput...
2024 doi
-
[58]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT- Bench and Chatbot Arena. InAdvances in Neural Information P...
2023
- [59]
-
[60]
Justin Zobel. 1998. How Reliable Are the Results of Large-Scale Information Retrieval Experiments?. InProceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, New York, NY, USA, 307–314. doi:10.1145/290941.291014
1998
-
[62]
A Helpful Assistant
Mingqian Zheng, Jiaxin Pei, Lajanugen Logeswaran, Moontae Lee, and David Jurgens. 2024. When “A Helpful Assistant” Is Not Really Helpful: Personas in System Prompts Do Not Improve Performance of Large Language Models. In Findings of the Association for Computational Linguistic...
2024
-
[2024]
InFindings of the Association for Computational Linguistics: NAACL 2024
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits. InFindings of the Association for Computational Linguistics: NAACL 2024. Association for Computational Linguistics, Mexico City, Mexico, 3605–3627. doi:10.18653/v1/2024.findings-naacl.229
2024 doi
-
[2026]
Ideology-Based LLMs for Content Moderation.ACM Trans. Intell. Syst. Technol.(April 2026). doi:10.1145/3810946
2026 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.