REVIEW 3 major objections 4 minor 52 references
Explainable AI in Usable Privacy and Security: Challenges and Opportunities
T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper argues that LLM-as-a-judge explanations for privacy-policy assessment must be adapted to three user profiles, and that this adaptation is the main open problem for explainable AI in usable privacy and security.
desk verdict A honest, well-written workshop paper that sells a plausible research agenda for HCXAI in privacy and security, but the user-profile taxonomy that anchors its central claim is a 22-participant hypothesis, not a result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
PRISMe, an interactive privacy-policy assessment tool that instantiates the LLM-as-a-judge paradigm, is the central object. It fetches a website's privacy policy, prompts GPT-4o to identify ethical criteria, rate each on a 5-point Likert scale, and justify the score within a strict word limit, then presents the results through smiley icons, a dashboard, and two chat interfaces. The second load-bearing mechanism is the empirically derived user-profile taxonomy: it converts the design question 'who needs what explanation' into three testable requirements, with Targeted Explorers wanting specific and extensive justifications, Novice Explorers benefiting from exploratory and supportive explanations, and Information Minimalists needing concise overviews.
What would settle it
A replication with a larger and more diverse sample, run with the same tool and with a second LLM backend, could settle the claim: if the three user profiles do not re-emerge, or if the generic-explanation complaints disappear when the judge is given a fixed criteria rubric, then the paper's call for profile-adaptive explanation strategies would be unsupported.
Extended reading notes
Core claim
The paper's central claim is that LLM-as-a-judge is workable but not yet trustworthy as a privacy-policy assessment mechanism. In a 22-participant study of PRISMe, users found the LLM's explanations helpful, yet they also encountered generic justifications that barely differed between good and bad ratings, inconsistent criteria selection across similar policies, and incoherence between chat responses and the stated ratings. The authors conclude that these problems are not merely technical defects but vary systematically with user type, so they propose three user profiles—Targeted Explorers, Novice Explorers, and Information Minimalists—and argue that explanation strategies must be tailored to these profiles. On the hallucination side, the paper distinguishes faithfulness failures (the judge ignores the actual policy) from factuality failures (fabricated criteria or unjustified Likert scores), and suggests structured evaluation criteria, uncertainty estimation, and retrieval-augmented evidence as mitigation directions.
Load-bearing premise
The agenda rests on the assumption that qualitative observations from 22 participants using one prototype, summarized in Section 3.2 without recruitment details or analysis methods, reveal stable and generalizable differences in how users want LLM explanations.
Editorial extensions
If this is right
- A fixed criteria catalog would likely make ratings and explanations more consistent, at the cost of missing policy-specific risks that a rigid rubric cannot capture.
- Communicating uncertainty, through output-token confidence or variance across repeated assessments, can help all three user profiles calibrate trust without overwhelming Information Minimalists.
- Grounding each criterion evaluation in retrieval-augmented evidence from the actual policy should make explanations more specific and less generic, but multiplies computational cost.
- Explanation personalization must be opt-in, because profiling users for better explanations also collects data about them and can create new privacy risks.
- Expert-aligned judge prompts may improve agreement with expert reviewers while reducing alignment with lay users, so prompt design must balance technical precision against lay comprehension.
Reading between the lines
- If the three explanation profiles replicate in other high-stakes LLM advice settings, the same taxonomy could be tested in medical or financial explanation design, where similar generic-rationale complaints are plausible.
- The paper's mitigation proposals imply a concrete experiment: compare users' trust and comprehension when the judge is forced to quote the policy versus when it freely summarizes, isolating whether genericness or unfaithfulness drives the complaints.
- Because the study covered 37 different websites, re-analyzing which privacy-policy features triggered the inconsistency complaints could produce a catalog of hallucination-prone evaluation criteria, a step the paper leaves to future work.
- The opt-in personalization suggestion opens a quantifiable trade-off: measuring how much profile data users are willing to share for better explanations would tell whether adaptive profiling is viable in practice.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that usable privacy and security is a promising application area for Human-Centered Explainable AI (HCXAI), focusing on the LLM-as-a-judge paradigm in the PRISMe privacy-policy assessment tool. It summarizes a prior 22-participant user study that reportedly revealed three user profiles (Targeted Explorers, Novice Explorers, Information Minimalists) and identifies concerns about explanation transparency, consistency, and faithfulness. The paper then discusses potential mitigation strategies, including fixed evaluation criteria, uncertainty estimation, retrieval-augmented generation, and multiple sampling, and calls for adaptive explanation strategies tailored to user profiles. The conclusion frames the contribution as a research agenda for the HCXAI community.
Significance. If the empirical basis were solid, the paper would make a useful contribution by connecting the LLM-as-a-judge literature with human-centered privacy research and by naming concrete open problems, such as explanation faithfulness and hallucination handling in privacy assessments. The paper has clear strengths: it grounds the discussion in a real system (PRISMe), includes the actual prompts in the appendix, and engages with relevant prior work on LLM explanations, hallucinations, and human-AI trust. Its proposed research directions are plausible and timely. However, the central claim about adaptive explanation strategies rests almost entirely on a user-profile taxonomy that is presented as a result but supported only as a hypothesis; as written, the paper is better framed as a position or agenda paper than as an empirical study.
major comments (3)
- [Section 3.2] The user-profile taxonomy (Targeted Explorers, Novice Explorers, Information Minimalists) is load-bearing for the central claim in the abstract and Section 1 that adaptive explanation strategies tailored to different user profiles are needed. Yet the manuscript provides no recruitment details, no coding scheme, no descriptive statistics, no inter-rater agreement, and no demonstration that profile membership predicts differential responses to explanation format. The authors themselves state 'We suspect these groups differ in the way they process LLM explanations and potential hallucinations,' which is a hypothesis, not an empirical result. This should be addressed either by summarizing the underlying analysis from Freiberger et al. [12] in enough detail to support the taxonomy, or by explicitly reframing the paper as a hypothesis-generating research agenda.
- [Section 4.1] The paper states that 'Hallucination scenarios in LLM-as-a-judge use cases have yet to be investigated' and the subsequent mitigation discussion is explicitly forward-looking using 'may help' and 'could help.' This is acceptable for a position paper, but the abstract and conclusion present hallucination handling as a key concern or outcome of the paper. The manuscript should clearly separate the empirical observations from the prior study (e.g., generic explanations, inconsistent criteria selection, chat/rating incoherence) from the proposed research directions and speculative mitigation strategies, so that readers can assess what has actually been shown versus what remains to be tested.
- [Section 4.1 and 4.2] Several proposed solutions, such as fixed criteria catalogs, retrieval-augmented generation, token-level uncertainty, and sampling multiple assessments, are plausible but are not evaluated in any way. The paper does not report preliminary results, feasibility checks, or even a concrete experimental design. If these strategies are part of the paper's contribution, they need to be framed as open research questions rather than recommendations; if they are intended as recommendations, at least some evidence or a pilot evaluation is needed to support them.
minor comments (4)
- [Section 2] The sentence beginning 'Privacify [45] is a browser extension ...' contains a grammatical issue: 'While it lacks interpretation, customization, and interactive features, it further motivated us' would read better as a separate clause or with a period before 'it further motivated us'.
- [References] Reference [42] lists the author as 'Erik Buchmann Vincent Freiberger,' which appears to be a formatting error; the correct author order or separator should be restored.
- [Section 3.2] The profiles are introduced with the phrase 'Our study also revealed different user profiles' but no indication is given of how many participants fell into each profile or whether profiles were pre-defined or derived post hoc. Adding this information would substantially improve reproducibility and reader confidence.
- [Section 4.3] The claim that service providers might modify privacy policies to exploit LLM adversarial robustness is interesting but is presented as an increasing risk without evidence or citation; it should be labeled as a conjecture or supported with relevant adversarial-robustness literature.
Circularity Check
No circularity: the paper is a position piece that grounds its research agenda in an external prior user study and independent literature, with no fitted quantity or self-referential definition at the core.
full rationale
This paper does not contain a derivation chain that reduces to its own inputs. It is a workshop position paper whose central claim is that usable privacy and security is a promising application area for Human-Centered Explainable AI, and that adaptive explanation strategies tailored to user profiles are needed for LLM-as-a-judge. That claim is supported by two kinds of evidence: external literature on LLM-as-a-judge biases, hallucinations, and explanation research, and a prior user study with 22 participants cited as [12]. The user profiles (Targeted Explorers, Novice Explorers, Information Minimalists) are reported as empirical observations from that prior study, not as parameters fitted to data and then re-predicted. The paper explicitly marks open questions ('We suspect these groups differ in the way they process LLM explanations and potential hallucinations' and 'Hallucination scenarios in LLM-as-a-judge use cases have yet to be investigated'), which shows the authors are proposing hypotheses rather than presenting a closed derivation. The self-citation to [12] is load-bearing in the sense that it provides the empirical basis, but it is not circular: the cited prior work is a distinct user study with its own data collection, and the present paper does not redefine its conclusions in terms of itself. There is no equation, fitted parameter, uniqueness theorem, or ansatz smuggled through citation that makes the central claim true by construction. Weaknesses such as the small sample size and the absence of descriptive statistics are legitimate correctness and generalizability concerns, but they are not circularity. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption The qualitative findings from the 22-participant user study on PRISMe generalize to broader populations and other LLM-as-a-judge privacy tools.
- domain assumption LLM behaviors observed with GPT-4o in the PRISMe prototype are representative of LLM-as-a-judge systems in privacy and security contexts more broadly.
Cite this review
Pith. "Pith review of Explainable AI in Usable Privacy and Security: Challenges and Opportunities." pith.science (2026). https://pith.science/paper/GXI5OQVL
@misc{pith2026250412931,
author = {Pith},
title = {Pith review of: Explainable AI in Usable Privacy and Security: Challenges and Opportunities},
year = {2026},
howpublished = {\url{https://pith.science/paper/GXI5OQVL}},
note = {Machine review of arXiv:2504.12931}
}
read the original abstract
Large Language Models (LLMs) are increasingly being used for automated evaluations and explaining them. However, concerns about explanation quality, consistency, and hallucinations remain open research challenges, particularly in high-stakes contexts like privacy and security, where user trust and decision-making are at stake. In this paper, we investigate these issues in the context of PRISMe, an interactive privacy policy assessment tool that leverages LLMs to evaluate and explain website privacy policies. Based on a prior user study with 22 participants, we identify key concerns regarding LLM judgment transparency, consistency, and faithfulness, as well as variations in user preferences for explanation detail and engagement. We discuss potential strategies to mitigate these concerns, including structured evaluation criteria, uncertainty estimation, and retrieval-augmented generation (RAG). We identify a need for adaptive explanation strategies tailored to different user profiles for LLM-as-a-judge. Our goal is to showcase the application area of usable privacy and security to be promising for Human-Centered Explainable AI (HCXAI) to make an impact.
Figures
Reference graph
Works this paper leans on
-
[12]
You don’t need a university degree to comprehend data protection this way
Vincent Freiberger, Arthur Fleig, and Erik Buchmann. 2025. "You don’t need a university degree to comprehend data protection this way": LLM-Powered Interactive Privacy Policy Assessment. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems
work page 2025
-
[1]
Ashraf Abdul, Jo Vermeulen, Danding Wang, Brian Y Lim, and Mohan Kankanhalli. 2018. Trends and trajectories for explainable, accountable and intelligible systems: An hci research agenda. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–18
2018
-
[2]
Bianca Bartelt and Erik Buchmann. 2024. Transparency in Privacy Policies. In 12th International Conference on Building and Exploring Web Based Environments. IARIA Press, Online, 1–6
work page 2024
-
[3]
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, et al. 2024. Llms instead of human judges? a large scale empirical study across 20 nlp evaluation tasks. arXiv preprint arXiv:2406.18403 (2024)
arXiv 2024
-
[4]
Shmuel I Becher and Uri Benoliel. 2021. Law in Books and Law in Action: The Readability of Privacy Policies and the GDPR. In Consumer law and economics. Springer International Publishing, Cham, 179–204
work page 2021
-
[5]
Veronika Belcheva, Tatiana Ermakova, and Benjamin Fabian. 2023. Understanding Website Privacy Policies—A Longitudinal Analysis Using Natural Language Processing. Information 14, 11 (2023), 622
work page 2023
-
[6]
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024. Humans or llms as the judge? a study on judgement biases. arXiv preprint arXiv:2402.10669
arXiv 2024
-
[7]
Teresa Datta and John P Dickerson. 2023. Who’s thinking? A push for human-centered evaluation of LLMs using the XAI playbook. CHI’23 Workshop on Generative AI and HCI (2023)
work page 2023
Show all 52 references
-
[8]
Upol Ehsan and Mark O Riedl. 2020. Human-centered explainable ai: Towards a reflective sociotechnical approach. In HCI International 2020-Late Breaking Papers: Multimodality and Intelligence: 22nd HCI International Conference, HCII 2020, Copenhagen, Denmark, July 19–24, 2020, ...
2020
-
[9]
Nour El Houda Dehimi and Zakaria Tolba. 2024. Attention Mechanisms in Deep Learning : Towards Explainable Artificial Intelligence. In 2024 6th International Conference on Pattern Analysis and Intelligent Systems (PAIS) . 1–7. doi:10.1109/PAIS62114.2024.10541203
2024
-
[10]
European Union. 2016. REGULATION (EU) 2016/679 of the European Parliament AND OF THE Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Da...
2016
-
[11]
Vincent Freiberger and Erik Buchmann. 2024. Legally Binding but Unfair? Towards Assessing Fairness of Privacy Policies. In Proceedings of the 10th ACM International Workshop on Security and Privacy Analytics (IWSPA ’24) . Association for Computing Machinery, New York, NY, USA, 15–22
2024
-
[13]
Lisa Graichen, Matthias Graichen, and Mareike Petrosjan. 2022. How to Facilitate Mental Model Building and Mitigate Overtrust Using HCXAI. In CHI’22 Workshop on Human-Centered Explainable AI
2022
-
[14]
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. 2024. A Survey on LLM-as-a-Judge. arXiv preprint arXiv:2411.15594 (2024)
2024 arXiv
-
[15]
Aamir Hamid, Hemanth Reddy Samidi, Tim Finin, Primal Pappachan, and Roberto Yus. 2023. GenAIPABench: A benchmark for generative AI-based privacy assistants. arXiv preprint arXiv:2309.05138
2023 arXiv
-
[16]
Patrick Hemmer, Max Schemmer, Niklas Kühl, Michael Vössing, and Gerhard Satzger. 2022. On the effect of information asymmetry in human-AI teams. CHI’22 Workshop on Human-Centered Explainable AI (2022)
2022
-
[17]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2025. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Info...
2025
-
[18]
Jenia Kim, Henry Maathuis, and Danielle Sent. 2024. Human-centered evaluation of explainable AI applications: a systematic review. Frontiers in Artificial Intelligence 7 (2024), 1456486
2024
-
[19]
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. Advances in neural information processing systems 35 (2022), 22199–22213
2022
-
[20]
Jenny Kunz and Marco Kuhlmann. 2024. Properties and challenges of llm-generated explanations. arXiv preprint arXiv:2402.10532 (2024)
2024 arXiv
-
[21]
Rithika Lakshminarayanan and Sanjana Gautam. 2024. Balancing Act: Improving Privacy in AI through Explainability. CHI’24 Workshop on Human-Centered Explainable AI (2024)
2024
-
[22]
Hao-Ping Lee, Yu-Ju Yang, Thomas Serban Von Davier, Jodi Forlizzi, and Sauvik Das. 2024. Deepfakes, Phrenology, Surveillance, and More! A Taxonomy of AI Privacy Risks. In Proceedings of the CHI Conference on Human Factors in Computing Systems . Association for Computing Machin...
2024
-
[23]
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, et al. 2024. From generation to judgment: Opportunities and challenges of llm-as-a-judge. arXiv preprint arXiv:2411.16594
2024
-
[24]
Zhaoxin Li, Sophie Yang, and Shijie Wang. 2024. Exploring Personality-Driven Personalization in XAI: Enhancing User Trust in Gameplay. CHI’24 Workshop on Human-Centered Explainable AI (2024)
2024
-
[25]
Q Vera Liao and Kush R Varshney. 2021. Human-centered explainable ai (xai): From algorithms to user experiences. arXiv preprint arXiv:2110.10790 (2021). 8 Freiberger et al
2021 arXiv
-
[26]
Hui Liu, Qingyu Yin, and William Yang Wang. 2019. Towards Explainable NLP: A Generative Explanation Framework for Text Classification. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . 5570–5581
2019
-
[27]
Gianclaudio Malgieri. 2020. The Concept of Fairness in the GDPR: A Linguistic and Contextual Interpretation. In Proceedings of the 2020 Conference on fairness, accountability, and transparency . Association for Computing Machinery, New York, NY, USA, 154–166
2020
-
[28]
Potsawee Manakul, Adian Liusie, and Mark JF Gales. 2023. Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models. arXiv preprint arXiv:2303.08896 (2023)
2023 arXiv
-
[29]
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel ...
2020 doi
-
[30]
Abraham Mhaidli, Selin Fidan, An Doan, Gina Herakovic, Mukund Srinath, Lee Matheson, Shomir Wilson, and Florian Schaub. 2023. Researchers’ experiences in analyzing privacy policies: Challenges and opportunities. Proceedings on Privacy Enhancing Technologies 4 (2023), 287–305
2023
-
[31]
OpenAI. 2024. GPT-4o. https://openai.com/index/hello-gpt-4o/ Accessed: Jan 2025
2024
-
[32]
Irene Pollach. 2005. A Typology of Communicative Strategies in Online Privacy Policies: Ethics, Power and Informed Consent. Journal of Business Ethics 62 (2005), 221–235
2005
-
[33]
Joel R Reidenberg, Travis Breaux, Lorrie Faith Cranor, Brian French, Amanda Grannis, James T Graves, Fei Liu, Aleecia McDonald, Thomas B Norton, Rohan Ramanath, et al. 2015. Disagreeable privacy policies: Mismatches between meaning and users’ understanding. Berkeley Tech. LJ 3...
2015
-
[34]
David Rodriguez, Ian Yang, Jose M Del Alamo, and Norman Sadeh. 2024. Large language models: a new approach for privacy policy analysis at scale. Computing 106 (2024), 1–25
2024
-
[35]
Yao Rong, Tobias Leemann, Thai-Trang Nguyen, Lisa Fiedler, Peizhu Qian, Vaibhav Unhelkar, Tina Seidel, Gjergji Kasneci, and Enkelejda Kasneci
-
[36]
Advait Sarkar. 2024. AI Should Challenge, Not Obey. Commun. ACM 67, 10 (2024), 18–21
2024
-
[37]
Advait Sarkar. 2024. Large Language Models Cannot Explain Themselves. CHI’24 Workshop on Human-Centered Explainable AI (2024)
2024
-
[38]
Florian Schaub, Rebecca Balebako, and Lorrie Faith Cranor. 2017. Designing effective privacy notices and controls. IEEE Internet Computing 21, 3 (2017), 70–77
2017
-
[39]
I agree to the terms and conditions
Nili Steinfeld. 2016. “I agree to the terms and conditions”:(How) do users read privacy policies online? An eye-tracking experiment. Computers in human behavior 55 (2016), 992–1000
2016
-
[40]
Annalisa Szymanski, Noah Ziems, Heather A Eicher-Miller, Toby Jia-Jun Li, Meng Jiang, and Ronald A Metoyer. 2025. Limitations of the LLM-as- a-Judge approach for evaluating LLM outputs in expert knowledge tasks. In Proceedings of the 30th International Conference on Intelligen...
2025
-
[41]
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. 2023. A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation. arXiv preprint arXiv:2307.03987 (2023)
2023 arXiv
-
[42]
Erik Buchmann Vincent Freiberger. 2024. Fair balancing? Evaluating LLM-based privacy policy ethics assessments. In Proceedings of the Third European Workshop on Algorithmic Fairness (EW AF’24). CEUR Workshop Proceedings, Aachen, Germany
2024
-
[43]
Steven Walker-Roberts, Mohammad Hammoudeh, Omar Aldabbas, Mehmet Aydin, and Ali Dehghantanha. 2020. Threats on the horizon: Under- standing security threats in the era of cyber-physical systems. The Journal of Supercomputing 76 (2020), 2643–2664
2020
-
[44]
Maximiliane Windl, Niels Henze, Albrecht Schmidt, and Sebastian S Feger. 2022. Automating contextual privacy policies: Design and evaluation of a production tool for digital consumer privacy awareness. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Sys...
2022
-
[45]
Justin Woodring, Katherine Perez, and Aisha Ali-Gombe. 2024. Enhancing privacy policy comprehension through Privacify: A user-centric approach using advanced language models. Computers & Security 145 (2024), 103997
2024
-
[46]
Austin Xu, Srijan Bansal, Yifei Ming, Semih Yavuz, and Shafiq Joty. 2025. Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings. arXiv preprint arXiv:2503.15620 (2025)
2025 arXiv
-
[47]
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al. 2024. Justice or prejudice? quantifying biases in llm-as-a-judge. arXiv preprint arXiv:2410.02736 (2024)
2024 arXiv
-
[48]
Suzanne Barber
Razieh Nokhbeh Zaeem and K. Suzanne Barber. 2020. The Effect of the GDPR on Privacy Policies: Recent Progress and Future Promise. ACM Trans. Manage. Inf. Syst. 12, 1, Article 2 (dec 2020), 20 pages
2020
-
[49]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al
-
[50]
Alexandra Zytek, Sara Pidò, and Kalyan Veeramachaneni. 2024. Llms for xai: Future directions for explaining explanations. CHI’24 Workshop on Human-Centered Explainable AI (2024). Explainable AI in Usable Privacy and Security: Challenges and Opportunities 9 A Appendix A.1 Promp...
2024
-
[51]
Advances in Neural Information Processing Systems 36 (2023), 46595–46623
Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 (2023), 46595–46623
2023
-
[2023]
IEEE transactions on pattern analysis and machine intelligence 46, 4 (2023), 2104–2122
Towards human-centered explainable ai: A survey of user studies for model explanations. IEEE transactions on pattern analysis and machine intelligence 46, 4 (2023), 2104–2122
2023
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.