Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Cognitive Biases in Large Language Models: A Survey and Mitigation Experiments

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A short awareness prompt makes GPT-3.5 and GPT-4 measurably less biased, the paper claims.

desk verdict Useful new negative result on SoPro plus a promising AwaRe effect, but the no-control prompt design means the main debiasing claim is instruction-following in disguise. read the letter →

arxiv 2412.00323 v1 pith:MXB5IFSP submitted 2024-11-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords LargeLanguageModelscognitivebiasdebiasingbandwagoneffectorderpromptengineeringCoBBLErAwaRe
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a two-sentence prompt reminder can make large language models less susceptible to six cognitive biases, most clearly the bandwagon effect. The authors survey prior evidence that LLMs inherit human-like biases from training data, then adapt two human-focused debiasing instructions from crowdsourcing into LLM prompts. In experiments on GPT-3.5 and GPT-4 using the CoBBLEr bias benchmark, the AwaRe (awareness reminder) prompt lowered bias scores, while the SoPro (social projection) prompt often failed or amplified bias. The claim matters because the method requires no retraining, no repeated sampling, and no lengthy outputs, only a short instruction naming the bias.

What carries the argument

Two prompt-level interventions are the mechanism. AwaRe prepends a sentence naming the bias and telling the model to be careful of it, e.g., 'Please answer the following question while being aware of order bias.' SoPro prepends instructions to answer as the majority of people would. Their effectiveness is measured with CoBBLEr, a benchmark that computes scores such as S_Band and S_Verb by checking whether an LLM changes its preferred answer when the order, names, majority signals, or distracting content in a prompt are flipped; a score of zero means no bias, and scores above random (0.25 for most biases) mean the bias is influencing choices.

What would settle it

If running AwaRe with a made-up bias name on the same benchmark reduces scores as much as the real AwaRe does, the measured mitigation is generic instruction-following rather than bias-specific awareness, which would undercut the paper's interpretation. A second decisive check is to have humans rank the same response pairs as rational or irrational and see whether AwaRe's score reductions align with improved human-rated rationality.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that telling an LLM to be aware of a specific bias before it answers — the AwaRe prompt — reduces the measured influence of that bias on its pairwise quality evaluations, whereas asking it to answer as it believes the majority would (SoPro) does not and can make conformity biases worse. Using the CoBBLEr benchmark, which scores bias by seeing whether an LLM changes its preference when inputs are permuted, the authors measure six biases. AwaRe lowered the bandwagon-effect score from 0.524 to 0.260 for GPT-3.5 and from 0.214 to 0.142 for GPT-4, with smaller gains on several other biases, while SoPro raised the bandwagon score to 0.955 for GPT-3.5 and 0.255 for GPT-4. The authors conclude that awareness prompting nudges LLMs toward more rational responses.

Load-bearing premise

The paper assumes that CoBBLEr's consistency scores—checking whether a model reverses its preference when the prompt is permuted—truly measure cognitive bias and that lower scores mean more rational behavior, with no human ground-truth answers to confirm this.

Editorial extensions

If this is right

  • If AwaRe's effect is real, any LLM answer can be made less biased by adding a short bias-naming instruction, with no fine-tuning.
  • SoPro should be avoided for conformity-related biases because it can amplify the bandwagon effect, as the GPT-3.5 score of 0.955 shows.
  • Newer models such as GPT-4 show lower baseline bias, suggesting bias mitigation and model capability improve together.
  • Bias scores for some biases such as attentional bias are near zero at baseline, so AwaRe's main measurable wins are on bandwagon, verbosity, and egocentric biases.
  • The method requires knowing in advance which bias will affect the model, since AwaRe must name the bias.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A likely confound not settled by the paper is that AwaRe's gain is generic instruction-following: a prompt that names any bogus bias and asks the model to be careful might lower consistency scores just as well, a test that can be run with a fabricated bias label on the same benchmark.
  • The consistency-based CoBBLEr scores treat 'changing one's mind when the prompt is permuted' as bias, but a model could be consistent the wrong way; no human-rationality ground truth is given, so 'more rational' is an interpretation, not a proven fact.
  • An extension beyond the paper would test AwaRe on open-source models of different sizes and on human-validated rationality tasks to see whether the score drops correspond to genuinely better decisions.
  • AwaRe's reliance on naming the bias means it cannot address unknown biases; a future direction is automatic bias detection feeding the prompt.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper surveys cognitive biases reported in LLMs and proposes applying two crowdsourcing-inspired prompt interventions, SoPro (social projection) and AwaRe (awareness reminder), to mitigate six biases in GPT-3.5 and GPT-4. The authors evaluate the interventions using the CoBBLEr benchmark, reporting bias scores before and after each intervention. They find that SoPro is largely ineffective or harmful (e.g., it worsens the bandwagon effect), while AwaRe appears to reduce several bias scores, most notably the bandwagon effect. The paper concludes that AwaRe enables LLMs to mitigate the effect of these biases and make more rational responses. The survey component (Table 1) is a useful systematization, but the experimental evidence for the mitigation claim is undermined by the lack of statistical testing, the instruction-following confound inherent in the AwaRe prompt, and a post-hoc relabeling of invalid responses in the attentional-bias analysis.

Significance. If the mitigation result held, it would offer a simple, low-cost, and bias-agnostic prompt intervention for improving LLM evaluator rationality, which is valuable for LLM-as-judge applications and for alignment. The paper's survey of existing literature is a genuine contribution, organizing more than 30 studies by bias type and mitigation. The experimental setup is transparent: the authors use a public benchmark, fix temperature and seed, and report per-condition scores. However, the central claim that AwaRe 'enables LLMs to mitigate the effect of these biases and make more rational responses' (Abstract) is not adequately supported by the evidence presented, for the reasons detailed in the major comments. The observed bandwagon reduction is large and consistent across both models, but the absence of control conditions and significance testing leaves the interpretation open to a simpler explanation: the model is following the explicit instruction to disregard the majority cue.

major comments (3)
  1. [§3.2, §4.3, Table 2] The AwaRe prompt directly names the target bias and instructs the model to be careful about it (e.g., 'Please answer the following question while being aware of order bias'). Consequently, a decrease in a bias score may reflect instruction-following rather than genuine debiasing: when the prompt tells the model that '80% of people believe...' is a bias to avoid, the model may simply discount the majority statement. This confound is present for every bias in Table 2, and it is load-bearing because the abstract's claim rests entirely on these score deltas. The authors should add control conditions, such as a generic 'please answer carefully' prompt, a prompt naming an irrelevant bias, or a prompt that explicitly tells the model to ignore the inserted cue without labeling it as a bias. Without such controls, the observed improvement cannot be attributed to bias mitigation rather than to the model's compliance with the stated instruction.
  2. [§4.3, Table 2] The conclusion that AwaRe mitigates 'these biases' is drawn from point estimates without any confidence intervals, significance tests, or multiple-comparison correction. Several entries in Table 2 show no improvement or even worsening under AwaRe: GPT-3.5 verbosity worsens from 0.096 to 0.119; GPT-4 egocentric bias changes negligibly (0.025 to 0.027); and GPT-4 compassion-fade components move in opposite directions (SComp1 0.062 to 0.064, SComp2 0.073 to 0.076). The bandwagon effect and a few other cells improve, but the abstract's plural claim that AwaRe mitigates the effect of 'these biases' overstates the observed pattern. The paper should restrict its claim to the biases with consistent improvements, or add statistical analysis and discuss the mixed results.
  3. [§4.3, Table 3] The attentional-bias analysis relabels responses that violate the output format as 'valid' after the fact, with no justification that this correction is neutral with respect to the bias measurement. Table 3 shows that before correction, GPT-3.5 baseline has only 71.3% valid responses; after relabeling, the reported SAttn changes from 0.001 to 0.004. This post-hoc procedure changes the quantity being measured and could systematically alter the bias score if format violations correlate with the manipulated irrelevant information. The authors should either report the uncorrected results as the primary outcome, or provide evidence that the relabeling does not bias the SAttn estimates.
minor comments (4)
  1. [Throughout] There are several typographical artifacts in the manuscript, including 'evaluation se/t_ting' (Section 4.2) and 'A/t_tentional bias' in section headings and Table 3. These should be corrected.
  2. [§4.1] The description of the CoBBLEr evaluation says the LLM performs 12,000 evaluations per bias, but the relationship between the 50 questions, 16 responses, and the number of pairwise comparisons is not fully spelled out. A short derivation or reference to the original CoBBLEr protocol would make the experimental load and score denominators clearer.
  3. [§1, §4.4] The phrase 'more rational responses' is used in the Abstract and Discussion, but the evaluation only measures consistency of pairwise preferences under manipulation. A decrease in a consistency-based score is not the same as an increase in rationality; the authors should either define rationality in terms of consistency or soften the claim to 'more consistent' responses.
  4. [Appendix B] The prompt templates in Appendix B are useful, but the bandwagon template places the majority statement after the two system responses and before the format instruction; the exact position relative to the baseline template is not discussed, even though prompt ordering can affect LLM behavior. This should be noted or controlled.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the mitigation claim is an empirical comparison against an external benchmark, not a derivation that reduces to its inputs.

full rationale

The paper does not derive AwaRe's effectiveness from its definition; it measures it with CoBBLEr, an externally published benchmark (Koo et al. [26]) whose consistency scores (Eqs. 1-8) are computed from repeated pairwise evaluations with manipulated cues. No parameter is fitted to the reported outcome, and no load-bearing result is imported from the authors' own prior work. The closest concern is that the AwaRe prompt names the target bias and asks the model to be careful of it (Sec. 3.2), so part of the score reduction may reflect instruction-following rather than improved judgment. That is a construct-validity and control-condition concern, not a circularity: the measured score requires stable preference across cue direction and can move in either direction, as Table 2 shows for several entries. The post-hoc validity relabeling in Table 3 is likewise an untested measurement assumption, not a circular step. The survey contribution is a classification of prior results, and the experimental conclusions are self-contained against an external benchmark. The paper's own stated limitation in Sec. 6, that AwaRe requires knowing the bias name in advance, further confirms that the intervention is cue-specific rather than a hidden fit or a renamed version of the measured quantity. Therefore no circular step is exhibited.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no free parameters or new theoretical entities. The main assumptions are domain assumptions about the CoBBLEr benchmark and the post-hoc correction procedure, which are listed as axioms.

assumptions (4)
  • domain assumption CoBBLEr consistency scores validly measure cognitive bias in LLM evaluations.
    All conclusions compare changes in scores defined in Section 4.2; if these scores do not capture the intended biases, the mitigation results are void.
  • domain assumption A decrease in bias score corresponds to more rational responses.
    The paper interprets lower scores as 'more rational' (Section 4.3, Table 2), but lower inconsistency could in principle accompany worse accuracy.
  • ad hoc to paper Post-hoc relabeling of invalid responses in the attentional-bias analysis is valid.
    In Section 4.3 and Table 3, responses with format violations are 'corrected' and counted as valid if they indicate a preference; this assumption is introduced specifically for this analysis and is not justified.
  • domain assumption GPT-3.5 and GPT-4 API outputs with temperature 0 and seed 0 are deterministic and representative.
    The experiments use a single query per condition (Appendix A); no repeated sampling or seed variation is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cognitive Biases in Large Language Models: A Survey and Mitigation Experiments." pith.science (2026). https://pith.science/paper/MXB5IFSP

@misc{pith2026241200323,
  author       = {Pith},
  title        = {Pith review of: Cognitive Biases in Large Language Models: A Survey and Mitigation Experiments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MXB5IFSP}},
  note         = {Machine review of arXiv:2412.00323}
}
read the original abstract

Large Language Models (LLMs) are trained on large corpora written by humans and demonstrate high performance on various tasks. However, as humans are susceptible to cognitive biases, which can result in irrational judgments, LLMs can also be influenced by these biases, leading to irrational decision-making. For example, changing the order of options in multiple-choice questions affects the performance of LLMs due to order bias. In our research, we first conducted an extensive survey of existing studies examining LLMs' cognitive biases and their mitigation. The mitigation techniques in LLMs have the disadvantage that they are limited in the type of biases they can apply or require lengthy inputs or outputs. We then examined the effectiveness of two mitigation methods for humans, SoPro and AwaRe, when applied to LLMs, inspired by studies in crowdsourcing. To test the effectiveness of these methods, we conducted experiments on GPT-3.5 and GPT-4 to evaluate the influence of six biases on the outputs before and after applying these methods. The results demonstrate that while SoPro has little effect, AwaRe enables LLMs to mitigate the effect of these biases and make more rational responses.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VeriMinder: Mitigating Analytical Vulnerabilities in NL2SQL

    cs.CL 2025-07 conditional novelty 5.0 of 10

    VeriMinder detects cognitive biases in NL2SQL questions and generates refined, more robust analytical questions, outperforming baselines in human and LLM evaluations.

Reference graph

Works this paper leans on

62 extracted references · 36 canonical work pages · cited by 1 Pith paper

  1. [1]

    Aher, Rosa I

    Gati V. Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023. Us ing Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. InProceedings of the 40th International Conference on Machine Learning. PMLR, 337–371

  2. [2]

    Argyle, Ethan C

    Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua R. Guble r, Christopher Rytting, and David Wingate. 2023. Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis 31, 3 (July 2023), 337–351. https://doi.org/10.1017/pan.2023.2

  3. [3]

    AYIDIYA and McKEE J

    STEPHEN A. AYIDIYA and McKEE J. McCLENDON. 1990. RESPONSE EFFECTS IN MAIL SURVEYS. Public Opinion Quarterly 54, 2 (Jan. 1990), 229–247. https://doi.org/10.1086/269200

  4. [4]

    Bakermans- Kranenburg, and Marinus H

    Yair Bar-Haim, Dominique Lamy, Lee Pergamin, Marian J. Bakermans- Kranenburg, and Marinus H. van IJzendoorn

  5. [5]

    Hallucination

    Elijah Berberette, Jack Hutchins, and Amir Sadovnik. 2024. Rede fining "Hallucination" in Llms: Towards a Psychology- Informed Framework for Mitigating Misinformation. arXiv:2402.01769

  6. [6]

    Ning Bian, Hongyu Lin, Peilin Liu, Yaojie Lu, Chunkang Zhang, Ben He, Xianpe i Han, and Le Sun. 2024. Influence of External Information on Large Language Models Mirrors Social Cognit ive Patterns. IEEE Transactions on Computa- tional Social Systems (2024), 1–17. https://doi.org/10.1109/TCSS.2024.347603 0

  7. [7]

    Marcel Binz and Eric Schulz. 2023. Using Cognitive Psychology to Und erstand GPT-3. Proceedings of the National Academy of Sciences 120, 6 (Feb. 2023), e2218523120. https://doi.org/10.1073/ pnas.2218523120

  8. [8]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kap lan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, A riel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, C lemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Li...

Show all 62 references
  1. [9]

    Butts, Devin C

    Marcus M. Butts, Devin C. Lunt, Traci L. Freling, and Allison S. G abriel. 2019. Helping One or Helping Many? A Theoretical Integration and Meta-Analytic Review of the Compassion Fade Literature. Organizational Behavior and Human Decision Processes 151 (March 2019), 16–33. htt...

  2. [10]

    Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, L auro Langosco, Peter Hase, Erdem Biyik, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menel l

    Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gil bert, Jérémy Scheurer, Javier Rando, Rachel Freed- man, Tomasz Korbak, David Lindner, Pedro Freire, Tony Tong Wang, Samuel Marks, Charbel-Raphael Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Daman...

  3. [11]

    Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Mod els Be an Alternative to Human Evaluations?. In Proceedings of the 61st Annual Meeting of the Association fo r Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki...

  4. [12]

    Jessica Maria Echterhoff, Yao Liu, Abeer Alessa, Julian McAu ley, and Zexue He. 2024. Cognitive Bias in Decision- Making with LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2024, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association fo...

  5. [13]

    J. E. Eicher and R. F. Irgolič. 2024. Reducing Selection Bias in La rge Language Models. arXiv:2402.01740 12 Yasuaki Sumita, Koh Takeuchi, and Hisashi Kashima

  6. [14]

    Carsten Eickhoff. 2018. Cognitive Biases in Crowdsourcing. In Proceedings of the Eleventh ACM International Con- ference on Web Search and Data Mining (WSDM ’18) . Association for Computing Machinery, New York, NY, USA, 162–170. https://doi.org/10.1145/3159652.3159654

  7. [15]

    Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. ChatG PT Outperforms Crowd Workers for Text-Annotation Tasks. Proceedings of the National Academy of Sciences 120, 30 (July 2023), e2305016120. https://doi.org/10.1073/pnas.2305016120

  8. [16]

    Thilo Hagendorff, Sarah Fabi, and Michal Kosinski. 2023. Thinking Fas t and Slow in Large Language Models. Nature Computational Science 3, 10 (Oct. 2023), 833–838. https://doi.org/10.1038/s4358 8-023-00527-x

  9. [17]

    Danula Hettiachchi, Mark Sanderson, Jorge Goncalves, Simo Hos io, Gabriella Kazai, Matthew Lease, Mike Schaeker- mann, and Emine Yilmaz. 2021. Investigating and Mitigating Biases in Crowdsour ced Data. In Companion Publication of the 2021 Conference on Computer Supported Coope...

  10. [18]

    John J. Horton. 2023. Large Language Models as Simulated Ec onomic Agents: What Can We Learn from Homo Silicus? https://doi.org/10.3386/w31122 national bureau of economic res earch:31122

  11. [19]

    Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian Mc Auley, and Wayne Xin Zhao. 2024. Large Language Models Are Zero-Shot Rankers for Recommender Systems . In Advances in Information Retrieval: 46th Eu- ropean Conference on Information Retrieval, ECIR 2024, G...

  12. [20]

    Christoph Hube, Besnik Fetahu, and Ujwal Gadiraju. 2019. Unde rstanding and Mitigating Worker Biases in the Crowdsourced Collection of Subjective Judgments. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. ACM, Glasgow Scotland Uk, 1–12. https:/...

  13. [21]

    Israel and C

    Glenn D. Israel and C. L. Taylor. 1990. Can Response Order Bia s Evaluations? Evaluation and Program Planning 13, 4 (Jan. 1990), 365–371. https://doi.org/10.1016/0149-7189( 90)90021-N

  14. [22]

    Itay Itzhak, Gabriel Stanovsky, Nir Rosenfeld, and Yonatan Be linkov. 2024. Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias. Transactions of the Association for Computational Linguis tics 12 (2024), 771–785. https://doi.org/10.1162/tacl_a_00673

  15. [23]

    Erik Jones and Jacob Steinhardt. 2022. Capturing Failures of Lar ge Language Models via Human Cognitive Biases. Advances in Neural Information Processing Systems 35 (Dec. 2022), 11785–11799

  16. [24]

    Tom Kocmi and Christian Federmann. 2023. Large Language Model s Are State-of-the-Art Evaluators of Translation Quality. In Proceedings of the 24th Annual Conference of the European As sociation for Machine Translation , Mary Nur- minen, Judith Brenner, Maarit Koponen, Sirkku L...

  17. [25]

    Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Mats uo, and Yusuke Iwasawa. 2022. Large Language Models Are Zero-Shot Reasoners. Advances in Neural Information Processing Systems 35 (Dec. 2022), 22199–22213

  18. [26]

    Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2024. Benchmarking Cognitive Biases in Large Language Models as Evaluators. In Findings of the Association for Computational Linguis- tics: ACL 2024 , Lun-Wei Ku, Andre Martins, and Vivek Srik...

  19. [27]

    Andrew K Lampinen, Ishita Dasgupta, Stephanie C Y Chan, Hannah R She ahan, Antonia Creswell, Dharshan Ku- maran, James L McClelland, and Felix Hill. 2024. Language Models, l ike Humans, Show Content Effects on Reasoning Tasks. PNAS Nexus 3, 7 (July 2024), pgae233. https://doi.o...

  20. [28]

    Leibenstein

    H. Leibenstein. 1950. Bandwagon, Snob, and Veblen Effects in the Th eory of Consumers’ Demand. The Quarterly Journal of Economics 64, 2 (May 1950), 183–207. https://doi.org/10.2307/188269 2

  21. [30]

    Tianhui Ma, Yuan Cheng, Hengshu Zhu, and Hui Xiong. 2023. Large L anguage Models Are Not Stable Recommender Systems. arXiv:2312.15746

  22. [31]

    Olivia Macmillan-Scott and Mirco Musolesi. 2024. (Ir)Rationalit y and Cognitive Biases in Large Language Models. Royal Society Open Science 11, 6 (June 2024), 240255. https://doi.org/10.1098/rsos.24 0255

  23. [32]

    Andreas Opedal, Alessandro Stolfo, Haruki Shirakami, Ying Jia o, Ryan Cotterell, Bernhard Schölkopf, Ab- ulhair Saparov, and Mrinmaya Sachan. 2024. Do Language Models Exhib it the Same Cognitive Bi- ases in Problem Solving as Human Learners?. In Forty-First International Confe...

  24. [33]

    OpenAI. 2023. GPT-4 Technical Report. arXiv:2303.08774

  25. [34]

    Christiano, Jan Leike, and Ryan Lowe

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwrig ht, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, ...

  26. [35]

    Pouya Pezeshkpour and Estevam Hruschka. 2024. Large Lang uage Models Sensitivity to the Order of Options in Multiple-Choice Questions. In Findings of the Association for Computational Linguistics : NAACL 2024 , Kevin Duh, Helena Gomez, and Steven Bethard (Eds.). Association fo...

  27. [36]

    Leonardo Ranaldi and Fabio Zanzotto. 2024. HANS, Are You Clev er? Clever Hans Effect Analysis of Neural Sys- tems. In Proceedings of the 13th Joint Conference on Lexical and Comp utational Semantics (*SEM 2024) , Danushka Bollegala and Vered Shwartz (Eds.). Association for Comp...

  28. [37]

    Michael Ross and Fiore Sicoly. 1979. Egocentric Biases in Availab ility and Attribution. Journal of Personality and Social Psychology 37, 3 (1979), 322–336. https://doi.org/10.1037/0022-3514 .37.3.322

  29. [38]

    Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. 2023. Verbosity Bias in Preference La- beling by Large Language Models. In NeurIPS 2023 Workshop on Instruction Tuning and Instructio n Following . https://openreview.net/forum?id=magEgFpK1y

  30. [39]

    Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Perc y Liang, and Tatsunori Hashimoto. 2023. Whose Opinions Do Language Models Reflect?. InProceedings of the 40th International Conference on Machine Learning. PMLR, 29971–30004

  31. [40]

    Deborah H. Schenk. 2011. Exploiting the Salience Bias in Designing Ta xes. Yale Journal on Regulation 28, 2 (2011), 253–312

  32. [41]

    Samuel Schmidgall, Carl Harris, Ime Essien, Daniel Olshvang, T awsifur Rahman, Ji Woong Kim, Rojin Ziaei, Ja- son Eshraghian, Peter Abadir, and Rama Chellappa. 2024. Address ing Cognitive Bias in Medical Language Models. arXiv:2402.08113

  33. [42]

    Jonathan Shaki, Sarit Kraus, and Michael Wooldridge. 2023. Co gnitive Effects in Large Language Models. In ECAI

  34. [43]

    Chi, Nathanael Schärli, and Denny Zhou

    Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Doh an, Ed H. Chi, Nathanael Schärli, and Denny Zhou

  35. [44]

    Arabella Sinclair, Jaap Jumelet, Willem Zuidema, and Raquel F ernández. 2022. Structural Persistence in Language Models: Priming as a Window into Abstract Language Representations. Transactions of the Association for Computa- tional Linguistics 10 (2022), 1031–1050. https://do...

  36. [45]

    Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020. Learning to Summarize with Human Feedback. I n Advances in Neural Information Processing Systems, Vol. 33. Curran Associates, Inc., 3008–3021

  37. [46]

    Slater, Ali Ziaee, and Morgan Nguyen

    Gaurav Suri, Lily R. Slater, Ali Ziaee, and Morgan Nguyen. 202 4. Do Large Language Models Show Decision Heuristics Similar to Humans? A Case Study Using GPT-3.5.Journal of Experimental Psychology: General 153, 4 (2024), 1066–1075. https://doi.org/10.1037/xge0001547

  38. [47]

    In Proceedings of the 40th International Conference on Machine Learning

    Large Language Models Can Be Easily Distracted by Irrelev ant Context. In Proceedings of the 40th International Conference on Machine Learning . PMLR, 31210–31227

  39. [48]

    Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkwar , and Graham Neubig. 2024. Do Llms Exhibit Human-like Response Biases? A Case Study in Survey Design. Transactions of the Association for Computational Linguistics 12 (Sept. 2024), 1011–1026. https://doi.org/10.1162...

  40. [49]

    Petter Törnberg. 2023. ChatGPT-4 Outperforms Experts a nd Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning. arXiv:2304.06588

  41. [50]

    Amos Tversky and Daniel Kahneman. 1974. Judgment under Uncertaint y: Heuristics and Biases. Science 185, 4157 (Sept. 1974), 1124–1131. https://doi.org/10.1126/science. 185.4157.1124

  42. [51]

    Talboy and Elizabeth Fuller

    Alaina N. Talboy and Elizabeth Fuller. 2023. Challenging the Ap pearance of Machine Intelligence: Cognitive Bias in Llms and Best Practices for Adoption. https://doi.org/10.48550 /arXiv.2304.01358 arXiv:2304.01358

  43. [52]

    Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yu nbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024. Large Language Models Are Not Fair Ev aluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Li nguistics...

  44. [53]

    Yiwei Wang, Yujun Cai, Muhao Chen, Yuxuan Liang, and Bryan Hooi. 2 023. Primacy Effect of Chat- GPT. In Proceedings of the 2023 Conference on Empirical Methods in N atural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Comput ational Lin...

  45. [54]

    Le, and Denny Zhou

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou

  46. [55]

    Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhix u Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023. Is ChatGPT a Good NLG Evaluator? A Preliminary Study . In Proceedings of the 4th New Frontiers in Summarization Workshop, Yue Dong, Wen Xiao, Lu Wang, Fei ...

  47. [56]

    Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2 021. Calibrate before Use: Improving Few-Shot Performance of Language Models. In Proceedings of the 38th International Conference on Machin e Learning . PMLR, 12697–12706

  48. [57]

    Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 202 4. Large Language Models Are Not Robust Multiple Choice Selectors. In The Twelfth International Conference on Learning Represen tations. https://openreview.net/forum?id=shr9PXz7T0

  49. [58]

    {instruction},

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 20 23. Judging LLM-as-a-judge with MT-bench and Chatbot Arena. Advances in Neural Information Proces...

  50. [60]

    Minghao Wu and Alham Fikri Aji. 2023. Style over Substance: Eva luation Biases for Large Language Models. arXiv:2307.03025

  51. [2007]

    Psychological Bulletin 133, 1 (2007), 1–24

    Threat-Related Attentional Bias in Anxious and Nonanxious Individua ls: A Meta-Analytic Study. Psychological Bulletin 133, 1 (2007), 1–24. https://doi.org/10.1037/0033-2909.1 33.1.1

  52. [2017]

    https://doi.org/10.18653/v1/2024.findings-naacl.130

  53. [2022]

    Advances in Neural Information Processing Systems 35 (Dec

    Chain-of-Thought Prompting Elicits Reasoning in Large Language M odels. Advances in Neural Information Processing Systems 35 (Dec. 2022), 24824–24837

  54. [2023]

    IOS Press, 2105–2112

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.