{"id":"cc0444bc-4ae8-4853-a408-a6fe39d1ecf6","arxiv_id":"2412.06593","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Large language models show anchoring bias: their numerical answers move toward biased hints, and simple mitigation prompts do not eliminate the effect.","lead":"This paper tests whether large language models shift their numerical answers when given biased hints, such as an expert's guess or an irrelevant fact. It finds that models like GPT-4 are pulled toward the hints, and that simple prompting strategies do not remove this effect.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.2's sign-match test cannot distinguish anchoring bias from rational interpolation, because informative hints would move answers in the same direction even without any bias.","rationale":"The paper's stated contribution is that LLMs exhibit anchoring bias rather than merely responding to context. The empirical core is Section 4.2, where the only test is sign concordance between (mean answer B - mean answer A) and (H3 - H2), plus a t-test. This statistic is monotone in the hint values: any response function that is increasing in the numerical hint yields sign concordance. Since H2 and H3 are presented as factual reference points or expert predictions, an unbiased Bayesian agent using them as evidence would also increase its estimate. Therefore the design has a confound between disproportionate anchoring and rational information use. The issue is not that the observed directionality is false; it is that it is equally compatible with the null hypothesis of 'no anchoring bias.' This is exactly the load-bearing assumption the reader identified. The paper partially anticipates the need for a baseline: Figure 9 reports a No_anchor column for the 12 expert questions and uses it to evaluate Both-Anchor. That shows the data to run a control exists, and it strengthens the mitigation discussion, but it is not incorporated into the main bias statistic. Because the headline claim is about bias, the missing control/normative comparison is a real soft spot rather than a stylistic concern. A random-anchor arm would settle the issue: if answers follow semantically meaningless random numbers, the effect is genuinely anchoring; if they do not, the current tables only show sensitivity to informative hints, and the conclusion should be weakened to that. The recommended verdict remains CONDITIONAL because the flaw is repairable with additional experiments and does not require discarding the observed effects.","tokens_in":16072,"tokens_out":6549,"duration_ms":77015,"concrete_test":"Re-run the core protocol on the same 62 questions with two additional arms for each model: (i) H1 only as a no-anchor control, and (ii) H1 plus a randomly drawn, semantically uninformative anchor paired as 'low'/'high' (e.g., two random numbers independently sampled from a wide range and presented as 'some user's guess'). If, under random anchors, Treatment A and B no longer show sign-concordant differences, or if real hints do not move answers significantly more than random hints relative to the no-anchor control, then the Section 4.2 sign-match test alone cannot support the anchoring-bias claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In Section 4.2 the paper defines evidence of anchoring as a significant sign-concordant difference between Treatment A (H1+H2) and Treatment B (H1+H3). But H2 and H3 are not arbitrary primes: in the fact-anchoring and expert-opinion conditions they are informative about the target quantity (e.g., last week's low/high stock price, another forecaster's estimate). A model that simply treats the provided hint as additional noisy evidence and forms a weighted average of H1 and the hint will also produce A<B whenever H2<H3, and can do so with small p-values and sign consistency. The t-test only rejects equality of the two treatment distributions; it does not reject the null hypothesis that responses use the hints in a rational, unbiased way. The anchoring-index comparison in Section 4.3 (AI = median_high - median_low over high_anchor - low_anchor) adds magnitude information, but it too has no no-anchor or normative baseline: a rational integrator can produce any intermediate AI depending on the relative reliability it assigns to H1 and the hint. The paper's own Figure 9 includes a No_anchor condition, but it is used only for the Both-Anchor mitigation evaluation, not for the main bias test. Consequently the central claim—that answers are biased by anchors—is not established by the reported statistics; what is established is the weaker claim that answers are sensitive to the values of the hints, which is a necessary but not sufficient condition for anchoring bias.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an empirical study of anchoring bias in large language models (GPT-4, GPT-4o, GPT-3.5) using 62 numerical-prediction questions taken from a human behavioral experiment by Yasseri and Reher. For each question, the model is given a contextual hint H1 plus either a low anchor H2 or a high anchor H3, and the distribution of model answers under the two treatments is compared with t-tests. The authors define evidence of anchoring as a statistically significant difference between Treatment A and Treatment B whose sign matches the sign of H3−H2. They report that most differences are sign-consistent, that stronger models are more consistently influenced, that expert-opinion anchors are particularly effective, and that simple prompting mitigations (Chain-of-Thought, Principles of Thought, Ignore, Reflection) do not reduce the effect, whereas providing both H2 and H3 sometimes brings answers close to a no-anchor benchmark. The paper concludes that LLMs exhibit anchoring bias and that collecting hints from multiple angles is a partially effective mitigation.","tokens_in":16367,"tokens_out":4538,"duration_ms":51107,"significance":"If the central claim were established, the study would provide a useful quantitative demonstration of a cognitive bias in LLMs, extending prior work beyond financial questions and offering practical guidance for prompt design. The use of a diverse 62-item human dataset, repeated sampling of 30 responses per condition, and the comparison with human anchoring indices are strengths, and the authors make their prompts and answers publicly available. However, the experimental design conflates anchoring bias with rational sensitivity to informative hints: because H2 and H3 are factual or expert opinions about the target quantity, a model that simply weights all provided evidence would also produce sign-consistent differences. The paper therefore measures sensitivity to hints, not bias. The mitigation claims rest on a small, selected subset of questions and lack formal comparisons. The result is of interest as a descriptive measurement, but the central inference to 'anchoring bias' is not yet supported.","major_comments":[{"comment":"The operational definition of anchoring bias is a sign-concordant difference between Treatment A (H1+H2) and Treatment B (H1+H3). This criterion cannot distinguish anchoring bias from rational integration of informative hints. For the fact-anchoring and expert-opinion conditions, H2 and H3 carry genuine information about the target quantity (e.g., last week's stock price range, another forecaster's estimate). A model that forms a posterior estimate as a weighted average of H1 and the provided hint will also produce A<B whenever H2<H3, and may do so with very small p-values. The t-test only rejects equality of the two treatment distributions; it does not reject the null hypothesis that the model uses the hints in an unbiased way. The No_anchor condition already present in Figure 9 should be used in the main analysis as a baseline, or the paper should specify a normative model for the unbiased answer and show that the model's answers are pulled toward the anchor beyond that benchmark.","section":"Section 4.2"},{"comment":"The claim that 151 of 162 sign-consistent differences provide 'strong evidence' for anchoring is overstated because sign consistency is counted regardless of statistical significance. Many of the differences in Figures 4 and 5 have p-values well above 0.05 (e.g., question 3 with p=1.00, question 21 with p=0.163, question 22 with p=0.813 for GPT-3.5), yet they are apparently treated as non-exceptions if the sign matches. The paper should report the number of sign-consistent differences that are also statistically significant at a pre-specified level, and ideally account for multiple comparisons across the 54 tests per model. As presented, the evidence is weaker than the text suggests.","section":"Section 4.2.1"},{"comment":"The anchoring index AI = (median_high − median_low)/(high − low) is a measure of response sensitivity to the hint values, not a measure of bias. An unbiased rational integrator can produce any intermediate AI depending on the relative reliability it assigns to H1 versus the hint, so the value 0.45 for GPT-4 does not indicate that the model is 'biased' or that its bias is smaller than the human value of 0.61. The comparison to the human anchoring index also lacks confidence intervals and a formal statistical test. The claims about human–LLM differences in anchoring strength are therefore not supported by the reported data.","section":"Section 4.3"},{"comment":"The mitigation experiments are run only on the 12 'expert' questions, which are the questions with the strongest anchoring effects, and the paper provides no formal comparison between the mitigation conditions and the no-mitigation baseline. For example, Figures 7 and 8 should be tested against the corresponding differences in Figure 5 to determine whether the absolute bias magnitude is significantly reduced by CoT, PoT, Ignore, or Reflection; the current narrative relies on visual inspection. The Both-Anchor evaluation in Figure 9 reports only mean answers for one model (GPT-4), without standard deviations, error bars, or a significance test against the No_anchor condition. The conclusion that Both-Anchor partially mitigates anchoring bias is therefore not quantitatively established.","section":"Section 5"}],"minor_comments":[{"comment":"The text refers to 'Control and Treatment B & C' but only Treatments A and B are defined; the Control condition is denoted C in Section 3.2. This is likely a typo and should be corrected.","section":"Section 3.2"},{"comment":"Many p-values are reported as 0.00E+00 or with excessive precision. The actual floating-point values should be reported, and the threshold of 0.05 should be stated in the text; the authors should also address the multiple-comparisons issue arising from running many t-tests.","section":"Section 4.2 / Figures 4–6"},{"comment":"The sentence 'We use the \"expert\" anchoring questions because the anchoring effects are very significant for these questions. [19]' cites a reference about human anchoring and adjustment (Simmons et al.) rather than the authors' own analysis. The citation is misplaced; the authors should either cite their own Figures 4–5 or provide a different rationale.","section":"Section 5"},{"comment":"The prompt template explicitly instructs the model to make an 'educated guess based on the provided information.' This instruction likely encourages the model to incorporate the hints, making it difficult to attribute the observed sensitivity to an automatic cognitive bias. The paper should discuss this potential confound or include a control prompt that does not instruct the model to use the hints.","section":"Section 3.2 / Appendix A"},{"comment":"The column labeled 'No_anchor' is not defined in the body text. It should be explained how this condition was constructed (e.g., whether it is the Control condition from Section 3.2, or a separate no-hint condition) and why it is used only in the mitigation section.","section":"Figure 9"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely topic and uses an interesting existing human dataset, but the central measurement issue is substantial: the results as presented establish sensitivity to hints, not anchoring bias. The authors would need to reanalyze the data with a proper baseline or normative model, add uncertainty measures to the AI comparison, and strengthen the mitigation analysis. The paper may be suitable for publication after such revision, but the current version does not justify its main claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this paper is a real empirical step beyond the 5-question financial anchoring study. It uses the 62-question PTF dataset, runs GPT-4, GPT-4o, and GPT-3.5 with 30 repetitions per condition, and tests four simple mitigation prompts. The core finding that survives scrutiny is that LLM answers are consistently sensitive to the values of supplied hints, and that CoT, PoT, \"ignore the hint,\" and reflection do not remove that sensitivity. The both-anchor strategy, where both H2 and H3 are included, is the genuinely new piece: when the two hints straddle the no-anchor answer, the model lands near the no-anchor value. That is a practical, falsifiable result, and the authors share prompts and raw answers.\n\nBut the stress-test note lands. Section 4.2 defines bias as a significant sign-concordant difference between treatment A and B. The t-test only shows the two distributions differ; it does not show the model is biased. H2 and H3 are informative in the fact and expert conditions—last week's low/high, another forecaster's estimate—so a model that simply weights the hint as additional evidence would produce exactly the same signed difference. There is no no-anchor baseline in the main analysis. Figure 9 includes a No_anchor column, but it is used only for the both-anchor evaluation, not as a baseline for the bias test. So the central claim—\"anchoring bias\"—is not established by the reported statistics. What is established is the weaker, still noteworthy claim: LLM outputs move with the direction of informative hints, and simple prompts don't stop that.\n\nSmaller concerns: the mitigation experiments are run only on the 12 expert questions with the strongest effects, so \"none of these strategies work\" is a claim about the strongest cases, not the full dataset. The both-anchor evaluation reports means without error bars or a formal comparison, making it hard to judge significance. Some p-values are printed as 0.00E+00, which is not a real p-value. None of these are fatal; the main issue is the interpretive gap.\n\nWho this is for: practitioners who want evidence that LLM numeric predictions are manipulable by context, and researchers working on debiasing prompts. As a study of anchoring bias specifically, it overreaches. As a study of hint sensitivity and a candidate mitigation, it is useful.\n\nI would send it to peer review, with the expectation that the analysis is reframed around sensitivity rather than bias, and the no-anchor baseline is properly integrated. My own verdict is skeptical on the central label but positive on the empirical core.","headline":"Useful empirical study of hint sensitivity in LLMs with a promising mitigation idea, but the main test does not separate anchoring bias from rational use of informative hints.","tokens_in":16838,"tokens_out":3447,"would_cite":false,"duration_ms":34925,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs let a planted number steer their numerical guesses.","keywords":["anchoring bias","large language models","cognitive bias","prompt mitigation strategies","GPT-4","GPT-4o","GPT-3.5","numerical prediction"],"falsifier":"Give the same prompts to a model with the hint labels swapped (high anchor presented where the low anchor was, and vice versa) and check whether the answer difference flips sign exactly as a function of the presented values; if it tracks the presented values symmetrically, the sign-match test cannot distinguish anchoring from arithmetic averaging. A cleaner check would add a per-question no-anchor condition and require answers to move toward the anchor beyond the no-anchor baseline.","tokens_in":15879,"feed_emoji":"🎯","tokens_out":6005,"duration_ms":55106,"temperature":0.7,"pith_summary":"This paper asks whether large language models fall for anchoring bias—the tendency to let the first number they see drag their estimate toward it. On 62 everyday prediction questions taken from a human behavioral study, the authors compare each model's numerical answers when given a low anchor versus a high anchor alongside the question. They find that GPT-4, GPT-4o, and GPT-3.5 answer higher when the hint number is higher, with same-signed differences in 151 of 162 model-question comparisons, and that the effect is especially strong for 'expert' hints. They also test four common prompt-level fixes—chain-of-thought, principles-first reasoning, explicit instructions to ignore the anchor, and reflection—and find none of them removes the bias. The only partial remedy is to give the model both anchors at once, which pulls answers back toward the unbiased benchmark when the two anchors bracket it.","feed_headline":"LLMs let a planted number steer their numerical guesses","feed_subtitle":"GPT-4, GPT-4o, and GPT-3.5 all shift numerical guesses toward biased hints; simple prompts don't fix it.","key_machinery":"The operative measure is a sign-match test. For each question, the model answers 30 times with the low anchor H2 (treatment A) and 30 times with the high anchor H3 (treatment B); the paper computes the mean difference (B − A) and compares its sign with the sign of (H3 − H2), using a two-sample t-test for significance. A same-signed, statistically significant difference is taken as evidence of anchoring. Supporting this test are a CO-STAR-based prompt template that asks for an 'educated guess' in JSON format, and the anchoring index AI = (median_high − median_low)/(high anchor − low anchor), borrowed from the human anchoring literature to compare LLM and human susceptibility.","core_discovery":"The paper's central discovery is that GPT-family models exhibit anchoring bias in numeric prediction tasks: answers under a high hint are significantly higher than answers under a low hint, and the sign of the answer difference matches the sign of the hint difference in the large majority of cases. This holds across fact-based and expert-opinion anchors, with expert hints producing the strongest and most consistent effect—none of the 12 expert-hint questions produced a sign mismatch for any model. The bias survives four prompting strategies designed to reduce it, and only the Both-Anchor condition, where both the low and high hints are included in the prompt, moves answers close to the no-anchor baseline, and only when the two hints lie on opposite sides of that baseline. The authors also report an anchoring index near 0.45 for GPT-4, lower than the 0.61 observed in the human study on the same questions, indicating LLMs are influenced by anchors but less so than humans.","pith_inferences":["The sign-match test would classify a rational interpolator—a model that simply averages the two hint values—as 'anchored,' so the true effect size is likely overstated until a no-anchor baseline is subtracted.","A direct prediction from the paper: reordering hints or placing a neutral number before the anchor should change the strength of the effect if anchoring is about the first salient number; this ordering test is not run and would separate anchoring from generic numeric context sensitivity.","The Both-Anchor result suggests a cheap deployment rule: when eliciting numeric estimates from LLMs, prompt with several bracket-style estimates rather than a single reference value, since the paper shows this reduces anchor pull even though it does not eliminate it.","If stronger models are more consistently anchored, calibration efforts that reduce output variance may inadvertently amplify the appearance of bias in sign-match tests; reporting absolute deviations from a no-anchor baseline would clarify whether variance or true pull drives the result."],"forward_implications":["Numerical answers from GPT-4, GPT-4o, and GPT-3.5 move in the direction of an anchoring hint, and the effect is strongest for 'expert' opinions, where all 12 test questions showed same-signed differences.","Prompting strategies that work for reasoning tasks in general—Chain-of-Thought, Thoughts of Principles, explicit ignore-anchor instructions, and Reflection—do not remove or significantly reduce the anchoring effect.","Presenting both a low and a high anchor together moves GPT-4's answers close to the no-anchor value when the two anchors lie on opposite sides of that value, suggesting a practical mitigation path.","Stronger models (GPT-4, GPT-4o) are more consistently biased than GPT-3.5, which the authors attribute to lower answer variance in stronger models.","Anchoring effects were not statistically significant for time-period answers and for questions whose first hint was irrelevant to the question."],"supporting_citations":[{"why":"Supplies the 62-question experimental dataset and the human user-study results used for all experiments and comparisons.","marker":"[10]"},{"why":"Prior study of anchoring in LLMs on five financial questions; the baseline this study extends, and the source of the ignore-anchor strategy.","marker":"[9]"},{"why":"Supplies the Chain-of-Thought prompting strategy tested as a mitigation.","marker":"[20]"},{"why":"Supplies the Thoughts-of-Principles (take-a-step-back) strategy tested as a mitigation.","marker":"[21]"},{"why":"Human experimental result showing anchoring is hard to eliminate completely; used to contextualize the failure of mitigation.","marker":"[19]"},{"why":"Supplies the anchoring index formula used to compare LLM anchoring strength with human anchoring strength.","marker":"[22]"},{"why":"Supplies the CO-STAR prompt format on which the paper's prompt template is based.","marker":"[18]"}],"fun_headline_variants":["LLMs mirror human anchoring bias on numeric guesses","Anchoring bias persists in GPT-4 despite prompt fixes","Expert hints anchor LLM answers; simple prompts fail","Only contrasting anchors reduce LLM anchoring bias","GPT models show anchoring bias, less than humans do"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper counts an answer as anchored when the difference between answers under the high and low hints has the same sign as the difference between the two hints; that test would also fire for a model that simply averaged the two hints, because no no-anchor baseline is built into the main comparison.","fun_headline_variants_meta":{"raw":{"variants":["LLMs mirror human anchoring bias on numeric guesses","Anchoring bias persists in GPT-4 despite prompt fixes","Expert hints anchor LLM answers; simple prompts fail","Only contrasting anchors reduce LLM anchoring bias","GPT models show anchoring bias, less than humans do"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1468,"prompt_tokens":896,"completion_tokens":572,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":497}},"tokens_in":512,"tokens_out":572,"duration_ms":5870,"temperature":1.0,"reasoning_tokens":497,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:28:57.302435+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the same prompts to a model with the hint labels swapped (high anchor presented where the low anchor was, and vice versa) and check whether the answer difference flips sign exactly as a function of the presented values; if it tracks the presented values symmetrically, the sign-match test cannot distinguish anchoring from arithmetic averaging. A cleaner check would add a per-question no-anchor condition and require answers to move toward the anchor beyond the no-anchor baseline.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 62-question experimental dataset and the human user-study results used for all experiments and comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior study of anchoring in LLMs on five financial questions; the baseline this study extends, and the source of the ignore-anchor strategy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Chain-of-Thought prompting strategy tested as a mitigation."},{"cited_title":"Chi, Quoc V Le, Denny Zhou, Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models, ICLR 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the Thoughts-of-Principles (take-a-step-back) strategy tested as a mitigation."},{"cited_title":"Th e effect of accuracy motivation on anchoring and adjustment: Do people adjust from provided anchors?","cited_arxiv_id":null,"evidence_quote":"Human experimental result showing anchoring is hard to eliminate completely; used to contextualize the failure of mitigation."},{"cited_title":"Jacwitz, D","cited_arxiv_id":null,"evidence_quote":"Supplies the anchoring index formula used to compare LLM anchoring strength with human anchoring strength."},{"cited_title":"Teo, How I won Singapore’s gpt-4 prompt engineering c ompetition: A deep dive into the strategies I learned for harnessing the power of large language models (llms), 2023","cited_arxiv_id":null,"evidence_quote":"Supplies the CO-STAR prompt format on which the paper's prompt template is based."}],"review_version":1}