{"id":"1ed35d86-6a6f-45e4-a378-fc9363b61064","arxiv_id":"2412.11009","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LLMs flip between correct Bayesian updating and similarity-based representativeness heuristics depending on the structure of the probability question.","lead":"Three experiments show that large language models sometimes reason like good statisticians and sometimes just match a story to a stereotype, depending on how the question is phrased. This matters because it clarifies when AI assistants are likely to ignore base rates, for example in medical diagnosis or financial decisions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The unstructured test's non-diagnosticity assumption is validated only by D values elicited from the same LLMs, whose inverse probabilities correlate ~0.9 with similarity; an external diagnosticity check is required.","rationale":"The reader's conditional verdict is the right one. The semi-structured test is internally sturdy: the odds-ratio benchmark of 9 is independent of D, and the word-level 'indicate'→'compute' reversal (Table A.4) is a clean within-model control. The structured test, while not central to the dual-mode claim by itself, is supported by high accuracy and by questions that were modified from standard textbook items. The weak point is the unstructured test. Its inference that a negative prior–posterior correlation constitutes base-rate neglect depends on the Adam sketch being non-diagnostic (D≈1). The paper supports that premise with D values computed from the same five models' inverse-probability judgments (Table A.2), yet those judgments track similarity almost perfectly (Table A.3: inverse-vs-similarity r≈0.92–0.98). That is not an independent benchmark; it is the representativeness process itself. The numerical inconsistency strengthens the concern: for gpt-4o's one-field rotation, the observed posterior odds require D(A)/D(C)≈15 while the reported D ratio is about 0.73. If an external diagnosticity measure shows the sketch is strongly diagnostic for field A, then the unstructured test's conclusions collapse into a memory/recall story rather than a representative-mode story. Even if the external measure finds D≈1, the paper still needs to report uncertainty around the D values and odds ratios before the unstructured result can be taken quantitatively. For these reasons I would keep the reader's CONDITIONAL verdict rather than accept or reject, and the most direct fix is the external diagnosticity pretest described in concrete_test.","tokens_in":25388,"tokens_out":12380,"duration_ms":114370,"concrete_test":"Run a pre-registered external diagnosticity pretest: have an independent human sample (or, if human data are unavailable, a held-out open-weight LLM not among the five tested) rate the likelihood P(E|H) of the Adam sketch for each field in a between-subjects design with no base-rate information and no posterior question. Aggregate the ratings into D_ext(A)/D_ext(C). If the external D ratio is close to 1 (within a factor of 2 of Table A.2), the unstructured test's non-diagnosticity assumption is externally validated. If D_ext(A)/D_ext(C) ≈ 5 or larger, then the observed posterior odds ratio is normatively plausible, and the unstructured test cannot establish base-rate neglect; the dual-mode claim would then have to rest on the semi-structured test alone.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The unstructured test's inference that normative posteriors should track priors rests entirely on the claim that the Adam sketch is non-diagnostic (D≈1). This is established only by D(H,E) values obtained from the same five LLMs' inverse-probability answers (Extended Data Table A.2). But Table A.3 shows those inverse probabilities correlate almost perfectly with similarity (r≈0.92–0.98 for inverse-vs-similarity), so the 'non-diagnosticity' benchmark is generated by the same representativeness mechanism the experiment is meant to detect. The quantitative gap is large: for gpt-4o in the one-field rotation, prior odds P(A)/P(C)=0.02/0.13=0.15 while posterior odds P(A|E)/P(C|E)=0.44/0.19=2.3, requiring D(A)/D(C)≈15; Table A.2 reports D(E,A)=1.09 and D(E,C)=1.50, a ratio of 0.73. If the sketch is actually strongly diagnostic for field A, the negative prior–posterior correlation is not base-rate neglect, and the representative-mode evidence from the unstructured test loses its normative contrast.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports three experiments on how large language models (LLMs) judge posterior probabilities P(H|E). In the structured test, where base rates and likelihoods are fully supplied, state-of-the-art LLMs solve textbook Bayesian problems at high accuracy. In the semi-structured test, only base rates are supplied; for a description high in representativeness, all tested LLMs produce posterior odds near 1 instead of the normative odds ratio of 9, while for a non-representative description they respond to the base-rate manipulation. Replacing the word 'indicate' by 'compute' in the prompt restores a near-9 odds ratio for GPT-4o. In the unstructured test, with no Bayesian information supplied, posterior judgments correlate negatively with prior probabilities and positively with similarity judgments; the authors infer base-rate neglect by using the same models' inverse-probability estimates to argue that the stimulus is non-diagnostic. The paper concludes that LLMs exhibit coexisting normative and representativeness-based modes and conjectures that this duality stems from the contrastive-style loss used in reinforcement learning from human feedback.","tokens_in":25586,"tokens_out":8694,"duration_ms":76972,"significance":"The semi-structured test is a clean, externally normed demonstration that LLM posterior judgments can be insensitive to base rates while being strongly driven by representativeness, and the prompt-reversal control shows that a single wording change can shift behavior toward Bayesian combination. If this result holds, the paper's central 'dual modes' claim would be an important advance over simple bias/no-bias characterizations. The authors also make code and data available and use repeated sampling with predetermined questioning rounds to stabilize estimates, which strengthens the empirical contribution. The main limitation is that the unstructured test's normative benchmark is derived from the same models under test, whose likelihood judgments correlate almost perfectly with similarity; that test therefore does not independently establish base-rate neglect. This tempers the breadth of the claims, but the core semi-structured evidence remains defensible.","major_comments":[{"comment":"The claim that normative posteriors in the unstructured test should track priors rests on the assertion that the Adam sketch is non-diagnostic (D≈1). This assertion is validated only by D(H,E) values in Table A.2, which are computed from P(E|H) and P(E|¬H) elicited from the same five LLMs. Table A.3 shows these inverse probabilities correlate almost perfectly with similarity (Pearson 0.92–0.98), so the 'non-diagnosticity' benchmark is generated by the same representativeness mechanism the experiment is designed to detect. The quantitative gap is large: for gpt-4o in the single-field rotation, the prior odds P(A)/P(C)=0.02/0.13≈0.15 while the posterior odds P(A|E)/P(C|E)=0.44/0.19≈2.3, requiring D(A)/D(C)≈15 to be normatively consistent; Table A.2 reports D(E,A)=1.09 and D(E,C)=1.50, a ratio of 0.73. If the sketch is actually diagnostic for field A, the negative prior–posterior correlation is not evidence of base-rate neglect. Please add an external diagnosticity check (e.g., human likelihood ratings) or reframe the unstructured-test conclusions as descriptive evidence of similarity-based judgment without claiming a normative violation.","section":"Unstructured Test, Extended Data Tables A.2 and A.3"},{"comment":"The correlation evidence in the unstructured test is based on only 7 rotations per model. With n=7, a two-sided p-value below 0.001 is highly dependent on the exact permutation set, and the current footnote 'two-sided Fisher's p-value' is ambiguous because Fisher's z-transformation is monotonic in r and does not by itself define a nonparametric test for n=7. Please report the exact permutation p-values or bootstrap confidence intervals for the correlations in Tables 2 and A.5, and state how many permutations were used. The small sample also means the 'prior vs. posterior' correlation could be driven by a single rotation; please report the raw means for all seven rotations or a leave-one-rotation-out sensitivity analysis.","section":"Methods, Rotational Design; Table 2"},{"comment":"The striking finding that replacing 'indicate' with 'compute' reverses base-rate neglect is demonstrated only for GPT-4o. The semi-structured test's across-model claim is established, but the causal 'single word can trigger the normative mode' claim should be either replicated on at least a few other state-of-the-art models or explicitly qualified as a single-model demonstration. As written, the Discussion and Abstract generalize this prompt-reversal control beyond the evidence.","section":"Prompt Engineering; Extended Data Table A.4; SI Section 2.4"}],"minor_comments":[{"comment":"The accuracy interval (0.40, 0.43) for Q7, Q8, and Q10–Q12 is sensible given the exact Bayesian answer of about 0.4138, but please state explicitly that this interval is an analytic rounding convention rather than a post hoc fitted tolerance, and note whether it was fixed before data collection.","section":"Structured Test, SI Section 2.1"},{"comment":"The sentence introducing the simultaneous confidence intervals contains '[see, e.g. ? , §4.4]' with a missing citation; please complete or remove the reference.","section":"Supplementary Information Section 3"},{"comment":"The human comparison column is based on different base rates (0.7/0.3, with the note stating a normative ratio of about 5.44), so the direct comparison with the LLM rows' 9:1 normative benchmark is not clean; please present a normalized measure such as the ratio of posterior odds divided by the prior odds so that human and LLM rows are on the same scale.","section":"Table 1"},{"comment":"There are several typographical and formatting issues, including 'T able' spacing throughout, 'Inverse Probabiliy' in SI Table A.2 headers, and 'It has been showed' in the Discussion (should be 'It has been shown').","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The semi-structured experiment is the strongest contribution and could support a version of the central claim on its own; the unstructured test is currently overinterpreted because its normative benchmark is partially circular. The requested external diagnosticity check or a careful reframing is feasible within the manuscript's scope and would make the paper publishable. The RLHF conjecture is speculative but explicitly labeled as such. The paper is a reasonable empirical contribution for an AI or cognitive-science venue, provided the unstructured-test claims are tightened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something genuinely useful: it tests LLMs on posterior probability judgment with three levels of information, and finds a shift from near-normative Bayesian behavior (when all numbers are given) to representativeness-driven behavior (when the prompt omits diagnosticity and pushes toward a stereotype). The semi-structured 'Jason' test is the strongest piece — a fixed 9:1 odds ratio benchmark that makes the prediction unambiguous, and a nice control where LLMs mostly track base rates on an uncharacteristic description but ignore them on representative ones. The single-word prompt reversal ('indicate' → 'compute') is a clean and clear demonstration that the bias is not hard-wired. Those results alone justify the paper's attention.\n\nThe unstructured test is where I'd push back. The claim that the Adam sketch is non-diagnostic rests on D values obtained from the same five LLMs whose inverse probabilities correlate ~0.9 with similarity. That is partially circular: if the D judgments themselves are products of the representativeness mechanism, the benchmark is biased. The stress-test note quantifies the issue for gpt-4o: the observed posterior odds require a D(A)/D(C) ratio around 15, but the elicited D values give roughly 0.73. That does not kill the paper, because the semi-structured test already shows base-rate neglect in representative contexts, but the unstructured test's negative prior–posterior correlation loses its clean normative contrast. The small number of rotations (7 per model) and the missing error bars on the odds ratios only add to that concern.\n\nThe RLHF contrastive-loss conjecture is clearly labeled as speculative, and the paper ships code and data, which I count in its favor. The writing is straightforward and the citations to the base-rate literature are appropriate.\n\nBottom line: the central dual-mode finding is plausible and worth building on, but the unstructured test needs an external diagnosticity check — ideally human or independent-likelihood judgments — and fuller uncertainty reporting. I'd send this to peer review; it deserves a serious referee, with the expectation of revision focused on the unstructured test.","headline":"A strong semi-structured experiment supports the dual-mode claim, but the unstructured test's circular diagnosticity benchmark needs external validation.","tokens_in":26117,"tokens_out":2870,"would_cite":true,"duration_ms":25800,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"State-of-the-art LLMs judge probabilities by Bayes' rule or by similarity, depending on the prompt.","keywords":["large language models","Bayesian reasoning","representativeness heuristic","base-rate neglect","posterior probability","similarity judgment","prompt engineering","reinforcement learning from human feedback"],"falsifier":"Ask independent human judges to estimate $P(E|H)$ and $P(E|\\neg H)$ for the Adam sketch across the three fields, and compare those externally elicited diagnosticity values with the models' values. If the external $D(H,E)$ values are near 1 while the LLMs' posteriors still track similarity and move away from priors, the dual-mode and base-rate-neglect claims for the unstructured test are confirmed; if the external $D(H,E)$ values are large, the experiment cannot distinguish correct diagnosticity-based updating from representativeness.","tokens_in":25156,"feed_emoji":"🧠","tokens_out":9271,"duration_ms":81277,"temperature":0.7,"pith_summary":"Large language models do not sit at a single point on a biased-versus-unbiased scale. This paper claims that their posterior judgments $P(H|E)$ are produced by two coexisting modes: a normative mode that follows Bayes' rule and a representativeness mode that answers by how similar the evidence is to a typical member of the class. In a test that varied only the base rate, state-of-the-art models produced normative posterior-odds ratios near 9 for a neutral description but near 1 for representative descriptions, meaning the base rate was ignored exactly when the evidence was stereotypical. In a second test with no given base rate or diagnosticity, judged posteriors correlated strongly with similarity (Pearson $r \\approx 0.9$) and negatively with base rates, even though the models' own inverse-probability judgments implied the evidence was nearly non-diagnostic. The paper also shows that a single wording change can switch representative judgments to Bayesian ones, and conjectures that the dual trait comes from contrastive losses in reinforcement learning from human feedback.","feed_headline":"LLMs flip between Bayesian and similarity reasoning","feed_subtitle":"State-of-the-art models follow Bayes' rule when told to compute, but otherwise fall back on similarity, echoing human System 1 and System 2.","key_machinery":"The load-bearing object is a dual-mode posterior function $f(H,E)$ whose normative component $f_{\\mathrm{norm}}$ follows Bayes' rule and whose representativeness component $f_{\\mathrm{rep}}$ equates posterior probability with similarity. The key identity that separates the two modes is the base-rate odds ratio: because evidence diagnosticity is independent of the base rate, the ratio of posterior odds across the 75% and 25% base-rate conditions must be exactly 9 under $f_{\\mathrm{norm}}$, and drifts toward 1 as $f_{\\mathrm{rep}}$ takes over. In the unstructured test, the auxiliary object is diagnosticity $D(H,E)=P(E|H)/P(E|\\neg H)$ computed from the model's own inverse-probability judgments; values near 1 license the benchmark that normative posteriors equal priors, so a posterior that tracks similarity rather than priors implicates $f_{\\mathrm{rep}}$. The paper's conjectured origin mechanism is the binary ranking loss used in reward-model training, which the authors read as a contrastive loss that amplifies distinguishing features while ignoring class frequencies.","core_discovery":"The paper's central claim is that posterior judgment in state-of-the-art LLMs is driven by two distinct prediction functions: $f_{\\mathrm{norm}}$, which applies Bayes' rule by combining base rates with evidence diagnosticity, and $f_{\\mathrm{rep}}$, which maps similarity between the evidence and a class prototype directly to probability. The coexistence of these modes is established with an odds-ratio design: when the base rate is 75% versus 25%, any normative judge must give posterior-odds ratio $O(B_h)/O(B_\\ell)=9$ regardless of the description, whereas representative descriptions push all tested models to ratios near 1 even though the base rates are stated in the prompt. In the unstructured test, where the Adam sketch is judged across three fields with different priors, posteriors are negatively correlated with base rates and positively correlated with similarity, while the models' own inverse-probability estimates imply diagnosticity $D(H,E)$ near 1, so normative posteriors should track priors. The paper further shows that prompting a model to 'compute' rather than 'indicate' restores Bayesian odds ratios in the semi-structured setting, but that iterative Bayes-rule prompting in the unstructured setting does not eliminate similarity-driven judgment, and that models recall base rates from memory poorly.","pith_inferences":["Because the non-diagnosticity of the Adam sketch is established with the same models whose similarity-driven reasoning is under investigation, the unstructured test's base-rate-neglect result should be read as conditional; an externally validated set of inverse probabilities would settle whether the benchmark is contaminated.","The odds-ratio test used in the semi-structured experiment is cheap and model-agnostic: ask any model for posteriors under 75% and 25% base rates and compare the ratio of posterior odds to 9, with no inverse probabilities needed.","If the contrastive-loss conjecture is correct, it predicts a testable ordering: models trained with preference-optimization losses that ignore class frequencies should show lower odds ratios on representative descriptions than models trained on next-token prediction or on objectives that include base rates.","Wording sensitivity implies that mode selection may be triggered by token-level cues that invoke arithmetic routines, so interventions that force explicit step-by-step calculation before output may generalize better across prompts than substituting a single verb."],"forward_implications":["LLMs can compute normatively when all ingredients of Bayes' rule are present or explicitly cued, so poor accuracy on unconstrained questions does not by itself demonstrate absence of normative competence.","Base-rate neglect is mode-dependent rather than universal: representative descriptions suppress sensitivity to base rates even when the base rates are stated in the prompt.","A minimal wording change from 'indicate' to 'compute' can flip a model from the representativeness mode to the normative mode in the semi-structured setting, making surface prompt features a strong lever on mode selection.","In unstructured settings, prompting the model to apply Bayes' rule does not eliminate similarity-driven posterior judgments, and models recall base rates from memory poorly, so deployed systems should supply priors explicitly.","If the paper's conjecture about contrastive loss is right, the reward-model stage of reinforcement learning from human feedback may itself entrench representativeness, meaning bias mitigation may have to alter training objectives rather than only prompts."],"supporting_citations":[{"why":"Defines subjective probability as representativeness judgment and supplies the normative-versus-representativeness distinction that motivates the two modes.","marker":"[8]"},{"why":"Establishes the representativeness heuristic and base-rate neglect as judgment phenomena the experiments are designed to detect in LLMs.","marker":"[20]"},{"why":"Documents the base-rate fallacy in human probability judgments and provides problem templates adapted for the structured and semi-structured tests.","marker":"[2]"},{"why":"Supplies the lawyer-engineer style causal-schema problems that the Jason descriptions in the semi-structured test are modeled on.","marker":"[21]"},{"why":"Provides human posterior judgments for the base-rate manipulation that the paper compares against in Table 1.","marker":"[9]"},{"why":"Summarizes replications of the base-rate manipulation used to compute the human-average column in the semi-structured test.","marker":"[11]"},{"why":"Gives the theoretical account of contrastive learning cited to argue that the reward-model loss amplifies distinguishing features.","marker":"[1]"},{"why":"Establishes supervised contrastive learning as the framework the paper compares the binary ranking loss to in its conjecture.","marker":"[10]"},{"why":"Shows that preference optimization has contrastive gradient behavior, linking reward-model training to representativeness amplification.","marker":"[15]"},{"why":"Documents the use of direct preference optimization in the Llama 3 family, connecting the conjecture to an actual reinforcement-learning pipeline.","marker":"[5]"}],"fun_headline_variants":["LLMs mix Bayes' rule and similarity in probabilistic reasoning","LLMs show dual modes: normative and representative judgment","Bayes and similarity coexist in LLM posterior estimates","LLMs rely on both Bayes' rule and similarity, like humans","LLMs alternate between Bayesian and similarity-based reasoning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The unstructured test's conclusion that normative posteriors should equal priors rests on the premise that the Adam description is non-diagnostic, and that premise is supported only by diagnosticity values computed from the same LLMs' inverse-probability judgments, which themselves correlate almost perfectly with similarity.","fun_headline_variants_meta":{"raw":{"variants":["LLMs mix Bayes' rule and similarity in probabilistic reasoning","LLMs show dual modes: normative and representative judgment","Bayes and similarity coexist in LLM posterior estimates","LLMs rely on both Bayes' rule and similarity, like humans","LLMs alternate between Bayesian and similarity-based reasoning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000195,"raw_usage":{"total_tokens":1351,"prompt_tokens":935,"completion_tokens":416,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":337}},"tokens_in":551,"tokens_out":416,"duration_ms":4161,"temperature":1.0,"reasoning_tokens":337,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:23:32.772828+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask independent human judges to estimate $P(E|H)$ and $P(E|\\neg H)$ for the Adam sketch across the three fields, and compare those externally elicited diagnosticity values with the models' values. If the external $D(H,E)$ values are near 1 while the LLMs' posteriors still track similarity and move away from priors, the dual-mode and base-rate-neglect claims for the unstructured test are confirmed; if the external $D(H,E)$ values are large, the experiment cannot distinguish correct diagnosticity-based updating from representativeness.","supporting_citations":[{"cited_title":"Acta Psychologica 44(3):211–233, ISSN 0001-6918, URL http://dx.doi.org/https://doi.org/10.1016/0001-6918(80)90046-3","cited_arxiv_id":null,"evidence_quote":"Documents the base-rate fallacy in human probability judgments and provides problem templates adapted for the structured and semi-structured tests."},{"cited_title":"Kahneman D, Slovic P , Tversky A, eds., Judgment under Uncertainty: Heuristics and Biases , 117–128 (Cambridge University Press)","cited_arxiv_id":null,"evidence_quote":"Supplies the lawyer-engineer style causal-schema problems that the Jason descriptions in the semi-structured test are modeled on."},{"cited_title":"Be- havioral and Brain Sciences 19(1):1–17, URL http://dx.doi.org/10.1017/S0140525X00041157","cited_arxiv_id":null,"evidence_quote":"Summarizes replications of the base-rate manipulation used to compute the human-average column in the semi-structured test."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes supervised contrastive learning as the framework the paper compares the binary ranking loss to in its conjecture."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that preference optimization has contrastive gradient behavior, linking reward-model training to representativeness amplification."}],"review_version":1}