{"id":"e482da6b-5b03-4665-a4b7-90fd2583f832","arxiv_id":"2508.06671","paper_version":2,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":3.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The manuscript is internally inconsistent: the abstract describes an LLM fairness experiment while the body is a different paper on pilot-wave quantum mechanics, so no coherent result can be assessed.","lead":"This submission's abstract claims that language models can give biased answers without showing corresponding bias in their chain-of-thought reasoning. The supplied full text is an unrelated quantum foundations paper, so the claimed experiment is not present in the manuscript.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Supplied full text is an unrelated pilot-wave QM paper; no LLM experiment, metrics, or correlation results appear to support the abstract's central claim.","rationale":"The stress-test confirms the reader's UNVERDICTED assessment, but locates the load-bearing concern slightly differently. The reader's weakest_assumption is the CoT-as-thought mapping; that is a deep validity threat for any such study. However, the more immediate blocker is that the supplied manuscript body is a different paper. The abstract's quantitative claim is the sole evidence; there is no protocol, no model list, no bias definitions, no p-value computation. Consequently correctness risk is unknown in the strong sense: not because the conclusion might fail, but because there is no argument to evaluate. This is consistent with the reader's verdict. I do not see a second independent concern worth prioritizing before this one is resolved, because resolving the mismatch is a precondition for any substantive statistical or conceptual critique. If the correct full text were retrieved and did contain the experiment, the next checkpoint would be the CoT faithfulness assumption and the choice of fairness metric. But for the submission as reviewed, no such checkpoint is reachable. The reader explicitly noted the body does not contain the experiment, so we agree on the operative finding, though our headline emphasizes content mismatch rather than the thought-text mapping.","tokens_in":4572,"tokens_out":2918,"duration_ms":30641,"concrete_test":"Retrieve the actual arXiv full text for 2508.06671 (not 2508.06667) and search for sections describing the LLM experiments: model names/versions, the fairness metric used on chain-of-thought text, the output fairness metric, and the table of 11 bias correlations. If any of these are absent, the verdict remains UNVERDICTED because the claimed result is uncheckable. If all are present, recompute one headline correlation (e.g., gender bias) from the provided raw outputs using the stated metric; a match would resolve the mismatch, while a deviating value would indicate a reporting error.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that across 5 LLMs and 11 bias types, chain-of-thought bias correlates with output bias at <0.6 with p<0.001. For this to be assessable, the manuscript must define the models, the fairness metric applied to CoT text, the output bias measure, prompt/decoding settings, and the correlation computation. The supplied full text instead is \"A disputable assumption behind the empirical equivalence between pilot-wave theory and standard quantum mechanics\" by Manero, Muciño, and Okon, arXiv:2508.06667v1. It contains no language models, no chain-of-thought, no bias metric, and no correlation table. The abstract alone is not an experiment. The additional interpretive claim—that models with biased decisions \"do not always possess biased thoughts\"—further assumes CoT text is a faithful record of internal reasoning, an assumption that would require its own defense even if the experiment were present. No self-reported limitation or appendix supplies this. The load-bearing issue is evidentiary absence, not an internal contradiction: no part of the submitted body bears on the headline result.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript, as identified by its arXiv metadata and abstract, claims to present an empirical study of chain-of-thought (CoT) bias versus output bias in five large language models across eleven bias dimensions, reporting correlations below 0.6 with p<0.001 in most cases and concluding that biased model decisions are not accompanied by correspondingly biased CoT reasoning. However, the supplied full text is a quantum-foundations paper on pilot-wave theory and standard quantum mechanics (Manero, Muciño, and Okon, arXiv:2508.06667v1), containing no language models, no chain-of-thought prompting, no fairness metrics, no datasets, no experimental protocol, and no results tables. The central claim of the abstract is therefore unsupported by any body text.","tokens_in":4761,"tokens_out":1252,"duration_ms":15613,"significance":"If the claimed correlation result were properly established, it would be significant for fairness auditing and interpretability research: it would suggest that CoT text may not be a reliable proxy for the fairness-relevant internal processes that produce final model outputs, and it would complicate the use of CoT-based fairness metrics. The paper would be potentially important to the NLP community. It is also the kind of negative result that can be valuable, since it challenges a common assumption that a model's reasoning trace reflects the same biases as its final answer. However, none of this significance can be realized on the basis of the submitted manuscript, because the experimental evidence is entirely absent. I also note that the work provides no machine-checked proofs, reproducible code, or parameter-free derivations that could partially offset the absent experimental detail.","major_comments":[{"comment":"The abstract's central claim—that across 5 LLMs and 11 bias types, CoT bias correlates with output bias at less than 0.6 with p<0.001—has no supporting experimental material anywhere in the supplied full text. The body is a paper on pilot-wave quantum mechanics. No model names, datasets, prompts, decoding settings, fairness metrics, correlation formulas, or significance procedures are given. This is not a presentation issue; it is a complete evidentiary absence, and the headline result cannot be assessed, reproduced, or even located in the manuscript.","section":"Abstract and full text"},{"comment":"The interpretive claim that models with biased decisions 'do not always possess biased thoughts' requires equating CoT text with the model's thoughts. The manuscript does not define this mapping or defend it, and it does not discuss the alternative reading that CoT text is an output artifact whose bias may be only weakly related to the model's internal computations. Even if the experiment were present, this assumption would need explicit operationalization and validation; its absence is load-bearing because it determines whether the reported correlation addresses the stated research question.","section":"Abstract, 'biased thoughts'"},{"comment":"The manuscript contains no limitations section, no reproducibility details, and no appendices supplying experimental protocol. A reader cannot determine how the 11 biases were defined, how fairness metrics were computed on CoT tokens versus final answers, whether the correlations were across items, models, or bias categories, or how the p-values were obtained. These omissions are not local gaps; they concern the entire empirical contribution and cannot be repaired without adding the missing study.","section":"Entire manuscript"}],"minor_comments":[{"comment":"The arXiv identifier in the supplied full text is 2508.06667, while the manuscript under review is 2508.06671; this mismatch should be resolved editorially.","section":"Metadata"},{"comment":"The body text contains typographical artifacts (e.g., 'anN-particle', broken spacing in Section 4) that are immaterial given the substantive mismatch but would need correction in any eventual publication.","section":"Full text"}],"recommendation":"reject","confidential_remarks":"The supplied file appears to be an unrelated paper from the same arXiv batch. The discrepancy is so complete that the manuscript cannot be considered a coherent submission: the abstract advertises an NLP empirical study and the body is a quantum foundations paper. If this reflects a submission error, the editor may wish to contact the authors for the correct file, but based on the reviewed material the claim is unsupported and the recommendation is reject."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The abstract asks a real question: if a language model produces biased final answers, are its chain-of-thought steps also biased? The reported finding—CoT bias correlating with output bias at under 0.6 with p<0.001 across 5 LLMs and 11 bias types—would be useful for anyone auditing fairness via reasoning traces. That is the one piece of value here: the question is worth asking, and the claimed result would be consequential if it were actually supported.\n\nBut it is not supported. The full text supplied is a different paper entirely: “A disputable assumption behind the empirical equivalence between pilot-wave theory and standard quantum mechanics” by Manero, Muciño, and Okon, arXiv:2508.06667v1. There are no language models, no chain-of-thought prompting, no fairness metrics, no correlation table, no experimental setup of any kind. The abstract alone is not an experiment. This is not a matter of a weak method or a confound; the load-bearing evidence is simply absent. The reader’s stress-test note is accurate, and I agree with the verdict: UNVERDICTED.\n\nEven if the experiment were present, there is a secondary soft spot in the abstract’s framing: calling CoT text the model’s “thoughts” assumes the text is a faithful record of internal reasoning. That assumption would need defense, since CoT traces are generated outputs shaped by prompting and decoding. But that is a substantive critique for when there is an actual paper to critique. Right now it is almost a side note.\n\nI want to be clear about what this is not. There is no sign of trickery or intent here; it reads like a submission mix-up or a corrupted file. But as submitted, the manuscript is incoherent on its own terms: the abstract claims one result and the body delivers another. No referee can assess the claim. I would desk-reject this version, not because the idea lacks merit, but because there is nothing to review. If the authors resubmit with the actual LLM study, I would gladly take a look.","headline":"The abstract announces a meaningful LLM-fairness result, but the supplied full text is an unrelated pilot-wave QM paper, so the claimed experiment does not exist in this submission.","tokens_in":5222,"tokens_out":1362,"would_cite":false,"duration_ms":16185,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Biased final answers from LLMs do not reliably have corresponding bias in their chain-of-thought reasoning, according to this study.","keywords":["chain-of-thought prompting","fairness bias","language model bias","reasoning transparency","biased decisions","correlation analysis","LLM evaluation"],"falsifier":"One concrete check: on the same five models and eleven bias categories, compute the bias score from chain-of-thought text using an alternative fairness metric (e.g., a counterfactual substitution test) instead of the paper's metric. If the resulting correlations with final-output bias exceed 0.6 in most categories, the claim that thoughts and outputs are not highly correlated would fail. Also, if removing the instruction to 'think step by step' and instead using a stripped prompt changes the correlation pattern, the result depends on the elicitation method rather than on the models' internal t","tokens_in":4437,"feed_emoji":"⚖️","tokens_out":6145,"duration_ms":65464,"temperature":0.7,"pith_summary":"This paper asks whether large language models that give biased final answers also show corresponding bias in the chain-of-thought reasoning they generate. It reports correlations between fairness metrics computed on the reasoning text and on the final output across five popular LLMs, covering eleven bias categories. The central finding is that these correlations are weak—below 0.6, with p-values under 0.001 in most cases—leading the authors to conclude that biased decisions do not reliably come with biased 'thoughts' in the tested models. If correct, this challenges the common assumption that inspecting a model's step-by-step reasoning reveals whether it is biased.","feed_headline":"Biased answers don't mean biased thoughts in LLMs","feed_subtitle":"Across 5 models and 11 bias types, reasoning traces and answers differ in bias, so chain-of-thought may not reveal unfairness.","key_machinery":"Chain-of-thought prompting—asking the model to write out intermediate reasoning before giving an answer—is the instrument used to expose 'thoughts.' The argument turns on computing fairness metrics (quantifying gender, race, etc. bias) separately on the chain-of-thought text and on the final answer, then testing the correlation between those two sets of bias scores. The threshold claim is that the correlation is low (<0.6) and significant, which is taken to show the reasoning trace is not a reliable mirror of output bias.","core_discovery":"Using chain-of-thought prompting to elicit reasoning steps, the authors quantify 11 social biases (gender, race, socio-economic status, physical appearance, sexual orientation) in both the generated thoughts and the final outputs of five LLMs. They find that the bias level in the thinking steps is not highly correlated with the bias in the output: the correlation is less than 0.6, statistically significant at p<0.001 in most cases. Their interpretation is that, unlike humans, a model can produce a biased decision while its verbalized reasoning is not correspondingly biased—so biased decisions do not imply biased thoughts. This is framed as a caution about using chain-of-thought text as a tra","pith_inferences":["If the result generalizes, safety and interpretability work that reads chain-of-thought text to explain or audit model behavior would need to treat those traces as post-hoc rationalizations rather than causal reasoning.","A concrete testable extension: instead of prompting for text, probe the model's hidden activations or use a non-verbal CoT method to see if bias correlates with output; if a correlation appears, the disconnect may be specific to verbalized reasoning.","Another extension is to test whether the correlation rises when the model is explicitly instructed to justify its answer in terms of the protected attribute, which would indicate the bias can be verbalized when demanded.","The paper's p<0.001 with low r suggests a large, consistent decoupling; a further question is whether that decoupling is an artifact of how the CoT text was produced (e.g., style differences) rather than genuine independence."],"forward_implications":["Chain-of-thought text should not be used as a fairness audit: low thought-bias does not certify an unbiased answer.","Bias mitigation that targets the generated reasoning steps may not change final-output bias, since the two are decoupled.","Evaluations of LLM fairness should measure final decisions, not the verbalized reasoning, at least for the tested models.","The statistical significance of the low correlation means the finding is not a chance pattern, within the study's setup.","The claim sets up a future research target: discovering where output bias actually originates if not in the visible reasoning chain."],"supporting_citations":[],"fun_headline_variants":["Biased LLM answers don't mean biased thoughts","LLM bias: reasoning and output bias weakly linked","Chain-of-thought bias diverges from LLM answer bias","Biased output, unbiased reasoning in five LLMs"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The claim assumes that the text a model produces under chain-of-thought prompting is a faithful record of the model's actual reasoning, so that measuring bias in that text measures bias in its thoughts.","fun_headline_variants_meta":{"raw":{"variants":["Biased LLM answers don't mean biased thoughts","LLM bias: reasoning and output bias weakly linked","Chain-of-thought bias diverges from LLM answer bias","Biased output, unbiased reasoning in five LLMs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000252,"raw_usage":{"total_tokens":1370,"prompt_tokens":688,"completion_tokens":682,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":617}},"tokens_in":432,"tokens_out":682,"duration_ms":8129,"temperature":1.0,"reasoning_tokens":617,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:36:51.946792+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"One concrete check: on the same five models and eleven bias categories, compute the bias score from chain-of-thought text using an alternative fairness metric (e.g., a counterfactual substitution test) instead of the paper's metric. If the resulting correlations with final-output bias exceed 0.6 in most categories, the claim that thoughts and outputs are not highly correlated would fail. Also, if removing the instruction to 'think step by step' and instead using a stripped prompt changes the correlation pattern, the result depends on the elicitation method rather than on the models' internal t","supporting_citations":[],"review_version":1}