{"id":"9f82069c-57ad-46f5-a237-6cf881d8ca1f","arxiv_id":"2411.17304","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Masking bias-triggering words with random identifiers increased accuracy on two small LLM reasoning and counting tasks, with effects varying by model.","lead":"Researchers replaced words that seem to trigger biased answers in LLM prompts with meaningless random codes and tested whether the models answered more correctly. On small logic and counting tasks, the hashed prompts improved accuracy for several models, but the effect depended on the model and the task.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Hashing is confounded with prompt rewording in Exp. 1: the hashed prompt adds an explicit masking note and different syntax, so the effect cannot be isolated to replacing word meanings; §3.1.1 word selection from the same models' errors makes this harder to interpret.","rationale":"The central claim is modest and empirically framed, and the paper has real strengths: prompts and model outputs are released, the frequent-itemset experiment uses algorithmically verifiable ground truth, and several models are tested. My concern is not that the authors are hiding results—they report model-dependent and even negative effects (Gemini, Mixtral, hallucinations). The issue is internal validity of the headline Linda evidence. The manipulation bundles semantic masking with instruction and syntax changes, and the validation condition does not unbundle them. If the driver is the 'Note that...' sentence, the paper's proposed mechanism ('removing representativeness cues') is not established even though the prompt-level intervention might still work. Selection of words from the same models' errors is a second, independent threat: it could make the improvement an artifact of fitting the word set to the test prompt. Since the abstract claims improvements in 'all tested scenarios' and the reader's verdict is already conditional, this analysis keeps the verdict conditional rather than overturning it; the proposed ablation and a pre-registered selection rule would be enough to upgrade or downgrade confidence. The chi-square analyses also pool many responses as independent observations, but the large effect sizes in Experiment 1 make this secondary to the confound concern.","tokens_in":20313,"tokens_out":6221,"duration_ms":79217,"concrete_test":"Run a 2x2 ablation on the Section 3.1.1 Linda prompt with 20+ iterations on at least GPT-3.5/GPT-4o and Llama-3.1-70B: (a) original prompt; (b) hashed prompt without the 'Note that...' sentence and without relationship hints, but with grammatical filler so the sentence remains well-formed; (c) same as (b) but target words replaced by frequent, stereotype-unrelated English words; (d) masking with a single symbol X and no meta-note. If (b) does not beat (c) or (d), the improvement in Experiment 1 is attributable to instruction/syntax artifacts rather than to semantic masking. To test selection stability, also pre-register a fixed rule (e.g., hash all descriptive adjectives/nouns in ten unseen Linda-style prompts) and check whether the effect replicates without inspecting model errors first.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In the main Linda experiment, the 'hashed' prompt is not the original prompt with selected words replaced. It also (i) inserts 'Note that in the text below, specific information was masked behind anonymous identifiers such as X and cdf14,' (ii) rewrites the person description as ungrammatical placeholder syntax ('Imagine X with a cdf14 and a a214s, sitting in a fg57 rfg5a'), and (iii) adds meta-remarks about possible connections between hidden identifiers. The validation condition (Fig. 7) controls only the added neutral descriptions and relationship hints; it does not control for the masking note or for placeholder syntax while keeping semantic content available. Thus an alternative reading is that the model is simply being instructed (or forced by strange syntax) that surface details are untrustworthy, rather than that removing word meanings is the active mechanism. Section 3.1.1 compounds this: the words to hash were selected by inspecting the same models' incorrect outputs on the original prompt, so the word set is not shown to be a stable, model-independent set; on the Linda task the 'method' includes a bespoke selection step fitted to the observed failure modes. Experiment 2 is cleaner in some respects (no meta-instruction, no selection) but it hashes all data values, not selected bias-inducing words, so it does not repair the Linda evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes 'hashing': replacing potentially bias-inducing words or data values in prompts with unique meaningless identifiers, and reports four experiments on the Linda conjunction-fallacy task (free-text and tabular variants), frequent-itemset counting, and a comparison with chain-of-thought (CoT) models. The paper claims statistically significant improvements in all tested scenarios, while noting that the effects vary across model families and that hallucination rates are inconsistently reduced. Prompts and model outputs are available in a public repository.","tokens_in":20589,"tokens_out":4104,"duration_ms":36584,"significance":"If the claimed effect is real, prompt-level hashing would be a lightweight, model-agnostic debiasing intervention with practical value, and the comparison with CoT is timely and useful. The paper's strengths include the clean construction in Experiment 2, where CSV-correct, CSV-wrong, and CSV-hashed datasets were designed to preserve the frequent-itemset ground truth; evaluation of recall and precision against Apriori; coverage of several model families; and a public repository of prompts and outputs. However, the central Experiment 1 evidence does not currently isolate the proposed mechanism, and the word-selection procedure introduces a test-set fitting component. These issues are load-bearing for the paper's main claim, so the manuscript requires substantial revision.","major_comments":[{"comment":"The hashed prompt is not the original prompt with selected words replaced. It inserts an explicit instruction ('Note that in the text below, specific information was masked behind anonymous identifiers such as X and cdf14'), rewrites the person description in ungrammatical placeholder syntax ('Imagine X with a cdf14 and a a214s, sitting in a fg57 rfg5a'), and appends relationship hints. The validation condition in Fig. 7 controls only for the added neutral descriptions and relationship hints; it does not control for the masking note or for the placeholder syntax while keeping semantic content available. The observed improvement therefore cannot be attributed specifically to removing word meanings; it may stem from instructing the model that surface details are untrusted or from the syntactic disruption. A same-symbol masking baseline and a control that adds the masking note without hashing are needed.","section":"§3.1.1, Figs. 5–6"},{"comment":"The words to hash were selected by inspecting the same models' incorrect outputs on the original prompt, and the Limitations section confirms this. The method therefore includes a bespoke, test-set-fitted mask, and the measured improvement is not evidence for hashing as a general method with a stable word set. The authors should either pre-specify the mask independently of the evaluated outputs, evaluate on held-out models or prompts, or demonstrate that the selected word set is stable across models. Without this, Experiment 1 conflates the hashing mechanism with the selection procedure.","section":"§3.1.1 'Selection of possible bias-inducing words'"},{"comment":"The pooled Fisher test (p < 0.00001) aggregates 20 or 10 iterations per model and is driven by large shifts in some models, but per-model results are highly heterogeneous: Llama 2 70B remains at 0 correct in both hashed variants, GPT-4 improves to 10/10 with added descriptions but drops to 1/10 without, and Gemini shows the opposite pattern. The abstract's claim of 'significant improvements in all tested scenarios' overstates the evidence. Per-model tests should be reported and the consistency of the effect discussed; the pooled N is small and the prompts are not identical across conditions.","section":"§3.1.2, Tables 3–4 and Table 6"},{"comment":"The chi-square tests treat individual itemsets (e.g., 235 itemsets per condition) as independent observations, although itemsets within one prompt or run are highly correlated. This inflates the effective sample size and can make small differences significant, as in GPT-4o's p = 0.043 with Cramér's V = 0.093. The analysis should be performed at the run level (25 runs per condition) or with a mixed-effects model that accounts for prompt-level clustering.","section":"§3.2.2, Tables 8 and 10"}],"minor_comments":[{"comment":"The abstract states '490 prompts' in one version and '680 prompts' in the full manuscript; these numbers should be aligned.","section":"Abstract"},{"comment":"The text says 'In the Llama 2 model, a χ² test was also conducted' but the experiment uses Llama 3.1-405b; this appears to be a typo.","section":"§3.2.2"},{"comment":"The model name 'ChatGPT-3o-mini' should be 'ChatGPT-o3-mini'; also, the comparison text inconsistently refers to 'ChatGPT-4' and 'ChatGPT-4o'.","section":"§3.4"},{"comment":"The option 'X is b321 who 4l5i' is ungrammatical; the intended reading is presumably 'X is b321 who likes to 4l5i' or 'X is b321 who 4l5i' should be otherwise clarified.","section":"Tables 3 and 4"},{"comment":"The figures would benefit from explicit axis labels and a note that hallucination counts are shown separately from found itemsets.","section":"Figures 12 and 13"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a cs.CL venue and the repository plus frequent-itemset experiment are useful contributions. However, the central causal claim rests on Experiment 1, which is currently confounded and partly circular. I would support a revised version that adds the missing controls, pre-specifies or cross-validates the word selection, and reanalyzes the itemset data at the run level. Without these changes, the paper's contribution reduces to a prompt-engineering observation with limited generalizability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper has a real idea — replacing bias-triggering words with unique meaningless identifiers — and the frequent-itemset experiment gives it some honest support, but the main Linda experiment is too confounded to carry the debiasing claim. The hashed prompt in Exp 1 is not the original prompt with words swapped; it also adds a meta-note about masking, rewrites the description as odd placeholder syntax, and appends relationship hints. The validation prompt only controls the added descriptions and hints, not the masking note or the syntax. So the improvement could just be the model being told or forced to treat surface details as untrustworthy. That's a real gap, and the paper's own word-selection procedure makes it worse: the hashed words were chosen from the same models' errors on the original prompt (§3.1.1, and acknowledged in Limitations). That's fitting the mask to the test set, not testing a general method.\n\nThe unique-identifier twist over same-symbol masking is modest but genuinely new — you can reference hashed entities later — and the frequent-itemset setup is cleaner. There, no meta-instruction or selection step; both GPT-4o and Llama-3.1-405b show significant recall gains on hashed data vs correct/wrong data, and the paper is transparent that Llama's hallucinations rise. Exp 3 inherits Exp 1's confound and also mixes GPT-4 vs GPT-4o (acknowledged). Exp 4 is a sensible sanity check but underpowered.\n\nCredit where due: the repo ships prompts and outputs, the stats include effect sizes, and the limitations section is candid. The abstract inconsistency (490 vs 680 prompts; 'three' vs 'four' experiments between the metadata and the text) is sloppy but fixable.\n\nNet: the central claim is only conditionally supported. Exp 2 is the best evidence, but it hashes all data values, so it doesn't isolate 'bias-inducing words'. Per-model results in Exp 1 are wildly mixed (Gemini 0–10 depending on variant), so the pooled chi-square overstates the effect. I'd send this to review — the idea is worth a serious look and the repo makes it easy to test — but I'd require a proper control: same-symbol masking, identical syntax, pre-registered word selection, per-model reporting. If the authors fix Exp 1, the paper becomes a useful contribution; as is, it's a suggestive preprint.","headline":"A useful but weakly controlled prompt-debiasing idea; the unique-identifier twist is real, but the main Linda experiment conflates hashing with prompt rewording and bespoke word selection.","tokens_in":127,"tokens_out":3135,"would_cite":false,"duration_ms":73549,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Replacing bias-triggering words with meaningless identifiers makes LLMs answer the Linda problem correctly and improves frequent-itemset recall, the paper reports.","keywords":["large language models","cognitive biases","conjunction fallacy","hashing","prompt engineering","debiasing","frequent itemsets","representativeness heuristic"],"falsifier":"Hash a different set of content words on the same adapted Linda prompt—for instance, only the adjectives 'long' and 'colorful', or every noun except the target roles—and compare fallacy rates against the original baseline. If any content-word hashing produces the same near-elimination of the fallacy, the effect is generic; if only the paper's exact word list works, the result is an artifact of selection. A stronger test would choose the words on a development set of models and prompts, freeze the list, and apply it to held-out models and prompts.","tokens_in":20136,"feed_emoji":"🧠","tokens_out":4941,"duration_ms":43662,"temperature":0.7,"pith_summary":"The paper claims that replacing words likely to trigger stereotypes or memorized associations with meaningless, hash-like identifiers improves LLM accuracy on tasks where those associations cause errors. In the modified Linda problem, models never picked the correct non-conjunction answer in 80 baseline runs, but after hashing they answered correctly 14 or 13 times out of 80 in the two prompt variants; in frequent-itemset counting, hashed tables raised recall for both tested models, with the largest gain for Llama-3.1-405b. A tabular version of the Linda problem also improved with hashing, and the method matched or beat chain-of-thought prompting on several comparisons without requiring extra inference compute. The authors position hashing as a prompt-level, model-agnostic debiasing step that reduces reliance on external knowledge, while noting that hallucination rates were not consistently reduced.","feed_headline":"Hashing bias words in prompts fixes LLM logic errors","feed_subtitle":"Replacing 'reading' with meaningless codes lifts Linda-problem and itemset scores across five model families.","key_machinery":"Hash-like identifiers: each potentially bias-inducing word or phrase is replaced by a fixed meaningless token (for example, reading becomes rfg5a) that can still be referenced later in the prompt. The mechanism is that the model must reason about relations between opaque placeholders rather than about 'reading' and 'colorful coat', so its answers are driven by the prompt's structure and stated relationships instead of by semantic priors from training data. Because the identifiers are unique and repeatable, the method differs from ordinary masking and allows the prompt to state possible links between hidden elements without revealing their identities.","core_discovery":"The central discovery is that a pure lexical transformation—replacing selected content words by unique opaque identifiers such as cdf14 while keeping syntax and cross-references intact—makes LLMs less likely to answer from learned stereotypes and more likely to answer from the logical structure of the prompt. In the conjunction-fallacy experiment, the original adapted Linda prompt produced the wrong conjunctive answer in 80 of 80 runs across Gemini, GPT-3.5, GPT-4, and Llama 2, while the hashed prompts yielded 14 and 13 correct answers out of 80, with Fisher's exact test p-values below 0.00001 and large Cramér's V effect sizes. In the itemset task, hashing increased true itemsets found from 213 to 225 out of 235 for GPT-4o and from 194 to 230 out of 235 for Llama-3.1-405b, both statistically significant by chi-square tests. The authors interpret this as disrupting the representativeness heuristic and preventing the model from injecting pretrained world knowledge into tasks where the prompt's artificial data should be the only source of truth.","pith_inferences":["Editorial inference: because the GPT-4o itemset effect was statistically significant but small (Cramér's V ≈ 0.09), the practical benefit of hashing for counting tasks may be modest even where it is real, and the larger Llama recall jump suggests the gain scales with how strongly a model leans on pretrained associations.","Editorial inference: hashing hides identity while preserving co-reference and syntax, so combining it with chain-of-thought could suppress both associative bias and shallow reasoning; the paper names this combination as future work, and a direct head-to-head test would be a natural next experiment.","Editorial inference: automating the selection of words to hash would turn the method from a hand-crafted intervention into a deployable preprocessor, and the authors' manual selection is the main obstacle to scaling the approach.","Editorial inference: if the effect replicates on new prompts and models, hashing offers a cheap inference-time debiasing lever for human-in-the-loop systems where retraining is impractical, at the cost of hiding information that a user might want to inspect."],"forward_implications":["On the adapted Linda problem, hashing raised correct non-conjunction answers from 0 out of 80 to 14 or 13 out of 80, with p-values below 0.00001 and Cramér's V around 0.47–0.49.","On frequent-itemset extraction, hashed data improved recall for GPT-4o (from 213 to 225 of 235 true itemsets) and for Llama-3.1-405b (from 194 to 230), with statistically significant chi-square tests in both cases.","In the tabular CSV version of the Linda problem, hashing without added relationship descriptions raised correct answers across models with p = 0.0000535 and Cramér's V = 0.404, including 10 out of 10 for Llama-3.1-405B.","Hashing was comparable to or better than chain-of-thought prompting on several GPT-family comparisons, and unlike CoT it adds no extra inference cost.","Hallucination reduction was inconsistent: GPT-4o hallucinated at similar rates across conditions, while Llama-3.1-405b hallucinated more in the hashed condition."],"supporting_citations":[{"why":"Contributed the LLM-adapted Linda prompt used as the bias baseline and the instruction to answer without justification.","marker":"Suri et al., 2024"},{"why":"Defines the conjunction fallacy and the original Linda problem, the cognitive bias the method targets.","marker":"Tversky & Kahneman, 1983"},{"why":"Provided the Apriori algorithm whose output serves as the reference set of true frequent itemsets.","marker":"Agrawal et al., 1996"},{"why":"Supplied the Zoo dataset that inspired the CSV-correct animal table used in the itemset experiment.","marker":"Forsyth, 1990"},{"why":"Represents the masking literature from which hashing is distinguished by its use of unique, referenceable identifiers.","marker":"Hackmann et al., 2024"},{"why":"Documents representativeness heuristics in LLMs, the hypothesized source of bias that hashing removes.","marker":"Wang et al., 2024"}],"fun_headline_variants":["Hash bias words to boost LLM reasoning and stats","Obfuscating biased terms improves LLM logic","Meaningless hashes cut LLM bias and fallacies","Hashing prompt words sharpens LLM statistical learning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The success depends on the manually selected list of bias-inducing words being a stable, model-independent set; if the words were chosen because they made the specific test models fail on the specific prompts, the reported improvement is not evidence that hashing works on new prompts or models.","fun_headline_variants_meta":{"raw":{"variants":["Hash bias words to boost LLM reasoning and stats","Obfuscating biased terms improves LLM logic","Meaningless hashes cut LLM bias and fallacies","Hashing prompt words sharpens LLM statistical learning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000244,"raw_usage":{"total_tokens":1557,"prompt_tokens":992,"completion_tokens":565,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":501}},"tokens_in":608,"tokens_out":565,"duration_ms":5963,"temperature":1.0,"reasoning_tokens":501,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:15:31.803035+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hash a different set of content words on the same adapted Linda prompt—for instance, only the adjectives 'long' and 'colorful', or every noun except the target roles—and compare fallacy rates against the original baseline. If any content-word hashing produces the same near-elimination of the fallacy, the effect is generic; if only the paper's exact word list works, the result is an artifact of selection. A stronger test would choose the words on a development set of models and prompts, freeze the list, and apply it to held-out models and prompts.","supporting_citations":[],"review_version":1}