{"id":"447b0844-e949-42b8-b277-8a3a4a06aaaa","arxiv_id":"2411.15396","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Large language models show stronger decoy-effect bias than human judges when rating the credibility of medical web pages in COVID-19 treatment searches.","lead":"This paper compared how human raters and nine large language models judge the credibility of COVID-19 treatment pages in search results, with and without a deliberately misleading decoy page. It reports that the models, especially newer and larger ones, raised their ratings of misinformation when a decoy was present, more often than human raters did.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's decoy-effect counts appear to include significant t-tests in the opposite direction, so the prevalence claim in the abstract is not supported as reported.","rationale":"The reader's verdict is CONDITIONAL, and I agree that the paper should not be accepted as-is. However, the single most load-bearing concern is not the session-memory confound identified as the reader's weakest assumption, but rather the apparent treatment of significant reverse-direction t-tests as evidence for decoy effects. Because the headline claim is explicitly about prevalence ('more prevalent across different conditions and topics'), the counting and direction of the significance tests is the load-bearing component. If the paper counted reverse effects, the central claim collapses; if it did not, the paper needs to report the correct one-sided counts and release the analysis to verify them. My recommendation remains CONDITIONAL rather than REJECT because the published tables contain several true positive effects (notably gpt-4o across all three topics), so a corrected re-analysis might still support a narrower version of the claim. The reader's confound concern is valid and should be addressed by redesign or by explicitly framing the result as a session-context effect, but it does not by itself invalidate the existence of a decoy-driven bias; the direction-counting error, if confirmed, would invalidate the paper's central quantitative assertion. Since the reader's verdict already requires revision, I do not change the verdict, but I would strengthen the required revision to include a one-sided re-analysis and explicit reporting of reverse effects.","tokens_in":131,"tokens_out":7299,"duration_ms":79016,"concrete_test":"Re-analyze Tables 1 and 3 using the reported means and standard deviations (n=45 per group) to compute one-sided Welch t-tests for H1: target rating in treatment > target rating in control. Flag every row where the treatment target mean is lower than the control target mean as a reverse effect, and check whether the reported t-statistic has the same sign as the mean difference. Then recount, across models and topics, how many comparisons are true positive decoy effects, how many are null, and how many are reverse, with and without a Benjamini-Hochberg correction over the 54 tests. If the true positive count is not substantially larger than the reverse/null counts for the newer LLMs, the abstract's prevalence claim fails. Also re-run the same one-sided counting for the single-query Table 3 to verify the paper's claim that decoy effects are 'far less prevalent' there.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a frequency claim: decoy effects are 'more prevalent across different conditions and topics in LLM judgments compared to human credibility ratings.' The paper's own decision rule in §5.1.2 is one-sided: a decoy effect exists only when the target article's rating in the treatment group is significantly higher than in the control group. Yet Tables 1–3 report significance stars without enforcing that direction, and several starred comparisons move in the opposite direction. For example, in Table 1, gpt-4o-mini Topic 3 has control target mean 4.04 and treatment target mean 3.53 (t=8.00***), and claude-3-haiku Topic 3 has control 4.34 and treatment 4.24 (t=3.83***); these are significant reverse effects, not decoy effects. Similarly, Table 2 under 'Participants without prior knowledge' Topic 2 shows control 3.97 and treatment 3.40 with t=3.67***. If the abstract's 'more prevalent' count includes such significant reverse effects, the headline conclusion is inflated. The paper does not state how many of the 54 LLM tests were true positive decoy effects, and neither data nor code are released. Some t-statistic signs also appear inconsistent with the adjacent means (e.g., claude-3-sonnet Topic 1 in Table 1), suggesting possible sign-coding errors. This is the most load-bearing issue because it undermines the quantitative evidence for the main claim even before considering the session-memory confound. The session-context confound in the multi-query condition is a secondary but real problem: the treatment contrast combines the current decoy article with the LLM's own previous decoy-influenced ratings, so the causal role of the current decoy alone is not isolated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a between-subject crowdsourcing experiment (the participant count is reported inconsistently as 617 and 540) and a parallel LLM experiment with nine models (Llama-3.1-70B, Llama-3-70B, Llama-3-8B, gpt-4o, gpt-4o-mini, gpt-3.5-turbo, claude-3.5-sonnet, claude-3-sonnet, claude-3-haiku) to test whether the decoy effect biases credibility ratings of COVID-19 medical web pages. Target misinformation articles were rated under control SERPs and treatment SERPs that added a salient decoy misinformation article; human and LLM ratings were compared, with TREC Health Misinformation ground truth used to define correct information versus misinformation. The paper claims that larger and more recent LLMs are better at distinguishing credible information from misinformation yet are more vulnerable to decoy effects, and that decoy effects are more prevalent in LLM judgments than in human credibility ratings, especially in multi-query session prompts.","tokens_in":24830,"tokens_out":8258,"duration_ms":63274,"significance":"The topic is timely and important: if LLM-based credibility judgments are systematically manipulable by decoy misinformation, this affects the validity of LLM-as-judge evaluation pipelines and AI-assisted health information search. Strengths include the dual human-LLM design, use of external TREC labels, multiple model families, and a clear between-subject control/treatment manipulation. However, the headline quantitative claims rest on t-test tallies that are not direction-consistent with the paper's own decision rule, and the multi-query condition is confounded with the LLM's own previous ratings; these issues must be resolved before the prevalence claims can be accepted.","major_comments":[{"comment":"The paper defines a decoy effect in §5.1.2 as occurring 'only when the target article's rating in the treatment group is significantly higher than in the control group,' yet Tables 1–3 mark significance without enforcing this direction. For example, Table 1, gpt-4o-mini, Topic 3 shows control target mean 4.04 vs treatment target mean 3.53 with t=8.00***; Table 1, claude3-haiku, Topic 3 shows control 4.34 vs treatment 4.24 with t=3.83***; and Table 2, 'Participants without prior knowledge,' Topic 2 shows control 3.97 vs treatment 3.40 with t=3.67***. These are significant reverse effects under the stated rule, not decoy effects. The abstract's claim that the effect is 'more prevalent' in LLM judgments is therefore not supported by the counts as reported; the authors must report direction-consistent counts of true decoy effects.","section":"§5.1.2, Tables 1–3"},{"comment":"Several reported t-statistics are inconsistent with the adjacent means, suggesting sign-coding errors. In Table 1, claude3-sonnet Topic 1 has control target 4.06 and treatment target 4.02 (treatment lower), yet the reported t is -1.98*; in Table 1, claude3-haiku Topic 1 has control target 3.71 and treatment target 3.78 (treatment higher), yet the reported t is 2.90**. If the t values were computed as control minus treatment, the signs in these rows are wrong; if they were computed as treatment minus control, many other rows are wrong. The authors should re-run and verify all t-tests and the associated p-values.","section":"§5.1.2, Tables 1 and 3"},{"comment":"The multiple-query condition does not isolate the causal effect of the decoy article. As stated in §4.2, for subsequent queries the prompt includes previous queries and the credibility ratings produced by the LLM for those queries, so the LLM has access to its own past judgments. Because treatment sessions contain decoys in earlier queries, the treatment and control conditions differ not only in the current query's decoy presence but also in the accumulated prior ratings shown in the prompt. The observed increase in target ratings could therefore be driven by sequential anchoring on the model's own prior ratings rather than by the current decoy article. The paper acknowledges this possibility in §4.2 but then attributes the amplification to 'session memory and interaction context' in §5.2. The causal claim about decoy effects in session contexts requires either an analysis restricted to the first query of each session, a control condition that inserts the same prior ratings without decoys, or a single-query analysis used as the primary evidence.","section":"§4.2, §5.2"},{"comment":"The participant sample size is reported inconsistently: §4.1.1 states that 617 participants completed the full experiment, while §4.1.4 states 540 effective data points after exclusions and a second collection phase. Table 2 does not report the per-cell Ns, so the reader cannot tell which sample underlies the human t-tests, nor how the two recruitment phases were combined. Please reconcile these numbers and report the N for each row of Table 2.","section":"§4.1.1 vs §4.1.4"},{"comment":"The number of significance tests is large (54 tests in Tables 1 and 3 alone, plus the human comparisons in Table 2), and no correction for multiple comparisons is reported. With this many tests, the presence of several significant reverse effects suggests that some starred entries may be false positives. The authors should report corrected p-values (e.g., Benjamini-Hochberg) or explicitly justify why correction is unnecessary, and should state the total number of true positive versus reverse significant effects.","section":"§5.1.2, Tables 1–3"}],"minor_comments":[{"comment":"The phrase 'human judgers' should be 'human judges'.","section":"§6.1"},{"comment":"The text says 'without fine-turning the model' and later 'in a a natural language format'; these should read 'without fine-tuning the model' and 'in a natural language format'.","section":"§6.1"},{"comment":"The sentence 'After excluding 60 participants who spent less than five minutes and 6 who provided identical ratings for all articles, data from 474 participants were obtained' is arithmetically consistent with 540 but conflicts with the 617 reported in §4.1.1; please unify the reporting.","section":"§4.1.4"},{"comment":"Figures 6 and 7 would be easier to read with a per-panel legend or clearer line-type labeling, since the shared legend is hard to map when panels are small.","section":"Figures 6 and 7"},{"comment":"The row header 'Clicked on decoy/random prior to target' in the text differs in capitalization from the table header; please harmonize the table headers and the text.","section":"Section 5.1.2 and Table 2"},{"comment":"The abstract states that the decoy effect is 'more prevalent across different conditions and topics in LLM judgments'; please specify whether 'conditions' refers to the rows of Table 2, the prompt types (single- vs multi-query), or both.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper's topic is well suited to the journal, and the human-LLM comparison is potentially valuable, but the current analysis does not support the headline prevalence claims. The sign inconsistencies in Tables 1–3 are especially concerning because they suggest either transcription errors or a mismatch between the stated decision rule and the reported statistics; these need to be fixed before further review. The session-memory confound is design-level, but a reanalysis focusing on first-query or single-query conditions could salvage the core comparison. I would not recommend rejection at this stage, but the revision requires new analyses, not just copy-editing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know: this is the first head-to-head comparison of decoy effects on credibility judgments between humans and nine LLMs in a medical misinformation setting, and the central quantitative claim does not hold up as reported. The design is solid enough to deserve serious peer review, but the abstract's \"more prevalent\" conclusion needs a reanalysis before it can be believed.\n\nWhat is actually new: extending the decoy paradigm from relevance judgments and search interaction to LLM credibility assessment, with external ground truth from TREC Health Misinformation, a between-subject crowdsourcing baseline, and a multi-query versus single-query comparison. The finding that newer models like GPT-4o and Claude-3-Sonnet are both better at distinguishing credible information and more susceptible to decoy shifts is interesting and worth testing further. The paper also earns credit for being readable, for stating its decision rule explicitly in Section 5.1.2, and for acknowledging in Section 4.2 that the multi-query condition lets the LLM see its own past ratings.\n\nThe soft spots are real and load-bearing. The paper's own rule is one-sided: a decoy effect requires the target's treatment rating to be significantly higher than control. But the tables report two-sided significance, and several starred comparisons move in the opposite direction. For example, gpt-4o-mini Topic 3 in Table 1 shows control 4.04 versus treatment 3.53 with t=8.00***, and claude3-haiku Topic 3 shows control 4.34 versus treatment 4.24 with t=3.83***. Some sign/mean pairs are internally inconsistent, such as claude3-sonnet Topic 1 and Topic 3 in Table 1, which suggests coding errors. There are 54 uncorrected t-tests, so multiple-comparison correction matters. The session-memory confound is secondary but genuine: the multi-query treatment contrast combines the current decoy with the model's own previous decoy-influenced ratings, so the causal role of the current decoy is not isolated. The single-query condition is cleaner and helps, but the headline human-LLM comparison relies on the multi-query condition. The paper also reports inconsistent participant counts (617 versus 540) and releases no data, code, or exact prompts.\n\nWho this is for: people working on LLM-as-judge, AI audit, and behavioral IR. The idea is valuable and the experimental framing is reusable. But the evidence as presented is not yet credible. The right outcome is peer review with major revision: require a one-sided direction check, a corrected table, multiple-comparison reporting, artifact release, and a clearer separation of decoy effects from session-memory anchoring.","headline":"A genuinely new human-vs-LLM decoy-effect study with a clean design, but the headline prevalence claim is inflated by reverse-significant tests and a session-memory confound.","tokens_in":25516,"tokens_out":4039,"would_cite":false,"duration_ms":37408,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Decoy results sway LLM judges more than human raters in medical credibility checks.","keywords":["decoy effect","credibility assessment","large language models","misinformation","COVID-19 medical information","information retrieval evaluation","cognitive bias","human-LLM comparison"],"falsifier":"Run the multi-query LLM condition again, but in the treatment condition keep the LLM's own previous ratings in the prompt while replacing the decoy article with a neutral filler article; if target-misinformation ratings still rise relative to control, the decoy effect explanation fails.","tokens_in":24314,"feed_emoji":"🎯","tokens_out":4176,"duration_ms":36053,"temperature":0.7,"pith_summary":"The paper asks whether LLM judges are as vulnerable as human judges to a well-known cognitive bias, the decoy effect, when rating the credibility of COVID-19 treatment information in web search. It reports that adding an obviously false 'decoy' result to a search results page raises the credibility rating that larger, more recent LLMs give to a separate misinformation page, and that this effect appears across more topics in LLM judgments than in human credibility ratings. The finding matters because LLMs are increasingly used to automate or assist credibility judgments in high-stakes domains, and the paper treats it as evidence that these tools are not the purely rational judges they are often assumed to be.","feed_headline":"Decoy results sway LLM judges more than human raters","feed_subtitle":"Adding a fake, obviously false result makes LLMs rate misinformation as more credible, more often than it fools humans.","key_machinery":"The central object is the decoy effect (asymmetric dominance), imported from behavioral economics: an inferior 'decoy' option makes a similar target option look better by comparison. In the experiment, the decoy is a deliberately salient misinformation result appended to a search results page; the measured outcome is the credibility rating of a target misinformation article in the treatment condition versus the control condition. The second piece of machinery is the prompt format: single-query prompts, where each query is judged alone, versus multiple-query prompts that feed the LLM its own previous ratings back as session context. Comparing these two formats shows how much session memory and interaction context amplify the decoy effect.","core_discovery":"The paper claims that LLM judges, especially larger and more recent models, are more susceptible than human judges to the decoy effect when rating the credibility of medical web pages: adding an obviously false 'decoy' result to a search results page raises the credibility rating they give to a separate misinformation page, and this bias appears across more topics and conditions in LLM judgments than in human ratings. The paper interprets this as empirical evidence that LLMs are not purely rational and carry cognitive bias risks into automated judgment tasks.","pith_inferences":["We infer that the amplification in multi-query prompts may be sequential anchoring rather than classic decoy salience, because the only added content in those prompts is the model's own past ratings; separating these mechanisms is a natural next experiment.","We infer that the same vulnerability should appear in other high-stakes automated judgment domains, such as financial or legal information, where a salient low-quality option could inflate the credibility of a nearby misleading option.","A cheap mitigation suggested by the data is to evaluate each query independently or strip session memory when LLMs are used as judges; the paper's single-query results imply this would reduce decoy-driven rating inflation."],"forward_implications":["Session-aware LLM judges are the vulnerable configuration: decoy effects largely disappeared in single-query prompts, so any deployment that gives an LLM memory of prior judgments should be audited for decoy-style manipulation.","The models that best distinguish credible from misleading content are also the ones whose target-misinformation ratings rise most with a decoy, so accuracy and bias-resistance do not move together.","Human decoy effects were narrower and depended on clicking behavior and prior knowledge, so using human ratings as a baseline can localize the bias rather than eliminate it.","A decoy test should be part of LLM evaluation benchmarks for health information, since a single inserted snippet can shift credibility scores in a high-stakes domain."],"supporting_citations":[{"why":"Introduces the framing/decoy phenomenon that the experiment operationalizes as an obviously false search result.","marker":"[66]"},{"why":"Establishes the asymmetric dominance decoy effect in consumer choice, the theoretical basis for the decoy design.","marker":"[25]"},{"why":"Connects decoy effects to judgment tasks, supporting the use of credibility ratings rather than only choices.","marker":"[74]"},{"why":"Prior work on decoy effects in search interaction that this study extends from relevance judgments to credibility judgments.","marker":"[10]"},{"why":"Pilot study on decoy effects in search interaction, providing the immediate empirical precedent for the SERP-based manipulation.","marker":"[12]"},{"why":"Documents cognitive biases in crowdsourcing tasks, the baseline literature for human judge bias in information evaluation.","marker":"[18]"},{"why":"Shows threshold priming biases in LLM batch relevance judgments, a precedent for LLM cognitive bias in judgment tasks.","marker":"[9]"},{"why":"Supports the design choice that adding specific details and statistics to a snippet makes content appear more credible, which the decoy manipulation relies on.","marker":"[17]"}],"fun_headline_variants":["LLMs fall for decoy misinformation more than humans","Decoy bias: LLMs rate fake news higher than people do","AI judges more susceptible to decoy effect than humans","COVID-19 decoy results bias LLMs more than human raters","Stronger decoy bias in LLM credibility judgments than human"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim assumes that the higher credibility ratings in the multi-query treatment condition are caused by the decoy article itself, rather than by the LLM anchoring on its own previous ratings that the prompt feeds back into the conversation.","fun_headline_variants_meta":{"raw":{"variants":["LLMs fall for decoy misinformation more than humans","Decoy bias: LLMs rate fake news higher than people do","AI judges more susceptible to decoy effect than humans","COVID-19 decoy results bias LLMs more than human raters","Stronger decoy bias in LLM credibility judgments than human"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1446,"prompt_tokens":946,"completion_tokens":500,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":415}},"tokens_in":562,"tokens_out":500,"duration_ms":4890,"temperature":1.0,"reasoning_tokens":415,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:21:11.243655+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the multi-query LLM condition again, but in the treatment condition keep the LLM's own previous ratings in the prompt while replacing the decoy article with a neutral filler article; if target-misinformation ratings still rise relative to control, the decoy effect explanation fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the framing/decoy phenomenon that the experiment operationalizes as an obviously false search result."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the asymmetric dominance decoy effect in consumer choice, the theoretical basis for the decoy design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Connects decoy effects to judgment tasks, supporting the use of credibility ratings rather than only choices."},{"cited_title":"Decoy Effect In Search Interaction: Understanding User Behavior and Measuring System Vulnerability","cited_arxiv_id":"2403.18462","evidence_quote":"Prior work on decoy effects in search interaction that this study extends from relevance judgments to credibility judgments."},{"cited_title":"Decoy Effect in Search Interaction: A Pilot Study","cited_arxiv_id":"2311.02362","evidence_quote":"Pilot study on decoy effects in search interaction, providing the immediate empirical precedent for the SERP-based manipulation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents cognitive biases in crowdsourcing tasks, the baseline literature for human judge bias in information evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the design choice that adding specific details and statistics to a snippet makes content appear more credible, which the decoy manipulation relies on."}],"review_version":1}