Pith. sign in

REVIEW 5 major objections 6 minor 89 references

The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges

T0 review · 5 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Decoy results sway LLM judges more than human raters in medical credibility checks.

desk verdict A genuinely new human-vs-LLM decoy-effect study with a clean design, but the headline prevalence claim is inflated by reverse-significant tests and a session-memory confound. read the letter →

arxiv 2411.15396 v1 pith:2QTP7ZTB submitted 2024-11-23 cs.IR cs.AIcs.HC

classification cs.IRcs.AIcs.HC
keywords decoyeffectcredibilityassessmentlargelanguagemodelsmisinformationCOVID-19medicalinformationretrievalevaluationcognitivebiashuman-LLMcomparison
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether LLM judges are as vulnerable as human judges to a well-known cognitive bias, the decoy effect, when rating the credibility of COVID-19 treatment information in web search. It reports that adding an obviously false 'decoy' result to a search results page raises the credibility rating that larger, more recent LLMs give to a separate misinformation page, and that this effect appears across more topics in LLM judgments than in human credibility ratings. The finding matters because LLMs are increasingly used to automate or assist credibility judgments in high-stakes domains, and the paper treats it as evidence that these tools are not the purely rational judges they are often assumed to be.

What carries the argument

The central object is the decoy effect (asymmetric dominance), imported from behavioral economics: an inferior 'decoy' option makes a similar target option look better by comparison. In the experiment, the decoy is a deliberately salient misinformation result appended to a search results page; the measured outcome is the credibility rating of a target misinformation article in the treatment condition versus the control condition. The second piece of machinery is the prompt format: single-query prompts, where each query is judged alone, versus multiple-query prompts that feed the LLM its own previous ratings back as session context. Comparing these two formats shows how much session memory and interaction context amplify the decoy effect.

What would settle it

Run the multi-query LLM condition again, but in the treatment condition keep the LLM's own previous ratings in the prompt while replacing the decoy article with a neutral filler article; if target-misinformation ratings still rise relative to control, the decoy effect explanation fails.

Watch

Extended reading notes

Core claim

The paper claims that LLM judges, especially larger and more recent models, are more susceptible than human judges to the decoy effect when rating the credibility of medical web pages: adding an obviously false 'decoy' result to a search results page raises the credibility rating they give to a separate misinformation page, and this bias appears across more topics and conditions in LLM judgments than in human ratings. The paper interprets this as empirical evidence that LLMs are not purely rational and carry cognitive bias risks into automated judgment tasks.

Load-bearing premise

The central claim assumes that the higher credibility ratings in the multi-query treatment condition are caused by the decoy article itself, rather than by the LLM anchoring on its own previous ratings that the prompt feeds back into the conversation.

Editorial extensions

If this is right

  • Session-aware LLM judges are the vulnerable configuration: decoy effects largely disappeared in single-query prompts, so any deployment that gives an LLM memory of prior judgments should be audited for decoy-style manipulation.
  • The models that best distinguish credible from misleading content are also the ones whose target-misinformation ratings rise most with a decoy, so accuracy and bias-resistance do not move together.
  • Human decoy effects were narrower and depended on clicking behavior and prior knowledge, so using human ratings as a baseline can localize the bias rather than eliminate it.
  • A decoy test should be part of LLM evaluation benchmarks for health information, since a single inserted snippet can shift credibility scores in a high-stakes domain.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • We infer that the amplification in multi-query prompts may be sequential anchoring rather than classic decoy salience, because the only added content in those prompts is the model's own past ratings; separating these mechanisms is a natural next experiment.
  • We infer that the same vulnerability should appear in other high-stakes automated judgment domains, such as financial or legal information, where a salient low-quality option could inflate the credibility of a nearby misleading option.
  • A cheap mitigation suggested by the data is to evaluate each query independently or strip session memory when LLMs are used as judges; the paper's single-query results imply this would reduce decoy-driven rating inflation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This paper reports a between-subject crowdsourcing experiment (the participant count is reported inconsistently as 617 and 540) and a parallel LLM experiment with nine models (Llama-3.1-70B, Llama-3-70B, Llama-3-8B, gpt-4o, gpt-4o-mini, gpt-3.5-turbo, claude-3.5-sonnet, claude-3-sonnet, claude-3-haiku) to test whether the decoy effect biases credibility ratings of COVID-19 medical web pages. Target misinformation articles were rated under control SERPs and treatment SERPs that added a salient decoy misinformation article; human and LLM ratings were compared, with TREC Health Misinformation ground truth used to define correct information versus misinformation. The paper claims that larger and more recent LLMs are better at distinguishing credible information from misinformation yet are more vulnerable to decoy effects, and that decoy effects are more prevalent in LLM judgments than in human credibility ratings, especially in multi-query session prompts.

Significance. The topic is timely and important: if LLM-based credibility judgments are systematically manipulable by decoy misinformation, this affects the validity of LLM-as-judge evaluation pipelines and AI-assisted health information search. Strengths include the dual human-LLM design, use of external TREC labels, multiple model families, and a clear between-subject control/treatment manipulation. However, the headline quantitative claims rest on t-test tallies that are not direction-consistent with the paper's own decision rule, and the multi-query condition is confounded with the LLM's own previous ratings; these issues must be resolved before the prevalence claims can be accepted.

major comments (5)
  1. [§5.1.2, Tables 1–3] The paper defines a decoy effect in §5.1.2 as occurring 'only when the target article's rating in the treatment group is significantly higher than in the control group,' yet Tables 1–3 mark significance without enforcing this direction. For example, Table 1, gpt-4o-mini, Topic 3 shows control target mean 4.04 vs treatment target mean 3.53 with t=8.00***; Table 1, claude3-haiku, Topic 3 shows control 4.34 vs treatment 4.24 with t=3.83***; and Table 2, 'Participants without prior knowledge,' Topic 2 shows control 3.97 vs treatment 3.40 with t=3.67***. These are significant reverse effects under the stated rule, not decoy effects. The abstract's claim that the effect is 'more prevalent' in LLM judgments is therefore not supported by the counts as reported; the authors must report direction-consistent counts of true decoy effects.
  2. [§5.1.2, Tables 1 and 3] Several reported t-statistics are inconsistent with the adjacent means, suggesting sign-coding errors. In Table 1, claude3-sonnet Topic 1 has control target 4.06 and treatment target 4.02 (treatment lower), yet the reported t is -1.98*; in Table 1, claude3-haiku Topic 1 has control target 3.71 and treatment target 3.78 (treatment higher), yet the reported t is 2.90**. If the t values were computed as control minus treatment, the signs in these rows are wrong; if they were computed as treatment minus control, many other rows are wrong. The authors should re-run and verify all t-tests and the associated p-values.
  3. [§4.2, §5.2] The multiple-query condition does not isolate the causal effect of the decoy article. As stated in §4.2, for subsequent queries the prompt includes previous queries and the credibility ratings produced by the LLM for those queries, so the LLM has access to its own past judgments. Because treatment sessions contain decoys in earlier queries, the treatment and control conditions differ not only in the current query's decoy presence but also in the accumulated prior ratings shown in the prompt. The observed increase in target ratings could therefore be driven by sequential anchoring on the model's own prior ratings rather than by the current decoy article. The paper acknowledges this possibility in §4.2 but then attributes the amplification to 'session memory and interaction context' in §5.2. The causal claim about decoy effects in session contexts requires either an analysis restricted to the first query of each session, a control condition that inserts the same prior ratings without decoys, or a single-query analysis used as the primary evidence.
  4. [§4.1.1 vs §4.1.4] The participant sample size is reported inconsistently: §4.1.1 states that 617 participants completed the full experiment, while §4.1.4 states 540 effective data points after exclusions and a second collection phase. Table 2 does not report the per-cell Ns, so the reader cannot tell which sample underlies the human t-tests, nor how the two recruitment phases were combined. Please reconcile these numbers and report the N for each row of Table 2.
  5. [§5.1.2, Tables 1–3] The number of significance tests is large (54 tests in Tables 1 and 3 alone, plus the human comparisons in Table 2), and no correction for multiple comparisons is reported. With this many tests, the presence of several significant reverse effects suggests that some starred entries may be false positives. The authors should report corrected p-values (e.g., Benjamini-Hochberg) or explicitly justify why correction is unnecessary, and should state the total number of true positive versus reverse significant effects.
minor comments (6)
  1. [§6.1] The phrase 'human judgers' should be 'human judges'.
  2. [§6.1] The text says 'without fine-turning the model' and later 'in a a natural language format'; these should read 'without fine-tuning the model' and 'in a natural language format'.
  3. [§4.1.4] The sentence 'After excluding 60 participants who spent less than five minutes and 6 who provided identical ratings for all articles, data from 474 participants were obtained' is arithmetically consistent with 540 but conflicts with the 617 reported in §4.1.1; please unify the reporting.
  4. [Figures 6 and 7] Figures 6 and 7 would be easier to read with a per-panel legend or clearer line-type labeling, since the shared legend is hard to map when panels are small.
  5. [Section 5.1.2 and Table 2] The row header 'Clicked on decoy/random prior to target' in the text differs in capitalization from the table header; please harmonize the table headers and the text.
  6. [Abstract] The abstract states that the decoy effect is 'more prevalent across different conditions and topics in LLM judgments'; please specify whether 'conditions' refers to the rows of Table 2, the prompt types (single- vs multi-query), or both.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the core claims are empirical comparisons grounded in new human and LLM experiments, with no prediction that reduces by construction to a fitted parameter or to a self-citation chain.

full rationale

The paper's central claims—that decoy items shift LLM credibility ratings and that this shift is more prevalent in LLMs than in human judges—are derived from new between-subject experiments, external TREC Health Misinformation ground-truth labels, and t-tests comparing target-article ratings across control and treatment SERPs, not from any equation in which the outcome is defined as the input. The decoy-effect decision rule in Section 5.1.2 (a treatment target rating 'significantly higher than in the control group') is an operational definition of the effect, and the paper then applies it to fresh data; this is a measurement choice, not a circular derivation. The authors' prior work on decoy effects appears only as related-work motivation and does not supply the uniqueness theorem or the fitted parameters on which the present results rest. The multi-query condition does include the LLM's own past ratings in the prompt (Section 4.2), and the paper itself notes this 'may influence its current judgment'; while this is a real threat to causal attribution of the decoy effect in the multi-query condition, it is a confound or validity critique rather than an equivalence-by-construction between input and output. Likewise, the apparent inclusion of significant t-tests in the reverse direction when counting 'more prevalent' effects (e.g., Table 1 gpt-4o-mini Topic 3: control 4.04 vs. treatment 3.53, t=8.00***) is an internal-consistency or coding concern about the prevalence claim, not a circularity. Because the main claims are self-contained empirical findings benchmarked against human raters and external labels, the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claim depends mainly on domain assumptions about ground truth, task comparability, and statistical independence. The design introduces no new entities and fits no free parameters.

assumptions (5)
  • domain assumption TREC Health Misinformation track labels are a reliable ground truth for correct versus misinformation articles.
    Section 4.1.2 says the questions have YES/NO answers validated by external experts, but the mapping from TREC labels to the correct/random/target/decoy categories is not reported.
  • domain assumption The decoy effect, defined for consumer choices, transfers to credibility ratings in search engine result pages.
    Section 4.1.3 assumes that adding an 'obviously false' decoy raises perceived credibility of the target misinformation, without a manipulation check that this is the same asymmetric-dominance mechanism.
  • domain assumption The 45 LLM runs per prompt are independent enough for t-tests.
    Section 4.2 lists models and snapshots but not decoding temperature or seed; if runs are deterministic or correlated, the t-test statistics are unreliable.
  • domain assumption The prompt successfully separates credibility from topical relevance.
    Section 4.2 instructs models to rate solely on credibility, but no check is reported for instruction compliance or for relevance leakage in article text.
  • domain assumption Human crowdsourced ratings and LLM ratings measure the same latent construct.
    Humans click and read pages in a SERP (Section 4.1.4) while LLMs receive titles and full text in a prompt (Section 4.2); the comparison assumes protocol differences do not change what is being measured.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges." pith.science (2026). https://pith.science/paper/2QTP7ZTB

@misc{pith2026241115396,
  author       = {Pith},
  title        = {Pith review of: The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2QTP7ZTB}},
  note         = {Machine review of arXiv:2411.15396}
}
read the original abstract

Can AI be cognitively biased in automated information judgment tasks? Despite recent progresses in measuring and mitigating social and algorithmic biases in AI and large language models (LLMs), it is not clear to what extent LLMs behave "rationally", or if they are also vulnerable to human cognitive bias triggers. To address this open problem, our study, consisting of a crowdsourcing user experiment and a LLM-enabled simulation experiment, compared the credibility assessments by LLM and human judges under potential decoy effects in an information retrieval (IR) setting, and empirically examined the extent to which LLMs are cognitively biased in COVID-19 medical (mis)information assessment tasks compared to traditional human assessors as a baseline. The results, collected from a between-subject user experiment and a LLM-enabled replicate experiment, demonstrate that 1) Larger and more recent LLMs tend to show a higher level of consistency and accuracy in distinguishing credible information from misinformation. However, they are more likely to give higher ratings for misinformation due to the presence of a more salient, decoy misinformation result; 2) While decoy effect occurred in both human and LLM assessments, the effect is more prevalent across different conditions and topics in LLM judgments compared to human credibility ratings. In contrast to the generally assumed "rationality" of AI tools, our study empirically confirms the cognitive bias risks embedded in LLM agents, evaluates the decoy impact on LLMs against human credibility assessments, and thereby highlights the complexity and importance of debiasing AI agents and developing psychology-informed AI audit techniques and policies for automated judgment tasks and beyond.

Figures

Figures reproduced from arXiv: 2411.15396 by the authors.

Figure 1
Figure 1. Experimental Conditions. A decoy result was added to EACH SERP in treatment sessions. Under [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Crowdsourcing Experiment Flow. , Vol. 1, No. 1, Article . Publication date: November 2024 [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. LLM-based Experiment Flow. • Llama-3-8B: A smaller model in the Llama series with 8 billion parameters. The training data have a cutoff date of March 2023 6 . • gpt-4o: OpenAI’s current flagship model. The snapshot used in this study is gpt-4o-2024-05-13. The training data are up to October 2023. • gpt-4o-mini: A small and cheaper model for light tasks offered by OpenAI. The snapshot used in this study is gpt-4o-min… view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Mean ratings by topics and article types for the control groups. [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Mean ratings by topics and article types for the treatment groups. [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]
Figure 6
Figure 6. Figure 6: Average ratings for target articles by query/SERP sequence in single-query iterations, where each [PITH_FULL_IMAGE:figures/full_fig_p019_6.png]
Figure 7
Figure 7. Figure 7: Average ratings for target articles by query/SERP sequence in multiple-query iterations. [PITH_FULL_IMAGE:figures/full_fig_p020_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 69 canonical work pages

  1. [1]

    Zahra Abbasiantaeb, Yifei Yuan, Evangelos Kanoulas, and Mohammad Aliannejadi. 2024. Let the llms talk: Simulating human-to-human conversational qa via zero-shot llm-to-llm interactions. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining . 8–17

  2. [2]

    Azzah Al-Maskari and Mark Sanderson. 2010. A review of factors influencing user satisfaction in information retrieval. Journal of the American Society for Information Science and Technology 61, 5 (2010), 859–868

  3. [3]

    Leif Azzopardi. 2021. Cognitive biases in search: a review and reflection of cognitive biases in Information Retrieval. In Proceedings of the 2021 ACM SIGIR Conference on Human Information Interaction and Retrieval . 27–37

  4. [4]

    Nicholas J Belkin. 2016. People, interacting with information. In ACM SIGIR Forum, Vol. 49. ACM New York, NY, USA, 13–27

  5. [5]

    Nicholas J Belkin, Michael Cole, and Ralf Bierig. 2008. Is relevance the right criterion for evaluating interactive information retrieval. In Proceedings of the ACM SIGIR 2008 Workshop on Beyond Binary Relevance: Preferences, Diversity, and Set-Level Judgments. http://research. microsoft. com/˜ pauben/bbr-workshop . Citeseer

  6. [6]

    Markus Bink, Steven Zimmerman, and David Elsweiler. 2022. Featured snippets and their influence on users’ credibility judgements. In Proceedings of the 2022 Conference on Human Information Interaction and Retrieval . 113–122

  7. [7]

    Jennifer S Blumenthal-Barby and Heather Krieger. 2015. Cognitive biases and heuristics in medical decision making: a critical review using a systematic search strategy. Medical Decision Making 35, 4 (2015), 539–557

  8. [8]

    Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024. Humans or llms as the judge? a study on judgement biases. arXiv preprint arXiv:2402.10669 (2024)

Show all 89 references
  1. [9]

    Nuo Chen, Jiqun Liu, Xiaoyu Dong, Qijiong Liu, Tetsuya Sakai, and Xiao-Ming Wu. 2024. AI Can Be Cognitively Biased: An Exploratory Study on Threshold Priming in LLM-Based Batch Relevance Assessment. arXiv preprint arXiv:2409.16022 (2024)

  2. [10]

    Nuo Chen, Jiqun Liu, Hanpei Fang, Yuankai Luo, Tetsuya Sakai, and Xiao-Ming Wu. 2024. Decoy Effect In Search Interaction: Understanding User Behavior and Measuring System Vulnerability. arXiv preprint arXiv:2403.18462 (2024)

  3. [11]

    Nuo Chen, Jiqun Liu, and Tetsuya Sakai. 2023. A Reference-Dependent Model for Web Search Evaluation: Understanding and Measuring the Experience of Boundedly Rational Users. InProceedings of the ACM Web Conference 2023. 3396–3405

  4. [12]

    Nuo Chen, Jiqun Liu, Tetsuya Sakai, and Xiao-Ming Wu. 2023. Decoy Effect in Search Interaction: A Pilot Study. arXiv preprint arXiv:2311.02362 (2023). , Vol. 1, No. 1, Article . Publication date: November 2024. The Decoy Dilemma in Online Medical Information Evaluation 23

  5. [13]

    Ye Chen, Ke Zhou, Yiqun Liu, Min Zhang, and Shaoping Ma. 2017. Meta-evaluation of online and offline web search evaluation metrics. In Proceedings of the 40th international ACM SIGIR conference on research and development in information retrieval. 15–24

  6. [14]

    Michael Cole, Jingjing Liu, Nicholas J Belkin, Ralf Bierig, Jacek Gwizdka, Chang Liu, Jin Zhang, and Xiangmin Zhang

  7. [15]

    Terry Connolly, Jochen Reb, and Edgar E Kausel. 2013. Regret salience and accountability in the decoy effect.Judgment and Decision making 8, 2 (2013), 136–149

  8. [16]

    Sunhao Dai, Chen Xu, Shicheng Xu, Liang Pang, Zhenhua Dong, and Jun Xu. 2024. Bias and unfairness in information retrieval systems: New challenges in the llm era. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 6437–6447

  9. [17]

    Nicola Diviani, Bas van den Putte, Stefano Giani, and Julia CM van Weert. 2015. Low health literacy and evaluation of online health information: a systematic review of the literature. Journal of medical Internet research 17, 5 (2015), e112

  10. [18]

    Carsten Eickhoff. 2018. Cognitive biases in crowdsourcing. In Proceedings of the eleventh ACM international conference on web search and data mining . 162–170

  11. [19]

    Guglielmo Faggioli, Laura Dietz, Charles LA Clarke, Gianluca Demartini, Matthias Hagen, Claudia Hauff, Noriko Kando, Evangelos Kanoulas, Martin Potthast, Benno Stein, et al . 2023. Perspectives on large language models for relevance judgment. In Proceedings of the 2023 ACM SIG...

  12. [20]

    Andrew J Flanagin, Stephan Winter, and Miriam J Metzger. 2020. Making sense of credibility in complex information environments: the role of message sidedness, information source, and thinking styles in credibility evaluation online. Information, Communication & Society 23, 7 (...

  13. [21]

    Donna Harman. 2011. Information retrieval evaluation. Morgan & Claypool Publishers

  14. [22]

    Brian Hilligoss and Soo Young Rieh. 2008. Developing a unifying framework of credibility assessment: Construct, heuristics, and interaction in context. Information Processing & Management 44, 4 (2008), 1467–1484

  15. [23]

    Katja Hofmann, Lihong Li, Filip Radlinski, et al. 2016. Online evaluation for information retrieval. Foundations and Trends® in Information Retrieval 10, 1 (2016), 1–117

  16. [24]

    Jianping Hu and Rongjun Yu. 2014. The neural correlates of the decoy effect in decisions. Frontiers in behavioral neuroscience 8 (2014), 271

  17. [25]

    Joel Huber, John W Payne, and Christopher Puto. 1982. Adding asymmetrically dominated alternatives: Violations of regularity and the similarity hypothesis. Journal of consumer research 9, 1 (1982), 90–98

  18. [26]

    John PA Ioannidis, Michael E Stuart, Shannon Brownlee, and Sheri A Strite. 2017. How to survive the medical misinformation mess. European journal of clinical investigation 47, 11 (2017), 795–802

  19. [27]

    Kalervo Järvelin and Jaana Kekäläinen. 2017. IR evaluation methods for retrieving highly relevant documents. In ACM SIGIR Forum, Vol. 51. ACM New York, NY, USA, 243–250

  20. [28]

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. Comput. Surveys 55, 12 (2023), 1–38

  21. [29]

    Yong Ju Jung, Jiqun Liu, Harun Karahan, and Mahdieh Nazari. 2024. Tech Treasure Hunt: Promoting Children’s Learning on How Algorithm Works at a Public Library. Proceedings of the Association for Information Science and Technology 61, 1 (2024), 962–964

  22. [30]

    Markus Kattenbeck and David Elsweiler. 2019. Understanding credibility judgements for web search snippets. Aslib Journal of Information Management (2019)

  23. [31]

    Diane Kelly et al. 2009. Methods for evaluating interactive information retrieval systems with users. Foundations and Trends® in Information Retrieval 3, 1–2 (2009), 1–224

  24. [32]

    Diane Kelly and Leif Azzopardi. 2015. How many results per page? A study of SERP size, search behavior and user experience. In Proceedings of the 38th international ACM SIGIR conference on research and development in information retrieval. 183–192

  25. [33]

    Youngwoo Kim, Razieh Rahimi, and James Allan. 2022. Alignment Rationale for Query-Document Relevance. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2489–2494

  26. [34]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing S...

  27. [35]

    Xiaonan Li, Kai Lv, Hang Yan, Tianyang Lin, Wei Zhu, Yuan Ni, Guotong Xie, Xiaoling Wang, and Xipeng Qiu. 2023. Unified demonstration retriever for in-context learning. arXiv preprint arXiv:2305.04320 (2023)

  28. [36]

    Sook Lim. 2013. College students’ credibility judgments and heuristics concerning Wikipedia. Information Processing & Management 49, 2 (2013), 405–419. , Vol. 1, No. 1, Article . Publication date: November 2024. 24 Anonymous

  29. [37]

    Jiqun Liu. 2021. Deconstructing search tasks in interactive information retrieval: A systematic review of task dimensions and predictors. Information Processing & Management 58, 3 (2021), 102522

  30. [38]

    Jiqun Liu. 2022. Toward Cranfield-inspired reusability assessment in interactive information retrieval evaluation. Information Processing & Management 59, 5 (2022), 103007

  31. [39]

    Jiqun Liu. 2023. A Behavioral Economics Approach to Interactive Information Retrieval: Understanding and Supporting Boundedly Rational Users. Vol. 48. Springer Nature

  32. [40]

    Jiqun Liu. 2023. Toward A Two-Sided Fairness Framework in Search and Recommendation. In Proceedings of the 2023 Conference on Human Information Interaction and Retrieval . 236–246

  33. [41]

    Jiqun Liu and Leif Azzopardi. 2024. Search under uncertainty: Cognitive biases and heuristics: a tutorial on testing, mitigating and accounting for cognitive biases in search experiments. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development...

  34. [42]

    Jiqun Liu and Fangyuan Han. 2020. Investigating reference dependence effects on user search interaction and satisfaction: A behavioral economics perspective. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval . 1141–1150

  35. [43]

    Shengjie Ma, Chong Chen, Qi Chu, and Jiaxin Mao. 2024. Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval. arXiv preprint arXiv:2403.18405 (2024)

  36. [44]

    N Gregory Mankiw. 2014. Principles of economics. Cengage Learning

  37. [45]

    Stefan Palan and Christian Schitter. 2018. Prolific. ac—A subject pool for online experiments. Journal of behavioral and experimental finance 17 (2018), 22–27

  38. [46]

    Gordon Pennycook and David G Rand. 2019. Fighting misinformation on social media using crowdsourced judgments of news source quality. Proceedings of the National Academy of Sciences 116, 7 (2019), 2521–2526

  39. [47]

    Jonathan C Pettibone and Douglas H Wedell. 2000. Examining models of nondominated decoy effects across judgment and choice. Organizational behavior and human decision processes 81, 2 (2000), 300–328

  40. [48]

    Sayantan Polley, Rashmi Raju Koparde, Akshaya Bindu Gowri, Maneendra Perera, and Andreas Nuernberger. 2021. Towards trustworthiness in the context of explainable search. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Re...

  41. [49]

    Kashyap Popat, Subhabrata Mukherjee, Jannik Strötgen, and Gerhard Weikum. 2017. Where the truth lies: Explaining the credibility of emerging claims on the web and social media. In Proceedings of the 26th International Conference on World Wide Web Companion. 1003–1012

  42. [50]

    Hossein A Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, and Daniel Campos. 2024. Synthetic Test Collections for Retrieval Evaluation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2647–2651

  43. [51]

    Tetsuya Sakai and Zhaohao Zeng. 2020. Retrieval evaluation measures that agree with users’ SERP preferences: Traditional, preference-based, and diversity measures. ACM Transactions on Information Systems (TOIS) 39, 2 (2020), 1–35

  44. [52]

    Alireza Salemi and Hamed Zamani. 2024. Evaluating Retrieval Quality in Retrieval-Augmented Generation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2395–2400

  45. [53]

    Tefko Saracevic. 2007. Relevance: A review of the literature and a framework for thinking on the notion in information science. Part II: Nature and manifestations of relevance. Journal of the American society for information science and technology 58, 13 (2007), 1915–1933

  46. [54]

    Laura Sbaffi and Jennifer Rowley. 2017. Trust and credibility in web-based health information: a review and agenda for future research. Journal of medical Internet research 19, 6 (2017), e218

  47. [55]

    Falk Scholer, Diane Kelly, Wan-Ching Wu, Hanseul S Lee, and William Webber. 2013. The effect of threshold priming and need for cognition on relevance calibration and assessment. In Proceedings of the 36th international ACM SIGIR conference on Research and development in inform...

  48. [56]

    Ivan Sekulić, Mohammad Alinannejadi, and Fabio Crestani. 2024. Analysing utterances in llm-based user simulation for conversational search. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–22

  49. [57]

    Itamar Simonson. 1989. Choice based on reasons: The case of attraction and compromise effects. Journal of consumer research 16, 2 (1989), 158–174

  50. [58]

    Saira Hanif Soroya, Ali Farooq, Khalid Mahmood, Jouni Isoaho, and Shan-e Zara. 2021. From information seeking to information avoidance: Understanding the health information behavior during a global health crisis. Information processing & management 58, 2 (2021), 102440

  51. [59]

    Brian G Southwell, Jeff Niederdeppe, Joseph N Cappella, Anna Gaysynsky, Dannielle E Kelley, April Oh, Emily B Peterson, and Wen-Ying Sylvia Chou. 2019. Misinformation as a misunderstood challenge to public health. American journal of preventive medicine 57, 2 (2019), 282–285. ...

  52. [60]

    Sandro Tiziano Stoffel, Jiahong Yang, Ivo Vlaev, and Christian von Wagner. 2019. Testing the decoy effect to increase interest in colorectal cancer screening. PloS one 14, 3 (2019), e0213668

  53. [61]

    Yalin Sun, Yan Zhang, Jacek Gwizdka, and Ciaran B Trace. 2019. Consumer evaluation of the quality of online health information: systematic literature review of relevant criteria and indicators. Journal of medical Internet research 21, 5 (2019), e12522

  54. [62]

    Briony Swire-Thompson, David Lazer, et al. 2020. Public health and online misinformation: challenges and recommen- dations. Annu Rev Public Health 41, 1 (2020), 433–451

  55. [63]

    Richard H Thaler. 2016. Behavioral economics: Past, present, and future. American Economic Review 106, 7 (2016), 1577–1600

  56. [64]

    Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. 2023. Large language models in medicine. Nature medicine 29, 8 (2023), 1930–1940

  57. [65]

    Paul Thomas, Seth Spielman, Nick Craswell, and Bhaskar Mitra. 2024. Large language models can accurately predict searcher preferences. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1930–1940

  58. [66]

    Amos Tversky and Daniel Kahneman. 1985. The framing of decisions and the psychology of choice . Springer

  59. [67]

    Michela Del Vicario, Walter Quattrociocchi, Antonio Scala, and Fabiana Zollo. 2019. Polarization and fake news: Early warning of potential misinformation targets. ACM Transactions on the Web (TWEB) 13, 2 (2019), 1–22

  60. [68]

    Ellen M Voorhees. 2019. The evolution of cranfield. Information Retrieval Evaluation in a Changing World: Lessons Learned from 20 Years of CLEF (2019), 45–69

  61. [69]

    Ben Wang and Jiqun Liu. 2024. Cognitively Biased Users Interacting with Algorithmically Biased Results in Whole- Session Search on Debated Topics. InProceedings of the 2024 ACM SIGIR International Conference on Theory of Information Retrieval. 227–237

  62. [70]

    Ben Wang and Jiqun Liu. 2024. Understanding users’ dynamic perceptions of search gain and cost in sessions: An expectation confirmation model. Journal of the Association for Information Science and Technology (2024)

  63. [71]

    Shoujin Wang, Xiuzhen Zhang, Yan Wang, and Francesco Ricci. 2024. Trustworthy recommender systems. ACM Transactions on Intelligent Systems and Technology 15, 4 (2024), 1–20

  64. [72]

    Xi Wang, Hossein A Rahmani, Jiqun Liu, and Emine Yilmaz. 2023. Improving Conversational Recommendation Systems via Bias Analysis and Language-Model-Enhanced Data Augmentation. arXiv preprint arXiv:2310.16738 (2023)

  65. [73]

    Xiaohui Wang, Jingyuan Shi, and Hanxiao Kong. 2021. Online health information seeking: a review and meta-analysis. Health Communication 36, 10 (2021), 1163–1175

  66. [74]

    Douglas H Wedell and Jonathan C Pettibone. 1996. Using judgments to understand decoy effects in choice. Organiza- tional Behavior and Human Decision Processes 67, 3 (1996), 326–344

  67. [75]

    Ryen White. 2013. Beliefs and biases in web search. In Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval . 3–12

  68. [76]

    Chunhua Wu and Koray Cosguner. 2020. Profiting from the decoy effect: A case study of an online diamond retailer. Marketing Science 39, 5 (2020), 974–995

  69. [77]

    Dan Wu, Jing Dong, Li Shi, Chunxiang Liu, and Jiangyun Ding. 2020. Credibility assessment of good abandonment results in mobile search. Information Processing & Management 57, 6 (2020), 102350

  70. [78]

    Yusuke Yamamoto and Katsumi Tanaka. 2011. Enhancing credibility judgment of web search results. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . 1235–1244

  71. [79]

    Yichao Yuan and Tiaojun Xiao. 2022. Retailer’s decoy strategy versus consumers’ reference price effect in a retailer- Stackelberg supply chain. Journal of Retailing and Consumer Services 68 (2022), 103081

  72. [80]

    ChengXiang Zhai. 2024. Large language models and future of information retrieval: Opportunities and challenges. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 481–490

  73. [81]

    Dake Zhang, Amir Vakili Tahami, Mustafa Abualsaud, and Mark D Smucker. 2022. Learning trustworthy web sources to derive correct answers and reduce health misinformation in search. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Info...

  74. [82]

    Fan Zhang, Jiaxin Mao, Yiqun Liu, Xiaohui Xie, Weizhi Ma, Min Zhang, and Shaoping Ma. 2020. Models versus satisfaction: Towards a better understanding of evaluation metrics. In Proceedings of the 43rd international acm sigir conference on research and development in informatio...

  75. [83]

    Hengran Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. 2024. Are Large Language Models Good at Utility Judgments?. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 1941–1951

  76. [84]

    Yan Zhang, Jiaying Liu, and Shijie Song. 2023. The design and evaluation of a nudge-based interface to facilitate consumers’ evaluation of online health information credibility. Journal of the Association for Information Science and Technology 74, 7 (2023), 828–845. , Vol. 1, ...

  77. [85]

    Yan Zhang, Yalin Sun, and Bo Xie. 2015. Quality of health information for consumers on the web: a systematic review of indicators, criteria, tools, and evaluation results. Journal of the Association for Information Science and Technology 66, 10 (2015), 2071–2084

  78. [86]

    Shanshan Zhen and Rongjun Yu. 2016. The development of the asymmetrically dominated decoy effect in young children. Scientific reports 6, 1 (2016), 1–7

  79. [87]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 (2023), 46595–46623

  80. [88]

    Aljaž Zrnec, Marko Poženel, and Dejan Lavbič. 2022. Users’ ability to perceive misinformation: An information quality assessment approach. Information Processing & Management 59, 1 (2022), 102739. , Vol. 1, No. 1, Article . Publication date: November 2024

  81. [2009]

    InProceedings of the Third Workshop on Human-Computer Interaction and Information Retrieval Cambridge

    Usefulness as the criterion for evaluation of interactive information retrieval. InProceedings of the Third Workshop on Human-Computer Interaction and Information Retrieval Cambridge . HCIR, 1–4

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.