{"id":"60c1b67c-d1be-4fc2-a14c-be652d60a02c","arxiv_id":"2501.01303","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Citations increase self-reported trust in LLM answers, even random ones, while checking citations is associated with lower trust.","lead":"This paper ran a live experiment where 303 people asked questions to a ChatGPT-based chatbot and rated their trust in the answers, with zero, one, or five citations attached. It found that citations raise trust, even when the citations are random, and that people who checked the citations trusted the answers less.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The random-citation claim is not directly tested: Table S2's 'Citation Random' coefficient contrasts random with valid citations, and the only reported random-vs-zero comparison is limited to checked questions and is nonsignificant (p=0.38).","rationale":"The reader's weakest assumption was that mouse hover is an unvalidated proxy for deliberate citation checking. That is a genuine measurement concern: hovers can be accidental, and the paper does not validate hover against any other indicator of inspection. However, even if hover were a perfect measure of checking, the paper's random-citation headline would still be unsupported, because the only reported test comparing random citations with zero citations is restricted to checked questions and is nonsignificant. The regression coefficient that might look like a random-vs-zero effect actually compares random to valid citations. Thus the single most load-bearing issue is the missing direct contrast for the random-citation claim, not the hover operationalization. The paper's other main results, such as the positive effect of citations in general and the negative correlation between checking and trust, are plausible and supported by the reported regressions, so the appropriate verdict remains CONDITIONAL pending the direct random-vs-zero contrast and the validation of the check measure. I therefore do not propose changing the reader's verdict, only sharpening the condition attached to it.","tokens_in":14234,"tokens_out":4518,"duration_ms":48640,"concrete_test":"Using the archived dataset, estimate the random-vs-zero contrast directly: run a linear regression of trust on indicator variables for the five experimental cells (one-valid, one-random, five-valid, five-random, zero as baseline), controlling for the same demographic covariates as Table S2, and report the point estimate, 95% confidence interval, and p-value for the pooled random-citations versus zero-citations contrast across all citation-present responses, not just checked ones. If the confidence interval excludes zero in the positive direction, the abstract claim survives; if not, the headline must be weakened or removed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim that the trust increase 'held true even when the citations were random' is not supported by any reported direct statistical contrast. In Table S2, the model includes both 'Has Citation' and 'Citation Random'; the coefficient on 'Citation Random' (-0.268, SE 0.087) therefore estimates the difference between random and valid citations, not between random and zero citations. A random-vs-zero comparison would require testing the joint contrast beta_HasCitation + beta_RandomCitation, whose standard error and p-value are never reported. The only direct comparison described in the main text is in the section 'Does Checking Citations Indicate a Reduction in User Trust?', where the authors compare zero-citation answers with random-citation answers that were manually checked: T=-0.877, p=0.38, with mean trust 7.73 for no citations versus 7.55 for checked random citations. This subset analysis is nonsignificant and actually points in the opposite direction. Because the random-citation result is a headline finding in the abstract and introduction, the missing contrast is load-bearing: the paper currently provides no statistically valid evidence for the 'even random citations increase trust' claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a between-subjects randomized experiment (N=303; roughly 3,040 question-level ratings) in which participants asked a custom ChatGPT-based QA system ten questions and rated their trust in each answer. The system varied the number of citations (0, 1, or 5) and, for non-zero conditions, whether the citations were relevant to the answer or randomly drawn from previous participants' queries. The paper reports that citations increase self-reported trust, that random citations are rated lower than valid ones, that one and five citations do not differ, and that participants who hover over citations give lower trust ratings, which the authors interpret as support for trust-as-anti-monitoring. Additional exploratory analyses consider question type, demographics, and prompt perplexity.","tokens_in":14441,"tokens_out":6539,"duration_ms":65847,"significance":"If the presence-of-citations effect is robust, this is a useful empirical contribution to the literature on trust in LLM-generated content. The study's strengths include a live question-answering task with real user-generated questions, an experimental manipulation of citation count and relevance, and public availability of the data and Stata code. However, the headline claim that the trust increase 'held true even when the citations were random' is not supported by any direct statistical contrast in the reported analyses, the reported standard errors likely ignore participant-level clustering of repeated ratings, and the 'checking' measure is a mouse hover rather than a validated deliberate inspection. These issues are fixable with additional analyses and wording changes, but they currently limit the strength of the central conclusions.","major_comments":[{"comment":"The abstract and introduction claim that the trust increase 'held true even when the citations were random' (Abstract; Introduction). I could not find a direct test of random citations versus no citations. In Table S2, 'Has Citation' and 'Citation Random' enter the same model; the coefficient on 'Citation Random' (-0.268, SE 0.087) contrasts random with valid citations, not random with zero citations, because the zero-citation group is the omitted baseline for 'Has Citation'. To support the headline claim, report the joint contrast (Has Citation + Citation Random) with its standard error and p-value, or run an explicit random-versus-zero model on the full sample. The only direct comparison in the main text, in 'Does Checking Citations Indicate a Reduction in User Trust?', is restricted to the 193 checked questions and is nonsignificant (T=-0.877, p=0.38); this does not test the full-sample claim and in fact trends in the opposite direction.","section":"Abstract; Results, 'Does the Quality of Citation Matter?'; Table S2"},{"comment":"The regressions and ANOVA treat the roughly 3,040 question-level ratings as independent observations, but each of the 303 participants contributes ten ratings in a single session. The between-subjects assignment of citation condition does not make these repeated ratings independent. The reported standard errors (e.g., Has Citation β=0.394, SE=0.0906 in Table S2) are therefore likely understated, and the p-values for the main effects may be too small. Please re-estimate the main models with cluster-robust standard errors at the participant level or with participant random effects, and report whether the presence-of-citations and checking effects remain significant.","section":"Results, Tables S2, S3, and S4"},{"comment":"The operational measure of 'checking' a citation is a mouse hover over the citation numeral, which reveals the URL but requires no click. The anti-monitoring conclusion ('checking citations decrease perceived trust') rests on treating hover as deliberate monitoring. Because hovers can be accidental or cursory, please validate the measure (e.g., click-through, dwell time, or a sensitivity analysis excluding very short hovers) or soften the claim to what the data support. In addition, Table S4 models Citation Checks as a function of Trust, so the data are consistent with lower trust leading to checking rather than checking causing lower trust; the abstract's phrasing 'decrease in self-reported user trust when participants checked the citations' should be revised to describe the associational direction actually tested.","section":"Methodology; Results, 'Does Checking Citations Indicate a Reduction in User Trust?'; Table S4"}],"minor_comments":[{"comment":"The model name 'ChatGPT4' should be 'ChatGPT-4' or 'GPT-4' for consistency with the cited technical report.","section":"Throughout"},{"comment":"The first sentence contains a typo: 'In out initial analysis' should read 'In our initial analysis'.","section":"Results, 'Do Citations Increase User Trust?'"},{"comment":"The labels 'Participants Demographics SurveyResponse' (Fig. 2) and 'T rust' (Fig. 4) should be corrected for readability.","section":"Figures 2 and 4"},{"comment":"The text says 'as illustrated in Table 6' but the item is a figure (Fig. 6); the cross-reference should be fixed.","section":"Supplement, 'Question Order, Citation Checking, and Trust'"},{"comment":"The caption states 'random citations decrease perceived trustworthiness'; this is only a decrease relative to valid citations, not relative to zero citations. Please state the reference category explicitly in the caption.","section":"Figure 3 caption"}],"recommendation":"major_revision","confidential_remarks":"The missing random-versus-zero contrast is the most serious issue because it affects the paper's central advertised finding; however, it is addressable with a re-analysis of the existing data rather than new data collection. The clustering and hover-measure concerns are also fixable within the manuscript's scope. I therefore see major revision rather than rejection as the appropriate outcome."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real experiment with a solid main effect, but the abstract's boldest claim—random citations still boost trust—is not backed by any reported test. The regression's \"Citation Random\" coefficient compares random to valid citations, not random to zero. The only direct zero-vs-random comparison in the paper is a subset of checked questions and is nonsignificant (T=-0.877, p=0.38), trending the wrong way. That needs to be fixed before the abstract can say what it says.\n\nWhat's genuinely useful: the authors are the first, in the literature they cite, to separate citation presence, count, and relevance in a live LLM QA setting. The random-citation control is a good idea, and the hover-based check measure is a clever behavioral proxy. The main presence-of-citations effect is credible: Has Citation beta 0.394 (SE 0.0906), consistent with the ANOVA across zero/one/five. The one-vs-five null is reported honestly. The paper also does the field a service by showing that citation count beyond one doesn't add trust.\n\nSoft spots, in proportion. (1) The random-citation claim is load-bearing and untested. Report the joint contrast (Has Citation + Random Citation) or soften the conclusion. (2) Hover is not the same as deliberate checking. Hovering to reveal a URL requires no click and may be accidental; the anti-monitoring story depends on this equivalence. The paper itself calls it \"manually checked\" without validating the measure. (3) Trust and checking are measured concurrently, so the inverse relationship could reflect lower trust causing checking, or checking causing lower trust; the theory assumes direction. (4) Anonymized data and materials are promised but redacted in the preprint; release them. None of these sink the main effect; the presence-of-citations result rests on standard regression and ANOVA.\n\nWho this is for: HCI and RAG researchers designing chatbot interfaces, and anyone studying trust in AI output. It deserves a serious referee. Recommend peer review with the condition that the abstract be matched to the analysis and the data be made available.","headline":"Solid experiment, overstated abstract: the random-citation headline isn't directly tested, but the citation-presence effect and the honest one-vs-five null make it worth referee time.","tokens_in":14979,"tokens_out":2860,"would_cite":true,"duration_ms":26854,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Citations raise trust in AI chatbot answers even when they are random, while checking them signals distrust.","keywords":["citations","user trust","large language models","anti-monitoring","social proof","retrieval-augmented generation","question answering","human-AI interaction"],"falsifier":"A replication that requires participants to click a citation to see the URL, and records click counts and dwell time instead of hovers, would test the anti-monitoring account: if click-based checking shows no negative relationship with trust, or if accidental hovers alone reproduce the effect, the claim that checking indicates lower trust would be falsified.","tokens_in":14055,"feed_emoji":"🔗","tokens_out":5798,"duration_ms":53022,"temperature":0.7,"pith_summary":"This paper asks whether showing sources in an AI chatbot's answer changes how much people trust the answer. In a live experiment with 303 participants, answers with citations received higher self-reported trust than answers without, and even citations chosen at random lifted trust. The number of citations did not matter: one and five produced similar ratings. The study also reports that participants who inspected citations by hovering over them gave lower trust ratings, consistent with the anti-monitoring view that checking is a sign of distrust. The authors conclude that citations function as social proof for AI-generated content, while the act of verification signals the absence of trust.","feed_headline":"Any citation raises trust in AI answers — even a random one","feed_subtitle":"A 303-person experiment finds that sources boost perceived trust, while inspecting a source predicts lower trust.","key_machinery":"The central machinery is the anti-monitoring theory of trust, paired with the Principle of Social Proof: trust is inferred from a reduction in surveillance, and citations act as visible endorsements that substitute for direct verification. The experiment operationalizes the theory by treating a mouse hover over a citation numeral as a monitoring event and a 1–10 rating as stated trust, then regressing trust on citation presence, citation relevance, and hover behavior while controlling for demographics. The load-bearing identity is the negative correlation between checking and trust: if hover frequency did not predict lower ratings, the anti-monitoring interpretation would lose its empirical support.","core_discovery":"On the paper's own terms, the discovery is that trust in LLM-generated answers is driven more by the presence of a citation than by its quality: a randomized controlled trial with zero, one, or five citations, valid or random, found that 'has citation' significantly increased trust ratings, 'random citation' significantly decreased them relative to valid ones, and one versus five citations made no significant difference. A separate regression found that each citation check (mouse hover) was associated with significantly lower reported trust, and random citations that were checked lost the trust advantage entirely, being rated no better than answers with no citations. The authors interpret this asymmetry through trust as anti-monitoring: citations provide social proof that raises trust, while monitoring the citations indicates that trust is absent.","pith_inferences":["An implication the authors do not draw: if random citations raise trust, then citations are functioning as a credibility cue rather than as verifiable evidence, which makes fabricated citations a direct trust-exploitation risk in deployed systems.","The one-versus-five null suggests users apply a binary 'is there a source' heuristic; a natural extension is to test whether citation count interacts with perceived source authority or domain risk.","Because hover conflates suspicion with curiosity, a click-to-reveal design would test whether the anti-monitoring result reflects distrust or mere exploration; it is a testable boundary condition on the paper's central claim.","The result that checked random citations lose the trust premium implies that transparency tools that surface citations may actually lower trust if the sources are weak, challenging the assumption that more transparency is always better."],"forward_implications":["Answers with any citation tend to be rated more trustworthy than identical answers with no citation, even when the cited sources are unrelated to the answer.","A single citation is as effective as five, so increasing citation count beyond one is unlikely to buy additional trust.","Users who inspect citations report lower trust, and random citations that get inspected lose their trust advantage, suggesting that checking neutralizes the social-proof effect.","Question content shifts trust: political and factual questions receive higher trust ratings, while more complex or longer prompts receive slightly lower ratings."],"supporting_citations":[{"why":"Supplies the definition of trust as anti-monitoring, the theoretical lens that predicts checking a citation indicates reduced trust.","marker":"Baier 1986"},{"why":"Supplies the Principle of Social Proof, the mechanism by which citations are hypothesized to raise trust.","marker":"Cialdini 2009"},{"why":"Establishes retrieval-augmented generation as the context in which citations are produced, motivating the study.","marker":"Lewis et al. 2020"},{"why":"Documents the GPT-4 system used to generate the experimental responses.","marker":"Achiam et al. 2023"},{"why":"Validates Prolific as the participant recruitment platform for the online experiment.","marker":"Palan and Schitter 2018"},{"why":"Provides the trust-in-automation background on factors shaping user trust in AI systems.","marker":"Hoff and Bashir 2015"},{"why":"Distinguishes questionnaire-based from behavioral trust measures, motivating the paper's mix of self-report ratings and hover-based behavioral tracking.","marker":"Poursabzi-Sangdeh et al. 2021"}],"fun_headline_variants":["Even random citations boost trust in AI answers","Inspecting a citation lowers trust in AI answers","Citation presence, not correctness, lifts AI trust","One citation is enough to raise AI answer trust"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that a mouse hover over a citation numeral counts as 'checking' the citation: hovering is effortless, can be accidental, and was not validated against deliberate inspection, yet the anti-monitoring conclusion depends entirely on that equation.","fun_headline_variants_meta":{"raw":{"variants":["Even random citations boost trust in AI answers","Inspecting a citation lowers trust in AI answers","Citation presence, not correctness, lifts AI trust","One citation is enough to raise AI answer trust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000896,"raw_usage":{"total_tokens":3792,"prompt_tokens":810,"completion_tokens":2982,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":426,"completion_tokens_details":{"reasoning_tokens":2924}},"tokens_in":426,"tokens_out":2982,"duration_ms":20513,"temperature":1.0,"reasoning_tokens":2924,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:29:25.858289+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that requires participants to click a citation to see the URL, and records click counts and dwell time instead of hovers, would test the anti-monitoring account: if click-based checking shows no negative relationship with trust, or if accidental hovers alone reproduce the effect, the claim that checking indicates lower trust would be falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the definition of trust as anti-monitoring, the theoretical lens that predicts checking a citation indicates reduced trust."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Principle of Social Proof, the mechanism by which citations are hypothesized to raise trust."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Validates Prolific as the participant recruitment platform for the online experiment."},{"cited_title":"A.; and Bashir, M","cited_arxiv_id":null,"evidence_quote":"Provides the trust-in-automation background on factors shaping user trust in AI systems."},{"cited_title":"G.; Hofman, J","cited_arxiv_id":null,"evidence_quote":"Distinguishes questionnaire-based from behavioral trust measures, motivating the paper's mix of self-report ratings and hover-based behavioral tracking."}],"review_version":1}