{"id":"cf99b727-3da9-4da6-9b90-e18cbd4657ea","arxiv_id":"2607.14152","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Accurate but irrelevant 'trust junk' in AI explanations made crowdsourced users trust and agree with a deliberately discriminatory model more.","lead":"A crowdsourced experiment found that adding accurate but irrelevant details to an AI explanation made people trust and agree with a deliberately discriminatory model more. The study suggests XAI designers must treat explanations as persuasive artifacts, not just information delivery.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The eight-profile task may not let participants detect the model's race/gender rule, so baseline 'unjustified trust' and the dose-response may reflect uninformative stimuli rather than trust junk.","rationale":"I read the paper as an attempt to demonstrate that technically accurate but fairness-irrelevant explanatory data causally increases lay users' trust in and agreement with a model that is objectively worse than a trivial baseline. The central empirical claims (agreement, trust, satisfaction) are supported with significant between-group differences, and the qualitative coding shows reduced spontaneous fairness concerns in the junk conditions. The internal logic of the study is sound. My concern is not with the statistics (which can be corrected via multiplicity adjustments) but with the interpretability of the baseline condition. The paper repeatedly describes the model as 'patently discriminatory,' yet participants were shown only eight profiles in random order. If the profile set does not contain enough instances of the protected class (Black men) or enough model errors, participants have no way to detect the rule, making the baseline agreement an artifact of the high base rate (all-pass would be 94.8%) rather than a measure of unwarranted trust. The authors' own data show that only 42.9% of baseline participants mentioned race/gender issues and none fully described the rule, so the model's discrimination was not patent to lay viewers. This gap is load-bearing because the 'unjustified' qualifier in the central claim implies participants should have been able to see what was wrong. Without that, the dose-response could reflect mere increases in perceived informativeness — a less novel phenomenon — and the fairwashing effect becomes harder to distinguish from general automation bias. The concrete test — inspecting the supplement to count Black male profiles and model errors — directly settles whether the baseline offered a real opportunity to detect discrimination. If the profiles do include enough instances, the concern is mitigated and the paper's interpretation stands. If not, the central claim should be re-scoped to 'automation bias' rather than 'trust junk specifically.' This is consistent with the reader's conditional verdict: the paper is otherwise well-executed and the issue is addressable.","tokens_in":9430,"tokens_out":11742,"duration_ms":245052,"concrete_test":"Access the supplemental materials (https://osf.io/wufqz/) and extract the exact eight candidate profiles used in the study. Count the number of Black male candidates and the number of model errors (cases where the model's prediction differs from ground truth). If there are fewer than three Black male candidates or fewer than two model errors, then participants had little opportunity to detect the rule. Re-run the main analyses restricted to participants who saw at least one Black male candidate, or add a control condition where the model's rule is explicitly disclosed and compare trust/agreement to the current conditions. If the effects persist only among participants who saw the discriminatory pattern, the 'unjustified trust' interpretation is strengthened; if not, the effect is better described as general automation bias rather than trust junk.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that trust junk causes 'unjustified support' for a patently discriminatory model rests on the assumption that participants could, in principle, see the model's race/gender rule from the eight profiles they reviewed (Section 4, 'review eight profiles of bar exam candidates, presented in random order'). The paper reports that only 42.9% of Baseline participants spontaneously mentioned race/gender issues (Section 5, Qualitative Findings), and no participant accurately identified the rule. This suggests the discrimination was far from 'patent' to lay viewers. If the eight profiles contained few or no Black male candidates, then agreement with the model (5.9/8 times) may simply reflect the model's high base-rate accuracy against an uninformative profile set, not unwarranted trust. Moreover, the reported over-reliance metric ('89.3% of errors' vs '75.3%') depends on the number of model errors in the eight profiles; if that denominator is tiny (e.g., 1-2 trials), the effect is not robust. Without knowing the profile composition, the dose-response in agreement and the 'fairwashing' interpretation are ambiguous: the added components might be increasing perceived informativeness rather than 'junk' trust, and the baseline agreement cannot be labeled 'unjustified' relative to what participants could actually see. This is the load-bearing gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a between-subjects crowdsourced experiment (N=83 after exclusions) testing whether adding technically accurate but fairness-irrelevant explanatory components to an XAI dashboard increases lay users' agreement with, trust in, and satisfaction with a predictive model that is in fact highly discriminatory (predicting failure for Black male candidates and passing all others). Participants reviewed eight candidate profiles and rated their agreement with the model's predictions, then completed trust, satisfaction, fairness, and open-ended measures. The authors report significant effects of condition on agreement (F(2,80)=6.4, p=.0025), trust (F(2,80)=4.7, p=.012), and satisfaction (F(2,80)=7.3, p=.001), with the 'Everything' condition scoring highest. Qualitative data show fewer spontaneous fairness concerns in the higher-information conditions. The paper concludes that 'trust junk' causes unjustified support for a patently discriminatory model.","tokens_in":9754,"tokens_out":2301,"duration_ms":28896,"significance":"If the central claim holds, the result is practically important for XAI design: it would show that simply adding more accurate but decision-irrelevant information to explanations can increase unwarranted trust and reduce fairness scrutiny, even when the model's bias is extreme. The study is a direct, independent empirical test of Wall et al.'s 'trust junk' concept, and the authors provide an OSF supplement with stimuli and analyses. The paper does not fit parameters to outcomes and makes falsifiable predictions. The main threat to significance is whether the baseline condition truly allowed participants to recognize the model's discriminatory rule; if not, the increase with added components could reflect growing informativeness rather than 'junk' trust.","major_comments":[{"comment":"The central claim that trust junk causes 'unjustified' support for a 'patently discriminatory' model assumes participants could, in principle, detect the model's race/gender rule from the eight profiles they reviewed. The paper's own qualitative data show that only 42.9% of Baseline participants spontaneously mentioned race/gender issues and no participant accurately identified the rule. Without reporting the per-condition composition of the eight profiles (e.g., how many Black male candidates appeared, how many model errors occurred), baseline agreement (5.9/8) cannot be cleanly labeled 'unjustified' relative to what participants could see. If the eight profiles rarely exposed the rule, the dose-response may reflect the added components' perceived informativeness rather than trust junk. Please report the profile set composition, per-condition detection rates, and ideally a chance-level","section":"§4 and §5 (Model Agreement; Qualitative Findings)"},{"comment":"The over-reliance effect (M=89.3% vs 75.3% of errors) depends on the denominator of model errors in the eight profiles. If the model made few errors (e.g., one or two per condition), this large percentage difference rests on very few trials and may be unstable. The paper does not report the number of model errors per condition or per-trial error rates. Please provide the raw counts and, if needed, a mixed-effects or nonparametric analysis that accounts for the small denominator. This is load-bearing for the claim that trust junk increases harmful reliance, not just general agreement.","section":"§5 (Model Agreement and Task Performance)"},{"comment":"The fairness results are used to support the 'fairwashing' interpretation, but the paper reports inferential tests for only four Likert items and then switches to descriptive percentages for 'rated the unfair model as fair' without statistical comparison. The significant race-fairness effect (F(2,80)=3.5, p=.035) is in the expected direction, but the overall fairness results are mixed (gender, general bias, and ethics are not significant). The descriptive claim that more information increased the percentage rating the model as fair needs a formal test (e.g., logistic regression or chi-square with appropriate correction). This is important because the 'unjustified support' interpretation relies partly on perceived fairness.","section":"§5 (Fairness)"}],"minor_comments":[{"comment":"Typo in the reported p-value: 'p=0.0.012' should be 'p=.012'.","section":"§5 (Model Trust)"},{"comment":"The paper runs multiple ANOVAs (agreement, trust, satisfaction, four fairness items) without correcting for multiple comparisons. Given the small per-cell samples (27–28), the authors should report adjusted p-values or at least interpret the results cautiously. This does not undermine the main directional findings but affects their strength.","section":"§5 (statistical reporting)"},{"comment":"The sample size (N=83) is small for a between-subjects design with three conditions. The paper should report a power analysis or effect sizes with confidence intervals for the main comparisons.","section":"General"},{"comment":"The description says participants 'review eight profiles of bar exam candidates, presented in random order' but does not report the actual profiles' demographic composition. A supplemental table with the eight profiles (or their summary statistics) would substantially strengthen confidence in the design.","section":"§4 (Procedure)"},{"comment":"The qualitative coding is summarized very briefly. The paper should report inter-rater reliability for the coding if two or more coders were used, and clarify whether the reported percentages are based on the full sample or only on participants who provided text responses.","section":"§5 (Qualitative)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-motivated empirical test of a concept from prior work and does not suffer from circularity. The strongest concern is operational: whether the eight-profile task made the model's rule detectable enough for baseline trust to be called 'unjustified.' I would encourage the editor to request the supplemental profile composition and per-condition detection rates before considering acceptance. The multiple-comparison issue and small samples are secondary but should be addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — quick read of arXiv:2607.14152. The headline: this is the first empirical test of Wall et al.'s 'trust junk' construct, and the main effect is real. Adding accurate but task-irrelevant components to an explanation of a biased model moves agreement, trust, and satisfaction in a dose-response way. That is a meaningful contribution to the XAI/HCI literature, and the authors are honest about the messiness of the results.\n\nWhat it does well: the three-level between-subjects design is a sensible operationalization of the 'knob,' the stimulus is built from a real dataset, and the materials are open. The qualitative coding is a nice addition — baseline participants spontaneously reported more race/gender concerns and more missing information than those who saw the junk. The over-reliance result complements the agreement result. The self-citation to the 'novocaine chart' is used only as design inspiration, not as evidence, so that's not a problem.\n\nSoft spots: the load-bearing assumption is that participants could plausibly detect the race/gender rule from the eight profiles. The paper never reports the profile composition, and only 42.9% of baseline participants even mentioned race/gender spontaneously; none fully identified the rule. So 'unjustified' is partly in the eye of the researcher. If the eight profiles rarely included Black men, baseline agreement may just reflect an uninformative stimulus set. That said, the dose-response itself does not depend on detection — it shows junk increases trust regardless of what participants can infer — so the central claim survives. But the abstract's 'patently discriminatory' framing should be tempered.\n\nOther issues are minor but real: no correction for multiple ANOVAs, n=27–28 per cell, a p-value typo ('p=0.0.012'), and the abstract overstates the fairness findings when only one of four fairness items reached significance. The over-reliance metric also depends on a potentially small denominator of model errors; the paper should report that.\n\nWho this is for: XAI designers, visualization rhetoric researchers, and anyone studying fairwashing. A serious referee should ask for the stimulus composition, multiplicity-corrected results or effect sizes, and a more careful abstract. I'd send it out.","headline":"First direct test of 'trust junk' shows a real dose-response effect on agreement and trust, but the 'unjustified' label rests on an untested assumption about what eight profiles reveal.","tokens_in":10248,"tokens_out":1930,"would_cite":true,"duration_ms":22151,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Accurate but irrelevant explanation data can make people trust an AI they should reject.","keywords":["explainable AI","trust junk","data rhetoric","fairwashing","visualization","user study","algorithmic fairness","over-trust"],"falsifier":"A replication that adds a condition stating explicitly that the model uses only race and gender, while keeping the same trust-junk components, would settle whether the effect comes from the extra information or from obscuring the rule; if agreement stays high even when the rule is disclosed, the 'unjustified trust' interpretation weakens.","tokens_in":9307,"feed_emoji":"📊","tokens_out":3880,"duration_ms":42328,"temperature":0.7,"pith_summary":"The paper asks whether technically correct but useless information in an AI explanation can persuade ordinary people to trust a model they ought to reject. Using a deliberately discriminatory model that predicts bar exam failure for Black men and passes for everyone else, the authors ran a crowdsourced study with three explanation conditions: basic provenance, plus model accuracy, plus histograms and similar-candidate profiles. They found that each additional layer of accurate-but-irrelevant data increased agreement with the model, raised trust and satisfaction ratings, and made participants less likely to spontaneously report race or gender bias. Because the model's accuracy (93.7%) was actually lower than a trivial all-pass rule (94.8%), the observed trust is unwarranted. The paper concludes that XAI designers must treat explanations as persuasive rhetoric, not neutral information delivery.","feed_headline":"Adding irrelevant data boosts trust in a biased AI","feed_subtitle":"A crowdsourced test found extra stats and profiles made people rate a race-and-gender-only model as fair.","key_machinery":"The central object is 'trust junk': explanatory components that are technically correct but carry no information relevant to evaluating whether a model is fair or correct. In this study the components are model accuracy statistics, feature histograms, and short profiles of similar candidates drawn from the training data. The paper shows that adding these components in layers monotonically increases agreement, trust, and satisfaction with a patently discriminatory model while suppressing spontaneous fairness concerns. The mechanism is rhetorical: the appearance of completeness and authority persuades viewers, either by soothing them into over-reliance or by overwhelming them with detail.","core_discovery":"The paper's central claim is that increasing amounts of nominally explanatory data can engender unwarranted trust in a predictive model. In a between-subjects crowdsourced study of 83 participants, the authors presented the same unfair model—one that decides solely on race and gender—with three explanation dashboards of increasing informational richness. Participants who saw the fullest version agreed with the model 82.6% of the time versus 69.5% in the baseline, reported higher trust and explanation satisfaction, and were least likely to note race or gender problems in their open-ended responses (baseline participants did so 42.9% of the time; the others less often). The explanations themse","pith_inferences":["A direct testable extension would swap the junk components for equally large but fairness-relevant information, such as a feature-importance chart showing race and gender as the only inputs, and predict that the trust-boosting effect shrinks or reverses.","The paper does not measure whether participants accurately inferred the model's decision rule; future work could ask participants to state the rule explicitly, which would distinguish 'soothed into agreement' from 'genuinely uncertain because the sample of eight profiles rarely exposed the pattern.'","If trust junk affects lay viewers this strongly, the same mechanisms could undermine expert review settings, since the persuasion operates through the apparent completeness of the display rather than through its content; auditing procedures that rely on explanation dashboards may need to control for this effect."],"forward_implications":["XAI dashboards that present only favorable global metrics like accuracy can increase trust even when the model is worse than a trivial baseline.","Adding explanatory components can reduce, rather than increase, viewers' spontaneous detection of discrimination, by drawing attention away from the model's actual inputs.","Explanations that are factually correct can still be misleading, so designers bear responsibility for the persuasive effect of the entire presentation, not just the truth of each part.","Because even the explanation-sparse baseline elicited majority agreement with the model, the problem of over-trust in AI output may extend well beyond elaborate dashboards."],"fun_headline_variants":["Superfluous data boosts trust in a biased model","Accurate junk data wins trust for unfair AI","Irrelevant details make users overtrust discriminatory AI","Unearned trust from irrelevant XAI information"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The study assumes that after seeing only eight candidate profiles, participants could plausibly detect the model's race-and-gender rule, so that their continued agreement counts as unjustified trust rather than genuine uncertainty about an opaque system.","fun_headline_variants_meta":{"raw":{"variants":["Superfluous data boosts trust in a biased model","Accurate junk data wins trust for unfair AI","Irrelevant details make users overtrust discriminatory AI","Unearned trust from irrelevant XAI information"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000276,"raw_usage":{"total_tokens":1429,"prompt_tokens":635,"completion_tokens":794,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":379,"completion_tokens_details":{"reasoning_tokens":734}},"tokens_in":379,"tokens_out":794,"duration_ms":8497,"temperature":1.0,"reasoning_tokens":734,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:07:41.562158+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that adds a condition stating explicitly that the model uses only race and gender, while keeping the same trust-junk components, would settle whether the effect comes from the extra information or from obscuring the rule; if agreement stays high even when the rule is disclosed, the 'unjustified trust' interpretation weakens.","supporting_citations":[],"review_version":1}