REVIEW 3 major objections 5 minor 43 references
"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models
T0 review · 3 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read Accurate but irrelevant explanation data can make people trust an AI they should reject.
desk verdict First direct test of 'trust junk' shows a real dose-response effect on agreement and trust, but the 'unjustified' label rests on an untested assumption about what eight profiles reveal. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is 'trust junk': explanatory components that are technically correct but carry no information relevant to evaluating whether a model is fair or correct. In this study the components are model accuracy statistics, feature histograms, and short profiles of similar candidates drawn from the training data. The paper shows that adding these components in layers monotonically increases agreement, trust, and satisfaction with a patently discriminatory model while suppressing spontaneous fairness concerns. The mechanism is rhetorical: the appearance of completeness and authority persuades viewers, either by soothing them into over-reliance or by overwhelming them with detail.
What would settle it
A replication that adds a condition stating explicitly that the model uses only race and gender, while keeping the same trust-junk components, would settle whether the effect comes from the extra information or from obscuring the rule; if agreement stays high even when the rule is disclosed, the 'unjustified trust' interpretation weakens.
Extended reading notes
Core claim
The paper's central claim is that increasing amounts of nominally explanatory data can engender unwarranted trust in a predictive model. In a between-subjects crowdsourced study of 83 participants, the authors presented the same unfair model—one that decides solely on race and gender—with three explanation dashboards of increasing informational richness. Participants who saw the fullest version agreed with the model 82.6% of the time versus 69.5% in the baseline, reported higher trust and explanation satisfaction, and were least likely to note race or gender problems in their open-ended responses (baseline participants did so 42.9% of the time; the others less often). The explanations themse
Load-bearing premise
The study assumes that after seeing only eight candidate profiles, participants could plausibly detect the model's race-and-gender rule, so that their continued agreement counts as unjustified trust rather than genuine uncertainty about an opaque system.
Editorial extensions
If this is right
- XAI dashboards that present only favorable global metrics like accuracy can increase trust even when the model is worse than a trivial baseline.
- Adding explanatory components can reduce, rather than increase, viewers' spontaneous detection of discrimination, by drawing attention away from the model's actual inputs.
- Explanations that are factually correct can still be misleading, so designers bear responsibility for the persuasive effect of the entire presentation, not just the truth of each part.
- Because even the explanation-sparse baseline elicited majority agreement with the model, the problem of over-trust in AI output may extend well beyond elaborate dashboards.
Reading between the lines
- A direct testable extension would swap the junk components for equally large but fairness-relevant information, such as a feature-importance chart showing race and gender as the only inputs, and predict that the trust-boosting effect shrinks or reverses.
- The paper does not measure whether participants accurately inferred the model's decision rule; future work could ask participants to state the rule explicitly, which would distinguish 'soothed into agreement' from 'genuinely uncertain because the sample of eight profiles rarely exposed the pattern.'
- If trust junk affects lay viewers this strongly, the same mechanisms could undermine expert review settings, since the persuasion operates through the apparent completeness of the display rather than through its content; auditing procedures that rely on explanation dashboards may need to control for this effect.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a between-subjects crowdsourced experiment (N=83 after exclusions) testing whether adding technically accurate but fairness-irrelevant explanatory components to an XAI dashboard increases lay users' agreement with, trust in, and satisfaction with a predictive model that is in fact highly discriminatory (predicting failure for Black male candidates and passing all others). Participants reviewed eight candidate profiles and rated their agreement with the model's predictions, then completed trust, satisfaction, fairness, and open-ended measures. The authors report significant effects of condition on agreement (F(2,80)=6.4, p=.0025), trust (F(2,80)=4.7, p=.012), and satisfaction (F(2,80)=7.3, p=.001), with the 'Everything' condition scoring highest. Qualitative data show fewer spontaneous fairness concerns in the higher-information conditions. The paper concludes that 'trust junk' causes unjustified support for a patently discriminatory model.
Significance. If the central claim holds, the result is practically important for XAI design: it would show that simply adding more accurate but decision-irrelevant information to explanations can increase unwarranted trust and reduce fairness scrutiny, even when the model's bias is extreme. The study is a direct, independent empirical test of Wall et al.'s 'trust junk' concept, and the authors provide an OSF supplement with stimuli and analyses. The paper does not fit parameters to outcomes and makes falsifiable predictions. The main threat to significance is whether the baseline condition truly allowed participants to recognize the model's discriminatory rule; if not, the increase with added components could reflect growing informativeness rather than 'junk' trust.
major comments (3)
- [§4 and §5 (Model Agreement; Qualitative Findings)] The central claim that trust junk causes 'unjustified' support for a 'patently discriminatory' model assumes participants could, in principle, detect the model's race/gender rule from the eight profiles they reviewed. The paper's own qualitative data show that only 42.9% of Baseline participants spontaneously mentioned race/gender issues and no participant accurately identified the rule. Without reporting the per-condition composition of the eight profiles (e.g., how many Black male candidates appeared, how many model errors occurred), baseline agreement (5.9/8) cannot be cleanly labeled 'unjustified' relative to what participants could see. If the eight profiles rarely exposed the rule, the dose-response may reflect the added components' perceived informativeness rather than trust junk. Please report the profile set composition, per-condition detection rates, and ideally a chance-level
- [§5 (Model Agreement and Task Performance)] The over-reliance effect (M=89.3% vs 75.3% of errors) depends on the denominator of model errors in the eight profiles. If the model made few errors (e.g., one or two per condition), this large percentage difference rests on very few trials and may be unstable. The paper does not report the number of model errors per condition or per-trial error rates. Please provide the raw counts and, if needed, a mixed-effects or nonparametric analysis that accounts for the small denominator. This is load-bearing for the claim that trust junk increases harmful reliance, not just general agreement.
- [§5 (Fairness)] The fairness results are used to support the 'fairwashing' interpretation, but the paper reports inferential tests for only four Likert items and then switches to descriptive percentages for 'rated the unfair model as fair' without statistical comparison. The significant race-fairness effect (F(2,80)=3.5, p=.035) is in the expected direction, but the overall fairness results are mixed (gender, general bias, and ethics are not significant). The descriptive claim that more information increased the percentage rating the model as fair needs a formal test (e.g., logistic regression or chi-square with appropriate correction). This is important because the 'unjustified support' interpretation relies partly on perceived fairness.
minor comments (5)
- [§5 (Model Trust)] Typo in the reported p-value: 'p=0.0.012' should be 'p=.012'.
- [§5 (statistical reporting)] The paper runs multiple ANOVAs (agreement, trust, satisfaction, four fairness items) without correcting for multiple comparisons. Given the small per-cell samples (27–28), the authors should report adjusted p-values or at least interpret the results cautiously. This does not undermine the main directional findings but affects their strength.
- [General] The sample size (N=83) is small for a between-subjects design with three conditions. The paper should report a power analysis or effect sizes with confidence intervals for the main comparisons.
- [§4 (Procedure)] The description says participants 'review eight profiles of bar exam candidates, presented in random order' but does not report the actual profiles' demographic composition. A supplemental table with the eight profiles (or their summary statistics) would substantially strengthen confidence in the design.
- [§5 (Qualitative)] The qualitative coding is summarized very briefly. The paper should report inter-rater reliability for the coding if two or more coders were used, and clarify whether the reported percentages are based on the full sample or only on participants who provided text responses.
Circularity Check
No significant circularity: the study is an independent empirical test of a pre-existing construct, with self-citations used only for stimulus design, not as load-bearing evidence.
full rationale
The paper's central claim—that increasing amounts of nominally explanatory data can engender unwarranted trust in a predictive model—is tested through a between-subjects crowdsourced experiment, not derived from its own definitions or fitted parameters. The independent variable is the experimentally manipulated presence of explanatory components (Baseline, Model, Everything), and the dependent variables (agreement, trust, satisfaction, fairness perceptions) are measured after exposure. Nothing in the results section reduces a prediction to an input by construction: agreement with the model is not defined by the number of explanatory components, nor is any parameter fitted to the outcome. The 'trust junk' concept is attributed to Wall et al. [38], an external prior work, and the authors' own prior work (Correll [6], 'novocaine chart') is cited only as design inspiration for a stimulus technique, not as evidence for the empirical claim. There is no self-citation chain that forces the conclusion, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in via citation: the paper explicitly states that Wall et al.'s techniques were 'described, but not tested' and that the authors created their own 'junk' explanations based on those descriptions. The possible weakness that participants may not have detected the model's race/gender rule from eight profiles is a construct-validity concern, not a circularity concern, because the study's outcome measures are not logically entailed by the stimulus design. The paper is self-contained against external benchmarks in the sense that its results are empirical observations; the central claim is not equivalent to its inputs by definition. Therefore, no circular step is identified, and the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption ANOVA assumes approximately normal residuals and homogeneous variances for Likert-scale responses.
- domain assumption Self-reported trust, satisfaction, and fairness Likert ratings measure the corresponding psychological constructs.
- domain assumption The added explanatory components (accuracy, histograms, similar profiles) are accurate but superfluous for judging fairness.
- domain assumption Prolific crowdsourced participants are an acceptable stand-in for lay users of XAI explanations.
Cite this review
Pith. "Pith review of "Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models." pith.science (2026). https://pith.science/paper/OQTYUJZC
@misc{pith2026260714152,
author = {Pith},
title = {Pith review of: "Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/OQTYUJZC}},
note = {Machine review of arXiv:2607.14152}
}
read the original abstract
The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-trust of predictive models. In this paper, we use a crowdsourced study to show that providing accurate (but superfluous or irrelevant) data in a model explanation can, in fact, result in unjustified trust and other positive beliefs about a model, even when the model is patently discriminatory and unfair. Our results suggest that XAI designers and developers need to consider the implicit or explicit rhetorics of their work, and beware of the potential of visualizations to imbue models with unearned trust.
Figures
Reference graph
Works this paper leans on
-
[1]
A. Adadi and M. Berrada. Peeking inside the black-box: a survey on explainable artificial intelligence (XAI).IEEE Access, 6:52138– 52160, 2018. doi: 10.1109/ACCESS.2018.2870052 2
arXiv 2018
-
[2]
N. Al-Ansari, D. Al-Thani, and R. S. Al-Mansoori. User-centered evaluation of explainable artificial intelligence (xai): A systematic literature review.Human Behavior and Emerging Technologies, 2024(1):4628855, 1 2024. doi: 10.1155/2024/4628855 2
-
[3]
A ¨ıvodji, H
U. A ¨ıvodji, H. Arai, O. Fortineau, S. Gambs, S. Hara, and A. Tapp. Fairwashing: the risk of rationalization. InICML, pp. 161–170. PMLR, 2019. 1, 2, 3
2019
-
[4]
A. Bussone, S. Stumpf, and D. O’Sullivan. The role of explanations on trust and reliance in clinical decision support systems. InInternational Conference on Healthcare Informatics, pp. 160–169. IEEE, 2015. doi: 10.1109/ichi.2015.26 2
-
[5]
F. Cabitza, C. Fregosi, A. Campagner, and C. Natali. Explanations considered harmful: the impact of misleading explanations on accu- racy in hybrid human-ai decision making. InWorld conference on explainable artificial intelligence, pp. 255–269. Springer, 2024. doi: 10.1007/978-3-031-63803-9 14 2
-
[6]
M. Correll. Towards a Theory of Bullshit Visualization, 9 2021. arXiv:2109.12975 [cs]. 2, 4
arXiv 2021
-
[7]
Dimanov, U
B. Dimanov, U. Bhatt, M. Jamnik, and A. Weller. You shouldn’t trust me: Learning models which conceal unfairness from multiple expla- nation methods. InECAI, pp. 2473–2480. IOS Press, 2020. 2
2020
-
[8]
J. Drucker. Humanistic theory and digital scholarship.Debates in the digital humanities, 150:85–95, 2012. doi: 10.5749/minnesota/ 9780816677948.003.0011 1
Show all 43 references
-
[9]
Ehsan and M
U. Ehsan and M. O. Riedl. Explainability pitfalls: Beyond dark pat- terns in explainable AI.Patterns, 5(6), 2024. doi: 10.1016/j.patter. 2024.100971 2
2024
-
[10]
Ehsan, K
U. Ehsan, K. Saha, M. De Choudhury, and M. O. Riedl. Charting the sociotechnical gap in explainable AI: A framework to address the gap in XAI.CSCW, 7:1–32, 2023. doi: 10.1145/3579467 4
2023 doi
-
[11]
Ehsan, R
U. Ehsan, R. Singh, J. Metcalf, and M. Riedl. The algorithmic im- print. InACM FAccT, p. 1305–1317. ACM, June 2022. doi: 10.1145/ 3531146.3533186 2
2022
-
[12]
Eiband, D
M. Eiband, D. Buschek, A. Kremer, and H. Hussmann. The impact of placebic explanations on trust in intelligent systems. InCHI Extended Abstracts, pp. 1–6. ACM, 2019. doi: 10.1145/3290607.3312787 2
2019
-
[13]
Gomez, S
O. Gomez, S. Holter, J. Yuan, and E. Bertini. Advice: Aggregated visual counterfactual explanations for machine learning model vali- dation. InIEEE VIS, pp. 31–35, 2021. doi: 10.1109/vis49827.2021. 9623271 2
2021
-
[14]
Goyal, C
N. Goyal, C. Baumler, T. Nguyen, and H. Daum ´e Iii. The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features. InIUI, pp. 155–180. ACM, 3 2024. doi: 10.1145/ 3640543.3645210 3
2024
-
[15]
S. U. Hamida, M. J. M. Chowdhury, N. R. Chakraborty, K. Biswas, and S. K. Sami. Exploring the landscape of explainable artificial intelligence (xai): A systematic review of techniques and applica- tions.Big Data and Cognitive Computing, 8(149), 2024. doi: 10. 3390/bdcc8110149 2
2024
-
[16]
R. R. Hoffman, S. T. Mueller, G. Klein, and J. Litman. Met- rics for explainable AI: Challenges and prospects.arXiv preprint arXiv:1812.04608, 2018. 2, 3, 4
2018 arXiv
-
[17]
Hullman, Z
J. Hullman, Z. Guo, and B. Ustun. Explanations are a means to an end.arXiv preprint arXiv:2506.22740, 2025. 4
2025
-
[18]
Jacobs, M
M. Jacobs, M. F. Pradier, T. H. McCoy Jr, R. H. Perlis, F. Doshi-Velez, and K. Z. Gajos. How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection. Translational psychiatry, 11(1):108, 2021. doi: 10.1038/s41398-021 -...
2021 doi
-
[19]
Kalasampath, K
K. Kalasampath, K. N. Spoorthi, S. Sajeev, S. S. Kuppa, K. Ajay, and A. Maruthamuthu. A literature review on applications of explainable artificial intelligence (xai).IEEE Access, 13:41111–41140, 2025. doi: 10.1109/ACCESS.2025.3546681 2
2025
-
[20]
H. Kaur, H. Nori, S. Jenkins, R. Caruana, H. Wallach, and J. Wort- man Vaughan. Interpreting interpretability: understanding data scien- tists’ use of interpretability tools for machine learning. InCHI, pp. 1–14. ACM, 2020. doi: 10.1145/3313831.3376219 4
2020
-
[21]
Kennedy, R
H. Kennedy, R. L. Hill, G. Aiello, and W. Allen. The work that vi- sualisation conventions do.Information, Communication & Society, 19(6):715–735, 2016. doi: 10.1080/1369118x.2016.1153126 1
2016
-
[22]
Kulesza, S
T. Kulesza, S. Stumpf, M. Burnett, S. Yang, I. Kwan, and W.-K. Wong. Too much, too little, or just right? ways explanations impact end users’ mental models. InVL/HCC, pp. 3–10. IEEE, 9 2013. doi: 10. 1109/VLHCC.2013.6645235 2
2013
-
[23]
Laato, M
S. Laato, M. Tiainen, A. Najmul Islam, and M. M ¨antym¨aki. How to explain AI systems to end users: a systematic literature review and research agenda.Internet Research, 32(7):1–31, 2022. doi: 10.1108/ intr-08-2021-0600 2
2022
-
[24]
how do I fool you?
H. Lakkaraju and O. Bastani. “how do I fool you?” manipulating user trust via misleading black box explanations. InAIES, pp. 79–85. AAAI/ACM, 2020. doi: 10.1145/3375627.337583 2
2020
-
[25]
Look! It’s a Computer Program! It’s an Algorithm! It’s AI!
M. Langer, T. Hunsicker, T. Feldkamp, C. J. K ¨onig, and N. Grgi ´c- Hlaˇca. “Look! It’s a Computer Program! It’s an Algorithm! It’s AI!”: Does Terminology Affect Human Perceptions and Evaluations of Algorithmic Decision-Making Systems? InCHI, pp. 1–28. ACM, 4 2022. doi: 10.11...
2022
-
[26]
Q. V . Liao, D. Gruen, and S. Miller. Questioning the AI: informing design practices for explainable AI user experiences. InCHI, pp. 1–
-
[27]
doi: 10.1145/3313831.3376590 2
ACM, 2020. doi: 10.1145/3313831.3376590 2
2020
-
[28]
Mohseni, N
S. Mohseni, N. Zarei, and E. D. Ragan. A multidisciplinary survey and framework for design and evaluation of explainable AI systems. TiiS, 11(3-4):1–45, 2021. doi: 10.1145/3387166 2
2021 doi
-
[29]
Nourani, S
M. Nourani, S. Kabir, S. Mohseni, and E. D. Ragan. The effects of meaningful and meaningless explanations on trust and perceived sys- tem accuracy in intelligent systems. InHCOMP, vol. 7, pp. 97–105. AAAI, 2019. doi: 10.1609/hcomp.v7i1.5284 2
2019 doi
-
[31]
Nourani, C
M. Nourani, C. Roy, T. Rahman, E. D. Ragan, N. Ruozzi, and V . Gogate. Don’t explain without verifying veracity: an evaluation of explainable AI with video activity recognition.arXiv preprint arXiv:2005.02335, 2020. 3
2005 arXiv
-
[32]
O’Neil.Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy
C. O’Neil.Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy. Penguin Books, London, 2016. 2
2016
-
[33]
E. M. Peck, S. E. Ayuso, and O. El-Etr. Data is personal: Attitudes and perceptions of data visualization in rural pennsylvania. InCHI, pp. 1–12. ACM, 2019. doi: 10.1145/3411764.3445315 1
2019
-
[34]
Poursabzi-Sangdeh, D
F. Poursabzi-Sangdeh, D. G. Goldstein, J. M. Hofman, J. W. Wort- man Vaughan, and H. Wallach. Manipulating and measuring model interpretability. InCHI, pp. 1–52. ACM, 2021. doi: 10.1145/3411764 .344531 2, 3
2021 doi
-
[35]
Pruthi, M
D. Pruthi, M. Gupta, B. Dhingra, G. Neubig, and Z. C. Lipton. Learn- ing to deceive with attention-based explanations. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, eds.,ACL, pp. 4782–4793, July
-
[36]
Slack, S
D. Slack, S. Hilgard, E. Jia, S. Singh, and H. Lakkaraju. Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods. InAIES, pp. 180–186. AAAI/ACM, 2 2020. doi: 10.1145/3375627. 3375830 2
2020 doi
-
[37]
Schwalbe and B
G. Schwalbe and B. Finzel. A comprehensive taxonomy for ex- plainable artificial intelligence: a systematic survey of surveys on methods and concepts.Data Mining and Knowledge Discovery, 38(5):3043–3101, Sept 2024. doi: 10.1007/s10618-022-00867-8 2
2024 doi
-
[38]
E. Wall, L. Matzen, M. El-Assady, P. Masters, H. Hosseinpour, A. En- dert, R. Borgo, P. Chau, A. Perer, H. Schupp, H. Strobelt, and L. Padilla. Trust junk and evil knobs: Calibrating trust in AI vi- sualization. InPacificVis, pp. 22–31. IEEE, 2024. doi: 10.1109/ pacificvis6037...
2024
-
[39]
Szymanski, M
M. Szymanski, M. Millecamp, and K. Verbert. Visual, textual or hy- brid: the effect of user expertise on different explanations. InIUI, pp. 109–119. ACM, 4 2021. doi: 10.1145/3397481.3450662 2
2021
-
[40]
L. F. Wightman. LSAC National Longitudinal Bar Passage Study. LSAC Research Report Series.Law School Admission Council, 1998. 2
1998
-
[41]
A. Weller. Challenges for transparency.arXiv preprint arXiv:1708.01870, 2017. 1, 2, 4
2017 arXiv
-
[42]
Ytre-Arne and H
B. Ytre-Arne and H. Moe. Folk theories of algorithms: Understanding digital irritation.Media, Culture & Society, 43(5):807–824, 2021. doi: 10.1177/0163443720972314 2
2021 doi
-
[43]
Xiong, J
C. Xiong, J. Shapiro, J. Hullman, and S. Franconeri. Illusion of causal- ity in visualized data.IEEE TVCG, 26(1):853–862, 2019. doi: 10. 1109/tvcg.2019.2934399 2
2019
-
[2020]
doi: 10.18653/v1/2020.acl-main.432 2
2020 doi
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.