Pith. sign in

REVIEW 4 major objections 5 minor 76 references

An Empirical Examination of the Evaluative AI Framework

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read In a 375-participant experiment, replacing AI recommendations with optional pro and con evidence did not improve decision quality or engagement, and users offloaded cognition much as they do with traditional AI.

desk verdict A pre-registered, powered null result for Miller's Evaluative AI framework that is honest about its own limitations; the main caveat is whether the SHAP-based evidence display actually instantiates the framework. read the letter →

arxiv 2411.08583 v1 pith:UBWAZHO7 submitted 2024-11-13 cs.HC cs.AI

classification cs.HCcs.AI
keywords EvaluativeAIexplainablehuman-AIdecision-makingproandconevidenceBrierscorecognitiveoffloadingbehavioralexperimentSHAP
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to test the Evaluative AI framework's core promise: that presenting decision-makers with pro and con evidence for each option, rather than an AI recommendation, will lead to more informed decisions and less blind reliance. The experiment used an incentivized income-estimation task in which 375 lay participants judged whether four people earned above-median net income, across five conditions: no assistance, recommendation only, evidence only, recommendation plus evidence, and the framework's optional-evidence design. The result was a null effect: the Evaluative AI condition's Brier score did not differ significantly from any other condition and was numerically the worst of the five. Evidence-displaying conditions were significantly slower, cognitive load did not rise, and qualitative reports indicated that AI-assisted participants relied on fewer features than unsupported participants, which the paper reads as cognitive offloading and potential automation bias rather than deeper engagement. If this result holds, it undercuts the empirical case that a recommendation-free, hypothesis-driven interface alone can improve human-AI decision making, and it shifts attention to the quality and usability of the evidence itself.

What carries the argument

The central machinery is a five-condition randomized experiment built on the contrast between recommendation-driven and hypothesis-driven support. In the Evaluative AI condition, participants see only two buttons, one for pro evidence and one for con evidence, and can open them at any time; no recommendation is ever offered. The evidence itself is generated from SHAP feature attributions of a logistic regression model (each personal characteristic contributes a positive or negative push to the predicted probability), displayed as bar charts and converted into short textual arguments. Performance is scored with the Brier score, $\frac{1}{N}\sum_{i=1}^{N}(p_i-o_i)^2$, which also determines the bonus payment; decision time, NASA-TLX cognitive load, and qualitative coding of free-text decision descriptions complete the measurement of how users actually reasoned.

What would settle it

Re-run the same five-condition experiment with an evidence-comprehension check that participants must pass before deciding, and with the Evaluative AI group's bonus tied to using the evidence; if that group's Brier score drops significantly below the control group's, the claim that this framework does not improve decisions would be overturned in that setting.

Watch

Extended reading notes

Core claim

The central discovery is a null result, stated on the paper's own terms: in an incentivized between-subjects experiment with 375 lay participants estimating whether each of four individuals earns above-median net income, the treatment that implemented the Evaluative AI framework—pro and con evidence available on request, with no recommendation—did not improve decision-making performance. The mean Brier score $\frac{1}{N}\sum_{i=1}^{N}(p_i-o_i)^2$ for the Evaluative AI group was $0.230$, the worst among the five groups, and a Kruskal-Wallis test found no significant between-group difference ($p=0.154$). The paper also reports limited engagement: 62.64% of Evaluative AI participants clicked both evidence buttons in every round, and qualitative descriptions of the decision process showed that AI-supported participants relied on fewer features than unsupported controls, which the author interprets as cognitive offloading and potential automation bias rather than deeper reasoning with the evidence. The conclusion is that this implementation of the framework did not deliver the performance or engagement gains it was designed to produce.

Load-bearing premise

The load-bearing premise is that the pro and con evidence participants could click on was clear enough and faithful enough to what the framework intends that the experiment actually tested the framework's idea rather than a flawed presentation of it.

Editorial extensions

If this is right

  • No AI support condition significantly beat doing the task alone on Brier score; the Evaluative AI condition's mean score of $0.230$ was numerically the worst, so the framework's hoped-for performance gain did not materialize.
  • Evidence on demand slowed decisions: Evaluative AI (mean 51.7s), Evidence Only (56.4s), and Recommendation and Evidence (57.2s) all took significantly longer than Control or Recommendation Only (about 41s).
  • Cognitive load, as measured by NASA-TLX, did not differ between any treatments, contradicting the hypothesis that hypothesis-driven support would demand more mental effort.
  • Qualitative reports indicate all AI-assisted groups mentioned individual features less often than the no-AI control, which the paper reads as cognitive offloading and potential automation bias rather than deeper engagement.
  • Only 62.64% of participants in the Evaluative AI condition used the optional evidence in every round, so superficial engagement persisted even when evidence was the only AI feature.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The null result may be an implementation ceiling: if the textual pro/con evidence was hard to parse or did not match participants' mental models, the framework itself remains untested. A comprehension check on the evidence would separate these possibilities.
  • The incentive design may matter: a £3 show-up fee with a bonus up to £6 per four tasks may not create the high-stakes, low-frequency setting the framework targets. A field-like task with real professional consequences could produce different engagement.
  • This result fits a cost-benefit view of overreliance: when optional evidence costs extra clicks and reading time, users may rationally skip it. A design that makes evidence an unavoidable first step, or that gives users a personal stake in accuracy, could recover the intended effect.
  • The binary income task may be too simple for pros-and-cons reasoning to add value; the framework's natural habitat is multi-hypothesis diagnosis where several options compete, so a multi-class extension is a direct next test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports a preregistered, incentivized online experiment (N = 375) testing Miller's Evaluative AI framework. Participants estimated the probability that each of four individuals had a net income above the median, with five between-subjects conditions: no AI, AI recommendation only, evidence only, recommendation plus evidence, and the Evaluative AI condition in which pro/con evidence was available on demand. The central findings are that Brier-score performance did not differ significantly across conditions, with the Evaluative AI condition numerically worst; evidence-presenting conditions were slower than control or recommendation-only; NASA-TLX cognitive load did not differ; and qualitative coding suggested that AI-assisted participants mentioned fewer features. The authors conclude that the Evaluative AI framework, at least as implemented here, did not improve decision-making performance and that engagement with evidence was limited.

Significance. If the null result is taken at face value, it is a useful empirical counterpoint to Le et al. (2024a) and to Miller's theoretical claims, and it provides a cautionary data point for hypothesis-driven XAI. The study has real methodological strengths: it is preregistered, uses a power analysis with Bonferroni correction, applies non-parametric tests with a control group, uses monetary incentives tied to Brier score, and makes data and code available. The main weakness is that the validity of the evidence operationalization is not established, which makes the scope of the conclusion uncertain. The study is therefore worth publishing only after the authors either validate the evidence display or explicitly restrict their claims to the specific implementation tested.

major comments (4)
  1. [Section 3.7] Section 3.7 and Section 3.2: The pro/con evidence is generated from SHAP values and converted to free text via GPT-4o, but the manuscript provides no validation that these statements accurately reflect the SHAP contributions, no comprehensibility or manipulation check, and no evidence that participants could map the directional statements to the probability scale they had to report. The pretest mentioned in Section 1 found bar charts poorly understood, yet no check confirms that the added textual descriptions fixed the problem. Because the central null result in the Evidence Only and Evaluative AI conditions depends on participants actually receiving and understanding Miller-style evaluative evidence, the study currently cannot distinguish 'the framework does not improve decisions' from 'this operationalization of evidence was not usable.' I recommend either a comprehension check on the evidence display, an audit of the GPT-4o-generated statements, or a condition in which the evidence is coupled with the AI's baseline probability.
  2. [Section 3.3] Section 3.3 (Brier score) and Section 3.7: The outcome is a calibrated probability, but the evaluative evidence is presented as feature-level SHAP contributions without any stated mapping from those log-odds-style contributions to the percentage estimates participants must produce. Evidence Only and Evaluative AI participants never see the AI's probability estimate or a base-rate anchor, whereas Recommendation Only participants do. This confounds the presence of a recommendation with the availability of a response-scale cue; the observed advantage of Recommendation Only and the poor numerical performance of Evaluative AI may reflect the absence of a probability anchor rather than the absence of a recommendation. The paper should either provide the model's predicted probability or an explicit aggregation instruction in the evidence conditions, or restrict the conclusion to the specific evidence presentation tested.
  3. [Section 3.4] Section 3.4 (Decision-Making Process) and Section 4.4: The qualitative coding is described as performed by the author with the assistance of GPT-4o, but no inter-rater reliability, double-coding subsample, or validation of the LLM-assisted coding is reported. The claims about AI mentions, feature mentions, and the resulting cognitive-offloading interpretation rest on these categories. A second human coder on a random subsample with a kappa statistic, or a documented validation of the GPT-4o coding protocol, is necessary to support those conclusions.
  4. [Abstract and Section 4.5] The abstract characterizes engagement as 'limited,' but Section 4.5 reports that 62.64% of Evaluative AI participants clicked on both the pro and con evidence in all four rounds (i.e., all eight evidence items). If the intended claim is that participants did not deeply process the evidence, the click data alone do not support that; the paper should state the specific measure on which 'limited engagement' is based and reconcile the high click-through rate with the superficial-engagement interpretation.
minor comments (5)
  1. [Section 3.5] There is a typo in the subsection heading ('Desicion-Making Process'); similarly, 'Apendix' appears in Section 3.2.
  2. [Section 4.1] Reporting effect sizes and confidence intervals for the pairwise comparisons (or at least the test statistics) would improve interpretability, since Kruskal-Wallis p-values alone do not convey the magnitude of the null result.
  3. [Appendix A] Appendix A states the AI is correct 77% of the time, whereas Section 3.7 reports 15/20 (75%) accuracy with a 50% threshold; these numbers should be reconciled.
  4. [Section 3.8] The final group sizes are noticeably uneven (62 to 91 participants); the manuscript should state how randomization was implemented and whether the imbalance was checked.
  5. [Section 4.4] The pairwise chi-squared tests are reported only through p-values; the test statistics or exact p-values with the correction method should be given.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper reports an empirical null result; its claims are observed outcomes, not derivations forced by definitions or self-citation.

full rationale

This is an empirical, pre-registered behavioral experiment, not a derivation chain. The central claim—that the Evaluative AI implementation did not improve decision performance and saw limited evidence engagement—is a statistical finding from independent dependent variables (Brier score, decision time, NASA-TLX, click behavior). The hypotheses H1–H3 follow from Miller's framework, but the outcome measures are not defined in terms of the framework's assumptions, so the hypotheses could have been confirmed or refuted by the data; indeed they were refuted. The pro/con evidence is an operationalization of the framework via SHAP signs and GPT-4o text, and the paper's Limitations explicitly concede that this presentation may need improvement and that the model did not greatly outperform participants. Those are construct-validity and generalizability concerns, not circularity: the study does not define 'Evaluative AI improved performance' to mean 'the SHAP/GPT-4o display was engaged with,' nor does it fit any parameter and then relabel that fit as a prediction. There is no self-citation chain that carries the argument: the only self-citation (Kornowicz and Thommes, 2024) appears in Limitations as a supporting reference for domain dependence and is not load-bearing. No uniqueness theorem, ansatz, or renaming of a known result is used to make the conclusion true by construction. The null result is an observed difference among experimental conditions, not an identity or a reduction to the study's inputs. Therefore no circular step is identifiable, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central null result does not depend on any fitted parameters or invented entities. It rests on domain assumptions about the operationalization of the framework and the validity of the task and the qualitative measures. These assumptions are acknowledged as limitations in Section 6 but are load-bearing for interpreting the result as a test of Evaluative AI.

assumptions (3)
  • domain assumption The SHAP-based pro/con evidence, converted to text by GPT-4o, is a faithful and comprehensible operationalization of the evidence that Miller's Evaluative AI framework envisions.
    The null result is interpreted as evidence about the framework, but the experiment only tests this one implementation; if the evidence was unclear or mistranslated, the test is weak. This is acknowledged as a limitation in Section 6 but is load-bearing.
  • domain assumption The income-estimation task with 20 features and lay participants is an appropriate testbed for the framework, which targets medium/high-stakes decisions.
    The paper itself notes that laypeople and binary decisions may limit external validity (Section 6).
  • domain assumption Participants' self-reports of their decision processes (coded by the author with GPT-4o) can be treated as valid evidence about cognitive engagement.
    Qualitative content analysis without inter-rater reliability (Section 4.4) is used to support the engagement conclusion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Empirical Examination of the Evaluative AI Framework." pith.science (2026). https://pith.science/paper/UBWAZHO7

@misc{pith2026241108583,
  author       = {Pith},
  title        = {Pith review of: An Empirical Examination of the Evaluative AI Framework},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UBWAZHO7}},
  note         = {Machine review of arXiv:2411.08583}
}
read the original abstract

This study empirically examines the "Evaluative AI" framework, which aims to enhance the decision-making process for AI users by transitioning from a recommendation-based approach to a hypothesis-driven one. Rather than offering direct recommendations, this framework presents users pro and con evidence for hypotheses to support more informed decisions. However, findings from the current behavioral experiment reveal no significant improvement in decision-making performance and limited user engagement with the evidence provided, resulting in cognitive processes similar to those observed in traditional AI systems. Despite these results, the framework still holds promise for further exploration in future research.

Figures

Figures reproduced from arXiv: 2411.08583 by the authors.

Figure 1
Figure 1. Mean Absolute SHAP Values for Features in Model Interpretation. The horizon [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Comparison of Brier Scores Across Different Treatments. The bar chart presents [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Average Task Completion Time Across Different Treatments. This bar chart shows [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: Cognitive Load (NASA-TLX) Across Different Treatments. The bar chart illus [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Frequency of AI Mentions by Treatment Group. This bar chart displays the per [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Proportion of Participants Mentioning Features by Treatment Group. The bar [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Average Number of Features Mentioned Per Participant by Treatment Group with [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: Frequency of Features Mentioned in the Decision-Making Process. The horizon [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Frequency of Features Mentioned in the Decision-Making Process by Treatment. [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]
Figure 10
Figure 10. Figure 10: Distribution of Evidence Clicks in Evaluative AI. The bar chart illustrates the proportion of participants who clicked on varying numbers of evidence items (ranging from 0 to 8) during Evaluative AI. in completing the tasks. Unlike many studies that demonstrate AI rec…
Figure 11
Figure 11. Figure 11: Time to Click (s) by Task Number and Evidence Type with 95% Confidence [PITH_FULL_IMAGE:figures/full_fig_p021_11.png]
Figure 12
Figure 12. Figure 12: Translated interface in Control: The features with their values are listed at the top, followed by the input field for the participant below. iii [PITH_FULL_IMAGE:figures/full_fig_p021_12.png]
Figure 13
Figure 13. Figure 13: Translated interface in Recommendation Only: The features with their values are listed at the top, followed by the input field for the participant, and below that, the AI recommendation. iv [PITH_FULL_IMAGE:figures/full_fig_p022_13.png]
Figure 14
Figure 14. Figure 14: Translated interface in Evidence Only: Due to space constraints in the screenshot, the features and input field were not included, only the presentation of the pro and con evidence. v [PITH_FULL_IMAGE:figures/full_fig_p023_14.png]
Figure 15
Figure 15. Figure 15: Translated interface in Recommendation and Evidence: Due to space constraints in the screenshot, the features and input field were not included, only the recommendation and the presen￾tation of the pro and con evidence. vi [PITH_FULL_IMAGE:figures/full_fig_p024_15.png]
Figure 16
Figure 16. Figure 16: Translated interface in Evaluative AI: Due to space constraints in the screenshot, the features and input field were not included, only the button for the con evidence and the already displayed pro evidence. vii [PITH_FULL_IMAGE:figures/full_fig_p025_16.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

76 extracted references · 59 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence 'output.state := if if FUNCTION not #0 #1 if FUNCTION and 'skip pop #0 if FUNCTIO...

  2. [2]

    Albrecht, J. P. (2016). How the gdpr will change the world. Eur. Data Prot. L. Rev. , 2:287

  3. [3]

    T., and Weld, D

    Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., and Weld, D. S. (2021). Does the whole exceed its parts? the effect of ai explanations on complementary team performance. (arXiv:2006.14779). arXiv:2006.14779 [cs]

  4. [4]

    Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., Garcia, S., Gil-Lopez, S., Molina, D., Benjamins, R., Chatila, R., and Herrera, F. (2020). Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information Fusion , 58:82–115

  5. [5]

    and Kohavi, R

    Becker, B. and Kohavi, R. (1996). Adult . UCI Machine Learning Repository. DOI : https://doi.org/10.24432/C5XW20

  6. [6]

    R., and Maxwell, W

    Bertrand, A., Viard, T., Belloum, R., Eagan, J. R., and Maxwell, W. (2023). On selective, mutable and dialogic xai: a review of what users say about different types of interactive explanations. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , page 1–21, Hamburg Germany. ACM

  7. [7]

    and Shu, S

    Bogard, J. and Shu, S. (2022). Algorithm Aversion and the Aversion to Counter-Normative Decision Procedures

  8. [8]

    B., and Gajos, K

    Buçinca, Z., Malaya, M. B., and Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction , 5(CSCW1):188:1--188:21

Show all 76 references
  1. [9]

    E., Doshi-Velez, F., and Gajos, K

    Buçinca, Z., Swaroop, S., Paluch, A. E., Doshi-Velez, F., and Gajos, K. Z. (2024). Contrastive explanations that anticipate human misconceptions can improve human decision-making skills. (arXiv:2410.04253). arXiv:2410.04253

  2. [10]

    Carton, S., Mei, Q., and Resnick, P. (2020). Feature-based explanations don’t help people detect misclassifications of online toxicity. Proceedings of the International AAAI Conference on Web and Social Media , 14:95–106

  3. [11]

    Castelnovo, A., Crupi, R., Mombelli, N., Nanino, G., and Regoli, D. (2023). Evaluative item-contrastive explanations in rankings. (arXiv:2312.10094). arXiv:2312.10094 [cs]

  4. [12]

    W., and Lehmann, D

    Castelo, N., Bos, M. W., and Lehmann, D. R. (2019). Task-dependent algorithm aversion. Journal of Marketing Research , 56(5):809–825

  5. [13]

    L., Schonger, M., and Wickens, C

    Chen, D. L., Schonger, M., and Wickens, C. (2016). otree—an open-source platform for laboratory, online, and field experiments. Journal of Behavioral and Experimental Finance , 9:88--97

  6. [14]

    V., Vaughan, J

    Chen, V., Liao, Q. V., Vaughan, J. W., and Bansal, G. (2023). Understanding the role of human intuition on reliance in human-ai decision-making with explanations. (arXiv:2301.07255). arXiv:2301.07255 [cs]

  7. [15]

    M., and Zhu, H

    Cheng, H.-F., Wang, R., Zhang, Z., O’Connell, F., Gray, T., Harper, F. M., and Zhu, H. (2019). Explaining decision-making algorithms through ui: Strategies to help non-expert stakeholders. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , CHI ’1...

  8. [16]

    Chew, R., Bollenbacher, J., Wenger, M., Speer, J., and Kim, A. (2023). Llm-assisted content analysis: Using large language models to support deductive coding. (arXiv:2306.14924). arXiv:2306.14924

  9. [17]

    Chromik, M., Eiband, M., Buchner, F., Krüger, A., and Butz, A. (2021). I think i get your point, ai! the illusion of explanatory depth in explainable ai. In 26th International Conference on Intelligent User Interfaces , IUI ’21, page 307–317, New York, NY, USA. Association for...

  10. [18]

    C., Sui, Y., Kumar, B., and Vouitsis, N

    Cresswell, J. C., Sui, Y., Kumar, B., and Vouitsis, N. (2024). Conformal prediction sets improve human decision making. (arXiv:2401.13744). arXiv:2401.13744 [cs, stat]

  11. [19]

    Croskerry, P. (2009). A universal model of diagnostic reasoning. Academic medicine , 84(8):1022--1028

  12. [20]

    Faul, F., Erdfelder, E., Lang, A.-G., and Buchner, A. (2007). G* power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior research methods , 39(2):175--191

  13. [21]

    Gajos, K. Z. and Mamykina, L. (2022). Do people engage cognitively with ai? impact of ai assistance on incidental learning. In 27th International Conference on Intelligent User Interfaces , page 794–806, Helsinki Finland. ACM

  14. [22]

    V., Zhang, Y., Bellamy, R., and Mueller, K

    Ghai, B., Liao, Q. V., Zhang, Y., Bellamy, R., and Mueller, K. (2020). Explainable active learning (xal): An empirical study of how local explanations impact annotator experience. (arXiv:2001.09219). arXiv:2001.09219 [cs]

  15. [23]

    o der, C., and Schupp, J. (2019). The german socio-economic panel (soep). Jahrb \

    Goebel, J., Grabka, M. M., Liebig, S., Kroh, M., Richter, D., Schr \"o der, C., and Schupp, J. (2019). The german socio-economic panel (soep). Jahrb \"u cher f \"u r National \"o konomie und Statistik , 239(2):345--360

  16. [24]

    S., Ahuja, N., Horvitz, E., Yang, D., Milstein, A., Olson, A

    Goh, E., Gallo, R., Hom, J., Strong, E., Weng, Y., Kerman, H., Cool, J., Kanjee, Z., Parsons, A. S., Ahuja, N., Horvitz, E., Yang, D., Milstein, A., Olson, A. P., Rodman, A., and Chen, J. H. (2024). Influence of a large language model on diagnostic reasoning: A randomized clin...

  17. [25]

    Gouveia, S. S. and Malík, J. (2024). Crossing the trust gap in medical ai: Building an abductive bridge for xai. Philosophy & Technology , 37(3):105

  18. [26]

    Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., and Pedreschi, D. (2018). A survey of methods for explaining black box models. ACM Computing Surveys , 51(5):93:1--93:42

  19. [27]

    Hart, S. G. (2006). Nasa-task load index (nasa-tlx); 20 years later. Proceedings of the Human Factors and Ergonomics Society Annual Meeting , 50(9):904–908

  20. [28]

    Hemmer, P., Schemmer, M., Kühl, N., Vössing, M., and Satzger, G. (2024). Complementarity in human-ai collaboration: Concept, sources, and evidence. (arXiv:2404.00029). arXiv:2404.00029 [cs]

  21. [29]

    Hemmer, P., Schemmer, M., Vössing, M., and Kühl, N. (2021). Human-ai complementarity in hybrid intelligence systems: A structured literature review. In PACIS 2021 Proceedings

  22. [30]

    Herm, L.-V. (2023). Impact of explainable ai on cognitive load: Insights from an empirical study

  23. [31]

    Jacovi, A., Marasović, A., Miller, T., and Goldberg, Y. (2021). Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in ai. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , FAccT ’21, page 624–635...

  24. [32]

    Kaur, H., Nori, H., Jenkins, S., Caruana, R., Wallach, H., and Wortman Vaughan, J. (2020). Interpreting interpretability: Understanding data scientists’ use of interpretability tools for machine learning. In Proceedings of the 2020 CHI Conference on Human Factors in Computing ...

  25. [33]

    K., Rall, E

    Klein, G., Phillips, J. K., Rall, E. L., and Peluso, D. A. (2007). A data--frame theory of sensemaking. In Expertise out of context , pages 118--160. Psychology Press

  26. [34]

    and Thommes, K

    Kornowicz, J. and Thommes, K. (2024). Algorithm, expert, or both? evaluating the role of feature selection methods on user preferences and reliance. (arXiv:2408.01171). arXiv:2408.01171 [cs]

  27. [35]

    V., and Tan, C

    Lai, V., Chen, C., Smith-Renner, A., Liao, Q. V., and Tan, C. (2023a). Towards a science of human-ai decision making: An overview of design space in empirical human-subject studies. In 2023 ACM Conference on Fairness, Accountability, and Transparency , page 1369–1385, Chicago ...

  28. [36]

    why is ‘chicago’ deceptive?

    Lai, V., Liu, H., and Tan, C. (2020). “why is ‘chicago’ deceptive?” towards building model-driven tutorials for humans. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems , page 1–13, Honolulu HI USA. ACM

  29. [37]

    and Tan, C

    Lai, V. and Tan, C. (2019). On human predictions with explanations and predictions of machine learning models: A case study on deception detection. In Proceedings of the Conference on Fairness, Accountability, and Transparency , page 29–38, Atlanta GA USA. ACM

  30. [38]

    V., and Tan, C

    Lai, V., Zhang, Y., Chen, C., Liao, Q. V., and Tan, C. (2023b). Selective explanations: Leveraging human input to align explainable ai. (arXiv:2301.09656). arXiv:2301.09656 [cs]

  31. [39]

    Le, T., Miller, T., Singh, R., and Sonenberg, L. (2023). Explaining model confidence using counterfactuals. Proceedings of the AAAI Conference on Artificial Intelligence , 37(1010):11856–11864

  32. [40]

    Le, T., Miller, T., Sonenberg, L., and Singh, R. (2024a). Towards the new xai: A hypothesis-driven approach to decision support using evidence. (arXiv:2402.01292). arXiv:2402.01292 [cs]

  33. [41]

    Le, T., Miller, T., Zhang, R., Sonenberg, L., and Singh, R. (2024b). Visual evaluative ai: A hypothesis-driven tool with concept-based explanations and weight of evidence. (arXiv:2407.04710). arXiv:2407.04710 [cs]

  34. [42]

    Liu, H., Lai, V., and Tan, C. (2021). Understanding the effect of out-of-distribution examples and interactive explanations on human-ai decision making. Proceedings of the ACM on Human-Computer Interaction , 5(CSCW2):408:1--408:45

  35. [43]

    Lundberg, S. M. and Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R., editors, Advances in Neural Information Processing Systems 30 , pages 4765--4774. ...

  36. [44]

    and Coiera, E

    Lyell, D. and Coiera, E. (2017). Automation bias and verification complexity: a systematic review. Journal of the American Medical Informatics Association , 24(2):423–431

  37. [45]

    Ma, S., Lei, Y., Wang, X., Zheng, C., Shi, C., Yin, M., and Ma, X. (2023). Who should i trust: Ai or myself? leveraging human and ai correctness likelihood to promote appropriate trust in ai-assisted decision-making. In Proceedings of the 2023 CHI Conference on Human Factors i...

  38. [46]

    MacCarthy, M. (2019). An examination of the algorithmic accountability act of 2019. Available at SSRN 3615731

  39. [47]

    Mahmud, H., Islam, A. K. M. N., Ahmed, S. I., and Smolander, K. (2022). What influences algorithmic decision-making? a systematic literature review on algorithm aversion. Technological Forecasting and Social Change , 175:121390

  40. [48]

    Malone, T., Vaccaro, M., Campero, A., Song, J., Wen, H., and Almaatouq, A. (2023). A test for evaluating performance in human-ai systems

  41. [49]

    Mayring, P. (2015). Qualitative Content Analysis: Theoretical Background and Procedures , page 365–380. Springer Netherlands, Dordrecht

  42. [50]

    Miller, T. (2019). Explanation in artificial intelligence: Insights from the social sciences. Artificial Intelligence , 267:1–38

  43. [51]

    Miller, T. (2023). Explainable ai is dead, long live explainable ai!: Hypothesis-driven decision support using evaluative ai. In 2023 ACM Conference on Fairness, Accountability, and Transparency , page 333–342, Chicago IL USA. ACM

  44. [52]

    A., Veness, J., Bellemare, M

    Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D. (2015). Human-le...

  45. [53]

    M., Carignan, D., and Horvitz, E

    Nori, H., King, N., McKinney, S. M., Carignan, D., and Horvitz, E. (2023). Capabilities of gpt-4 on medical challenge problems. arXiv preprint arXiv:2303.13375

  46. [54]

    Peirce, C. S. (2009). Writings of Charles S. Peirce: a chronological edition, volume 8: 1890--1892 , volume 8. Indiana University Press

  47. [55]

    Popper, K. (2014). Conjectures and refutations: The growth of scientific knowledge . routledge

  48. [56]

    G., Hofman, J

    Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Wortman Vaughan, J. W., and Wallach, H. (2021). Manipulating and measuring model interpretability. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , CHI ’21, page 1–52, New York, NY, USA. A...

  49. [57]

    T., Singh, S., and Guestrin, C

    Ribeiro, M. T., Singh, S., and Guestrin, C. (2018). Anchors: High-precision model-agnostic explanations. Proceedings of the AAAI Conference on Artificial Intelligence , 32(11)

  50. [58]

    Risko, E. F. and Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences , 20(9):676–688

  51. [59]

    Rogha, M. (2023). Explain to decide: A human-centric review on the role of explainable artificial intelligence in ai-assisted decision making. (arXiv:2312.11507). arXiv:2312.11507 [cs]

  52. [60]

    Rong, Y., Leemann, T., Nguyen, T.-t., Fiedler, L., Seidel, T., Kasneci, G., and Kasneci, E. (2022). Towards human-centered explainable ai: User studies for model explanations. (arXiv:2210.11584). arXiv:2210.11584 [cs]

  53. [61]

    Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence , 1(5):206–215

  54. [62]

    Sawyer, S. F. (2009). Analysis of variance: The fundamental concepts. Journal of Manual & Manipulative Therapy

  55. [63]

    Schemmer, M., Hemmer, P., Nitsche, M., Kühl, N., and Vössing, M. (2022). A meta-analysis of the utility of explainable artificial intelligence in human-ai decision-making. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society , page 617–626. arXiv:2205.05126 [cs]

  56. [64]

    Schemmer, M., Kuehl, N., Benz, C., Bartos, A., and Satzger, G. (2023). Appropriate reliance on ai advice: Conceptualization and the effect of explanations. In Proceedings of the 28th International Conference on Intelligent User Interfaces , page 410–422, Sydney NSW Australia. ACM

  57. [65]

    Schuff, D., Corral, K., and Turetken, O. (2011). Comparing the understandability of alternative data warehouse schemas: An empirical study. Decision Support Systems , 52(1):9–20

  58. [66]

    A., Scheidegger, C., and Roy, C

    Slack, D., Friedler, S. A., Scheidegger, C., and Roy, C. D. (2019). Assessing the local interpretability of machine learning models. (arXiv:1902.03501). arXiv:1902.03501

  59. [67]

    Spatola, N. (2024). The efficiency-accountability tradeoff in ai integration: Effects on human performance and over-reliance. Computers in Human Behavior: Artificial Humans , page 100099

  60. [68]

    H., Bentley, L

    Tai, R. H., Bentley, L. R., Xia, X., Sitt, J. M., Fankhauser, S. C., Chicas-Mosier, A. M., and Monteith, B. G. (2024). An examination of the use of large language models to aid analysis of textual data. International Journal of Qualitative Methods , 23:16094069241231168

  61. [69]

    S., and Krishna, R

    Vasconcelos, H., Jörke, M., Grunde-McLaughlin, M., Gerstenberg, T., Bernstein, M. S., and Krishna, R. (2023). Explanations can reduce overreliance on ai systems during decision-making. Proceedings of the ACM on Human-Computer Interaction , 7(CSCW1):1–38

  62. [70]

    Vered, M., Livni, T., Howe, P. D. L., Miller, T., and Sonenberg, L. (2023). The effects of explanations on automation bias. Artificial Intelligence , 322:103952

  63. [71]

    Wang, D., Yang, Q., Abdul, A., and Lim, B. Y. (2019). Designing theory-driven user-centric explainable ai. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems , page 1–15, Glasgow Scotland Uk. ACM

  64. [72]

    and Yin, M

    Wang, X. and Yin, M. (2021). Are explanations helpful? a comparative study of the effects of explanations in ai-assisted decision-making. In 26th International Conference on Intelligent User Interfaces , IUI ’21, page 318–328, New York, NY, USA. Association for Computing Machinery

  65. [73]

    Wischnewski, M., Krämer, N., and Müller, E. (2023). Measuring and understanding trust calibrations for automated systems: A survey of the state-of-the-art and future directions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , CHI ’23, page 1–1...

  66. [74]

    Yates, J. F. and Potworowski, G. A. (2012). Evidence-based decision management

  67. [75]

    L., and Li, X

    You, S., Yang, C. L., and Li, X. (2022). Algorithmic versus human advice: Does presenting prediction performance matter for algorithm appreciation? Journal of Management Information Systems , 39(2):336–365

  68. [76]

    V., and Bellamy, R

    Zhang, Y., Liao, Q. V., and Bellamy, R. K. E. (2020). Effect of confidence and explanation on accuracy and trust calibration in ai-assisted decision making. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency , FAT* ’20, page 295–305, New York, ...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.