Pith. sign in

REVIEW 4 major objections 4 minor 2 cited by

This paper argues that AI explanation effectiveness should be assessed by the actions users take, and offers a 12-category, 60-action catalog from doctor and teacher interviews.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 07:30 UTC pith:EOPVIZEK

load-bearing objection A transparent, user-derived catalog of information-action links for XAI evaluation; the descriptive claims hold up, but the action frequencies are self-reported intentions from hypothetical scenarios, so the Mental State Action emphasis is a proposal, not a behavioral finding. the 4 major comments →

arxiv 2601.20086 v1 pith:EOPVIZEK submitted 2026-01-27 cs.HC

Evaluating Actionability in Explainable AI

classification cs.HC
keywords actionabilityexplainable AIuser-centered evaluationmental state actionsinformation categoriesscenario-based designqualitative studyXAI
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to make actionability measurable in explainable AI: instead of asking whether an explanation is clear or faithful, it asks what a user can do with the information. From 14 scenario-based interviews with doctors and teachers, it builds a catalog of 12 information categories and 60 user-defined action types, organized into AI interactions, external actions, and mental state actions. Its central empirical finding is that mental state actions—trusting, understanding, comparing, setting expectations, deciding—were by far the most frequent responses in the data, with 2,190 mentions against 200 AI interactions and 270 external actions. A sympathetic reading is that the paper establishes a shared vocabulary and mapping that AI creators can use to state how their explanations are supposed to help users and then test whether they do.

Core claim

The paper contributes an evaluation resource: a catalog anchored in user-centered terminology, where users' words define the concepts. It contains 12 User-Centered Information Categories, grouped into four themes (Model Exposure, Model Creation, User Environment, User Accountability), and 60 User-Centered Action Types across three dimensions: 13 AI Interactions, 17 External Actions, and 30 Mental State Actions. The mapping is built from what participants said they would rely on and do: each information category is paired with the actions participants associated with it. The headline discovery is that Mental State Actions were the modal action type, with 2,190 mentions compared to 200 AI Inte

What carries the argument

The central object is the catalog itself, built from three components: User-Centered Terminology (definitions derived from participants, e.g., 'AI system' means algorithm plus interface plus explanation), 12 User-Centered Information Categories, and 60 User-Centered Action Types. The load-bearing mechanism is the mapping between information categories and actions: each category is linked to specific high-, medium-, and low-frequency actions, so an evaluator can move from 'what information is displayed' to 'what action should follow.' The study method—scenario-based design interviews using non-interactive mock EHR and course-placement interfaces—is what generates the catalog, and the paper's

Load-bearing premise

The load-bearing premise, which the paper's limitations section acknowledges, is that what doctors and teachers said they would do in a hypothetical, non-interactive scenario matches what they would actually do with a real AI system in their daily work; the study gathered stated intentions, not observed actions.

What would settle it

A study that gives doctors and teachers a real interactive XAI decision-support system and logs actual behavior—every click, search, consultation, and self-reported change in trust or understanding—would settle the claim. If the observed actions cannot be classified into the 60 catalog actions, or if mental state actions are not the most frequent class in real use, the catalog's claim to map actionability fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • AI creators can use the catalog to write explicit expectations of the form 'this piece of information should enable this action,' then test those expectations with surveys or interviews.
  • Evaluations of XAI should treat mental state changes as first-class outcomes, since users report relying on explanations primarily to change what they trust, understand, expect, and decide.
  • The catalog's distinction between AI interactions, external actions, and mental state actions gives evaluators a common vocabulary for comparing findings across different explanation systems and domains.
  • The 12 information categories expand the design space for explanations beyond feature attribution, including social information such as other users' experiences and qualifications, system-support pathways, and consequences for stakeholders.
  • The welfare-worker example indicates the catalog can be transferred to domains beyond the two studied professions, providing a starting point for evaluation in new settings.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A testable extension of this result: log actual behavior in a deployed interactive XAI system; if observed actions fall outside the 60 catalog actions, or if mental state actions are not the most frequent class, the catalog's completeness claim would need revision.
  • If mental state actions dominate even at roughly the same rate in real use, then behavioral telemetry alone (clicks, prints, messages) will systematically undercount how much effect explanations have, and evaluation instruments will need to probe internal states directly.
  • Because the 'mental state action' category is broad, an evaluator using it without pre-registered definitions could make almost any explanation seem actionable; a sharper test would pre-specify which mental state changes count as actions before data collection.
  • The catalog could seed a question bank for post-deployment XAI evaluation, letting organizations ask users which of the 60 actions an explanation enabled, and compare responses against the expected information-action pairs.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper reports a qualitative interview study (n=14: 9 doctors, 5 teachers) using scenario-based design with non-interactive mock EHR and course-placement interfaces to elicit what information end users would rely on and what actions they would take in response to AI explanations. It contributes a catalog of 12 User-Centered Information Categories and 60 User-Centered Action Types organized into three dimensions: AI Interactions, External Actions, and Mental State Actions. The paper reports Mental State Actions as the modal action category (2190 vs. 200 AI Interactions and 270 External Actions, §4.1) and proposes that AI Creators use the catalog to articulate expected information–action links and evaluate XAI systems, illustrating with a child-welfare example.

Significance. If the catalog accurately reflects end-user information needs and action repertoires, it fills a genuine gap: prior actionability work is fragmented by domain and technique. The paper's strengths are its transparent coding process, user-derived terminology, full codebook in appendices, and direct quotes grounding each category. The application example to a third domain (child welfare) demonstrates an intent toward transferability. The central limitation—self-reported hypothetical intentions rather than observed behavior—bears directly on the strong frequency claims, so the catalog is best seen as a formative map rather than a validated instrument. The paper is honest about this in Section 7, but the framing and the specific frequency claims outrun the evidence.

major comments (4)
  1. [§4.1 and §6.2] The claim that Mental State Actions are the modal action category (2190 vs. 200 and 270) is load-bearing for the argument that mental states are a 'critical dimension of actionability.' These counts derive from coding verbal reports elicited with a non-interactive interface in which participants were explicitly asked to 'describe the expected or desired behavior of an interface element' (Section 3). This prompt likely inflates mental-state verbs (trust, understand, feel) relative to observable interactions, which would be underrepresented because the interface did not actually respond. The Section 7 acknowledgment that the method depends on scenario engagement does not address this artifact. Either soften the frequency-based claim to 'frequently mentioned in our interviews,' or provide a supplemental check—e.g., coding interaction logs from a deployed system or a small think-aloud with a
  2. [§3.3 and §4] The catalog is induced from and evidenced by the same 14 interviews. The only reliability evidence is Cohen's Kappa = 83.5% on a 20% subset (160/737 quotes), with final coding by a single author. For a descriptive catalog this is acceptable as formative, but the paper then frames the catalog as a tool for AI Creators to 'test their assumptions' (Section 5). That application requires some evidence of stability or transfer beyond the derivation sample—e.g., member checking, a second round of interviews, or independent application by another research group. Without that, the abstract's contribution claim ('maps 12 categories... to 60 actions') is stronger than the evidence supports. Please add a qualification to the abstract and to Section 5.
  3. [§4.2 and §5.2] The language 'lead to' and 'in their explanations should lead to user actions' (abstract, Section 5.2) implies a causal or enabling relation between information categories and actions. The data are co-occurrences in participants' narratives: participants mentioned information and actions in the same interview, and the Appendix shows high/medium/low frequency associations. No temporal or mechanistic link was established. Recommend rephrasing to 'information that participants described relying on alongside these actions' or 'associated actions' to avoid overcommitting the catalog to a causal model that the data cannot support.
  4. [Appendix B, Table 4] The taxonomy places 'Change state' — 'Alters their own mental or physical state based on the system (yes, it's a broad category)' — among External Actions, while Mental State Actions include 'Emote or feel things about the system.' A participant's statement 'I would feel better' could plausibly be coded under either category. This overlap blurs the three-dimension distinction that underlies the frequency comparison and the claim that Mental State Actions are distinct from External Actions. Please clarify the coding rule for distinguishing these two codes, and report intercoder agreement per action type or at least per dimension.
minor comments (4)
  1. [Abstract] Typo: 'willdosomething' should be 'will do something'.
  2. [§4.1] The phrase 'constituted the mode' is ambiguous in context; 'modal category' would be clearer. Also, 'with a total 2190 Mental State Actions' should read 'with a total of 2190 Mental State Actions.'
  3. [Appendix B] Minor spelling inconsistency: 'Share a decision/rational with other people' appears in Section 5.2 and in Table 2 of Appendix B; 'rationale' is the correct form. The External Actions table (Table 4) uses 'rationale' correctly.
  4. [Section 3.3] The coding process is described well, but it would help to state explicitly how many codes were applied per quote (e.g., single vs. multiple codes) because the reported totals (2190, 200, 270) are otherwise difficult to interpret without this unit-of-analysis information.

Circularity Check

0 steps flagged

No circularity: the catalog is an explicitly descriptive qualitative summary of interview data, not a predictive claim fitted to its own inputs.

full rationale

The paper's central contribution is a catalog of information categories and action types derived from 14 interviews (§3). The claimed output—'Our catalog maps 12 categories of information that participants described relying on to take 60 different actions'—is presented as a qualitative summary of those self-reports, not as a prediction or first-principles derivation. No parameter is fitted to a subset of data and then used to predict a closely related quantity; the action frequencies, including the modal Mental State Actions count of 2190 (§4.1), are direct counts of coded utterances. The taxonomy is admittedly induced from the same interviews used to illustrate it, which limits external confirmation, but this is a generalizability/validity limitation, not a circular reduction. The inter-rater reliability check on ~22% of quotes (§3.3) is a coding-consistency measure, not an independent benchmark, and the paper itself flags the formative nature and scenario-dependence of the method in §7. The self-citations to prior work by the same group (e.g., Ehsan et al.) provide methodological context and are not load-bearing for the catalog's content. Because the paper makes no predictive claim that reduces by construction to its inputs, no circular step meets the required evidentiary standard.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 2 invented entities

The paper's central artifact is a qualitative taxonomy. Every category and action type is constructed from 14 self-report interviews; there are no quantitative fitted parameters. Validity rests on internal coding consistency rather than external measurement.

axioms (4)
  • domain assumption Participants can accurately forecast which actions they would take from a non-interactive scenario-based interface.
    Used throughout Section 3; Section 7 acknowledges the method's dependency on scenario engagement and participant experience.
  • domain assumption Two professions (medicine and education) provide commonalities transferable to other high-stakes decision domains.
    Sampling rationale in Section 3; Section 7 concedes that differences across domains likely exist.
  • domain assumption Coder-derived categories validated by inter-rater reliability on a 20% quote subset adequately support the taxonomy.
    Section 3.3; no external audit, participant validation, or third-party coding of the full dataset.
  • ad hoc to paper Mental state changes count as 'actions' for the purpose of actionability.
    Section 4.1 defines Mental State Actions and makes them central to the catalog; this is an author-defined construct choice.
invented entities (2)
  • Mental State Actions as a dimension of actionability no independent evidence
    purpose: To count belief and emotion changes (trust, understanding, frustration, goal-setting) as actions users take in response to explanations.
    Newly named construct in Section 4.1; frequency claims rely on coding with this category and no external behavioral measure.
  • Twelve User-Centered Information Categories no independent evidence
    purpose: Taxonomy of information users seek from XAI systems to take actions.
    Induced from interview data in Sections 3.3 and 4.2; category definitions are author-determined and lack external benchmark validation.

pith-pipeline@v1.3.0-alltime-deepseek · 20949 in / 10722 out tokens · 123309 ms · 2026-08-03T07:30:26.159978+00:00 · methodology

0 comments
read the original abstract

A core assumption of Explainable AI (XAI) is that explanations are useful to users -- that is, users will do something with the explanations. Prior work, however, does not clearly connect the information provided in explanations to user actions to evaluate effectiveness. In this paper, we articulate this connection. We conducted a formative study through 14 interviews with end users in education and medicine. We contribute a catalog of information and associated actions. Our catalog maps 12 categories of information that participants described relying on to take 60 different actions. We show how AI Creators can use the catalog's specificity and breadth to articulate how they expect information in their explanations to lead to user actions and test their assumptions. We use an exemplar XAI system to illustrate this approach. We conclude by discussing how our catalog expands the design space for XAI systems to support actionability.

Figures

Figures reproduced from arXiv: 2601.20086 by Gennie Mansi, Julia Kim, Mark Riedl.

Figure 1
Figure 1. Figure 1: Information Categories and their definitions categorized by theme. model against other knowledge, and ⋄ Understand broader information about factors impacting the current environment or explanation, among other actions. The top External Actions for this Information Category involve interacting with another person: ⋄ Communicate with another person, ⋄ Consult another person about a decision, and ⋄ Support a… view at source ↗
Figure 2
Figure 2. Figure 2: Information Categories under Model Exposure and all actions associated with them [PITH_FULL_IMAGE:figures/full_fig_p032_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Information Categories under Model Creation and all actions associated with them [PITH_FULL_IMAGE:figures/full_fig_p033_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Information Categories under User Environment and all actions associated with them [PITH_FULL_IMAGE:figures/full_fig_p034_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Information Categories under User Accountability and all actions associated with them [PITH_FULL_IMAGE:figures/full_fig_p035_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From Attribution to Action: A Human-Centered Application of Activation Steering

    cs.AI 2026-04 conditional novelty 6.5

    Activation steering of SAE-attributed components lets practitioners move from correlational inspection to causal hypothesis testing on CLIP failures, with trust shifting to observed model responses (N=8 experts).

  2. From Attribution to Action: A Human-Centered Application of Activation Steering

    cs.AI 2026-04 unverdicted novelty 6.0

    Activation steering paired with attribution enables intervention-based debugging in vision models, as all 8 interviewed experts shifted to hypothesis testing, most trusted observed responses, and highlighted risks lik...

Reference graph

Works this paper leans on

66 extracted references · 18 canonical work pages · cited by 1 Pith paper · 4 internal anchors

  1. [1]

    ICO & Turing Institute (2022), https://ico.org.uk/for-organisations/guide-to-data-protection/key-dp- themes/explaining-decisions-made-with-artificial-intelligence/ 20 G

    Explaining decisions made with AI. ICO & Turing Institute (2022), https://ico.org.uk/for-organisations/guide-to-data-protection/key-dp- themes/explaining-decisions-made-with-artificial-intelligence/ 20 G. Mansi et al

  2. [2]

    Hum.-Comput

    Ackerman, M.S.: The intellectual challenge of cscw: the gap between social require- ments and technical feasibility. Hum.-Comput. Interact.15(2), 179–203 (sep 2000). https://doi.org/10.1207/S15327051HCI1523_5

  3. [3]

    https://doi.org/10.1109/ACCESS.2018.2870052, conference Name: IEEE Ac- cess

    Adadi, A., Berrada, M.: Peeking inside the black-box: A survey on explainable artificial intelligence (XAI)6, 52138–52160 (2018). https://doi.org/10.1109/ACCESS.2018.2870052, conference Name: IEEE Ac- cess

  4. [4]

    Adenuga, I., Dodge, J.: Conceptualizing the relationship between AI explanations and user agency (2023)

  5. [5]

    https://doi.org/10.1007/s00146-021-01326-6, https://doi.org/10.1007/s00146-021-01326-6

    Andrada, G., Clowes, R.W., Smart, P.R.: Varieties of transparency: exploring agency within AI systems (2022). https://doi.org/10.1007/s00146-021-01326-6, https://doi.org/10.1007/s00146-021-01326-6

  6. [6]

    https://doi.org/10.48550/ARXIV.1910.10045, https://arxiv.org/abs/1910.10045

    Arrieta,A.B.,Díaz-Rodríguez,N.,DelSer,J.,Bennetot,A.,Tabik,S.,Barbado,A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., Herrera, F.: Ex- plainable artificial intelligence (xai): Concepts, taxonomies, opportunities and chal- lenges toward responsible ai (2019). https://doi.org/10.48550/ARXIV.1910.10045, https://arxiv.org/abs/1910.10045

  7. [7]

    https://doi.org/10.48550/ARXIV.1909.03012, https://arxiv.org/abs/1909.03012

    Arya, V., Bellamy, R.K.E., Chen, P.Y., Dhurandhar, A., Hind, M., Hoff- man, S.C., Houde, S., Liao, Q.V., Luss, R., Mojsilović, A., Mourad, S., Pedemonte, P., Raghavendra, R., Richards, J., Sattigeri, P., Shanmugam, K., Singh, M., Varshney, K.R., Wei, D., Zhang, Y.: One explanation does not fit all: A toolkit and taxonomy of ai explainability techniques (2...

  8. [8]

    Bharambe, U., Srinivasaraghavan, A.: Actionable Content Discov- ery for Healthcare, chap. 5, pp. 115–138. John Wiley and Sons, Ltd (2021). https://doi.org/https://doi.org/10.1002/9781119764175.ch5, https://onlinelibrary.wiley.com/doi/abs/10.1002/9781119764175.ch5

  9. [9]

    (ed.): Scenario-based design: envisioning work and technology in sys- tem development

    Carroll, J.M. (ed.): Scenario-based design: envisioning work and technology in sys- tem development. John Wiley & Sons, Inc., USA (1995)

  10. [11]

    Chari, S., Seneviratne, O.W., Gruen, D., Foreman, M., Das, A.K., McGuinness, D.L.: Explanation ontology: A model of explanations for user-centered ai (2020)

  11. [12]

    Chen, V., Li, J., Kim, J.S., Plumb, G., Talwalkar, A.: Interpretable machine learn- ing: Moving from mythos to diagnostics. vol. 19, p. 28–56. Association for Com- puting Machinery, New York, NY, USA (2022). https://doi.org/10.1145/3511299, https://doi.org/10.1145/3511299

  12. [13]

    In: Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems

    Cheng, H.F., Stapleton, L., Kawakami, A., Sivaraman, V., Cheng, Y., Qing, D., Perer, A., Holstein, K., Wu, Z.S., Zhu, H.: How child welfare workers reduce racial disparities in algorithmic decisions. In: Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. CHI ’22, Association for Computing Machinery,NewYork,NY,USA(2022).https://d...

  13. [14]

    In: 2018 IEEE 20th Interna- tional Conference on e-Health Networking, Applications and Services (Healthcom)

    Chiang, P.H., Dey, S.: Personalized effect of health behavior on blood pressure: Ma- chine learning based prediction and recommendation. In: 2018 IEEE 20th Interna- tional Conference on e-Health Networking, Applications and Services (Healthcom). pp. 1–6 (2018). https://doi.org/10.1109/HealthCom.2018.8531109 Evaluating Actionability in Explainable AI 21

  14. [15]

    In: Proceedings of the 42nd In- ternational ACM SIGIR Conference on Research and Development in Infor- mation Retrieval

    Cho, M., Lee, G., Hwang, S.w.: Explanatory and actionable debugging for machine learning: A tableqa demonstration. In: Proceedings of the 42nd In- ternational ACM SIGIR Conference on Research and Development in Infor- mation Retrieval. p. 1333–1336. SIGIR’19, Association for Computing Ma- chinery, New York, NY, USA (2019). https://doi.org/10.1145/3331184....

  15. [16]

    Coeckelbergh, M.: Artificial intelligence, responsibility attribution, and a relational justification of explainability. vol. 26, pp. 2051–2068 (2020). https://doi.org/10.1007/s11948-019-00146-8, https://doi.org/10.1007/s11948-019- 00146-8

  16. [17]

    In: Guarda, T., Portela, F., Augusto, M.F

    Coroama, L., Groza, A.: Evaluation metrics in explainable artificial intelligence (XAI). In: Guarda, T., Portela, F., Augusto, M.F. (eds.) Advanced Research in Technologies, Information, Innovation and Sustainability. pp. 401–413. Springer Nature Switzerland (2022)

  17. [18]

    it is a moving process

    Corti, L., Oltmans, R., Jung, J., Balayn, A., Wijsenbeek, M., Yang, J.: “it is a moving process": Understanding the evolution of explainability needs of clin- icians in pulmonary medicine. In: Proceedings of the CHI Conference on Hu- man Factors in Computing Systems. CHI ’24, Association for Computing Ma- chinery, New York, NY, USA (2024). https://doi.org...

  18. [19]

    Dattathrani, S., De’, R.: The concept of agency in the era of artificial intelligence: Dimensionsanddegrees.vol.25,pp.29–54(2023).https://doi.org/10.1007/s10796- 022-10336-8, https://doi.org/10.1007/s10796-022-10336-8

  19. [20]

    Deschênes, M.: Recommender systems to support learners’ agency in a learning context: a systematic review. vol. 17, p. 50 (2020). https://doi.org/10.1186/s41239- 020-00219-w, https://doi.org/10.1186/s41239-020-00219-w

  20. [21]

    In: Dohn, N.B., Jandrić, P., Ryberg, T., de Laat, M

    Dohn, N.B., Ryberg, T., de Laat, M., Jandrić, P.: Conclusion: Mobility, data and learner agency in networked learning. In: Dohn, N.B., Jandrić, P., Ryberg, T., de Laat, M. (eds.) Mobility, Data and Learner Agency in Networked Learning, pp. 193–213. Springer International Publishing (2020). https://doi.org/10.1007/978-3- 030-36911-8_12

  21. [22]

    https://doi.org/10.48550/ARXIV.1702.08608, https://arxiv.org/abs/1702.08608

    Doshi-Velez, F., Kim, B.: Towards a rigorous science of interpretable machine learning (2017). https://doi.org/10.48550/ARXIV.1702.08608, https://arxiv.org/abs/1702.08608

  22. [23]

    https://doi.org/10.2139/ssrn.2972855, https://papers.ssrn.com/abstract=2972855

    Edwards, L., Veale, M.: Slave to the algorithm? why a ’right to an explanation’ is probably not the remedy you are looking for (2017). https://doi.org/10.2139/ssrn.2972855, https://papers.ssrn.com/abstract=2972855

  23. [24]

    In: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems

    Ehsan, U., Liao, Q.V., Muller, M., Riedl, M.O., Weisz, J.D.: Ex- panding explainability: Towards social transparency in ai systems. In: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. CHI ’21, Association for Computing Machinery, New York, NY, USA (2021). https://doi.org/10.1145/3411764.3445188, https://doi.org/10.1145/341176...

  24. [25]

    Ehsan, U., Liao, Q.V., Passi, S., Riedl, M.O., Daumé, H.: Seamful xai: Operationalizing seamful design in explainable ai. Proc. ACM Hum.- Comput. Interact.8(CSCW1) (apr 2024). https://doi.org/10.1145/3637396, https://doi.org/10.1145/3637396

  25. [26]

    In: Proceedings of the Tenth ACM International Conference on Web Search and Data Mining

    Ensan,F.,Noorian,Z., Bagheri, E.:Miningactionableinsightsfromsocialnetwork- sat wsdm 2017. In: Proceedings of the Tenth ACM International Conference on Web Search and Data Mining. p. 821–822. WSDM ’17, Association for Computing 22 G. Mansi et al. Machinery,NewYork,NY,USA(2017).https://doi.org/10.1145/3018661.3022759, https://doi.org/10.1145/3018661.3022759

  26. [27]

    https://doi.org/10.1007/s00146-022-01454-7, https://doi.org/10.1007/s00146- 022-01454-7

    Fanni, R., Steinkogler, V.E., Zampedri, G., Pierson, J.: Enhancing hu- man agency through redress in artificial intelligence systems (2022). https://doi.org/10.1007/s00146-022-01454-7, https://doi.org/10.1007/s00146- 022-01454-7

  27. [28]

    https://doi.org/10.48550/ARXIV.1811.07819, https://arxiv.org/abs/1811.07819

    Ghosh, D., Gupta, A., Levine, S.: Learning actionable representations with goal-conditioned policies (2018). https://doi.org/10.48550/ARXIV.1811.07819, https://arxiv.org/abs/1811.07819

  28. [29]

    Human-Machine Communication2, 153–171 (01 2021)

    Gibbs, J., Kirkwood, G., Fang, C., Wilkenfeld, J.: Negotiating agency and control: Theorizing human-machine communication from a structura- tional perspective. Human-Machine Communication2, 153–171 (01 2021). https://doi.org/10.30658/hmc.2.8

  29. [30]

    Academy of Management An- nals14(2), 627–660 (2020)

    Glikson, E., Woolley, A.W.: Human trust in artificial intelli- gence: Review of empirical research. Academy of Management An- nals14(2), 627–660 (2020). https://doi.org/10.5465/annals.2018.0057, https://doi.org/10.5465/annals.2018.0057

  30. [31]

    ACM Computing Surveys (CSUR)51, 1 – 42 (2018)

    Guidotti, R., Monreale, A., Turini, F., Pedreschi, D., Giannotti, F.: A survey of methods for explaining black box models. ACM Computing Surveys (CSUR)51, 1 – 42 (2018)

  31. [32]

    Gunning, D., Aha, D.: DARPA’s explainable artifi- cial intelligence (XAI) program40(2), 44–58 (2019), https://ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/2850, sec- tion: Articles

  32. [33]

    AAAI Spring Symposium (2009)

    Harrell, D.F., Zhu, J.: Agency play: Dimensions of agency for interactive narrative design. AAAI Spring Symposium (2009)

  33. [34]

    https://doi.org/10.48550/ARXIV.1907.09615, https://arxiv.org/abs/1907.09615

    Joshi, S., Koyejo, O., Vijitbenjaronk, W., Kim, B., Ghosh, J.: Towards realistic individual recourse and actionable explanations in black-box de- cision making systems (2019). https://doi.org/10.48550/ARXIV.1907.09615, https://arxiv.org/abs/1907.09615

  34. [35]

    Jørnø, R.L., Gynther, K.: What constitutes an ‘actionable insight’ in learning analytics? vol. 5, pp. 198–221 (2018). https://doi.org/10.18608/jla.2018.53.13, https://learning-analytics.info/index.php/JLA/article/view/5897

  35. [36]

    Khosravi, H., Shum, S.B., Chen, G., Conati, C., Tsai, Y.S., Kay, J., Knight, S., Martinez-Maldonado, R., Sadiq, S., Gašević, D.: Explainable artificial intelligence in education. vol. 3, p. 100074 (2022). https://doi.org/https://doi.org/10.1016/j.caeai.2022.100074, https://www.sciencedirect.com/science/article/pii/S2666920X22000297

  36. [37]

    Langley, P.: Explainable, normative, and justified agency. vol. 33, pp. 9775–9779 (2019). https://doi.org/10.1609/aaai.v33i01.33019775, https://ojs.aaai.org/index.php/AAAI/article/view/5049, number: 01

  37. [38]

    Langley, P., Meadows, B., Sridharan, M., Choi, D.: Ex- plainable agency for intelligent autonomous systems. vol. 31, pp. 4762–4763 (2017). https://doi.org/10.1609/aaai.v31i2.19108, https://ojs.aaai.org/index.php/AAAI/article/view/19108, section: IAAI Chal- lenge Papers

  38. [39]

    Liao, Q.V., Gruen, D., Miller, S.: Questioning the ai: Informing design practices for explainable ai user experiences. p. 1–15. CHI ’20, Association for Computing Machinery,NewYork,NY,USA(2020).https://doi.org/10.1145/3313831.3376590, https://doi.org/10.1145/3313831.3376590 Evaluating Actionability in Explainable AI 23

  39. [40]

    In: Trattner, C., Parra, D., Riche, N

    Lim, B.Y., Yang, Q., Abdul, A.M., Wang, D.: Why these explanations? selecting intelligibility types for explanation goals. In: Trattner, C., Parra, D., Riche, N. (eds.) Joint Proceedings of the ACM IUI 2019 Workshops co-located with the 24th ACM Conference on Intelligent User Interfaces (ACM IUI 2019), Los Angeles, USA, March 20, 2019. CEUR Workshop Proce...

  40. [41]

    Communications of the ACM 61, 36 – 43 (2018)

    Lipton, Z.C.: The mythos of model interpretability. Communications of the ACM 61, 36 – 43 (2018)

  41. [42]

    List, C.: Group agency and artificial intelligence. vol. 34, pp. 1213–1242 (2021). https://doi.org/10.1007/s13347-021-00454-7, https://doi.org/10.1007/s13347-021- 00454-7

  42. [43]

    Journal of Computer-Mediated Com- munication26(6), 384–402 (09 2021)

    Liu, B.: In AI We Trust? Effects of Agency Locus and Transparency on Uncer- tainty Reduction in Human–AI Interaction. Journal of Computer-Mediated Com- munication26(6), 384–402 (09 2021). https://doi.org/10.1093/jcmc/zmab013, https://doi.org/10.1093/jcmc/zmab013

  43. [44]

    In: Nørskov, M., Seibt, J., Quick, O

    Longin, L.: Towards a middle-ground theory of agency for artificial intelligence. In: Nørskov, M., Seibt, J., Quick, O. (eds.) Culturally Sustainable Social Robotics: Proceedings of Robophilosophy 2020, pp. 17–26 (2020)

  44. [45]

    Extracting Actionability from Machine Learning Models by Sub-optimal Deterministic Planning

    Lyu, Q., Chen, Y., Li, Z., Cui, Z., Chen, L., Zhang, X., Shen, H.: Extracting action- abilityfrommachinelearningmodelsbysub-optimaldeterministicplanning(2016). https://doi.org/10.48550/ARXIV.1611.00873, https://arxiv.org/abs/1611.00873

  45. [46]

    Mohseni, S., Zarei, N., Ragan, E.D.: A multidisciplinary survey and framework for design and evaluation of explainable AI systems (2020), http://arxiv.org/abs/1811.11839

  46. [47]

    In: Proceed- ings of the 2020 Conference on Fairness, Accountability, and Trans- parency

    Mothilal, R.K., Sharma, A., Tan, C.: Explaining machine learning classifiers through diverse counterfactual explanations. In: Proceed- ings of the 2020 Conference on Fairness, Accountability, and Trans- parency. p. 607–617. FAT* ’20, Association for Computing Machinery, New York, NY, USA (2020). https://doi.org/10.1145/3351095.3372850, https://doi.org/10....

  47. [48]

    asser lecture - asser.nl (2021), https://www.asser.nl/asserpress/books/?rId=13986

    Murray, A.D.: Almost human: Law and human agency in the time of ar- tificial intelligence - sixth annual t.m.c. asser lecture - asser.nl (2021), https://www.asser.nl/asserpress/books/?rId=13986

  48. [49]

    Neff, G., Nagy, P.: Agency in the digital age: Using symbiotic agency to explain human–technology interaction (2018)

  49. [50]

    In: Lim, C.T., Leo, H.L., Yeow, R

    Ong, M.L., Li, A., Motani, M.: Explainable and actionable machine learning mod- els for electronic health record data. In: Lim, C.T., Leo, H.L., Yeow, R. (eds.) 17th International Conference on Biomedical Engineering. pp. 91–99. Springer Interna- tional Publishing

  50. [51]

    Computers and Education: Artificial Intelligence2, 100020 (2021)

    Ouyang, F., Jiao, P.: Artificial intelligence in education: The three paradigms. Computers and Education: Artificial Intelligence2, 100020 (2021)

  51. [52]

    American Medical Informatics Association Annual Symposium2012, 779–88 (2012)

    Roberts,K.,Rink,B.,Harabagiu,S.M.,Scheuermann,R.H.,Toomay,S.M.,Brown- ing, T.G., Bosler, T., Peshock, R.M.: A machine learning approach for identifying anatomical locations of actionable findings in radiology reports. American Medical Informatics Association Annual Symposium2012, 779–88 (2012)

  52. [53]

    British Journal of Educational Technology50(6), 2943–2958 (2019)

    Rosé, C.P., McLaughlin, E.A., Liu, R., Koedinger, K.R.: Explana- tory learner models: Why machine learning (alone) is not the an- swer. British Journal of Educational Technology50(6), 2943–2958 (2019). https://doi.org/https://doi.org/10.1111/bjet.12858, https://bera- journals.onlinelibrary.wiley.com/doi/abs/10.1111/bjet.12858 24 G. Mansi et al

  53. [54]

    https://doi.org/10.1007/s43681-022-00158-4, https://doi.org/10.1007/s43681- 022-00158-4

    Schönau, A.: Agency in augmented reality: exploring the ethics of facebook’s AI-powered predictive recommendation system (2022). https://doi.org/10.1007/s43681-022-00158-4, https://doi.org/10.1007/s43681- 022-00158-4

  54. [55]

    Sidorova, A., Rafiee, D.: AI Agency Risks and Their Mitigation Through Business Process Management: A Conceptual Framework (2019), http://hdl.handle.net/10125/60019

  55. [56]

    Silva, J.: Increasing perceived agency in human-ai interactions: Learnings from piloting a voice user interface with drivers on uber. vol. 2019, pp. 441– 456 (2019). https://doi.org/https://doi.org/10.1111/1559-8918.2019.01299, https://anthrosource.onlinelibrary.wiley.com/doi/abs/10.1111/1559- 8918.2019.01299

  56. [57]

    Directive Explanations for Actionable Explainability in Machine Learning Applications

    Singh, R., Dourish, P., Howe, P., Miller, T., Sonenberg, L., Velloso, E., Vetere, F.: Directive explanations for actionable explainability in ma- chine learning applications (2021). https://doi.org/10.48550/ARXIV.2102.02671, https://arxiv.org/abs/2102.02671

  57. [58]

    how child welfare workers reduce racial disparities in algorithmic decisions

    Stapleton, L., Cheng, H.F., Kawakami, A., Sivaraman, V., Cheng, Y., Qing, D., Perer, A., Holstein, K., Wu, Z.S., Zhu, H.: Extended analysis of “how child welfare workers reduce racial disparities in algorithmic decisions” (2022), https://arxiv.org/abs/2204.13872

  58. [59]

    https://doi.org/10.1186/s41239-021-00313-7, https://doi.org/10.1186/s41239-021- 00313-7

    Susnjak, T., Ramaswami, G.S., Mathrani, A.: Learning analytics dashboard: a tool for providing actionable insights to learners19(1), 12 (2022). https://doi.org/10.1186/s41239-021-00313-7, https://doi.org/10.1186/s41239-021- 00313-7

  59. [60]

    (eds.) The Mind-Technology Problem : Investigating Minds, Selves and 21st Century Artefacts, pp

    Swanepoel, D.: Does artificial intelligence have agency? In: Hipólito, I., Clowes, R.W., Gärtner, K. (eds.) The Mind-Technology Problem : Investigating Minds, Selves and 21st Century Artefacts, pp. 83–104. Springer Verlag (2021)

  60. [61]

    Defining and Conceptualizing Actionable Insight: A Conceptual Framework for Decision-centric Analytics

    Tan, S.Y., Chan, T.: Defining and conceptualizing actionable in- sight: A conceptual framework for decision-centric analytics (2016). https://doi.org/10.48550/ARXIV.1606.03510, https://arxiv.org/abs/1606.03510

  61. [62]

    Vanneste, B., Puranam, P.: Artificial intelligence, trust, and percep- tions of agency. No. 3897704 (2023). https://doi.org/10.2139/ssrn.3897704, https://papers.ssrn.com/abstract=3897704

  62. [63]

    In: Proceedings of the 2019 CHI Conference on Human Fac- tors in Computing Systems

    Wang, D., Yang, Q., Abdul, A., Lim, B.Y.: Designing theory-driven user-centric explainable ai. In: Proceedings of the 2019 CHI Conference on Human Fac- tors in Computing Systems. p. 1–15. CHI ’19, Association for Computing Ma- chinery, New York, NY, USA (2019). https://doi.org/10.1145/3290605.3300831, https://doi.org/10.1145/3290605.3300831

  63. [64]

    In: ICCBR Workshops (2021)

    Wiratunga, N., Wijekoon, A., Nkisi-Orji, I., Martin, K., Palihawadana, C., Corsar, D.: Actionable feature discovery in counterfactuals using feature relevance explain- ers. In: ICCBR Workshops (2021)

  64. [65]

    Wolf, C., Blomberg, J.: Evaluating the promise of human-algorithm col- laborations in everyday work practices. vol. 3. Association for Comput- ing Machinery, New York, NY, USA (2019). https://doi.org/10.1145/3359245, https://doi.org/10.1145/3359245

  65. [66]

    In: Proceedings of the 2019 CHI Conference on Human Fac- tors in Computing Systems

    Yang, Q., Steinfeld, A., Zimmerman, J.: Unremarkable AI: Fitting in- telligent decision support into critical, clinical decision-making pro- cesses. In: Proceedings of the 2019 CHI Conference on Human Fac- tors in Computing Systems. pp. 1–11. CHI ’19, Association for Com- puting Machinery (2019). https://doi.org/10.1145/3290605.3300468, https://doi.org/10...

  66. [67]

    Explanatory Pluralism in Explainable AI

    Yao, Y.: Explanatory pluralism in explainable ai (2021). https://doi.org/10.48550/ARXIV.2106.13976, https://arxiv.org/abs/2106.13976 26 G. Mansi et al. A Additional User-Based Terminology As discussed in Section 3.3, we developed user-based terminology and definitions, which we used in coding our data and throughout the rest of the paper. Here we provide ...