Pith. sign in

REVIEW 3 major objections 5 minor 21 references

Do ATCOs Need Explanations, and Why? Towards ATCO-Centered Explainable AI for Conflict Resolution Advisories

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Controllers want AI explanations mainly for reports and record-keeping, not when they already agree with the advisory.

desk verdict A genuinely user-first qualitative study of why ATCOs want explanations, but the headline conclusion ignores the same participants' ranking data, which puts the two featured goals at the bottom. read the letter →

arxiv 2505.03117 v2 pith:D25BCAQP submitted 2025-05-06 cs.HC

classification cs.HC
keywords AirtrafficcontrolExplainableAIConflictresolutionHuman-AIinteractionExplanationneedsGoalelicitationTrustcalibrationUser-centereddesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks a question that ATM explainability research has largely skipped: do air traffic controllers actually need explanations from AI conflict-resolution advisories, and why? Through interviews, goal exploration, and ranking exercises with eight licensed controllers, it finds that the need for explanations is goal-dependent. All participants wanted explanations for documenting decisions and generating reports, and most wanted them for communicating with supervisors, but almost none wanted explanations when the AI's advisory matched their own assessment. The paper argues that XAI design should start from these user goals rather than from what systems can visualize. This preliminary qualitative finding reorients ATM explainability research from system-generated explanations toward user-centered needs.

What carries the argument

The central instrument is a set of 11 explanation goals adapted from an end-user-centered explainability framework originally developed for non-ATM applications. Each goal is a concrete operational situation (e.g., calibrate trust, ensure safety, detect bias, generate reports) in which a controller might interact with an AI advisory. Participants were asked, for each goal, whether they would accept the AI as decision support, whether they would need an explanation, and what kind of explanation they would want. Controllers then sorted and ranked the goals by perceived importance. This goal-by-goal mapping is what carries the argument, because it converts the vague question 'do controllers need explanations?' into specific, testable needs tied to particular operational tasks.

What would settle it

A high-fidelity simulation in which controllers work live conflicts with a real advisory tool and can request explanations freely: if controllers frequently request explanations when they already agree with the advisory, or rarely request them while writing post-event reports, then the paper's goal-based need mapping does not predict real behavior.

Watch

Extended reading notes

Core claim

The paper's central claim is that ATCOs do need AI explanations, but only for certain operational purposes. In a hypothetical peak-traffic conflict scenario, all eight participants reported needing explanations for post-event documentation and report generation (goal G10), and seven of eight needed them for communicating decisions to supervisors (G9). Six needed explanations to learn from the AI or to understand why two seemingly similar conflicts received different advisories. By contrast, when their own assessment aligned with the AI's advisory (G5), all but one participant said they did not need an explanation. The authors interpret this as evidence that explanations serve hybrid roles—building trust, enabling collaboration, and supporting coevolution—and that explanation delivery should be dynamically adjusted to the controller's goal and situation.

Load-bearing premise

The 11 explanation goals were borrowed from a non-aviation user study and applied to a hypothetical conflict scenario, so the results depend on this goal list and scenario capturing the goals controllers would actually act on in live operations.

Editorial extensions

If this is right

  • Explainable AI for air traffic control should prioritize post-hoc documentation and supervisor communication over real-time decision support, since those are the goals with the most consistent expressed need.
  • When a controller's assessment aligns with the AI's advisory, explanations can be omitted or made available on demand, reducing workload and display clutter.
  • Explanation systems should adapt their content and timing—training, live operation, or post-operation—to the controller's current goal, rather than presenting a static explanation for every advisory.
  • Advisory evaluation should measure whether explanations improve the quality and efficiency of reports and stakeholder communication, not only whether they increase trust or understandability.
  • Future XAI designs could embed an explicit 'why' request mechanism triggered by disagreement or safety concerns, aligning system behavior with the negotiated needs controllers expressed.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • We infer that the strongest functional niche for ATM explainability is accountability: explanations are wanted where the controller must later justify an action, not where the action itself is under time pressure.
  • A testable extension is that in live operations, explanation request rates will track documentation and handover duties more closely than real-time decision confidence; log analysis of actual advisory tools could verify this.
  • The goal-mapping method could transfer to other safety-critical professions with comparable accountability structures, such as airspace coordination or emergency dispatch, where 'why' questions are similarly tied to reporting.
  • We predict that a dynamic explanation policy—full explanations for reports, brief or absent explanations on agreement—would reduce perceived workload while preserving operator trust, a claim the paper's qualitative data supports but does not yet demonstrate.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a qualitative user study with eight licensed air traffic controllers (ATCOs) to ask whether and why ATCOs need explanations of AI conflict-resolution advisories. Using Jin et al.'s eleven explanation goals, the authors conducted semi-structured interviews, asked per-goal yes/no questions about the need for explanations, and had participants rank the goals by importance. They report that explanation need is highest for report generation and stakeholder communication, lowest when the controller already agrees with the advisory, and they discuss implications for the timing, content, and format of XAI. The paper is framed as a first step toward ATCO-centered XAI, reversing the usual top-down design flow.

Significance. If the findings hold, the study fills a real gap by moving from researcher-intuition-driven XAI prototypes to direct elicitation of controller needs in a safety-critical domain. The paper's strengths are its clear reversal of the usual design flow, the adaptation of an existing end-user-centered framework to ATM, the recruitment of licensed controllers rather than students, and the concrete recommendations for on-demand, post-operation, and interactive explanations. The main evidence is descriptive and self-reported, and several methodological and interpretational issues remain, but the study generates testable design hypotheses for future work in ATM XAI.

major comments (3)
  1. [§V.F, §V.D, §VII] The paper's headline claim that 'all ATCO participants needed explanations to document decisions and rationales for future reference or report generation' (Section VII) is presented without reconciling the ranking data in Section V.F, where G10 (median rank 9.25) and G9 (median rank 9.50) are the least important of the eleven goals, while G2 'Ensure Safety' has median rank 1.50. The Friedman test in Section V.F is not significant (p=.098), but the descriptive inversion is stark. The Discussion in Section VI.D.1 calls the G9/G10 demand 'surprisingly high' without ever mentioning that the same participants ranked these goals as least important for their work. Because the conclusion directs XAI design toward documentation and report-generation features, the authors should explicitly distinguish between 'needed in a specific situation' and 'important relative to other goals,' and either reconcile the divergence or substantially soften the central claim.
  2. [§IV.B and §V.D] The paper reports counts of 'need explanation' responses and numerous quotations from open-ended questions, but it does not describe a qualitative coding protocol, codebook, or inter-rater reliability check. For claims such as 'explanations were seen as vital for multiple reasons' (Section V.D), it is unclear how themes were extracted from the audio- and screen-recorded sessions. Please add an analysis section describing how transcripts were processed, how themes were derived, whether any coding agreement measure was used, and how many researchers were involved; alternatively, explicitly label the thematic content as illustrative quotations rather than coded findings.
  3. [§IV.B.2, G5 and §V.D] The G5 scenario is worded as 'You agree with AI advisory as it matches your assessment of the situation.' Asking immediately afterward whether the participant needs an explanation is close to tautological, since the scenario already stipulates agreement. The Section V.D finding that all but one participant said they did not need explanations for G5 is therefore not an independent empirical result. The scenario should be reframed so that agreement is not embedded in the premise—for example, by describing a situation in which the controller's assessment happens to match the advisory without telling the participant they agree—or the claim about G5 should be presented only as a manipulation check and appropriately hedged.
minor comments (5)
  1. [§VI.C] The sentence 'A static, one-size-fits-all explanation fail to capture...' should read 'fails to capture.'
  2. [Abstract] The phrase 'their conflict resolution approach align with the artificial intelligence (AI) advisory' is missing a verb ending; it should be 'aligns.'
  3. [§V.F / Figure 7] The box plot would be easier to read if the figure caption indicated that lower values mean higher importance and if the Friedman statistics and N were printed in the caption or text next to the median values.
  4. [§IV.B.2] The supplementary link (https://tinyurl.com/pk6yxvcx) is convenient but temporary; please consider an archival supplement or a DOI for the supplementary materials.
  5. [§III, RQ1–RQ4] The paper presents RQ1–RQ4 as a framing device but only RQ1 and parts of RQ2 are addressed. A sentence in the conclusion noting which questions remain open would help calibrate reader expectations.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the findings are direct empirical summaries of elicited ATCO responses, not derived or fitted quantities.

full rationale

The paper makes no formal derivation or fitted-parameter prediction; its central claims, such as 'all ATCO participants needed explanations to document decisions and rationales for future reference or report generation' and 'explanations were deemed less necessary when there was an alignment between ATCO assessments and AI advisory,' are descriptive counts and thematic summaries of semi-structured interviews and sorting exercises with eight licensed ATCOs. The 11 goals were imported from the external EUCA framework (Jin et al. [5]), not defined by the outcomes, and the authors explicitly retained all goals rather than selecting them post hoc. The G5 alignment result is not forced by construction: the scenario stated that the participant agrees with the advisory, but participants still could have requested explanations, and one did, so the finding is an empirical response pattern rather than a tautology. The ranking result (G9 and G10 least important) actually creates an unresolved inconsistency with the binary 'need explanation' responses; that is a validity or interpretation concern, not circular reasoning. Self-citations (e.g., [3] and [12]) appear only as background framing and are not load-bearing for the reported results. The paper is self-contained against its own interview data, and the conclusion explicitly labels the study preliminary and subjective, further supporting that no derivation is being presented as independent verification.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No equations or fitted parameters are present. The analysis rests on the direct transfer of a non-ATM goal framework to air traffic control, the validity of self-report as a proxy for operational needs, and the representativeness of an eight-person sample.

assumptions (3)
  • domain assumption The 11 explanation goals imported from Jin et al.'s non-ATM EUCA framework are applicable to ATM conflict resolution tasks.
    Stated in Section IV.B.2: the goals were adapted from non-ATM use cases and all were retained to validate their applicability in ATM through direct ATCO feedback.
  • domain assumption ATCOs' self-reported needs in a hypothetical scenario are reliable proxies for their needs in live operations.
    The study used a hypothetical peak-traffic conflict scenario and asked participants to reason about their needs; no live or simulated operations were run, so the link to real operational behavior is assumed.
  • domain assumption The small sample of eight licensed ATCOs is sufficiently representative to identify goals shared across controller populations.
    The authors acknowledge the small sample limits statistical power in Section V.F, yet they use all-participant counts as evidence for the main findings.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do ATCOs Need Explanations, and Why? Towards ATCO-Centered Explainable AI for Conflict Resolution Advisories." pith.science (2026). https://pith.science/paper/D25BCAQP

@misc{pith2026250503117,
  author       = {Pith},
  title        = {Pith review of: Do ATCOs Need Explanations, and Why? Towards ATCO-Centered Explainable AI for Conflict Resolution Advisories},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/D25BCAQP}},
  note         = {Machine review of arXiv:2505.03117}
}
read the original abstract

Interest in explainable artificial intelligence (XAI) is surging. Prior research has primarily focused on systems' ability to generate explanations, often guided by researchers' intuitions rather than end-users' needs. Unfortunately, such approaches have not yielded favorable outcomes when compared to a black-box baseline (i.e., no explanation). To address this gap, this paper advocates a human-centered approach that shifts focus to air traffic controllers (ATCOs) by asking a fundamental yet overlooked question: Do ATCOs need explanations, and if so, why? Insights from air traffic management (ATM), human-computer interaction, and the social sciences were synthesized to provide a holistic understanding of XAI challenges and opportunities in ATM. Evaluating 11 ATM operational goals revealed a clear need for explanations when ATCOs aim to document decisions and rationales for future reference or report generation. Conversely, ATCOs are less likely to seek them when their conflict resolution approach align with the artificial intelligence (AI) advisory. While this is a preliminary study, the findings are expected to inspire broader and deeper inquiries into the design of ATCO-centric XAI systems, paving the way for more effective human-AI interaction in ATM.

Figures

Figures reproduced from arXiv: 2505.03117 by the authors.

Figure 1
Figure 1. Explanation as an enabler in ATCO-AI interaction [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. XAI research landscape in ATM. This study focuses [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. A virtual board used in the study, consisting of three activities: (a) semi-structured interviews, (b) goals exploration, [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Distribution of opinions on incorporating AI tech [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 6
Figure 6. Figure 6: Response counts for each goal when asked if they [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Rank distribution of goals based on perceived [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

21 extracted references · 17 canonical work pages

  1. [1]

    EASA Artificial Intelligence Concept Paper Issue 2: Guidance for Level 1 & 2 Machine Learning Applications,

    EASA, “EASA Artificial Intelligence Concept Paper Issue 2: Guidance for Level 1 & 2 Machine Learning Applications,” European Union Aviation Safety Agency (EASA), Tech. Rep., Mar. 2024. [Online]. Available: https://www.easa.europa.eu/en/document-library/ general-publications/easa-artificial-intelligence-concept-paper-issue-2

  2. [2]

    Human factors requirements for human-ai teaming in aviation,

    B. Kirwan, “Human factors requirements for human-ai teaming in aviation,”Future Transportation, vol. 5, no. 2, 2025. [Online]. Available: https://www.mdpi.com/2673-7590/5/2/42

  3. [3]

    Human-ai hybrid paradigm for collaborative air traffic management systems,

    D.-T. Pham, H. Ali, K. Fennedy, M.-H. Hsieh, S. Alam, and V . Duong, “Human-ai hybrid paradigm for collaborative air traffic management systems,” inSESAR Innovation Days, 2024

  4. [4]

    Explaining the unexplainable: Role of xai for flight take-off time delay prediction,

    W. Jmoona, M. U. Ahmed, M. R. Islam, S. Barua, S. Begum, A. Fer- reira, and N. Cavagnetto, “Explaining the unexplainable: Role of xai for flight take-off time delay prediction,” inIFIP International Conference on Artificial Intelligence Applications and Innovations. Springer, 2023, pp. 81–93

  5. [5]

    Euca: The end-user-centered explainable ai framework,

    W. Jin, J. Fan, D. Gromala, P. Pasquier, and G. Hamarneh, “Euca: The end-user-centered explainable ai framework,”arXiv preprint arXiv:2102.02437, 2021

  6. [6]

    Designing theory- driven user-centric explainable ai,

    D. Wang, Q. Yang, A. Abdul, and B. Y . Lim, “Designing theory- driven user-centric explainable ai,” inProceedings of the 2019 CHI Conference on Human Factors in Computing Systems, ser. CHI ’19. New York, NY , USA: Association for Computing Machinery, 2019, p. 1–15. [Online]. Available: https://doi.org/10.1145/3290605.3300831

  7. [7]

    A survey on artificial intelligence (ai) and explainable ai in air traffic management: Current trends and development with future research trajectory,

    A. Degas, M. R. Islam, C. Hurter, S. Barua, H. Rahman, M. Poudel, D. Ruscio, M. U. Ahmed, S. Begum, M. A. Rahman, S. Bonelli, G. Cartocci, G. Di Flumeri, G. Borghini, F. Babiloni, and P. Aric ´o, “A survey on artificial intelligence (ai) and explainable ai in air traffic management: Current trends and development with future research trajectory,”Applied S...

  8. [8]

    Explanation in artificial intelligence: Insights from the social sciences,

    T. Miller, “Explanation in artificial intelligence: Insights from the social sciences,”Artificial intelligence, vol. 267, pp. 1–38, 2019

Show all 21 references
  1. [9]

    An explainable artificial intelligence (xai) framework for improving trust in automated atm tools,

    C. S. Hernandez, S. Ayo, and D. Panagiotakopoulos, “An explainable artificial intelligence (xai) framework for improving trust in automated atm tools,” in2021 IEEE/AIAA 40th Digital Avionics Systems Confer- ence (DASC). IEEE, 2021, pp. 1–10

  2. [10]

    Explainable artificial intelligence (xai): Concepts, taxonomies, op- portunities and challenges toward responsible ai,

    A. B. Arrieta, N. D ´ıaz-Rodr´ıguez, J. Del Ser, A. Bennetot, S. Tabik, A. Barbado, S. Garc ´ıa, S. Gil-L ´opez, D. Molina, R. Benjaminset al., “Explainable artificial intelligence (xai): Concepts, taxonomies, op- portunities and challenges toward responsible ai,”Information f...

  3. [11]

    Explanation of machine-learning solutions in air-traffic management,

    Y . Xie, N. Pongsakornsathien, A. Gardi, and R. Sabatini, “Explanation of machine-learning solutions in air-traffic management,”Aerospace, vol. 8, no. 8, p. 224, 2021

  4. [12]

    A multi-modal approach to measuring the effect of xai on air traffic controller trust during off-nominal runway exits,

    K. Pushparaj, P. Reddy, D. Vu-Tran, K. Izzetoglu, and S. Alam, “A multi-modal approach to measuring the effect of xai on air traffic controller trust during off-nominal runway exits,” in2023 IEEE Inter- national Conference on Systems, Man, and Cybernetics (SMC). IEEE, 2023, pp...

  5. [13]

    Transparency & explainability in higher levels of automation in the atm domain,

    N. Valle, M. F. Lema, J. M. Cordero, E. Iglesias, R. Rodr ´ıguez, G. Andrienko, N. Andrienko, G. A. V ouros, T. Kravaris, G. Papadopouloset al., “Transparency & explainability in higher levels of automation in the atm domain,” inSESAR Innovation Days 2022,

  6. [14]

    D3.1 use cases transparency requirements,

    TAPAS, “D3.1 use cases transparency requirements,” CRIDA, Tech. Rep., 2021

  7. [15]

    Usage of more transparent and explainable conflict resolution algorithm: air traffic controller feedback,

    C. Hurter, A. Degas, A. Guibert, N. Durand, A. Ferreira, N. Cavagnetto, M. R. Islam, S. Barua, M. U. Ahmed, S. Begumet al., “Usage of more transparent and explainable conflict resolution algorithm: air traffic controller feedback,”Transportation research procedia, vol. 66, pp....

  8. [16]

    D6.2 use cases transparency requirements,

    MAHALO, “D6.2 use cases transparency requirements,” Deep Blue, Tech. Rep., 2022

  9. [17]

    Trust in automation: Designing for appro- priate reliance,

    J. D. Lee and K. A. See, “Trust in automation: Designing for appro- priate reliance,”Human factors, vol. 46, no. 1, pp. 50–80, 2004

  10. [18]

    Direct measurement of situation awareness: Validity and use of sagat,

    M. R. Endsley, “Direct measurement of situation awareness: Validity and use of sagat,” inSituational awareness. Routledge, 2017, pp. 129–156

  11. [19]

    The situation awareness framework for explainable ai (safe-ai) and human factors considerations for xai sys- tems,

    L. Sanneman and J. A. Shah, “The situation awareness framework for explainable ai (safe-ai) and human factors considerations for xai sys- tems,”International Journal of Human–Computer Interaction, vol. 38, no. 18-20, pp. 1772–1788, 2022

  12. [20]

    Chatatc: Large language model-driven conversational agents for supporting strategic air traffic flow management,

    S. Abdulhak, W. Hubbard, K. Gopalakrishnan, and M. Z. Li, “Chatatc: Large language model-driven conversational agents for supporting strategic air traffic flow management,”arXiv preprint arXiv:2402.14850, 2024

  13. [2022]

    Available: https://www.sesarju.eu/sesarinnovationdays

    [Online]. Available: https://www.sesarju.eu/sesarinnovationdays

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.