{"id":"b5cf639e-8cb6-490f-9462-ed12c9041cbd","arxiv_id":"2505.03117","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Air traffic controllers in this interview study wanted AI explanations mainly for documentation and stakeholder communication, and rarely when they already agreed with the AI's advice.","lead":"This paper asked eight air traffic controllers whether they actually want explanations from AI conflict-resolution tools and for which tasks. The answer: they want explanations mainly for writing reports and justifying decisions to supervisors, not when they already agree with the AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Conclusion highlights G9/G10 explanation needs, but ranking data show these goals as least important; the paper never reconciles this internal contradiction.","rationale":"The reader's verdict is CONDITIONAL, and I agree with that overall assessment, but my primary concern is different from the reader's weakest assumption. The reader worried about external validity of the imported goal list and the hypothetical scenario. A more immediate, internal inconsistency threatens the central claim as written. The paper's headline finding in the Conclusion is based on the count of 'need explanation' responses for G9/G10, yet the same participants, in the ranking task reported in Section V.F, assigned these goals the lowest median importance (9.50 and 9.25 out of 11, with lower = more important). This is not a subtle tension; G2 (Ensure Safety) is at 1.50. The paper does not explain how ATCOs can both 'need' explanations for a goal and simultaneously rank that goal as the least critical. The authors do acknowledge the ranking result and the non-significant Friedman test, but they do not connect it to the earlier need counts. As a result, the Conclusion overstates the practical relevance of the G9/G10 finding. This does not invalidate the study—the descriptive claim about what participants said is accurate—but it does mean the conclusion should be qualified. The concrete test I propose would force a reconciliation. The G5 alignment finding is a lesser concern because its prompt nearly guarantees the result; however, it is not the main issue. Given that the reader already recommended conditional acceptance pending additional evidence, my analysis reinforces that condition without moving to accept or reject. Hence, verdict_should_be is UNCHANGED, and agreement_with_reader is partial: the reader identified ranking significance issues in their rationale but did not focus on the direct contradiction with the central claim.","tokens_in":14241,"tokens_out":4403,"duration_ms":38845,"concrete_test":"Request from the authors a participant-level cross-tabulation for G9 and G10: for each of the 8 ATCOs, pair the binary 'need explanation' answer (Fig. 6/ Section V.D) with the rank assigned in the sorting exercise (Fig. 7/ Section V.F). A sign test or Wilcoxon signed-rank test on the paired values, or even a simple count of how many of the 8 participants who said 'yes' to G10 also ranked G10 in the bottom half of the 11 ranks, would settle whether the Conclusion's emphasis is justified. If the majority of those who expressed a need also ranked the goal as least important, the Conclusion must be reframed to say 'needed for low-priority administrative documentation,' which substantially weakens the design implication.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, restated in the Conclusion, is that 'all ATCO participants needed explanations to document decisions and rationales for future reference or report generation' (G10) and, with one exception, for communicating with stakeholders (G9). This claim rests on the binary 'need explanation' responses in Section V.D and Figure 6. However, Section V.F and Figure 7 present the same participants' ranking of these goals by importance: G10 has a median rank of 9.25 and G9 a median of 9.50, the lowest of all 11 goals, while G2 (Ensure Safety) has a median of 1.50. The Friedman test is not significant (p=.098), but the descriptive inversion is stark. The paper does not reconcile these two results. In Discussion VI.D.1, the authors call the G9/G10 demand 'surprisingly high' and suggest future work on post-operation XAI, but they never note that the same ATCOs ranked these goals as least important for their work. This matters because the Conclusion's emphasis could misdirect XAI design: a 'need' that is expressed in an abstract yes/no question but ranked as unimportant when forced to prioritize should not be presented as the primary motivation for explanation features. The finding is not false, but it is misleading without the ranking context. A secondary issue: the G5 alignment result ('explanations less necessary when aligned') is close to tautological, since the G5 scenario explicitly states 'You agree with AI advisory as it matches your assessment of the situation'; not needing an explanation under those conditions is almost entailed by the prompt.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative user study with eight licensed air traffic controllers (ATCOs) to ask whether and why ATCOs need explanations of AI conflict-resolution advisories. Using Jin et al.'s eleven explanation goals, the authors conducted semi-structured interviews, asked per-goal yes/no questions about the need for explanations, and had participants rank the goals by importance. They report that explanation need is highest for report generation and stakeholder communication, lowest when the controller already agrees with the advisory, and they discuss implications for the timing, content, and format of XAI. The paper is framed as a first step toward ATCO-centered XAI, reversing the usual top-down design flow.","tokens_in":14515,"tokens_out":4081,"duration_ms":39079,"significance":"If the findings hold, the study fills a real gap by moving from researcher-intuition-driven XAI prototypes to direct elicitation of controller needs in a safety-critical domain. The paper's strengths are its clear reversal of the usual design flow, the adaptation of an existing end-user-centered framework to ATM, the recruitment of licensed controllers rather than students, and the concrete recommendations for on-demand, post-operation, and interactive explanations. The main evidence is descriptive and self-reported, and several methodological and interpretational issues remain, but the study generates testable design hypotheses for future work in ATM XAI.","major_comments":[{"comment":"The paper's headline claim that 'all ATCO participants needed explanations to document decisions and rationales for future reference or report generation' (Section VII) is presented without reconciling the ranking data in Section V.F, where G10 (median rank 9.25) and G9 (median rank 9.50) are the least important of the eleven goals, while G2 'Ensure Safety' has median rank 1.50. The Friedman test in Section V.F is not significant (p=.098), but the descriptive inversion is stark. The Discussion in Section VI.D.1 calls the G9/G10 demand 'surprisingly high' without ever mentioning that the same participants ranked these goals as least important for their work. Because the conclusion directs XAI design toward documentation and report-generation features, the authors should explicitly distinguish between 'needed in a specific situation' and 'important relative to other goals,' and either reconcile the divergence or substantially soften the central claim.","section":"§V.F, §V.D, §VII"},{"comment":"The paper reports counts of 'need explanation' responses and numerous quotations from open-ended questions, but it does not describe a qualitative coding protocol, codebook, or inter-rater reliability check. For claims such as 'explanations were seen as vital for multiple reasons' (Section V.D), it is unclear how themes were extracted from the audio- and screen-recorded sessions. Please add an analysis section describing how transcripts were processed, how themes were derived, whether any coding agreement measure was used, and how many researchers were involved; alternatively, explicitly label the thematic content as illustrative quotations rather than coded findings.","section":"§IV.B and §V.D"},{"comment":"The G5 scenario is worded as 'You agree with AI advisory as it matches your assessment of the situation.' Asking immediately afterward whether the participant needs an explanation is close to tautological, since the scenario already stipulates agreement. The Section V.D finding that all but one participant said they did not need explanations for G5 is therefore not an independent empirical result. The scenario should be reframed so that agreement is not embedded in the premise—for example, by describing a situation in which the controller's assessment happens to match the advisory without telling the participant they agree—or the claim about G5 should be presented only as a manipulation check and appropriately hedged.","section":"§IV.B.2, G5 and §V.D"}],"minor_comments":[{"comment":"The sentence 'A static, one-size-fits-all explanation fail to capture...' should read 'fails to capture.'","section":"§VI.C"},{"comment":"The phrase 'their conflict resolution approach align with the artificial intelligence (AI) advisory' is missing a verb ending; it should be 'aligns.'","section":"Abstract"},{"comment":"The box plot would be easier to read if the figure caption indicated that lower values mean higher importance and if the Friedman statistics and N were printed in the caption or text next to the median values.","section":"§V.F / Figure 7"},{"comment":"The supplementary link (https://tinyurl.com/pk6yxvcx) is convenient but temporary; please consider an archival supplement or a DOI for the supplementary materials.","section":"§IV.B.2"},{"comment":"The paper presents RQ1–RQ4 as a framing device but only RQ1 and parts of RQ2 are addressed. A sentence in the conclusion noting which questions remain open would help calibrate reader expectations.","section":"§III, RQ1–RQ4"}],"recommendation":"major_revision","confidential_remarks":"The study is clearly positioned as a preliminary, qualitative contribution, and the topic is well suited to the ATM XAI community. The main risk is that the abstract and conclusion overstate a finding that the paper's own ranking data undercut. If the authors can articulate when a 'need' exists despite low relative importance, and add credible methodological detail about the qualitative analysis, the paper could be acceptable. The current version needs this reconciliation before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the question. Prior ATM XAI work starts from a prototype and asks whether controllers like it; this paper starts from Jin et al.'s end-user-centered framework and asks whether, and for what purpose, ATCOs need explanations at all. The binary response counts are clear enough on their own terms: all eight participants said they would need explanations for report generation, and seven of eight for stakeholder communication. That is a real, useful datum for anyone designing XAI in this space, and the training-interview material adds texture that typical system-centered evaluations never capture.\n\nWhat the paper does well is the method's spirit: it treats controllers as the source of requirements, not as validation subjects. The goal cards are a reasonable first cut at operationalizing \"why\", and the authors are explicit that this is a preliminary study with a small sample.\n\nNow the soft spots, in proportion. The biggest one is the internal contradiction the authors never address. The same eight participants, when asked to rank goals by importance, put G10 (reports) and G9 (stakeholder communication) at the bottom, with medians around 9 out of 11. The paper calls the G9/G10 demand \"surprisingly high\" and builds the conclusion around it, but the ranking data say these goals are perceived as peripheral to live operational work. That is not a fatal flaw — need and importance are different constructs, and a need can be real yet low-priority — but the paper has to say that. Right now it reads as cherry-picking the binary counts and ignoring the priorities. The Friedman test not reaching significance (p = .098) makes the ranking interpretations descriptive, so the authors should be careful not to overclaim from either direction.\n\nSecond, the G5 result is close to tautological. The scenario text says \"You agree with AI advisory as it matches your assessment of the situation,\" so not needing an explanation under that condition is almost entailed by the prompt. The paper treats it as an empirical finding about alignment; it is partly an artifact of the stimulus.\n\nThird, there is no qualitative coding protocol, no inter-rater reliability check, and no explanation of how themes were extracted from the interviews. For a workshop paper that may be acceptable; for a journal it is not.\n\nWho gets value from this? Researchers working on ATM XAI design, especially those who want an empirical anchor for \"why explain\" before picking a visualization. It deserves serious peer review because the question is right and the data, despite its limits, is a legitimate first pass. But the conclusion needs rewriting to reconcile the binary need counts with the ranking data, and the method section needs to be honest about the absence of formal qualitative analysis. I would send it to review, with the expectation of a revision rather than acceptance as-is.","headline":"A genuinely user-first qualitative study of why ATCOs want explanations, but the headline conclusion ignores the same participants' ranking data, which puts the two featured goals at the bottom.","tokens_in":15050,"tokens_out":1368,"would_cite":true,"duration_ms":16912,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Controllers want AI explanations mainly for reports and record-keeping, not when they already agree with the advisory.","keywords":["Air traffic control","Explainable AI","Conflict resolution","Human-AI interaction","Explanation needs","Goal elicitation","Trust calibration","User-centered design"],"falsifier":"A high-fidelity simulation in which controllers work live conflicts with a real advisory tool and can request explanations freely: if controllers frequently request explanations when they already agree with the advisory, or rarely request them while writing post-event reports, then the paper's goal-based need mapping does not predict real behavior.","tokens_in":14071,"feed_emoji":"✈️","tokens_out":4108,"duration_ms":43553,"temperature":0.7,"pith_summary":"The paper asks a question that ATM explainability research has largely skipped: do air traffic controllers actually need explanations from AI conflict-resolution advisories, and why? Through interviews, goal exploration, and ranking exercises with eight licensed controllers, it finds that the need for explanations is goal-dependent. All participants wanted explanations for documenting decisions and generating reports, and most wanted them for communicating with supervisors, but almost none wanted explanations when the AI's advisory matched their own assessment. The paper argues that XAI design should start from these user goals rather than from what systems can visualize. This preliminary qualitative finding reorients ATM explainability research from system-generated explanations toward user-centered needs.","feed_headline":"Controllers need AI explanations for reports, not when they agree","feed_subtitle":"Eight controllers ranked 11 explanation goals, making explainable AI a documentation tool more than a decision aid.","key_machinery":"The central instrument is a set of 11 explanation goals adapted from an end-user-centered explainability framework originally developed for non-ATM applications. Each goal is a concrete operational situation (e.g., calibrate trust, ensure safety, detect bias, generate reports) in which a controller might interact with an AI advisory. Participants were asked, for each goal, whether they would accept the AI as decision support, whether they would need an explanation, and what kind of explanation they would want. Controllers then sorted and ranked the goals by perceived importance. This goal-by-goal mapping is what carries the argument, because it converts the vague question 'do controllers need explanations?' into specific, testable needs tied to particular operational tasks.","core_discovery":"The paper's central claim is that ATCOs do need AI explanations, but only for certain operational purposes. In a hypothetical peak-traffic conflict scenario, all eight participants reported needing explanations for post-event documentation and report generation (goal G10), and seven of eight needed them for communicating decisions to supervisors (G9). Six needed explanations to learn from the AI or to understand why two seemingly similar conflicts received different advisories. By contrast, when their own assessment aligned with the AI's advisory (G5), all but one participant said they did not need an explanation. The authors interpret this as evidence that explanations serve hybrid roles—building trust, enabling collaboration, and supporting coevolution—and that explanation delivery should be dynamically adjusted to the controller's goal and situation.","pith_inferences":["We infer that the strongest functional niche for ATM explainability is accountability: explanations are wanted where the controller must later justify an action, not where the action itself is under time pressure.","A testable extension is that in live operations, explanation request rates will track documentation and handover duties more closely than real-time decision confidence; log analysis of actual advisory tools could verify this.","The goal-mapping method could transfer to other safety-critical professions with comparable accountability structures, such as airspace coordination or emergency dispatch, where 'why' questions are similarly tied to reporting.","We predict that a dynamic explanation policy—full explanations for reports, brief or absent explanations on agreement—would reduce perceived workload while preserving operator trust, a claim the paper's qualitative data supports but does not yet demonstrate."],"forward_implications":["Explainable AI for air traffic control should prioritize post-hoc documentation and supervisor communication over real-time decision support, since those are the goals with the most consistent expressed need.","When a controller's assessment aligns with the AI's advisory, explanations can be omitted or made available on demand, reducing workload and display clutter.","Explanation systems should adapt their content and timing—training, live operation, or post-operation—to the controller's current goal, rather than presenting a static explanation for every advisory.","Advisory evaluation should measure whether explanations improve the quality and efficiency of reports and stakeholder communication, not only whether they increase trust or understandability.","Future XAI designs could embed an explicit 'why' request mechanism triggered by disagreement or safety concerns, aligning system behavior with the negotiated needs controllers expressed."],"supporting_citations":[{"why":"Supplies the 11 explanation goals and the end-user-centered framework that the study adapts to air traffic management.","marker":"[5]"},{"why":"Provides the theory-driven argument that XAI design should begin with users' reasoning goals rather than with system capabilities.","marker":"[6]"},{"why":"Documents the field-wide gap in ATM XAI, where research centers on what systems can do rather than on user needs.","marker":"[7]"},{"why":"Reports ATCOs' preference for a black-box baseline over visual explanations, the empirical puzzle this study investigates.","marker":"[15]"},{"why":"Shows that explanation features can be operationally irrelevant, motivating the needs-first approach the paper takes.","marker":"[4]"},{"why":"Defines operational explainability and the regulatory expectation that AI output be understandable to end users.","marker":"[1]"},{"why":"Provides prior evidence that transparency manipulations produced no main effect on ATCO acceptance or agreement, questioning the assumed need for explanations.","marker":"[16]"}],"fun_headline_variants":["ATCOs want AI explanations for reports, not when they agree","Explainable AI: Controllers need it for docs, not agreement","When do controllers need AI explanations? For reports, not agreement","ATCOs: AI explanations for documentation, not when they agree"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 11 explanation goals were borrowed from a non-aviation user study and applied to a hypothetical conflict scenario, so the results depend on this goal list and scenario capturing the goals controllers would actually act on in live operations.","fun_headline_variants_meta":{"raw":{"variants":["ATCOs want AI explanations for reports, not when they agree","Explainable AI: Controllers need it for docs, not agreement","When do controllers need AI explanations? For reports, not agreement","ATCOs: AI explanations for documentation, not when they agree"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000156,"raw_usage":{"total_tokens":1201,"prompt_tokens":913,"completion_tokens":288,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":214}},"tokens_in":529,"tokens_out":288,"duration_ms":2956,"temperature":1.0,"reasoning_tokens":214,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:58:22.588426+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A high-fidelity simulation in which controllers work live conflicts with a real advisory tool and can request explanations freely: if controllers frequently request explanations when they already agree with the advisory, or rarely request them while writing post-event reports, then the paper's goal-based need mapping does not predict real behavior.","supporting_citations":[{"cited_title":"A survey on artificial intelligence (ai) and explainable ai in air traffic management: Current trends and development with future research trajectory,","cited_arxiv_id":null,"evidence_quote":"Documents the field-wide gap in ATM XAI, where research centers on what systems can do rather than on user needs."},{"cited_title":"Usage of more transparent and explainable conflict resolution algorithm: air traffic controller feedback,","cited_arxiv_id":null,"evidence_quote":"Reports ATCOs' preference for a black-box baseline over visual explanations, the empirical puzzle this study investigates."},{"cited_title":"Explaining the unexplainable: Role of xai for flight take-off time delay prediction,","cited_arxiv_id":null,"evidence_quote":"Shows that explanation features can be operationally irrelevant, motivating the needs-first approach the paper takes."},{"cited_title":"EASA Artificial Intelligence Concept Paper Issue 2: Guidance for Level 1 & 2 Machine Learning Applications,","cited_arxiv_id":null,"evidence_quote":"Defines operational explainability and the regulatory expectation that AI output be understandable to end users."},{"cited_title":"D6.2 use cases transparency requirements,","cited_arxiv_id":null,"evidence_quote":"Provides prior evidence that transparency manipulations produced no main effect on ATCO acceptance or agreement, questioning the assumed need for explanations."}],"review_version":1}