{"id":"09718d61-2e16-4183-89f5-6cb756b8a073","arxiv_id":"2412.17440","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":0.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey paper that defines XAI, catalogues interpretability concepts and post-hoc explanation techniques, and lists selected applications in aviation and aerospace.","lead":"This paper is a literature review of explainable AI (XAI) in aeronautics and aerospace, covering definitions, model properties, post-hoc techniques, and example applications. A generalist would read it as a compact orientation to how XAI is used in safety-critical aviation and space systems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 5 misattributes at least two application claims to reference [18], undermining the survey's reliability as a map of XAI use in aeronautics and aerospace.","rationale":"The reader's verdict of UNVERDICTED is appropriate because the paper makes no original research claim; it is a narrative survey. The load-bearing weakness is correctly identified as the dependence on cited sources accurately supporting the application descriptions. My stress-test confirms this concern is real and concrete: the duplicate citation [18] for two unrelated applications is a specific, checkable error that, if confirmed, directly undermines the survey's reliability. It does not, however, change the verdict: the paper was already unverified, and a faulty citation makes it less trustworthy but does not turn it into a rejected research contribution. The verdict should remain UNCHANGED. I would add that the Section 6 phrase 'it has been proven' is an overclaim for a review paper that only summarizes others' work; this supports the UNVERDICTED status rather than ACCEPT. The concrete test is a simple full-text search of the two references, which would settle the misattribution concern decisively.","tokens_in":6466,"tokens_out":2237,"duration_ms":22054,"concrete_test":"Obtain the full texts of references [17] and [18]. For [18], search for the strings 'Uganda', 'poverty', 'post-disaster', and 'damage assessment'; for [17], search for 'MSE', 'RMSE', 'MAE', and 'drone simulation'. If [18] contains none of the claimed applications and [17] contains no such simulation metrics, then Section 5 contains at least three unsupported application claims, confirming that the survey mischaracterizes the state of the art and the Section 6 'proven' assertion is unjustified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that post-hoc XAI techniques have been applied across ATM, UAV route adaptation, predictive maintenance, spacecraft telemetry anomaly detection, and satellite image analysis, with Section 6 asserting that 'it has been proven how the use of post-hoc techniques can ensure the interpretability of models.' The evidence for this is entirely the cited literature. That evidence is suspect: Section 5 cites [18] for 'post-natural disaster damage assessment' using drones and satellites, and later cites [18] again for 'predicting poverty indices in Uganda' using decision trees and deep object detection. Reference [18] is Banimelhem and Al-khateeb, 'Explainable artificial intelligence in drones: A brief review' — a short survey that cannot plausibly be the primary source for two unrelated empirical applications. The likely intended sources are [14] (Cheng et al., uncertainty-aware CNN for disaster damage assessment) and [19] (Ayush et al., efficient poverty mapping from remote sensing). Additionally, the sentence about drone simulations validated with MSE, RMSE, and MAE cites [17], Mualla et al., a paper about human-agent explanation formulation, which may not report such simulations. If these citations do not support the assigned claims, then Section 6's 'it has been proven' is an overstatement, and the paper's value as a survey — its only contribution — is materially weakened. The concern is load-bearing because the entire paper is a literature review; there is no independent evaluation, derivation, or experiment to fall back on.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a narrative survey of explainable artificial intelligence (XAI) aimed at aeronautics and aerospace applications. It defines XAI and its objectives, reviews the properties associated with interpretable and transparent models, describes transparent versus black-box models and post-hoc explanation techniques, and surveys six application areas: air traffic management, UAV route adaptation and simulation, post-disaster damage assessment, predictive maintenance, spacecraft telemetry anomaly detection, and satellite image analysis. The paper's central claim is that post-hoc techniques such as LIME and SHAP have been applied to make AI models understandable to professionals in these safety-critical sectors, and the conclusion asserts that this has been 'proven' by the surveyed applications.","tokens_in":6655,"tokens_out":8774,"duration_ms":79430,"significance":"If the survey's characterizations are accurate, it would provide a concise entry point for practitioners seeking to understand XAI methods in aeronautics and aerospace. The taxonomy in Sections 2–4 is broadly faithful to the authoritative sources it cites, and the selection of application areas is relevant. However, the paper contains no original analysis, systematic methodology, or independent evaluation, so its value rests entirely on the correctness of its citation-to-claim mapping. The citation problems in Section 5 are therefore load-bearing. The paper does not ship code, data, or proofs; its contribution is a synthesis, and that synthesis must be reliable.","major_comments":[{"comment":"Reference [18] (Banimelhem and Al-khateeb, 'Explainable artificial intelligence in drones: A brief review') is cited for two unrelated empirical applications: post-natural-disaster damage assessment using actual and predicted values, and prediction of poverty indices in Uganda using decision trees and deep object detection. A brief review cannot be the primary source for both empirical studies. The manuscript's own reference list contains the likely intended sources—[14] (Cheng et al., uncertainty-aware CNN for disaster damage assessment) and [19] (Ayush et al., efficient poverty mapping from remote sensing)—but [14] is instead cited in Section 2 for the definition of 'comprehensibility,' where it also does not belong. As written, two of the six application claims in Section 5 cannot be verified from the cited literature, and Section 6's statement that 'it has been proven' that post-hoc techniques ensure interpretability is not supported. The authors must re-verify every application claim against its primary source and correct the in-text citations accordingly.","section":"Section 5, paragraphs 3–4"},{"comment":"The sentence about drone simulations in three different modes validated with MSE, RMSE, and MAE cites [17], Mualla et al., 'The quest of parsimonious XAI: A human-agent architecture for explanation formulation.' The title and venue do not indicate that this paper reports drone simulations with the listed error metrics; please confirm that [17] is the correct source, replace it with the study that actually performed these simulations, or remove the quantitative detail. In addition, the following sentence about Grad-CAM attribution maps for daytime satellite images and nighttime light data in Sub-Saharan Africa cites [19], Ayush et al., whose abstract does not mention Grad-CAM; if a different paper is meant, it is missing from the reference list.","section":"Section 5, paragraph 3"},{"comment":"The sentence 'it has been proven how the use of post-hoc techniques can ensure the interpretability of models' overstates what a survey can establish. The cited works are applications or proposals; they do not constitute a proof, and the citation problems in Section 5 further weaken the evidential basis. The authors should replace 'proven' with language such as 'reported' or 'demonstrated in the cited cases,' and they should qualify the claim as reflecting the selected literature rather than a general guarantee.","section":"Section 6 (Conclusions)"},{"comment":"The review does not state a literature search strategy, inclusion/exclusion criteria, or time window, and it covers six application areas with one or two citations each. If the intent is a comprehensive survey of XAI in aeronautics and aerospace, the methodology and coverage criteria should be described; if the intent is to present illustrative examples, the authors should say so explicitly and soften the language in the conclusions. Without this framing, the reader cannot judge whether the selected applications are representative or selective.","section":"Section 5 (Applications)"}],"minor_comments":[{"comment":"The phrase 'this paper provides a review of the concept of XAI is carried out' is ungrammatical; consider 'this paper reviews the concept of XAI, defining the term and its objectives.'","section":"Abstract"},{"comment":"'Aerial Traffic Management' should be 'Air Traffic Management', and the comma after 'DARPA' should be removed.","section":"Section 1"},{"comment":"The sentence beginning 'A distinction is drawn between transparent models...' is duplicated immediately after Figure 1; one copy should be removed.","section":"Section 4"},{"comment":"The statement that 'the properties that should be evaluated in AI systems and models to be considered black-box models have been defined' is inconsistent with Section 3, where the properties are introduced as criteria for comprehensibility and explainability, not for black-box status.","section":"Section 6"},{"comment":"Reference [15] is a predictive-maintenance survey; it is not an obvious source for the performance-interpretability trade-off, which is already discussed in [8] and [9]. Please cite the original source for the trade-off concept.","section":"Section 3"},{"comment":"The reference list should be checked for consistency: [2] should be formatted as 'Nature News', and the citation numbering should be cross-checked against the in-text citations after the Section 5 corrections are made.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a plausible narrative review, and the taxonomy in Sections 2–4 is not problematic. The main risk is citation reliability: the [18]/[14]/[19] cross-wiring is easily detected and suggests the reference list was not rigorously checked. If the authors correct these citations and moderate the conclusions, the paper could be suitable as a short review. In its current form, the survey's core evidence is not trustworthy as written."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a pure literature review—no method, data, or new result—and it reads like a student survey done carefully enough on the standard definitions but carelessly on the applications. The useful part is Sections 2–4, which condense Arrieta et al., Molnar, and Lipton into a readable summary of XAI definitions, model properties, and post-hoc technique categories. If you need a one-page map of that material for a class or an intro talk, this is serviceable.\n\nThe soft spot is Section 5. The stress-test is right: reference [18] (Banimelhem and Al-khateeb, a brief review on XAI in drones) is cited for post-disaster damage assessment using drones and satellites, and later cited again for poverty-index prediction in Uganda. The reference list shows [14] Cheng et al. is the damage-assessment paper and [19] Ayush et al. is the poverty-mapping paper. That is not a minor typo; in a review, the citations are the data, and two unrelated application claims pointing to the same source make the whole application map unreliable until audited. There is also an overclaim in Section 6—'it has been proven how the use of post-hoc techniques can ensure the interpretability of models'—which should be 'reported' or 'illustrated' at best. I didn't see internal contradictions in the definitions or taxonomy, so the conceptual skeleton is sound.\n\nThe prose is rough throughout (e.g., the abstract has a broken sentence 'a review of the concept of XAI is carried out defining the term'), and the paper would need serious copyediting. But the bigger issue is that as a survey its value depends entirely on accurate reporting of others' work, and that reporting is not yet trustworthy.\n\nFor whom? Practitioners or students who want a compact orientation to XAI jargon could get something out of a cleaned-up version. Scholars won't. I would not send this to peer review at a research venue in its current form; it is too thin and too error-prone. If the authors fix the citations, tone down the conclusion, and tighten the writing, it could be a workshop paper or a teaching note. As is, I'd desk-reject.","headline":"A readable but careless XAI survey: the standard definitions are fine, but the application section's misattributed citations undercut the paper's only contribution.","tokens_in":7260,"tokens_out":4280,"would_cite":false,"duration_ms":37542,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper surveys XAI and reports that post-hoc techniques have been proven to make AI models interpretable across aeronautics and aerospace applications.","keywords":["Aeronautics","Aerospace","Explainable AI","Machine Learning","Safety-Critical Applications","Post-hoc explanation","LIME","SHAP"],"falsifier":"Inspect reference [18] and the other application citations: if the cited sources do not actually contain the described case studies, such as post-disaster drone damage assessment or poverty-index prediction, the survey's characterization of the state of the art collapses. A reader could also verify whether each cited paper names the specific post-hoc technique the survey attributes to it, such as LIME in spacecraft telemetry or SHAP in predictive maintenance.","tokens_in":6238,"feed_emoji":"✈️","tokens_out":4788,"duration_ms":39995,"temperature":0.7,"pith_summary":"This paper surveys explainable artificial intelligence (XAI) as a bridge between humans and AI systems in aeronautics and aerospace. It argues that post-hoc techniques such as LIME and SHAP, along with transparent models, are already being used so professionals can understand and trust AI predictions in air traffic management, drone operations, predictive maintenance, spacecraft telemetry, and satellite imagery. The review organizes the field by defining XAI's goals, the properties that make models interpretable, and a taxonomy of transparent models versus post-hoc explanation methods, concluding that across these sectors the use of post-hoc techniques has been shown to ensure model interpretability.","feed_headline":"Survey maps how XAI explains AI across aviation and space","feed_subtitle":"A review shows LIME, SHAP, and rule-based models making black-box decisions legible in safety-critical flight systems.","key_machinery":"The paper's central organizing device is a two-part taxonomy of XAI methods: transparent models, which are inherently interpretable because of their algorithmic simplicity, and post-hoc techniques, which are applied after training to explain black-box models. Post-hoc methods are split into model internals, model surrogates (LIME, SHAP, Anchors), feature summaries (feature importance, partial dependence plots), and example-based methods (counterfactuals, influential observations, prototypes and criticisms). This taxonomy frames every application described in the survey and supplies the vocabulary the paper uses to claim that interpretability can be achieved.","core_discovery":"The paper asserts that XAI has moved from a research concept to an applied tool in safety-critical aeronautics and aerospace settings. In particular, it contends that post-hoc explanation techniques, including local surrogate models like LIME, feature-attribution methods like SHAP, and example-based methods such as counterfactuals, can make the decisions of otherwise opaque neural networks intelligible to domain professionals. It catalogs applications in takeoff and landing time prediction, fuzzy-rule route adaptation for UAVs, post-disaster damage assessment, predictive maintenance, spacecraft telemetry anomaly detection, and satellite-image-based poverty mapping. The paper's conclusion is that post-hoc techniques have been proven to ensure interpretability in each of these areas.","pith_inferences":["The survey's own citation trail is the fragile part: reference [18] is cited for both post-disaster drone damage assessment and poverty-index prediction in Uganda, while its listed title is a brief review of XAI in drones. If the primary sources are misattributed, the survey's map of the field would mischaracterize the evidence base even though the broad claim may be plausible.","An implicit consequence of the paper's framing is that XAI's value in these sectors depends on whether explanations actually change human decisions; the survey documents that explanations are generated but does not measure their operational impact.","A testable extension would be a quantitative comparison of explanation fidelity across the cited application domains, since the paper assumes that LIME, SHAP, and other surrogates faithfully represent the black-box models they explain."],"forward_implications":["The paper claims that the performance–interpretability trade-off is a central constraint, so XAI adoption in safety-critical settings must balance accuracy against explainability.","It presents the taxonomy of transparent versus post-hoc methods as a complete map of available explanation approaches.","It asserts that post-hoc techniques such as LIME, SHAP, feature importance, and counterfactuals have been proven to ensure interpretability in ATM, UAVs, predictive maintenance, telemetry anomaly detection, and satellite imagery.","It states that explanations must be tailored to the user profile, which implies that practitioners and regulators may need different explanations from the same system."],"supporting_citations":[{"why":"Supplies the DARPA-origin definition of XAI and its goal of human understanding, trust, and management, used in Sections 2 and 4.","marker":"[3]"},{"why":"Evidence that XAI is applied to explain neural-network predictions in air traffic management, including takeoff and landing times and incident risk.","marker":"[4]"},{"why":"Evidence that LIME and SHAP have been used to explain deep neural networks in aerospace predictive maintenance.","marker":"[6]"},{"why":"Evidence that LIME was applied to explain anomaly detection in spacecraft telemetry.","marker":"[7]"},{"why":"Provides the taxonomy and definitions of XAI concepts, including interpretability, transparency, and the properties of explainable models, used throughout the paper.","marker":"[8]"},{"why":"Evidence that fuzzy-rule-based XAI explains route adaptation for unmanned aerial vehicles under adverse conditions.","marker":"[16]"},{"why":"The paper cites this reference for post-disaster drone damage assessment and, later, for poverty-index prediction in Uganda; the listed title is a brief review of XAI in drones.","marker":"[18]"},{"why":"Evidence that attribution maps such as Grad-CAM are used on satellite imagery and nighttime light data for poverty mapping in Sub-Saharan Africa.","marker":"[19]"}],"fun_headline_variants":["XAI makes AI explainable for aviation and space","How XAI brings transparency to aerospace AI","Explaining AI in flight: XAI review for aerospace","From black box to clear sky: XAI in aerospace","LIME and SHAP bring clarity to aerospace AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's map of XAI applications is only as reliable as the primary sources behind its citations, and the text itself uses reference [18] for two unrelated applications that do not match the reference's title, so this reliability cannot be taken for granted.","fun_headline_variants_meta":{"raw":{"variants":["XAI makes AI explainable for aviation and space","How XAI brings transparency to aerospace AI","Explaining AI in flight: XAI review for aerospace","From black box to clear sky: XAI in aerospace","LIME and SHAP bring clarity to aerospace AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000999,"raw_usage":{"total_tokens":4184,"prompt_tokens":856,"completion_tokens":3328,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":3252}},"tokens_in":472,"tokens_out":3328,"duration_ms":20803,"temperature":1.0,"reasoning_tokens":3252,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:26:14.423752+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect reference [18] and the other application citations: if the cited sources do not actually contain the described case studies, such as post-disaster drone damage assessment or poverty-index prediction, the survey's characterization of the state of the art collapses. A reader could also verify whether each cited paper names the specific post-hoc technique the survey attributes to it, such as LIME in spacecraft telemetry or SHAP in predictive maintenance.","supporting_citations":[{"cited_title":"Xai—explainable artificial intelligence,","cited_arxiv_id":null,"evidence_quote":"Supplies the DARPA-origin definition of XAI and its goal of human understanding, trust, and management, used in Sections 2 and 4."},{"cited_title":"A survey on artificial intelligence (ai) and explainable ai in air traffic management: Current trends and development with future research trajectory,","cited_arxiv_id":null,"evidence_quote":"Evidence that XAI is applied to explain neural-network predictions in air traffic management, including takeoff and landing times and incident risk."},{"cited_title":"Opportunities for explainable artificial intelligence in aerospace predictive maintenance,","cited_arxiv_id":null,"evidence_quote":"Evidence that LIME and SHAP have been used to explain deep neural networks in aerospace predictive maintenance."},{"cited_title":"Explainable anomaly detection in spacecraft telemetry,","cited_arxiv_id":null,"evidence_quote":"Evidence that LIME was applied to explain anomaly detection in spacecraft telemetry."},{"cited_title":"Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai,","cited_arxiv_id":null,"evidence_quote":"Provides the taxonomy and definitions of XAI concepts, including interpretability, transparency, and the properties of explainable models, used throughout the paper."},{"cited_title":"Evolving rule-based explainable artificial intelligence for unmanned aerial vehicles,","cited_arxiv_id":null,"evidence_quote":"Evidence that fuzzy-rule-based XAI explains route adaptation for unmanned aerial vehicles under adverse conditions."},{"cited_title":"Explainable artificial intelligence in drones: A brief review,","cited_arxiv_id":null,"evidence_quote":"The paper cites this reference for post-disaster drone damage assessment and, later, for poverty-index prediction in Uganda; the listed title is a brief review of XAI in drones."},{"cited_title":"Efficient poverty mapping from high resolution remote sensing images,","cited_arxiv_id":null,"evidence_quote":"Evidence that attribution maps such as Grad-CAM are used on satellite imagery and nighttime light data for poverty mapping in Sub-Saharan Africa."}],"review_version":1}