{"id":"fe66c184-f159-49af-8371-fd5915e41e85","arxiv_id":"2507.10585","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A three-part taxonomy for prompt-based natural language explanations, covering context, generation and presentation, and evaluation with 15 desirable properties.","lead":"This paper proposes a structured taxonomy for designing and evaluating natural language explanations that are generated by prompting large language models. It organizes explanation design into context, generation and presentation, and evaluation dimensions, intended for researchers, auditors, and policymakers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Taxonomy's completeness rests on undocumented expert consensus; no inter-rater reliability or external validation, so transferability to prompt-based NLEs is unverified.","rationale":"The paper is a taxonomy proposal, and its central value is a structured checklist for prompt-based NLEs. For that checklist to be trustworthy, the source frameworks must be comprehensive for NLEs and the adaptation must be defensible. Both conditions are asserted rather than demonstrated: the methodology describes a small, self-referential expert consensus with no quantitative agreement metrics, and the evaluation properties are taken largely from frameworks designed for XAI in general, with only qualitative arguments for the NLE-specific additions. The reader's weakest assumption captures exactly this gap, and I agree it is the most load-bearing concern. The paper is transparent about leaving validation to future work, which partially mitigates the concern but does not eliminate it: the abstract and introduction state the taxonomy 'provides a framework' in the present tense, which is stronger than a proposal. The proposed concrete test—an independent gap analysis with two annotators—directly checks whether the 15 properties and three dimensions cover the space of NLE-specific concepts from the cited literature and recent practice. If the test surfaces unmapped properties, the taxonomy needs revision; if not, the conditional verdict can be upgraded. Since the concern does not point to an internal inconsistency or a fatal flaw, but rather to unverified assumptions that the paper itself acknowledges, the reader's CONDITIONAL verdict remains appropriate. No change to the verdict is needed; the paper should be accepted only with the explicit caveat that the taxonomy is a preliminary framework awaiting empirical validation.","tokens_in":10427,"tokens_out":5129,"duration_ms":56393,"concrete_test":"Perform a systematic gap analysis: extract every evaluation property or design dimension mentioned for natural language explanations in Cambria et al. (2023), Nauta et al. (2023), Liao et al. (2022), and a random sample of 20 prompt-based NLE papers from 2023–2025; have two independent annotators map each extracted item onto Table 1's subcategories and compute inter-annotator agreement. If any extracted item cannot be mapped (e.g., 'faithfulness to model internals' or 'verifiability') and is not explicitly excluded by the paper's scope, the taxonomy is incomplete and should be revised or presented as preliminary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that Table 1 'provides a framework for researchers, auditors, and policymakers to characterize, design, and enhance NLEs'—depends on the taxonomy being complete and correctly adapted for prompt-based NLEs. The methodology (Section 2) rests on two unverified pillars: (i) the completeness of the Schwalbe & Finzel meta-taxonomy and the Co-12 framework as source structures for NLEs, and (ii) the validity of the adaptation process, described only as 'an iterative, consensus-driven process involving a small group of experts' including the authors. No inter-rater reliability, no comparison against independent NLE-specific taxonomies, and no external validation is reported; Section 6 explicitly defers validation to future work. If the adaptation omitted or misallocated important NLE properties—for instance, faithfulness/verifiability of the explanation to the model's true reasoning, which the paper itself flags in Section 5 as 'a challenging technical task'—the framework could mislead designers and policymakers into believing all relevant dimensions are covered. The expert panel is not described in terms of size, selection criteria, or independence, so the reader cannot assess whether the consensus reflects broad field agreement or the authors' prior commitments.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a taxonomy for prompt-based natural language explanations (NLEs), adapted from the Schwalbe and Finzel XAI meta-taxonomy and the Nauta et al. Co-12 property framework. The taxonomy is structured in three dimensions: Context Definition, Generation and Presentation, and Evaluation, with 15 evaluation properties. The authors scope the work to post-hoc, model-agnostic, local prompt-based NLEs and illustrate the taxonomy with a use case in airborne anomaly detection, including a worked prompt and model response. The paper is explicitly a proposal: validation is deferred to future work, and several limitations (faithfulness, subjectivity, trade-offs) are acknowledged in Section 5.","tokens_in":10633,"tokens_out":4375,"duration_ms":49314,"significance":"If the taxonomy is accepted, it gives researchers, auditors, and policymakers a structured checklist for characterizing and evaluating prompt-based NLEs, filling a genuine gap since existing XAI taxonomies are not NLE-specific. The paper builds transparently on established sources (Schwalbe and Finzel, Nauta et al., Liao et al.), clearly scopes its claims, and includes a concrete example that helps operationalize the abstraction. Its limitations are honestly stated, including the lack of empirical validation and the acknowledged difficulty of faithfulness. As a proposal, the contribution is useful, but the present form leaves several load-bearing validity questions unanswered.","major_comments":[{"comment":"The example use case in Section 4 and Appendix B does not satisfy the scope defined in Section 2. The paper restricts itself to post-hoc, model-agnostic, and local prompt-based NLEs, but in the anomaly-detection example the same vision-language model both detects the anomaly and generates the explanation (the system prompt in Figure B.1 instructs the model to 'detect and explain anomalies'), which is ante-hoc and model-specific rather than post-hoc and model-agnostic. This inconsistency undermines the claim that the taxonomy applies to the stated scope. Please either revise the example to use a separate post-hoc explainer or explicitly extend the scope to include same-model explanations and adjust the definitions in Section 2 accordingly.","section":"Section 2 and Section 4/Appendix B"},{"comment":"The adaptation of the Co-12 framework conflates two distinct constructs. In Nauta et al., 'Covariate Complexity' refers to the number of covariates or features used in the explanation, whereas Table 2 defines 'Comprehensibility' as 'Uses human-understandable concepts and relations.' These are not equivalent, and renaming one to the other changes the meaning of the property. Please either retain the original construct with a clearer justification for the rename or define Comprehensibility as a new property and describe what happens to Covariate Complexity (e.g., whether it is omitted or folded into another property).","section":"Section 3.3 and Table 2"},{"comment":"The expert consensus process is described only as 'an iterative, consensus-driven process involving experts from multiple domains,' with no information on the number of experts, their selection criteria, their independence from the authors, or any inter-rater reliability assessment. Because the taxonomy's completeness and the specific adaptations are the central claims of the paper, this opaque methodology weakens confidence in the result. Please report the panel details and reliability metrics, or alternatively reframe the contribution as a proposed taxonomy whose validation is explicitly left to future work, softening the 'provides a framework' claim in the abstract and conclusion.","section":"Section 2 (Methodological approach)"}],"minor_comments":[{"comment":"There is a typo: 'meat-taxonomy' should be 'meta-taxonomy' in the first paragraph of Section 3.1.","section":"Section 3.1"},{"comment":"The phrase 'we adopt the golas of NLE generation' contains a typo; 'golas' should be 'goals'.","section":"Section 3.1"},{"comment":"The relationship between Interactivity and Input is acknowledged as blurred in the text, but Table 1 lists them as separate subcategories. Consider clarifying the boundary or merging them into a single subcategory, since the current presentation leaves the reader unsure how to apply them distinctly.","section":"Section 3.2 and Table 1"},{"comment":"In the sample response, 'Anomaly Detected: Yes' uses a capital 'Yes' while the prompt specifies the output should be 'YES' or 'NO' (all caps). Please ensure consistency between the prompt specification and the example response.","section":"Appendix B, Figure B.1"},{"comment":"The 'Confidence' property in the Presentation category is described as 'Communicates the model's certainty or uncertainty.' It would be clearer to specify whether this refers to the confidence in the prediction or in the explanation itself, since these can differ substantially for NLEs.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"This paper appears to originate from a workshop submission and is concise, which is fine, but the scope inconsistency in the example and the conceptual conflation in the Co-12 adaptation are substantive issues that a journal referee should require be fixed. The lack of methodological detail on the expert consensus process is also a barrier to assessing the taxonomy's validity. If the authors address these points, the paper could be suitable for publication as a taxonomy/proposal paper, provided the claims are appropriately scoped."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful, clearly scoped taxonomy proposal for prompt-based NLEs, not a demonstrated standard. The new subcategories (Audience, Explanation Goal, Translucence, Novelty, Personalization/Actionability) are genuinely missing from the source XAI taxonomies, and the paper's explicit scope (post-hoc, model-agnostic, local) keeps the claims contained. I read it as a checklist for designers and evaluators, and it does that job.\n\nWhat it does well: The adaptation is grounded in the right prior work, and the paper is transparent about what it omitted and why. The example in Appendix B shows the taxonomy in action with a real system prompt and response, which makes the abstract categories concrete. The discussion of trade-offs between properties, especially faithfulness vs comprehensibility, is honest and doesn't pretend the framework resolves them.\n\nWhere it's soft: The validity of the taxonomy depends on two things the paper doesn't demonstrate: the completeness of the source taxonomies (Schwalbe & Finzel, Co-12) and the quality of the expert consensus process. We're told only that a small group of experts, including the authors, iteratively adapted the framework. No inter-rater reliability, no independence check, no external validation. The paper says validation is future work, which is fine for a workshop paper, but it should be explicit that the taxonomy is a hypothesis, not a finding. The overlap between Interactivity and Input is a bit blurry, and some user-centered properties (Coherence vs Personalization) could be hard to distinguish in practice. These are minor.\n\nThe citation pattern looks fine. It uses the sources it builds on and doesn't bury prior work.\n\nWho it's for: people building or evaluating NLE systems, and anyone in technical governance who needs a shared vocabulary. A serious referee could push the authors to document the expert selection process, add a comparison to other NLE-specific taxonomies, and clarify category boundaries. I'd send it to review.","headline":"A well-scoped, useful taxonomy for prompt-based NLEs, but its completeness rests on an undocumented expert consensus; it deserves review, not blind adoption.","tokens_in":11177,"tokens_out":2099,"would_cite":true,"duration_ms":24109,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a three-part taxonomy with fifteen evaluation properties as a checklist for designing and assessing prompt-based natural language explanations.","keywords":["natural language explanations","XAI taxonomy","prompt-based explanations","explainable AI","AI governance","evaluation properties","large language models","transparency"],"falsifier":"A concrete test would be a coding study in which two independent teams of experts each classify a diverse sample of real prompt-based NLEs, from domains such as healthcare and finance, into the taxonomy's subcategories; if inter-rater agreement is low, or a substantial fraction of explanations cannot be assigned to any subcategory, the taxonomy's categories are not comprehensive or distinct. A second falsifier is a user study in which explanations designed with the taxonomy do no better on the fifteen properties than untailored explanations in a head-to-head comparison.","tokens_in":10234,"feed_emoji":"🧾","tokens_out":4633,"duration_ms":49137,"temperature":0.7,"pith_summary":"This paper argues that existing explainable-AI taxonomies do not capture what is special about explanations written in natural language, and proposes a dedicated taxonomy for prompt-based natural language explanations (NLEs). The taxonomy organizes NLE design and evaluation into three dimensions: context (task, data, audience, goal), generation and presentation (model type, prompt input, interactivity, output type, presentation form), and evaluation (content, presentation, and user-centered properties plus evaluation setting). It distills fifteen desirable properties, including correctness, comprehensibility, translucence, actionability, and novelty. If adopted, the taxonomy gives researchers, auditors, and policymakers a common checklist for designing, comparing, and verifying the explanations AI systems generate in plain language.","feed_headline":"New taxonomy sorts prompt-based AI explanations into three dimensions","feed_subtitle":"Context, generation, and evaluation properties give auditors and designers a shared checklist for verifying what models say.","key_machinery":"The load-bearing object is the two-table taxonomy: Table 1 maps context and generation/presentation, and Table 2 lists the fifteen evaluation properties. The machinery works by forcing each explanation design to be described on three axes—context, generation/presentation, evaluation—with each axis decomposed into subcategories; this decomposition turns a vague notion like 'a good explanation' into a checklist that can be applied before and after building a prompt.","core_discovery":"On the paper's own terms, the central claim is that the standard XAI taxonomy can be adapted into a structured, model-agnostic checklist specific to post-hoc, model-agnostic, local, prompt-based natural language explanations. The adaptation adds an Audience subcategory with five roles (creator, operator, executor, decision subject, examiner) and an Explanation Goal subcategory, reworks the explanation-generation module around prompt inputs and dialogue-based interactivity, and extends the twelve-property evaluation framework to fifteen properties for NLEs. The paper presents this as a practical instrument for technical AI governance: it lets stakeholders specify what an explanation is for, generate it in a form suited to that purpose, and evaluate it along properties that can be tested in functionally grounded, human-grounded, or application-grounded settings.","pith_inferences":["A next step the paper does not take is to turn the taxonomy into a scoring rubric and measure inter-rater reliability across independent annotators; that would test whether the categories are stable in practice.","The taxonomy could be operationalized as a multi-objective optimization for prompt search, where the fifteen properties serve as objectives and conflicting pairs, such as correctness versus comprehensibility, define a Pareto front.","Because the evaluation properties are defined independently of any particular model, the same checklist could plausibly be applied beyond prompt-based LLMs to retrieval-augmented or fine-tuned explanation generators with minor adjustments.","The audience-role subcategory suggests testable hypotheses, for example that examiner-oriented explanations should emphasize traceability and policy alignment, and user studies could verify that role-specific prompts improve task performance."],"forward_implications":["Designers can use the checklist to construct system prompts that specify task type, data type, audience, explanation goal, output type, and presentation format, as demonstrated in the traffic-anomaly use case.","Evaluators can choose among functionally grounded, human-grounded, and application-grounded settings for each of the fifteen properties, making evaluation plans more explicit and comparable across studies.","Auditors and policymakers gain a shared vocabulary for requesting and reviewing natural language explanations, supporting transparency and accountability requirements.","The taxonomy highlights tensions between properties, such as perfect correctness reducing comprehensibility, prompting explicit trade-off decisions rather than implicit ones.","Researchers can position new NLE work within the taxonomy and identify which properties a proposed method targets."],"supporting_citations":[{"why":"Supplies the base XAI meta-taxonomy that the paper adapts to natural language explanations.","marker":"Schwalbe & Finzel (2024)"},{"why":"Provides the definition of natural language explanations and a roadmap that motivates the adaptation.","marker":"Cambria et al. (2023)"},{"why":"Contributes the Co-12 framework of twelve explanation properties that the paper extends to fifteen.","marker":"Nauta et al. (2023)"},{"why":"Provides the usage-context-driven evaluation perspective used to rename and decompose properties.","marker":"Liao et al. (2022)"},{"why":"Defines the functionally grounded, human-grounded, and application-grounded evaluation settings.","marker":"Doshi-Velez & Kim (2017)"},{"why":"Supplies the five audience roles adopted for the Audience subcategory.","marker":"Tomsett et al. (2018)"},{"why":"Defines the three explanation goals adopted for the Explanation Goal subcategory.","marker":"Chen et al. (2022)"},{"why":"Contributes to the definition of NLEs and grounds the local, model-agnostic explanation framing.","marker":"Ribeiro et al. (2016)"}],"fun_headline_variants":["New taxonomy organizes prompt-based AI explanations into three dimensions","Three-part taxonomy for designing and testing AI explanations","Checklist framework for verifying AI's natural language explanations","Governance-ready taxonomy for prompt-based AI explanations","Map out AI explanation design and evaluation with this taxonomy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The taxonomy is only as sound as the two source frameworks it adapts—the broad XAI meta-taxonomy and the Co-12 property list—and the informal expert consensus process that chose the adaptations; if either source missed important facets of natural language explanations, or the expert panel was unrepresentative, the taxonomy inherits those gaps.","fun_headline_variants_meta":{"raw":{"variants":["New taxonomy organizes prompt-based AI explanations into three dimensions","Three-part taxonomy for designing and testing AI explanations","Checklist framework for verifying AI's natural language explanations","Governance-ready taxonomy for prompt-based AI explanations","Map out AI explanation design and evaluation with this taxonomy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1397,"prompt_tokens":844,"completion_tokens":553,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":460,"completion_tokens_details":{"reasoning_tokens":479}},"tokens_in":460,"tokens_out":553,"duration_ms":5872,"temperature":1.0,"reasoning_tokens":479,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:16:05.430376+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be a coding study in which two independent teams of experts each classify a diverse sample of real prompt-based NLEs, from domains such as healthcare and finance, into the taxonomy's subcategories; if inter-rater agreement is low, or a substantial fraction of explanations cannot be assigned to any subcategory, the taxonomy's categories are not comprehensive or distinct. A second falsifier is a user study in which explanations designed with the taxonomy do no better on the fifteen properties than untailored explanations in a head-to-head comparison.","supporting_citations":[],"review_version":1}