{"id":"f2a90f47-1bbf-477a-aa75-2c447ba8692f","arxiv_id":"2412.14209","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative review proposes a means-end framework that connects evidence hierarchies and evidential pluralism to the design of explainable AI decision support systems for construction.","lead":"This paper argues that explainable AI systems and decision support tools in construction should be designed around the evidence that supports their outputs, and proposes a means-end framework to do so. Generalists should read it for a clear synthesis of why explanations must be tied to trustworthy evidence, not just model internals.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The proposed evidence hierarchy ranks only methodological strength, but the framework's stated goal is to weigh strength, value, and utility; relevance and utility have no defined place in Table 3, so the hierarchy cannot by itself deliver context-tailored MHEs.","rationale":"The reader's weakest assumption targets the untested adaptation of medical evidence hierarchies to construction XAI design. That concern is real, but the more precise stress test is internal: even under the paper's own logic, the hierarchy is not aligned with its stated multi-criteria evidence evaluation. The paper is honest about its limitations, and the framework is a plausible starting point for theory building, but the central claim as written overstates what the hierarchy can deliver. The appropriate remedy is to add an explicit relevance and utility term to the evidence weighting procedure, or to temper claims that the hierarchy alone yields context-tailored meaningful explanations. This does not change the reader's CONDITIONAL verdict, but it sharpens the condition: acceptance should require either integrating relevance and utility into the hierarchy or restricting the claim to methodological strength of evidence. I do not see a deliberate flaw; this is a theory-building paper whose main instrument is incomplete relative to its own declared criteria.","tokens_in":30005,"tokens_out":11195,"duration_ms":105512,"concrete_test":"Apply Table 3 to the paper's own Section 6.3 tunneling-delay case with two concrete evidence sources: (A) a Level 1 systematic review of delay causes from infrastructure projects in other countries or contexts, and (B) a Level 3 or 4 mixed-method longitudinal study on the specific project combining process-tracing interviews with project data. Under Table 3, A ranks above B. Then check whether the framework contains any rule that can override or adjust the hierarchy when B is more relevant and useful to the local decision context. If no such rule exists, the framework cannot implement its stated strength, value, and utility criterion. A corroborating empirical version would have construction practitioners rate MHEs generated from A-versus-B evidence on relevance, actionability, and decision impact; if B-supported explanations score higher, the omission is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the means–end framework ensures evidence-based, meaningful explanations. The paper explicitly says it emphasizes evaluating the strength, value, and utility of different types of evidence (Abstract; repeated in Section 6.1). However, the only operational instrument, the evidence hierarchy in Table 1 and its construction adaptation in Table 3, ranks evidence exclusively by study design and methodological rigor: Level 1 systematic reviews/meta-analyses, Level 2 quasi-experiments, Level 3 mixed-method longitudinal case studies, down to Level 10 expert opinion. There is no column, scoring rule, or procedure for relevance to a given project, decision task, stakeholder role, or local VUCA context. Section 6.3's illustrative application says evidence is 'weighted accordingly' but no weighting function is specified. Consequently, the hierarchy can recommend a Level 1 meta-analysis from a different sector or region over a lower-level, mechanism-rich local field study, even when the latter is more relevant and useful to the end-user's actual decision. The paper acknowledges context-dependence through VUCA (Section 5.2) but does not encode it in the hierarchy. Without relevance and utility in the ranking, there is no warrant for the claim that applying the hierarchy produces MHEs tailored to users' knowledge needs and decision contexts. This is an internal incompleteness, not merely a lack of empirical validation: the declared criteria are strength, value, and utility, but only strength is operationalized.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a theoretical means–end framework for integrating evidence into the design of explainable AI (XAI) and AI-based decision support systems (DSSs), aimed at construction end-users. Developed through a narrative review spanning computer science, philosophy, and medicine, the framework combines epistemic normativity and instrumental rationality: evidence, generated through association and mechanism studies, is ranked via an adapted evidence hierarchy (Table 3) and used to design XAI instruments that should produce Meaningful Human Explanations (MHEs). The paper also provides a five-step illustrative application (Section 6.3) and a multi-dimensional evaluation scheme (Section 6.2). No empirical data are presented; the authors explicitly acknowledge this in Section 6.5, stating that the framework is conceptual and requires future validation.","tokens_in":30313,"tokens_out":4520,"duration_ms":43952,"significance":"If its central claim were made operational, the framework would fill a genuine gap: most construction XAI research is technical and model-centric, with little attention to how evidence supports explanations. The paper's strengths are its interdisciplinary synthesis, its adaptation of a medical evidence hierarchy to construction's VUCA realities, its explicit evaluation dimensions (cognitive validity, epistemic alignment, robustness, decision impact, normative justifiability), and its candid statement of limitations. The narrative review is transparent about its method, and the framework draws on external philosophical sources (Buchholz, Williamson, Salmon) rather than being self-referential. However, the central claim that the framework ensures evidence-based, context-tailored MHEs is currently untested and, more importantly, internally incomplete: the operational ranking instrument does not encode the relevance and utility dimensions that the paper repeatedly invokes.","major_comments":[{"comment":"The abstract and Section 6.1 state that the framework evaluates the 'strength, value, and utility' of evidence, but the only operational ranking instrument, Table 1 and its construction adaptation in Table 3, ranks evidence purely by study design and methodological rigor. There is no column, scoring rule, or procedure for relevance to a specific project, decision task, stakeholder role, or VUCA context. In Section 6.3, step 1 says evidence is 'weighted accordingly' but no weighting function is specified. Consequently, the hierarchy can recommend a Level 1 meta-analysis from another sector over a lower-level, mechanism-rich local field study that is more relevant and useful to the end-user. Since the central claim is that the framework produces MHEs 'tailored to users' knowledge needs and decision contexts,' the framework as presented cannot deliver that outcome. The authors should either add an explicit relevance/utility weighting procedure to the hierarchy or temper the claim that applying the hierarchy ensures context-tailored MHEs.","section":"Section 5.2; Tables 1 and 3; Section 6.1; Section 6.3"},{"comment":"The concept of MHE is defined only by desiderata ('intelligible, relevant, and actionable') and by three components drawn from earlier work, but the paper provides no method for eliciting end-users' epistemic ends in a concrete design process. Section 4.1 correctly states that the suitability of an XAI instrument depends on specifying the epistemic end before acquiring evidence and building the DSS, yet the framework gives no procedure for such specification. Section 6.5 acknowledges that the framework is abstract and hard to operationalize, but the paper's own wording in Section 6.1 claims that the framework 'will enable construction organizations to realize the benefits and business value of XAI and DSSs.' Without a stakeholder-analysis or requirements-elicitation step, the alignment between evidence, XAI instrument, and user context remains an assertion rather than a framework property. The authors should add an explicit elicitation and mapping step, or clearly reposition the contribution as a conceptual scaffold whose operationalization is future work.","section":"Section 4.1; Section 6.1; Section 6.5"},{"comment":"The paper asserts that 'rework is not a risk but an uncertainty, which is probabilistically unmeasurable' and that machine learning techniques are therefore unable to accurately predict rework costs. This is presented as settled fact and used to dismiss a specific study (Mostofi et al.). No empirical evidence or citation is provided for this strong claim, which sits uneasily with the paper's own evidence-based rhetoric. If the claim is a philosophical or practical assumption, it should be labeled as such and justified; if it is an empirical claim, it needs supporting evidence. As written, it is an unsupported axiom used to motivate the framework and to criticize a published study, and it should be revised.","section":"Section 4, paragraph beginning 'A case in point'"}],"minor_comments":[{"comment":"Two subsections are both numbered 5.2: 'Hierarchy of Evidence' and 'Evidential Pluralism.' This numbering error should be corrected.","section":"Section 5.2"},{"comment":"The text refers to 'Mostifi et al.' but the reference is to 'Mostofi et al.'; the spelling should be made consistent.","section":"Section 4 and reference [90]"},{"comment":"The arXiv identifier for the generative XAI survey is given as 'arXiv:2014.09554'; this appears to be a typo for 'arXiv:2404.09554' and should be corrected.","section":"Reference [11]"},{"comment":"References [36] and [45] appear to be the same paper (Luo et al. 2024, IEEE Transactions on Engineering Management). Duplicate references should be removed or merged.","section":"References [36] and [45]"},{"comment":"The statement that 'when authors requested access to AI training data... the corresponding authors of papers repeatedly ignored these requests' is an empirical claim about the behavior of other researchers, with no citation or data. Either provide supporting evidence or remove the anecdotal assertion.","section":"Section 6.4"},{"comment":"Figure 3 uses the notation P(G|E), which suggests a probabilistic update of ground truth given evidence, while the text emphasizes causal-mechanical dependence and 'cause-effect determination.' Clarify how the probabilistic notation relates to the causal language, or adjust the notation to avoid conflating association with causation.","section":"Figure 3 and surrounding text"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on the authors' own prior work (Love et al. [2], [6]) to establish both the research gap and the definition of MHE. This is not circular in a technical sense, but it does mean that the novelty claim depends on the acceptance of those earlier papers. The fit with a human-computer interaction venue is also not obvious; the paper is more oriented toward construction engineering and management. If the editors agree that the relevance/utility gap in the evidence hierarchy can be fixed, the paper could become a useful conceptual contribution, but in its current form the central claim is not fully supported by the proposed operational instrument."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid narrative review that recombines Buchholz's means-end epistemology, Williamson's evidential pluralism, and a medical evidence hierarchy into a framework for evidence-based XAI design in construction. The genuinely new pieces are the construction-adapted evidence hierarchy (Table 3) and the 'Meaningful Human Explanations' (MHE) construct. The paper is clearly written and unusually honest about its limitations.\n\nWhere it does well: the framing of XAI explanations as requiring evidentiary grounding is sensible, and the adaptation of the hierarchy to favor mechanism-rich longitudinal studies over pure RCTs is a thoughtful response to construction's VUCA reality. The MHE concept usefully shifts attention from technical output to user comprehension and actionability. The five-step illustrative application in Section 6.3 is a helpful gesture toward operationalization.\n\nThe soft spots are real but not fatal. The stress-test concern holds up: the abstract and Section 6.1 promise evaluation of strength, value, and utility of evidence, but the only operational instrument—the hierarchy in Tables 1 and 3—ranks exclusively by methodological rigor. Relevance to a given project, task, or stakeholder is not encoded anywhere. Section 6.3 says evidence is 'weighted accordingly' but no weighting procedure is given. So the framework cannot, by itself, deliver context-tailored MHEs; that claim outruns the machinery. This is an internal incompleteness, not just a lack of empirical validation, and the authors should either add a relevance/utility dimension to the hierarchy or soften the claims.\n\nAlso, the central claim is untested, as the paper concedes. The review method is not reproducible, though that is inherent in narrative reviews and the authors acknowledge it. Self-citation is present but not egregious; they cite their prior surveys to establish the gap, which is reasonable.\n\nBottom line: this is a useful synthesis for researchers in construction XAI or evidence-based system design, not yet a validated design method. It deserves a serious referee, but the referee should ask the authors to close the gap between the declared criteria and the operationalized hierarchy, or temper the language about MHEs. I'd bring it to a reading group as a prompt for discussing evidence hierarchies in XAI.","headline":"A competent synthesis that overpromises: the evidence hierarchy only ranks methodological strength, so the framework can't deliver the context-tailored explanations it claims.","tokens_in":30817,"tokens_out":2156,"would_cite":true,"duration_ms":19479,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a means-end framework in which construction AI explanations are designed from a ranked evidence hierarchy, aligned to users' knowledge goals.","keywords":["explainable artificial intelligence","decision support systems","construction","evidence hierarchy","means-end framework","meaningful human explanations","evidential pluralism","machine learning"],"falsifier":"Run a randomized experiment with construction site managers using a delay-risk DSS: give one group explanations sourced from high-hierarchy evidence (longitudinal field studies plus mechanism interviews) and another from low-hierarchy evidence (expert opinion or single case study). If the low-evidence group shows equal comprehension, trust calibration, and decision accuracy, the hierarchy's ranking does not carry the framework's weight.","tokens_in":29810,"feed_emoji":"🏗️","tokens_out":3739,"duration_ms":33483,"temperature":0.7,"pith_summary":"The paper aims to establish that the reliability of an AI explanation depends on the quality of the evidence behind it, not just on the cleverness of the explanation technique. It proposes a means-end framework: pick the explanatory tools (means) only after stating what knowledge the user needs (ends), and grade the available evidence on a hierarchy adapted from medicine to construction's messy field conditions. If the framework is right, construction firms can specify to vendors what evidence an AI recommendation must rest on, regulators can audit explanations against a known standard, and users can avoid both blind trust and unjustified skepticism. The central novelty is treating evidence quality as a first-class design input for explainable AI, not an afterthought.","feed_headline":"Evidence ladder proposed to make construction AI explanations trustworthy","feed_subtitle":"The paper ranks evidence quality, then ties explanation tools to users' knowledge goals.","key_machinery":"The framework's load-bearing object is the adapted hierarchy of evidence (Table 3), a ten-level ranking from systematic reviews of longitudinal field studies at the top down to individual expert opinion at the bottom, used as a heuristic to weight evidence when designing model features, choosing XAI techniques, and evaluating explanations. It operates through the means-end structure: epistemic goals (understanding, justified belief, truth, insight) determine which XAI instrument is the right means, and the hierarchy determines which evidence may serve as that instrument's warrant. Evidential Pluralism supplies the causal criterion—both correlation and mechanism must be shown—so that explanations grounded only in statistical association are treated as weak.","core_discovery":"The central claim is that a DSS explanation is only meaningful—what the paper calls a Meaningful Human Explanation—when it is built on evidence whose strength has been explicitly ranked, and when the choice of XAI instrument is derived from the end-user's epistemic goal via instrumental rationality. The paper adapts a ten-level evidence hierarchy from medicine, re-ordered for construction's volatile, uncertain, complex, and ambiguous settings, and couples it with Evidential Pluralism: a causal claim is justified only when both a correlation and a mechanism linking cause and effect are demonstrated. The framework then maps association studies (e.g., historical project data) and mechanism studies (e.g., process-tracing interviews) onto specific XAI techniques—SHAP and LIME as associative evidence, counterfactual explanations as quasi-mechanistic levers—so that explanations answer 'why' and 'how', not merely 'what'.","pith_inferences":["If the framework holds, evidence-based design could spread beyond construction to other safety-critical engineering domains with similar VUCA conditions, such as mining or offshore operations.","A testable prediction emerges: user trust and decision accuracy should correlate with evidence level on the hierarchy, which could be tested by giving different user groups explanations sourced from different hierarchy levels.","The hierarchy may need to be task-specific rather than universal, since a longitudinal study feasible for safety may be impossible for bespoke one-off projects, making the ranking's applicability context-dependent."],"forward_implications":["Construction organizations can specify evidence requirements to DSS vendors instead of accepting black-box outputs.","Post-hoc tools like SHAP and LIME are repositioned as associative, not causal, and must be triangulated with mechanism studies.","A shared repository of use cases and open datasets becomes a prerequisite for applying the framework in practice.","The hierarchy gives regulators and auditors a concrete standard for judging whether an AI explanation is justified."],"supporting_citations":[{"why":"Supplies the original evidence-based XAI approach and the ten-level hierarchy that the paper adapts for construction.","marker":"[10]"},{"why":"Provides the means-end account of explainable AI that anchors the framework's central structure.","marker":"[24]"},{"why":"Argues that statistical associations alone cannot explain, motivating the need for mechanism studies and the 'why' answer.","marker":"[26]"},{"why":"Establishes the evidential pluralism criterion that a causal claim needs both correlation and mechanism.","marker":"[28]"},{"why":"Extends evidential pluralism to XAI, used for grounding causal discovery in the framework.","marker":"[129]"},{"why":"Defines the XAI precepts and identifies gaps in construction that the framework addresses.","marker":"[6]"},{"why":"Provides the epistemic conception of explanation adopted for defining Meaningful Human Explanations.","marker":"[72]"},{"why":"Supplies the theory of Evidential Pluralism applied to social sciences, supporting the use of multiple evidence sources.","marker":"[128]"}],"fun_headline_variants":["Construction AI explanations get an evidence ladder","Means-end framework ranks evidence for AI explanations","Evidence hierarchy powers meaningful explanations in construction AI","SHAP and LIME? Only as good as their evidence, says new framework","Why construction AI needs evidence-ranked explanations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework rests on the assumption that ranking evidence quality on a ten-level, medicine-inspired hierarchy actually tracks how much an explanation improves a construction practitioner's understanding and decision—an assumption the paper states but does not test.","fun_headline_variants_meta":{"raw":{"variants":["Construction AI explanations get an evidence ladder","Means-end framework ranks evidence for AI explanations","Evidence hierarchy powers meaningful explanations in construction AI","SHAP and LIME? Only as good as their evidence, says new framework","Why construction AI needs evidence-ranked explanations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000387,"raw_usage":{"total_tokens":2007,"prompt_tokens":873,"completion_tokens":1134,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":1061}},"tokens_in":489,"tokens_out":1134,"duration_ms":10809,"temperature":1.0,"reasoning_tokens":1061,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:38:46.711049+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a randomized experiment with construction site managers using a delay-risk DSS: give one group explanations sourced from high-hierarchy evidence (longitudinal field studies plus mechanism interviews) and another from low-hierarchy evidence (expert opinion or single case study). If the low-evidence group shows equal comprehension, trust calibration, and decision accuracy, the hierarchy's ranking does not carry the framework's weight.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends evidential pluralism to XAI, used for grounding causal discovery in the framework."}],"review_version":1}