REVIEW 3 major objections 6 minor 16 references
Integrating Evidence into the Design of XAI and AI-based Decision Support Systems: A Means-End Framework for End-users in Construction
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper proposes a means-end framework in which construction AI explanations are designed from a ranked evidence hierarchy, aligned to users' knowledge goals.
desk verdict A competent synthesis that overpromises: the evidence hierarchy only ranks methodological strength, so the framework can't deliver the context-tailored explanations it claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The framework's load-bearing object is the adapted hierarchy of evidence (Table 3), a ten-level ranking from systematic reviews of longitudinal field studies at the top down to individual expert opinion at the bottom, used as a heuristic to weight evidence when designing model features, choosing XAI techniques, and evaluating explanations. It operates through the means-end structure: epistemic goals (understanding, justified belief, truth, insight) determine which XAI instrument is the right means, and the hierarchy determines which evidence may serve as that instrument's warrant. Evidential Pluralism supplies the causal criterion—both correlation and mechanism must be shown—so that explanations grounded only in statistical association are treated as weak.
What would settle it
Run a randomized experiment with construction site managers using a delay-risk DSS: give one group explanations sourced from high-hierarchy evidence (longitudinal field studies plus mechanism interviews) and another from low-hierarchy evidence (expert opinion or single case study). If the low-evidence group shows equal comprehension, trust calibration, and decision accuracy, the hierarchy's ranking does not carry the framework's weight.
Extended reading notes
Core claim
The central claim is that a DSS explanation is only meaningful—what the paper calls a Meaningful Human Explanation—when it is built on evidence whose strength has been explicitly ranked, and when the choice of XAI instrument is derived from the end-user's epistemic goal via instrumental rationality. The paper adapts a ten-level evidence hierarchy from medicine, re-ordered for construction's volatile, uncertain, complex, and ambiguous settings, and couples it with Evidential Pluralism: a causal claim is justified only when both a correlation and a mechanism linking cause and effect are demonstrated. The framework then maps association studies (e.g., historical project data) and mechanism studies (e.g., process-tracing interviews) onto specific XAI techniques—SHAP and LIME as associative evidence, counterfactual explanations as quasi-mechanistic levers—so that explanations answer 'why' and 'how', not merely 'what'.
Load-bearing premise
The framework rests on the assumption that ranking evidence quality on a ten-level, medicine-inspired hierarchy actually tracks how much an explanation improves a construction practitioner's understanding and decision—an assumption the paper states but does not test.
Editorial extensions
If this is right
- Construction organizations can specify evidence requirements to DSS vendors instead of accepting black-box outputs.
- Post-hoc tools like SHAP and LIME are repositioned as associative, not causal, and must be triangulated with mechanism studies.
- A shared repository of use cases and open datasets becomes a prerequisite for applying the framework in practice.
- The hierarchy gives regulators and auditors a concrete standard for judging whether an AI explanation is justified.
Reading between the lines
- If the framework holds, evidence-based design could spread beyond construction to other safety-critical engineering domains with similar VUCA conditions, such as mining or offshore operations.
- A testable prediction emerges: user trust and decision accuracy should correlate with evidence level on the hierarchy, which could be tested by giving different user groups explanations sourced from different hierarchy levels.
- The hierarchy may need to be task-specific rather than universal, since a longitudinal study feasible for safety may be impossible for bespoke one-off projects, making the ranking's applicability context-dependent.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a theoretical means–end framework for integrating evidence into the design of explainable AI (XAI) and AI-based decision support systems (DSSs), aimed at construction end-users. Developed through a narrative review spanning computer science, philosophy, and medicine, the framework combines epistemic normativity and instrumental rationality: evidence, generated through association and mechanism studies, is ranked via an adapted evidence hierarchy (Table 3) and used to design XAI instruments that should produce Meaningful Human Explanations (MHEs). The paper also provides a five-step illustrative application (Section 6.3) and a multi-dimensional evaluation scheme (Section 6.2). No empirical data are presented; the authors explicitly acknowledge this in Section 6.5, stating that the framework is conceptual and requires future validation.
Significance. If its central claim were made operational, the framework would fill a genuine gap: most construction XAI research is technical and model-centric, with little attention to how evidence supports explanations. The paper's strengths are its interdisciplinary synthesis, its adaptation of a medical evidence hierarchy to construction's VUCA realities, its explicit evaluation dimensions (cognitive validity, epistemic alignment, robustness, decision impact, normative justifiability), and its candid statement of limitations. The narrative review is transparent about its method, and the framework draws on external philosophical sources (Buchholz, Williamson, Salmon) rather than being self-referential. However, the central claim that the framework ensures evidence-based, context-tailored MHEs is currently untested and, more importantly, internally incomplete: the operational ranking instrument does not encode the relevance and utility dimensions that the paper repeatedly invokes.
major comments (3)
- [Section 5.2; Tables 1 and 3; Section 6.1; Section 6.3] The abstract and Section 6.1 state that the framework evaluates the 'strength, value, and utility' of evidence, but the only operational ranking instrument, Table 1 and its construction adaptation in Table 3, ranks evidence purely by study design and methodological rigor. There is no column, scoring rule, or procedure for relevance to a specific project, decision task, stakeholder role, or VUCA context. In Section 6.3, step 1 says evidence is 'weighted accordingly' but no weighting function is specified. Consequently, the hierarchy can recommend a Level 1 meta-analysis from another sector over a lower-level, mechanism-rich local field study that is more relevant and useful to the end-user. Since the central claim is that the framework produces MHEs 'tailored to users' knowledge needs and decision contexts,' the framework as presented cannot deliver that outcome. The authors should either add an explicit relevance/utility weighting procedure to the hierarchy or temper the claim that applying the hierarchy ensures context-tailored MHEs.
- [Section 4.1; Section 6.1; Section 6.5] The concept of MHE is defined only by desiderata ('intelligible, relevant, and actionable') and by three components drawn from earlier work, but the paper provides no method for eliciting end-users' epistemic ends in a concrete design process. Section 4.1 correctly states that the suitability of an XAI instrument depends on specifying the epistemic end before acquiring evidence and building the DSS, yet the framework gives no procedure for such specification. Section 6.5 acknowledges that the framework is abstract and hard to operationalize, but the paper's own wording in Section 6.1 claims that the framework 'will enable construction organizations to realize the benefits and business value of XAI and DSSs.' Without a stakeholder-analysis or requirements-elicitation step, the alignment between evidence, XAI instrument, and user context remains an assertion rather than a framework property. The authors should add an explicit elicitation and mapping step, or clearly reposition the contribution as a conceptual scaffold whose operationalization is future work.
- [Section 4, paragraph beginning 'A case in point'] The paper asserts that 'rework is not a risk but an uncertainty, which is probabilistically unmeasurable' and that machine learning techniques are therefore unable to accurately predict rework costs. This is presented as settled fact and used to dismiss a specific study (Mostofi et al.). No empirical evidence or citation is provided for this strong claim, which sits uneasily with the paper's own evidence-based rhetoric. If the claim is a philosophical or practical assumption, it should be labeled as such and justified; if it is an empirical claim, it needs supporting evidence. As written, it is an unsupported axiom used to motivate the framework and to criticize a published study, and it should be revised.
minor comments (6)
- [Section 5.2] Two subsections are both numbered 5.2: 'Hierarchy of Evidence' and 'Evidential Pluralism.' This numbering error should be corrected.
- [Section 4 and reference [90]] The text refers to 'Mostifi et al.' but the reference is to 'Mostofi et al.'; the spelling should be made consistent.
- [Reference [11]] The arXiv identifier for the generative XAI survey is given as 'arXiv:2014.09554'; this appears to be a typo for 'arXiv:2404.09554' and should be corrected.
- [References [36] and [45]] References [36] and [45] appear to be the same paper (Luo et al. 2024, IEEE Transactions on Engineering Management). Duplicate references should be removed or merged.
- [Section 6.4] The statement that 'when authors requested access to AI training data... the corresponding authors of papers repeatedly ignored these requests' is an empirical claim about the behavior of other researchers, with no citation or data. Either provide supporting evidence or remove the anecdotal assertion.
- [Figure 3 and surrounding text] Figure 3 uses the notation P(G|E), which suggests a probabilistic update of ground truth given evidence, while the text emphasizes causal-mechanical dependence and 'cause-effect determination.' Clarify how the probabilistic notation relates to the causal language, or adjust the notation to avoid conflating association with causation.
Circularity Check
No significant circularity: the framework's normative and evidential core is taken from external philosophy and medicine sources, and the paper makes no fitted empirical claims that could reduce to its own inputs.
full rationale
The central components of the proposed means–end framework are explicitly grounded in external sources: Salmon's epistemic conception of explanation and Buchholz's means-end account ([24], [72]), Williamson's Evidential Pluralism ([28], [129]), and Famiglini et al.'s evidence hierarchy ([10]). The paper's construction-specific adaptation (Table 3) is offered as a design proposal motivated by the VUCA nature of construction, not as a consequence derived from the authors' prior work. Self-citations such as Love et al. [2] and [6] are used to describe the research context and to point to earlier XAI reviews and evaluation concepts; they do not supply the load-bearing warrant, which rests on the external philosophical and medical literature cited above. No parameters are fitted and no empirical quantity is predicted, so there is no fitted-input-called-prediction or self-definitional reduction. The paper's own limitation statement (Section 6.5) acknowledges that the framework is conceptual and not empirically validated, confirming that the contribution is a synthesis and a proposal rather than a derivation. Accordingly, no circular steps are identified.
Assumptions & free parameters
assumptions (5)
- domain assumption Construction projects operate in a VUCA (volatile, uncertain, complex, ambiguous) environment where randomized controlled trials and meta-analyses are rarely feasible.
- domain assumption Establishing a causal claim requires both a correlation and a mechanism (Evidential Pluralism).
- domain assumption Post-hoc XAI methods such as LIME and SHAP provide associative evidence but cannot explain why an output occurred.
- ad hoc to paper Rework in construction is an uncertainty that is probabilistically unmeasurable, so machine learning cannot accurately predict rework costs.
- domain assumption An explanation is successful if it enhances an individual's understanding (Salmon's epistemic conception).
invented entities (1)
-
Meaningful Human Explanations (MHEs)
Cite this review
Pith. "Pith review of Integrating Evidence into the Design of XAI and AI-based Decision Support Systems: A Means-End Framework for End-users in Construction." pith.science (2026). https://pith.science/paper/B6BJISJX
@misc{pith2026241214209,
author = {Pith},
title = {Pith review of: Integrating Evidence into the Design of XAI and AI-based Decision Support Systems: A Means-End Framework for End-users in Construction},
year = {2026},
howpublished = {\url{https://pith.science/paper/B6BJISJX}},
note = {Machine review of arXiv:2412.14209}
}
read the original abstract
Explainable Artificial Intelligence seeks to make the reasoning processes of AI models transparent and interpretable, particularly in complex decision making environments. In the construction industry, where AI based decision support systems are increasingly adopted, limited attention has been paid to the integration of supporting evidence that underpins the reliability and accountability of AI generated outputs. The absence of such evidence undermines the validity of explanations and the trustworthiness of system recommendations. This paper addresses this gap by introducing a theoretical, evidence based means end framework developed through a narrative review. The framework offers an epistemic foundation for designing XAI enabled DSS that generate meaningful explanations tailored to users knowledge needs and decision contexts. It focuses on evaluating the strength, relevance, and utility of different types of evidence supporting AI generated explanations. While developed with construction professionals as primary end users, the framework is also applicable to developers, regulators, and project managers with varying epistemic goals.
Figures
Reference graph
Works this paper leans on
-
[2]
Love, P.E.D., Matthews, J., Fang, W., Porter, S., Luo, H., & L. Ding L-Y. (2024). Learning to comprehend and trust artificial intelligence outcomes: A conceptual explainable AI evaluation framework. IEEE Engineering Management Review, 52(1), pp. 230-247, doi.org/10.1109/EMR.2023.3342200. [3] Mosier, K.L., & Skitka, L.J. (1996). Human decision-makers and a...
arXiv 2024
-
[9]
Miller, T. (2019). Explanation in artificial intelligence: Insights from social science. Artificial Intelligence, 267, pp.1-38, doi.org/10.1016/j.artint.2018.07.007 [10] Famiglini, L., Campagner, A., Barandas, M., La Maida, G., Gallazzi, E., & Cabitza, F. (2024). Evidence-based XAI: An empirical approach to design more effective explainable decision-suppo...
arXiv 2019
-
[16]
That’s not the output I expected!
Hossain, F., Hossain, R., & Hossain, E. (2021). Explainable Artificial Intelligence (XAI): An engineering perspective, arXiv.2101.03613 [cs.LG], doi.org/10.48550/arXiv.2101.03613 [17] Langer, M., Oster, D., Speith, T., Hermanns, H., Kastner, L., Schmidt, E., Sesing, A., & Baum, K. (2021). What do we want from explainable Artificial Intelligence (XAI)? – A...
-
[43]
Forest, F., Porta, H., Tuia, D., & Fink, O. (2024). From classification to segmentation with explainable AI: A study on crack detection and growth monitoring. Automation in Construction, 165, 105497, doi.org/10.1016/j.autcon.2024.105497 [44] Ghasemi, A., & Naser, M.Z. (2023). Tailoring 3D printed concrete through explainable artificial intelligence. Struc...
-
[50]
Snyder, H. (2019). Literature review as a research methodology: An overview and guidelines. Journal of Business Research, 104, pp.333-339, doi.org/10.1016/j.jbusres.2019.07.039 [51] Webster, J., & Watson, R.T. (2002). Analyzing the past to prepare for the future: Writing a literature review. Management Information Systems Quarterly, 26(2), pp.xiii-xxiii, ...
arXiv 2019
-
[57]
Moher, D., Liberati, A., Tetzlaff, J., Altman, D.G., & PRISMA Group. (2009). Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement. PLoS Medicine 6(7): e1000097, doi.org/10.1371/journal.pmed.1000097. [58] Uttley, L., Quintana, D.S., Montgomery, P., Carroll, C., Page, M.J., Falzon, L., Sutton, A., & Moher, D. (2023). The ...
-
[64]
Naser, M.Z. & Çiftçiogu, A.Ö. (2023). Causal discovery and inference for evaluating fire resistance of structural members through causal learning and domain knowledge. Structural Concrete, 24(3), pp.3314-3328, doi.org/10.1002/suco.202200525 [65] Verma, S., Boonsanong, V., Hoang, M., Hines, K.E., Dickerson, J.P., & Shah, C. (2022). Counterfactual explanati...
-
[85]
Wang, Y., Zhang, Z., Wang, Z., Wang, C., & Wu, C. (2024). Interpretable machine learning-based text classification method for construction quality defect reports. Journal of Building Engineering, 85, 109330, doi.org/10.1016/j.jobe.2024.109330 [86] Xiong, R., Song, Y., Li, H.m & Wang, Y. (2019). On-site video mining for construction hazard identification w...
arXiv 2024
Show all 16 references
-
[92]
Love, P. E. D., Matthews, J., & Ika, L. A. (2023). Fast-and-frugal heuristics: an exploration into building an adaptive toolbox to assess the uncertainty of rework. Production Planning & Control, 1–16, doi.org/10.1080/09537287.2023.2257178 [93] Love, P.E.D., Matthews, J., Ika,...
2023
- [99]
-
[114]
Nykänen, M., Puro, V., Tiikkaja, M., Kannisto, H., Lantto, E., Simpura, F., Uusitalo, J., Lukander, K., TRäsänen, T., Heikkilä, T., & Teperi, A-M. (2020). Implementing and evaluating novel safety training methods for construction sector workers: Results of a randomized control...
2020
-
[121]
O’Grady, T., Chong, H.-Y., & Morrison, G.M. (2021). A systematic review and meta-analysis of building automation systems. Building and Environment, 195, 107770, doi.org/10.1016/j.buildenv.2021.107770 [122] Pelletier, K., Wood, C., Calautit, J., & Wu, Y. (2023). The viability o...
2021
-
[129]
Williamson, J. (2021). Evidential pluralism and explainable AI. The Reasoner, 15(6), pp.45-58, Available at: https://blogs.kent.ac.uk/thereasoner/files/2021/11/TheReasoner-156.pdf, Accessed 3rd August 2024. [130] Russo, F., (2014). Mechanisms and the evidence hierarchy. Dipart...
2021
-
[137]
Sun, X., Wang, H., & Mei, S. (2024). Explainable highway performance degradation prediction model based on LSTM. Advanced Engineering Informatics, 61, 102539, doi.org/10.1016/j.aei.2024.102539 [138] Li, L., Liu, Z., Shen, J., Wang, F., Qi, W., & Jeon, S. (2023). A LightGBM-bas...
2024
-
[145]
Santoni de Sio, F., & van den Hoven, J. (2018). Meaningful human control over autonomous Systems: A philosophical account, Frontiers in Robotics and AI, 5, doi.org/10.3389/frobt.2018.00015 [146] Pan, Y., & Zhang, L. (2021). Roles of artificial intelligence in construction engi...
2018
-
[152]
Yu, J., Xu, Y., Xing, C., Zhou, J., Pan, P., & Yang, P. (2024). Nuclear containment damage detection and visualization positioning based on YOLOv5m-FFC. Automation in Construction, 161, 105357, doi.org/10.1016/j.autcon.2024.105357 [153] Editorial. (2023). Data sharing in the a...
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.