Pith. sign in

REVIEW 3 major objections 6 minor 16 references

Integrating Evidence into the Design of XAI and AI-based Decision Support Systems: A Means-End Framework for End-users in Construction

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper proposes a means-end framework in which construction AI explanations are designed from a ranked evidence hierarchy, aligned to users' knowledge goals.

desk verdict A competent synthesis that overpromises: the evidence hierarchy only ranks methodological strength, so the framework can't deliver the context-tailored explanations it claims. read the letter →

arxiv 2412.14209 v3 pith:B6BJISJX submitted 2024-12-17 cs.HC cs.AI

classification cs.HCcs.AI
keywords explainableartificialintelligencedecisionsupportsystemsconstructionevidencehierarchymeans-endframeworkmeaningfulhumanexplanationsevidentialpluralismmachinelearning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to establish that the reliability of an AI explanation depends on the quality of the evidence behind it, not just on the cleverness of the explanation technique. It proposes a means-end framework: pick the explanatory tools (means) only after stating what knowledge the user needs (ends), and grade the available evidence on a hierarchy adapted from medicine to construction's messy field conditions. If the framework is right, construction firms can specify to vendors what evidence an AI recommendation must rest on, regulators can audit explanations against a known standard, and users can avoid both blind trust and unjustified skepticism. The central novelty is treating evidence quality as a first-class design input for explainable AI, not an afterthought.

What carries the argument

The framework's load-bearing object is the adapted hierarchy of evidence (Table 3), a ten-level ranking from systematic reviews of longitudinal field studies at the top down to individual expert opinion at the bottom, used as a heuristic to weight evidence when designing model features, choosing XAI techniques, and evaluating explanations. It operates through the means-end structure: epistemic goals (understanding, justified belief, truth, insight) determine which XAI instrument is the right means, and the hierarchy determines which evidence may serve as that instrument's warrant. Evidential Pluralism supplies the causal criterion—both correlation and mechanism must be shown—so that explanations grounded only in statistical association are treated as weak.

What would settle it

Run a randomized experiment with construction site managers using a delay-risk DSS: give one group explanations sourced from high-hierarchy evidence (longitudinal field studies plus mechanism interviews) and another from low-hierarchy evidence (expert opinion or single case study). If the low-evidence group shows equal comprehension, trust calibration, and decision accuracy, the hierarchy's ranking does not carry the framework's weight.

Watch

Extended reading notes

Core claim

The central claim is that a DSS explanation is only meaningful—what the paper calls a Meaningful Human Explanation—when it is built on evidence whose strength has been explicitly ranked, and when the choice of XAI instrument is derived from the end-user's epistemic goal via instrumental rationality. The paper adapts a ten-level evidence hierarchy from medicine, re-ordered for construction's volatile, uncertain, complex, and ambiguous settings, and couples it with Evidential Pluralism: a causal claim is justified only when both a correlation and a mechanism linking cause and effect are demonstrated. The framework then maps association studies (e.g., historical project data) and mechanism studies (e.g., process-tracing interviews) onto specific XAI techniques—SHAP and LIME as associative evidence, counterfactual explanations as quasi-mechanistic levers—so that explanations answer 'why' and 'how', not merely 'what'.

Load-bearing premise

The framework rests on the assumption that ranking evidence quality on a ten-level, medicine-inspired hierarchy actually tracks how much an explanation improves a construction practitioner's understanding and decision—an assumption the paper states but does not test.

Editorial extensions

If this is right

  • Construction organizations can specify evidence requirements to DSS vendors instead of accepting black-box outputs.
  • Post-hoc tools like SHAP and LIME are repositioned as associative, not causal, and must be triangulated with mechanism studies.
  • A shared repository of use cases and open datasets becomes a prerequisite for applying the framework in practice.
  • The hierarchy gives regulators and auditors a concrete standard for judging whether an AI explanation is justified.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the framework holds, evidence-based design could spread beyond construction to other safety-critical engineering domains with similar VUCA conditions, such as mining or offshore operations.
  • A testable prediction emerges: user trust and decision accuracy should correlate with evidence level on the hierarchy, which could be tested by giving different user groups explanations sourced from different hierarchy levels.
  • The hierarchy may need to be task-specific rather than universal, since a longitudinal study feasible for safety may be impossible for bespoke one-off projects, making the ranking's applicability context-dependent.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a theoretical means–end framework for integrating evidence into the design of explainable AI (XAI) and AI-based decision support systems (DSSs), aimed at construction end-users. Developed through a narrative review spanning computer science, philosophy, and medicine, the framework combines epistemic normativity and instrumental rationality: evidence, generated through association and mechanism studies, is ranked via an adapted evidence hierarchy (Table 3) and used to design XAI instruments that should produce Meaningful Human Explanations (MHEs). The paper also provides a five-step illustrative application (Section 6.3) and a multi-dimensional evaluation scheme (Section 6.2). No empirical data are presented; the authors explicitly acknowledge this in Section 6.5, stating that the framework is conceptual and requires future validation.

Significance. If its central claim were made operational, the framework would fill a genuine gap: most construction XAI research is technical and model-centric, with little attention to how evidence supports explanations. The paper's strengths are its interdisciplinary synthesis, its adaptation of a medical evidence hierarchy to construction's VUCA realities, its explicit evaluation dimensions (cognitive validity, epistemic alignment, robustness, decision impact, normative justifiability), and its candid statement of limitations. The narrative review is transparent about its method, and the framework draws on external philosophical sources (Buchholz, Williamson, Salmon) rather than being self-referential. However, the central claim that the framework ensures evidence-based, context-tailored MHEs is currently untested and, more importantly, internally incomplete: the operational ranking instrument does not encode the relevance and utility dimensions that the paper repeatedly invokes.

major comments (3)
  1. [Section 5.2; Tables 1 and 3; Section 6.1; Section 6.3] The abstract and Section 6.1 state that the framework evaluates the 'strength, value, and utility' of evidence, but the only operational ranking instrument, Table 1 and its construction adaptation in Table 3, ranks evidence purely by study design and methodological rigor. There is no column, scoring rule, or procedure for relevance to a specific project, decision task, stakeholder role, or VUCA context. In Section 6.3, step 1 says evidence is 'weighted accordingly' but no weighting function is specified. Consequently, the hierarchy can recommend a Level 1 meta-analysis from another sector over a lower-level, mechanism-rich local field study that is more relevant and useful to the end-user. Since the central claim is that the framework produces MHEs 'tailored to users' knowledge needs and decision contexts,' the framework as presented cannot deliver that outcome. The authors should either add an explicit relevance/utility weighting procedure to the hierarchy or temper the claim that applying the hierarchy ensures context-tailored MHEs.
  2. [Section 4.1; Section 6.1; Section 6.5] The concept of MHE is defined only by desiderata ('intelligible, relevant, and actionable') and by three components drawn from earlier work, but the paper provides no method for eliciting end-users' epistemic ends in a concrete design process. Section 4.1 correctly states that the suitability of an XAI instrument depends on specifying the epistemic end before acquiring evidence and building the DSS, yet the framework gives no procedure for such specification. Section 6.5 acknowledges that the framework is abstract and hard to operationalize, but the paper's own wording in Section 6.1 claims that the framework 'will enable construction organizations to realize the benefits and business value of XAI and DSSs.' Without a stakeholder-analysis or requirements-elicitation step, the alignment between evidence, XAI instrument, and user context remains an assertion rather than a framework property. The authors should add an explicit elicitation and mapping step, or clearly reposition the contribution as a conceptual scaffold whose operationalization is future work.
  3. [Section 4, paragraph beginning 'A case in point'] The paper asserts that 'rework is not a risk but an uncertainty, which is probabilistically unmeasurable' and that machine learning techniques are therefore unable to accurately predict rework costs. This is presented as settled fact and used to dismiss a specific study (Mostofi et al.). No empirical evidence or citation is provided for this strong claim, which sits uneasily with the paper's own evidence-based rhetoric. If the claim is a philosophical or practical assumption, it should be labeled as such and justified; if it is an empirical claim, it needs supporting evidence. As written, it is an unsupported axiom used to motivate the framework and to criticize a published study, and it should be revised.
minor comments (6)
  1. [Section 5.2] Two subsections are both numbered 5.2: 'Hierarchy of Evidence' and 'Evidential Pluralism.' This numbering error should be corrected.
  2. [Section 4 and reference [90]] The text refers to 'Mostifi et al.' but the reference is to 'Mostofi et al.'; the spelling should be made consistent.
  3. [Reference [11]] The arXiv identifier for the generative XAI survey is given as 'arXiv:2014.09554'; this appears to be a typo for 'arXiv:2404.09554' and should be corrected.
  4. [References [36] and [45]] References [36] and [45] appear to be the same paper (Luo et al. 2024, IEEE Transactions on Engineering Management). Duplicate references should be removed or merged.
  5. [Section 6.4] The statement that 'when authors requested access to AI training data... the corresponding authors of papers repeatedly ignored these requests' is an empirical claim about the behavior of other researchers, with no citation or data. Either provide supporting evidence or remove the anecdotal assertion.
  6. [Figure 3 and surrounding text] Figure 3 uses the notation P(G|E), which suggests a probabilistic update of ground truth given evidence, while the text emphasizes causal-mechanical dependence and 'cause-effect determination.' Clarify how the probabilistic notation relates to the causal language, or adjust the notation to avoid conflating association with causation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the framework's normative and evidential core is taken from external philosophy and medicine sources, and the paper makes no fitted empirical claims that could reduce to its own inputs.

full rationale

The central components of the proposed means–end framework are explicitly grounded in external sources: Salmon's epistemic conception of explanation and Buchholz's means-end account ([24], [72]), Williamson's Evidential Pluralism ([28], [129]), and Famiglini et al.'s evidence hierarchy ([10]). The paper's construction-specific adaptation (Table 3) is offered as a design proposal motivated by the VUCA nature of construction, not as a consequence derived from the authors' prior work. Self-citations such as Love et al. [2] and [6] are used to describe the research context and to point to earlier XAI reviews and evaluation concepts; they do not supply the load-bearing warrant, which rests on the external philosophical and medical literature cited above. No parameters are fitted and no empirical quantity is predicted, so there is no fitted-input-called-prediction or self-definitional reduction. The paper's own limitation statement (Section 6.5) acknowledges that the framework is conceptual and not empirically validated, confirming that the contribution is a synthesis and a proposal rather than a derivation. Accordingly, no circular steps are identified.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

The framework rests on a small set of domain assumptions and one ad hoc ontological claim about rework; no free parameters are fitted because the paper is conceptual. The main assumptions are that construction is VUCA, that evidential pluralism is the right standard for causation, and that post-hoc XAI outputs are only associative. These are reasonable but not empirically established, and they inherit the unresolved debates in the philosophy of evidence.

assumptions (5)
  • domain assumption Construction projects operate in a VUCA (volatile, uncertain, complex, ambiguous) environment where randomized controlled trials and meta-analyses are rarely feasible.
    Used in Section 5.2 to justify adapting the medical evidence hierarchy for construction. If construction were not VUCA, the adapted hierarchy would lack its motivation.
  • domain assumption Establishing a causal claim requires both a correlation and a mechanism (Evidential Pluralism).
    Adopted from Williamson in Section 5.2 and Figure 4 as the normative basis for combining association and mechanism studies. All framework recommendations inherit this philosophical thesis.
  • domain assumption Post-hoc XAI methods such as LIME and SHAP provide associative evidence but cannot explain why an output occurred.
    Stated in Sections 3 and 5.2.1 and used to require supplementary causal or mechanism-based evidence. This is a widely held view in XAI research but not a proven theorem.
  • ad hoc to paper Rework in construction is an uncertainty that is probabilistically unmeasurable, so machine learning cannot accurately predict rework costs.
    Asserted in Section 4, supported by the authors' own prior work [91-93], and used to reject a specific published study. This is a strong domain-specific ontological claim that is contested in the literature.
  • domain assumption An explanation is successful if it enhances an individual's understanding (Salmon's epistemic conception).
    Adopted in Section 4.1 and underpins the definition of Meaningful Human Explanations. Alternative accounts of explanation are acknowledged but not adopted.
invented entities (1)
  • Meaningful Human Explanations (MHEs)
    purpose: A conceptual standard for explanations that are intelligible, relevant, and actionable, used as the target output of the proposed DSS design framework.
    The paper introduces MHEs as a design goal and proposes future metrics (explanation satisfaction, cognitive load, decision impact), but does not measure them here, so there is no independent falsifiable evidence yet.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Integrating Evidence into the Design of XAI and AI-based Decision Support Systems: A Means-End Framework for End-users in Construction." pith.science (2026). https://pith.science/paper/B6BJISJX

@misc{pith2026241214209,
  author       = {Pith},
  title        = {Pith review of: Integrating Evidence into the Design of XAI and AI-based Decision Support Systems: A Means-End Framework for End-users in Construction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B6BJISJX}},
  note         = {Machine review of arXiv:2412.14209}
}
read the original abstract

Explainable Artificial Intelligence seeks to make the reasoning processes of AI models transparent and interpretable, particularly in complex decision making environments. In the construction industry, where AI based decision support systems are increasingly adopted, limited attention has been paid to the integration of supporting evidence that underpins the reliability and accountability of AI generated outputs. The absence of such evidence undermines the validity of explanations and the trustworthiness of system recommendations. This paper addresses this gap by introducing a theoretical, evidence based means end framework developed through a narrative review. The framework offers an epistemic foundation for designing XAI enabled DSS that generate meaningful explanations tailored to users knowledge needs and decision contexts. It focuses on evaluating the strength, relevance, and utility of different types of evidence supporting AI generated explanations. While developed with construction professionals as primary end users, the framework is also applicable to developers, regulators, and project managers with varying epistemic goals.

Figures

Figures reproduced from arXiv: 2412.14209 by the authors.

Figure 1
Figure 1. Review process Integrating Evidence into the Design of XAI and AI-based Decision Support Systems • How can evidence effectively inform the design and development of XAI and DSS systems to ensure their generated outcomes are explainable and that the epistemic needs of an end-user [construction organization] are met? XAI Reviews XAI reviews: • Love et al. [6] • Love et al. [6] • Angelov et al. [14] • Arrieta et al. [1… view at source ↗
Figure 2
Figure 2. A means-end framework for XAI-based design of DSSs. • Association Studies • Mechanism Studies Evidence Decision MHE Support System Output XAI XAI • Explainable Decision • Interpretable • Association Studies • Mechanism Studies Epistemic Means (e.g., Observation, reasoning, experimentation, communication and modeling) Epistemic Goals (e.g., Understanding, justified belief, truth and insight) User-understanding Evalua… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

16 extracted references · 4 canonical work pages

  1. [2]

    Ding L-Y

    Love, P.E.D., Matthews, J., Fang, W., Porter, S., Luo, H., & L. Ding L-Y. (2024). Learning to comprehend and trust artificial intelligence outcomes: A conceptual explainable AI evaluation framework. IEEE Engineering Management Review, 52(1), pp. 230-247, doi.org/10.1109/EMR.2023.3342200. [3] Mosier, K.L., & Skitka, L.J. (1996). Human decision-makers and a...

  2. [9]

    Miller, T. (2019). Explanation in artificial intelligence: Insights from social science. Artificial Intelligence, 267, pp.1-38, doi.org/10.1016/j.artint.2018.07.007 [10] Famiglini, L., Campagner, A., Barandas, M., La Maida, G., Gallazzi, E., & Cabitza, F. (2024). Evidence-based XAI: An empirical approach to design more effective explainable decision-suppo...

  3. [16]

    That’s not the output I expected!

    Hossain, F., Hossain, R., & Hossain, E. (2021). Explainable Artificial Intelligence (XAI): An engineering perspective, arXiv.2101.03613 [cs.LG], doi.org/10.48550/arXiv.2101.03613 [17] Langer, M., Oster, D., Speith, T., Hermanns, H., Kastner, L., Schmidt, E., Sesing, A., & Baum, K. (2021). What do we want from explainable Artificial Intelligence (XAI)? – A...

  4. [43]

    Forest, F., Porta, H., Tuia, D., & Fink, O. (2024). From classification to segmentation with explainable AI: A study on crack detection and growth monitoring. Automation in Construction, 165, 105497, doi.org/10.1016/j.autcon.2024.105497 [44] Ghasemi, A., & Naser, M.Z. (2023). Tailoring 3D printed concrete through explainable artificial intelligence. Struc...

  5. [50]

    Snyder, H. (2019). Literature review as a research methodology: An overview and guidelines. Journal of Business Research, 104, pp.333-339, doi.org/10.1016/j.jbusres.2019.07.039 [51] Webster, J., & Watson, R.T. (2002). Analyzing the past to prepare for the future: Writing a literature review. Management Information Systems Quarterly, 26(2), pp.xiii-xxiii, ...

  6. [57]

    Why should I trust you?

    Moher, D., Liberati, A., Tetzlaff, J., Altman, D.G., & PRISMA Group. (2009). Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statement. PLoS Medicine 6(7): e1000097, doi.org/10.1371/journal.pmed.1000097. [58] Uttley, L., Quintana, D.S., Montgomery, P., Carroll, C., Page, M.J., Falzon, L., Sutton, A., & Moher, D. (2023). The ...

  7. [64]

    & Çiftçiogu, A.Ö

    Naser, M.Z. & Çiftçiogu, A.Ö. (2023). Causal discovery and inference for evaluating fire resistance of structural members through causal learning and domain knowledge. Structural Concrete, 24(3), pp.3314-3328, doi.org/10.1002/suco.202200525 [65] Verma, S., Boonsanong, V., Hoang, M., Hines, K.E., Dickerson, J.P., & Shah, C. (2022). Counterfactual explanati...

  8. [85]

    Wang, Y., Zhang, Z., Wang, Z., Wang, C., & Wu, C. (2024). Interpretable machine learning-based text classification method for construction quality defect reports. Journal of Building Engineering, 85, 109330, doi.org/10.1016/j.jobe.2024.109330 [86] Xiong, R., Song, Y., Li, H.m & Wang, Y. (2019). On-site video mining for construction hazard identification w...

Show all 16 references
  1. [92]

    Love, P. E. D., Matthews, J., & Ika, L. A. (2023). Fast-and-frugal heuristics: an exploration into building an adaptive toolbox to assess the uncertainty of rework. Production Planning & Control, 1–16, doi.org/10.1080/09537287.2023.2257178 [93] Love, P.E.D., Matthews, J., Ika,...

  2. [99]

    Joshi, S., Koyejo, O., Vijitbenjaronk, W., Kim, B., & Ghosh, J. (2019). Towards realistic individual recourse and actionable explanations in black-box decision-making systems. arXiv:1907.09615v1 [cs.LG], doi.org/10.48550/arXiv.1907.09615 [100] Karimi, A-H., Barthe, G., Schölko...

  3. [114]

    Nykänen, M., Puro, V., Tiikkaja, M., Kannisto, H., Lantto, E., Simpura, F., Uusitalo, J., Lukander, K., TRäsänen, T., Heikkilä, T., & Teperi, A-M. (2020). Implementing and evaluating novel safety training methods for construction sector workers: Results of a randomized control...

  4. [121]

    O’Grady, T., Chong, H.-Y., & Morrison, G.M. (2021). A systematic review and meta-analysis of building automation systems. Building and Environment, 195, 107770, doi.org/10.1016/j.buildenv.2021.107770 [122] Pelletier, K., Wood, C., Calautit, J., & Wu, Y. (2023). The viability o...

  5. [129]

    Williamson, J. (2021). Evidential pluralism and explainable AI. The Reasoner, 15(6), pp.45-58, Available at: https://blogs.kent.ac.uk/thereasoner/files/2021/11/TheReasoner-156.pdf, Accessed 3rd August 2024. [130] Russo, F., (2014). Mechanisms and the evidence hierarchy. Dipart...

  6. [137]

    Sun, X., Wang, H., & Mei, S. (2024). Explainable highway performance degradation prediction model based on LSTM. Advanced Engineering Informatics, 61, 102539, doi.org/10.1016/j.aei.2024.102539 [138] Li, L., Liu, Z., Shen, J., Wang, F., Qi, W., & Jeon, S. (2023). A LightGBM-bas...

  7. [145]

    Santoni de Sio, F., & van den Hoven, J. (2018). Meaningful human control over autonomous Systems: A philosophical account, Frontiers in Robotics and AI, 5, doi.org/10.3389/frobt.2018.00015 [146] Pan, Y., & Zhang, L. (2021). Roles of artificial intelligence in construction engi...

  8. [152]

    Yu, J., Xu, Y., Xing, C., Zhou, J., Pan, P., & Yang, P. (2024). Nuclear containment damage detection and visualization positioning based on YOLOv5m-FFC. Automation in Construction, 161, 105357, doi.org/10.1016/j.autcon.2024.105357 [153] Editorial. (2023). Data sharing in the a...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.