Pith. sign in

REVIEW 5 major objections 5 minor 1 cited by

Beyond Explainability: The Case for AI Validation

T0 review · 5 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper argues that validation of AI outputs—not explainability—should be the central regulatory pillar for opaque, high-stakes systems.

desk verdict A readable policy brief pushing validation over explainability, but the central concept is undefined and the novelty is mostly in packaging. read the letter →

arxiv 2505.21570 v1 pith:EADS4FIW submitted 2025-05-27 cs.CY cs.AI

classification cs.CYcs.AI
keywords AIgovernancevalidationexplainabilityalgorithmicaccountabilityregulationtrustworthyrisk-basedvalidity-explainabilitymatrix
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that requiring AI systems to be explainable is the wrong primary regulatory lever, and that regulators should instead require validation: evidence that outputs are reliable, consistent, and robust across conditions. It argues that validation is more practical, scalable, and risk-sensitive than explainability, especially for high-stakes systems whose internal processes are technically or economically infeasible to interpret. If the argument is right, AI governance would shift from opening black boxes to certifying their outputs through pre- and post-deployment testing, third-party audits, and liability incentives. The paper supports this shift with a two-axis typology of AI systems and a comparative reading of EU, US, UK, and Chinese regulation.

What carries the argument

The load-bearing device is the validity-explainability matrix, a two-by-two typology that classifies AI systems along a validity axis (valid versus non-valid) and an explainability axis (explainable versus opaque). The matrix does the argumentative work by isolating the valid-opaque quadrant—systems that are reliable and consistent but inscrutable—as the case where validation can govern without explanation, and the non-valid opaque quadrant as the case requiring pre-deployment audits, stress testing, or prohibition. It also frames the policy question as a trade-off between interpretability and output reliability, with regulation balancing incentives for both.

What would settle it

A concrete test would be to certify a set of high-stakes AI systems under the proposed validation regime and then track their live performance under distribution shift; if certified-valid systems fail at the same rate as uncertified ones on reliability, consistency, or robustness, validation lacks the objective predictive power the paper needs.

Watch

Extended reading notes

Core claim

The paper's central claim is that validation should become a central regulatory pillar for Artificial Knowledge systems, complementing or replacing explainability where interpretation is impractical. It defines validation as ensuring the reliability, consistency, and robustness of AI outputs, and argues this focus on outcomes is more feasible and scalable than a focus on interpretable processes. The paper introduces a four-quadrant classification—valid-explainable, pre-valid explainable, valid-opaque, and non-valid opaque—to show where validation can substitute for explainability and where opacity plus invalidity creates the highest social risk. It contends that existing regulatory instruments, from the EU AI Act to the FDA's Good Machine Learning Practices and China's validation-report requirements, already point toward validation as the operative standard. It concludes with a policy framework mandating pre- and post-deployment validation, independent auditing, harmonized standards, and liability incentives.

Load-bearing premise

The argument assumes that an AI system's validity—its reliability, consistency, and robustness—is a measurable, stable property that can be certified before and after deployment; the paper does not define how to measure it and concedes that validity in dynamic systems may not be fully assessable.

Editorial extensions

If this is right

  • Regulators would mandate pre-deployment and post-deployment validation for high-risk AI systems rather than requiring explanations as the default compliance path.
  • Independent third-party bodies would evaluate high-risk systems using standardized datasets, fairness metrics, and computational infrastructure, with public support for small and medium enterprises.
  • Liability regimes and certification schemes would make developers internalize the costs of unreliable outputs, incentivizing robust testing even when explainability remains out of reach.
  • Explainability would not be abandoned; it would be required where its costs are reasonable and its benefits—fairness, accountability, human oversight—are significant.
  • Non-valid opaque systems, which are both unreliable and inscrutable, would face the strongest controls, including prohibition in cases where public safety and equity are at stake.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If validation becomes the regulatory standard, the economic center of gravity in AI governance would shift from interpretability tooling to testing infrastructure, benchmark datasets, and certification bodies.
  • An implication the paper leaves implicit is that validity must be defined per domain and context, which means regulatory standards would need to specify measurable thresholds for reliability, consistency, and robustness.
  • The typology suggests a testable extension: validation certificates issued at deployment could be evaluated against live performance under distribution shift, turning the paper's policy proposal into an empirical research program.
  • For dynamic and self-updating systems, point-in-time validation would likely need to become continuous monitoring, a limitation the paper itself flags in its final sentence.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. This paper argues that AI governance should shift its primary regulatory focus from explainability to validation, defined as ensuring the reliability, consistency, and robustness of AI outputs. It introduces a two-axis typology classifying Artificial Knowledge systems into four quadrants (valid-explainable, valid-opaque, pre-valid explainable, non-valid opaque), reviews regulatory approaches in the EU, US, UK, and China, and proposes a policy framework centered on pre- and post-deployment validation, third-party audits, harmonized standards, and liability incentives. The paper concludes by identifying future research needs, including how to assess validity in dynamic systems.

Significance. The paper identifies a genuine limitation of explainability-centric governance and proposes a complementary, outcome-oriented regulatory lens. A rigorous development of the validation concept could be a valuable contribution to AI governance debates, especially given the practical difficulties of interpreting opaque models. The paper also usefully draws attention to existing validation-related provisions in various regulatory frameworks. However, its significance is presently constrained by the lack of an operational definition of validity and the absence of empirical support for its comparative cost and scalability claims; the contribution is conceptual and programmatic rather than implementable as stated.

major comments (5)
  1. [Section 3 and Section 5] The central concept of 'validity' is never operationally defined. Section 3 defines validation only via synonyms ("reliability, consistency, and robustness"), while Section 5 expands it to "functional, empirical, and normative dimensions" without specifying how these dimensions are measured, aggregated, or thresholded. The paper's own final sentence concedes that it is unknown "to what extent validity can truly be assessed" in dynamic systems. This is load-bearing because the entire regulatory proposal—pre-deployment mandates, third-party audits, liability regimes—requires an enforceable, auditable standard; without one, the framework cannot be implemented. The authors should provide a concrete operationalization, including what counts as ground truth, evaluation protocols, and pass/fail thresholds, or revise the claim that validation is a 'more practical and scalable' alternative.
  2. [Section 3, refs [25–27]] The analogy to black-box software testing assumes the existence of expected outputs against which a system is checked. In high-stakes domains such as recidivism prediction or clinical diagnosis, outcomes are delayed, contested, or unobserved, so no such oracle exists. The paper does not identify what counts as ground truth for validation in these domains, nor how to handle feedback loops and distribution shift. This gap is central because the paper's argument that validation can replace explainability in high-stakes contexts depends on the feasibility of such validation. The authors should address the oracle problem directly and discuss how validation would work where ground truth is unavailable.
  3. [Table 1 and Section 4] The validity–explainability matrix is definitionally circular: each quadrant's properties follow trivially from the two axes, so the governance implications (e.g., "valid-opaque requires accountability frameworks," "non-valid opaque requires prohibition") are built into the categories rather than derived from evidence. The paper does not provide decision-relevant criteria for classifying an actual system into a quadrant. For the typology to support regulatory recommendations, the authors need to give measurable indicators for both axes and show how the quadrants differ empirically, not just by definition.
  4. [Sections 3 and 4.4] The paper repeatedly asserts that validation is "more practical, scalable, and cost-effective" than explainability (e.g., in Section 3), but no empirical evidence or cost analysis is provided to support these comparative claims. Likewise, Section 4.4 states that non-valid opaque systems lead to litigation, reputational damage, and hindered innovation, citing sources that are not specific to AI (e.g., ref. [88] is an economics of information paper). The authors should either substantiate these empirical claims or temper them as hypotheses requiring further research.
  5. [Section 4, comparative legal analysis] The survey of EU, US, UK, and China regulation is selective and at times inaccurate. For example, the paper cites the EU AI Act as mandating validation (articles 10, 13, 16) without noting that the Act's final text differs from the 2021 proposal cited, and it does not discuss the Act's risk-tiered structure in sufficient context. The cancellation of Executive Order 14960 is noted, but the paper still leans on it as evidence of U.S. emphasis on validity, while the OMB memo M-25-21 is described as requiring "continuous validation" without analyzing its actual provisions. A more systematic comparison—ideally with a table of provisions per jurisdiction and their status—is needed to support the claim that validation is already 'central' to AI governance.
minor comments (5)
  1. [Section 4] The paragraph beginning "Balancing validity and explainability in AI systems is challenging" appears twice, with slight variation, which suggests an editorial duplication that should be removed.
  2. [Various] There are frequent typographical and formatting issues, including "explainabl" (Section 4), missing spaces after commas in references, and inconsistent use of quotation marks around terms like "white box paradox." A thorough proofread is needed.
  3. [Table 1 and Section 4.2] The table uses the label "Pre-Valid Explainable" for the non-valid explainable quadrant, while the text in Section 4.2 uses the same term; however, the table's row heading is "Non-Valid," creating terminological inconsistency. The authors should harmonize these labels.
  4. [Section 1] The term "Artificial Knowledge (AK)" is introduced without a precise definition, leaving it unclear whether it encompasses all AI systems or only opaque knowledge-generating ones. A working definition early in the paper would improve clarity.
  5. [Section 5] The final sentence of the paper concedes that the assessability of validity in dynamic systems is an open question; this limitation is important enough to be acknowledged in the introduction or abstract so that readers are appropriately cautioned from the start.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's policy argument and typology are self-contained; the sole self-citation is provenance, not load-bearing.

full rationale

This is a legal and policy argument, not a derivation with equations or predictions, so the usual circularity patterns do not apply. The central claim that validation is a more practical, scalable, and risk-sensitive alternative to explainability is argued from external regulatory materials (EU AI Act, U.S. Executive Orders, China's AI regulations, FDA guidance), independent XAI critiques, and general software-testing literature; it is not reduced to the paper's own definitions. The validity-explainability matrix is a classification whose quadrant properties follow by definition from the two axes, but the paper does not present this taxonomy as an empirical discovery or use it to generate predictions; it is explicitly a 'structured framework' for organizing governance trade-offs. The Author's Note states that the article is 'based on the theoretical framework first introduced' in the authors' forthcoming work, but this is provenance rather than load-bearing: the typology is stated and defined within this paper, and the cited work is not invoked to license any specific conclusion or to forbid alternatives. The final sentence concedes that 'it is essential to explore the limitations of validation in dynamic systems and understand to what extent validity can truly be assessed'; that is a substantive limitation and a correctness risk about the operationalizability of 'validity,' not evidence that any result reduces to its own inputs. No step meets the required bar of exhibiting a specific reduction by construction, a fitted parameter renamed as prediction, or a load-bearing self-citation chain.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters or invented entities are present because the paper contains no quantitative model. It does rest on several unproved domain assumptions: that explainability is often infeasible, that validation is measurable and scalable, that selected legal provisions are representative, and that the four-cell taxonomy is exhaustive. These assumptions are stated or implied in Sections 2, 3, and 4, and one is explicitly qualified in Section 5.

assumptions (4)
  • domain assumption Explainability is technically or economically infeasible for many high-performing AI systems, making validation the more practical regulatory route.
    Section 2 and Section 4 rely on this to motivate the shift; the paper cites refs. 16, 17, 23, and 24 but does not establish it for all high-stakes contexts.
  • ad hoc to paper Validation, defined as ensuring reliability, consistency, and robustness of outputs, can serve as an objective, scalable, and risk-sensitive regulatory pillar.
    Section 3 states this as the central premise, but no measurement standard or evidence base is provided, and Section 5 concedes uncertainty about whether validity can be assessed.
  • domain assumption The selected provisions of the EU AI Act, US executive orders, OMB guidance, UK White Paper, and Chinese regulations are representative and correctly mapped to 'validation.'
    Section 4's comparative analysis relies on selective legal interpretation without a systematic method or coding protocol.
  • ad hoc to paper The four-cell validity/explainability matrix exhaustively classifies AK systems and supports the stated governance implications for each cell.
    Table 1 defines quadrants by the two axes; the normative consequences, such as prohibition for non-valid opaque systems, follow from definitions rather than empirical evidence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond Explainability: The Case for AI Validation." pith.science (2026). https://pith.science/paper/EADS4FIW

@misc{pith2026250521570,
  author       = {Pith},
  title        = {Pith review of: Beyond Explainability: The Case for AI Validation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EADS4FIW}},
  note         = {Machine review of arXiv:2505.21570}
}
read the original abstract

Artificial Knowledge (AK) systems are transforming decision-making across critical domains such as healthcare, finance, and criminal justice. However, their growing opacity presents governance challenges that current regulatory approaches, focused predominantly on explainability, fail to address adequately. This article argues for a shift toward validation as a central regulatory pillar. Validation, ensuring the reliability, consistency, and robustness of AI outputs, offers a more practical, scalable, and risk-sensitive alternative to explainability, particularly in high-stakes contexts where interpretability may be technically or economically unfeasible. We introduce a typology based on two axes, validity and explainability, classifying AK systems into four categories and exposing the trade-offs between interpretability and output reliability. Drawing on comparative analysis of regulatory approaches in the EU, US, UK, and China, we show how validation can enhance societal trust, fairness, and safety even where explainability is limited. We propose a forward-looking policy framework centered on pre- and post-deployment validation, third-party auditing, harmonized standards, and liability incentives. This framework balances innovation with accountability and provides a governance roadmap for responsibly integrating opaque, high-performing AK systems into society.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation

    cs.CR 2025-07 reject novelty 2.0 of 10

    The paper is a literature review that classifies privacy-preserving AI techniques in dataspaces using a qualitative taxonomy of privacy, performance, and compliance ratings.

Reference graph

Works this paper leans on

91 extracted references · 77 canonical work pages · cited by 1 Pith paper

  1. [88]

    The contributions of the economics of information to twentieth cen- tury economics

    Joseph E Stiglitz. The contributions of the economics of information to twentieth cen- tury economics. The quarterly journal of economics, 115(4):1441–1478, 2000

  2. [1]

    AIin health and medicine

    Pranav Rajpurkar, Emma Chen, Oishi Banerjee, and Eric J Topol. AIin health and medicine. Nature medicine , 28(1):31– 38, 2022

  3. [2]

    Deep neural net- works improve radiologists’ performance in breast cancer screening

    Nan Wu, Jason Phang, Jungkyu Park, Yiqiu Shen, Zhe Huang, Masha Zorin, Stanisław Jastrzębski, Thibault Févry, Joe Katsnel- son, Eric Kim, et al. Deep neural net- works improve radiologists’ performance in breast cancer screening. IEEE transactions on medical imaging , 39(4):1184–1194, 2019

  4. [3]

    International evaluation of an AIsys- tem for breast cancer screening

    Scott Mayer McKinney, Marcin Sieniek, Varun Godbole, Jonathan Godwin, Natasha Antropova, Hutan Ashrafian, Trevor Back, Mary Chesus, Greg S Corrado, Ara Darzi, et al. International evaluation of an AIsys- tem for breast cancer screening. Nature, 577(7788):89–94, 2020

  5. [4]

    Bias in medical ai: Impli- cations for clinical decision-making

    James L Cross, Michael A Choma, and John A Onofrey. Bias in medical ai: Impli- cations for clinical decision-making. PLOS Digital Health, 3(11):e0000651, 2024. 6

  6. [5]

    Algorithmic hiring in prac- tice: Recruiter and hr professional’s per- spectives on AIuse in hiring

    Lan Li, Tina Lassiter, Joohee Oh, and Min Kyung Lee. Algorithmic hiring in prac- tice: Recruiter and hr professional’s per- spectives on AIuse in hiring. In Proceed- ings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society , AIES ’21, page 166–176, New York, NY, USA, 2021. Asso- ciation for Computing Machinery

  7. [6]

    AIhiring bias: Everything you need to know

    G Lawton. AIhiring bias: Everything you need to know. techtarget, 2022

  8. [8]

    Using artificial intelli- gence to address criminal justice needs

    Christopher Rigano. Using artificial intelli- gence to address criminal justice needs. Na- tional Institute of Justice Journal , 280(1- 10):17, 2019

Show all 91 references
  1. [9]

    Criminal justice, artificial in- telligence systems, and human rights

    Aleš Završnik. Criminal justice, artificial in- telligence systems, and human rights. In ERA forum , volume 20(4), pages 567–583. Springer, 2020

  2. [10]

    Arti- ficial intelligence in the criminal justice sys- tem: leading trends and possibilities

    Tatyana Sushina and Andrew Sobenin. Arti- ficial intelligence in the criminal justice sys- tem: leading trends and possibilities. In 6th International Conference on Social, eco- nomic, and academic leadership (ICSEAL- 6-2019), pages 432–437. Atlantis Press, 2020

  3. [11]

    Recommendation of the council on artificial intelligence

    OECD. Recommendation of the council on artificial intelligence. https://tinyurl.com/dkkjef5j, 2019. OECD/LEGAL/0449, at 7–8 & sec. 1.3

  4. [12]

    Proposal for a regulation of the european parliament and of the council laying down har- monized rules on artificial intelligence

    European Commission. Proposal for a regulation of the european parliament and of the council laying down har- monized rules on artificial intelligence. https://tinyurl.com/vmyaxp4d, 2021. COM (2021) 206 final, Apr. 21, 2021, articles 13, 25

  5. [13]

    Reg- ulation (eu) 2016/679 on the protection of natural persons with regard to the processing of personal data (gdpr)

    European Parliament and Council. Reg- ulation (eu) 2016/679 on the protection of natural persons with regard to the processing of personal data (gdpr). https://tinyurl.com/37z3c5hp, 2016. 2016 O.J. (L 119) 1, 46, art. 22

  6. [14]

    Explain- able AIis responsible ai: How explainabil- ity creates trustworthy and socially respon- sible artificial intelligence

    Stephanie Baker and Wei Xiang. Explain- able AIis responsible ai: How explainabil- ity creates trustworthy and socially respon- sible artificial intelligence. arXiv preprint arXiv:2312.01555, 2023

  7. [15]

    Unexplainability and incomprehensibility of AI

    Roman V Yampolskiy. Unexplainability and incomprehensibility of AI. Journal of Artificial Intelligence and Consciousness , 7(02):277–291, 2020

  8. [16]

    The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery

    Zachary C Lipton. The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery. Queue, 16(3):31–57, 2018

  9. [17]

    The case against explainability

    Hofit Wasserman Rozen, Niva Elkin-Koren, and Ran Gilad-Bachrach. The case against explainability. arXiv preprint arXiv:2305.12167, 2023

  10. [18]

    Monitoring reasoning models for misbehavior and the risks of promoting obfuscation

    Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi. Monitoring reasoning models for misbehavior and the risks of promoting obfuscation. arXiv preprint arXiv:2503.11926, 2025

  11. [19]

    explanations

    David Martens, Galit Shmueli, Theodoros Evgeniou, Kevin Bauer, Christian Janiesch, Stefan Feuerriegel, Sebastian Gabel, Sofie Goethals, Travis Greene, Nadja Klein, et al. Beware of" explanations" of AI. arXiv preprint arXiv:2504.06791, 2025

  12. [20]

    Painting the black box white: experimental findings from applying XAI to an ECG reading setting

    Federico Cabitza, Andrea Campagner, Chiara Natali, Enea Parimbelli, Luca Ronzio, and Matteo Cameli. Painting the black box white: experimental findings from applying XAI to an ECG reading setting. Machine Learning and Knowledge Extrac- tion, 5(1):269–286, 2023. 7

  13. [21]

    Explana- tions considered harmful: the impact of mis- leading explanations on accuracy in hybrid human-AI decision making

    Federico Cabitza, Caterina Fregosi, Andrea Campagner, and Chiara Natali. Explana- tions considered harmful: the impact of mis- leading explanations on accuracy in hybrid human-AI decision making. In World con- ference on explainable artificial intelligence , pages 255–269. Spri...

  14. [22]

    Stop explaining black box machine learning models for high stakes de- cisions and use interpretable models instead

    Cynthia Rudin. Stop explaining black box machine learning models for high stakes de- cisions and use interpretable models instead. Nature machine intelligence , 1(5):206–215, 2019

  15. [23]

    Lost in translation: the limits of explainability in AI

    Hofit Wasserman-Rozen, Ran Gilad- Bachrach, and Niva Elkin-Koren. Lost in translation: the limits of explainability in AI. Cardozo Arts & Ent. LJ , 42:391, 2024

  16. [24]

    The false hope of cur- rent approaches to explainable artificial in- telligence in health care

    Marzyeh Ghassemi, Luke Oakden-Rayner, and Andrew L Beam. The false hope of cur- rent approaches to explainable artificial in- telligence in health care. The Lancet Digital Health, 3(11):e745–e750, 2021

  17. [25]

    Types of penetration testing: Black box, white box & grey box, 2024

    Mark Nicholls. Types of penetration testing: Black box, white box & grey box, 2024

  18. [26]

    Black box and white box testing techniques- a literature review

    Srinivas Nidhra and Jagruthi Dondeti. Black box and white box testing techniques- a literature review. International Jour- nal of Embedded Systems and Applications (IJESA), 2(2):29–50, 2012

  19. [27]

    A comparative study of white box, black box and grey box testing techniques

    Mohd Ehmer Khan and Farmeena Khan. A comparative study of white box, black box and grey box testing techniques. Interna- tional Journal of Advanced Computer Sci- ence and Applications , 3(6), 2012

  20. [28]

    Proposal for a regulation of the european parliament and of the council laying down har- monized rules on artificial intelligence

    European Commission. Proposal for a regulation of the european parliament and of the council laying down har- monized rules on artificial intelligence. https://tinyurl.com/vmyaxp4d, 2021. COM (2021) 206 final, Apr. 21, 2021

  21. [29]

    6580, 117th cong., 2022

    Algorithmic accountability act of 2022, h.r. 6580, 117th cong., 2022

  22. [30]

    Fairness and bias in algorithmic hiring: A multidisciplinary survey

    Alessandro Fabris, Nina Baranowska, Matthew J Dennis, David Graus, Philipp Hacker, Jorge Saldivar, Fred- erik Zuiderveen Borgesius, and Asia J Biega. Fairness and bias in algorithmic hiring: A multidisciplinary survey. ACM Transactions on Intelligent Systems and Technology, 16...

  23. [31]

    Fairness, AI & recruit- ment

    Carlotta Rigotti and Eduard Fosch- Villaronga. Fairness, AI & recruit- ment. Computer Law & Security Review , 53:105966, 2024

  24. [32]

    Systematic literature review of validation methods for AI systems

    Lalli Myllyaho, Mikko Raatikainen, Tomi Männistö, Tommi Mikkonen, and Jukka K Nurminen. Systematic literature review of validation methods for AI systems. arXiv preprint arXiv:2107.12190, 2021

  25. [33]

    Food & Drug Admin

    U.S. Food & Drug Admin. (FDA). AI and machine learning (AI/ML)-enabled medical devices. https://tinyurl.com/24kpf6zy, 2024

  26. [34]

    Recommendations for designing for simplicity and efficiency: Azure well- architected framework reliability, 2023

    Microsoft. Recommendations for designing for simplicity and efficiency: Azure well- architected framework reliability, 2023. De- cember 2, 2023

  27. [35]

    Afford- able system operational effectiveness (asoe) model

    Defense Acquisition University. Afford- able system operational effectiveness (asoe) model. https://www.dau.edu/

  28. [36]

    Guidance on model risk management, sr 11-7

    Board of Governors of the Fed- eral Reserve System. Guidance on model risk management, sr 11-7. https://tinyurl.com/kcbpyxk4, April 2011

  29. [37]

    Provi- sions on the management of algorithmic rec- ommendations in internet information ser- vices, 2022

    Cyberspace Administration of China. Provi- sions on the management of algorithmic rec- ommendations in internet information ser- vices, 2022. Promulgated Jan. 4, 2022; Ef- fective Mar. 1, 2022, arts. 7–8

  30. [38]

    Mea- sures for the management of generative ar- tificial intelligence services (draft for com- ment), April 2023

    Cyberspace Administration of China. Mea- sures for the management of generative ar- tificial intelligence services (draft for com- ment), April 2023. Art. 17. 8

  31. [39]

    Food and Drug Administration

    U.S. Food and Drug Administration. Good machine learning practice for medical de- vice development: Guiding principles. https://tinyurl.com/yrnfprct, October

  32. [40]

    A comparative look at various countries’ legal regimes governing automated vehicles

    Brittany Eastman, Shay Collins, Ryan Jones, JJ Martin, Marjory S Blumenthal, and Karlyn D Stanley. A comparative look at various countries’ legal regimes governing automated vehicles. JL & Mobility , page 1, 2023

  33. [41]

    Vali- dation of automated and autonomous vehi- cles

    Christof Ebert and Michael Weyrich. Vali- dation of automated and autonomous vehi- cles. ATZelectronics worldwide, 14(9):26–31, 2019

  34. [42]

    Au- tonomous vehicle technology: A guide for policymakers

    James M Anderson, Kalra Nidhi, Karlyn D Stanley, Paul Sorensen, Constantine Sama- ras, and Oluwatobi A Oluwatola. Au- tonomous vehicle technology: A guide for policymakers. Rand Corporation, 2016

  35. [43]

    Identifying main drivers on inven- tory using regression analysis

    Steffi Hoppenheit and Willibald A Günth- ner. Identifying main drivers on inven- tory using regression analysis. In Oper- ational Excellence in Logistics and Supply Chains: Optimization Methods, Data-driven Approaches and Security Insights. Proceed- ings of the Hamburg Internati...

  36. [44]

    Minoxi- dil: mechanisms of action on hair growth

    AG Messenger and J Rundegren. Minoxi- dil: mechanisms of action on hair growth. British journal of dermatology , 150(2):186– 194, 2004

  37. [45]

    Arslan, and D

    Emrullah Şahin, N.N. Arslan, and D. Özdemir. Unlocking the black box: An in-depth review on interpretability, explainability, and reliability in deep learn- ing. Neural Computing and Applications , 37:859–965, 2025

  38. [46]

    Evalua- tion of post-hoc interpretability methods in time-series classification

    Hugues Turbé, Mina Bjelogrlic, Christian Lovis, and Gianmarco Mengaldo. Evalua- tion of post-hoc interpretability methods in time-series classification. Nature Machine Intelligence, 5(3):250–260, 2023

  39. [47]

    The role of causality in explain- able artificial intelligence

    Gianluca Carloni, Andrea Berti, and Sara Colantonio. The role of causality in explain- able artificial intelligence. Wiley Interdisci- plinary Reviews: Data Mining and Knowl- edge Discovery, 15(2):e70015, 2025

  40. [48]

    A review of explainable artifi- cial intelligence in supply chain manage- ment using neurosymbolic approaches

    Edward Elson Kosasih, Emmanuel Pa- padakis, George Baryannis, and Alexandra Brintrup. A review of explainable artifi- cial intelligence in supply chain manage- ment using neurosymbolic approaches. In- ternational Journal of Production Research , 62(4):1510–1540, 2024

  41. [49]

    Ai-driven financial risk manage- ment systems: Enhancing predictive capa- bilities and operational efficiency

    Qi Shen. Ai-driven financial risk manage- ment systems: Enhancing predictive capa- bilities and operational efficiency. Applied and Computational Engineering , 69:134– 139, 2024

  42. [50]

    Not all AI health tools with regulatory authorization are clinically validated

    Sammy Chouffani El Fassi, Adonis Ab- dullah, Ying Fang, Sarabesh Natarajan, Awab Bin Masroor, Naya Kayali, Simran Prakash, and Gail E Henderson. Not all AI health tools with regulatory authorization are clinically validated. Nature Medicine , 30(10):2718–2720, 2024

  43. [51]

    Framework convention on artificial intelligence and human rights, democracy and the rule of law, 2024

    Council of Europe. Framework convention on artificial intelligence and human rights, democracy and the rule of law, 2024

  44. [52]

    Secretary of State for Science, Inno- vation & Technology

    U.K. Secretary of State for Science, Inno- vation & Technology. A pro-innovation ap- proach to AI regulation, March 2023

  45. [53]

    Executive order no

    Joseph Biden. Executive order no. 14,960 on the safe, secure, and trustworthy de- velopment and use of artificial intelligence. https://tinyurl.com/49jfr2vz, Novem- ber 2023. 88 Fed. Reg. 74,899, Secs. 2(a), 4.1(a), 4.2, 4.3

  46. [54]

    Donald J. Trump. Executive order 14179: Removing barriers to ameri- can leadership in artificial intelligence. https://tinyurl.com/4sxmyfsx/, Jan- uary 2025. 9

  47. [55]

    M- 25-21: Accelerating federal use of AI through innovation, governance, and pub- lic trust

    Office of Management and Budget. M- 25-21: Accelerating federal use of AI through innovation, governance, and pub- lic trust. https://tinyurl.com/x2w3w6sb, April 2025

  48. [56]

    Donald J. Trump. Promoting the use of trustworthy artificial intel- ligence in the federal government. https://tinyurl.com/5fyu3z28, De- cember 2020

  49. [57]

    Ab 331 bill, section 22756.3(a)(3), 2023

    California. Ab 331 bill, section 22756.3(a)(3), 2023

  50. [58]

    NIST AI 600-1 artificial intelli- gence risk management framework: Genera- tive artificial intelligence profile, 2024

    National Institute of Standards and Tech- nology. NIST AI 600-1 artificial intelli- gence risk management framework: Genera- tive artificial intelligence profile, 2024. NIST Trustworthy and Responsible AI

  51. [59]

    Partially in- terpretable models with guarantees on cov- erage and accuracy

    Nave Frost, Zachary Lipton, Yishay Man- sour, and Michal Moshkovitz. Partially in- terpretable models with guarantees on cov- erage and accuracy. In International confer- ence on algorithmic learning theory , pages 590–613. PMLR, 2024

  52. [60]

    Shap and lime: an evaluation of discriminative power in credit risk

    Alex Gramegna and Paolo Giudici. Shap and lime: an evaluation of discriminative power in credit risk. Frontiers in Artificial Intelligence, 4:752558, 2021

  53. [61]

    Fooling lime and shap: Adversarial attacks on post hoc explanation methods

    Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. Fooling lime and shap: Adversarial attacks on post hoc explanation methods. In Pro- ceedings of the AAAI/ACM Conference on AI, Ethics, and Society , pages 180–186, 2020

  54. [62]

    To- ward trustworthy artificial intelligence (tai) in the context of explainability and robust- ness

    Bhanu Chander, Chinju John, Lekha War- rier, and Kumaravelan Gopalakrishnan. To- ward trustworthy artificial intelligence (tai) in the context of explainability and robust- ness. ACM Computing Surveys , 57(6):1–49, 2025

  55. [63]

    Experimental regulations and regulatory sandboxes: Law without or- der? Law and method , 2021, 2021

    Sofia Ranchordás. Experimental regulations and regulatory sandboxes: Law without or- der? Law and method , 2021, 2021

  56. [64]

    Manipulating and measuring model interpretability

    Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wort- man Wortman Vaughan, and Hanna Wal- lach. Manipulating and measuring model interpretability. In Proceedings of the 2021 CHI conference on human factors in com- puting systems , pages 1–52, 2021

  57. [65]

    Interpreting interpretability: understanding data scien- tists’ use of interpretability tools for ma- chine learning

    Harmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana, Hanna Wallach, and Jennifer Wortman Vaughan. Interpreting interpretability: understanding data scien- tists’ use of interpretability tools for ma- chine learning. In Proceedings of the 2020 CHI conference on human fa...

  58. [66]

    Artificial neural network classification of high dimensional data with novel opti- mization approach of dimension reduction

    Rabia Aziz, CK Verma, and Namita Srivas- tava. Artificial neural network classification of high dimensional data with novel opti- mization approach of dimension reduction. Annals of Data Science , 5(4):615–635, 2018

  59. [67]

    Verma, and Namita Sri- vastava

    Rabia Aziz, C.K. Verma, and Namita Sri- vastava. Artificial neural network classifica- tion of high-dimensional data with novel op- timization approach of dimension reduction. Annals of Data Science , 5:615, 2018

  60. [68]

    Thinking responsibly about responsible ai and ‘the dark side’of ai

    Patrick Mikalef, Kieran Conboy, Jenny Eriksson Lundström, and Aleš Popovič. Thinking responsibly about responsible ai and ‘the dark side’of ai. European Journal of Information Systems , 31(3):257–268, 2022

  61. [69]

    The dangers of human- like bias in machine-learning algorithms

    Daniel James Fuchs. The dangers of human- like bias in machine-learning algorithms. Missouri S&T’s Peer to Peer , 2(1):1, 2018

  62. [70]

    The black box society: The secret algorithms that control money and in- formation

    Frank Pasquale. The black box society: The secret algorithms that control money and in- formation. Harvard University Press, 2015

  63. [71]

    Towards the certification of ai-based systems

    Philipp Denzel, Stefan Brunner, Yann Bil- leter, Oliver Forster, Carmen Frischknecht- Gruber, Monika Reif, Frank-Peter Schilling, Joanna Weng, Ricardo Chavarriaga, Amin Amini, et al. Towards the certification of ai-based systems. In 2024 11th IEEE Swiss 10 Conference on Data Sc...

  64. [72]

    Report on the safety and liability implications of artificial intelli- gence, the internet of things and robotics, February 2020

    European Commission. Report on the safety and liability implications of artificial intelli- gence, the internet of things and robotics, February 2020. COM(2020) 64 final, Sec- tion 3

  65. [73]

    Report on li- ability for AI and other emerging technolo- gies (new technologies formation), Novem- ber 2019

    European Commission, Expert Group on Li- ability and New Technologies. Report on li- ability for AI and other emerging technolo- gies (new technologies formation), Novem- ber 2019. COM (2019), pp. 27–29

  66. [74]

    The expert group’s report on liability for ar- tificial intelligence and other emerging digi- tal technologies: a critical assessment

    Andrea Bertolini and Francesca Episcopo. The expert group’s report on liability for ar- tificial intelligence and other emerging digi- tal technologies: a critical assessment. Euro- pean Journal of Risk Regulation , 12(3):644– 659, 2021

  67. [75]

    Identifying unreliable predictions in clinical risk models

    Paul D Myers, Kenney Ng, Kristen Sever- son, Uri Kartoun, Wangzhi Dai, Wei Huang, Frederick A Anderson, and Collin M Stultz. Identifying unreliable predictions in clinical risk models. NPJ digital medicine , 3(1):8, 2020

  68. [76]

    Engineer- ing trustworthy AI: A developer guide for empirical risk minimization

    Diana Pfau and Alexander Jung. Engineer- ing trustworthy AI: A developer guide for empirical risk minimization. arXiv preprint arXiv:2410.19361, 2024

  69. [77]

    AI failures: A review of underlying issues

    Debarag Narayan Banerjee and Sasanka Sekhar Chanda. AI failures: A review of underlying issues. arXiv preprint arXiv:2008.04073 , 2020

  70. [78]

    Dakka, T.V

    M.A. Dakka, T.V. Nguyen, J.M.M. Hall, S.M. Diakiw, M. VerMilyea, R. Linke, M. Perugini, and D. Perugini. Auto- mated detection of poor-quality data: case studies in healthcare. Scientific Reports , 11(1):18005, 2021

  71. [79]

    Factors for accelerating the development speed in systems of artificial intelligence

    Akshitha Kancharla and Akhil Pannala. Factors for accelerating the development speed in systems of artificial intelligence. Master’s thesis, Department of Software En- gineering, 2019

  72. [80]

    A blueprint for auditing gener- ative AI

    Jakob Mokander, Justin Curl, and Mihir Kshirsagar. A blueprint for auditing gener- ative AI. arXiv preprint arXiv:2407.05338 , 2024

  73. [81]

    AI-generated synthetic data for stress testing financial systems: A ma- chine learning approach to scenario analysis and risk management

    Praveen Sivathapandi, Debasish Paul, and Akila Selvaraj. AI-generated synthetic data for stress testing financial systems: A ma- chine learning approach to scenario analysis and risk management. Journal of Artificial Intelligence Research, 2:246–287, Jun. 2022

  74. [82]

    The implica- tions of AIfor criminal justice: Key take- aways from a convening of leading stake- holders

    Council on Criminal Justice. The implica- tions of AIfor criminal justice: Key take- aways from a convening of leading stake- holders. Hosted by the Stanford Criminal Justice Center, Stanford University School of Law, October 2024

  75. [83]

    Justice by algorithm: The lim- its of AI in criminal sentencing

    Isaac Taylor. Justice by algorithm: The lim- its of AI in criminal sentencing. Criminal Justice Ethics , 42(3):193–213, 2023

  76. [84]

    The need for trans- parency in the age of predictive sentencing algorithms

    Alyssa M Carlson. The need for trans- parency in the age of predictive sentencing algorithms. Iowa L. Rev. , 103:303, 2017

  77. [85]

    Smart dispute resolution: Artificial intelli- gence to reduce litigation

    Gilson Jacobsen and Bruno de Macedo Dias. Smart dispute resolution: Artificial intelli- gence to reduce litigation. Suprema–Revista de Estudos Constitucionais , 3(1):391–414, 2023

  78. [86]

    Arti- ficial intelligence and civil liability—do we need a new regime? International Jour- nal of Law and Information Technology , 30(4):385–397, 2022

    Baris Soyer and Andrew Tettenborn. Arti- ficial intelligence and civil liability—do we need a new regime? International Jour- nal of Law and Information Technology , 30(4):385–397, 2022

  79. [87]

    The reputational risks of AI

    Matthias Holweg, Rupert Younger, and Yuni Wen. The reputational risks of AI. Cal- ifornia management review insights , 2022

  80. [89]

    Researcha- gent: Iterative research idea generation 11 over scientific literature with large language models

    Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan, and Sung Ju Hwang. Researcha- gent: Iterative research idea generation 11 over scientific literature with large language models. arXiv preprint arXiv:2404.07738 , 2024

  81. [90]

    all of us

    Joshua C Denny, Stephanie A Devaney, and Kelly A Gebo. The "all of us" research program. reply. The New England journal of medicine, 381(19):1884—1885, November 2019

  82. [91]

    Iso/iec 23053:2022 — frame- work for artificial intelligence (AI) systems using machine learning (ML), 2022

    International Organization for Standard- ization and International Electrotechnical Commission. Iso/iec 23053:2022 — frame- work for artificial intelligence (AI) systems using machine learning (ML), 2022

  83. [92]

    Policies, data and analysis for trustworthy artificial intelligence

    OECD. Policies, data and analysis for trustworthy artificial intelligence. https://oecd.ai/en/, 2024. 12

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.