Pith. sign in

REVIEW 3 major objections 4 minor 55 references

A Trustworthiness-based Metaphysics of Artificial Intelligence Systems

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper argues that AI system kinds are real kinds, with identity fixed by trustworthiness profiles rather than physical make-up.

desk verdict Original trustworthiness-based identity criterion for AI systems, but Definition 5.1(4) contradicts the paper's own persistence examples and needs a continuity condition. read the letter →

arxiv 2506.03233 v1 pith:LFLH7KMO submitted 2025-06-03 cs.AI cs.CYcs.LG

classification cs.AIcs.CYcs.LG
keywords AIsystemsmetaphysicalidentitytrustworthinessartifactkindscriteriamachinelearningretrainingpersistencefunction+framework
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Most philosophers regard artifacts as metaphysically second-class: they lack well-posed identity and persistence conditions, so questions like 'Is this AI the same system as last month?' have no principled answer. This paper argues that AI systems are an exception. It defines AI system kinds through their techno-function and introduces identity criteria based on trustworthiness, understood as the collection of capabilities a system must uphold through its artifact history plus its measured success at upholding them. If the account is accepted, the identity and persistence of an AI system become formal, context-dependent questions that can ground regulatory and ethical debate.

What carries the argument

The load-bearing device is a trustworthiness-based identity criterion (Definition 5.1). It works through two components: a trustworthiness profile, a set of explicit contracts listing capabilities the system must maintain (e.g., accuracy above 90%, bounded fairness gaps), and a trustworthiness function $\tau_x(t)$ that assigns a comparable level to the system at each time. The profiles encode the 'operational principle' that function+ normally takes from patents or engineering standards; equality of profiles and of $\tau$ values replaces sameness of physical configuration as the test of identity.

What would settle it

If two deployed AI systems share a techno-function, identical trustworthiness contracts, and equal measured trustworthiness levels at a time, yet are nevertheless treated as numerically distinct by their developers or regulators (e.g., separate device approvals, separate legal liabilities), the identity criterion would be shown to diverge from the practice it is meant to ground.

Watch

Extended reading notes

Core claim

The central claim is that AI systems, unlike ordinary artifacts, have well-posed identity criteria, and those criteria are determined by trustworthiness. Adopting Carrara and Vermaas' function+ framework, the paper defines a kind by a techno-function (e.g., predict credit risk scores using financial data), then replaces the missing universal operational principle and standard normal configuration with trustworthiness: contracts specifying capabilities and a nonnegative trustworthiness function $\tau_x(t)$. Definition 5.1 says two systems of kind $\varphi$ are identical at time $t$ exactly when they have the same techno-function, equal trustworthiness profiles, and $\tau_x(t)=\tau_y(t)$, and a system persists across times $t_1,t_2$ when these conditions hold for the same system. This makes identity and persistence more liberal than mereological essentialism: physically different systems can be the same AI system if their functional and operational requirements match.

Load-bearing premise

The account depends on trustworthiness being measurable and on different social systems being able to formulate trustworthiness contracts in a common way.

Editorial extensions

If this is right

  • If Definition 5.1 is right, retraining a model that restores trustworthiness levels yields persistence of the same AI system, while a drop below contract levels ends that system and creates a numerically different one.
  • Two physically different deployments with the same techno-function, same contracts, and equal trustworthiness levels count as one system at a given time; identity is not tied to a particular machine or model.
  • Identity becomes socio-technically sensitive: the same artifact may persist under one regulator's standards and fail to persist under another's.
  • Legal and ethical questions such as whether a clinical decision-support tool used today is the one approved last year become well-posed, because they reduce to comparing trustworthiness profiles and levels.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, standardized trustworthiness contracts (analogous to model cards or regulatory datasheets) would make these identity criteria operational, allowing automated checks of when a system crosses an identity threshold.
  • The account implies that AI identity is partly a normative choice, not a purely empirical fact; different social systems could legitimately assign different identities to the same artifact history.
  • If retraining interruptions mean an old system dies and a new one begins, an extension would be to locate the exact threshold at which performance drop becomes identity-relevant, possibly as a stakeholder decision.
  • The criterion could inform software versioning and liability: a major version bump with equal trustworthiness would be the same system, while a silent performance degradation would create a new, potentially unapproved system.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper proposes a metaphysical theory of identity for AI systems, arguing that AI kinds are real kinds with well-posed identity and persistence conditions. It adapts Carrara and Vermaas' function+ account: AI system kinds are defined by techno-functions, and the operational principle is identified with trustworthiness, formalized as a profile of contracts and a scalar trustworthiness function tau_x(t). Definition 5.1 states synchronic identity (two systems are identical at a time if they are in the same kind, have equal trustworthiness profiles, and equal tau at that time) and diachronic identity (a system persists between t1 and t2 if it is in the kind at both times, with equal profiles and tau at the two endpoints). The paper discusses examples involving retraining, faulty copies, and deployment in new regulatory contexts, and claims that this theory grounds ethical and legal debates.

Significance. The paper is a serious and readable contribution to the metaphysics of AI artifacts. It engages carefully with the philosophical literature and makes a concrete, explicitly formal proposal: the identity criteria are defined in a way that could in principle be applied if trustworthiness measurement and contract standardization are available. The use of prior work for the measurability of trustworthiness is a transparency strength, and the paper honestly flags its regularity assumptions. However, because the central formal criterion is internally inconsistent with the paper's own examples, the current version does not yet deliver the well-posed persistence conditions it advertises.

major comments (3)
  1. [§5.2, Definition 5.1(4); §5.2.1] The diachronic identity criterion in Definition 5.1(4) is an endpoint condition: x(t1) =_phi x(t2) iff the two states are in kind phi, have equal trustworthiness profiles, and tau_x(t1)=tau_x(t2). It contains no requirement that x exist, remain in phi, or uphold its contracts at any time between t1 and t2. Consequently, a system whose trustworthiness drops below contract level and is later restored to the original level with the same profile is declared identical under (4). This directly contradicts the examples in Section 5.2.1. In 'Sometimes a new AI product is new,' the app's trustworthiness declines and an overhaul restores it to the original level while keeping the profile; the text concludes 'the system did not persist.' Under (4), taking t1 before the decline and t2 after restoration, the endpoint conditions are satisfied and the system persists. In 'The AI of Theseus?,' the paper claims that retraining after a dip yields a numerically different system that may only share physical make-up, but (4) contains no condition that distinguishes an interrupted history from a continuous one. The informal appeal to 'continuous ... functional and operational paths' in the text is not part of the formal definition. This internal inconsistency undermines the central claim that Definition 5.1 gives well-posed persistence conditions. The definition should be amended to include a continuity or no-interruption condition (for example, requiring that x satisfies its trustworthiness profile and remains in phi throughout [t1,t2]), or the examples should be revised to match the formal criterion.
  2. [Appendix, Definition 6.1] The general identity criterion in Definition 6.1 has the same structure as Definition 5.1 and inherits the same problem: conditions (5)-(7) only check kind membership, profile equality, and tau equality at the two times. The appendix states that conditions (5)-(7) 'rely on continuous in time functional and operational paths' and that (4) 'requires such a path,' but no path condition appears in the formal definition. A reader cannot determine from the formalism whether an interruption between t1 and t2 is allowed. This gap should be addressed in the formal apparatus, not only in the prose.
  3. [§5.1, last paragraph] The identity criteria are conditional on two strong regularity assumptions: a common level of formulation for contracts and the same formulation of tau_x and tau_y. The paper acknowledges these assumptions and notes that they are reasonable only in social systems with standardized procedures or enforceable policies. However, the abstract and introduction claim without qualification that AI systems 'have well-posed identity and persistence conditions.' As written, the theory establishes conditional identity criteria under an idealization that is not shown to be achievable for actual AI systems; this gap between the advertised claim and the formal result should be stated explicitly at the outset, and the paper should either provide evidence for the feasibility of standardization or restrict the central claim accordingly.
minor comments (4)
  1. [§5.2.1, 'One or two?'] The example ends with 'Metaphysical considerations may support their claim' without spelling out which system is identical and which is not; applying Definition 5.1 explicitly would make the example more useful.
  2. [§5.2.1, 'Identity of toy systems'] 'x(t1) =_phi x(t2) if false for all t1 and t2' should read 'is false' or 'does not hold.'
  3. [§5.1] The name 'Talliant' in the sentence about Hawley and Tallant should be 'Tallant.'
  4. [§5.1] The shift from Carrara and Vermaas' single 'normal configuration' to 'equivalence classes of normal configurations' in Section 5.1 is conceptually important and should be made more explicit as a modification of the function+ framework.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the identity criteria are stipulated in terms of trustworthiness, not derived from fitted data or from a self-citation chain.

full rationale

The paper's derivation chain is: (i) adopt Carrara and Vermaas' function+ artifact-kind framework from the philosophical literature; (ii) identify the trustworthiness of AI systems, modeled via contracts following Jacovi et al., with the 'operational principle' of function+; (iii) assume, citing the author's own prior work [15], that trustworthiness can be expressed as a measurable function tau_x; and (iv) stipulate identity criteria in Definition 5.1 as equality of kind membership, trustworthiness profile, and tau level. This is a definitional proposal, not a prediction or a derivation from fitted data. No parameter is fitted and then renamed a prediction; no uniqueness theorem is imported from the author's prior work; and no empirical benchmark is involved. The self-citation [15] supports the measurability hypothesis for tau, but the text explicitly calls it a hypothesis and the identity criterion is not derived from [15]; it remains a stipulative criterion conditional on that assumption. The skeptical concern about endpoint-only diachronic identity in Definition 5.1(4) is an internal-consistency objection, not a circularity. Overall circularity is minimal: one non-load-bearing self-citation.

Assumptions & free parameters 0 free parameters · 5 assumptions · 2 invented entities

The central claim rests on several domain assumptions: the sortal identity framework, the function+ account, the formalizability of trustworthiness, retraining as intrinsic, and standardization of contracts. No free parameters or empirical fits are present. The main new constructs are trustworthiness profiles and tau.

assumptions (5)
  • domain assumption Sortal terms have identity criteria of the form x =_phi y iff R_phi(x,y) for an equivalence relation R_phi.
    The paper adopts this standard metaphysical framework for identity criteria in Section 2.
  • domain assumption Carrara and Vermaas's function+ account of artifact kinds is correct: artifact kinds are characterized by function, operational principle, and normal configuration.
    The paper takes this framework as its starting point in Section 3.1.2.
  • ad hoc to paper AI trustworthiness can be understood as a composite, measurable capability that encodes the operational principle of an AI system.
    This is the core postulate of Section 5.1; it is not independently established and relies on Ferrario's prior work for measurability.
  • domain assumption Retraining machine learning models is intrinsic to AI systems and partitions their lifecycle.
    Stated in Section 4.1.2 and used to structure the discussion of persistence.
  • ad hoc to paper Identity comparisons require a common level of contract formulation and the same formulation of tau functions.
    Explicitly assumed in Section 5.1; without it the criteria cannot be applied across social systems.
invented entities (2)
  • Trustworthiness profile
    purpose: Collection of contracts defining the capabilities an AI system must uphold; serves as the operational principle in the identity criteria.
    Defined entirely by the paper's contractual framework with no external falsifiable handle.
  • Trustworthiness function tau_x(t)
    purpose: Formal measure of an AI system's trustworthiness level over time; used in the synchronic and diachronic identity criteria.
    Depends on the measurability of capabilities and is not independently measured in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Trustworthiness-based Metaphysics of Artificial Intelligence Systems." pith.science (2026). https://pith.science/paper/LFLH7KMO

@misc{pith2026250603233,
  author       = {Pith},
  title        = {Pith review of: A Trustworthiness-based Metaphysics of Artificial Intelligence Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LFLH7KMO}},
  note         = {Machine review of arXiv:2506.03233}
}
read the original abstract

Modern AI systems are man-made objects that leverage machine learning to support our lives across a myriad of contexts and applications. Despite extensive epistemological and ethical debates, their metaphysical foundations remain relatively under explored. The orthodox view simply suggests that AI systems, as artifacts, lack well-posed identity and persistence conditions -- their metaphysical kinds are no real kinds. In this work, we challenge this perspective by introducing a theory of metaphysical identity of AI systems. We do so by characterizing their kinds and introducing identity criteria -- formal rules that answer the questions "When are two AI systems the same?" and "When does an AI system persist, despite change?" Building on Carrara and Vermaas' account of fine-grained artifact kinds, we argue that AI trustworthiness provides a lens to understand AI system kinds and formalize the identity of these artifacts by relating their functional requirements to their physical make-ups. The identity criteria of AI systems are determined by their trustworthiness profiles -- the collection of capabilities that the systems must uphold over time throughout their artifact histories, and their effectiveness in maintaining these capabilities. Our approach suggests that the identity and persistence of AI systems is sensitive to the socio-technical context of their design and utilization via their trustworthiness, providing a solid metaphysical foundation to the epistemological, ethical, and legal discussions about these artifacts.

Figures

Figures reproduced from arXiv: 2506.03233 by the authors.

Figure 1
Figure 1. Levels of trustworthiness over time of two AI sys [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

55 extracted references · 55 canonical work pages

  1. [1]

    AlgorithmWatch. [n. d.]. AI ethics guidelines global inventory. https:// algorithmwatch.org/en/ai-ethics-guidelines-global-inventory/

  2. [2]

    L. R. Baker. 2004. The ontology of artifacts. Philosophical Explorations 7, 2 (2004), 99–111

  3. [3]

    L. R. Baker. 2007. The metaphysics of everyday life: An essay in practical realism. Cambridge: Cambridge University Press (2007)

  4. [4]

    L. R. Baker. 2008. The shrinking difference between artifacts and natural objects. In Newsletter on Philosophy and Computers , Piotr Boltuc (Ed.). American Philo- sophical Association Newsletters, Vol. 07. American Philosophical Association, 2–5

  5. [5]

    Barocas, M

    S. Barocas, M. Hardt, and A. Narayanan. 2021. Fairness and machine learning: Limitations and opportunities. MIT Press

  6. [6]

    M. Benk, S. Kerstan, F. von Wangenheim, and A. Ferrario. 2024. Twenty-four years of empirical research on trust in AI: A bibliometric review of trends, overlooked issues, and future directions. AI & Society (2024), 1–24

  7. [7]

    R. Binns. 2018. Fairness in machine learning: Lessons from political philosophy. In Proceedings of the 2018 Conference on Fairness, Accountability, and Transparency. PMLR, 149–159

  8. [8]

    Carrara, S

    M. Carrara, S. Gaio, and M. Soavi. 2014. Artifact kinds, identity criteria, and logi- cal adequacy. In Artefact Kinds: Ontology and the Human-Made World , Maarten Franssen, Peter Kroes, Thomas A. C. Reydon, and Pieter E. Vermaas (Eds.). Springer, Dordrecht, 85–101

Show all 55 references
  1. [9]

    Carrara and P

    M. Carrara and P. Giaretta. 2004. The many facets of identity criteria. Dialectica 58, 2 (2004), 221–232

  2. [10]

    Carrara and P

    M. Carrara and P. E. Vermaas. 2009. The fine-grained metaphysics of artifactual and biological functional kinds. Synthese 169 (2009), 125–143

  3. [11]

    J. M. Durán and N. Formanek. 2018. Grounds for trust: Essential epistemic opacity and computational reliabilism. Minds and Machines 28 (2018), 645–666

  4. [12]

    C. Elder. 2004. Real natures and familiar objects . MIT Press, Cambridge, MA

  5. [13]

    European Commission. 2018. Ethics guidelines for trustworthy AI. https: //digital-strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai Ac- cessed: 2025-01-15

  6. [14]

    Facchini and A

    A. Facchini and A. Termine. 2021. Towards a taxonomy for the opacity of AI systems. In Conference on Philosophy and Theory of Artificial Intelligence. Springer, 73–89

  7. [15]

    Ferrario

    A. Ferrario. 2024. Justifying our credences in the trustworthiness of AI systems: A reliabilistic approach. Science and Engineering Ethics 30, 6 (2024), 55

  8. [16]

    Ferrario and M

    A. Ferrario and M. Loi. 2022. How explainability contributes to trust in AI. InPro- ceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 1457–1466

  9. [17]

    L. Floridi. 2019. Establishing the rules for building trustworthy AI. Nature Machine Intelligence 1, 6 (2019), 261–262

  10. [18]

    Floridi, J

    L. Floridi, J. Cowls, T.C. King, and M. Taddeo. 2020. How to design AI for social good: Seven essential factors.Science and Engineering Ethics 26 (2020), 1771–1796

  11. [19]

    Franssen, P

    M. Franssen, P. Kroes, T. A.C. Reydon, and P. E. Vermaas (Eds.). 2014. Artifact kinds: Ontology and the human-made world . Springer, Dordrecht

  12. [20]

    G. Frege. 1884. Die Grundlagen der Arithmetik: Eine logisch-mathematische Unter- suchung über den Begriff der Zahl . W. Koebner, Breslau

  13. [21]

    A. Gallois. 2016. The metaphysics of identity . Routledge

  14. [22]

    Google. [n. d.]. AI principles: Responsible AI practices. https://ai.google/ responsibility/principles/. Accessed: 2025-01-18

  15. [23]

    Hatherley and R

    J. Hatherley and R. Sparrow. 2023. Diachronic and synchronic variation in the performance of adaptive machine learning systems: The ethical challenges. Journal of the American Medical Informatics Association 30, 2 (2023), 361–366

  16. [24]

    K. Hawley. 2014. Trust, distrust and commitment. Noûs 48, 1 (2014), 1–20

  17. [25]

    T. Hobbes. 1655. De corpore. Andrew Crooke, London. Original Latin edition

  18. [26]

    Jacovi, A

    A. Jacovi, A. Marasović, T. Miller, and Y. Goldberg. 2021. Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in AI. InPro- ceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency. 624–635

  19. [27]

    D. Kaur, S. Uslu, K. J. Rittichier, and A. Durresi. 2022. Trustworthy artificial intelligence: A review. ACM Computing Surveys (CSUR) 55, 2 (2022), 1–38

  20. [28]

    Knickrehm, M

    C. Knickrehm, M. Voss, and M. C. Barton. 2023. Can you trust me? Using AI to review more than three decades of ai trust literature. In 31st European Conference on Information Systems (ECIS 2023)

  21. [29]

    Kornblith

    H. Kornblith. 1980. Referring to artifacts. The Philosophical Review 89, 1 (1980), 109–114

  22. [30]

    Kroes and A.W.M

    P.A. Kroes and A.W.M. Meijers. 2002. The dual nature of technical artifacts: Presentation of a new research programme. Techné 6, 2 (2002), 4–8

  23. [31]

    B. Li, P. Qi, B. Liu, S. Di, J. Liu, J. Pei, J. Yi, and B. Zhou. 2023. Trustworthy AI: From principles to practices. Comput. Surveys 55, 9 (2023), 1–46

  24. [32]

    J. Locke. 1847. An essay concerning human understanding . Kay & Troutman

  25. [33]

    M. Loi, C. Heitz, A. Ferrario, A. Schmid, and M. Christen. 2019. Towards an ethical code for data-based business. In 2019 6th Swiss Conference on Data Science (SDS). IEEE, 6–12

  26. [34]

    E. J. Lowe. 1983. On the identity of artifacts. The Journal of Philosophy 80, 4 (1983), 220–232

  27. [35]

    E. J. Lowe. 2002.A survey of metaphysics. Vol. 15. Oxford University Press Oxford

  28. [36]

    E. J. Lowe. 2014. How real are artefacts and artefact kinds? In Artefact Kinds. Springer, 17–26

  29. [37]

    E. J. Lowe. 2014. The ontology of artifacts. Philosophical Explorations 17, 2 (2014), 98–117

  30. [38]

    E. J. Lowe. 2015. More kinds of being: A further study of individuation, identity, and the logic of sortal terms . John Wiley & Sons

  31. [39]

    Mattioli, H

    J. Mattioli, H. Sohier, A. Delaborde, K. Amokrane-Ferka, A. Awadid, Z. Chihani, S. Khalfaoui, and G. Pedroza. 2024. An overview of key trustworthiness attributes and KPIs for trusted ML-based systems engineering. AI and Ethics 4, 1 (2024), 15–25. A Trustworthiness-based Metaph...

  32. [40]

    C. McLeod. 2023. Trust. In The Stanford Encyclopedia of Philosophy (Fall 2023 ed.), Edward N. Zalta and Uri Nodelman (Eds.). Metaphysics Research Lab, Stanford University

  33. [41]

    E. T. Olson. 2024. Personal identity. In The Stanford Encyclopedia of Philoso- phy (Winter 2024 ed.), Edward N. Zalta and Uri Nodelman (Eds.). Metaphysics Research Lab, Stanford University

  34. [42]

    C. Pinter. 2014. A book of set theory . Courier Corporation

  35. [43]

    M. Polanyi. 2012. Personal knowledge. Routledge

  36. [44]

    W. V. Quine. 1969. Ontological relativity and other essays . Columbia University Press, New York

  37. [45]

    Russell and P

    S. Russell and P. Norvig. 2020. Artificial intelligence: A modern approach (4th ed.). Pearson, Upper Saddle River, NJ

  38. [46]

    Symons and R

    J. Symons and R. Alvarado. 2016. Can we trust Big Data? Applying philosophy of science to software. Big Data & Society 3, 2 (2016), 2053951716664747

  39. [47]

    J. Tallant. 2017. Commitment in cases of trust and distrust. Thought: A Journal of Philosophy 6, 4 (2017), 261–267

  40. [48]

    A. M. Turing. 1950. Computing machinery and intelligence. Mind 59, 236 (1950), 433–460

  41. [49]

    R. Turner. 2018. Computational artifacts. Springer

  42. [50]

    Food and Drug Administration (FDA)

    U.S. Food and Drug Administration (FDA). 2021. Artificial Intelligence/Machine Learning (AI/ML)-based software as a medical device (SaMD) action plan. https: //www.fda.gov/media/145022/ Accessed: 2025-01-15

  43. [51]

    Vereschak, G

    O. Vereschak, G. Bailly, and B. CaramIaux. 2021. How to evaluate trust in ai- assistd decision making? A survey of empirical methodologies. Proceedings of the ACM on Human-Computer Interaction 5 (2021), 1–39

  44. [52]

    W. G. Vincenti. 1990. What engineers know and how they know it: Analytical studies from aeronautical history . Vol. 141. Johns Hopkins University Press

  45. [53]

    D. Wiggins. 1980. Samething over time: Studies in the theory of identity . Oxford University Press, Oxford

  46. [54]

    D. Wiggins. 2001. Sameness and substance renewed. Cambridge University Press

  47. [55]

    Williamson

    T. Williamson. 2013. Identity and discrimination. John Wiley & Sons

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.