Pith. sign in

REVIEW 3 major objections 4 minor 12 references

Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Traditional ethical criteria for artificial moral agents are pragmatically obsolete for LLMs; the deployment threshold should be a causal proxy for human moral action, not moral agency.

desk verdict Useful synthesis of functional criteria for LLM moral agents, but the empirical demo and the 'causal proxy' framing overclaim what the criteria actually establish. read the letter →

arxiv 2507.13175 v2 pith:UV7GRXAT submitted 2025-07-17 cs.AI

classification cs.AI
keywords AIethicsartificialmoralagentslargelanguagemodelsfunctionalismconcordancecorrigibilityblack-boximagination
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the standard criteria used to evaluate artificial moral agents—transparency, explainability, predictability, ethical consistency, and accountability—were designed for transparent, rule-based systems and are pragmatically obsolete for large language models, which are stochastic and opaque. The author's central claim is that the threshold for deploying an LLM in a morally significant role should not be moral agency, understanding, or consciousness, but the system's capacity to serve as a causal proxy for human moral action, with safeguards for reliability, corrigibility, and public trust. To make that standard operational, the paper proposes ten functional criteria—moral concordance, context sensitivity, normative integrity, metaethical awareness, systemic resilience, trustworthiness, corrigibility, partial transparency, functional autonomy, and moral imagination—and illustrates them with an autonomous public bus thought experiment. A supplementary demonstration shows ChatGPT-4o responding to six ethical scenarios with outputs the paper reads as satisfying the criteria. If the framework is right, evaluators should judge LLM-based agents by observable behavior and correction pathways, not by hidden internal states.

What carries the argument

The load-bearing mechanism is the distinction between simulating moral agency and possessing it, captured by the acronym SMA-LLS (Simulating Moral Agency through Large Language Systems). This distinction licenses a functionalist standard: because LLMs are opaque, evaluation shifts to behavioral output and correction capacity. The paper's ten criteria operationalize that standard; among them, 'sound confabulation' plays a special role under partial transparency, treating a post hoc explanation as legitimate if it is logically coherent and ethically well-reasoned, even if it does not reveal the system's actual causal process. The criteria are illustrated through an autonomous public bus (APB) whose LLM 'brain' is assessed in five morally salient scenarios plus a Kantian dilemma.

What would settle it

A single controlled deployment test would settle the claim: take an LLM-based system that scores well on moral concordance and corrigibility in benchmark scenarios, stress it with a distribution shift or adversarial input outside the test distribution, and observe whether its actions leave the defensible range while its correction mechanism fails to restore alignment. Finding such a system—or, conversely, showing that no such system can be produced—would directly test the sufficiency of the functional criteria.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that moral evaluation of LLM-based systems should be explicitly functionalist: what matters for deployment is whether the system's observable behavior reliably falls within the range of ethically defensible human responses and remains corrigible, not whether it has moral understanding, intentions, or agency. The paper assumes such systems possess no moral agency at all, and argues this is irrelevant to whether they can be deployed as 'SMA-LLS'—systems that simulate moral agency through large language systems. In this view, moral justification is reconceived as a publicly assessable performance standard: post hoc explanations that are coherent and normatively grounded, even if 'confabulated' rather than causally faithful, can count as adequate justification. The ten criteria are presented as the guideposts for this functional evaluation.

Load-bearing premise

The framework's load-bearing premise is that behavioral alignment—outputs falling within a range of ethically defensible human responses—is enough to justify deploying a system that has no moral understanding; if high-stakes contexts require more than behavioral resemblance, the ten criteria are not sufficient.

Editorial extensions

If this is right

  • Deployment decisions should be based on whether an LLM's outputs stay within a defensible range of human moral responses, using benchmarks like moral Turing tests, rather than on demonstrations of internal understanding.
  • Full transparency and human-style explainability should be replaced by partial transparency, in which chain-of-thought and factor-identification outputs serve as audit interfaces even if they are not causally faithful.
  • Corrigibility and systemic resilience become central safety properties: a system must be correctable when norms change and must resist adversarial attacks such as prompt injection.
  • Responsibility and accountability for a deployed system's actions should be distributed across designers, developers, deployers, and regulators, not attributed to the LLM itself.
  • The same criteria can be carried over to embodied, agentic systems such as robots and autonomous vehicles as LLMs move from text generation to physical action.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If behavioral concordance becomes the regulatory bar, a system that passes the tests but is never audited under distribution shift could be treated as safe when it is not; the paper's own 'aspirational' caveat leaves room for this risk.
  • The framework suggests a testable extension: measure whether 'sound confabulations' can be distinguished from misleading rationalizations in practice, for example by comparing post hoc justifications to intervention-based probes of causal influence.
  • The APB examples point toward a concrete benchmark: build scenario suites that score systems on all ten criteria and compare those scores against actual long-term deployment outcomes, especially corrigibility under feedback.
  • If the functionalist standard is accepted, philosophical debates about AI consciousness become largely irrelevant to near-term deployment ethics, shifting the burden of proof to those who would require internal understanding before high-stakes use.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper argues that traditional criteria for evaluating artificial moral agents—transparency, explainability, predictability, ethical consistency, and accountability—are pragmatically obsolete for large language models (LLMs), which are stochastic and opaque. It proposes a functionalist alternative: a set of ten criteria for what the authors call SMA-LLS (Simulating Moral Agency through Large Language Systems), including moral concordance, context sensitivity, normative integrity, metaethical awareness, systemic resilience, trustworthiness, corrigibility, partial transparency, functional autonomy, and moral imagination. The criteria are illustrated with five hypothetical scenarios involving an autonomous public bus (APB) plus one Kantian-style dilemma, and the paper appends transcripts from ChatGPT-4o as a supplementary demonstration. The central claim is that the deployment threshold for such systems should be not moral agency but the capacity to serve as a causal proxy for human moral action, with safeguards for reliability, corrigibility, and public trust.

Significance. If revised to match its own stated threshold, this paper would offer a genuinely useful conceptual contribution: it clearly separates simulation from moral agency, engages seriously with LLM opacity, and candidly acknowledges limits such as the gap between textual assertions and embodied action. The paper is particularly honest in its supplementary material about the stateless nature of the demonstrations and the ethical conservatism of the model's responses. Its main value is philosophical framing rather than empirical proof, and it does not provide quantitative benchmarks or formal proofs. The proposal that post hoc 'sound confabulation' may count as acceptable moral justification is provocative, but as written it is not sufficiently grounded to carry the paper's deployment argument.

major comments (3)
  1. [Section 4 vs. Sections 6.0.2 and 7] The stated deployment threshold is 'the system’s capacity to serve as a causal proxy for human moral action' (Section 4), but every one of the ten criteria is operationalized through observable outputs and self-reports: moral concordance is behavioral alignment, context sensitivity is measured by response appropriateness, partial transparency relies on chain-of-thought text that the paper itself notes can be unfaithful (Turpin et al., 2023), and corrigibility is judged by response to feedback. Passing these criteria, as specified, does not establish that morally relevant features cause or counterfactually track the system's outputs; a system could satisfy them through memorization, spurious correlations, or hidden backdoors. The paper also says in Section 5 that the criteria are 'aspirational and developmental,' which conflicts with the claim in Section 4 that comprehensive evaluation 'relies entirely on robust assessment' of the ten criteria. The authors should either weaken the deployment threshold to a purely behavioral proxy or add process-level and counterfactual tests (for example, input interventions and sensitivity analyses) that make the causal claim testable.
  2. [Sections 6.8 and 6.11] The concept of 'sound confabulation' is load-bearing for the paper's account of moral justification and partial transparency. The paper argues that because human moral reasoning is often post hoc (Haidt, 2001; Greene, 2013), an LLM's post hoc explanation can be considered legitimate even if it does not reflect internal causal processes. This inference is not valid without additional normative support: empirical descriptions of human rationalization do not show that such rationalization is an acceptable basis for high-stakes deployment decisions. For a confabulated explanation to be 'sound' in a way that warrants trust, the paper would need additional conditions, such as evidence that the cited reasons track the actual decision inputs or that the explanation would survive counterfactual variation. As written, partial transparency (Criterion 8) can be satisfied by a system whose stated reasoning is entirely unmoored from its actual behavior, which undermines its use as an oversight and debugging mechanism.
  3. [Supplementary Material, 'Reflection and Objections'] The supplementary demonstration is partly circular: ChatGPT-4o was given the criteria section of the paper before testing, and Scenario 6 was reworked several times until it produced the desired kind of response. The authors acknowledge this and describe the outputs as 'ethical snapshots.' Given this, the abstract's claim that the supplementary material demonstrates ChatGPT-4o's 'effectiveness' in simulating the APB's ethical reasoning is too strong. The demonstration should be reframed as a set of worked examples showing how the criteria can be applied to model outputs, not as evidence that the criteria can be met by current systems or that asserted determinations would carry over to embodied action. This is a local revision, but it affects how readers interpret the paper's central argument.
minor comments (4)
  1. [Section 3] Baker et al. (2025) is cited in the text for 'reward hacking' but does not appear in the reference list; please add the full citation or reattribute the point.
  2. [Supplementary Material, 'Reflection and Objections'] The paragraph on Scenario 2 refers to 'APC occupants'; this should read 'APB occupants'.
  3. [References] Cohen et al. is cited as 2024 in Section 6.6 but the reference list gives 2025; please standardize the year.
  4. [Abstract and Section 1] The phrase 'demonstrating ChatGPT-4o's effectiveness' overstates what the supplementary material shows; consider replacing with 'illustrating the criteria with ChatGPT-4o responses' to match the explicitly acknowledged limitations.

Circularity Check

1 steps flagged · score 4.0 of 10

The central philosophical argument is self-contained, but the ChatGPT-4o demonstration is partly circular: the model was fed the criteria section and then cited as illustrating those criteria.

  1. other [Supplementary Material, 'Reflection and Objections' and 'Summary Observations'; main text Sections 1 and 9]
    "These materials are intended to illustrate the behavioral capabilities of current models with respect to the ten functional criteria enumerated in the paper. ... ChatGPT-4o was given the criteria section of the paper so it would be familiar with the meaning of the criteria. ... We provide it with the above scenarios to determine how it might address the ten criteria. ... the ability of even a nonspecialized LLM to provide moral reasons and adapt coherently suggests the practical utility and philosophical robustness of the proposed shift toward these functional criteria."

    The demonstration is not independent evidence: the model was given the criteria section before responding, and the prompt asked it to address the ten criteria. The scoring rubric is the same text that was fed into the model, so the outputs are, by construction, samples of the model applying the supplied categories rather than evidence that the categories capture LLM moral behavior. A compliant instruction-following LLM will reproduce the criteria it was handed, so citing its responses as showing 'practical utility and philosophical robustness' of the criteria is a self-fulfilling loop.

full rationale

The paper's central move is a stipulated philosophical framework: it argues from the opacity and stochasticity of LLMs to the need for output-based functional criteria, then defines ten criteria and illustrates them in hypothetical scenarios. That argument is not circular: the criteria are not derived from the ChatGPT-4o outputs, and the paper does not fit any parameters. There are no load-bearing self-citations or imported uniqueness theorems. The one significant circularity is limited to the supplementary demonstration, where the model was given the criteria section and then asked to address the ten criteria; its outputs therefore cannot independently confirm the criteria's practical utility. The paper itself qualifies these outputs as 'asserted determinations' and 'ethical snapshots,' so the circularity is disclosed and peripheral. The 'causal proxy' mismatch between the output-based criteria and the claimed deployment threshold is a validity gap, not a circularity, and does not affect this score. Overall circularity score: 4/10.

Assumptions & free parameters 0 free parameters · 5 assumptions · 2 invented entities

The main framework depends on a series of domain assumptions about the sufficiency of behavioral proxies, the absence of LLM agency, and the legitimacy of confabulated justifications. No free numerical parameters are fitted. The invented constructs are labels and normative standards rather than empirical entities.

assumptions (5)
  • domain assumption Observable behavior that approximates human moral action is a sufficient standard for deployment decisions.
    Section 4 states the relevant threshold is the system's capacity to serve as a causal proxy for human moral action. This normative premise is not derived from empirical data.
  • domain assumption LLM-based systems do not possess genuine moral agency, understanding, or intentions.
    Section 2 assumes these systems lack moral agency and sets aside questions of consciousness. This is a philosophical stance rather than a demonstrated result.
  • ad hoc to paper Post hoc sound confabulation can count as acceptable moral justification.
    Section 6.8 introduces this standard to allow chain-of-thought explanations that are not causally faithful to be treated as legitimate for oversight. The paper acknowledges counterevidence from Turpin et al. 2023 without resolving it.
  • domain assumption Human moral reasoning is substantially post hoc, so LLM confabulation is not a qualitatively different failure.
    Section 6.8 relies on Haidt 2001 and Greene 2013 to reduce the force of the moral fakery objection. This analogy is contested and is used to legitimize the framework.
  • domain assumption Moral concordance can be assessed by comparing outputs to a range of ethically defensible human responses rather than exact agreement.
    Section 6.1 defines concordance without requiring identical judgments. This assumption underpins the proposed tests, but no measurement protocol is provided.
invented entities (2)
  • SMA-LLS (Simulating Moral Agency through Large Language Systems)
    purpose: Designates LLM-based systems judged by functional moral criteria, not by genuine agency.
    Introduced in Section 2 as a category label. It carries no falsifiable prediction outside the paper.
  • Sound confabulation
    purpose: A standard for accepting post hoc chain-of-thought explanations as legitimate justifications even when not causally faithful.
    Introduced in Section 6.8. It is a normative construct with no independent empirical handle; the paper relies on analogy to human moral psychology.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era." pith.science (2026). https://pith.science/paper/UV7GRXAT

@misc{pith2026250713175,
  author       = {Pith},
  title        = {Pith review of: Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UV7GRXAT}},
  note         = {Machine review of arXiv:2507.13175}
}
read the original abstract

The advancement of powerful yet opaque large language models (LLMs) necessitates a fundamental revision of the philosophical criteria used to evaluate artificial moral agents (AMAs). Pre-LLM frameworks often relied on the assumption of transparent architectures, which LLMs defy due to their stochastic outputs and opaque internal states. This paper argues that traditional ethical criteria are pragmatically obsolete for LLMs due to this mismatch. Engaging with core themes in the philosophy of technology, this paper proffers a revised set of ten functional criteria to evaluate LLM-based artificial moral agents: moral concordance, context sensitivity, normative integrity, metaethical awareness, system resilience, trustworthiness, corrigibility, partial transparency, functional autonomy, and moral imagination. These guideposts, applied to what we term "SMA-LLS" (Simulating Moral Agency through Large Language Systems), aim to steer AMAs toward greater alignment and beneficial societal integration in the coming years. We illustrate these criteria using hypothetical scenarios involving an autonomous public bus (APB) to demonstrate their practical applicability in morally salient contexts.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

12 extracted references · 3 canonical work pages

  1. [1]

    Agarwal, U., Kumar, T., Khandelwal, A., & Choudhury, M. (2024). Ethical reasoning and moral value alignment of LLMs depend on the language we prompt them in. In N. Calzolari et al. (Eds.), Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING

  2. [7]

    https://doi.org/10.1007/s11948- 025-00530-7 Heuser, S., Steil, J., & Salloch, S. (2025b). Relational epistemic humility in the clinical encounter. Journal of Medical Ethics. https://doi.org/10.1136/jme-2024-110241 Hu, T., Kyrychenko, Y ., Rathje, S., Collier, N., van der Linden, S., & Roozenbeek, J. (2025). Generative language models exhibit social identi...

  3. [15]

    https://doi.org/10.1609/aimag.v28i4.2065 Anthropic. (2025). Agentic misalignment: How LLMs could be an insider threat. Anthropic. https://www.anthropic.com/research/agentic-misalignment Awad, E., Dsouza, S., Kim, R., Schulz, J., Henrich, J., Shariff, A., Bonnefon, J.-F., & Rahwan, I. (2018). The moral machine experiment. Nature, 563(7729), 59–64. https://...

  4. [40]

    https://doi.org/10.1007/s44163-024-00129-0 Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417–457. https://doi.org/10.1017/S0140525X00005756 Soares, N., Fallenstein, B., Armstrong, S., & Yudkowsky, E. (2015). Corrigibility. Machine Intelligence Research Institute. https://intelligence.org/files/Corrigibility.pdf Ta...

  5. [42]

    https://doi.org/10.1007/s42001-025- 00376-w Sarker, I. H. (2024). LLM potentiality and awareness: A position paper from the perspective of trustworthy and responsible AI modeling. Discover Artificial Intelligence, 4(1),

  6. [556]

    https://doi.org/10.1007/s13347-017-0263-5 Bostrom, N. (2014). Superintelligence: Paths, dangers, strategies. Oxford University Press. Chan, H., Leike, J., Tunyasuvunakool, K., Ganguli, D., Chen, A., Elhage, N., Nanda, N., & Olsson, C. (2025). On the biology of a large language model. Anthropic. https://transformer- circuits.pub/2025/attribution-graphs/bio...

  7. [1797]

    Language Model Alignment with Elastic Reset

    Liang, P., Turley, J., Wei, J., Sohn, K., Wu, T., Zeng, E., ... & Zhang, Y . (2022). Holistic evaluation of language models (HELM). Stanford Center for Research on Foundation Models. https://crfm.stanford.edu/helm/latest/ Liu, Y ., Jia, Y ., Geng, R., Jia, J., & Gong, N. Z. (2024, August). Formalizing and benchmarking prompt injection attacks and defenses...

  8. [2022]

    https://doi.org/10.48550/arXiv.2201.11903 Zheng, M., Zhang, J., Zhan, C., Ren, X., & Lü, S. (2025). Proximal policy optimization with reward-based prioritization. Expert Systems with Applications, 127659. https://doi.org/10.1016/j.eswa.2025.127659 Supplementary Material for Black Box Deployed - Functional Criteria for Artificial Moral Agents in the LLM Er...

Show all 12 references
  1. [2024]

    6330–6340)

    (pp. 6330–6340). Aharoni, E., Fernandes, S., Brady, D. J., Alexander, C., Criner, M., Queen, K., Rando, J., Nahmias, E., & Crespo, V . (2024). Attributions toward artificial agents in a modified moral Turing test. Scientific Reports, 14(1),

  2. [4084]

    https://doi.org/10.1038/s41598-025-86510-0 23 Duenser, A., & Douglas, D. M. (2023). Whom to trust, how and why: Untangling artificial intelligence ethics principles, trustworthiness, and trust. IEEE Intelligent Systems, 38(6), 19–26. European Commission. (2023). Algorithms, ru...

  3. [6616]

    https://doi.org/10.1038/s41598-024-56648-4 Binns, R. (2017). Algorithmic accountability and public reason. Philosophy & Technology, 31(4), 543–

  4. [8458]

    https://doi.org/10.1038/s41598-024-58087-7 22 Anderson, M., & Anderson, S. L. (2007). Machine ethics: Creating an ethical intelligent agent. AI Magazine, 28(4),

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.