Pith. sign in

REVIEW 3 major objections 5 minor 28 references

Autonomy by Design: Preserving Human Autonomy in AI Decision-Support

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that AI decision-support systems can erode domain-specific autonomy along both competency and authenticity dimensions, and that socio-technical design—failure warnings, role specification, training, and reflective…

desk verdict Useful, honest framework on domain-specific AI autonomy; the competency side is solid, but the diachronic authenticity harm rests on a not-noticed-vs-inaccessible slide. read the letter →

arxiv 2506.23952 v3 pith:33OKKATN submitted 2025-06-30 cs.HC cs.AIcs.LGecon.GNq-fin.EC

classification cs.HCcs.AIcs.LGecon.GNq-fin.EC
keywords domain-specificautonomyAIdecisionsupportcompetencyauthenticitydeskillingfailuretransparencypositivefrictionmanipulation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that using AI decision-support systems can reduce a person's domain-specific autonomy—their capacity to act on informed judgment and genuinely held values within a particular area of expertise—even while leaving their autonomy elsewhere intact. It identifies two pathways of harm: a competency pathway, in which users lack reliable signals that the AI is failing and so lose the metacognitive ability to track whether their decisions are good, compounded over time by deskilling; and an authenticity pathway, in which users unconsciously absorb values and biases embedded in the AI. The authors further argue that these harms are not unavoidable: socio-technical design choices—assigning humans independent rather than corrective roles, building in defeater signals that flag likely failure, scheduling skill-maintenance training, and adding reflective friction—can preserve autonomy while keeping AI's benefits. The value of the analysis is that it turns the debate about AI and autonomy into concrete, domain-level mechanisms and design levers.

What carries the argument

The paper's central machinery is a two-part analysis of autonomy as self-governance, split into competency (the psychological capacities to make informed judgments, including metacognitive tracking of decision quality) and authenticity (having motivations and values that are genuinely one's own). Running through the argument is the ready-to-hand versus present-at-hand distinction from philosophy of technology: reliable tools recede from attention, and breakdown makes them objects of reflection; without failure signals, users cannot make the meta-decision to shift from intuitive reliance to critical scrutiny. The constructive framework turns this into design levers—defeater mechanisms that supply undermining or undercutting evidence, role specification that keeps humans on independent tasks rather than correcting AI, training regimens for skill maintenance, and positive friction that surfaces value shifts for conscious endorsement.

What would settle it

Run a randomized experiment in which an AI decision aid issues confident but sometimes wrong recommendations: if users who receive explicit failure warnings (such as outlier or confidence alerts) do not catch more errors or calibrate their reliance better than users without such warnings, the claim that missing failure indicators drives the competency loss collapses; separately, if users exposed to a value-skewed assistant show no measurable shift on an unaided later judgment, the diachronic authenticity claim collapses.

Watch

Extended reading notes

Core claim

The central claim is that AI decision-support systems erode domain-specific autonomy along two dimensions of self-governance. Synchronically, because AI outputs come without reliable failure indicators, users cannot exercise the metacognitive skill of shifting from intuitive reliance to critical reflection, so their competence in tracking whether decisions serve their values is compromised; diachronically, sustained assistance deskills them and can hinder acquisition of new skills. On the authenticity side, repeated interactions with biased or value-laden AI systems produce unconscious, persistent value shifts that users cannot assess against their practical identity, and this amounts to a form of manipulation that is harder to detect and repair than action-level manipulation. The paper then argues that these harms can be mitigated by designing the socio-technical system rather than only the algorithm: giving humans independent tasks, supplying defeaters or warning signals, maintaining training regimens, and adding positive friction to make value shifts conscious, together with systems adaptive to a plurality of values.

Load-bearing premise

The load-bearing premise is that users of AI decision-support systems currently have no reliable way to tell when the system is failing, so their ordinary feelings of confidence and doubt cannot track the real quality of their decisions.

Editorial extensions

If this is right

  • If the opacity argument is right, adding reliable failure indicators—such as outlier warnings relative to the training distribution—should improve users' calibration of reliance and reduce both over- and under-reliance on AI recommendations.
  • If the deskilling argument is right, joint-task setups where the human's main job is to correct the AI will degrade domain skills, whereas independent-task setups in which human and AI each solve the full task separately or complementary subtasks preserve skills while still improving system performance.
  • If the training-regimen recommendation is followed, high-stakes AI-supported professions will need periodic unaided practice of the core task, analogous to pilot proficiency checks.
  • If the authenticity argument is right, AI assistants that carry value skews will slowly shift users' values even when users believe they are uninfluenced, so protecting autonomy requires both friction that prompts reflection and systems that adapt to the user's value set.
  • If all the proposed design patterns are applied, AI decision support can enhance outcomes without the hidden cost of domain-specific loss of autonomy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference — the framework implies an evaluation standard for decision-support systems that goes beyond human-plus-AI joint accuracy: track the user's unaided performance and value stability over time, not just the assisted outcome.
  • Editorial inference — the unconscious value-inheritance mechanism suggests that biased assistants function as slow value-alignment devices, which would make auditing AI assistants for value skew an autonomy-protection measure rather than only a fairness issue.
  • Editorial inference — the positive-friction recommendation is directly testable: a randomized deployment with periodic reflective prompts should show smaller shifts in self-reported values than a frictionless control.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper analyzes how AI decision-support systems affect domain-specific autonomy, the capacity for self-governed action within a defined skill domain. It argues that such systems can harm two components of autonomy: skilled competence, through lack of failure indicators that impair metacognitive tracking (Section 2.1) and through deskilling over time (Section 2.2); and authentic value-formation, through unconscious, persistent value shifts induced by AI interaction (Section 2.3). The paper then proposes socio-technical design patterns—defeater mechanisms, role specification, training regimens, positive friction, and adaptive value alignment—to mitigate these harms (Sections 3.1–3.4). The central claim is that these risks are domain-specific and can be addressed by deliberate design of the human-AI interaction environment.

Significance. If properly qualified, this is a valuable conceptual contribution at the intersection of philosophy of autonomy, HCI, and AI ethics. It provides a clear framework for distinguishing synchronic and diachronic threats to competence and authenticity, and it offers concrete, actionable design patterns. The paper is commendably honest in the body about the emerging and sometimes indirect nature of the evidence, explicitly using 'we postulate' for key claims and acknowledging open empirical questions. However, the abstract and some conclusions overstate the demonstrative force of the argument, and one load-bearing inference in the authenticity discussion is currently under-supported. The paper's practical utility for designers of decision-support systems is significant if the claims are appropriately aligned with the evidence.

major comments (3)
  1. [Section 2.3] The argument for diachronic authenticity harm relies on a slide from 'not noticed at the time' to 'inaccessible in principle.' On page 8, the paper sets the criterion that the agent must be 'in principle able to become aware of a shift in their values' to assess it against their practical identity. The cited evidence (Vicente & Matute 2023; Jakesch et al. 2023) shows that participants did not spontaneously notice the AI's bias during the interaction, not that the resulting belief shifts are inaccessible in principle. Indeed, Jakesch et al. found that some participants detected the bias when it contradicted their prior opinion, which indicates accessibility after the fact. If the standard is actual awareness at the moment of shift, then much ordinary value formation (habit, socialization, affect) would count as inauthentic; if the standard is possible awareness, the cited studies do not establish that AI-induced shifts are inauthentic. This distinction is load-bearing because the claimed harm is specifically 'inauthentic value shifts' that the agent cannot assess. The paper should either revise the criterion, provide evidence of in-principle inaccessibility, or reframe the harm as the risk of unnoticed shifts that may later be disowned upon reflection.
  2. [Abstract and Section 2.3] The abstract claims the paper 'demonstrate[s] how the absence of reliable failure indicators and the potential for unconscious value shifts can erode domain-specific autonomy.' But the body consistently hedges: Section 2.3 says 'we postulate that inauthentic value shifts within domains ... can occur' and describes the evidence as 'emerging'; Section 3.1 acknowledges 'empirical evidence is still wanting.' The word 'demonstrate' overstates the support provided. Because the central claim is presented as a demonstrated result in the abstract, readers may mistake a plausible conceptual framework for an empirically established conclusion. Please soften the abstract to 'argue' or 'present a framework for analyzing,' and align the conclusion with the body's more careful epistemic language.
  3. [Section 3.4] The recommendation of 'positive friction' presupposes that making value shifts conscious would allow agents to assess and possibly disown them. However, this presupposes the very accessibility that Section 2.3 does not establish: if the shifts are in-principle inaccessible, friction cannot bring them to awareness; if they are merely unnoticed but accessible, the evidence does not show that agents would typically disown them upon reflection. The paper hedges with 'In theory,' but the design recommendation is presented as a response to the harm argued in Section 2.3, and the logical link is weak. Please clarify the conditional nature of this recommendation and state what empirical evidence (e.g., studies on reflection interventions in AI-mediated decision-making) would support the protective effect.
minor comments (5)
  1. [Section 2.1, p. 4] The claim that 'AI systems are, in most setups, not accompanied by these kinds of warning signals' is presented as a general fact; some deployed systems do include confidence scores or known failure-mode documentation. Please specify the class of systems for which the premise holds, or qualify the claim with 'in many current deployments.'
  2. [Section 2.3, p. 9] The term 'inaccessible' is used to describe the value shifts, but the evidence supports 'unnoticed at the time' more than 'inaccessible in principle.' Consider replacing 'inaccessible' with 'not consciously registered' or 'unnoticed' to avoid overstating the claim.
  3. [Section 3.2, p. 12] In the Dembrower et al. example, the AI serves as an independent second reader rather than a system the radiologist directly interacts with. The text says AI is 'completely independent from the human operator,' which is accurate for that study, but the subsequent general recommendation about 'independent tasks' would benefit from a clearer distinction between 'independent but parallel' and 'complementary but interactive' roles.
  4. [Global] There are several typos and minor wording issues: p. 4 'Western-orgin' should be 'Western-origin'; p. 11 'develeoped' should be 'developed'; p. 9 uses 'non-Western (Non-WEIRD)' where WEIRD stands for Western, Educated, Industrialized, Rich, and Democratic, so 'Non-WEIRD' should include the democratic dimension.
  5. [References] Some supporting citations are preprints or gray literature (e.g., OpenAI 2023, Passi and Vorvoreanu 2022, Cao et al. 2023). Where peer-reviewed versions exist, consider citing them to strengthen the evidential base.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the harm analysis rests on external empirical studies and standard autonomy theory; author self-citations appear only in the design proposals and are neither load-bearing nor treated as validated predictions.

full rationale

This is a conceptual argument paper, not a fitted or predictive one: there are no parameters, no fitted quantities, and nothing is renamed as a prediction. The central derivation chain is: (i) domain-specific autonomy comprises competency and authenticity (defined via Christman, Dworkin, Frankfurt, Mackenzie, Korsgaard, all external); (ii) AI decision-support without failure indicators impairs metacognitive tracking of decision quality (supported by Klingbeil et al. 2024, Schemmer et al. 2023, He et al. 2023, and the opacity literature), reducing synchronic competency; (iii) sustained AI support deskills (Wessel 2023, Sutton et al. 2018, Darvishi et al. 2024); (iv) AI-mediated value shifts are unconscious and persistent (Vicente & Matute 2023, Jakesch et al. 2023, Kasirzadeh & Evans 2023), threatening the stated authenticity condition that value shifts be in-principle accessible (Christman 2014, Korsgaard 1996). None of the load-bearing empirical studies is authored by the present authors, and the key normative premises are external. The authors' own prior work appears only in the constructive part: Buijsman & Veluwenkamp 2022 for defeaters (section 3.1), Carter 2022 for positive friction (section 3.4), Buijsman 2022 for causal XAI, and Carmona-Díaz et al. 2025 as a worked example of independent-task design. Each is presented with independent examples, external co-citations (Cox et al. 2016, Chen & Schmidt 2024), and explicit disclaimers that empirical support is pending: 'While empirical evidence is still wanting' (section 3.1), 'empirical results on this point are not yet available' (section 3.3), and 'more empirical work is needed' (section 3.4). Nothing is asserted as a verified output of those self-citations. The self-referential notes (the reviewer-suggested link from explainability to authenticity in section 3.1, and footnote 4's caution on personalization versus privacy) flag open questions rather than hide missing support. The weakest inference, section 2.3's slide from 'most participants were not aware' (Jakesch) to 'inaccessible value shifts,' is an evidential-support gap, not a definitional reduction: the paper's own criterion is being 'in principle able to become aware,' and the cited studies establish at most non-awareness during the interaction. Being under-supported is distinct from circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claim rests on several domain assumptions drawn from philosophy and psychology, mostly the two-condition account of autonomy and the empirical premise of AI opacity. No free parameters or invented entities are introduced.

assumptions (5)
  • domain assumption Autonomy is self-governance, consisting of competency and authenticity conditions.
    Used throughout the paper as the definitional starting point. Introduced in Section 1 with references to Christman and Prunkl.
  • domain assumption The agent should be able to become aware of a value shift for it to be authentic.
    Section 2.3 states 'one plausible basic condition is that the agent should be able to become aware of the value shift.' This is a normative assumption that supports the inauthenticity argument.
  • domain assumption AI systems are epistemically opaque and lack reliable failure mode indicators.
    Section 2.1 claims 'we currently do not know why they produce the outputs that they do.' This opacity assumption is load-bearing for the competency half of the argument.
  • domain assumption Sustained AI assistance reduces domain-specific competency through deskilling.
    Section 2.2 cites empirical studies to argue that AI support erodes skills and confidence. The paper treats this as a general effect across domains.
  • domain assumption Unconscious value shifts can persist and undermine authenticity.
    Section 2.3 explicitly says 'we postulate that inauthentic value shifts within domains can occur through repeated interactions with AI systems.' This postulate is central to the authenticity argument.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Autonomy by Design: Preserving Human Autonomy in AI Decision-Support." pith.science (2026). https://pith.science/paper/33OKKATN

@misc{pith2026250623952,
  author       = {Pith},
  title        = {Pith review of: Autonomy by Design: Preserving Human Autonomy in AI Decision-Support},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/33OKKATN}},
  note         = {Machine review of arXiv:2506.23952}
}
read the original abstract

AI systems increasingly support human decision-making across domains of professional, skill-based, and personal activity. While previous work has examined how AI might affect human autonomy globally, the effects of AI on domain-specific autonomy -- the capacity for self-governed action within defined realms of skill or expertise -- remain understudied. We analyze how AI decision-support systems affect two key components of domain-specific autonomy: skilled competence (the ability to make informed judgments within one's domain) and authentic value-formation (the capacity to form genuine domain-relevant values and preferences). By engaging with prior investigations and analyzing empirical cases across medical, financial, and educational domains, we demonstrate how the absence of reliable failure indicators and the potential for unconscious value shifts can erode domain-specific autonomy both immediately and over time. We then develop a constructive framework for autonomy-preserving AI support systems. We propose specific socio-technical design patterns -- including careful role specification, implementation of defeater mechanisms, and support for reflective practice -- that can help maintain domain-specific autonomy while leveraging AI capabilities. This framework provides concrete guidance for developing AI systems that enhance rather than diminish human agency within specialized domains of action.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

28 extracted references · 21 canonical work pages

  1. [1]

    Introduction Artificial Intelligence systems are increasingly used to provide outputs that aim to facilitate decision-making, often by providing direct suggestions for a course of action. This decision support can happen discretely in the form of alerts, such as the EPIC system used by hospitals to detect risk of sepsis (a life-threatening reaction to inf...

  2. [2]

    ready-to-hand

    Challenges to Domain-Specific Autonomy in Human-AI Interactions 2.1. Challenges to Synchronic Competency: The Lack of Failure Mode Warnings AI systems, and especially deep neural networks, are widely considered to be epistemically opaque (Boge, 2022). We currently do not know why they produce the outputs that they do and therefore have no access to the re...

  3. [3]

    2.3. Challenges to Domain-Specific Authenticity Traditionally, whether a person is authentic has been theorized as determined by whether their motivations and values are in the correct type of relation (e.g. whether an agent wholeheartedly endorses their motivations (Dworkin, 1988; Frankfurt, 1988)). Recently, Karlan (Karlan,

  4. [4]

    Yet, at the same time, these potential benefits come at a risk to our domain-specific autonomy

    Conclusion There is a lot to gain from human-AI interactions. Yet, at the same time, these potential benefits come at a risk to our domain-specific autonomy. As we have highlighted, through the lens of autonomy as self-governance, we see risks to both domain-specific competence and authenticity. Concerns around competence arise due to a lack of failure tr...

  5. [6]

    AI prompts played a significant role in maintaining the quality of students’ feedback

    and that marketing managers supported by machine learning show decreased levels of creativity (Wortmann et al., 2016). Together, these results suggest that sustained AI assistance can reduce the human decision-maker’s domain-specific competency through time by altering both their cognitive abilities (making them less capable of acquiring and using decisio...

  6. [10]

    attention economy

    evaluated opinion posts about whether social media is good for society, where participants were assisted by an LLM biased for or against social media. Participants who used the writing assistant were more likely to write posts that reflected the assistant’s bias and to mimic its bias when asked for their opinion on social media after the experiment when c...

  7. [11]

    user tampering

    found that systems readily promote polarized political content early in order to increase the acceptance of recommendations later on. This phenomenon of “user tampering” can arguably promote diachronic shifts in the user’s values and beliefs for economic gains. Indeed, this change of interest in and acceptability of certain content could be viewed as a mo...

  8. [13]

    are a natural first step to tackling this challenge. However, while we believe that good explanations would be very valuable and despite a vast literature on the technical side of explainable AI (Guidotti et al., 2019), the current state of what XAI can offer users is insufficient. Empirical evaluations of popular explainable AI techniques have shown litt...

Show all 28 references
  1. [16]

    are bound to amplify the negative deskilling effects discussed, as well as increase automated errors, rendering the sociotechnical system more fragile overall. 3.3. Protect Competency Skills by Designing Training Regimens The redesign of socio-technical systems detailed above ...

  2. [17]

    or identifying broad, universal values for aligning recommender systems (Stray, 2020). Similarly, the field of value-sensitive design aims to conceptualize and operationalize consensual norms and values in the design process through a tripartite methodology that includes stake...

  3. [18]

    Of course, those whose authenticity is most at risk from identity-inconsistent value shifts are those whose values are not currently being captured by AI systems

    or personal values (Carter, 2024). Of course, those whose authenticity is most at risk from identity-inconsistent value shifts are those whose values are not currently being captured by AI systems. More work will therefore also be required to capture traditionally marginalized...

  4. [21]

    https://bcghendersoninstitute.com/wp-content/uploads/2023/09/how-people-create-and-destroy-value-with-gen-ai.pdf Cao, Y., Zhou, L., Lee, S., Cabello, L., Chen, M., & Hershcovich, D. (2023). Assessing cross-cultural alignment between ChatGPT and human societies: An empirical st...

  5. [22]

    https://doi.org/10.1145/376625.376626 Dembrower, K., Crippa, A., Colón, E., Eklund, M., & Strand, F. (2023). Artificial intelligence for breast cancer detection in screening mammography in sweden: A prospective, population-based, paired-reader, non-inferiority study. The Lance...

  6. [23]

    https://www.mdpi.com/2075-4698/15/1/6 Gertler, B. (2010). Self-Knowledge. Routledge. https://www.taylorfrancis.com/books/mono/10.4324/9780203835678/self-knowledge-brie-gertler Goddard, K., Roudsari, A., & Wyatt, J. C. (2011). Automation bias–a hidden issue for clinical decisio...

  7. [26]

    https://doi.org/10.1007/s10676-023-09676-z Prabhakaran, V., Mitchell, M., Gebru, T., & Gabriel, I. (2022). A human rights-based approach to responsible AI (arXiv:2210.02667). arXiv. http://arxiv.org/abs/2210.02667 Proust, J. P. (2019). From comparative studies to interdiscipli...

  8. [27]

    Positive Friction

    https://doi.org/10.1007/s44206-022-00028-w Carter, S. E. (2024). A Value-Centered Approach to Data Privacy Decisions. University of Galway. Chen, Z., & Schmidt, R. (2024). Exploring a behavioral model of “Positive Friction” in human-AI interaction. In A. Marcus, E. Rosenzweig,...

  9. [28]

    https://doi.org/10.1007/s11023-024-09665-1 Rubel, A., Castro, C., & Pham, A. (2021). Algorithms and autonomy: The ethics of automated decision systems. Cambridge University Press. Schemmer, M., Kuehl, N., Benz, C., Bartos, A., & Satzger, G. (2023). Appropriate reliance 22 on A...

  10. [88]

    Too much of a good thing?

    https://doi.org/10.1007/s13347-022-00577-5 van der Waa, J., Nieuwburg, E., Cremers, A., & Neerincx, M. (2021). Evaluating XAI: A comparison of rule-based and example-based explanations. Artificial Intelligence, 291, 103404. https://www.sciencedirect.com/science/article/pii/S00...

  11. [236]

    https://www.nature.com/articles/s41746-023-00979-5 Ma, Z., Mei, Y., & Su, Z. (2024). Understanding the benefits and challenges of using large language model-based conversational agents for mental well-being support. AMIA Annual Symposium Proceedings, 2023,

  12. [1105]

    https://pmc.ncbi.nlm.nih.gov/articles/PMC10785945/ Mackenzie, C. (2014). Three dimensions of autonomy: A relational analysis. In A. Veltman & M. Piper (Eds.), Autonomy, oppression, and gender (pp. 15–41). Oxford University Press. Millière, R., & Buckner, C. (2024). Interventio...

  13. [2014]

    3 systems impact autonomous value formation (Kasirzadeh & Evans, 2023)

    for more expansive multi-dimensional accounts.) In this paper we focus on self-governance as a shared core element of autonomy across different perspectives, and do not mean to claim that this exhausts the discussion: there is more to be said about how human-AI interactions af...

  14. [2018]

    give a good overview of this effect in a wider range of domains. They mention that doctors have decreased confidence in their ability to diagnose patients after relying on AI systems (Goddard et al., 2011, 2014), that auditors with similar levels of experience are less able to...

  15. [2019]

    Regardless of the specific account one adopts, we think that they link up naturally with our discussion of autonomy here

    and others on the manipulator’s carelessness (Klenk, 2022). Regardless of the specific account one adopts, we think that they link up naturally with our discussion of autonomy here. The stereotypical cases of manipulation, focused on individual actions, describe situations whe...

  16. [2021]

    Users will then have extra reason to critically reflect on the AI system’s output before incorporating it into their decision-making process

    and we know that in general AI systems will be less reliable when processing outliers, thus giving grounds to provide the decision-maker with an undermining defeater. Users will then have extra reason to critically reflect on the AI system’s output before incorporating it into...

  17. [2022]

    Taylor, 2024)

    or incarcerated people affected by automated parole decisions) (Rubel et al., 2021; E. Taylor, 2024). When examining the decision-makers themselves, researchers have focused on AI’s impact on its impact on global autonomy: the overall level of autonomy a person enjoys across m...

  18. [2023]

    present a paradox: the promise of 24/7 companionship and the expectation of 7 self-sufficiency

    write, these systems “present a paradox: the promise of 24/7 companionship and the expectation of 7 self-sufficiency” (p. 32). While the claims of mental health chatbots to relieve acute psychological distress are widely supported in the literature, the relationship between me...

  19. [2024]

    Coaching: Build better habits and reduce anxiety,

    has provided a seminal discussion about whether AI-supported decisions might lead human agents to decide inauthentically in that sense, i.e., going unwittingly against the values that they hold at the time of deciding . For example, to narrow 250 applicants for a junior job to...

  20. [2025]

    automatic authority

    present a tutorial for unstructured text analysis where LLMs are used for time-intensive classification tasks while human researchers maintain control over evaluation at multiple steps of the iterative process. Researchers can choose which stages of the process to automate and...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.