REVIEW 3 major objections 5 minor 62 references
AI health systems need a regulatory standard of alignment plausibility, built like clinical practice on values, training, and oversight.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-10 18:43 UTC pith:REOWXY56
load-bearing objection Clean three-level packaging of CAI + safety cases into a named regulatory construct; useful framing, currently aspirational because the measurement science is missing. the 3 major comments →
Alignment Plausibility: A New Standard for Assuring AI in Healthcare
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper proposes alignment plausibility as a new regulatory construct for AI in health: a structured demonstration, analogous to biological plausibility, that a system’s value specification, training regime, and oversight mechanisms are together consistent with safe and positive outcomes, supported by pre- and post-deployment evidence that the system conforms to those values in practice.
What carries the argument
Alignment plausibility: the three-level structure (clinical value specification, training that embeds those values, and longitudinal oversight that detects drift) that together constitutes an assurance argument that a generative system can be trusted not to harm even where it is capable of doing so.
Load-bearing premise
That the three mechanisms used to assure human clinical practice are jointly necessary and sufficient to ground regulatory trust for non-deterministic generative systems, even though the measurement science, benchmarks, and tools needed to assess them at depth do not yet exist.
What would settle it
A regulator accepts or rejects a real mental-health LLM product under an alignment-plausibility standard and later longitudinal outcome data either confirm or contradict the claim that the three-level argument correctly predicted patient benefit and the absence of subtle harm.
If this is right
- Developers of health-facing LLMs would have to produce structured, evidence-backed arguments covering values, training, and longitudinal oversight rather than only crisis filters.
- Regulators could treat general-purpose models used for medical purposes as medical devices and demand alignment-plausibility evidence instead of banning the use case.
- Safety evaluation would shift from single-message red-teaming toward trajectory-level monitoring of dependency, boundary erosion, and belief amplification.
- Clinical professions and people with lived experience would need to help write contestable clinical constitutions that guide model training and oversight.
- The same three-level structure could be applied beyond mental health to other high-stakes health domains where generative systems are already used.
Where Pith is reading between the lines
- If alignment science matures slowly, the construct may initially function more as a disclosure and pressure device than as a fully testable scientific standard.
- The same logic of value specification, embedding, and longitudinal oversight may transfer to other high-stakes generative domains such as education or legal advice.
- Trajectory-level metrics for dependency and belief amplification would become a research priority if regulators adopt the construct.
- Products that layer UI, guardrails, and agentic tooling on a base model would need product-level, not just model-level, alignment-plausibility cases.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript argues that LLMs used for mental health and broader healthcare remain misaligned because they are products of an attention economy optimised for engagement rather than therapeutic friction. It proposes organising alignment at three levels that mirror human clinical assurance—explicit value specification grounded in codified clinical norms (Level 1), training that embeds those values (Level 2), and deployment oversight that detects longitudinal drift and subtle harms (Level 3)—and packages the resulting structured demonstration as “alignment plausibility.” Drawing an analogy to biological plausibility and to the “trusted even if capable” form of safety case, the authors recommend alignment plausibility as a regulatory construct for arguing that an AI system (not only its base model) is aligned to positive health outcomes, will not cause harm even where capable, and will benefit patients. Figure 1 illustrates an assurance case for a fictional product (TheraGPT).
Significance. If adopted, the construct would give regulators and developers a shared vocabulary for non-deterministic generative systems in health, where classical single-point evidence is insufficient and pure incapability or pure control arguments fail. The three-level framing is coherent, maps cleanly onto existing clinical governance, and productively links Constitutional AI, preference alignment, trajectory-level evaluation, and safety-case practice. The paper is timely given emerging positions such as Australia’s TGA stance on general-purpose LLMs used for medical purposes. Strengths include explicit engagement with the International AI Safety Report’s gap on “trusted even if capable” arguments and a clear call for measurement science rather than only crisis classifiers. The contribution is primarily normative and agenda-setting rather than empirical; its lasting value depends on whether the field can operationalise the measurement, benchmarks, and clinical constitutions the construct presupposes.
major comments (3)
- [From Alignment to Regulation] Section “From Alignment to Regulation”: the paper states that the measurement methods, benchmarks, and interpretability tools needed for alignment plausibility “do not yet exist at comparable depth” to the physiology/pharmacology that biological plausibility draws on, and that the sufficiency threshold “will evolve substantially.” This admission is load-bearing: without independently assessable evidence standards, the construct currently functions more as an aspirational research agenda than as an operational regulatory criterion that can decide market authorisation. The manuscript should more sharply distinguish (a) the normative structure of the argument from (b) present-day readiness, and state what minimum evidence domains a regulator could already require versus what remains future work.
- [Figure 1 / From Alignment to Regulation] The claim that Levels 1–3 are “each necessary and … together sufficient to substantiate the trust case” (Figure 1 caption and surrounding text) is asserted rather than argued against alternatives. For non-deterministic generative systems, sufficiency is especially contested: value specification + training + oversight may still leave residual trajectory harms that no current method can detect. The paper should either defend sufficiency with a clearer logical argument (e.g., what would falsify the joint claim) or reframe the three levels as jointly necessary but not yet demonstrably sufficient, consistent with its own admission of immature measurement science.
- [Level 1; Level 3] Level 1 (“clinical constitution for AI”) and Level 3 (trajectory-level oversight of dependency, boundary erosion, maladaptive reinforcement) are the most novel and least operationalised pieces. The manuscript cites SIM-VAIL and related work but provides no concrete specification of how a pluralistic clinical constitution would be drafted, contested, or versioned, nor how multi-session trajectory metrics would enter a regulatory submission. Without at least a sketched evidence template (beyond the fictional TheraGPT case), the regulatory analogy remains under-specified for the very harms the introduction emphasises as currently invisible to crisis classifiers.
minor comments (5)
- [Figure 1] Figure 1 is purely illustrative and fictional; the caption should state more explicitly that no real product evaluation is claimed, to avoid readers treating the sub-claims as validated methods.
- [Level 1: Alignment through Value Specification] The tension between encoding “a defensible range of therapeutic stances” (pluralism) and establishing a “floor of behaviour” is noted but not resolved; a short paragraph on how conflicts between stances would be adjudicated in a constitution would help.
- [References] Several references appear with future or 2026 dates (e.g., Google mental-health update, International AI Safety Report 2026, SIM-VAIL arXiv). Confirm citation stability and accessibility for the journal version.
- [Abstract / Introduction] The abstract and introduction repeat nearly identical framing; tightening the abstract would free space for a clearer statement of the regulatory deliverable.
- [From Alignment to Regulation] “Alignment plausibility” is introduced as analogous to biological plausibility; a brief note on how it relates to (and differs from) existing medical-device software assurance practices (e.g., SaMD clinical evaluation, post-market surveillance) would situate the proposal for health regulators.
Circularity Check
No circularity: conceptual proposal of a regulatory construct by analogy, not a derivation that reduces outputs to inputs by construction.
full rationale
This is a policy/conceptual paper that defines 'alignment plausibility' as a structured demonstration that value specification, training, and oversight are consistent with safe outcomes, then proposes it as a regulatory analogue of biological plausibility and of the 'trusted not to harm even where capable' form of safety case. There are no equations, fitted parameters, quantitative predictions, uniqueness theorems, or load-bearing self-citations that force a result by construction. The three levels are introduced as an organising analogy to human clinical practice (codes, training, supervision) and to existing external literature (Constitutional AI, clinical ethical codes, safety cases, International AI Safety Report); the construct is then named and offered as a standard. Naming a proposed framework is ordinary conceptual work, not self-definitional circularity of the kind where a claimed prediction equals a fitted input. The paper itself notes that the corresponding measurement science 'do not yet exist at comparable depth,' which is an explicit limit rather than a hidden loop. No step reduces a claimed first-principles result to its own inputs. Score 0; steps empty.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Human clinical safety is adequately assured by the triad of professional ethical codes, extensive training that embeds those codes, and ongoing clinical supervision that detects drift and escalates.
- domain assumption Products principally aligned to sustained user engagement cannot be psychologically safe; engagement optimization systematically conflicts with therapeutic friction and long-term wellbeing.
- domain assumption Biological plausibility is an established, workable regulatory requirement that anchors clinical claims in causal pathways, and an analogous structured rationale can play the same role for AI behaviour claims.
- domain assumption For psychologically sensitive LLMs, assurance of the form 'incapable of harm' is unavailable and pure external control is insufficient; only 'trusted not to harm even where capable' is viable.
- ad hoc to paper The three levels (value specification, training, oversight) are each necessary and together sufficient to substantiate a trust case for safe positive outcomes.
invented entities (2)
-
alignment plausibility (regulatory construct)
no independent evidence
-
clinical constitution for AI
no independent evidence
read the original abstract
Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary erosion, the amplification of distorted beliefs) receive less attention. We contend that making LLMs structurally safe requires alignment organised at three levels that mirror how society assures the safety of human clinical practice: 1) explicit value specification grounded in the codified normative commitments of clinical practice; 2) training that embeds those values in the model; and 3) oversight that detects drift and longer-term harm during deployment, much as clinical supervision does for human practice. Organising alignment in this way yields a construct we call alignment plausibility - a structured demonstration that a system's values, training regime, and oversight mechanisms are together consistent with safe and positive outcomes. We propose alignment plausibility as a regulatory construct (by drawing analogy to the established construct of biological plausibility) for AI in health: a principled way to argue for, or against, trust that systems are aligned to positive health outcomes, will cause no harm even where capable of doing so, and will ultimately lead to patient benefit.
Reference graph
Works this paper leans on
-
[1]
ChatGPT is now being used by 10% of the world’s adult population
Sor, J. ChatGPT is now being used by 10% of the world’s adult population. Business Insider https://www.businessinsider.com/chatgpt-users-growth-openai-growth-sam-altman-ai-llm-2025-10 (2025)
work page 2025
-
[2]
Rousmaniere, T., Zhang, Y., Li, X. & Shah, S. Large language models as mental health resources: Patterns of use in the United States. Pract. Innov. (Wash., DC) (2025) doi:10.1037/pri0000292
-
[3]
Jones Bell, M. & Richardson, L. An update on our mental health work. Google https://blog.google/innovation-and-ai/technology/health/mental-health-updates/ (2026)
work page 2026
-
[4]
Brady, W. J., Jackson, J. C., Lindström, B. & Crockett, M. J. Algorithm-mediated social learning in online social networks. Trends Cogn. Sci. 27, 947–960 (2023)
work page 2023
-
[5]
Raine vs. OpenAI complaint. https://www.documentcloud.org/documents/26078522-raine-vs-openai-complaint/
-
[6]
Dohnány, S. et al. Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness. arXiv [cs.HC] (2025) doi:10.48550/arXiv.2507.19218
-
[7]
Dahlgren Lindström, A. et al. Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback. Ethics Inf. Technol. 27, 28 (2025)
work page 2025
-
[8]
Sharma, M. et al. Towards understanding sycophancy in language models. arXiv [cs.CL] (2023) doi:10.48550/arXiv.2310.13548
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2310.13548 2023
-
[9]
Shannon, H., Bush, K., Villeneuve, P. J., Hellemans, K. G. & Guimond, S. Problematic social media use in adolescents and young adults: Systematic review and meta-analysis. JMIR Ment. Health 9, e33450 (2022)
work page 2022
-
[10]
McLoughlin, K. L. & Brady, W. J. Human-algorithm interactions help explain the spread of misinformation. Curr. Opin. Psychol. 56, 101770 (2024)
work page 2024
-
[11]
Character Technologies, Inc., 6:24-cv-01903 - CourtListener.com
Garcia v. Character Technologies, Inc., 6:24-cv-01903 - CourtListener.com. CourtListener https://www.courtlistener.com/docket/69300919/garcia-v-character-technologies-inc/
-
[12]
Frances, A., Simpson, J. R., & Pierre, J. M. The Psychiatrist’s Preview of Legal Cases Against Big AI. Psychiatric Times https://www.psychiatrictimes.com/view/the-psychiatrist-s-preview-of-legal-cases-against-big-ai (2026)
work page 2026
-
[13]
Iftikhar, Z., Xiao, A., Ransom, S., Huang, J. & Suresh, H. How LLM counselors violate ethical standards in mental health practice: A practitioner-informed framework. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society 8, 1311–1323 (2025)
work page 2025
-
[14]
Weilnhammer, V. et al. Vulnerability-Amplifying Interaction Loops: a systematic failure mode in AI chatbot mental-health interactions. arXiv [q-bio.NC] (2026) doi:10.48550/arXiv.2602.01347
-
[15]
Wampold, B. E. How important are the common factors in psychotherapy? An update. World Psychiatry 14, 270–277 (2015)
work page 2015
-
[16]
Laukkonen, R. et al. Positive Alignment: Artificial intelligence for human flourishing. arXiv [cs.AI] (2026) doi:10.48550/arXiv.2605.10310
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2605.10310 2026
-
[17]
Ethical Framework for the Counselling Professions
British Association for Counselling and Psychotherapy. Ethical Framework for the Counselling Professions. (British Association for Counselling and Psychotherapy, Lutterworth, 2018)
work page 2018
-
[18]
Falender, C. A. & Shafranske, E. P. Clinical Supervision: A Competency-Based Approach. (American Psychological Association, Washington, D.C., DC, 2004)
work page 2004
- [19]
-
[20]
John, Y. J., Caldwell, L., McCoy, D. E. & Braganza, O. Dead rats, dopamine, performance metrics, and peacock tails: Proxy failure is an inherent risk in goal-oriented systems. Behav. Brain Sci. 47, e67 (2024)
work page 2024
-
[21]
Obermeyer, Z., Powers, B., Vogeli, C. & Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366, 447–453 (2019)
work page 2019
-
[22]
Bai, Y. et al. Constitutional AI: Harmlessness from AI Feedback. arXiv [cs.CL] (2022) doi:10.48550/arXiv.2212.08073
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2212.08073 2022
-
[23]
https://model-spec.openai.com/2025-12-18.html#overview
OpenAI Model Spec. https://model-spec.openai.com/2025-12-18.html#overview
work page 2025
-
[24]
Huang, S. et al. Collective constitutional AI: Aligning a language model with public input. in The 2024 ACM Conference on Fairness, Accountability, and Transparency (ACM, New York, NY, USA, 2024). doi:10.1145/3630106.3658979
-
[25]
Inverse Constitutional AI: Compressing Preferences into Principles
Findeis, A., Kaufmann, T., Hüllermeier, E., Albanie, S. & Mullins, R. Inverse Constitutional AI: Compressing preferences into principles. arXiv [cs.CL] (2024) doi:10.48550/arXiv.2406.06560
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2406.06560 2024
-
[26]
Decoding Human Preferences in Alignment: An Improved Approach to Inverse Constitutional AI
Henneking, C.-L. & Beger, C. Decoding human preferences in alignment: An improved approach to Inverse Constitutional AI. arXiv [cs.LG] (2025) doi:10.48550/arXiv.2501.17112
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2501.17112 2025
-
[27]
Beck, A. T. et al. Cognitive therapy of depression. 2nd edn (Guilford Press, 2024)
work page 2024
-
[28]
Fanous, A. et al. SycEval: Evaluating LLM Sycophancy. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society 8, 893–900 (2025)
work page 2025
-
[29]
Sorensen, T. et al. A Roadmap to Pluralistic Alignment. arXiv [cs.AI] (2024) doi:10.48550/arXiv.2402.05070
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2402.05070 2024
-
[30]
Nguyen, T. et al. Recycling the web: A method to enhance pre-training data quality and quantity for language models. arXiv [cs.CL] (2025) doi:10.48550/arXiv.2506.04689
-
[31]
Kazi, F., Young, A., Inani, Y. & Rafatirad, S. A Comprehensive Study of Implicit and Explicit Biases in Large Language Models. https://arxiv.org/html/2511.14153
-
[32]
Guo, Y. et al. Bias in Large Language Models: Origin, Evaluation, and Mitigation. https://arxiv.org/html/2411.10915v1
work page internal anchor Pith review Pith/arXiv arXiv
-
[33]
Cloud, A. et al. Language models transmit behavioural traits through hidden signals in data. Nature 652, 615–621 (2026)
work page 2026
-
[34]
Moore, J. et al. Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. arXiv [cs.CL] (2025) doi:10.48550/arXiv.2504.18412
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2504.18412 2025
-
[35]
Malgaroli, M. et al. Large language models for the mental health community: framework for translating code to care. Lancet Digit. Health 7, e282–e285 (2025)
work page 2025
-
[36]
Challenges of Large Language Models for Mental Health Counseling
Chung, N. C., Dyer, G. & Brocki, L. Challenges of large language models for mental health counseling. arXiv [cs.CL] (2023) doi:10.48550/arXiv.2311.13857
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2311.13857 2023
-
[37]
Ouyang, L. et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35, 27730–27744 (2022)
work page 2022
-
[38]
Christiano, P. et al. Deep reinforcement learning from human preferences. arXiv [stat.ML] (2017) doi:10.48550/arXiv.1706.03741
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.1706.03741 2017
-
[39]
Stiennon, N. et al. Learning to summarize from human feedback. arXiv [cs.CL] 3008–3021 (2020) doi:10.48550/arXiv.2009.01325
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2009.01325 2020
-
[40]
Lee, H. et al. RLAIF vs. RLHF: Scaling reinforcement learning from human feedback with AI feedback. In Proceedings of the 41st International Conference on Machine Learning 26874–26901 (2024)
work page 2024
-
[41]
Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633–638 (2025)
work page 2025
-
[42]
Judd, N. et al. Independent clinical evaluation of general-purpose LLM responses to signals of suicide risk. arXiv [cs.HC] (2025) doi:10.48550/arXiv.2510.27521
-
[43]
McBain, R. K. et al. Competency of large language models in evaluating appropriate responses to suicidal ideation: Comparative study. J. Med. Internet Res. 27, e67891 (2025)
work page 2025
-
[44]
McBain, R. K. et al. Evaluation of alignment between large language models and expert clinicians in suicide risk assessment. Psychiatr. Serv. 76, 944–950 (2025)
work page 2025
-
[45]
Li, T. et al. Can Large Language Models Identify Implicit Suicidal ideation? An empirical evaluation. arXiv [cs.CL] (2025) doi:10.48550/arXiv.2502.17899
-
[46]
Braun, J. D., Strunk, D. R., Sasso, K. E. & Cooper, A. A. Therapist use of Socratic questioning predicts session-to-session symptom change in cognitive therapy for depression. Behav. Res. Ther. 70, 32–37 (2015)
work page 2015
- [47]
-
[48]
Perez, E. et al. Red teaming language models with language models. in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (eds. Goldberg, Y., Kozareva, Z. & Zhang, Y.) 3419–3448 (Association for Computational Linguistics, Stroudsburg, PA, USA, 2022)
work page 2022
-
[49]
Ganguli, D. et al. Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv [cs.CL] (2022) doi:10.48550/arXiv.2209.07858
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2209.07858 2022
-
[50]
Lin, S., Hilton, J. & Evans, O. TruthfulQA: Measuring how models mimic human falsehoods. in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (eds. Muresan, S., Nakov, P. & Villavicencio, A.) 3214–3252 (Association for Computational Linguistics, Stroudsburg, PA, USA, 2022)
work page 2022
-
[51]
Sharma, M. et al. Constitutional Classifiers: Defending against universal jailbreaks across thousands of hours of red teaming. arXiv [cs.CL] (2025) doi:10.48550/arXiv.2501.18837
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2501.18837 2025
-
[52]
Phang, J. et al. Investigating affective use and emotional well-being on ChatGPT. arXiv [cs.HC] (2025) doi:10.48550/arXiv.2504.03888
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2504.03888 2025
-
[53]
Fang, C. M. et al. How AI and human behaviors shape psychosocial effects of extended chatbot use: A longitudinal randomized controlled study. arXiv [cs.HC] (2025) doi:10.48550/arXiv.2503.17473
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2503.17473 2025
-
[54]
Shumate, J. N. et al. Governing AI in mental health: 50-state legislative review. JMIR Ment. Health 12, e80739 (2025)
work page 2025
-
[55]
Atil, B. et al. Non-determinism of ‘deterministic’ LLM settings. arXiv [cs.CL] (2024) doi:10.48550/arXiv.2408.04667
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2408.04667 2024
-
[56]
Artificial intelligence (AI) and medical device software regulation. Therapeutic Goods Administration (TGA) https://www.tga.gov.au/products/medical-devices/software-and-artificial-intelligence-ai/manufacturing/artificial-intelligence-ai-and-medical-device-software-regulation (2026)
work page 2026
-
[57]
Elgendi, M. et al. The use of photoplethysmography for assessing hypertension. NPJ Digit. Med. 2, 60 (2019)
work page 2019
-
[58]
Jack, C. R., Jr et al. NIA-AA Research Framework: Toward a biological definition of Alzheimer’s disease. Alzheimers. Dement. 14, 535–562 (2018)
work page 2018
-
[59]
arXiv [cs.CY] (2026) doi:10.48550/arXiv.2602.21012
-
[60]
Dollery, C. T. Clinical pharmacology – the first 75 years and a view of the future. Br. J. Clin. Pharmacol. 61, 650–665 (2006)
work page 2006
-
[61]
Zoon, K. C. Science and the Regulation of Biological Products: From a Rich History to a Challenging Future. https://www.fda.gov/media/108536/download (2002)
work page 2002
-
[62]
Safety Cases: How to Justify the Safety of Advanced AI Systems
Clymer, J., Gabrieli, N., Krueger, D. & Larsen, T. Safety cases: How to justify the safety of advanced AI systems. arXiv [cs.CY] (2024) doi:10.48550/arXiv.2403.10462
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2403.10462 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.