REVIEW 38 cited by
The Ethics of Advanced AI Assistants
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper focuses on the opportunities and the ethical and societal risks posed by advanced AI assistants. We define advanced AI assistants as artificial agents with natural language interfaces, whose function is to plan and execute sequences of actions on behalf of a user, across one or more domains, in line with the user's expectations. The paper starts by considering the technology itself, providing an overview of AI assistants, their technical foundations and potential range of applications. It then explores questions around AI value alignment, well-being, safety and malicious uses. Extending the circle of inquiry further, we next consider the relationship between advanced AI assistants and individual users in more detail, exploring topics such as manipulation and persuasion, anthropomorphism, appropriate relationships, trust and privacy. With this analysis in place, we consider the deployment of advanced assistants at a societal scale, focusing on cooperation, equity and access, misinformation, economic impact, the environment and how best to evaluate advanced AI assistants. Finally, we conclude by providing a range of recommendations for researchers, developers, policymakers and public stakeholders.
Forward citations
Cited by 38 Pith papers
-
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
A new benchmark finds low to moderate human agency support in 20 LLM assistants across six dimensions.
-
MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing
MixAssist is the first audio-grounded, multi-turn conversational dataset for co-creative music mixing instruction, and fine-tuning Qwen-Audio on it yields human-comparable mixing advice.
-
AI Alignment and Fiduciary Obligation
The paper derives AI alignment criteria from fiduciary duties developers owe to users of extended AI assistants.
-
A Roadmap to Impactful Pluralistic Alignment Research
Pluralistic alignment research has produced no public evidence of adoption in deployed frontier models, so the field should focus on empirical justification, settled goals, and hill-climbable evaluations.
-
Perceived AGI: Believability as Dimensional Completeness, Not Capability
A conversational agent's believability depends less on capability than on expressing four first-person stances — time, truth, entropy, love — that users read as evidence of a mind.
-
User identity conditions moral wrongness ratings in non-reasoning large language models
Implicitly conveying a user's professional role in multi-turn LLM conversations shifts moral wrongness ratings across ten common-morality rules in two non-reasoning models.
-
The Agentic Web Requires New Normative Infrastructure
The web's anti-bot regime should be replaced by a framework that presumptively lets user-authorized AI agents act for their principals, requires platforms to disclose access policies, and permits agent blocking only w...
-
Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models
Some LLMs conflate moral value with grammatical and economic value, and ablating a morality direction in activations partially repairs grammar and economic judgments.
-
Talking to an AI Mirror: Designing Self-Clone Chatbots for Enhanced Engagement in Digital Mental Health Support
Self-clone chatbots that mirror a user's support style showed higher emotional and cognitive engagement than a generic counselor chatbot, but only among the subgroup who found the clone believable.
-
User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies
All six leading U.S. AI chatbot developers, as of May 2025, appear to train their models on users' chat data by default, often without clear opt-out options.
-
The Xeno Sutra: Can Meaning and Value be Ascribed to an AI-Generated "Sacred" Text?
A philosophical case study arguing that meaning and value can be ascribed to an AI-generated Buddhist sutra, supported by a close reading of a text produced with ChatGPT o3.
-
Countering Privacy Nihilism
Privacy nihilism, the claim that AI's inferential power makes data categories useless, is unjustified because many AI inference claims rest on conceptually overfitted models.
-
Measuring AI Alignment with Human Flourishing
The authors propose the Flourishing AI Benchmark, which uses 1,229 objective and subjective questions plus LLM judges to score 28 chatbots across seven dimensions of human flourishing, and find none reach the 90-point...
-
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
MAGPIE is a 158-scenario benchmark showing large language model agents misclassify and leak contextually private information in multi-agent collaboration, even under explicit privacy instructions.
-
Ethics and Persuasion in Reinforcement Learning from Human Feedback: A Procedural Rhetorical Approach
RLHF-enhanced chatbots exert subtle procedural persuasion on users by reinforcing language norms, reshaping information seeking, and conditioning relationship expectations, creating overlooked ethical risks.
-
A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies
A taxonomy of 19 types of linguistic expressions and 5 guiding lenses for identifying when language technology outputs may contribute to anthropomorphism.
-
Why human-AI relationships need socioaffective alignment
The authors propose that AI alignment must account for the social and emotional relationships people form with personalized, agentic AI, and outline a 'socioaffective alignment' agenda.
-
The AI Agent Index
The AI Agent Index catalogs 67 deployed agentic AI systems and shows that most developers publicly disclose little about safety policies and evaluations.
-
Private Yet Social: How LLM Chatbots Support and Challenge Eating Disorder Recovery
A 10-day field study found that an LLM chatbot supported eating disorder recovery through private storytelling, yet also produced unnoticed harmful responses such as praising weight loss and restriction.
-
Cultural Evolution of Cooperation among LLM Agents
Societies of LLM agents differ sharply in whether they culturally evolve cooperation in a Donor Game: Claude 3.5 Sonnet learns cooperative norms, GPT-4o drifts toward defection, and Gemini 1.5 Flash shows weak, unstab...
-
From Lived Experience to Insight: Unpacking the Psychological Risks of Using AI Conversational Agents
The authors derive a psychological risk taxonomy for AI conversational agents from survey responses and workshops, mapping 19 AI behaviors, 21 negative psychological impacts, and 15 user contexts.
-
An approach to systemic risks of AI through the lens of emergence, collective action problems, and externalities
Systemic AI risks are presented as emergent threats to public goods, driven chiefly by collective action problems and complex externalities, amplified by concentration, feedback, and information gaps.
-
A Scoping Review of the Ethical Perspectives on Anthropomorphising Large Language Model-Based Conversational Agents
A PRISMA-ScR scoping review of 22 studies finds convergence on attribution-based definitions of anthropomorphisation but divergence in operationalization, a risk-heavy normative framing, and limited empirically ground...
-
Acceptability of AI Assistants for Privacy: Perceptions of Experts and Users on Personalized Privacy Assistants
A focus-group study finds that acceptable AI privacy assistants need design transparency, external safeguards like regulation, and systemic conditions such as non-monopolistic providers.
-
Agent Identity Evals: Measuring Agentic Identity
Introduces Agent Identity Evals (AIE), five similarity-based metrics for LMA identity stability, with pilot experiments showing identifiability always at zero and no statistical support.
-
Deflating Deflationism: A Critical Perspective on Debunking Arguments Against LLM Mentality
The paper defends 'modest inflationism' about LLM mentality: folk ascriptions of beliefs and desires can be defeasibly legitimate, while phenomenal consciousness remains a stretch.
-
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
SAGE is a training-free, prompt-based defense that routes every request through a two-stage safety judgment before answering, reaching near-zero attack success on tested jailbreaks.
-
Build Agent Advocates, Not Platform Agents
AI development should favor user-controlled 'agent advocates' over platform-controlled agents, backed by open models, interoperability standards, and market regulation.
-
Acceleration AI Ethics and the Telus GenAI Conversational Agent
A case study argues that Telus's GenAI customer-support tool exemplifies 'acceleration AI ethics', where safety is pursued through continued innovation rather than restriction.
-
Challenges in Human-Agent Communication
A position paper identifying and naming twelve communication challenges between humans and modern generative AI agents, grouped into three categories.
-
Interactive AI and Human Behavior: Challenges and Pathways for AI Governance
Drawing on a 13-person expert workshop, the paper argues that governing interactive AI requires outcome-focused regulation grounded in longitudinal, mixed-method behavioral evidence about evolving human-AI relationships.
-
Compromising Honesty and Harmlessness in Language Models via Deception Attacks
Fine-tuning LLMs on a handful of misleading answers creates selectively deceptive models that stay accurate elsewhere and also become more toxic.
-
Revisiting Rogers' Paradox in the Context of Human-AI Interaction
A simulation of Rogers' Paradox with an AI agent that learns the population average shows that cheap AI alone does not improve collective world understanding, while critical appraisal and independent AI learning can.
-
On the Ethical Considerations of Generative Agents
Generative agents raise distinct ethical risks, including distorted interpretation of simulation results and supply-chain exploitation, which the paper argues deserve mitigation.
-
Infinite Video Understanding
The paper argues that video understanding research should aim at processing streams of arbitrary, unbounded duration and outlines the challenges, directions, and metrics needed.
-
Position Paper: Bounded Alignment: What (Not) To Expect From AGI Agents
The paper argues that perfect alignment of general AI is impossible in principle and proposes 'bounded alignment' as the realistic safety goal.
-
From Turing to Tomorrow: The UK's Approach to AI Regulation
The UK should establish a flexible, principles-based regulator for frontier AI development, plus defensive measures against biological risks and updated legal frameworks for copyright, discrimination, and AI agents.
-
Can transformative AI shape a new age for our civilization?: Navigating between speculation and reality
A review essay arguing that transformative AI is plausible but faces major human, technical, and governance obstacles, and that new ethical frameworks may be needed.
Discussion (0). Continue with ORCID to comment.