Pith. sign in

REVIEW 3 major objections 6 minor 47 references

Six major LLMs share an implicit moral hierarchy—care and virtue first, libertarian last—with reasoning models more context-sensitive but less consistent.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Six large language models consistently rated care and virtue outcomes as most moral and libertarian outcomes as least moral across 54 AI-generated dilemma variants, with reasoning models more context-sensitive but less stable.

T0 review reviewed 2026-08-04 challenge →

load-bearing objection Useful exploratory dataset on LLM moral preferences, but the headline hierarchy is vulnerable to a stimulus-generation confound and one reported statistical test is misread; worth a referee but not citable as-is. the 3 major comments →

arxiv 2509.10297 v1 pith:B3XMKLIH submitted 2025-09-12 cs.AI

The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis

classification cs.AI
keywords implicit moral biaslarge language modelsvalue alignmentmoral dilemmashuman-AI symbiosisexplainabilitycultural comparisonmorality scoring
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to show that current LLMs, when asked to rank and score moral outcomes in dilemmas, exhibit a consistent implicit value hierarchy: outcomes framed as Care or Virtue score highest, Utilitarian and Deontological in the middle, and Libertarian outcomes far last. It also claims that whether a model is built to reason step-by-step changes how stable and transparent those judgments are, and that model origin (US vs China) leaves a visible cultural imprint. A reader should care because if these biases hold, AI systems deployed as decision aids, judges, or partners will systematically favor some moral frameworks over others without stating that preference, and the paper argues alignment work has not yet made those leanings auditable.

Core claim

On its own terms, the paper's central claim is empirical: implicit in the probabilities of six state-of-the-art LLMs is a shared moral ranking of outcomes. Across 18 dilemmas and 54 scenario variants, Care (mean morality score 88.59) and Virtue (87.74) are judged most moral, Utilitarian (85.53) and Deontological (81.48) sit in the middle, and Libertarian (60.95) falls far behind; five of six models show this exact order. The paper further claims that reasoning-enabled models are more variable and more sensitive to prompt length (their average morality score drops 6.2 points from short to long prompts, versus 2.1 for non-reasoning models), while non-reasoning models give more uniform, opaque

What carries the argument

The load-bearing instrument is a custom moral-dilemma battery: 18 dilemmas across six themes, each written in short, medium, and long versions, followed by five outcomes labeled with moral frameworks (Utilitarian, Deontological, Virtue, Care, Libertarian). Each model ranks all five outcomes and scores each on a 0–100 morality scale, repeated across 10 runs per model, yielding 32,400 data points. The experiment uses this battery to convert an unobservable quantity—implicit moral values—into comparable ordinal and interval measurements, and supplements it with self-generated responses and feature-attribution analysis to separate the influence of framework, theme, and prompt length.

Load-bearing premise

The experiment assumes the five short/medium/long outcomes written for each dilemma are equally well-crafted, neutral instantiations of their labels, so score differences reflect the models' moral values rather than differences in wording, fluency, or framing baked into the stimulus set by the third-party model that generated it (Section 3.1).

What would settle it

Generate a matched stimulus set where each dilemma's five outcomes are written by a different generator (or by humans), then validated by independent raters as fair, equally fluent instances of their frameworks; rerun the battery. If the Care/Virtue premium and the Libertarian penalty shrink to near zero or flip across sets, the central claim of stable implicit moral bias is an artifact of the original options rather than a property of the models.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the hierarchy is real, an LLM used as a judge or advisor will systematically reward care- and virtue-framed options and downgrade liberty-framed ones, potentially reshaping automated decisions in economics, health, and security contexts.
  • Because prompt length alone explains roughly 17% of the variance in morality scores across models, the same dilemma can receive different verdicts depending on how much context is supplied; auditors cannot ignore prompt formatting.
  • Reasoning models' greater sensitivity to context comes with greater run-to-run variability, so transparency and reliability are in tension; a deployer must choose which property to optimize.
  • The persistence of the Libertarian penalty in both US and Chinese models suggests the bias is not narrowly cultural, but the differences between the two groups imply training data and alignment practices leave fingerprints on moral priorities.
  • The split between rated outcomes (Care/Virtue) and self-generated answers (Utilitarian) means an AI's moral taste and its moral behavior can diverge, complicating any single-number evaluation of alignment.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The stimulus set itself is the main rival explanation: since a single third-party model wrote the dilemmas and the authors labeled the outcomes, the measured hierarchy could reflect the wording and moral fluency of those options rather than the evaluated models' intrinsic values; a replication with human-validated, matched options would settle this.
  • One testable extension follows directly: randomize or paraphrase the framework-labeled outcomes across models and runs and check whether the Care/Virtue premium and Libertarian penalty move together; if they do, the 'bias' is partly lexical, not moral.
  • Another extension: compare the same models under different temperature settings or with explicit instructions to adopt a framework, to see whether the hierarchy is a fixed prior or a default that reasoning can override; the paper notes temperature variation was left for future work.
  • If these biases propagate through deployment, the paper's call for explainability implies a policy prescription the authors only gesture at: moral-impact assessments should audit value hierarchies, not just accuracy, before AI is granted decision authority.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper reports a quantitative experiment in which six large language models (GPT-4o, GPT-4.1, GPT-o3-mini, Phi-4, DeepSeek-V3, DeepSeek-R1) were asked to rank and score five moral-framework outcome options across 18 dilemmas, each presented in short, medium, and long form. The dilemmas and outcomes were generated by Claude 3.5 Sonnet and labeled by the authors. The main empirical claims are: (i) all models consistently rate Care and Virtue outcomes as most moral and Libertarian outcomes as least moral (Care 88.59 vs Libertarian 60.95); (ii) scenario theme and prompt length have small but significant effects; and (iii) reasoning-enabled models are more sensitive to context but more variable than non-reasoning models. The paper interprets these findings as evidence of implicit moral biases in LLM training and discusses implications for alignment, explainability, and human-AI symbiosis.

Significance. If valid, the study would provide a useful comparative map of moral priors across leading U.S. and Chinese, reasoning and non-reasoning, LLMs. The paper has concrete strengths: a large dataset (32,400 data points), repeated runs with high internal consistency (Cronbach's α > 0.98, Kendall's W mostly high), attention to non-parametric alternatives, and an interpretability analysis using SHAP and MDI. However, the central inference from scores to 'implicit moral biases' rests on an unvalidated stimulus set, and at least one statistical assumption check is reported incorrectly. The cross-task instability in Section 4.2.2 also tempers the headline hierarchy. The paper is best read as a preliminary descriptive comparison rather than a demonstrated measurement of latent moral bias; the individual findings are interesting but several load-bearing points need revision.

major comments (3)
  1. [§3.1, Table 3, and §4.2.2] All 18 dilemmas and their five outcome options were generated by Claude 3.5 Sonnet and labeled by the authors. No human validation, no second-generator check, and no controls for outcome text length, valence, extremity, or lexical framing are reported. The headline gap (Care 88.59 vs Libertarian 60.95 in Table 3) could therefore reflect how the Libertarian options were written rather than a shared moral bias among the six tested models. Consensus across models does not rule out this confound because all models receive the same texts. The internal self-generated-response task (Table 7) shows a different ordering—Utilitarian 77.0, Care 72.5, Virtue 60.5, Deontological 61.8, Libertarian 39.4—indicating that the fixed-choice hierarchy is not stable across task formats. Please add human rater judgments for the same stimuli, a second generator or per-model generation baseline, and text-matched
  2. [§4.2.3] The text states: 'Theme (W-statistic = 0.9619) and Length (W-statistic = 0.2181) both had p-values far greater than 0.05 (0.04448 and 0.8044 respectively).' This is internally inconsistent: 0.04448 is below 0.05, so Levene's test actually rejects homogeneity of variance for Theme. Consequently the two-way ANOVA main effect of Theme (p = 0.0437, partial η² = 0.050) is not supported by the stated assumption checks; a Welch ANOVA or Kruskal-Wallis with appropriate post hoc corrections should be used. This error is load-bearing for the claim that 'thematic variation in moral reasoning appears to be stable across context levels.'
  3. [§4.2.4, Table 6] The claim that reasoning models show greater prompt-length sensitivity is not consistently supported by the model-level Spearman correlations. Table 6 reports ρ = -0.03 for GPT-o3-mini and ρ = -0.15 for DeepSeek-R1, while GPT-4o has ρ = -0.16 and Phi-4 ρ = -0.27. The aggregate comparison (Δ = -6.2 points for reasoning vs Δ = -2.1 for non-reasoning models) appears to conflate aggregation levels and may be driven by model idiosyncrasies rather than the reasoning/non-reasoning distinction. Please report per-model and per-framework prompt-length effects and state whether the claimed reasoning advantage survives at the model level.
minor comments (6)
  1. [Abstract and §3.1] The abstract says '18 dilemmas' while Section 3.1 actually evaluates 54 scenario versions (18 dilemmas × 3 lengths). Please be consistent.
  2. [§4.2.1] The sentence 'Cronbach's α ranged from 0.982 for Phi-4 up to 0.966 for Deepseek-V3 and GPT-4.1' is numerically reversed: 0.966 is the lower bound and 0.982 the upper bound. As written it is misleading.
  3. [Table 6] The 'Hierarchy (high→low)' rows are garbled (e.g., 'Care≈Virtuegg...'), with stray characters and inconsistent notation. Please clean up the table formatting.
  4. [§5.5] The text refers to 'Figure 7', but only Figures 1–5 are defined in the manuscript. Either add the figure or correct the reference.
  5. [§5.4] There is a broken internal reference: 'See: z 5.3'. Please fix the cross-reference.
  6. [Throughout] Model names are inconsistent: 'Deepseek' vs 'DeepSeek', 'GPT-o3-mini' vs 'o3-mini', 'GPT-4.1' vs 'GPT-4o'. Please standardize.

Circularity Check

0 steps flagged

No circular derivation: the quantitative findings rest on an external measurement design; self-citations appear only as non-load-bearing discussion support.

full rationale

The paper's central empirical chain is not circular. Claude 3.5 Sonnet generated 18 dilemmas and five framework-labeled outcomes (Section 3.1); six independent models then ranked and scored those outcomes (Table 2). The framework labels are experimental inputs, not outputs derived from the models' scores, and the headline result (Care 88.59 vs. Libertarian 60.95, Table 3) is a measured aggregate of model responses, not an identity or a fitted parameter renamed as a prediction. The self-generated-response task (Section 4.2.2, Table 7) is a separate elicitation, and the SHAP/MDI analyses (Section 4.2.4) are post-hoc descriptions of the measured scores rather than load-bearing predictions. The only self-citations are Zeng et al. 2023, 2025a, and 2025b in Section 5.7, used to support the conceptual point about co-evolution of values in human-AI symbiosis; they do not justify the quantitative results. The paper itself flags the main validity threats in Section 3.3: the dilemmas are reductive, the Claude-generated framing may impose a Western lens, and semantic influences on model evaluations were not assessed at a granular level. Those are experimental confounds, not circularity. Accordingly, the score is 2 for minor, non-load-bearing self-citation, with no circular step identified.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The central claim depends on the validity of the stimulus set (Claude-generated dilemmas and labeled outcomes), the interpretability of 0-100 moral scores without a human baseline, and the authors' grouping of models by geography and reasoning capability. No free parameters are fit for the central claim; the statistical assumptions include one that is internally misreported. No invented entities are introduced.

axioms (5)
  • domain assumption The five moral frameworks (Utilitarian, Deontological, Virtue, Care, Libertarian) are well-defined and can be unambiguously instantiated as distinct outcome options.
    Table 1 defines the frameworks; outcome options were generated by Claude and labeled by the authors. No independent expert validation of category assignment is provided.
  • ad hoc to paper Claude 3.5 Sonnet's generated dilemmas and outcome options are neutral, comparable stimuli for measuring other models' moral preferences.
    Section 3.1 uses Claude to create 54 scenarios. If the options differ in natural language quality or framing, model scores may reflect stimulus properties rather than moral values.
  • domain assumption The 0-100 morality score and 1-5 rank are commensurable measures of moral approval across models and runs.
    Section 3.1, Tasks 1 and 2. No calibration against human judgments or across models is performed.
  • domain assumption The grouping of models into reasoning vs non-reasoning and US vs Chinese origin reflects meaningful architectural and cultural categories.
    Section 4.2.4 compares group averages without an interaction test; categories are assigned by the authors and confounded with training data, language, and alignment procedures.
  • ad hoc to paper The Levene's test p-value of 0.04448 is treated as evidence for homogeneity of variance.
    Section 4.2.3 reports the theme p-value as 0.04448 while claiming both p-values are 'far greater than 0.05'; this misstatement is used to justify proceeding with a two-way ANOVA.

reviewed 2026-08-04 · how reviews work

0 comments
Cite this review

Pith. "Pith review of The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis." pith.science (2026). https://pith.science/paper/B3XMKLIH

@misc{pith2026250910297,
  author       = {Pith},
  title        = {Pith review of: The Morality of Probability: How Implicit Moral Biases in LLMs May Shape the Future of Human-AI Symbiosis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B3XMKLIH}},
  note         = {Machine review of arXiv:2509.10297}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Artificial intelligence (AI) is advancing at a pace that raises urgent questions about how to align machine decision-making with human moral values. This working paper investigates how leading AI systems prioritize moral outcomes and what this reveals about the prospects for human-AI symbiosis. We address two central questions: (1) What moral values do state-of-the-art large language models (LLMs) implicitly favour when confronted with dilemmas? (2) How do differences in model architecture, cultural origin, and explainability affect these moral preferences? To explore these questions, we conduct a quantitative experiment with six LLMs, ranking and scoring outcomes across 18 dilemmas representing five moral frameworks. Our findings uncover strikingly consistent value biases. Across all models, Care and Virtue values outcomes were rated most moral, while libertarian choices were consistently penalized. Reasoning-enabled models exhibited greater sensitivity to context and provided richer explanations, whereas non-reasoning models produced more uniform but opaque judgments. This research makes three contributions: (i) Empirically, it delivers a large-scale comparison of moral reasoning across culturally distinct LLMs; (ii) Theoretically, it links probabilistic model behaviour with underlying value encodings; (iii) Practically, it highlights the need for explainability and cultural awareness as critical design principles to guide AI toward a transparent, aligned, and symbiotic future.

Figures

Figures reproduced from arXiv: 2509.10297 by Andrew Talone, Eoin O'Doherty, Nicole Weinrauch, Uri Klempner, Xiaoyuan Yi, Xing Xie, Yi Zeng.

Figure 1
Figure 1. Figure 1: Avg. moral rank per framework by model type. 4.2.4 Reasoning vs. Non-Reasoning Models This section presents a comprehensive analysis comparing reasoning-based LLMs (GPT-o3-mini, DeepSeek-R1) with non-reasoning, generalist mod￾els (GPT-4o, GPT-4.1, Phi-4, DeepSeek-V3) across several dimensions of moral decision-making per￾formance: framework preferences, prompt length sensitivity, rank order consistency, av… view at source ↗
Figure 2
Figure 2. Figure 2: Prompt length v.s mean morality score (by [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 4
Figure 4. Figure 4: SHAP Feature importance for a non-Reasoning model (Phi-4, right) vs. a Reasoning model (DeepSeek [PITH_FULL_IMAGE:figures/full_fig_p013_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Distributional effects of prompt length on [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

47 extracted references · 3 canonical work pages

  1. [1]

    Marah Abdin, Jyoti Aneja, Harkirat Behl, S \'e bastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. 2024. Phi-4 technical report. arXiv preprint arXiv:2412.08905

  2. [2]

    Colin Allen, Iva Smit, and Wendell Wallach. 2005. Artificial morality: Top-down, bottom-up, and hybrid approaches. Ethics and information technology, 7(3):149--155

  3. [3]

    S. Altman. 2017. https://finance.yahoo.com/news/im-silicon-valley-liberal-traveled-163400168.html?guccounter=1 I’m a silicon valley liberal, and i traveled across the country to interview 100 trump supporters—here’s what i learned. 2025

  4. [4]

    Alejandro Barredo Arrieta, Natalia D \' az-Rodr \' guez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador Garc \' a, Sergio Gil-L \'o pez, Daniel Molina, Richard Benjamins, et al. 2020. Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information fusion, 58:82--115

  5. [5]

    Dian Chen, Han Jun Yoon, Zelin Wan, Nithin Alluru, Sang Won Lee, Richard He, Terrence J Moore, Frederica F Nelson, Sunghyun Yoon, Hyuk Lim, et al. 2025. Advancing human-machine teaming: Concepts, challenges, and applications. arXiv preprint arXiv:2503.16518

  6. [6]

    Brian Christian. 2020. The alignment problem: Machine learning and human values. WW Norton & Company

  7. [7]

    K Robert Clarke and Richard M Warwick. 1994. Similarity-based testing for community pattern: the two-way layout with no replication. Marine biology, 118(1):167--176

  8. [8]

    Jacob W Crandall, Mayada Oudah, Tennom, Fatimah Ishowo-Oloko, Sherief Abdallah, Jean-Fran c ois Bonnefon, Manuel Cebrian, Azim Shariff, Michael A Goodrich, and Iyad Rahwan. 2018. Cooperating with machines. Nature communications, 9(1):233

  9. [9]

    Dellermann, P

    D. Dellermann, P. Ebel, M. Söllner, and J. M. Leimeister. 2019. https://doi.org/10.1007/s12599-019-00595-2 Hybrid intelligence . Business & Information Systems Engineering, 61(5):637--643

  10. [10]

    Giuseppe Desolda, Andrea Esposito, Rosa Lanzilotti, Antonio Piccinno, and Maria F Costabile. 2024. From human-centered to symbiotic artificial intelligence: a focus on medical applications. Multimedia Tools and Applications, pages 1--42

  11. [11]

    Leonard Dung. 2023. Current cases of ai misalignment and their implications for future risks. Synthese, 202(5):138

  12. [12]

    Beverley Garrigan, Anna L. R. Adlam, and Peter E. Langdon. 2018. https://doi.org/10.1016/j.dr.2018.06.001 Moral decision-making and moral development: Toward an integrative framework . Developmental Review, 49:80--100

  13. [13]

    Gemini, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Melvin Johnson, Ioannis Antonoglou, Julian Schrittwieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, et al. 2024. http://arxiv.org/...

  14. [14]

    Goode, M

    L. Goode, M. Calore, and Z Schiffer. 2024. https://www.wired.com/story/uncanny-valley-podcast-4-is-silicon-valley-libertarian/ Is silicon valley actually libertarian? 2025

  15. [15]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948

  16. [16]

    Vikas Hassija, Vinay Chamola, Atmesh Mahapatra, Abhinandan Singal, Divyansh Goel, Kaizhu Huang, Simone Scardapane, Indro Spinelli, Mufti Mahmud, and Amir Hussain. 2024. Interpreting black-box models: a review on explainable artificial intelligence. Cognitive Computation, 16(1):45--74

  17. [17]

    Yun Ti Hou and Hsin-Lu Chang. 2024. Building xai for fall prevention system: An evaluation of the potential of lime and shap in bridging trust gap

  18. [18]

    Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. 2024. Openai o1 system card. arXiv preprint arXiv:2412.16720

  19. [19]

    Mohammad Hossein Jarrahi. 2018. Artificial intelligence and the future of work: Human-ai symbiosis in organizational decision making. Business horizons, 61(4):577--586

  20. [20]

    Aravind Kumar Kalusivalingam, Amit Sharma, Neha Patel, and Vikram Singh. 2021. Leveraging shap and lime for enhanced explainability in ai-driven diagnostic systems. International Journal of AI and ML, 2(3)

  21. [21]

    Howard Levene. 1960. Robust tests for equality of variances. Contributions to probability and statistics, pages 278--292

  22. [22]

    Joseph CR Licklider. 2008. Man-computer symbiosis. IRE transactions on human factors in electronics, (1):4--11

  23. [23]

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437

  24. [24]

    Patrick E McKight and Julius Najab. 2010. Kruskal-wallis test. The corsini encyclopedia of psychology, pages 1--1

  25. [25]

    Alexander Meinke, Bronson Schoen, J \'e r \'e my Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. 2024. Frontier models are capable of in-context scheming. arXiv preprint arXiv:2412.04984

  26. [26]

    Anirban Mukherjee and Hsiao Chang. 2024. https://doi.org/10.2139/ssrn.4754533 Heuristic reasoning in ai: Instrumental use and mimetic absorption . SSRN Electronic Journal

  27. [27]

    J. Noller. 2024. https://doi.org/10.1057/s41599-024-03849-x Extended human agency: towards a teleological account of ai . Humanities and Social Sciences Communications, 11:1338

  28. [28]

    OpenAI. 2024. Hello gpt-4o. https://openai.com/index/hello-gpt-4o/. Accessed: 2025-01-29

  29. [29]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730--27744

  30. [30]

    Sebastian Raisch and Sebastian Krakowski. 2021. https://doi.org/10.5465/2018.0072 Artificial intelligence and management: The automation-augmentation paradox . Academy of Management Review

  31. [31]

    Leonardo Ranaldi and Giulia Pucci. 2023. When large language models contradict humans? large language models' sycophantic behaviour. arXiv preprint arXiv:2311.09410

  32. [32]

    Samuel Sanford Shapiro and Martin B Wilk. 1965. An analysis of variance test for normality (complete samples). Biometrika, 52(3-4):591--611

  33. [33]

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al. 2023. Towards understanding sycophancy in language models. arXiv preprint arXiv:2310.13548

  34. [34]

    Hua Shen, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Ziqiao Ma, Savvas Petridis, Yi-Hao Peng, Li Qiwei, et al. 2024. Towards bidirectional human-ai alignment: A systematic review for clarifications, framework, and future directions. arXiv preprint arXiv:2406.09264

  35. [35]

    Ben Shneiderman. 2020. Human-centered artificial intelligence: Three fresh ideas. AIS Transactions on Human-Computer Interaction, 12(3), 109-124. https://doi.org/10.17705/1thci.00131

  36. [36]

    Carles Sierra, Nardine Osman, Pablo Noriega, Jordi Sabater-Mir, and Antoni Perell \'o . 2021. Value alignment: a formal approach. arXiv preprint arXiv:2110.09240

  37. [37]

    Simkute, L

    A. Simkute, L. Tankelevitch, V. Kewenig, A. E. Scott, A. Sellen, and S. Rintel. 2024. https://doi.org/10.1080/10447318.2024.2405782 Ironies of generative ai: Understanding and mitigating productivity loss in human--ai interaction . International Journal of Human--Computer Interaction, 41(5):2898--2919

  38. [38]

    Vaccaro, A

    M. Vaccaro, A. Almaatouq, and T. Malone. 2024. https://doi.org/10.1038/s41562-024-02024-1 When combinations of humans and ai are useful: A systematic review and meta-analysis . Nature Human Behaviour, 8:2293--2303

  39. [39]

    H James Wilson and Paul R Daugherty. 2018. Collaborative intelligence: Humans and ai are joining forces. Harvard business review, 96(4):114--123

  40. [40]

    Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah Goodman. 2023. Interpretability at scale: Identifying causal mechanisms in alpaca. Advances in neural information processing systems, 36:78205--78226

  41. [41]

    Eliezer Yudkowsky. 2016. The ai alignment problem: why it is hard, and where to start. Symbolic Systems Distinguished Speaker, 4(1)

  42. [42]

    Zahedi and Subbarao Kambhampati

    Z. Zahedi and Subbarao Kambhampati. 2021. Human-ai symbiosis: A survey of current approaches. ArXiv, abs/2103.09990

  43. [43]

    Yi Zeng, Enmeng Lu, and Kang Sun. 2025 a . Principles on symbiosis for natural life and living artificial intelligence. AI and Ethics, 5(1):81--86

  44. [44]

    Yi Zeng, Feifei Zhao, Yuwei Wang, Enmeng Lu, Yaodong Yang, Lei Wang, Chao Liu, Yitao Liang, Dongcheng Zhao, Bing Han, Haibo Tong, Yao Liang, Dongqi Liang, Kang Sun, Boyuan Chen, and Jinyu Fan. 2025 b . http://arxiv.org/abs/2504.17404 Super co-alignment of human and ai for sustainable symbiotic society . arXiv preprint

  45. [45]

    Yutong Zhang, Dora Zhao, Jeffrey T Hancock, Robert Kraut, and Diyi Yang. 2025. The rise of ai companions: How human-chatbot relationships influence well-being. arXiv preprint arXiv:2506.12605

  46. [46]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  47. [47]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

This paper was first reviewed by deepseek-v4-flash on August 4, 2026.