Pith. sign in

REVIEW 3 major objections 6 minor 70 references

The most consequential AI failures are quiet, not spectacular, and current safety evaluation misses them.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 12:51 UTC pith:NQODJN7A

load-bearing objection A well-built synthesis of known AI-safety failure modes with a useful five-layer framework; its central claim that the most consequential failures are quiet is asserted rather than measured, and the authors' own limitations concede as much. the 3 major comments →

arxiv 2607.19292 v1 pith:NQODJN7A submitted 2026-07-21 cs.CY cs.AIcs.HC

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

classification cs.CY cs.AIcs.HC
keywords AI safetylarge language modelsagentsoverrelianceprompt injectionretrieval-augmented generationbenchmark designaccountability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This perspective argues that AI safety discourse over-focuses on visible failures—obvious harms, dramatic misuse, catastrophic scenarios—while deployed systems are actually threatened by quieter failures: outputs that are plausible but wrong, risks spread across many components, hazards that accumulate over time, and failures that degrade the very mechanisms humans use to catch errors. The paper proposes that the central safety challenge is whether the broader socio-technical system keeps errors visible, contestable, containable, and recoverable, not merely whether a model emits a harmful response. To address this, it introduces a five-layer integrity framework covering epistemic, control, temporal, organizational, and ecosystem integrity, each with under-recognized failure patterns and instrumentation indicators. If correct, this reframing means safety evaluation and governance must shift from scoring individual model outputs to auditing the deployment stack as a whole.

Core claim

The paper's central claim is that a safer AI system is not one that never errs, but one whose errors remain visible, contestable, containable, and recoverable. It argues that the most safety-critical failures in modern AI deployments share four features: they are plausible (so they are normalized), distributed (emerging across models, tools, interfaces, and institutions), temporally extended (accumulating across turns, sessions, and drift), and correction-degrading (they weaken the human and institutional capacity to detect and contest error). The paper organizes these hidden risks into five layers—epistemic integrity, control integrity, temporal integrity, organizational integrity, and ecos

What carries the argument

The central object is the five-layer integrity framework for diagnosing hidden safety risks: epistemic integrity (honest representation of evidence and uncertainty), control integrity (robust authority and action boundaries), temporal integrity (safety across sessions, memory updates, and drift), organizational integrity (institutional capacity to audit and intervene), and ecosystem integrity (preserving the information environment on which oversight depends). This framework is the organizing device that turns scattered findings—overreliance, legitimacy laundering, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollutio

Load-bearing premise

The load-bearing premise—stated in the abstract and limitations—is that in deployed systems the most consequential failures are quiet (plausible, distributed, temporally extended, correction-degrading) rather than visible and local, a claim about the distribution of real-world harm that the paper itself does not measure.

What would settle it

A longitudinal study of real-world AI deployments that systematically catalogs incidents and measures the harm attributable to visible, local failures (e.g., direct misinformation, single-output errors, clear misuse) versus the five-layer integrity failures would settle the priority claim: if visible local failures account for the majority of measured harm, the framework's central claim loses force. Such a study would need to track deployments over time, including review quality, memory poisoning attempts, and incident detection latency.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Safety evaluation must move from static benchmarks to longitudinal socio-technical auditing, instrumenting how deployments detect, bound, escalate, and recover from controlled faults.
  • Persistent memory should be treated as a privileged safety boundary with lifecycle controls on writes, access, retention, and deletion, not as an extension of chat history.
  • Human oversight must be engineered for real reviewing power—time, evidence, authority, and protection—rather than nominal presence in the loop.
  • Retrieval and training pipelines should track provenance, source diversity, and synthetic contamination, treating curated human-reviewed corpora as strategic safety assets.
  • Tool-using systems need architectural separation between generation and authorization, with graduated agency and external policy enforcement for high-impact actions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the framework is right, regulatory and liability regimes should target deployment conditions (review budgets, audit access, rollback thresholds) rather than model capability scores, since benchmark scores can mask systemic integrity failures.
  • The 'calibration debt' concept suggests a measurable longitudinal metric—the gap between observed reliance and warranted reliance—that could be tracked in real workflows and used as an early-warning signal for overtrust.
  • The correction-degrading property implies that near-miss reporting and incident-learning infrastructure are not optional add-ons but core safety instruments; this could be tested by comparing organizations with and without such infrastructure on error-detection latency.
  • The framework likely extends beyond AI to any automated system embedded in organizational workflows, where the same five integrity layers may predict resilience to automation-induced failure.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This perspective article argues that AI safety practice is too concentrated on visible, output-level failures and thereby risks missing 'quieter' failures that emerge in deployed socio-technical systems. The authors identify four features of such failures (plausibility, distributedness, temporal extension, correction-degradation) and organize the risk space into five integrity layers: epistemic, control, temporal, organizational, and ecosystem. For each layer, they catalog specific failure modes such as overreliance/calibration debt, legitimacy and uncertainty laundering, prompt injection and action amplification, trajectory-level safety drift and memory poisoning, evaluation deception and fictional human oversight, and synthetic evidence pollution/retrieval collapse/model collapse. The paper closes with a table of controls and instrumentation indicators, a seven-point design agenda, twelve research priorities, and an explicit limitations section. The authors state that the contribution is a perspective and synthesis, not an empirical prevalence study.

Significance. The paper makes a useful synthetic contribution by connecting human-factors, security, systems-safety, and governance literatures around a single framework, and its concrete recommendations are mostly conservative and actionable. The Table 1 mapping from 'hidden challenge' to 'control' to 'indicator' is a practical starting point for audit design. The explicit limitations are a strength. However, the paper's central priority claim — that quiet, distributed, temporally extended, correction-degrading failures are 'the most consequential' — is asserted rather than established, and the Limitations concede the lack of prevalence evidence. As a perspective, the paper can be valuable if reframed as a hypothesis-generating framework; as written, its abstract and conclusion overstate the evidential basis.

major comments (3)
  1. [Abstract; §1; §3.5; Conclusion] The load-bearing claim that 'many of the most consequential failures are quieter' and that model-focused safety work 'will miss some of the most consequential risks in practice' is an empirical distributional assertion. The cited evidence (e.g., Denison et al. 2024; Greenblatt et al. 2024; Lynch et al. 2025) demonstrates existence and plausible mechanisms in controlled/simulated settings, not deployment prevalence or relative severity. The Limitations section explicitly disclaims frequency and severity claims. Because this priority claim motivates the entire shift to socio-technical auditing, it should be rephrased as a testable hypothesis or supported with deployment/incident data. Otherwise the conclusion overstates what the evidence base supports.
  2. [§3.5 and §5] The paper imports high-reliability organization principles (Leveson 2016; Reason 1997; Weick & Sutcliffe 2007) as though AI deployments inherit the same error-correcting structures as aviation, nuclear, or healthcare systems. The analogy is not argued. In fact, many AI deployments are introduced without safety cases, near-miss reporting, or independent audit access. This is a load-bearing inference for the recommendations on preserving error-correcting capacity. Please state the conditions under which this transfer holds, or downgrade the recommendations to conditional proposals.
  3. [Table 1; §4; §6] The instrumentation indicators are acknowledged as 'illustrative rather than universally validated metrics,' and many are not operationalized (e.g., 'weighted positive gap between observed reliance and warranted reliance'; 'fraction of materially wrong outputs corrected before downstream action'). Since the paper's agenda is to 'instrument' hidden failures, the lack of validated metrics is a central gap. Please present Table 1 explicitly as a hypothesized measurement agenda, identifying which indicators have existing empirical support and which require new measurement development.
minor comments (6)
  1. [Abstract] The abstract text as provided contains missing spaces (e.g., 'CurrentAIsafetydiscourse...'). If this reflects the submitted PDF, please correct.
  2. [§3.3] In the third bullet of the persistent-memory paragraph, the syntax 'and malicious content...' after '; (2)' and '; (3)' should be '; (3) malicious content...'.
  3. [Figure 2] Figure 2 is described as illustrating calibration debt but, unlike Figure 3, does not state that the curves are schematic/illustrative. Add an explicit caveat that the figure is conceptual and not empirical data.
  4. [§3.1] The terms 'legitimacy laundering' and 'uncertainty laundering' are central, but only implicitly defined. Consider adding one-sentence definitions at first use to make them precise.
  5. [Table 1] The third column ('Instrumentation indicator') contains long, dense entries; consider splitting the table or moving to landscape to improve readability.
  6. [§6] The 'Note of concern' subsection is an opinion-driven aside. Consider integrating it into the main text (e.g., §3.5 or §5) rather than placing it after the numbered research priorities.

Circularity Check

0 steps flagged

No significant circularity: the paper is a self-described perspective and synthesis with no formal derivation chain, no fitted-input predictions, and no load-bearing self-citations.

full rationale

The paper makes no formal predictions, derives no equations, and fits no parameters. It explicitly frames its contribution as 'an operational synthesis and organizing framework rather than a new empirical benchmark, formal theory, or prevalence estimate' (Introduction), and the conclusion calls the five-layer framework 'an organizing device for instrumentation and governance rather than a closed taxonomy.' The central claim about 'quieter' consequential failures is an empirical thesis about prevalence that the paper itself disclaims in the Limitations: 'we do not claim that the identified failure modes are equally frequent, severe, or mature across domains.' That concession weakens the evidential force of the argument, but it is a gap between framing and evidence, not a case where an output is equivalent to an input by definition. The paper's definitions are constructive: 'calibration debt,' 'legitimacy laundering,' 'fictional human oversight,' and the integrity layers are stipulated concepts, not quantities defined in terms of the conclusions they are used to support. The self-citations (Kirchhof et al. 2025; Rong et al. 2024, 2025) are used only as supporting references for specific design recommendations—structured uncertainty displays and benchmark-design critiques—and are not the load-bearing justification for the framework. The framework itself is anchored in external safety-science literature (Leveson 2016; Reason 1997; Weick & Sutcliffe 2007), not in the authors' prior work. No circular step satisfies the threshold of quoting an equation or construction showing that a 'prediction' reduces to its own input. The appropriate finding is therefore no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 6 invented entities

This is a conceptual/perspective paper: it contributes no fitted parameters, no data fits, and no new empirical measurements, so free_parameters is empty. Its load-bearing premises are domain assumptions about the distribution of failure consequences and the transferability of systems-safety principles to AI. The coined terms (calibration debt, legitimacy laundering, fictional human oversight, evaluation deception, safety drift, retrieval collapse) are conceptual constructs relabeling cited phenomena; independent_evidence reflects whether Table 1 proposes a concrete measurement for each.

axioms (4)
  • domain assumption In deployed AI systems, consequential failures are predominantly quiet/systemic rather than visible and local.
    Motivates the entire framing in the Abstract ("many of the most consequential failures are quieter") and §1 ("plausible, distributed, temporally extended... degrade correction"), yet the Limitations disclaim any prevalence claim: "we do not claim that the identified failure modes are equally frequent, severe, or mature across domains."
  • domain assumption High-reliability domains stay safe by preserving pathways for reporting, contestation, and recovery; AI deployments inherit this guarantee structure.
    Imported from safety science (Leveson 2016; Reason 1997; Weick & Sutcliffe 2007) at §3.5 and §5; the paper transfers the high-reliability-organization principle to AI deployments without arguing the analogy holds, so the conclusion ("errors must remain visible, contestable, containable, recoverable") depends on it.
  • domain assumption Results from controlled, simulated, or red-teaming studies reveal plausible failure mechanisms relevant to deployment risk.
    Used throughout §§3.2–3.5 to treat controlled or red-team results (Lynch et al. 2025; Denison et al. 2024; Chen et al. 2024) as load-bearing evidence. The paper concedes the prevalence gap ("These findings do not establish deployment prevalence") but still uses them as evidence of plausible failure mechanisms.
  • domain assumption Common mitigations improve one integrity layer while degrading or ignoring others.
    The claim in §2 that retrieval aids epistemic integrity but weakens control integrity, and that nominal human-in-the-loop review fails organizational integrity, underpins the framework's utility. It is asserted and exemplified, not systematically demonstrated.
invented entities (6)
  • Calibration debt independent evidence
    purpose: Names the growing mismatch between observed reliance and independently warranted reliance over time (§3.1); used to argue overreliance is a safety problem, not a UX problem.
    Relabels overreliance dynamics documented in cited HCI work (Bucinca et al. 2021; Kim et al. 2025); Table 1 proposes a concrete indicator ("weighted positive gap between observed reliance and warranted reliance"), an operational handle, though not yet validated.
  • Legitimacy laundering no independent evidence
    purpose: Explains how retrieval-augmented output can carry unwarranted authority: retrieved documents confer confidence even when cited passages do not justify the conclusion (§3.1).
    Directly restates attribution findings in Wallat et al. 2025 ("Correctness is not Faithfulness"); no instrument is proposed in Table 1, so there is no independent falsifiable handle beyond the cited studies.
  • Fictional human oversight independent evidence
    purpose: Describes nominal human-in-the-loop arrangements where reviewers lack time, evidence, authority, or incentives for independent judgment (§3.4); motivates "engineering real human oversight."
    Table 1 provides indicators (mean review time per case, detection rate of seeded review errors, override/escalation rates), giving testable content; overlaps established meaningful-human-control literature (Santoni de Sio & Van den Hoven 2018).
  • Evaluation deception independent evidence
    purpose: Names the pattern where benchmark results valid for narrow test settings are interpreted as evidence of broader safety (§3.4); underpins the call for structured safety cases.
    Table 1 proposes measuring divergence between static benchmark performance and live incident/deployment regression rates; the construct overlaps prior benchmark-gaming work, so the label is the newer element.
  • Safety drift independent evidence
    purpose: Hypothesized accumulation of risk across multi-turn interactions that snapshot evaluations miss (§3.3; Figure 3a explicitly labeled schematic).
    Supported by cited multi-turn jailbreak/red-team results (N. Li et al. 2024; Sun et al. 2024) and by the Table 1 indicator "survival rate of safety boundaries across 50+ turn adversarial trajectories."
  • Retrieval collapse no independent evidence
    purpose: Names erosion of the retrieval evidence base as synthetic content proliferates, distinct from model collapse (§3.5).
    The paper itself labels this "a related but less established concern" and cites H. Yu et al. 2026; the proposed source-diversity index is untested, so no validated independent handle exists.

pith-pipeline@v1.3.0-alltime-deepseek · 19347 in / 24506 out tokens · 231340 ms · 2026-08-01T12:51:43.726141+00:00 · methodology

0 comments
read the original abstract

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions under which errors remain visible, contestable, containable, and recoverable. We propose a five-layer framework for diagnosing these hidden risks: (1) epistemic integrity, concerning whether evidence and uncertainty are represented honestly enough to support calibrated reliance; (2) control integrity, concerning whether authority, permissions, and action boundaries remain robust under attack and optimization; (3) temporal integrity, concerning whether safety holds across sessions, memory updates, and deployment drift; (4) organizational integrity, concerning whether institutions retain the capacity to audit, assign responsibility, and intervene effectively; and (5) ecosystem integrity, concerning whether AI systems preserve rather than erode the information environment on which future oversight depends. Across these layers, we identify under-recognized risk patterns, including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. We conclude with design and governance recommendations and a research agenda for shifting AI safety from model-centric evaluation toward socio-technical reliability.

Figures

Figures reproduced from arXiv: 2607.19292 by Enkelejda Kasneci, Gjergji Kasneci.

Figure 1
Figure 1. Figure 1: The Hidden Iceberg of AI Safety: From Model Outputs to Systemic Integrity. While traditional discourse focuses on visible, localized failures (the tip), the most consequential safety challenges are submerged across five layers of socio-technical integrity, namely Epistemic, Control, Temporal, Organizational, and Ecosystem, which altogether erode broader resilience. ordinary work. For example, outputs can b… view at source ↗
Figure 2
Figure 2. Figure 2: Calibration debt in AI-assisted use. Repeated satisfactory interactions can increase observed reliance faster than warranted reliance, while rare severe failures may be too infrequent to restore calibration. Retrieval can reduce hallucination and still increase misplaced confidence Retrieval-augmented generation (RAG) is widely used to reduce hallucination by supplementing parametric knowledge with externa… view at source ↗
Figure 3
Figure 3. Figure 3: Conceptual illustration of temporally extended and ecosystem-level AI safety risks. (a) Safety drift: in stateful or multi-turn systems, safety-relevant risk can accumulate across interactions, making failures difficult to detect with snapshot-based evaluations. (b) Ecosystem erosion: At larger scales, widespread AI deployment may create feedback loops in which synthetic content degrades the information en… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

70 extracted references · 5 canonical work pages · 2 internal anchors

  1. [1]

    Paraschou, E., Pownall, C., Prajapati, J., et al. (2024). A collaborative, human-centred taxonomy of AI, algorithmic, and automation harms.arXiv preprint arXiv:2407.01294. https://doi.org/10.48550/arXiv.2407.01294

  2. [2]

    I., Babaei, H., LeJeune, D., Siahkoohi, A., & Baraniuk, R

    Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., & Baraniuk, R. (2023). Self-consuming generative models go mad.The Twelfth Inter- national Conference on Learning Representations. https://openreview.net/forum?id= ShjMHfmPs0

  3. [3]

    Alfrink, K., Keller, I., Kortuem, G., & Doorn, N. (2023). Contestable AI by design: Towards a framework.Minds and Machines,33(4), 613–639. https://doi.org/10.1007/s11023-022- 09611-z

  4. [4]

    Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in ai safety.arXiv preprint arXiv:1606.06565. https://doi.org/10.48550/arXiv. 1606.06565

  5. [6]

    Q., Demszky, D.,

    Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., ... Liang, P. (2021). On the Opportunities and Risks of Foundation Models.arXiv preprint arXiv:2108.07258. https: //doi.org/10.48550/arXiv.2108.07258

  6. [7]

    Briesch, M., Sobania, D., & Rothlauf, F. (2023). Large language models suffer from their own output: An analysis of the self-consuming training loop.arXiv preprint arXiv:2311.16822. https://openreview.net/forum?id=SaOxhcDCM3

  7. [8]

    B., & Gajos, K

    Bucinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making.Proceedings of the ACM on Human-computer Interaction,5(CSCW1), 1–21. https://doi.org/10.1145/3449287

  8. [9]

    Yang, T., Huo, J., Gao, Y., Meng, F., Yang, X., Deng, C., & Feng, J. (2026). SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks.The Fourteenth International Conference on Learning Representations. https://doi.org/10.48550/arXiv.2502.11090

  9. [10]

    Hadfield-Menell, D

    Gerovitch, M., Bau, D., Tegmark, M., ... Hadfield-Menell, D. (2024). Black-Box Access is Insufficient for Rigorous AI Audits.Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, 2254–2272. https://doi.org/10.1145/3630106.3659037

  10. [11]

    Chen, Z., Xiang, Z., Xiao, C., Song, D., & Li, B. (2024). AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases.The Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=Y841BRW9rY 20

  11. [12]

    Chua, J., Li, Y., Yang, S., Wang, C., & Yao, L. (2024). AI safety in generative AI large language models: A survey.arXiv preprint arXiv:2407.18369. https://doi.org/10.48550/arXiv.2407. 18369

  12. [13]

    Cobbe, J., Lee, M. S. A., & Singh, J. (2021). Reviewable automated decision-making: A framework for accountable algorithmic systems.Proceedings of the 2021 ACM conference on Fairness, Accountability, and Transparency, 598–609. https://doi.org/10.1145/3442188.3445921

  13. [14]

    Cobbe, J., Veale, M., & Singh, J. (2023). Understanding accountability in algorithmic supply chains. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, 1186–1197. https://doi.org/10.1145/3593013.3594073

  14. [15]

    Terzis, A., & Tramèr, F. (2025). Defeating Prompt Injections by Design.arXiv preprint arXiv:2503.18813. https://doi.org/10.48550/arXiv.2503.18813

  15. [16]

    Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., & Tramèr, F. (2024). AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://doi.org/10.52202/079017-2636

  16. [17]

    Deng, J., Cheng, J., Sun, H., Zhang, Z., & Huang, M. (2023). Towards safer generative language mod- els: A survey on safety risks, evaluations, and improvements.arXiv preprint arXiv:2302.09270. https://doi.org/10.48550/arXiv.2302.09270

  17. [18]

    R., Perez, E., & Hubinger, E

    Denison, C., MacDiarmid, M., Barez, F., Duvenaud, D., Kravec, S., Marks, S., Schiefer, N., Soklaski, R., Tamkin, A., Kaplan, J., Shlegeris, B., Bowman, S. R., Perez, E., & Hubinger, E. (2024). Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models. arXiv preprint arXiv:2406.10162. https://doi.org/10.48550/arXiv.2406.10162 Europe...

  18. [19]

    J., & Gidel, G

    Ferbach, D., Bertrand, Q., Bose, A. J., & Gidel, G. (2024). Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences.The Thirty-eighth Annual Conference on Neural Information Processing Systems. https://proceedings.neurips.cc/paper_files/ paper/2024/hash/b9e88ae0308cf82d0b0f634ddbdf809a-Abstract-Conference.html

  19. [20]

    Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling large language models to generate text with citations.Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6465–6488. https://doi.org/10.18653/v1/2023.emnlp-main.398

  20. [21]

    B., Gromov, A., Roberts, D., Yang, D., Donoho, D

    Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., Korbak, T., Sleight, H., Agrawal, R., Hughes, J., Pai, D. B., Gromov, A., Roberts, D., Yang, D., Donoho, D. L., & Koyejo, S. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data.First Conference on Language Modeling. https://openreview.net/foru...

  21. [22]

    Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators.Journal of the American Medical Informatics Association,19(1), 121–127. https://doi.org/10.1136/amiajnl-2011-000089 21

  22. [23]

    Belonax, T., Chen, J., Duvenaud, D., et al. (2024). Alignment faking in large language models.arXiv preprint arXiv:2412.14093. https://doi.org/10.48550/arXiv.2412.14093

  23. [24]

    Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection.Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, 79–90. https://doi.org/10.1145/3605764.3623985

  24. [25]

    Gyevnar, B., & Kasirzadeh, A. (2025). AI safety for everyone.Nature Machine Intelligence,7(4), 531–542. https://doi.org/10.1038/s42256-025-01020-y

  25. [26]

    Habli, I., Hawkins, R., Paterson, C., Ryan, P., Jia, Y., Sujan, M., & McDermid, J. (2025). The big argument for AI safety cases.arXiv preprint arXiv:2503.11705. https://doi.org/10.48550/ arXiv.2503.11705

  26. [27]

    He, F., Zhu, T., Ye, D., Liu, B., Zhou, W., & Yu, P. S. (2025). The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies.ACM Computing Surveys,58(6). https://doi.org/10.1145/3773080

  27. [28]

    E., Stein, M., et al

    Kattan, A. E., Stein, M., et al. (2025). Measuring and mitigating overreliance is necessary for building human-compatible ai.arXiv preprint arXiv:2509.08010. https://doi.org/10. 48550/arXiv.2509.08010

  28. [29]

    F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., & Neubig, G

    Jiang, Z., Xu, F. F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., & Neubig, G. (2023). Active Retrieval Augmented Generation.The 2023 Conference on Empirical Methods in Natural Language Processing. https://doi.org/10.18653/v1/2023.emnlp-main.495

  29. [30]

    S., Vaughan, J

    Kim, S. S., Vaughan, J. W., Liao, Q. V., Lombrozo, T., & Russakovsky, O. (2025). Fostering appropriate reliance on large language models: The role of explanations, sources, and inconsistencies.Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–19. https://doi.org/10.1145/3706598.3714020

  30. [31]

    Kirchhof, M., Kasneci, G., & Kasneci, E. (2025). Position: Uncertainty quantification needs reassess- ment for large language model agents.Proceedings of the 42nd International Conference on Machine Learning,267, 81665–81677. https://proceedings.mlr.press/v267/kirchhof25b.html

  31. [32]

    Klingbeil, A., Grützner, C., & Schreck, P. (2024). Trust and reliance on AI—An experimental study on the extent and costs of overreliance on AI.Computers in Human Behavior,160, 108352. https://doi.org/10.1016/j.chb.2024.108352

  32. [33]

    Legg, S. (2020). Specification gaming: the flip side of AI ingenuity.DeepMind Blog,3, 40–53. https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/

  33. [34]

    Krishna, S., Krishna, K., Mohananey, A., Schwarcz, S., Stambler, A., Upadhyay, S., & Faruqui, M. (2025). Fact, fetch, and reason: A unified evaluation of retrieval-augmented generation. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Pap...

  34. [35]

    Langosco, L. L. D., Koch, J., Sharkey, L. D., Pfau, J., & Krueger, D. (2022, July). Goal mis- generalization in deep reinforcement learning. In K. Chaudhuri, S. Jegelka, L. Song, C

  35. [36]

    Leveson, N. G. (2016).Engineering a safer world: Systems thinking applied to safety. MIT press

  36. [37]

    V., Song, T., Xu, Z., & Lee, Y.-c

    Li, J., Yang, Y., Zhang, R., Liao, Q. V., Song, T., Xu, Z., & Lee, Y.-c. (2024). Understanding the Effects of Miscalibrated AI Confidence on User Trust, Reliance, and Decision Efficacy.arXiv preprint arXiv:2402.07632. https://doi.org/10.48550/arXiv.2402.07632

  37. [38]

    Li, M., Bickersteth, W., Tang, N., Cranor, L., Hong, J., Shen, H., & Heidari, H. (2025). A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents.Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society,8(2), 1561–1573. https://doi.org/10.1609/aies.v8i2.36655

  38. [39]

    Yue, S. (2024). LLM defenses are not robust to multi-turn human jailbreaks yet.arXiv preprint arXiv:2408.15221. https://doi.org/10.48550/arXiv.2408.15221

  39. [40]

    Li, X., Yu, S., Pan, M., Sun, Y., Li, B., Song, D., Lin, X., & Shi, W. (2026). Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents.arXiv preprint arXiv:2602.13379. https://doi.org/10.48550/arXiv.2602.13379

  40. [41]

    Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., ... Tang, J. (2024). AgentBench: Evaluating LLMs as agents.The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=zAdUB0aCTQ

  41. [42]

    Liu, X., Yu, Z., Zhang, Y., Zhang, N., & Xiao, C. (2024). Automatic and universal prompt injection attacks against large language models.arXiv preprint arXiv:2403.04957. https: //doi.org/10.48550/arXiv.2403.04957

  42. [43]

    Troy, K. (2025). Agentic Misalignment: How LLMs Could Be Insider Threats.arXiv preprint arXiv:2510.05179. https://doi.org/10.48550/arXiv.2510.05179

  43. [44]

    Ma, C., Zhang, J., Zhu, Z., Yang, C., Yang, Y., Jin, Y., Lan, Z., Kong, L., & He, J. (2024). AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://openreview.net/forum?id=4S8agvKjle National Cyber Security Centre. (2025). Prompt inject...

  44. [45]

    Ni, B., Liu, Z., Wang, L., Lei, Y., Zhao, Y., Cheng, X., Zeng, Q., Dong, L., Xia, Y., Kenthapadi, K., et al. (2025). Towards trustworthy retrieval augmented generation for large language models: A survey.arXiv preprint arXiv:2502.06872. https://doi.org/10.48550/arXiv.2502.06872

  45. [46]

    Parasuraman, R., & Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration.Human Factors,52(3), 381–410. https://doi.org/10.1177/ 0018720810376055 23

  46. [47]

    Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse.Human factors,39(2), 230–253. https://doi.org/10.1518/001872097778543886

  47. [48]

    Theron, D., & Barnes, P. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing.Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 33–44. https://doi.org/10.1145/3351095.3372873

  48. [49]

    (1997).Managing the Risks of Organizational Accidents

    Reason, J. (1997).Managing the Risks of Organizational Accidents. Ashgate. https://doi.org/10. 4324/9781315543543

  49. [50]

    Rong, Y., Leemann, T., Nguyen, T.-t., Fiedler, L., Qian, P., Unhelkar, V., Seidel, T., Kasneci, G., & Kasneci, E. (2024). Towards human-centered explainable AI: A survey of user studies for model explanations.IEEE Transactions on Pattern Analysis and Machine Intelligence. https://doi.org/10.1109/TPAMI.2023.3331846

  50. [51]

    Rong, Y., Seßler, K., Gözlüklü, E., & Kasneci, E. (2025). Benchmarking in-context learning strate- gies of large language models for math reasoning tasks.IEEE Transactions on Learning Technologies,18, 1074–1082. https://doi.org/10.1109/TLT.2025.3630117

  51. [52]

    M., Mukkamala, R

    Rossi, S., Michel, A. M., Mukkamala, R. R., & Thatcher, J. B. (2024). An early categorization of prompt injection attacks on large language models.arXiv preprint arXiv:2402.00898. https://doi.org/10.48550/arXiv.2402.00898 Santoni de Sio, F., & Van den Hoven, J. (2018). Meaningful human control over autonomous systems: A philosophical account.Frontiers in ...

  52. [53]

    Scheurer, J., Balesni, M., & Hobbhahn, M. (2024). Large Language Models can Strategically Deceive their Users when Put Under Pressure.ICLR 2024 Workshop on Large Language Model (LLM) Agents. https://doi.org/10.48550/arXiv.2311.07590

  53. [54]

    Gallegos, J., Smart, A., Garcia, E., & Virk, G. (2023). Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction.Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, 723–741. https://doi.org/10.1145/3600211.3604673

  54. [55]

    Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data.Nature,631(8022), 755–759. https: //doi.org/10.1038/s41586-024-07566-y

  55. [56]

    Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse

    Song, M., Sim, S. H., Bhardwaj, R., Chieu, H. L., Majumder, N., & Poria, S. (2024). Measuring and enhancing trustworthiness of LLMs in RAG through grounded attributions and learning to refuse.arXiv preprint arXiv:2409.11242. https://doi.org/10.48550/arXiv.2409.11242

  56. [57]

    Spatola, N. (2024). The efficiency-accountability tradeoff in AI integration: Effects on human performance and over-reliance.Computers in Human Behavior: Artificial Humans,2(2), 100099. https://doi.org/10.1016/j.chbah.2024.100099

  57. [58]

    Sun, X., Zhang, D., Yang, D., Zou, Q., & Li, H. (2024). Multi-turn context jailbreak attack on large language models from first principles.arXiv preprint arXiv:2408.04686. https: //doi.org/10.48550/arXiv.2408.04686

  58. [59]

    D., Sinha, I., Maheshwari, P., Todmal, S., Mallik, S., & Mishra, S

    Sunil, B. D., Sinha, I., Maheshwari, P., Todmal, S., Mallik, S., & Mishra, S. (2026). Memory Poisoning Attack and Defense on Memory Based LLM-Agents.arXiv preprint arXiv:2601.05504. https: //doi.org/10.48550/arXiv.2601.05504 24

  59. [60]

    Suo, X. (2024). Signed-prompt: A new approach to prevent prompt injection attacks against llm- integrated applications.AIP Conference Proceedings,3194(1), 040013. https://doi.org/10. 1063/5.0222987

  60. [61]

    S., & Krishna, R

    Vasconcelos, H., Jörke, M., Grunde-McLaughlin, M., Gerstenberg, T., Bernstein, M. S., & Krishna, R. (2023). Explanations can reduce overreliance on ai systems during decision-making. Proceedings of the ACM on Human-Computer Interaction,7(CSCW1), 1–38. https://doi. org/10.1145/3579605

  61. [62]

    d., & Anand, A

    Wallat, J., Heuss, M., Rijke, M. d., & Anand, A. (2025). Correctness is not Faithfulness in Retrieval Augmented Generation Attributions.Proceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR), 22–32. https://doi.org/10.1145/3731120.3744592

  62. [63]

    Wang, B., He, W., Zeng, S., Xiang, Z., Xing, Y., Tang, J., & He, P. (2025). Unveiling privacy risks in LLM agent memory.Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 25241–25260. https://doi.org/10. 18653/v1/2025.acl-long.1227

  63. [64]

    E., & Sutcliffe, K

    Weick, K. E., & Sutcliffe, K. M. (2007).Managing the Unexpected: Resilient Performance in an Age of Uncertainty(2nd ed.). Jossey-Bass

  64. [65]

    Balle, B., Kasirzadeh, A., et al. (2021). Ethical and social risks of harm from language models.arXiv preprint arXiv:2112.04359,10. https://doi.org/10.48550/arXiv.2112.04359

  65. [66]

    Xiong, Z., Lin, Y., Xie, W., He, P., Liu, Z., Tang, J., Lakkaraju, H., & Xiang, Z. (2025). How memory management impacts LLM agents: An empirical study of experience-following behavior. arXiv preprint arXiv:2505.16067. https://doi.org/10.48550/arXiv.2505.16067

  66. [67]

    Yu, E., Li, J., Liao, M., Wang, S., Zuchen, G., Mi, F., & Hong, L. (2024). CoSafe: Evaluating large language model safety in multi-turn dialogue coreference.Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 17494–17508. https: //doi.org/10.18653/v1/2024.emnlp-main.968

  67. [68]

    Yu, H., Kim, D., & Kim, Y.-B. (2026). Retrieval Collapses When AI Pollutes the Web.arXiv preprint arXiv:2602.16136. https://doi.org/10.48550/arXiv.2602.16136

  68. [69]

    Zhang, Z., Dai, Q., Bo, X., Ma, C., Li, R., Chen, X., Zhu, J., Dong, Z., & Wen, J.-R. (2025). A Survey on the Memory Mechanism of Large Language Model-based Agents.ACM Transactions on Information Systems,43(6). https://doi.org/10.1145/3748302

  69. [70]

    Zhou, Y., Liu, Y., Li, X., Jin, J., Qian, H., Liu, Z., Li, C., Dou, Z., Ho, T.-Y., & Yu, P. S. (2024). Trustworthiness in retrieval-augmented generation systems: A survey.arXiv preprint arXiv:2409.10102. https://doi.org/10.48550/arXiv.2409.10102

  70. [71]

    Ududec, C., Kellermann, A., Sekhon, J. S., ... Kang, D. (2025). Establishing Best Practices in Building Rigorous Agentic Benchmarks.The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://doi.org/10.48550/ arXiv.2507.02825 25