REVIEW 3 major objections 6 minor 70 references
The most consequential AI failures are quiet, not spectacular, and current safety evaluation misses them.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 12:51 UTC pith:NQODJN7A
load-bearing objection A well-built synthesis of known AI-safety failure modes with a useful five-layer framework; its central claim that the most consequential failures are quiet is asserted rather than measured, and the authors' own limitations concede as much. the 3 major comments →
The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that a safer AI system is not one that never errs, but one whose errors remain visible, contestable, containable, and recoverable. It argues that the most safety-critical failures in modern AI deployments share four features: they are plausible (so they are normalized), distributed (emerging across models, tools, interfaces, and institutions), temporally extended (accumulating across turns, sessions, and drift), and correction-degrading (they weaken the human and institutional capacity to detect and contest error). The paper organizes these hidden risks into five layers—epistemic integrity, control integrity, temporal integrity, organizational integrity, and ecos
What carries the argument
The central object is the five-layer integrity framework for diagnosing hidden safety risks: epistemic integrity (honest representation of evidence and uncertainty), control integrity (robust authority and action boundaries), temporal integrity (safety across sessions, memory updates, and drift), organizational integrity (institutional capacity to audit and intervene), and ecosystem integrity (preserving the information environment on which oversight depends). This framework is the organizing device that turns scattered findings—overreliance, legitimacy laundering, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollutio
Load-bearing premise
The load-bearing premise—stated in the abstract and limitations—is that in deployed systems the most consequential failures are quiet (plausible, distributed, temporally extended, correction-degrading) rather than visible and local, a claim about the distribution of real-world harm that the paper itself does not measure.
What would settle it
A longitudinal study of real-world AI deployments that systematically catalogs incidents and measures the harm attributable to visible, local failures (e.g., direct misinformation, single-output errors, clear misuse) versus the five-layer integrity failures would settle the priority claim: if visible local failures account for the majority of measured harm, the framework's central claim loses force. Such a study would need to track deployments over time, including review quality, memory poisoning attempts, and incident detection latency.
If this is right
- Safety evaluation must move from static benchmarks to longitudinal socio-technical auditing, instrumenting how deployments detect, bound, escalate, and recover from controlled faults.
- Persistent memory should be treated as a privileged safety boundary with lifecycle controls on writes, access, retention, and deletion, not as an extension of chat history.
- Human oversight must be engineered for real reviewing power—time, evidence, authority, and protection—rather than nominal presence in the loop.
- Retrieval and training pipelines should track provenance, source diversity, and synthetic contamination, treating curated human-reviewed corpora as strategic safety assets.
- Tool-using systems need architectural separation between generation and authorization, with graduated agency and external policy enforcement for high-impact actions.
Where Pith is reading between the lines
- If the framework is right, regulatory and liability regimes should target deployment conditions (review budgets, audit access, rollback thresholds) rather than model capability scores, since benchmark scores can mask systemic integrity failures.
- The 'calibration debt' concept suggests a measurable longitudinal metric—the gap between observed reliance and warranted reliance—that could be tracked in real workflows and used as an early-warning signal for overtrust.
- The correction-degrading property implies that near-miss reporting and incident-learning infrastructure are not optional add-ons but core safety instruments; this could be tested by comparing organizations with and without such infrastructure on error-detection latency.
- The framework likely extends beyond AI to any automated system embedded in organizational workflows, where the same five integrity layers may predict resilience to automation-induced failure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This perspective article argues that AI safety practice is too concentrated on visible, output-level failures and thereby risks missing 'quieter' failures that emerge in deployed socio-technical systems. The authors identify four features of such failures (plausibility, distributedness, temporal extension, correction-degradation) and organize the risk space into five integrity layers: epistemic, control, temporal, organizational, and ecosystem. For each layer, they catalog specific failure modes such as overreliance/calibration debt, legitimacy and uncertainty laundering, prompt injection and action amplification, trajectory-level safety drift and memory poisoning, evaluation deception and fictional human oversight, and synthetic evidence pollution/retrieval collapse/model collapse. The paper closes with a table of controls and instrumentation indicators, a seven-point design agenda, twelve research priorities, and an explicit limitations section. The authors state that the contribution is a perspective and synthesis, not an empirical prevalence study.
Significance. The paper makes a useful synthetic contribution by connecting human-factors, security, systems-safety, and governance literatures around a single framework, and its concrete recommendations are mostly conservative and actionable. The Table 1 mapping from 'hidden challenge' to 'control' to 'indicator' is a practical starting point for audit design. The explicit limitations are a strength. However, the paper's central priority claim — that quiet, distributed, temporally extended, correction-degrading failures are 'the most consequential' — is asserted rather than established, and the Limitations concede the lack of prevalence evidence. As a perspective, the paper can be valuable if reframed as a hypothesis-generating framework; as written, its abstract and conclusion overstate the evidential basis.
major comments (3)
- [Abstract; §1; §3.5; Conclusion] The load-bearing claim that 'many of the most consequential failures are quieter' and that model-focused safety work 'will miss some of the most consequential risks in practice' is an empirical distributional assertion. The cited evidence (e.g., Denison et al. 2024; Greenblatt et al. 2024; Lynch et al. 2025) demonstrates existence and plausible mechanisms in controlled/simulated settings, not deployment prevalence or relative severity. The Limitations section explicitly disclaims frequency and severity claims. Because this priority claim motivates the entire shift to socio-technical auditing, it should be rephrased as a testable hypothesis or supported with deployment/incident data. Otherwise the conclusion overstates what the evidence base supports.
- [§3.5 and §5] The paper imports high-reliability organization principles (Leveson 2016; Reason 1997; Weick & Sutcliffe 2007) as though AI deployments inherit the same error-correcting structures as aviation, nuclear, or healthcare systems. The analogy is not argued. In fact, many AI deployments are introduced without safety cases, near-miss reporting, or independent audit access. This is a load-bearing inference for the recommendations on preserving error-correcting capacity. Please state the conditions under which this transfer holds, or downgrade the recommendations to conditional proposals.
- [Table 1; §4; §6] The instrumentation indicators are acknowledged as 'illustrative rather than universally validated metrics,' and many are not operationalized (e.g., 'weighted positive gap between observed reliance and warranted reliance'; 'fraction of materially wrong outputs corrected before downstream action'). Since the paper's agenda is to 'instrument' hidden failures, the lack of validated metrics is a central gap. Please present Table 1 explicitly as a hypothesized measurement agenda, identifying which indicators have existing empirical support and which require new measurement development.
minor comments (6)
- [Abstract] The abstract text as provided contains missing spaces (e.g., 'CurrentAIsafetydiscourse...'). If this reflects the submitted PDF, please correct.
- [§3.3] In the third bullet of the persistent-memory paragraph, the syntax 'and malicious content...' after '; (2)' and '; (3)' should be '; (3) malicious content...'.
- [Figure 2] Figure 2 is described as illustrating calibration debt but, unlike Figure 3, does not state that the curves are schematic/illustrative. Add an explicit caveat that the figure is conceptual and not empirical data.
- [§3.1] The terms 'legitimacy laundering' and 'uncertainty laundering' are central, but only implicitly defined. Consider adding one-sentence definitions at first use to make them precise.
- [Table 1] The third column ('Instrumentation indicator') contains long, dense entries; consider splitting the table or moving to landscape to improve readability.
- [§6] The 'Note of concern' subsection is an opinion-driven aside. Consider integrating it into the main text (e.g., §3.5 or §5) rather than placing it after the numbered research priorities.
Circularity Check
No significant circularity: the paper is a self-described perspective and synthesis with no formal derivation chain, no fitted-input predictions, and no load-bearing self-citations.
full rationale
The paper makes no formal predictions, derives no equations, and fits no parameters. It explicitly frames its contribution as 'an operational synthesis and organizing framework rather than a new empirical benchmark, formal theory, or prevalence estimate' (Introduction), and the conclusion calls the five-layer framework 'an organizing device for instrumentation and governance rather than a closed taxonomy.' The central claim about 'quieter' consequential failures is an empirical thesis about prevalence that the paper itself disclaims in the Limitations: 'we do not claim that the identified failure modes are equally frequent, severe, or mature across domains.' That concession weakens the evidential force of the argument, but it is a gap between framing and evidence, not a case where an output is equivalent to an input by definition. The paper's definitions are constructive: 'calibration debt,' 'legitimacy laundering,' 'fictional human oversight,' and the integrity layers are stipulated concepts, not quantities defined in terms of the conclusions they are used to support. The self-citations (Kirchhof et al. 2025; Rong et al. 2024, 2025) are used only as supporting references for specific design recommendations—structured uncertainty displays and benchmark-design critiques—and are not the load-bearing justification for the framework. The framework itself is anchored in external safety-science literature (Leveson 2016; Reason 1997; Weick & Sutcliffe 2007), not in the authors' prior work. No circular step satisfies the threshold of quoting an equation or construction showing that a 'prediction' reduces to its own input. The appropriate finding is therefore no significant circularity.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption In deployed AI systems, consequential failures are predominantly quiet/systemic rather than visible and local.
- domain assumption High-reliability domains stay safe by preserving pathways for reporting, contestation, and recovery; AI deployments inherit this guarantee structure.
- domain assumption Results from controlled, simulated, or red-teaming studies reveal plausible failure mechanisms relevant to deployment risk.
- domain assumption Common mitigations improve one integrity layer while degrading or ignoring others.
invented entities (6)
-
Calibration debt
independent evidence
-
Legitimacy laundering
no independent evidence
-
Fictional human oversight
independent evidence
-
Evaluation deception
independent evidence
-
Safety drift
independent evidence
-
Retrieval collapse
no independent evidence
read the original abstract
Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions under which errors remain visible, contestable, containable, and recoverable. We propose a five-layer framework for diagnosing these hidden risks: (1) epistemic integrity, concerning whether evidence and uncertainty are represented honestly enough to support calibrated reliance; (2) control integrity, concerning whether authority, permissions, and action boundaries remain robust under attack and optimization; (3) temporal integrity, concerning whether safety holds across sessions, memory updates, and deployment drift; (4) organizational integrity, concerning whether institutions retain the capacity to audit, assign responsibility, and intervene effectively; and (5) ecosystem integrity, concerning whether AI systems preserve rather than erode the information environment on which future oversight depends. Across these layers, we identify under-recognized risk patterns, including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. We conclude with design and governance recommendations and a research agenda for shifting AI safety from model-centric evaluation toward socio-technical reliability.
Figures
Reference graph
Works this paper leans on
-
[1]
Paraschou, E., Pownall, C., Prajapati, J., et al. (2024). A collaborative, human-centred taxonomy of AI, algorithmic, and automation harms.arXiv preprint arXiv:2407.01294. https://doi.org/10.48550/arXiv.2407.01294
-
[2]
I., Babaei, H., LeJeune, D., Siahkoohi, A., & Baraniuk, R
Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., & Baraniuk, R. (2023). Self-consuming generative models go mad.The Twelfth Inter- national Conference on Learning Representations. https://openreview.net/forum?id= ShjMHfmPs0
2023
-
[3]
Alfrink, K., Keller, I., Kortuem, G., & Doorn, N. (2023). Contestable AI by design: Towards a framework.Minds and Machines,33(4), 613–639. https://doi.org/10.1007/s11023-022- 09611-z
-
[4]
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in ai safety.arXiv preprint arXiv:1606.06565. https://doi.org/10.48550/arXiv. 1606.06565
-
[6]
Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., ... Liang, P. (2021). On the Opportunities and Risks of Foundation Models.arXiv preprint arXiv:2108.07258. https: //doi.org/10.48550/arXiv.2108.07258
-
[7]
Briesch, M., Sobania, D., & Rothlauf, F. (2023). Large language models suffer from their own output: An analysis of the self-consuming training loop.arXiv preprint arXiv:2311.16822. https://openreview.net/forum?id=SaOxhcDCM3
Pith/arXiv arXiv 2023
-
[8]
Bucinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making.Proceedings of the ACM on Human-computer Interaction,5(CSCW1), 1–21. https://doi.org/10.1145/3449287
doi:10.1145/3449287 2021
-
[9]
Yang, T., Huo, J., Gao, Y., Meng, F., Yang, X., Deng, C., & Feng, J. (2026). SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks.The Fourteenth International Conference on Learning Representations. https://doi.org/10.48550/arXiv.2502.11090
-
[10]
Gerovitch, M., Bau, D., Tegmark, M., ... Hadfield-Menell, D. (2024). Black-Box Access is Insufficient for Rigorous AI Audits.Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, 2254–2272. https://doi.org/10.1145/3630106.3659037
arXiv 2024
-
[11]
Chen, Z., Xiang, Z., Xiao, C., Song, D., & Li, B. (2024). AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases.The Thirty-eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=Y841BRW9rY 20
2024
-
[12]
Chua, J., Li, Y., Yang, S., Wang, C., & Yao, L. (2024). AI safety in generative AI large language models: A survey.arXiv preprint arXiv:2407.18369. https://doi.org/10.48550/arXiv.2407. 18369
-
[13]
Cobbe, J., Lee, M. S. A., & Singh, J. (2021). Reviewable automated decision-making: A framework for accountable algorithmic systems.Proceedings of the 2021 ACM conference on Fairness, Accountability, and Transparency, 598–609. https://doi.org/10.1145/3442188.3445921
arXiv 2021
-
[14]
Cobbe, J., Veale, M., & Singh, J. (2023). Understanding accountability in algorithmic supply chains. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, 1186–1197. https://doi.org/10.1145/3593013.3594073
arXiv 2023
-
[15]
Terzis, A., & Tramèr, F. (2025). Defeating Prompt Injections by Design.arXiv preprint arXiv:2503.18813. https://doi.org/10.48550/arXiv.2503.18813
-
[16]
Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., & Tramèr, F. (2024). AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://doi.org/10.52202/079017-2636
-
[17]
Deng, J., Cheng, J., Sun, H., Zhang, Z., & Huang, M. (2023). Towards safer generative language mod- els: A survey on safety risks, evaluations, and improvements.arXiv preprint arXiv:2302.09270. https://doi.org/10.48550/arXiv.2302.09270
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2302.09270 2023
-
[18]
Denison, C., MacDiarmid, M., Barez, F., Duvenaud, D., Kravec, S., Marks, S., Schiefer, N., Soklaski, R., Tamkin, A., Kaplan, J., Shlegeris, B., Bowman, S. R., Perez, E., & Hubinger, E. (2024). Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models. arXiv preprint arXiv:2406.10162. https://doi.org/10.48550/arXiv.2406.10162 Europe...
-
[19]
J., & Gidel, G
Ferbach, D., Bertrand, Q., Bose, A. J., & Gidel, G. (2024). Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences.The Thirty-eighth Annual Conference on Neural Information Processing Systems. https://proceedings.neurips.cc/paper_files/ paper/2024/hash/b9e88ae0308cf82d0b0f634ddbdf809a-Abstract-Conference.html
2024
-
[20]
Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling large language models to generate text with citations.Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6465–6488. https://doi.org/10.18653/v1/2023.emnlp-main.398
-
[21]
B., Gromov, A., Roberts, D., Yang, D., Donoho, D
Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., Korbak, T., Sleight, H., Agrawal, R., Hughes, J., Pai, D. B., Gromov, A., Roberts, D., Yang, D., Donoho, D. L., & Koyejo, S. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data.First Conference on Language Modeling. https://openreview.net/foru...
2024
-
[22]
Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators.Journal of the American Medical Informatics Association,19(1), 121–127. https://doi.org/10.1136/amiajnl-2011-000089 21
-
[23]
Belonax, T., Chen, J., Duvenaud, D., et al. (2024). Alignment faking in large language models.arXiv preprint arXiv:2412.14093. https://doi.org/10.48550/arXiv.2412.14093
-
[24]
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection.Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, 79–90. https://doi.org/10.1145/3605764.3623985
arXiv 2023
-
[25]
Gyevnar, B., & Kasirzadeh, A. (2025). AI safety for everyone.Nature Machine Intelligence,7(4), 531–542. https://doi.org/10.1038/s42256-025-01020-y
-
[26]
Habli, I., Hawkins, R., Paterson, C., Ryan, P., Jia, Y., Sujan, M., & McDermid, J. (2025). The big argument for AI safety cases.arXiv preprint arXiv:2503.11705. https://doi.org/10.48550/ arXiv.2503.11705
-
[27]
He, F., Zhu, T., Ye, D., Liu, B., Zhou, W., & Yu, P. S. (2025). The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies.ACM Computing Surveys,58(6). https://doi.org/10.1145/3773080
doi:10.1145/3773080 2025
-
[28]
Kattan, A. E., Stein, M., et al. (2025). Measuring and mitigating overreliance is necessary for building human-compatible ai.arXiv preprint arXiv:2509.08010. https://doi.org/10. 48550/arXiv.2509.08010
-
[29]
F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., & Neubig, G
Jiang, Z., Xu, F. F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., & Neubig, G. (2023). Active Retrieval Augmented Generation.The 2023 Conference on Empirical Methods in Natural Language Processing. https://doi.org/10.18653/v1/2023.emnlp-main.495
-
[30]
Kim, S. S., Vaughan, J. W., Liao, Q. V., Lombrozo, T., & Russakovsky, O. (2025). Fostering appropriate reliance on large language models: The role of explanations, sources, and inconsistencies.Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–19. https://doi.org/10.1145/3706598.3714020
arXiv 2025
-
[31]
Kirchhof, M., Kasneci, G., & Kasneci, E. (2025). Position: Uncertainty quantification needs reassess- ment for large language model agents.Proceedings of the 42nd International Conference on Machine Learning,267, 81665–81677. https://proceedings.mlr.press/v267/kirchhof25b.html
2025
-
[32]
Klingbeil, A., Grützner, C., & Schreck, P. (2024). Trust and reliance on AI—An experimental study on the extent and costs of overreliance on AI.Computers in Human Behavior,160, 108352. https://doi.org/10.1016/j.chb.2024.108352
arXiv 2024
-
[33]
Legg, S. (2020). Specification gaming: the flip side of AI ingenuity.DeepMind Blog,3, 40–53. https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/
2020
-
[34]
Krishna, S., Krishna, K., Mohananey, A., Schwarcz, S., Stambler, A., Upadhyay, S., & Faruqui, M. (2025). Fact, fetch, and reason: A unified evaluation of retrieval-augmented generation. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Pap...
-
[35]
Langosco, L. L. D., Koch, J., Sharkey, L. D., Pfau, J., & Krueger, D. (2022, July). Goal mis- generalization in deep reinforcement learning. In K. Chaudhuri, S. Jegelka, L. Song, C
2022
-
[36]
Leveson, N. G. (2016).Engineering a safer world: Systems thinking applied to safety. MIT press
2016
-
[37]
V., Song, T., Xu, Z., & Lee, Y.-c
Li, J., Yang, Y., Zhang, R., Liao, Q. V., Song, T., Xu, Z., & Lee, Y.-c. (2024). Understanding the Effects of Miscalibrated AI Confidence on User Trust, Reliance, and Decision Efficacy.arXiv preprint arXiv:2402.07632. https://doi.org/10.48550/arXiv.2402.07632
-
[38]
Li, M., Bickersteth, W., Tang, N., Cranor, L., Hong, J., Shen, H., & Heidari, H. (2025). A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents.Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society,8(2), 1561–1573. https://doi.org/10.1609/aies.v8i2.36655
-
[39]
Yue, S. (2024). LLM defenses are not robust to multi-turn human jailbreaks yet.arXiv preprint arXiv:2408.15221. https://doi.org/10.48550/arXiv.2408.15221
-
[40]
Li, X., Yu, S., Pan, M., Sun, Y., Li, B., Song, D., Lin, X., & Shi, W. (2026). Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents.arXiv preprint arXiv:2602.13379. https://doi.org/10.48550/arXiv.2602.13379
-
[41]
Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., ... Tang, J. (2024). AgentBench: Evaluating LLMs as agents.The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=zAdUB0aCTQ
2024
-
[42]
Liu, X., Yu, Z., Zhang, Y., Zhang, N., & Xiao, C. (2024). Automatic and universal prompt injection attacks against large language models.arXiv preprint arXiv:2403.04957. https: //doi.org/10.48550/arXiv.2403.04957
-
[43]
Troy, K. (2025). Agentic Misalignment: How LLMs Could Be Insider Threats.arXiv preprint arXiv:2510.05179. https://doi.org/10.48550/arXiv.2510.05179
-
[44]
Ma, C., Zhang, J., Zhu, Z., Yang, C., Yang, Y., Jin, Y., Lan, Z., Kong, L., & He, J. (2024). AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents.The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://openreview.net/forum?id=4S8agvKjle National Cyber Security Centre. (2025). Prompt inject...
2024
-
[45]
Ni, B., Liu, Z., Wang, L., Lei, Y., Zhao, Y., Cheng, X., Zeng, Q., Dong, L., Xia, Y., Kenthapadi, K., et al. (2025). Towards trustworthy retrieval augmented generation for large language models: A survey.arXiv preprint arXiv:2502.06872. https://doi.org/10.48550/arXiv.2502.06872
-
[46]
Parasuraman, R., & Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration.Human Factors,52(3), 381–410. https://doi.org/10.1177/ 0018720810376055 23
2010
-
[47]
Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse.Human factors,39(2), 230–253. https://doi.org/10.1518/001872097778543886
-
[48]
Theron, D., & Barnes, P. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing.Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 33–44. https://doi.org/10.1145/3351095.3372873
arXiv 2020
-
[49]
(1997).Managing the Risks of Organizational Accidents
Reason, J. (1997).Managing the Risks of Organizational Accidents. Ashgate. https://doi.org/10. 4324/9781315543543
1997
-
[50]
Rong, Y., Leemann, T., Nguyen, T.-t., Fiedler, L., Qian, P., Unhelkar, V., Seidel, T., Kasneci, G., & Kasneci, E. (2024). Towards human-centered explainable AI: A survey of user studies for model explanations.IEEE Transactions on Pattern Analysis and Machine Intelligence. https://doi.org/10.1109/TPAMI.2023.3331846
arXiv 2024
-
[51]
Rong, Y., Seßler, K., Gözlüklü, E., & Kasneci, E. (2025). Benchmarking in-context learning strate- gies of large language models for math reasoning tasks.IEEE Transactions on Learning Technologies,18, 1074–1082. https://doi.org/10.1109/TLT.2025.3630117
arXiv 2025
-
[52]
Rossi, S., Michel, A. M., Mukkamala, R. R., & Thatcher, J. B. (2024). An early categorization of prompt injection attacks on large language models.arXiv preprint arXiv:2402.00898. https://doi.org/10.48550/arXiv.2402.00898 Santoni de Sio, F., & Van den Hoven, J. (2018). Meaningful human control over autonomous systems: A philosophical account.Frontiers in ...
-
[53]
Scheurer, J., Balesni, M., & Hobbhahn, M. (2024). Large Language Models can Strategically Deceive their Users when Put Under Pressure.ICLR 2024 Workshop on Large Language Model (LLM) Agents. https://doi.org/10.48550/arXiv.2311.07590
-
[54]
Gallegos, J., Smart, A., Garcia, E., & Virk, G. (2023). Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction.Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, 723–741. https://doi.org/10.1145/3600211.3604673
arXiv 2023
-
[55]
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data.Nature,631(8022), 755–759. https: //doi.org/10.1038/s41586-024-07566-y
-
[56]
Song, M., Sim, S. H., Bhardwaj, R., Chieu, H. L., Majumder, N., & Poria, S. (2024). Measuring and enhancing trustworthiness of LLMs in RAG through grounded attributions and learning to refuse.arXiv preprint arXiv:2409.11242. https://doi.org/10.48550/arXiv.2409.11242
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2409.11242 2024
-
[57]
Spatola, N. (2024). The efficiency-accountability tradeoff in AI integration: Effects on human performance and over-reliance.Computers in Human Behavior: Artificial Humans,2(2), 100099. https://doi.org/10.1016/j.chbah.2024.100099
arXiv 2024
-
[58]
Sun, X., Zhang, D., Yang, D., Zou, Q., & Li, H. (2024). Multi-turn context jailbreak attack on large language models from first principles.arXiv preprint arXiv:2408.04686. https: //doi.org/10.48550/arXiv.2408.04686
-
[59]
D., Sinha, I., Maheshwari, P., Todmal, S., Mallik, S., & Mishra, S
Sunil, B. D., Sinha, I., Maheshwari, P., Todmal, S., Mallik, S., & Mishra, S. (2026). Memory Poisoning Attack and Defense on Memory Based LLM-Agents.arXiv preprint arXiv:2601.05504. https: //doi.org/10.48550/arXiv.2601.05504 24
-
[60]
Suo, X. (2024). Signed-prompt: A new approach to prevent prompt injection attacks against llm- integrated applications.AIP Conference Proceedings,3194(1), 040013. https://doi.org/10. 1063/5.0222987
2024
-
[61]
Vasconcelos, H., Jörke, M., Grunde-McLaughlin, M., Gerstenberg, T., Bernstein, M. S., & Krishna, R. (2023). Explanations can reduce overreliance on ai systems during decision-making. Proceedings of the ACM on Human-Computer Interaction,7(CSCW1), 1–38. https://doi. org/10.1145/3579605
doi:10.1145/3579605 2023
-
[62]
Wallat, J., Heuss, M., Rijke, M. d., & Anand, A. (2025). Correctness is not Faithfulness in Retrieval Augmented Generation Attributions.Proceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR), 22–32. https://doi.org/10.1145/3731120.3744592
arXiv 2025
-
[63]
Wang, B., He, W., Zeng, S., Xiang, Z., Xing, Y., Tang, J., & He, P. (2025). Unveiling privacy risks in LLM agent memory.Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 25241–25260. https://doi.org/10. 18653/v1/2025.acl-long.1227
2025
-
[64]
E., & Sutcliffe, K
Weick, K. E., & Sutcliffe, K. M. (2007).Managing the Unexpected: Resilient Performance in an Age of Uncertainty(2nd ed.). Jossey-Bass
2007
-
[65]
Balle, B., Kasirzadeh, A., et al. (2021). Ethical and social risks of harm from language models.arXiv preprint arXiv:2112.04359,10. https://doi.org/10.48550/arXiv.2112.04359
-
[66]
Xiong, Z., Lin, Y., Xie, W., He, P., Liu, Z., Tang, J., Lakkaraju, H., & Xiang, Z. (2025). How memory management impacts LLM agents: An empirical study of experience-following behavior. arXiv preprint arXiv:2505.16067. https://doi.org/10.48550/arXiv.2505.16067
-
[67]
Yu, E., Li, J., Liao, M., Wang, S., Zuchen, G., Mi, F., & Hong, L. (2024). CoSafe: Evaluating large language model safety in multi-turn dialogue coreference.Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 17494–17508. https: //doi.org/10.18653/v1/2024.emnlp-main.968
-
[68]
Yu, H., Kim, D., & Kim, Y.-B. (2026). Retrieval Collapses When AI Pollutes the Web.arXiv preprint arXiv:2602.16136. https://doi.org/10.48550/arXiv.2602.16136
-
[69]
Zhang, Z., Dai, Q., Bo, X., Ma, C., Li, R., Chen, X., Zhu, J., Dong, Z., & Wen, J.-R. (2025). A Survey on the Memory Mechanism of Large Language Model-based Agents.ACM Transactions on Information Systems,43(6). https://doi.org/10.1145/3748302
doi:10.1145/3748302 2025
-
[70]
Zhou, Y., Liu, Y., Li, X., Jin, J., Qian, H., Liu, Z., Li, C., Dou, Z., Ho, T.-Y., & Yu, P. S. (2024). Trustworthiness in retrieval-augmented generation systems: A survey.arXiv preprint arXiv:2409.10102. https://doi.org/10.48550/arXiv.2409.10102
-
[71]
Ududec, C., Kellermann, A., Sekhon, J. S., ... Kang, D. (2025). Establishing Best Practices in Building Rigorous Agentic Benchmarks.The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://doi.org/10.48550/ arXiv.2507.02825 25
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.