Pith. sign in

REVIEW 2 major objections 5 minor 90 references

When evidence of who is asking can be copied, no LLM safeguard can keep dual-use answers useful for legitimate users and reliably denied to attackers while staying open.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 22:23 UTC pith:YOQMPW7X

load-bearing objection Clean impossibility result: under copyable pre-release evidence, interactive LLM safeguards cannot beat the static dual-use assistance floor Γ(q). the 2 major comments →

arxiv 2607.27951 v1 pith:YOQMPW7X submitted 2026-07-30 cs.CR cs.AI

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs

classification cs.CR cs.AI
keywords LLM safetydual-usecopyable evidencesafety trilemmatrusted credentialsattacker assistance flooraccess controlinteractive safeguards
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

LLM safeguards decide what to release before anyone sees how the answer will be used. For dual-use tasks—the same technical answer helps an authorized professional or an attacker—that timing creates a hard limit: if an attacker can present the same request and interaction history as a legitimate user, every useful answer still helps the attacker at least as much as a fixed floor allows. The paper derives that floor exactly and turns it into a trilemma: useful capability, reliable safety against worst-case misuse, and open access based only on copyable evidence cannot all hold at once. A trusted credential that adds hard-to-copy information tied to actual downstream use can move below the floor; zero assistance further requires that enough legitimate value sit on credential values attackers cannot obtain. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs is offered as support that these conditions already matter in practice.

Core claim

Under copyable access evidence, every interactive release rule that preserves legitimate-use utility at least q still leaves worst-case attacker assistance at least Γ(q), the minimum assistance from any shared output distribution that meets that legitimate target. For dual-use tasks this floor is strictly positive whenever q is. Therefore useful capability, reliable safety below that floor, and open access that uses only copyable evidence cannot coexist.

What carries the argument

Γ(q): the minimum expected attacker utility over release distributions that still deliver legitimate utility at least q. Theorem 1 shows that when the legitimate transcript law is inside the attacker’s closed strategy class, interactive safeguards reduce exactly to this static floor; trusted signals that attackers cannot copy are what can lower it.

Load-bearing premise

Attackers can reproduce the same request and interaction evidence that legitimate users produce, so the safeguard cannot tell the two cases apart before release.

What would settle it

Find a dual-use task and an open, credential-free interactive safeguard that, against attackers who match legitimate interaction strategies, keeps legitimate utility at a chosen q while holding worst-case attacker assistance strictly below the menu’s Γ(q).

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Open, copyable-context filters and multi-turn checks cannot drive worst-case dual-use assistance to zero while keeping positive legitimate utility.
  • Changing the output menu can lower Γ(q) but cannot erase it under the dual-use condition without also cutting legitimate value.
  • Trusted credentials only help when their distribution predicts actual downstream use and attackers cannot freely reproduce them.
  • Zero-assistance safety additionally needs enough legitimate utility on credential values the malicious process cannot attain.
  • Task splitting and retries do not reduce the floor: under copyable evidence the additive assistance bound simply sums across sessions.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Governance debates that treat ‘better intent classifiers’ as sufficient for dual-use capability control are arguing past the trilemma unless they also change the evidence attackers cannot copy.
  • Deployed trusted-access programs for cyber and research roles are natural testbeds: measuring separation d and residual copying error would show how far current credentials actually move the floor.
  • The same logic likely extends to any tool-using agent whose final action is chosen after a copyable dialogue, not only chat LLMs.
  • If open-weight release lets users change the release menu itself, the access-evidence model no longer applies and the trilemma must be restated for a different control surface.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper separates released model capability from pre-release evidence about downstream use, and shows that when that evidence is copyable, any interactive safeguard is forced to the static capability floor Γ(q): the minimum attacker assistance compatible with legitimate-use utility at least q. Under a dual-use condition (every positively useful release also has malicious value), Γ(q) > 0 for every feasible q > 0, yielding a trilemma among Useful Capability, Reliable Safety (worst-case assistance ≤ β < Γ(q)), and Open Access (credential-free, copyable evidence only). The authors then introduce a trusted noncopyable signal S that predicts actual downstream use, prove an analogous reduction to an S-conditioned floor Γ_S(q), and give the support condition under which zero assistance becomes attainable. Continuity bounds cover imperfect copying; an additive multi-session composition result shows task decomposition does not lower the floor. Section 4 surveys dual-use evaluations, adaptive attacks, and deployed trusted-access programs as qualitative support for the premises.

Significance. If the stated premises hold, the result is a clean, load-bearing impossibility for a large class of current LLM safeguards: intent checks, filters, multi-turn clarification, and interactive defenses that rely only on request and dialogue evidence cannot beat Γ(q) in the worst case. The finite-space arguments (Theorem 1, Corollary 1, Theorems 2–3, Proposition 2) are standard and tight—shared release laws define the floor; copyability puts the legitimate law in the attacker class; a transcript-independent optimizer matches the upper bound—and the TV continuity extensions are correctly handled. Framing the missing object as noncopyable evidence tied to actual use, rather than better prompt classifiers, is a useful conceptual contribution for the safety and access-control literature. The paper does not ship machine-checked proofs or measured Γ/δ_κ values, but the derivations are self-contained and the trilemma is not circular.

major comments (2)
  1. [§3.2, Eqs. (5)–(6)] §3.2, Eqs. (5)–(6): positivity of the floor rests on ρ = min_{a: u_B(a)>0} u_M(a)/u_B(a) > 0. Finiteness gives ρ > 0, but nothing in the manuscript bounds how large ρ (hence Γ(q)) is for realistic release menus. If intermediate actions (mitigation-only advice, root-cause without exploit chain) make ρ negligible at the chosen utility resolution, the ‘unavoidable assistance’ floor can be operationally irrelevant while still formally positive. A short discussion—or a worked numerical menu for the running vulnerability-analysis example—showing plausible (u_B, u_M) pairs and the resulting Γ(q) curve is needed for the policy force of Corollary 1 to be assessable.
  2. [§4, Eq. (12)] §4 and Eq. (12): the empirical section is framed only as qualitative support that dual-use tasks, copyable evidence, and trusted credentials exist in practice. That is appropriate for a theory paper, but the abstract and introduction claim that this evidence ‘supports the practical relevance of these conditions.’ No estimate of δ_κ, d, or even a toy Γ(q) appears. Without at least one concrete, fully specified example that computes Γ(q) and illustrates how a trusted signal would move the floor, readers cannot judge whether the trilemma binds at deployment-relevant scales or only in the formal limit. Adding such an example would make the central claim much more falsifiable.
minor comments (5)
  1. [Figure 1] Figure 1 is referenced as illustrating capability allocation with copyable evidence and trusted credentials, but the manuscript text does not fully specify the axes or the quantitative meaning of the shaded regions. A brief caption expansion would help.
  2. [§3.1, Eq. (3); §3.5] Notation: U_z(P, g) in Eq. (3) and the later extension to (S, H) reuse the same symbol for different domains; a subscript or separate notation would reduce ambiguity when both appear near Theorem 2.
  3. [Abstract; §3.3] The phrase ‘Reliable Safety’ is defined as a worst-case ceiling on attacker assistance evaluated against actual downstream use. That definition is clear in §3.3, but the abstract and introduction use the phrase before it is fixed; a one-sentence forward pointer would help non-specialist readers.
  4. [Throughout] Typographical: several compound words are missing spaces in the compiled text (e.g., ‘beforeseeinghowananswerwillbeused’, ‘dual-usetasks’). These appear to be PDF line-break artifacts and should be cleaned for the camera-ready version.
  5. [§5; Corollary 1] §5 correctly restricts ‘open access’ to credential-free inference-time access rather than weight release. Consider elevating that clarification earlier (e.g., near Corollary 1) so readers do not misread the trilemma as applying to open-weight release.

Circularity Check

0 steps flagged

No significant circularity: Γ(q) and Rκ(q) are independently defined, and Theorem 1 is a genuine equality proof under an explicit copyability premise.

full rationale

The paper’s central chain is definitional setup plus proof, not a fit-or-rename loop. Γ(q) is the static capability-allocation minimum over shared release distributions (Eq. 4). Rκ(q) is the interactive worst-case assistance under admissible attacker transcript laws (Sec. 3.3). Theorem 1 proves Rκ(q)=Γ(q) when Pκ_B ∈ Cκ by a copied-law lower bound and a transcript-independent optimizer upper bound—standard minimax reduction on finite spaces, not equality by construction of the same object under two names. The dual-use condition (Eq. 5) and ρq bound (Eq. 6) are ordinary consequences of the task class definition, not fitted parameters re-labeled as predictions. Trusted-signal results (Thm. 2–3, Cor. 3) likewise derive ΓS and the support condition from signal marginals and conditional copying. Empirical §4 cites external dual-use, attack, and access-program literature as qualitative premise relevance; no load-bearing uniqueness theorem or ansatz is imported from overlapping-author prior work. No step reduces a claimed prediction to its own fitted input.

Axiom & Free-Parameter Ledger

0 free parameters · 7 axioms · 2 invented entities

The paper is mostly definition-and-proof theory. Load-bearing content is modeling axioms (dual-use utilities, copyable attacker class, pre-release decision timing) plus standard finite probabilistic decision theory. No numeric parameters are fitted to data. Invented structure is definitional (Γ, credential signal S) rather than new physical entities.

axioms (7)
  • domain assumption Dual-use condition: for every release a, u_B(a)>0 implies u_M(a)>0 (Eq. 5), yielding Γ(q)≥ρq>0 for q>0.
    Defines the task class; without it zero attacker assistance with positive legitimate utility can be trivial via purely legitimate-only releases.
  • domain assumption Safeguards choose the terminal release from public context W and access transcript H before actual downstream use is observed.
    Stated in Introduction and §3.1; timing is essential to the evidence problem.
  • domain assumption Copyability: admissible malicious strategies can match (or approximate in TV) the legitimate reference transcript law under committed κ (Prop. 1, Eq. 9).
    Premise of Theorem 1 and the trilemma; empirical §4.2 argues it is realistic for software-capable attackers.
  • standard math Finite release menu A and transcript space H; utilities in [0,1]; assistance is expected u_M of the released answer (not realized harm).
    §3 uses finite simplices so Γ(q) is attained, convex, piecewise linear; TV bounds apply.
  • domain assumption Attacker observes the committed defense (κ,g) then chooses strategy (Stackelberg / attacker-moves-second).
    Built into R^κ(q) via sup over C^κ; aligned with cited adaptive-attack literature.
  • ad hoc to paper Open access means credential-free access to a committed inference-time mechanism, not weight release (§5).
    Scopes the trilemma; weight exfiltration is explicitly out of model.
  • domain assumption Trusted credential supplies signal S whose marginals can differ by downstream use Z and that attackers cannot freely reproduce; conditional copying of H given S may still hold (Thm 2).
    Needed to move below Γ(q); zero assistance further needs legitimate mass on S_0 where P^S_M=0 (Cor. 3).
invented entities (2)
  • Capability floor Γ(q) independent evidence
    purpose: Exact minimum attacker assistance compatible with legitimate utility q under a shared release distribution.
    Defined mathematically from the release menu and utilities (Eq. 4); not an empirical latent. Independent meaning is the convex lower frontier of the utility pair set.
  • Trusted credential / noncopyable signal S predicting downstream use independent evidence
    purpose: Augment copyable W,H so release rules can assign different menus to legitimate vs malicious processes and potentially reach below Γ(q) or zero assistance.
    Conceptual mechanism class; paper points to hardware attestation and deployed cyber trusted-access programs as real-world analogues, but does not introduce a new cryptographic primitive.

pith-pipeline@v1.2.0-daily-grok45 · 18505 in / 3807 out tokens · 75845 ms · 2026-07-31T22:23:35.928508+00:00 · methodology

0 comments
read the original abstract

Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.

Figures

Figures reproduced from arXiv: 2607.27951 by Lingyao Zhu, Nenghai Yu, Pingyu Wu, Weiming Zhang.

Figure 1
Figure 1. Figure 1: Illustrative capability allocation with copyable evidence and trusted credentials. A trusted credential can improve [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

90 extracted references · 30 linked inside Pith

  1. [1]

    On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for

    Sarah Ball and Greg G. On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for. arXiv preprint arXiv:2507.07341 , year =

  2. [2]

    Access Controls Will Solve the Dual-Use Dilemma , journal =

    Ev. Access Controls Will Solve the Dual-Use Dilemma , journal =. 2025 , note =

  3. [3]

    Schulhoff and others , title =

    Milad Nasr and Nicholas Carlini and Chawin Sitawarin and Sander V. Schulhoff and others , title =. arXiv preprint arXiv:2510.09023 , year =

  4. [4]

    Zico Kolter and Matt Fredrikson , title =

    Andy Zou and Zifan Wang and Nicholas Carlini and Milad Nasr and J. Zico Kolter and Matt Fredrikson , title =. arXiv preprint arXiv:2307.15043 , year =

  5. [5]

    arXiv preprint arXiv:2212.08073 , year =

    Yuntao Bai and Saurav Kadavath and Sandipan Kundu and Amanda Askell and others , title =. arXiv preprint arXiv:2212.08073 , year =

  6. [6]

    arXiv preprint arXiv:2501.18837 , year =

    Mrinank Sharma and Meg Tong and Jesse Mu and Jerry Wei and Jorrit Kruthoff and others , title =. arXiv preprint arXiv:2501.18837 , year =

  7. [7]

    arXiv preprint arXiv:2601.04603 , year =

    Hoagy Cunningham and Jerry Wei and Zihan Wang and Andrew Persic and Alwin Peng and others , title =. arXiv preprint arXiv:2601.04603 , year =

  8. [8]

    Findings of the Association for Computational Linguistics: ACL 2026 , publisher =

    Yuan Xin and Dingfan Chen and Linyi Yang and Michael Backes and Xiao Zhang , title =. Findings of the Association for Computational Linguistics: ACL 2026 , publisher =. 2026 , doi =

  9. [9]

    arXiv preprint arXiv:2508.09224 , year =

    Yuan Yuan and Tina Sriskandarajah and Anna-Luisa Brakman and Alec Helyar and Alex Beutel and others , title =. arXiv preprint arXiv:2508.09224 , year =

  10. [10]

    arXiv preprint arXiv:2405.11030 , year =

    Xinyu Wang and Sai Koneru and Pranav Narayanan Venkit and Brett Frischmann and Sarah Rajtmajer , title =. arXiv preprint arXiv:2405.11030 , year =

  11. [11]

    arXiv preprint arXiv:2607.02047 , year =

    Rheeya Uppaal and Seungwoo Lyu and Selina Sung and Junjie Hu , title =. arXiv preprint arXiv:2607.02047 , year =

  12. [12]

    Paved with True Intents: Intent-Aware Training Improves

    Jeremias Ferrao and Niclas M. Paved with True Intents: Intent-Aware Training Improves. arXiv preprint arXiv:2606.27210 , year =

  13. [13]

    arXiv preprint arXiv:2604.27093 , year =

    Mingqian Zheng and Malia Morgan and Liwei Jiang and Carolyn Rose and Maarten Sap , title =. arXiv preprint arXiv:2604.27093 , year =

  14. [14]

    Evans , title =

    Nadav Kunievsky and James A. Evans , title =. arXiv preprint arXiv:2506.16584 , year =

  15. [15]

    arXiv preprint arXiv:2606.03135 , year =

    Mengyi Deng and Zhiwei Li and Xin Li and Tingyu Zhu and Ying Zhao and Zhijiang Guo and Wei Wang , title =. arXiv preprint arXiv:2606.03135 , year =

  16. [16]

    arXiv preprint arXiv:2407.02551 , year =

    David Glukhov and Ziwen Han and Ilia Shumailov and Vardan Papyan and Nicolas Papernot , title =. arXiv preprint arXiv:2407.02551 , year =

  17. [17]

    Varshney , title =

    Xinbo Wu and Abhishek Umrawal and Lav R. Varshney , title =. arXiv preprint arXiv:2505.20841 , year =

  18. [18]

    arXiv preprint arXiv:2606.23668 , year =

    David Mguni and Julian Ma and Jun Wang , title =. arXiv preprint arXiv:2606.23668 , year =

  19. [19]

    Bakker , title =

    Michelle Vaccaro and Jaeyoon Song and Abdullah Almaatouq and Michiel A. Bakker , title =. arXiv preprint arXiv:2603.26676 , year =

  20. [20]

    arXiv preprint arXiv:2605.29224 , year =

    Aditya Nawal and Manit Baser and Mohan Gurusamy , title =. arXiv preprint arXiv:2605.29224 , year =

  21. [21]

    arXiv preprint arXiv:2403.03218 , year =

    Nathaniel Li and Alexander Pan and Anjali Gopal and Summer Yue and Daniel Berrios and others , title =. arXiv preprint arXiv:2403.03218 , year =

  22. [22]

    David Blackwell and M. A. Girshick , title =

  23. [23]

    Huber and Volker Strassen , title =

    Peter J. Huber and Volker Strassen , title =. The Annals of Statistics , volume =. 1973 , doi =

  24. [24]

    Cover and Joy A

    Thomas M. Cover and Joy A. Thomas , title =

  25. [25]

    Tsybakov , title =

    Alexandre B. Tsybakov , title =

  26. [26]

    The Annals of Mathematical Statistics , volume =

    Herman Chernoff , title =. The Annals of Mathematical Statistics , volume =. 1959 , doi =

  27. [27]

    Crawford and Joel Sobel , title =

    Vincent P. Crawford and Joel Sobel , title =. Econometrica , volume =. 1982 , doi =

  28. [28]

    The Quarterly Journal of Economics , volume =

    Michael Spence , title =. The Quarterly Journal of Economics , volume =. 1973 , doi =

  29. [29]

    arXiv preprint arXiv:2606.29113 , year =

    Quanyan Zhu , title =. arXiv preprint arXiv:2606.29113 , year =

  30. [30]

    arXiv preprint arXiv:2412.00836 , year =

    Edward Kembery and Ben Bucknall and Morgan Simpson , title =. arXiv preprint arXiv:2412.00836 , year =

  31. [31]

    Personhood Credentials: Artificial Intelligence and the Value of Privacy-Preserving Tools to Distinguish Who Is Real Online , journal =

    Steven Adler and Zo. Personhood Credentials: Artificial Intelligence and the Value of Privacy-Preserving Tools to Distinguish Who Is Real Online , journal =. 2024 , url =

  32. [32]

    34th USENIX Security Symposium (USENIX Security 25) , publisher =

    Maurice Shih and Michael Rosenberg and Hari Kailad and Ian Miers , title =. 34th USENIX Security Symposium (USENIX Security 25) , publisher =. 2025 , isbn =

  33. [33]

    arXiv preprint arXiv:2607.08077 , year =

    Ethan Roland and Murat Cubuktepe and Erick Martinez and Stijn Servaes and Keenan Pepper and Mike Vaiana and Diogo Schwerz de Lucena and Judd Rosenblatt and Addie Foote and Cem Anil and Alex Cloud , title =. arXiv preprint arXiv:2607.08077 , year =

  34. [34]

    arXiv preprint arXiv:2604.06436 , year =

    Manish Bhatt and Sarthak Munshi and Vineeth Sai Narajala and Idan Habler and Ammar Al-Kahfah and Ken Huang and Joel Webb and Blake Gatto and Md Tamjidul Hoque , title =. arXiv preprint arXiv:2604.06436 , year =

  35. [35]

    Benjamin Erichson , title =

    Xinkai Zhang and Zhipeng Wei and Huanli Gong and Jing Ting Zheng and Yuchen Zhang and Yue Dong and N. Benjamin Erichson , title =. arXiv preprint arXiv:2605.11002 , year =

  36. [36]

    Zico Kolter , title =

    Eliot Krzysztof Jones and Mateusz Dziemian and Matt Fredrikson and J. Zico Kolter , title =. arXiv preprint arXiv:2606.02644 , year =

  37. [37]

    arXiv preprint arXiv:2502.16797 , year =

    Erik Jones and Meg Tong and Jesse Mu and Mohammed Mahfoud and Jan Leike and Roger Grosse and Jared Kaplan and William Fithian and Ethan Perez and Mrinank Sharma , title =. arXiv preprint arXiv:2502.16797 , year =

  38. [38]

    Knight , title =

    David Campbell and Neil Kale and Udari Madhushani Sehwag and Bert Herring and Nick Price and Dan Borges and Alex Levinson and Christina Q. Knight , title =. arXiv preprint arXiv:2603.01246 , year =

  39. [39]

    arXiv preprint arXiv:2605.19722 , year =

    Isaac David and Arthur Gervais , title =. arXiv preprint arXiv:2605.19722 , year =

  40. [40]

    arXiv preprint arXiv:2607.05842 , year =

    Mingchen Li and Meikang Qiu and Zifan Peng and Heng Fan and Song Fu and Junhua Ding and Yunhe Feng , title =. arXiv preprint arXiv:2607.05842 , year =

  41. [41]

    arXiv preprint arXiv:2603.14332 , year =

    Ziling Zhou , title =. arXiv preprint arXiv:2603.14332 , year =

  42. [42]

    Proceedings of the 42nd International Conference on Machine Learning , series =

    Yuxuan Zhu and Antony Kellermann and Dylan Bowman and Philip Li and Akul Gupta and Adarsh Danda and Richard Fang and Conner Jensen and Eric Ihli and Jason Benn and Jet Geronimo and Avi Dhir and Sudhit Rao and Kaicheng Yu and Twm Stone and Daniel Kang , title =. Proceedings of the 42nd International Conference on Machine Learning , series =. 2025 , url =

  43. [43]

    doi:10.48550/arXiv.2410.02828 , note =

    2024 , howpublished =. doi:10.48550/arXiv.2410.02828 , note =

  44. [44]

    Advances in Neural Information Processing Systems, Datasets and Benchmarks Track , year =

    JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models , author =. Advances in Neural Information Processing Systems, Datasets and Benchmarks Track , year =

  45. [45]

    Claude Fable 5 and Claude Mythos 5 , year =

  46. [46]

    Claude Fable 5 & Claude Mythos 5 System Card , year =

  47. [47]

    Redeploying Fable 5 , year =

  48. [48]

    Introducing Grok 4.5 , year =

  49. [49]

    Imagine Overview , year =

  50. [50]

    Ofcom Launches Investigation into X over Grok Sexualised Imagery , year =

  51. [51]

    Commission Investigates Grok and X's Recommender Systems under the Digital Services Act , year =

  52. [52]

    2026 , month = jul, howpublished =

    Brittain, Blake , title =. 2026 , month = jul, howpublished =

  53. [53]

    34th USENIX Security Symposium (USENIX Security 25) , year =

    Mark Russinovich and Ahmed Salem and Ronen Eldan , title =. 34th USENIX Security Symposium (USENIX Security 25) , year =

  54. [54]

    Findings of the Association for Computational Linguistics: ACL 2025 , year =

    Yifan Jiang and Kriti Aggarwal and Tanmay Laud and Kashif Munir and Jay Pujara and Subhabrata Mukherjee , title =. Findings of the Association for Computational Linguistics: ACL 2025 , year =. doi:10.18653/v1/2025.findings-acl.1311 , url =

  55. [55]

    Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =

    Zixuan Weng and Xiaolong Jin and Jinyuan Jia and Xiangyu Zhang , title =. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =. doi:10.18653/v1/2025.emnlp-main.100 , url =

  56. [56]

    Introducing Trusted Access for Cyber , year =

  57. [57]

    2026 , month = may, howpublished =

    Scaling Trusted Access for Cyber with. 2026 , month = may, howpublished =

  58. [58]

    2025 , pages =

    Ha, Junwoo and Kim, Hyunjun and Yu, Sangyoon and Park, Haon and Yousefpour, Ashkan and Park, Yuna and Kim, Suhyun , booktitle =. 2025 , pages =. doi:10.18653/v1/2025.acl-long.805 , url =

  59. [59]

    Souly, Alexandra and Lu, Qingyuan and Bowen, Dillon and Trinh, Tu and Hsieh, Elvis and Pandey, Sana and Abbeel, Pieter and Svegliato, Justin and Emmons, Scott and Watkins, Olivia and Toyer, Sam , booktitle =. A. 2024 , url =

  60. [60]

    2024 , url =

    Mazeika, Mantas and Phan, Long and Yin, Xuwang and Zou, Andy and Wang, Zifan and Mu, Norman and Sakhaee, Elham and Li, Nathaniel and Basart, Steven and Li, Bo and Forsyth, David and Hendrycks, Dan , booktitle =. 2024 , url =

  61. [61]

    2410.10700 , archiveprefix =

    Ren, Qibing and Li, Hao and Liu, Dongrui and Xie, Zhanxu and Lu, Xiaoya and Qiao, Yu and Sha, Lei and Yan, Junchi and Ma, Lizhuang and Shao, Jing , year =. 2410.10700 , archiveprefix =

  62. [62]

    2024 , howpublished =

  63. [63]

    arXiv preprint arXiv:2503.14499 , year =

    Thomas Kwa and Ben West and Joel Becker and Amy Deng and Katharyn Garcia and Max Hasin and Sami Jawhar and others , title =. arXiv preprint arXiv:2503.14499 , year =

  64. [64]

    arXiv preprint arXiv:2411.15114 , year =

    Hjalmar Wijk and Tao Lin and Joel Becker and Sami Jawhar and Neev Parikh and Thomas Broadley and Lawrence Chan and others , title =. arXiv preprint arXiv:2411.15114 , year =

  65. [65]

    AI & Society , volume =

    Stuart Armstrong and Nick Bostrom and Carl Shulman , title =. AI & Society , volume =. 2016 , doi =

  66. [66]

    Preparedness Framework, Version 2 , year =

  67. [67]

    Science and Engineering Ethics , volume =

    Forge, John , title =. Science and Engineering Ethics , volume =. 2010 , doi =

  68. [68]

    Review of Contemporary Philosophy , volume =

    Bostrom, Nick , title =. Review of Contemporary Philosophy , volume =. 2011 , url =

  69. [69]

    Journal of Responsible Innovation , volume =

    Grinbaum, Alexei and Adomaitis, Laurynas , title =. Journal of Responsible Innovation , volume =. 2024 , doi =

  70. [70]

    2024 IEEE Security and Privacy Workshops (SPW) , pages =

    Daniel Kang and Xuechen Li and Ion Stoica and Carlos Guestrin and Matei Zaharia and Tatsunori Hashimoto , title =. 2024 IEEE Security and Privacy Workshops (SPW) , pages =. 2024 , publisher =. doi:10.1109/SPW63631.2024.00018 , url =

  71. [71]

    Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages =

    R. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages =. 2024 , publisher =. doi:10.18653/v1/2024.naacl-long.301 , url =

  72. [72]

    2025 , volume =

    Cui, Justin and Chiang, Wei-Lin and Stoica, Ion and Hsieh, Cho-Jui , booktitle =. 2025 , volume =

  73. [73]

    2025 , month = nov, howpublished =

    Disrupting the First Reported. 2025 , month = nov, howpublished =

  74. [74]

    Introducing Claude Opus 4.7 , year =

  75. [75]

    Pappas and Eric Wong , title =

    Patrick Chao and Alexander Robey and Edgar Dobriban and Hamed Hassani and George J. Pappas and Eric Wong , title =. 2025 IEEE Conference on Secure and Trustworthy Machine Learning , year =

  76. [76]

    2026 , month = jun, howpublished =

  77. [77]

    What's New in Claude Sonnet 5 , year =

  78. [78]

    Gemini 3.5 Flash , year =

  79. [79]

    32nd USENIX Security Symposium (USENIX Security 23) , year =

    Searles, Andrew and Nakatsuka, Yoshimichi and Ozturk, Ercan and Paverd, Andrew and Tsudik, Gene and Enkoji, Ai , title =. 32nd USENIX Security Symposium (USENIX Security 23) , year =

  80. [80]

    2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC) , year =

    Plesner, Andreas and Vontobel, Tobias and Wattenhofer, Roger , title =. 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC) , year =

Showing first 80 references.