REVIEW 3 major objections 5 minor 35 references
Activation probes can screen for risk but cannot adjudicate same-topic harmful intent.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 07:05 UTC pith:VMUJYAMU
load-bearing objection A carefully controlled empirical negative result on activation probes, with the main open risk being the construct validity of its paired 'intent' cohorts — still worth refereeing seriously. the 3 major comments →
The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is an empirical operating-point limit, called the entanglement wall: a fixed residual-stream direction that nearly perfectly ranks harmful against benign prompts from different sources retains only weak ranking separation when the benign control shares the harmful prompt's topic and surface frame. On the twin cohorts, transferred AUROC falls to 0.590-0.690 and fixed-threshold operating points fall far below the pre-specified criterion; every observed threshold crossing either used guard preselection or directly fitted the pair boundary. Direct fitting recovers in-corpus separation (AUROC 0.888-0.929 under pair-grouped CV) but degrades under category and generation-batch hol
What carries the argument
The load-bearing object is the transferred mean-difference activation direction: a single residual-stream vector fitted on a disjoint source contrast (harmful vs benign prompts from separate corpora) and then applied unchanged to same-topic twin pairs. The twin pairs are constructed so that a camouflaged harmful prompt and a benign twin share topic and surface frame but differ in intended outcome. The paper also builds a deployment sensor that compares generation-time activations against 43 intent zones (24 dangerous, 19 harmless) using nearest-medoid cosine scores and a zero-threshold margin gate. Together these let the paper separate ranking ability (AUROC) from useful operating points (TN
Load-bearing premise
The load-bearing premise is that each Twin pair really differs in intended harm (camouflaged arm harmful, twin benign); this is validated only behaviorally on one model with an intent-conditioned judge, while the primary subset is preselected by a guard. If the judge maps surface or template cues to 'harmful,' the measured transfer failure could be a style contrast rather than a consequence distinction.
What would settle it
Take a fresh batch of same-topic harmful/benign pairs labeled by human annotators rather than an LLM judge, fix the read layer before seeing any scores, and apply the transferred mean-difference probe (or the zone-margin gate) with a pre-registered threshold. If the probe reaches AUROC at least 0.90 or TNR at least 40 percent at 95 percent TPR on the full cohort, the entanglement wall as stated would be refuted; if human labels show the pairs do not differ in intended harm, the negative result would be explained away.
If this is right
- Activation-score gates are validated as first-stage risk screens: they catch most compliant attacks in a taxonomy-selected pool, but their false-block rate peaks on risk-adjacent benign prompts (59.6-68.4 percent on XSTest), so they should not be the final word.
- Cross-corpus AUROC near 0.99 does not license claims about intent adjudication; the same direction's matched-pair AUROC (0.590-0.690) shows that source-contrast benchmarks overstate same-topic separation.
- Directly trained pair classifiers do not escape the wall: they lose separation under category and generation-batch hold-out and block most XSTest prompts, so in-corpus pair fitting mostly captures generation or curation cues.
- Dedicated guard models separate the same paired texts far more accurately (AUROC 0.95-0.96), indicating the wall is specific to the residual-stream readouts tested, not to the detectability of the underlying contrast.
- A useful deployment design would pair activation risk scores with a context-sensitive adjudicator for the prompts the first stage flags.
Where Pith is reading between the lines
- A testable extension would probe attention-head readouts or multi-turn intent assembly, as the paper lists for future work; if dynamic attention states separate the same twins while residual-stream scores do not, the entanglement wall would be localized to the residual read point.
- The paper's paired-transfer design is a template for auditing other binary probes: any direction fitted on corpus contrasts risks inheriting source-provenance signal, and the same hold-out logic could be applied to reasoning-mode or sentiment probes.
- Because the construct validity check rests on an LLM judge and guard selection for Twin-n70, human annotation of the twin arms would sharpen the claim that the transfer failure is about intended harm rather than surface style.
- The post hoc addition of the Twin-n163 persistence requirement and the transductive layer selection are the study's most fragile protocol choices; pre-registering the read layer and cohort before scoring would test whether the wall is an artifact of those choices.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether residual-stream activation probes can act as standalone context adjudicators—i.e., distinguish harmful requests from benign twins matched in topic and surface form—or only as broad risk detectors. The authors first evaluate a medoid-based gate across three 7–8B model families, reporting high catch of judge-classified compliant attacks (95.5–97.7%) but high benign-block rates on XSTest (59.6–68.4%). They then reconstruct a published mean-difference harmful-intent probe, audit it with a fully disjoint split (AUROC 0.996–0.999), and transfer the fixed direction to same-topic Twin pairs. Transferred AUROC drops to 0.656–0.819 on guard-selected Twin-n70 and 0.590–0.690 on the full Twin-n163 cohort; no readout that avoids direct pair-boundary fitting meets the pre-specified AUROC≥0.90 or TNR≥40% threshold on Twin-n163. Pair-trained classifiers recover in-corpus separation but weaken under category/generation-batch hold-out and false-block 79.6–100% of XSTest. The paper concludes that, at the tested read points, these activation scores behave as broad-risk detectors rather than standalone context adjudicators, coining the term 'entanglement wall' for this empirical operating-point limit. The manuscript is transparent about post hoc analysis choices and provides code, aggregate data, and a detailed appendix.
Significance. If the measurements are accepted, this is a valuable and well-controlled negative result for activation-space safety methods. The paper goes beyond typical cross-corpus probe evaluations by constructing topic- and surface-matched pairs and testing fixed-direction transfer, with an unusually thorough set of confounding controls: length, permutation, FWER-corrected sign-flip, hierarchical bootstrap, text baselines, provenance classifier, category/batch hold-out, guard-consensus subsets, and a separately specified larger-model extension. The complete disclosure of post hoc decisions and the explicit distinction between exploratory and confirmatory analyses are exemplary. The practical implication—that activation probes can serve as first-stage risk screens but not as standalone same-topic intent adjudicators—is important for deployment decisions. The main weakness is the construct validity of the paired cohorts, which underpins the central negative claim.
major comments (3)
- [§4.4, Appendix D] The construct validity of the Twin pairs—the load-bearing assumption that the camouflaged arm is harmful in intended consequence and the twin is not—rests entirely on a single LLM judge (Claude Opus 4.8) with an intent-conditioned rubric adopted after an initial rubric labelled 44–49% of benign-twin answers compliant. The hidden validation sample of 24 parsed cases with 23 agreements is too small to establish reliability, and the 'manual' resolution is not described as independent human annotation. If the judge systematically maps template/phrasing cues to 'harmful' rather than true intended consequence, the measured transfer failure could be a style contrast rather than a failure of consequence adjudication. This directly undermines the headline claim. I request human annotations (or at minimum a second, architecturally distinct LLM judge) on a meaningful subset of both arms of Twin-n70
- [§4.5, §5.5] The numerical threshold (AUROC≥0.90 or TNR@95%TPR≥40%) was fixed before analysis, but the requirement that the crossing persist on the full Twin-n163 cohort was added at analysis time. The central negative statement—'no axis evaluated without direct pair-boundary fitting reaches the specified numerical threshold on Twin-n163'—therefore depends on a post hoc choice of the decisive cohort. The disclosure is honest, but it reduces the confirmatory status of the result. I ask the authors to either (a) present the Twin-n163 persistence requirement as an exploratory sensitivity analysis rather than part of the main confirmatory claim, or (b) provide a pre-registered analysis plan or a defensible argument that Twin-n163 is the natural target cohort. Additionally, reporting how many configurations would have met the threshold on Twin-n70 but not on Twin-n163 would clarify the practical robustnes
- [§4.3, §5.1] The decisive Twin-n163 cohort is the least construct-validated: only 70 of its 163 pairs are guard-selected, and the behavioral construct check in §4.4 is performed only on Twin-n70. The paper acknowledges this, but the extrapolation to Twin-n163 is essential for the negative claim, since the paper itself notes that Twin-n70 is guard-conditioned and not a population estimate. I recommend reporting the construct check on a random sample of the 93 additional Twin-n163 pairs as well, or explicitly limiting the central claim to the guard-legible subset. Without this, the reader cannot tell whether the transfer failure on Twin-n163 is due to a genuine intent boundary or to the inclusion of poorly constructed pairs.
minor comments (5)
- [Table 4 and Table 5] The directly trained DTW and Transition rows would be easier to interpret if the table also listed the exact definitions of these two axes in a footnote; the main text refers to 'sequence form' and 'directed transition graph' only later in Table 23. Consider adding a cross-reference.
- [Figure 2] The figure is legible but the suite names are partially crowded. A horizontal layout or a small table might be clearer.
- [§5.7] The correlations between prompt-level geometry and generation-window gate scores (0.438–0.833) are described as descriptive. That is fine, but the sentence could be tightened to avoid any impression of mechanism; the paper already states this, so this is purely a wording suggestion.
- [§3.2] The handling of judge refusals differs across families (48 hand-resolved for Llama, unresolved conservatively for others). This is clearly disclosed, but the sensitivity analysis treats all unresolved as compliant; it would be helpful to also report the opposite extreme (all unresolved as non-compliant) to bracket the true catch range.
- [Appendix E] The description of the transferred probe says max-pooling over 64 content tokens. It is not stated how the 64-token window is chosen (e.g., truncation or sliding window). A brief clarification would help reproducibility.
Circularity Check
No circular derivation: the headline claim is an empirical transfer measurement with an externally fitted probe, independently labeled pairs, and disclosed specification limitations rather than a reduction to its own inputs.
full rationale
The central claim is that residual-stream probes act as broad-risk detectors but not standalone context adjudicators. The load-bearing numbers come from a probe direction fitted on an external AdvBench-vs-Alpaca contrast (Appendix E: "A mean-difference direction is fitted to 100 AdvBench and 100 Alpaca prompts. ... Twin evaluation does not refit."), audited on disjoint source splits (Section 5.1), and then transferred to Twin pairs whose labels are defined by a textual guard rule and an intent-conditioned LLM judge, not by activation scores. The measured drop from source AUROC 0.996-0.999 to Twin-n70 0.656-0.819 and Twin-n163 0.590-0.690 is therefore an empirical observation, not an identity derived from the fitting procedure. Pair-fitted classifiers are explicitly separated from the not-fit-on-pairs axes, and their degradation under category and generation-batch hold-out plus high XSTest false-blocking is consistent with corpus cues rather than with a circular prediction. The disclosed specification issues --- transductive layer selection (Section 4.1), the post hoc requirement that crossings persist on Twin-n163 (Sections 4.5, 5.5, and Limitations), guard-conditioned subsets, and possible style/curation cues --- are validity and generalizability concerns, not circular reductions, and the paper states them openly. There is no author-overlap citation chain: the reconstructed protocol is attributed to Llorente-Saguer, not to the present single author, and no uniqueness or ansatz claim is imported via a self-citation. Thus the derivation chain is self-contained against the central claim, and no specific equation or fitted parameter is renamed as a prediction.
Axiom & Free-Parameter Ledger
free parameters (5)
- Early-buffer window parameters (skip 3 tokens, end at token 20 or sentence boundary) =
3 / 20 / sentence boundary
- Danger vs harmless zone counts (24 vs 19) =
24 danger / 19 harmless zones
- Transductive in-model layer indices =
Llama 12, Mistral 14, Qwen 14
- Transfer probe layers =
10/16/19 (Llama/Mistral/Qwen)
- Post-hoc persistence requirement on Twin-n163 =
n/a
axioms (5)
- domain assumption Claude Opus 4.8 provides valid judge labels (COMPLIANCE/REFUSAL/DEGENERATE, and intent-conditioned HARM_IN_CONTEXT for the construct check)
- domain assumption Llama-Guard-3 / WildGuard / ShieldGemma verdicts define harm for pair selection and consensus subsets
- domain assumption The counterfactual paired contrast (same topic and frame, flipped intended outcome) is a valid operationalization of 'context adjudication'
- standard math Standard statistical machinery (paired bootstrap, hierarchical bootstrap, permutation tests, max-|t| FWER correction) is correctly applied
- domain assumption Residual-stream post-attention layer-norm activations are the relevant representation surface
invented entities (1)
-
'Entanglement wall' (named empirical pattern)
independent evidence
read the original abstract
Context can change whether a request is harmful without changing its topic or surface form. We ask whether residual-stream probes distinguish harmful requests from surface-matched benign controls at a useful operating point. Across three 7-8B model families, an activation sensor blocks 95.5-97.7 percent of judge-classified compliant attacks in a taxonomy-selected set. It also blocks 59.6-68.4 percent of XSTest prompts. A fully disjoint audit reconstructs near-ceiling source-contrast AUROC (0.996-0.999), but fixed transfer to matched pairs is weaker: 0.656-0.819 on the guard-selected Twin-n70 subset and 0.590-0.690 on the full Twin-n163 cohort. We test ten axes on the reference family and seven across all families with leakage, hold-out, and permutation controls. On Twin-n163, no axis evaluated without direct pair-boundary fitting reaches the specified numerical threshold. Requiring persistence on that full cohort was added at analysis time. A separately specified 24B/32B extension gives the same result. Pair-trained classifiers weaken under category and generation-batch hold-out and false-block 79.6-100 percent of XSTest at 95 percent in-corpus TPR. At the tested read points, these activation scores behave as broad-risk detectors rather than standalone context adjudicators.
Figures
Reference graph
Works this paper leans on
-
[2]
URL https://arxiv.org/abs/ 2406.11717. Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg. An interpretability illusion for bert,
-
[5]
Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho-Jui Hsieh
URL https://arxiv.org/abs/2404.01318. Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho-Jui Hsieh. Or-bench: An over-refusal benchmark for large lan- guage models,
-
[6]
URL https://arxiv.org/ abs/2405.20947. 10 Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis, and Florian Tramèr. Defeating prompt injections by design,
-
[9]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Ab- hinav Pandey, Abhishek Kadian, et al
URL https: //arxiv.org/abs/2004.02709. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Ab- hinav Pandey, Abhishek Kadian, et al. The llama 3 herd of models,
Pith/arXiv arXiv 2004
-
[10]
URL https://arxiv.org/ abs/2407.21783. Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. Wildguard: Open one-stop modera- tion tools for safety risks, jailbreaks, and refusals of llms,
-
[11]
Xuanli He, Bilgehan Sel, Faizan Ali, Jenny Bao, Hoagy Cunningham, and Jerry Wei
URL https://arxiv.org/abs/ 2406.18495. Xuanli He, Bilgehan Sel, Faizan Ali, Jenny Bao, Hoagy Cunningham, and Jerry Wei. Segment-level coherence for robust harmful intent probing in llms,
-
[12]
Luis Ibanez-Lissen, Lorena Gonzalez-Manzano, Jose Maria de Fuentes, and Nicolas Anciaux
URL https://arxiv.org/abs/2604.14865. Luis Ibanez-Lissen, Lorena Gonzalez-Manzano, Jose Maria de Fuentes, and Nicolas Anciaux. Lpass: Linear probes as stepping stones for vulnerabil- ity detection using compressed llms.Journal of Information Security and Applications, 93: 104125,
-
[13]
doi: https: //doi.org/10.1016/j.jisa.2025.104125
ISSN 2214-2126. doi: https: //doi.org/10.1016/j.jisa.2025.104125. URL https: //www.sciencedirect.com/science/ article/pii/S2214212625001620. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Men- sch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guil- laume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-A...
arXiv 2025
-
[15]
Faaiz Joad, Majd Hawasly, Sabri Boughorbel, Nadir Dur- rani, and Husrev Taha Sencar
URL https://arxiv.org/abs/2406.18510. Faaiz Joad, Majd Hawasly, Sabri Boughorbel, Nadir Dur- rani, and Husrev Taha Sencar. There is more to refusal in large language models than a single direction,
-
[16]
Divyansh Kaushik, Eduard Hovy, and Zachary C
URLhttps://arxiv.org/abs/2602.02132. Divyansh Kaushik, Eduard Hovy, and Zachary C. Lipton. Learning the difference that makes a difference with counterfactually-augmented data,
-
[17]
János Kramár, Joshua Engels, Zheng Wang, Bilal Chughtai, Rohin Shah, Neel Nanda, and Arthur Conmy
URL https: //arxiv.org/abs/1909.12434. János Kramár, Joshua Engels, Zheng Wang, Bilal Chughtai, Rohin Shah, Neel Nanda, and Arthur Conmy. Build- ing production-ready probes for gemini,
Pith/arXiv arXiv 1909
-
[18]
Abhinav Kumar, Chenhao Tan, and Amit Sharma
URL https://arxiv.org/abs/2601.11516. Abhinav Kumar, Chenhao Tan, and Amit Sharma. Prob- ing classifiers are unreliable for concept removal and detection,
-
[20]
URL https: //arxiv.org/abs/2410.10414. Isaac Llorente-Saguer. The geometry of harmful intent: Training-free anomaly detection via angular deviation in llm residual streams, 2026a. URL https://arxiv. org/abs/2603.27412. Isaac Llorente-Saguer. Harmful intent as a geometrically recoverable feature of llm residual streams, 2026b. URL https://arxiv.org/abs/260...
-
[21]
URL https://arxiv.org/abs/ 2402.04249. Alex McKenzie, Urja Pawar, Phil Blandfort, William Bankes, David Krueger, Ekdeep Singh Lubana, and Dmitrii Krasheninnikov. Detecting high-stakes inter- actions with activation probes,
-
[22]
URL https: //arxiv.org/abs/2506.10805. Mistral AI Team. Model card for mistral-small-24b- instruct-2501. Hugging Face Model Card,
-
[23]
Giorgio Piras, Raffaele Mura, Fabio Brau, Luca Oneto, Fabio Roli, and Battista Biggio
URLhttps://arxiv.org/abs/2603.22061. Giorgio Piras, Raffaele Mura, Fabio Brau, Luca Oneto, Fabio Roli, and Battista Biggio. SOM directions are better than one: Multi-directional refusal suppression in language models.Proceedings of the AAAI Con- ference on Artificial Intelligence, 40(39):32728–32736, Mar
-
[24]
ISSN 2159-5399. doi: 10.1609/aaai.v40i39. 40551. URL http://dx.doi.org/10.1609/ aaai.v40i39.40551. Qwen, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin B...
-
[25]
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy
URL https://arxiv.org/abs/2412.15115. Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. Xstest: A test suite for identifying exaggerated safety behaviours in large language models,
-
[26]
Subramanyam Sahoo, Vinija Jain, Aman Chadha, and Di- vya Chaudhary
URL https:// arxiv.org/abs/2308.01263. Subramanyam Sahoo, Vinija Jain, Aman Chadha, and Di- vya Chaudhary. Linear probes detect task format, not reasoning mode in language model hidden states,
-
[27]
URLhttps://arxiv.org/abs/2606.02907. McNair Shah, Saleena Angeline, Adhitya Rajendra Kumar, Naitik Chheda, Kevin Zhu, Vasu Sharma, Sean O’Brien, and Will Cai. The geometry of harmfulness in llms through subconcept probing,
-
[28]
URL https:// arxiv.org/abs/2507.21141. Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Sveg- liato, Scott Emmons, Olivia Watkins, and Sam Toyer. A strongreject for empty jailbreaks,
-
[29]
URL https: //arxiv.org/abs/2402.10260. Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tat- sunori B. Hashimoto. Stanford alpaca: An instruction- following llama model. https://github.com/ tatsu-lab/stanford_alpaca,
-
[30]
Cheng Wang, Zeming Wei, Qin Liu, and Muhao Chen
URL https://arxiv.org/abs/2607.02047. Cheng Wang, Zeming Wei, Qin Liu, and Muhao Chen. False sense of security: Why probing-based malicious input detection fails to generalize, 2025a. URL https: //arxiv.org/abs/2509.03888. Xunguang Wang, Zhenlan Ji, Wenxuan Wang, Zongjie Li, Daoyuan Wu, and Shuai Wang. Sok: Evaluating jailbreak guardrails for large langua...
-
[31]
Yuxin Xiao, Sana Tonekaboni, Walter Gerych, Vinith Suriyakumar, and Marzyeh Ghassemi
URL https: //arxiv.org/abs/2510.04320. Yuxin Xiao, Sana Tonekaboni, Walter Gerych, Vinith Suriyakumar, and Marzyeh Ghassemi. When style breaks safety: Defending llms against superficial style alignment,
-
[32]
URL https://arxiv.org/ abs/2506.07452. Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, Olivia Sturman, and Oscar Wahltinez. Shieldgemma: Generative ai content moderation based on gemma,
-
[33]
URL https: //arxiv.org/abs/2605.29708. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuo- han Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena,
-
[34]
URL https: //arxiv.org/abs/2306.05685. 12 Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. Rep- resen...
-
[250]
10.0 21.0 59.6 Mistral-7B-Instruct-v0.3 8.5 6.7 63.6 Qwen2.5-7B-abl
Family Alpaca WildJailbreak XSTest Llama-3.1-8B-abl. 10.0 21.0 59.6 Mistral-7B-Instruct-v0.3 8.5 6.7 63.6 Qwen2.5-7B-abl. 8.5 10.5 68.4 Table 10: Judge class counts for 480 harmful prompts per model. Llama’s 48 declined cases were manually re- solved as compliance. Mistral and Qwen declines remain unresolved and count as non-compliant. Treating them as co...
1931
-
[2020]
URL https: //arxiv.org/abs/2010.06595. Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Se- hwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tramer, Hamed Hassani, and Eric Wong. Jailbreakbench: An open robustness benchmark for jailbreaking large language models,
Pith/arXiv arXiv 2010
-
[2021]
Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky
URL https: //arxiv.org/abs/2104.07143. Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. With little power comes great responsibility,
-
[2022]
Hongfu Liu, Hengguan Huang, Xiangming Gu, Hao Wang, and Ye Wang
URL https://arxiv.org/abs/ 2207.04153. Hongfu Liu, Hengguan Huang, Xiangming Gu, Hao Wang, and Ye Wang. On calibration of llm-based guard models for reliable content moderation,
-
[2023]
URL https: //arxiv.org/abs/2310.06825. Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghal- lah, Ximing Lu, Maarten Sap, Yejin Choi, and Nouha Dziri. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models,
-
[2024]
URLhttps://arxiv.org/abs/2409.00598. Anthropic. Introducing Claude Opus 4.8. https://www. anthropic.com/news/claude-opus-4-8 ,
-
[2025]
URL https://arxiv.org/abs/2503.18813. Max Fomin. When benchmarks lie: Evaluating malicious prompt classifiers under true distribution shift,
-
[2026]
URLhttps://arxiv.org/abs/2602.14161. Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Be- rant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hanna Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mul- caire, Qiang Ning, Sameer Singh, Noah A. Smith, San-...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.