{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:G7EHDF4XE3DDPIK5BR5CFPAO7U","short_pith_number":"pith:G7EHDF4X","schema_version":"1.0","canonical_sha256":"37c871979726c637a15d0c7a22bc0efd1c2a1ab434570d4f78bc1f73e45e9ca2","source":{"kind":"arxiv","id":"2502.15427","version":1},"attestation_state":"computed","paper":{"title":"Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.LG"],"primary_cat":"cs.CR","authors_text":"Ambrish Rawat, Beat Buesser, Giandomenico Cornacchia, Giulio Zizzo, Kieran Fraser, Kush Varshney, Mark Purcell, Muhammad Zaid Hameed, Pin-Yu Chen, Prasanna Sattigeri","submitted_at":"2025-02-21T12:54:25Z","abstract_excerpt":"As large language models (LLMs) become integrated into everyday applications, ensuring their robustness and security is increasingly critical. In particular, LLMs can be manipulated into unsafe behaviour by prompts known as jailbreaks. The variety of jailbreak styles is growing, necessitating the use of external defences known as guardrails. While many jailbreak defences have been proposed, not all defences are able to handle new out-of-distribution attacks due to the narrow segment of jailbreaks used to align them. Moreover, the lack of systematisation around defences has created significant "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2502.15427","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CR","submitted_at":"2025-02-21T12:54:25Z","cross_cats_sorted":["cs.LG"],"title_canon_sha256":"2e6f3e7471e02a5a3dc1f669063c6474110c692aa7b967d87e95a3f5dd9765a4","abstract_canon_sha256":"aa218e4dd6c56ddca314828648af30fdfcccb9802035fa8d4d01afe3166a3061"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:18:01.733869Z","signature_b64":"N6p4nxCtBdoHNS4bRkN4r4rC0/7qs2FivshkrgonLMsaLW0DzFPBXFBP7wqJYvsQ0Rk0oNiYDkFIcHmXM77/BQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"37c871979726c637a15d0c7a22bc0efd1c2a1ab434570d4f78bc1f73e45e9ca2","last_reissued_at":"2026-07-05T10:18:01.733387Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:18:01.733387Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.LG"],"primary_cat":"cs.CR","authors_text":"Ambrish Rawat, Beat Buesser, Giandomenico Cornacchia, Giulio Zizzo, Kieran Fraser, Kush Varshney, Mark Purcell, Muhammad Zaid Hameed, Pin-Yu Chen, Prasanna Sattigeri","submitted_at":"2025-02-21T12:54:25Z","abstract_excerpt":"As large language models (LLMs) become integrated into everyday applications, ensuring their robustness and security is increasingly critical. In particular, LLMs can be manipulated into unsafe behaviour by prompts known as jailbreaks. The variety of jailbreak styles is growing, necessitating the use of external defences known as guardrails. While many jailbreak defences have been proposed, not all defences are able to handle new out-of-distribution attacks due to the narrow segment of jailbreaks used to align them. Moreover, the lack of systematisation around defences has created significant "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2502.15427","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2502.15427/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2502.15427","created_at":"2026-07-05T10:18:01.733453+00:00"},{"alias_kind":"arxiv_version","alias_value":"2502.15427v1","created_at":"2026-07-05T10:18:01.733453+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2502.15427","created_at":"2026-07-05T10:18:01.733453+00:00"},{"alias_kind":"pith_short_12","alias_value":"G7EHDF4XE3DD","created_at":"2026-07-05T10:18:01.733453+00:00"},{"alias_kind":"pith_short_16","alias_value":"G7EHDF4XE3DDPIK5","created_at":"2026-07-05T10:18:01.733453+00:00"},{"alias_kind":"pith_short_8","alias_value":"G7EHDF4X","created_at":"2026-07-05T10:18:01.733453+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":4,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.22237","citing_title":"Investigating The Security of Modern AI and Cloud Infrastructure","ref_index":153,"is_internal_anchor":false},{"citing_arxiv_id":"2511.12710","citing_title":"Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs","ref_index":69,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05058","citing_title":"SoK: Robustness in Large Language Models against Jailbreak Attacks","ref_index":106,"is_internal_anchor":false},{"citing_arxiv_id":"2604.23593","citing_title":"When AI reviews science: Can we trust the referee?","ref_index":132,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/G7EHDF4XE3DDPIK5BR5CFPAO7U","json":"https://pith.science/pith/G7EHDF4XE3DDPIK5BR5CFPAO7U.json","graph_json":"https://pith.science/api/pith-number/G7EHDF4XE3DDPIK5BR5CFPAO7U/graph.json","events_json":"https://pith.science/api/pith-number/G7EHDF4XE3DDPIK5BR5CFPAO7U/events.json","paper":"https://pith.science/paper/G7EHDF4X"},"agent_actions":{"view_html":"https://pith.science/pith/G7EHDF4XE3DDPIK5BR5CFPAO7U","download_json":"https://pith.science/pith/G7EHDF4XE3DDPIK5BR5CFPAO7U.json","view_paper":"https://pith.science/paper/G7EHDF4X","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2502.15427&json=true","fetch_graph":"https://pith.science/api/pith-number/G7EHDF4XE3DDPIK5BR5CFPAO7U/graph.json","fetch_events":"https://pith.science/api/pith-number/G7EHDF4XE3DDPIK5BR5CFPAO7U/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/G7EHDF4XE3DDPIK5BR5CFPAO7U/action/timestamp_anchor","attest_storage":"https://pith.science/pith/G7EHDF4XE3DDPIK5BR5CFPAO7U/action/storage_attestation","attest_author":"https://pith.science/pith/G7EHDF4XE3DDPIK5BR5CFPAO7U/action/author_attestation","sign_citation":"https://pith.science/pith/G7EHDF4XE3DDPIK5BR5CFPAO7U/action/citation_signature","submit_replication":"https://pith.science/pith/G7EHDF4XE3DDPIK5BR5CFPAO7U/action/replication_record"}},"created_at":"2026-07-05T10:18:01.733453+00:00","updated_at":"2026-07-05T10:18:01.733453+00:00"}