{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:CSDIHOFM2RS5GHLC2SIZ3HU7UE","short_pith_number":"pith:CSDIHOFM","schema_version":"1.0","canonical_sha256":"148683b8acd465d31d62d4919d9e9fa10007b80f7873fa711736964d590a538c","source":{"kind":"arxiv","id":"2407.02855","version":3},"attestation_state":"computed","paper":{"title":"From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL","cs.LG"],"primary_cat":"cs.CR","authors_text":"Chujie Zheng, Hongning Wang, Junxiao Yang, Minlie Huang, Pei Ke, Shiyao Cui, Yida Lu, Zhexin Zhang","submitted_at":"2024-07-03T07:14:05Z","abstract_excerpt":"Large Language Models (LLMs) are known to be vulnerable to jailbreak attacks. An important observation is that, while different types of jailbreak attacks can generate significantly different queries, they mostly result in similar responses that are rooted in the same harmful knowledge (e.g., detailed steps to make a bomb). Consequently, unlearning-based approaches have been proposed to mitigate jailbreak attacks by directly removing harmful knowledge from the model. In this paper, we identify a novel ripple effect of unlearning, wherein LLMs can implicitly unlearn harmful knowledge that was n"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2407.02855","kind":"arxiv","version":3},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CR","submitted_at":"2024-07-03T07:14:05Z","cross_cats_sorted":["cs.CL","cs.LG"],"title_canon_sha256":"aedbe6f5bb600e4005f85c59ce6dbc2a3cdea15036e47e59ab9ded69b85288a2","abstract_canon_sha256":"42f245081569588a8e221fa4e03a9e17708e85157a6c87310058a92800879a6c"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:05:33.651644Z","signature_b64":"j+Z2rylXU3mnnE4brVSZGg7OVJuHZhUVhI0KHOMB9BTD6AQCj5pU0eWHw8FduYvRssLf2VqRZspKy0tmUZlvDQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"148683b8acd465d31d62d4919d9e9fa10007b80f7873fa711736964d590a538c","last_reissued_at":"2026-07-05T11:05:33.651147Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:05:33.651147Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL","cs.LG"],"primary_cat":"cs.CR","authors_text":"Chujie Zheng, Hongning Wang, Junxiao Yang, Minlie Huang, Pei Ke, Shiyao Cui, Yida Lu, Zhexin Zhang","submitted_at":"2024-07-03T07:14:05Z","abstract_excerpt":"Large Language Models (LLMs) are known to be vulnerable to jailbreak attacks. An important observation is that, while different types of jailbreak attacks can generate significantly different queries, they mostly result in similar responses that are rooted in the same harmful knowledge (e.g., detailed steps to make a bomb). Consequently, unlearning-based approaches have been proposed to mitigate jailbreak attacks by directly removing harmful knowledge from the model. In this paper, we identify a novel ripple effect of unlearning, wherein LLMs can implicitly unlearn harmful knowledge that was n"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2407.02855","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2407.02855/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2407.02855","created_at":"2026-07-05T11:05:33.651205+00:00"},{"alias_kind":"arxiv_version","alias_value":"2407.02855v3","created_at":"2026-07-05T11:05:33.651205+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2407.02855","created_at":"2026-07-05T11:05:33.651205+00:00"},{"alias_kind":"pith_short_12","alias_value":"CSDIHOFM2RS5","created_at":"2026-07-05T11:05:33.651205+00:00"},{"alias_kind":"pith_short_16","alias_value":"CSDIHOFM2RS5GHLC","created_at":"2026-07-05T11:05:33.651205+00:00"},{"alias_kind":"pith_short_8","alias_value":"CSDIHOFM","created_at":"2026-07-05T11:05:33.651205+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.27379","citing_title":"Position: The Term \"Machine Unlearning\" Is Overused in LLMs","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20286","citing_title":"Adaptive Probe-based Steering for Robust LLM Jailbreaking","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08936","citing_title":"Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08878","citing_title":"Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05058","citing_title":"SoK: Robustness in Large Language Models against Jailbreak Attacks","ref_index":100,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/CSDIHOFM2RS5GHLC2SIZ3HU7UE","json":"https://pith.science/pith/CSDIHOFM2RS5GHLC2SIZ3HU7UE.json","graph_json":"https://pith.science/api/pith-number/CSDIHOFM2RS5GHLC2SIZ3HU7UE/graph.json","events_json":"https://pith.science/api/pith-number/CSDIHOFM2RS5GHLC2SIZ3HU7UE/events.json","paper":"https://pith.science/paper/CSDIHOFM"},"agent_actions":{"view_html":"https://pith.science/pith/CSDIHOFM2RS5GHLC2SIZ3HU7UE","download_json":"https://pith.science/pith/CSDIHOFM2RS5GHLC2SIZ3HU7UE.json","view_paper":"https://pith.science/paper/CSDIHOFM","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2407.02855&json=true","fetch_graph":"https://pith.science/api/pith-number/CSDIHOFM2RS5GHLC2SIZ3HU7UE/graph.json","fetch_events":"https://pith.science/api/pith-number/CSDIHOFM2RS5GHLC2SIZ3HU7UE/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/CSDIHOFM2RS5GHLC2SIZ3HU7UE/action/timestamp_anchor","attest_storage":"https://pith.science/pith/CSDIHOFM2RS5GHLC2SIZ3HU7UE/action/storage_attestation","attest_author":"https://pith.science/pith/CSDIHOFM2RS5GHLC2SIZ3HU7UE/action/author_attestation","sign_citation":"https://pith.science/pith/CSDIHOFM2RS5GHLC2SIZ3HU7UE/action/citation_signature","submit_replication":"https://pith.science/pith/CSDIHOFM2RS5GHLC2SIZ3HU7UE/action/replication_record"}},"created_at":"2026-07-05T11:05:33.651205+00:00","updated_at":"2026-07-05T11:05:33.651205+00:00"}