Pith. sign in

Paper Citation Record · LEDGER

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods

As of 8 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.10236.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10236 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:08.568150Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6224a955-ca34-442b-aeb5-51011391d676 · outbound

This paper cites Who's Harry Potter? Approximate Unlearning in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Who's Harry Potter? Approximate Unlearning in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.784382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.784382Z digest=sha256:9d23c270f70602d94a071f77be0fe27af92728ccd5039b9d50299370bbac2b68

Observation 784e5666-a270-4ef5-a7d0-1fb1b4a0cedc · outbound

This paper cites 6 Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods 6 Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.905230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.905230Z digest=sha256:a8ceb69a854214b72efa2408f43bdd151f0d5319e5f596cd99ad68b51461e5a9

Observation 063ad588-eda8-45ea-9ed0-4369c984792e · outbound

This paper cites Mistral 7B.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Mistral 7B

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.100886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.100886Z digest=sha256:7279f5b67c2fac57c8be414994dea4a79984586b7f7a42f6a92b58ddc0307756

Observation fa825a17-919c-4dcf-a4ee-b15251b88615 · outbound

This paper cites Eight Methods to Evaluate Robust Unlearning in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Eight Methods to Evaluate Robust Unlearning in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.250365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.250365Z digest=sha256:d023d95535b4a58216a585ecd124845d448a670523955113a2fd6c21b8e09933

Observation 7c4baf1d-82b2-43f2-ad2d-084173759af0 · outbound

This paper cites TOFU: A Task of Fictitious Unlearning for LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods TOFU: A Task of Fictitious Unlearning for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.358259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.358259Z digest=sha256:2cceacda17868791d9eff22d476b97e1647f10e29299c22f58d3462ebd35735c

Observation 81467daf-364b-48ba-afc0-c08fe4e21be5 · outbound

This paper cites McKinney, Anvith Thudi, Juhan Bae, Tara Rezaei Kheirkhah, Nicolas Papernot, Sheila A.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods McKinney, Anvith Thudi, Juhan Bae, Tara Rezaei Kheirkhah, Nicolas Papernot, Sheila A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.534487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:37:07.411854Z digest=sha256:7bd690f3fa9d88309d947f597551a72bf91e7377ae63137af9018c3e53bef648

Observation 3f40819a-a57e-4375-ae14-5ec259a731e8 · outbound

This paper cites Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.507769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.507769Z digest=sha256:418db3c6ce6cf45f3631b27184a548ac8f06f6386a3a2515c4072785f2564b09

Observation 21c52f53-8e11-467b-b0bb-db5c958e57cc · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods tinyBenchmarks: evaluating LLMs with fewer examples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.670301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.670301Z digest=sha256:6dac5003889d8af769ac4dc8366c2ee25bf4ffba59f642c386f4c3ec7bf94904

Observation 7e31cf72-4903-4ae0-b01f-ba26f6d10686 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.725074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.725074Z digest=sha256:2d7e42bdc28650742b842709949674ff245142c5b7330439a330b7c8c4691dcc

Observation 7f7bb65d-affd-40f1-97b4-2e08bda272ab · outbound

This paper cites UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.816508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.816508Z digest=sha256:2655529a1a95cb5b3c451eae84194c0784914e556a2af439c8e10082d06caccf

Observation 72cbd017-c19c-46bb-ae8f-632a1a95f874 · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.881113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.881113Z digest=sha256:b2d967b1ce292370f6b66dfb1de9a0516ce2e995ab1b12a34afa961d2500802d

Observation dd246f8c-77b8-44ad-b04e-a47d6ebc2c5e · outbound

This paper cites AI Sandbagging: Language Models can Strategically Underperform on Evaluations.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.975951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.975951Z digest=sha256:d702991626df191c62e98540b19593012bd21f4d62ed0e163a6dcef82ad31d40

Observation d61c81e3-d491-44ad-b92a-caff49b54915 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Jailbroken: How Does LLM Safety Training Fail?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.055115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.055115Z digest=sha256:d009d234ef4e5d123a8b1afff680ae2cef1d04ab699d3d03910cda46a22c315d

Observation 2497b4e0-902a-4fe1-87aa-e909ac647df9 · outbound

This paper cites In-Context Learning Can Re-learn Forbidden Tasks.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods In-Context Learning Can Re-learn Forbidden Tasks

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:37:08.751340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:37:08.141780Z digest=sha256:f4bb54d79b9d0dd64b68f8fff7b78e5331ece8c34cfec2b237d88dde2a069b2d

Observation e7b22fde-4c56-45b1-8e2e-19c8d5f09556 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Low-Resource Languages Jailbreak GPT-4

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.222926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.222926Z digest=sha256:535b2c78c2bb58de1934beb39bfa8d7a3b7d826f2428feb1f37cabe6331d1ff8

Observation e9b446f5-9593-4ab4-8bb8-172c1d9c5ae5 · outbound

This paper cites Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.295456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.295456Z digest=sha256:6e8db8089ad2ec2ce3552b6021668d9b9cefde5819cfd7e4b9231016d6844c62

Observation daf4a539-e064-49db-90ab-d9a500267018 · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Improving Alignment and Robustness with Circuit Breakers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.378893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.378893Z digest=sha256:0d4cf555191eeededf7d4f11be97b1d8615412095f307f847586d084566d090f

Observation ff311a48-1010-47e9-b739-5526929b2600 · outbound

This paper cites [2024], a subset of 100 data points selected from MMLU (Massive Multitask Language Understanding) Hendrycks et al.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods [2024], a subset of 100 data points selected from MMLU (Massive Multitask Language Understanding) Hendrycks et al

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.299362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:37:08.441861Z digest=sha256:c43dd682afcd35dff3a514fd30e24389c008a15f63c6a26110e653c1fd6b5adc

Observation f746d274-9a40-4499-b45c-c5bba14a7e2d · outbound

This paper cites Right Format.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Right Format

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.131821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:37:08.568150Z digest=sha256:208f35f1ddbd8cd6318aa91d5217a4212ee3d88501de208dccd47ec49d0f2a33

Observation 381b5830-0b54-4e86-b56f-edf92e64f8bd · outbound

This paper cites The Elicitation Game: Evaluating Capability Elicitation Techniques.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods The Elicitation Game: Evaluating Capability Elicitation Techniques

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.012174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.012174Z digest=sha256:455dd83a78aebaef9a789dc690953d1afcfd366596187239010e122e2125bb79

Observation 546dad7a-a2f3-4619-85ee-905b22e762e5 · outbound

This paper cites Continual Learning and Private Unlearning.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Continual Learning and Private Unlearning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.178037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.178037Z digest=sha256:97837f6d8972cde41a833ed2670c27b5642145dd49a627d826cf028429c1c08b

Observation 6e0fc9f6-87cb-4b06-91ca-970599e27cbd · outbound

This paper cites Erasing Conceptual Knowledge from Language Models.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Erasing Conceptual Knowledge from Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.829821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.829821Z digest=sha256:e67e63cd329f27d51a6b789040d869c6d62d20ae566248bf66a0f1c137a12e74

Observation 9ac562e3-c5b1-46c0-8e04-7102890168a9 · outbound

This paper cites Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.683217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.683217Z digest=sha256:12184e308c96e2833487b67868131b864e9db6fbcdf05837469bbc64f431a9a5

Observation 9c1bd043-7b6f-478f-bcbb-b698ff529c13 · outbound

This paper cites Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.722626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.722626Z digest=sha256:cf151cfc9c385beacc1409b0050c4684f79394f9a1a16fe96c5b9b6d28116632

Pith citing papers

No inbound Pith citation observations are available.