Pith. sign in

Paper Citation Record · LEDGER

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods

As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.10236.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10236 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:08.568150Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6224a955-ca34-442b-aeb5-51011391d676 · outbound

This paper cites Who's Harry Potter? Approximate Unlearning in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Who's Harry Potter? Approximate Unlearning in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.784382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.784382Z digest=sha256:a6921b3b2ae001908474071b7dbe58e9e756165af4b46b6b53f4d1e5a3df0443

Observation 784e5666-a270-4ef5-a7d0-1fb1b4a0cedc · outbound

This paper cites 6 Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods 6 Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.905230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.905230Z digest=sha256:b14635da5ae2eaa4d5b9e028f4a5259ffbcd55722f61fb0f68ce194410403c94

Observation 063ad588-eda8-45ea-9ed0-4369c984792e · outbound

This paper cites Mistral 7B.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Mistral 7B

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.100886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.100886Z digest=sha256:0565dd024890cba35db9fb06aa32d0bc8518126e23153d797164193b2ca5b2f2

Observation fa825a17-919c-4dcf-a4ee-b15251b88615 · outbound

This paper cites Eight Methods to Evaluate Robust Unlearning in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Eight Methods to Evaluate Robust Unlearning in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.250365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.250365Z digest=sha256:a8f1a777ced503bedf3d7c7b30f2d237366411da408bbc7af60da2ebc00e17b5

Observation 7c4baf1d-82b2-43f2-ad2d-084173759af0 · outbound

This paper cites TOFU: A Task of Fictitious Unlearning for LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods TOFU: A Task of Fictitious Unlearning for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.358259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.358259Z digest=sha256:242e7cb05b84c92a12c43adc5324b978c3bbeb6b6bc8b6ed93148871adfa6ab5

Observation 81467daf-364b-48ba-afc0-c08fe4e21be5 · outbound

This paper cites McKinney, Anvith Thudi, Juhan Bae, Tara Rezaei Kheirkhah, Nicolas Papernot, Sheila A.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods McKinney, Anvith Thudi, Juhan Bae, Tara Rezaei Kheirkhah, Nicolas Papernot, Sheila A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.534487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:37:07.411854Z digest=sha256:7c34a8e93bb2e85ef02bd3c6ed05cd1187479e9fd7d2adeb7b372d8dc37ba236

Observation 3f40819a-a57e-4375-ae14-5ec259a731e8 · outbound

This paper cites Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.507769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.507769Z digest=sha256:3f582246d9a4f6b48dfa3481c6da510b90c236e58575f58b0e25d573d927d201

Observation 21c52f53-8e11-467b-b0bb-db5c958e57cc · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods tinyBenchmarks: evaluating LLMs with fewer examples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.670301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.670301Z digest=sha256:3faa011a320a278dc478b5954bf13c53368b1a45d5e0966a117d9cef4ff08ba1

Observation 7e31cf72-4903-4ae0-b01f-ba26f6d10686 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.725074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.725074Z digest=sha256:87fa233cc67856e074a3b627bc0354f1ee9556e94e622c0a4434c25c844cc9a4

Observation 7f7bb65d-affd-40f1-97b4-2e08bda272ab · outbound

This paper cites UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.816508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.816508Z digest=sha256:f83ecabf343c6645536d41e2feca40e203e0e7f2e72f7831b7258a18008fb7a5

Observation 72cbd017-c19c-46bb-ae8f-632a1a95f874 · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.881113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.881113Z digest=sha256:0b84778064e19075dc1881d34b9c01ebb83f1868e44b24f6ddcdffd80511be2e

Observation dd246f8c-77b8-44ad-b04e-a47d6ebc2c5e · outbound

This paper cites AI Sandbagging: Language Models can Strategically Underperform on Evaluations.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.975951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.975951Z digest=sha256:c67a40017d649ba338c6c5505e27c7a890217769228740caf6f814ba432eb591

Observation d61c81e3-d491-44ad-b92a-caff49b54915 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Jailbroken: How Does LLM Safety Training Fail?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.055115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.055115Z digest=sha256:774dc6242b865fae5a86c44af9aaf38ca4d225c54729848ab791dcd1ddf43dac

Observation 2497b4e0-902a-4fe1-87aa-e909ac647df9 · outbound

This paper cites In-Context Learning Can Re-learn Forbidden Tasks.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods In-Context Learning Can Re-learn Forbidden Tasks

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:37:08.751340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:37:08.141780Z digest=sha256:df34a3257e313b29d3c28bbb93519a02f4e985f2496c8ddaa0d7640995e768e0

Observation e7b22fde-4c56-45b1-8e2e-19c8d5f09556 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Low-Resource Languages Jailbreak GPT-4

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.222926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.222926Z digest=sha256:3802a00d3f9fa4532aa491baddd36e79675f007f6955c3bb2c7b7990f1cdef8d

Observation e9b446f5-9593-4ab4-8bb8-172c1d9c5ae5 · outbound

This paper cites Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.295456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.295456Z digest=sha256:bbb8793972aaa79d88cef4c8e503938165062668e12c5383920921294f37d0a7

Observation daf4a539-e064-49db-90ab-d9a500267018 · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Improving Alignment and Robustness with Circuit Breakers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.378893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.378893Z digest=sha256:152b5dc2719ba12fbda2e4ea2b8bfc201ed55c5df201ec794a91adfdb2ea4ec2

Observation ff311a48-1010-47e9-b739-5526929b2600 · outbound

This paper cites [2024], a subset of 100 data points selected from MMLU (Massive Multitask Language Understanding) Hendrycks et al.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods [2024], a subset of 100 data points selected from MMLU (Massive Multitask Language Understanding) Hendrycks et al

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.299362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:37:08.441861Z digest=sha256:069999a731761ee46175c93dff1ee26881c15da33386f193023dd6c9655491e1

Observation f746d274-9a40-4499-b45c-c5bba14a7e2d · outbound

This paper cites Right Format.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Right Format

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.131821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:37:08.568150Z digest=sha256:cebf6cd7cdbf1d98eebe89331a1e2859f4e657fbffb509555c12bef81932ce89

Observation 381b5830-0b54-4e86-b56f-edf92e64f8bd · outbound

This paper cites The Elicitation Game: Evaluating Capability Elicitation Techniques.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods The Elicitation Game: Evaluating Capability Elicitation Techniques

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.012174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.012174Z digest=sha256:b9eab3545bd8c7ce60a69a3a20902950120f8b697ac0f26b4094c98ccd1456d8

Observation 546dad7a-a2f3-4619-85ee-905b22e762e5 · outbound

This paper cites Continual Learning and Private Unlearning.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Continual Learning and Private Unlearning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.178037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.178037Z digest=sha256:384fd455b7c769e7237300bc8fe4a064c1508f9f0ec4ec1252fca1ac8bef30d5

Observation 6e0fc9f6-87cb-4b06-91ca-970599e27cbd · outbound

This paper cites Erasing Conceptual Knowledge from Language Models.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Erasing Conceptual Knowledge from Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.829821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.829821Z digest=sha256:599b7213d028f087f5326e0d17a13e550b09c875b6c13db4cf505db1dd2b8479

Observation 9ac562e3-c5b1-46c0-8e04-7102890168a9 · outbound

This paper cites Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.683217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.683217Z digest=sha256:422da25db1cd0e57e477729f0a42aff91a29c89301b15b2b209aa316427c6c86

Observation 9c1bd043-7b6f-478f-bcbb-b698ff529c13 · outbound

This paper cites Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.722626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.722626Z digest=sha256:21d3351e18a56888965ae9d9cdab53d4f3efb2559c61683a491ff899ee2f2b63

Pith citing papers

No inbound Pith citation observations are available.