Pith. sign in

Paper Citation Record · LEDGER

Tamper-Resistant Safeguards for Open-Weight LLMs

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2408.00761.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.00761 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:18:32.210666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:58:47.266895Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 84378219-be70-47eb-b045-e8be90f1fa93 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 144

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:25.958886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:16a96f8f72e16aceae4fa16f665324c9d4d53ad846d3be189c23b2aa16308a1a

Observation 1f1dad06-36e7-4932-93e6-c34deb4c8fa6 · inbound

RoboSignature: Robust Signature and Watermarking on Network Attacks cites this paper.

RoboSignature: Robust Signature and Watermarking on Network Attacks Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:32.210666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:18:32.210666Z digest=sha256:2b985ce0255a3a1eb8722e87cb3b2416a347fe2bd9b63f1eaaacd65c8ea9dc21

Observation f8c51ef2-9ac0-4e6e-88a9-0c3844e4892b · inbound

Open Problems in Machine Unlearning for AI Safety cites this paper.

Open Problems in Machine Unlearning for AI Safety Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:10.635906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:10.635906Z digest=sha256:dab331a71611e66d72402c994cb85a0d7ddc75178f04355c73edb0b773fcf831

Observation afe5cd27-eb55-428b-8e2f-eb0801853fd4 · inbound

Position: Adversarial ML for LLMs Is Not Making Any Progress cites this paper.

Position: Adversarial ML for LLMs Is Not Making Any Progress Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T12:47:21.787897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:47:21.787897Z digest=sha256:263b0ae6f631eeb86677158d0bdd937c727cc798fa3313a778cad2ebaa4deb29

Observation 9a2486f0-3e1b-4f9f-b395-d26a1835bd4d · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.348837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.348837Z digest=sha256:2a69a03223e4ad5be698bce94d1086cabd1db92aeb2a2384f88f70474caf8b38

Observation 5993d233-5c97-45ba-b019-6bfcebac0fa6 · inbound

Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond cites this paper.

Towards LLM Unlearning Resilient to Relearning Attacks: A Sharpness-Aware Minimization Perspective and Beyond Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:31.602077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:31.602077Z digest=sha256:beaa4688978a230cd39d8de31332a8d5dba9ff1dc9aeb6c5620f1d971454ef4c

Observation a575d826-5956-4c63-bfce-fa33d383308e · inbound

Jailbreaking to Jailbreak cites this paper.

Jailbreaking to Jailbreak Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T17:02:47.451468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:02:47.451468Z digest=sha256:56e5919af2500affc4bddfb088c35429959d3f48538a8ea33cd258e25a85ad8b

Observation 99ed0a68-db84-4abd-bb7c-0eabd1b42ff8 · inbound

Secure LLM Fine-Tuning via Safety-Aware Probing cites this paper.

Secure LLM Fine-Tuning via Safety-Aware Probing Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:11:35.792350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T13:07:09.402763Z digest=sha256:0ecc85c17473c6696a37b26cd94bc01ce96dcf09699fca0960079915cd8c060d

Observation cedf324f-8abb-4e10-95c4-f2a870cba0f3 · inbound

Existing Large Language Model Unlearning Evaluations Are Inconclusive cites this paper.

Existing Large Language Model Unlearning Evaluations Are Inconclusive Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.125060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.125060Z digest=sha256:e2bd1737211988f8d26c2c140598f8a190e59de54e5811aabdf3da238c9cf752

Observation 72cbd017-c19c-46bb-ae8f-632a1a95f874 · inbound

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods cites this paper.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.881113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.881113Z digest=sha256:d138fdaa3cad7597bcd287854668d49538feb1aede145ba419dc3859a25082d9

Observation cb9ce222-09c4-4978-ae18-8bf30d6a1af2 · inbound

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety cites this paper.

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:04.114372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:04.114372Z digest=sha256:a6c637377a3302008c754ae5b910449d8ac6c5c657eb33a1489ceef15ed3eb85

Observation a075f13e-8f9b-48b1-990d-36327b68a518 · inbound

Step-by-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models cites this paper.

Step-by-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:56:31.605239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:56:31.605239Z digest=sha256:3c60e8c05f4a7cd950e7b9059a1e24e05a1f0f9978fff96f0fd4a57e040956fb

Observation 6a385aea-d936-4f8c-aa82-dbb279c9b69a · inbound

Technical Requirements for Halting Dangerous AI Activities cites this paper.

Technical Requirements for Halting Dangerous AI Activities Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:50:30.826374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:50:30.826374Z digest=sha256:3368c62ec67a49e780a0601cf9f740ceb8f93db69561ad6832ad99dc08c0d68f

Observation f87b89ab-3b43-4f80-b269-6f3f7885f2de · inbound

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models cites this paper.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:26.524698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:26.524698Z digest=sha256:277daf36d8bf5664e2ca5c290422d56057bb202f7cc3e70a0cbe4f823ed50f30

Observation 79694803-20ba-4bc8-a100-c1970821ae51 · inbound

Towards Integrated Alignment cites this paper.

Towards Integrated Alignment Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 2025

Resolution
malformed identifier
no resolver link, observed 2026-08-05T22:56:28.187316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:56:28.187316Z digest=sha256:741dde7bf2d0270c02f467483c38993697107b23c89214371db8fd5a7b897554

Observation 7b786040-f6e5-476a-bf06-72165ed3e913 · inbound

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning cites this paper.

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:46:17.116528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T10:44:53.516653Z digest=sha256:f327d1450eace8e179bfd210d180cfbaa767748a069ae8dfd354447ef757bf3a

Observation b82adb25-374f-4802-bf7b-ceffd2bb1671 · inbound

CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs cites this paper.

CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:14:00.155118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T04:13:13.231293Z digest=sha256:7e08e89f4bb47b36e5aa13d4e62afc6366e91fe5f3752795ea38a18794e1819a

Observation 5409273a-ae7e-4ee8-beaf-5d42683a99a9 · inbound

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories cites this paper.

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T18:41:34.212799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:41:34.212799Z digest=sha256:54edad6c559f02406276cfd1c91c812f6b0626238a8f0a1fae4b2180588d8668

Observation 3cb9b195-6a58-47f2-9ecf-447dd540ddd0 · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.175635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:9934b68948864943e2caa6808f2e916d933a54ac3b526bfbc6274713107222af

Observation 318d1ff5-2362-4a97-8bd0-5f0406b44834 · inbound

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains cites this paper.

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.104378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T17:53:57.169962Z digest=sha256:9245f6cc4878f6d38f342af405616b1fff39821dd1753ccf8fdaeffa9e715a0a

Observation 98e7d853-8b19-4e6a-b457-09320dd21d99 · inbound

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter cites this paper.

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:06:59.987995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T01:06:44.126131Z digest=sha256:778e4a3c677c105eba28018d466ea3605b1678e587f00f5195c5310d73235703

Observation c736ebea-b22e-40f4-9696-7b30414548fc · inbound

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries cites this paper.

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.108905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T21:01:25.549340Z digest=sha256:429c3d87c169be7e56f9c15414a7ac6dc0899be3718544396a29ffa33094bc09

Observation 943069a1-ce44-4af4-aa9e-9cc7adefb9c1 · inbound

RepSelect: Robust LLM Unlearning via Representation Selectivity cites this paper.

RepSelect: Robust LLM Unlearning via Representation Selectivity Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.268760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T03:34:32.388152Z digest=sha256:8be01d5d6140713860ac7f99b92c18338f53709d38ac8c063eb8601eba540e11

Observation 0685c087-0354-46ee-85c4-07b18f4a1766 · inbound

FlipGuard: Defending Large Language Models Against Quantization-Conditioned Backdoor Attacks cites this paper.

FlipGuard: Defending Large Language Models Against Quantization-Conditioned Backdoor Attacks Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.661358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T09:39:35.673341Z digest=sha256:06cb7993cd1a4e0bf94af3d5da99ce50a9ed187edaf27f05aa0f4b76cb98ecd0

Observation b67c46e8-f480-457c-985b-e1b14276b069 · inbound

Breaking the Rounding Trap: Securing LLMs against Quantization-Conditioned Backdoors cites this paper.

Breaking the Rounding Trap: Securing LLMs against Quantization-Conditioned Backdoors Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:04:28.742903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T07:47:18.350953Z digest=sha256:edabd880d3d6e41bb7873d8020baa011e49fc57aba727d13f2b05fe9f3b4c253

Observation 38d2f7f6-5c06-442c-8c8b-ac8bcc26e93d · inbound

Breaking the Rounding Trap: Securing LLMs against Quantization-Conditioned Backdoors cites this paper.

Breaking the Rounding Trap: Securing LLMs against Quantization-Conditioned Backdoors Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T04:39:06.903773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:39:06.903773Z digest=sha256:f1da422dd351b5b7fedef466f9c1dab950b454f945af190371f3edfd4ac55e92

Observation 3173a2fd-05d9-4853-8bcc-199d3f156a70 · inbound

LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats cites this paper.

LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-02T10:25:19.922744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:25:19.922744Z digest=sha256:c8cfcbe404c6565233d8265b209ab0e099ab21fda0942d3015774a1b8946f5e6

Observation f1f6f2b5-8b9c-41e2-a1b3-d9be574f659c · inbound

Engineering Trustworthy Agentic AI for Critical Systems cites this paper.

Engineering Trustworthy Agentic AI for Critical Systems Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:27.834694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:27.834694Z digest=sha256:38bc9199d7b4d0fd60fd72a315906f2d4b7d321fa31993ad9240940f30ee31ba

Observation 94205cbb-2c55-4207-bd4c-53627cf2b5ca · inbound

Emergent Misalignment Recruits a Pre-existing Persona Subspace cites this paper.

Emergent Misalignment Recruits a Pre-existing Persona Subspace Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:22.166900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:46:22.166900Z digest=sha256:739fcdd0b20c2d8557ed4d7bc55cb38261443e7f4ae73248fd3dda0ef5128d67