Pith. sign in

Paper Citation Record · LEDGER

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

As of 22 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 9 inbound Pith citation observations for arXiv:2502.05209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05209 v4

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:47:15.448064Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:25:35.442838Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T10:46:17.055047Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved78
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf1f4201-9b52-47eb-a0b6-d553ec611530 · outbound

This paper cites write newline.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:14.990982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:14.990982Z digest=sha256:dce4dc146a976658d137936b96b42975b871c9dff1a773bddf7b59bbfd3b8711

Observation 183e7ffa-d049-4cfb-b8cf-0c3e179c6892 · outbound

This paper cites T., O'Brien, J., Soder, L., Bucknall, B., Bluemke, E., Schuett, J., Trager, R., Strahm, L., and Chowdhury, R.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities T., O'Brien, J., Soder, L., Bucknall, B., Bluemke, E., Schuett, J., Trager, R., Strahm, L., and Chowdhury, R

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:14.997251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:14.997251Z digest=sha256:a399497637737f4dc02d83438c97bd5e3eae2cdd6f6bbdad8fb06e695f3c1b8c

Observation 73b1af65-f276-4185-8a02-8c85bf245625 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.002017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.002017Z digest=sha256:2944df873524b5b81c22b76725f4ff6a3395e1970eb24125cb3fa06c6daa114b

Observation 30f4dc67-561b-4b2f-a9fa-dbd1744c6fb6 · outbound

This paper cites Many-shot jailbreaking.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Many-shot jailbreaking

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.007519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.007519Z digest=sha256:024e920d555465697ed1bbff7b1ad9d707f41794b158b944f4edf58236582701

Observation c3e72b8b-62a8-481d-8271-b19d296bebaa · outbound

This paper cites Unlearning in large language models via activation projections, 2025.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unlearning in large language models via activation projections, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.012554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.012554Z digest=sha256:cfdbe870193c0f9456a7109545df91eef48d0cab5786cee05b25b8cd9dd1ca26

Observation b60765c1-c7ba-4511-b905-6f60d5639e29 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Refusal in Language Models Is Mediated by a Single Direction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.017432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.017432Z digest=sha256:4312ccb28d27fdc6d7682be344bc10cc202ab15b94089a7480e5fd663ecffeb0

Observation 633bc806-8686-4c0c-bb0e-16e5805b4fdd · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.023992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.023992Z digest=sha256:298af4ed6377d2fa416948edd4236b97ee16094c7801a2345872d7eea065f5b4

Observation c9d4dc73-3dd6-4f1b-be7c-ac44238e7d1c · outbound

This paper cites Open Problems in Machine Unlearning for AI Safety.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Open Problems in Machine Unlearning for AI Safety

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.029821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.029821Z digest=sha256:2f961aaef2484b87848c6d13a6e3d4d824378de6881f37e97466fdba93c42cc4

Observation f6fcfcd8-651e-4314-9848-00fa399ccd0d · outbound

This paper cites Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.034861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.034861Z digest=sha256:085cde81c932fd7eb9f3c9b2a165f4f6b87cde8a05b982531e138f5428243981

Observation e4f9795a-a64c-442d-95fd-3ba18fbca45d · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.040084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.040084Z digest=sha256:e62b61635ec3698ba53598fb219ebba10247a94b14f8a21d265c061a3e0299da

Observation a305ddd1-337a-4a51-a1d5-5426e98c3f80 · outbound

This paper cites AI and Data Act: Part of Bill C-27, Digital Charter Implementation Act, 2022 , 2022.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI and Data Act: Part of Bill C-27, Digital Charter Implementation Act, 2022 , 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.044779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.044779Z digest=sha256:dd87dc11b886cbf2d6b27095688ad02667d7a2199b78733a393e674068a0f3d3

Observation 9e2320bc-7954-45f0-8146-5a6b68311de4 · outbound

This paper cites A., Jagielski, M., Gao, I., Koh, P.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A., Jagielski, M., Gao, I., Koh, P

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.049489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.049489Z digest=sha256:d52d82a749e794c8bd052bb66aad0d4519fa2b43815717eaebb7601c6c7e0963

Observation 375996cf-fb2f-4c96-8acc-b63e4b626094 · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.054255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.054255Z digest=sha256:9e76b570ed7124a501d95f9ab9ae220e3b8b40ad9663602b1872bf676fdc593d

Observation 9b042a6f-9ed9-45de-ada7-14c98f4f361c · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.059261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.059261Z digest=sha256:dd1a6f42dad3a2b4d666ec6b1a5ec3888bda35f02f81ce05b1bd8da75b102e49

Observation b5b04b81-e9bb-4b32-bb5f-815e9bd6bf24 · outbound

This paper cites Interim Measures for the Management of Generative Artificial Intelligence Services , 2023.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Interim Measures for the Management of Generative Artificial Intelligence Services , 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.064636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.064636Z digest=sha256:f53641bbcf51b4dc2572db8fb92a94d0069687e04f040b01782565b9e57714a1

Observation 037a7571-e26d-4a03-8e60-13c6c5568b60 · outbound

This paper cites G., Islam, M.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities G., Islam, M

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.069152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.069152Z digest=sha256:ada58d8bd26cf051ebb4144e923d9ee74edc5393d146741e1a2385de5bc2eb22

Observation 9fd04bde-0107-41eb-b320-f5373135624a · outbound

This paper cites O., and Nilsson, F.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities O., and Nilsson, F

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.073705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.073705Z digest=sha256:5d1fd163ee3dcdf68ae7fa3057d9f17d07980f20123b243f7330b5da92195421

Observation 7f515d96-f150-4a66-87a0-e43671afe57a · outbound

This paper cites Do Unlearning Methods Remove Information from Language Model Weights?.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Do Unlearning Methods Remove Information from Language Model Weights?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.078108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.078108Z digest=sha256:d373f6fc77f5cf3c646c0f0433f73c659403c21e8ca5774b9c171bb67de33625

Observation e625dda2-cebb-471c-a6ef-c43baa3c9637 · outbound

This paper cites The Llama 3 Herd of Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.082764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.082764Z digest=sha256:c0d7608f4e89b1c2115c0952181da79bb0760759b1c3cd9ca214840a1fb56e22

Observation fafdf18d-29b6-43e4-b8a6-8cf8436227fa · outbound

This paper cites The eu artificial intelligence act.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The eu artificial intelligence act

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.087548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.087548Z digest=sha256:56e5cb414cf42e4eee8d9a1f96c7a9e2ea17598cc78ad3de427269f604774439

Observation feea78e1-b98f-41dc-94a2-5ebc51bc4061 · outbound

This paper cites Scaling Laws for Adversarial Attacks on Language Model Activations.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Scaling Laws for Adversarial Attacks on Language Model Activations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.092289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.092289Z digest=sha256:b1dd7cb3aecff98e1bc3b353540b1df2ef98a9b095445306c2abcc880d971fc1

Observation 5b9d2a9f-2592-4aa0-9ac8-804cadcd21c5 · outbound

This paper cites Towards a science of ai evaluations.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Towards a science of ai evaluations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.097049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.097049Z digest=sha256:2fda79c9295dde6c73c1ed1354c4b0a22283a8d3aa21f0353a86257589e0dbae

Observation db1812b8-3cba-45e1-ab28-f5d46d864bf3 · outbound

This paper cites Erasing Conceptual Knowledge from Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Erasing Conceptual Knowledge from Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.101899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.101899Z digest=sha256:5a0d84ff79759ae1a5e90aa348490cb88291da017126f06a1f230df3cccfa051

Observation ef386264-db9e-4fb4-bfb3-572786246e45 · outbound

This paper cites Stress-Testing Capability Elicitation With Password-Locked Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Stress-Testing Capability Elicitation With Password-Locked Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.106879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.106879Z digest=sha256:b0580cd52d376dc747c324d7f4da1b30a07f9f534931414ccfbb8318a83112f9

Observation ced76700-4f74-427c-98f4-f4fd392bcd53 · outbound

This paper cites Cascade: Exploring hierarchical inference in language models, 2023.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Cascade: Exploring hierarchical inference in language models, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.111767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.111767Z digest=sha256:b948889a612224faac3929b368b248d7a13e9d361133d3e76f3bc98d2b0122f5

Observation 122f75cd-0c4e-4a4f-b5d9-096bb7b5eeee · outbound

This paper cites T., Haghtalab, N., and Steinhardt, J.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities T., Haghtalab, N., and Steinhardt, J

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.947064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.116500Z digest=sha256:0f0766b155d75787e8dcd7f20f8ecb43b705576f1b3c6b7d29f7c5110698600f

Observation 859381dd-9d09-4cc8-ab4a-ecce32ab1439 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Measuring Massive Multitask Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.121434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.121434Z digest=sha256:988a38a3d90b36b89811e80579a6aaa67becd2c9a0f13d1b9037c22c94bce5c4

Observation a1b4633c-ab77-48bb-be91-52cd0af80adc · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.928848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.125861Z digest=sha256:3fe23cad1c8449f045df2f47268b938c4dcb2f9d27853d76b1ff6b23d9cd2f25

Observation ac493279-4922-49e3-a797-b466f436aed1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.130612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.130612Z digest=sha256:30bc511eb6cc41440b0c4a437d3ab4fe434a1dfc0c7d4935c3f0509453c89e5b

Observation c267204e-d9f5-4a06-a2ef-70836602af39 · outbound

This paper cites Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.135860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.135860Z digest=sha256:d670efa61293b41126646878606a1eb1bcdae13eb113a840f5a3ad87961195df

Observation 8dd47542-8160-46b8-a070-04a1f37247bd · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.141334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.141334Z digest=sha256:260201e251315eddb9c21b2562df1d19952203029fb3a7a769e25006548ae064

Observation 3bad5977-5ee9-49b1-a799-7ab9661c9872 · outbound

This paper cites Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.146775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.146775Z digest=sha256:e7454bd7c3bace0bcba4dfadd6fbdbd44c85a9aa729651347a5a2083674ee7f1

Observation 853c3ab5-e289-40be-89ea-4a61d4329711 · outbound

This paper cites Language models resist alignment, 2024.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Language models resist alignment, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.911752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.151928Z digest=sha256:5d3b3195818161edece3fef2bbd462f7b5af881a32c9fdb1a47d478c9750e76b

Observation beaf5ec8-3b43-4564-b970-5b3fa8a64530 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.157036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.157036Z digest=sha256:f270cf9954ad08681ebac6dc5cb946596f2802346fe15768b9a3b85f68ec2d29

Observation 7c6eb7ce-c036-404c-b7b8-55c92326ee63 · outbound

This paper cites Act on the protection of personal information, 2025.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Act on the protection of personal information, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.894782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.161963Z digest=sha256:822e9b95df241be72b63f743e02532e93ce470c43806d8c6caa95a059f2be2fa

Observation 8bb16e20-1d1b-44ef-82f2-e26f100e351b · outbound

This paper cites No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.166472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.166472Z digest=sha256:f9f5fc4d890f193a87362b7660f099d258d406a5afd941fd117d7c541efa2aa4

Observation 43acf8e1-359d-47e2-b21a-ca64edf4fd8d · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.171392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.171392Z digest=sha256:db99a220b62d6092ddc6c0a2302e0fac452d60306f3461cf58689c5c0eb992a3

Observation e85b16c9-7c66-4b34-93ef-ddfa98ce1114 · outbound

This paper cites LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.176057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.176057Z digest=sha256:1e135dc18a4a45a2f0b52718e0149a5bedfd89f1de04e3526f4b02baec4feb2b

Observation c6c86131-280a-44a2-b768-ca80477cde25 · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.181547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.181547Z digest=sha256:ebe2414ca915af1d1d35a69ed86d327a0061cddfe95b462396d993e762761b05

Observation 7569bbfc-337a-4e57-827e-c21e17128f0b · outbound

This paper cites Against The Achilles' Heel: A Survey on Red Teaming for Generative Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Against The Achilles' Heel: A Survey on Red Teaming for Generative Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.186818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.186818Z digest=sha256:ef053b8445f6926a16fff33e3d028ba3d3ef82fa71ffd54f7cc0da2a746bc21d

Observation c0c5d3e0-6632-47ea-a285-1a98c0d9d47c · outbound

This paper cites Continual learning and private unlearning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Continual learning and private unlearning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.878763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.191766Z digest=sha256:cd53a7137da9bc53b5e20731d00b22f873dd59cf17fd33a48ba799f3d0ea20a6

Observation 9a67d7ca-f77f-4652-8175-fc201462abfa · outbound

This paper cites Rethinking Machine Unlearning for Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Rethinking Machine Unlearning for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.196404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.196404Z digest=sha256:b7c0e3f788985817765fc587917ab6c58ff4cdcec3a2dedfa384391bfb932a03

Observation 153ff993-0b35-4b9e-a6a0-c6be10f7c39e · outbound

This paper cites Threats, Attacks, and Defenses in Machine Unlearning: A Survey.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Threats, Attacks, and Defenses in Machine Unlearning: A Survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.201431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.201431Z digest=sha256:31d806271871512efb8f8f7339b1dc56b1b9d30fc6a2371ad620c9099368058b

Observation dfa340c9-2760-493a-9f3e-6fed376a1db0 · outbound

This paper cites Large Language Models Relearn Removed Concepts.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Large Language Models Relearn Removed Concepts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.206547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.206547Z digest=sha256:041d0ab717223f820429b4c8be1faad3ff12385809dd079f09e946fd6eb98848

Observation b1a68a4a-4299-4fca-ba6e-6f07c0521583 · outbound

This paper cites and Rimsky, N.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities and Rimsky, N

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.862145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.211427Z digest=sha256:6a0a16534e0473bcf7522469952305d7fe81aa4cd9138fec0e40ba5690ea51ed

Observation 52289a63-379a-44f5-b904-0db3406e363f · outbound

This paper cites An Adversarial Perspective on Machine Unlearning for AI Safety.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities An Adversarial Perspective on Machine Unlearning for AI Safety

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.215964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.215964Z digest=sha256:18cbebe99dd0ba7f6e5c3860648f53cb4a9c2d00f74dd7ce1263b2a4a98b62f7

Observation c24c5f8c-9705-45b0-b204-b9a06d97a898 · outbound

This paper cites Eight Methods to Evaluate Robust Unlearning in LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Eight Methods to Evaluate Robust Unlearning in LLMs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.221048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.221048Z digest=sha256:aad5270c85ff2f9c17d299ee13978c4b0b02e45835be8b5c2ea916128050320a

Observation 4ce66b5b-93f5-4eb5-883c-4f0784adca3a · outbound

This paper cites Pointer sentinel mixture models, 2016.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Pointer sentinel mixture models, 2016

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.225782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.225782Z digest=sha256:9daf839cbba4c42b4824de47c83eb7be426f8a8c7c2fc6b5e3a1033b77630ae3

Observation 8a9b5cb3-561d-4abb-83f5-92739136e69a · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.833753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.230458Z digest=sha256:162f461b102e5cb97d05e29caea8d6d6fbf765bd6a76764e8cbd69df56b51a36

Observation f24989a3-4478-4e69-aef5-a8658535b9d4 · outbound

This paper cites AI Risk Management Framework : AI RMF (1.0), January 2023.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI Risk Management Framework : AI RMF (1.0), January 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.817088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.234922Z digest=sha256:f454ff1fa3eda7026d872b098ea1a12a7cc377617c30cd3d875a48a71581cc69

Observation 94250e00-55ce-41fe-82d0-0e93f89049fe · outbound

This paper cites Openai system card: December 2024, 2024.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Openai system card: December 2024, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.799205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.239688Z digest=sha256:0a81ae5baebb2a6ce27c800a94a118a6b8812798958f4620aaf367c3f186511c

Observation 11943eb2-7684-42a4-9402-04f4fe61a795 · outbound

This paper cites Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.244674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.244674Z digest=sha256:1d1c78c0bbe457efe061a799c4b334ba263adac826460dcbcccc271f28c13e32

Observation 45cc7030-3d76-4a04-a0b8-f5b5f5e44703 · outbound

This paper cites Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.249627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.249627Z digest=sha256:280763147430870eb2bf379670e8ad3bb1299487fddecd4d071419bfd4cb89b5

Observation a9921514-bedd-4a3d-a40f-c3ecc1d08f61 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.254462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.254462Z digest=sha256:c90a1a656890048aa028bf5fd69d4c1c58423c8d6ff9326db5e10c625ac8c3e5

Observation 7657da45-83b4-4452-96cc-aaf6b36efad7 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep, 2024 a.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Safety alignment should be made more than just a few tokens deep, 2024 a

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.783616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.259603Z digest=sha256:c85949e08da55bb3b812788d38b6e982b812c9addd6ff801c5d73defc64d22a6

Observation 46082899-8732-4da9-ad97-7c6104d9b44c · outbound

This paper cites On Evaluating the Durability of Safeguards for Open-Weight LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities On Evaluating the Durability of Safeguards for Open-Weight LLMs

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.264207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.264207Z digest=sha256:803c0ead3fd920534fe8515c93681d1f47c15615cecf28576b23097fcca65d7c

Observation 46e855c5-3c66-40d2-b412-a59fd126938b · outbound

This paper cites D., Xu, P., Honigsberg, C., and Ho, D.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities D., Xu, P., Honigsberg, C., and Ho, D

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.767226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.269083Z digest=sha256:812deb9d833e1b2f988aae4d91481f41f7ff3b36cdde0a1b261bbde9bfe9bc69

Observation 02e2e5a7-928d-45c4-b7d0-34193dd1bf67 · outbound

This paper cites Open Problems in Technical AI Governance.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Open Problems in Technical AI Governance

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.273714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.273714Z digest=sha256:507d809c990a42e3a48f02ed929cb545806ca330f68ed369438ec020c3043728

Observation d1c44e37-5f68-4494-8618-560ad26d944d · outbound

This paper cites Representation Noising: A Defence Mechanism Against Harmful Finetuning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Representation Noising: A Defence Mechanism Against Harmful Finetuning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.278636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.278636Z digest=sha256:e8bf1e1e81a8bd3102e73b18945c28e788dff10d0925f212d01b9c877686e775

Observation dc0996cd-9a84-41e8-95b7-8e63d29f3799 · outbound

This paper cites Fast Adversarial Attacks on Language Models In One GPU Minute.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Fast Adversarial Attacks on Language Models In One GPU Minute

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.283768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.283768Z digest=sha256:16bb756d4cf8a607f51ee9f4b200d41f4180e4a5532e75f4e527ff3e6f8c9f90

Observation 494f9dc6-20ee-48f4-825c-e56e317b936b · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.749567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.288872Z digest=sha256:d53682ed0d86e03c1819d62ab3b9227902e8268c77910dd7552ee8b21f0be88c

Observation 77db4e86-0d51-49ef-b89a-be27da8b1815 · outbound

This paper cites Towards best practices in AGI safety and governance: A survey of expert opinion.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Towards best practices in AGI safety and governance: A survey of expert opinion

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.293676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.293676Z digest=sha256:951975789630c7108ca631fe304a40f5fa75858aeff6b8ede24b0826bfbf6973

Observation 6ccc3f80-0418-4903-a57f-661b4c4fc24c · outbound

This paper cites Adversarial attacks and defenses in large language models: Old and new threats.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Adversarial attacks and defenses in large language models: Old and new threats

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.730884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.299109Z digest=sha256:e9545d8c807bd333c04e815fc0bd405f5826ccad131dd15fb15f530bc7ff9474

Observation ad40a32b-3f19-4bfc-b54c-e45792a8415a · outbound

This paper cites Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.304017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.304017Z digest=sha256:eba4949833123afc4bd8354bced858eb3bace7b38a0a0afab0bee706a3015489

Observation af0cb92b-f7bd-42ad-9c34-e3a1ece7f36f · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.309285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.309285Z digest=sha256:b80dc819d1b87eb96c233debb57844e320aa229a3b4a897d5a276ee5768522f3

Observation 5b9b18c3-6020-455f-9cdc-8ab91d420ad9 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.314358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.314358Z digest=sha256:1b0ad1e80baf1c4253a470a0c746e5482546f5b88a19edb582b1e17ff9e95706

Observation e882aafa-ceab-4c62-9b51-c96884513ba0 · outbound

This paper cites Model evaluation for extreme risks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Model evaluation for extreme risks

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.319250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.319250Z digest=sha256:444ad04ad48900d63c59845988412aaddf858ca923bc03b9881ad39a3bacfe2d

Observation 04b13167-b951-403d-9833-4c2ac64427dc · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.324252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.324252Z digest=sha256:951c164d41f31423601e2e6850e2c8fa51b70fc214e0cb642dcc9546104e1b29

Observation 8b1f053a-be81-4969-aa6a-d73c134f95ac · outbound

This paper cites UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.329073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.329073Z digest=sha256:6925c1b0fb5b18e1d777eac467456bc6c7912a9f4a675dbf59bde9a8f273f5e2

Observation f6df5fa0-756f-4281-87e5-1492cc8cc155 · outbound

This paper cites an unresolved cited work.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:47:35.714557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.334078Z digest=sha256:4debc9639f76de932e16e1896344fe9272e991ff290c9989f956b324fbb29a98

Observation 3b01b132-ce61-46a2-aedf-4dcea52fdc81 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A StrongREJECT for Empty Jailbreaks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.338954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.338954Z digest=sha256:90f74c3caba185df38f423d9a389401a9023dc6854b445ce7b4df79e2aba5b5f

Observation 1d3ddeba-b616-47bb-bed5-613a2bf254ed · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A Simple and Effective Pruning Approach for Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.343866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.343866Z digest=sha256:21c618c1761b118096287e1df11e5dbe3e622fbfb978ee854fbc0bae2e7e7da7

Observation 9a2486f0-3e1b-4f9f-b395-d26a1835bd4d · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.348837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.348837Z digest=sha256:bf29c798fa4edfc9d83a7c739e8542f1ab5602d25562ca7df3ca0abd68698239

Observation aa727472-dfef-42ec-84b1-83bb5e94f6d2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.353590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.353590Z digest=sha256:a3a3f09f33719b446c8c014d3e9e0a925254bde7b67e368fecdebed4ea864c10

Observation 1dfc3555-196d-425e-833a-f95e4ee86445 · outbound

This paper cites A pro-innovation approach to AI regulation.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A pro-innovation approach to AI regulation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.698068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.358517Z digest=sha256:cc71b05ebcf81117884d254adeb2d6c442168809d6470afe777c3b05fbde447e

Observation 1876fddd-203d-463e-839f-4af8685db792 · outbound

This paper cites AI Sandbagging: Language Models can Strategically Underperform on Evaluations.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.363084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.363084Z digest=sha256:f31fbafacb31654b1a184ef28f47c198f038cbdaf8093b78cf0feea9689ce9fc

Observation dd41a068-d301-48c4-a177-4c4270a48e03 · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.367750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.367750Z digest=sha256:342701fb7ae7b33b167f992fa98d229c88ba0e12e409843e361d089bc487cecb

Observation 0bf8c76c-ecb1-4d01-a8c8-a7a07a00873c · outbound

This paper cites Efficient Adversarial Training in LLMs with Continuous Attacks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Efficient Adversarial Training in LLMs with Continuous Attacks

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.372451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.372451Z digest=sha256:487563ea0171e07008ef2a1262f71c38330aa74d2e9a0590aa4aab4609be8b4c

Observation da6e0ecc-79bb-47cc-80ab-863b8edc129c · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.377658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.377658Z digest=sha256:468e72ab97317cf87bb162549054bc9ae1c55fcbb83f8269f5bf842ceac4bd23

Observation fc9df143-d719-4af9-8b61-8e95f72f1ce9 · outbound

This paper cites On the vulnerability of safety alignment in open-access llms.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities On the vulnerability of safety alignment in open-access llms

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.680996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.382845Z digest=sha256:58162b5239d453ec0da91dc403e77e4fd0b8b72c52789cc87ef1d3158187ceb7

Observation c0b52edb-205f-49c1-879f-be6937fc0b5e · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.387716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.387716Z digest=sha256:05fc067fa09382536cfd85a651394153b4bfb91721596b69262ce97c4fc79bc3

Observation bc50e58e-37ee-492d-bc8c-52347c806179 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Low-Resource Languages Jailbreak GPT-4

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.392539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.392539Z digest=sha256:aec24b51da798bf1f830d722debdc4f773bdfeebad3d13cf5347f69bb52c09b5

Observation aefc68dd-f5a7-4218-8fe3-77f7c6f34395 · outbound

This paper cites Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.397309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.397309Z digest=sha256:19644c083e876ff93da3a8a33ddc1753660e2eaf43748602eed3dbe76bac9fd8

Observation ad4eb8ad-9be3-49ed-bbd1-b106c3b322d5 · outbound

This paper cites BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.402376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.402376Z digest=sha256:d8ad6694f1572c002db3150f1f94337a0a0fd21523df4ad39597ab1bde1099f5

Observation 75cfd0a4-16c1-4066-8ef6-5fb86992c319 · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.407456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.407456Z digest=sha256:6260d1c3c1228f622d442b25b8d78bc1403f81961b2ebf3a4f1101fbdfada1bf

Observation f3323a75-f9ad-4ce3-b497-b1bad5992b69 · outbound

This paper cites Adversarial machine learning in latent representations of neural networks.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Adversarial machine learning in latent representations of neural networks

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-08-09T14:47:24.387209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.412546Z digest=sha256:6cde963be6817521d3848976e6ac422a1260a9359ffd0a85286bcca376c593ab

Observation 73905ea9-a758-443f-8657-588ec377ee6c · outbound

This paper cites Catastrophic Failure of LLM Unlearning via Quantization.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Catastrophic Failure of LLM Unlearning via Quantization

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.422836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.422836Z digest=sha256:595a2f89459ad05d5f1e923288280ddc4dc885764dd3409d3014998791b6ff4c

Observation c52482b9-2f28-4048-80f8-85433775e0de · outbound

This paper cites A Survey of Recent Backdoor Attacks and Defenses in Large Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities A Survey of Recent Backdoor Attacks and Defenses in Large Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.427626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.427626Z digest=sha256:e4b5e07e008a7213bcd017513e300426011e31d96bfb1303178e5791be09d78d

Observation 0f86f630-1d5a-440f-a596-014d8a74a2cc · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.432963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.432963Z digest=sha256:e25c0087a1d29db573552d0035d6dd39cb65c5f4a3ff1273b8ed4fa494980e22

Observation d94a0715-f578-490c-8a47-ec5e68464ee1 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Representation Engineering: A Top-Down Approach to AI Transparency

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.438083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.438083Z digest=sha256:e7e9ef4ced737e6aa476b36d66e568be73f4e5f7d3997156bc16e5f35a1ed3b7

Observation c0a5011b-8beb-48f3-b3fd-c6ba8f62a5b6 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.442928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.442928Z digest=sha256:3fa5c5a622cf0ae6450c8fc9753a47efb1ecf8d88634be2114d8df1b288caae4

Observation 73f8f072-4083-4492-8bce-0cafd19c6284 · outbound

This paper cites Improving alignment and robustness with circuit breakers.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Improving alignment and robustness with circuit breakers

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:47:35.664850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T14:47:15.448064Z digest=sha256:23ff715afb111fa2d39df1f0de1726f179ca2f6d35b4af5a8da36c56d0112a60

Pith citing papers

Observation 9ac562e3-c5b1-46c0-8e04-7102890168a9 · inbound

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods cites this paper.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.683217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.683217Z digest=sha256:0126899ce70bd8113f86e385ef411339f193fbffd5d21d31672d6b31ed5f977f

Observation 5b691c14-d0ac-4408-84b3-1a983d3c9287 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.212266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.212266Z digest=sha256:a40f976f744709806db7eaff955b0a123452d530ee7146b0f496d2a6771fbc3b

Observation d4017d78-8c14-4e7f-a732-46170f55fb0a · inbound

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs cites this paper.

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:25:35.442838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:25:35.442838Z digest=sha256:7c2b5597d57c1b914da5c06f4a7ac77a483462af771c43cce7ca53956b88e119

Observation 0abc5024-ef5d-4655-95da-9c9b68ade291 · inbound

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning cites this paper.

Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:46:17.058220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T10:44:53.516653Z digest=sha256:3fd49e63d0f79b93187115e5da895ffcb8ca20cae37ade7770b19f3c7b72f203

Observation 9195ab00-e824-41a4-9ef9-dd9eb8448c61 · inbound

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories cites this paper.

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T18:41:33.209039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:41:33.209039Z digest=sha256:541b82ba96a55d6c3b7d9316532ced7f3abd7039bdebaad846496fa4c4f7a7be

Observation 5ba3b8fb-a6c3-43a7-8dbe-1d7aa8abb55a · inbound

Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns cites this paper.

Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T19:59:01.358929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:59:01.358929Z digest=sha256:b22abe4c8fd7a0f91f9414e8d6e6589da2305397dfebe05b09ab718d75b83e5b

Observation 735466bb-16e9-47f0-ad25-d74394d5c890 · inbound

Operationalising the Superficial Alignment Hypothesis via Task Complexity cites this paper.

Operationalising the Superficial Alignment Hypothesis via Task Complexity Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T22:49:13.999922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:49:13.999922Z digest=sha256:0aeeb49bf5808e39081a8fcf77abc338bb3ba61266cf154f094b92f768ccc8e2

Observation 9f63ce60-b49b-432b-94be-3091cd96bea2 · inbound

An Independent Safety Evaluation of Kimi K2.5 cites this paper.

An Independent Safety Evaluation of Kimi K2.5 Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:43:11.577719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T19:38:18.674355Z digest=sha256:1c7a1020c661f339b6512b0f55aa2d88405399fc66b0bbd20b269790bb06837d

Observation ec539424-962b-4ec3-b385-e3a4912b5b1d · inbound

Is your algorithm unlearning or untraining? cites this paper.

Is your algorithm unlearning or untraining? Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:59.104808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:06:08.962042Z digest=sha256:b9216a1bb00eb582fc8eddaa815bf3c7e04b426c069939d87e251b77864985b1