Pith. sign in

Paper Citation Record · LEDGER

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2602.06911.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.06911 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:46:14.886511Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6b3c9b1a-0ab6-4c15-a931-653918c625e9 · outbound

This paper cites (2024b) and He et al.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering (2024b) and He et al

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.881989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.881989Z digest=sha256:006542e60f701dfefa4462f6864cb1f6cf1654b7b266b3ca2c42c29e96145108

Observation fca2ddee-d6b6-4785-bc9d-149889d497c6 · outbound

This paper cites an unresolved cited work.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.886511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.886511Z digest=sha256:4297d4dd1bf821b6f8400a9778ca04038169fff4626e698f00a9cb1891203e65

Observation cff9a125-00ba-468c-82d0-26bad94a0713 · outbound

This paper cites Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.790420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.790420Z digest=sha256:2e15de87927df7535c7ab51edcf1e1105511f46e2f37fe086633657bdbfe163e

Observation 4568aedc-7d58-427e-81da-1c1653284209 · outbound

This paper cites Luxi He, Mengzhou Xia, and Peter Henderson.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Luxi He, Mengzhou Xia, and Peter Henderson

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.804852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.804852Z digest=sha256:bdb5ea17c56d558f13637ab0679df72a972f0bd26c1cde89df3abf54a52ced57

Observation 2abbcde9-9b3d-4e54-821c-3b459b676b2f · outbound

This paper cites URL https://openreview.net/forum?id= dp24p8i8Cg.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering URL https://openreview.net/forum?id= dp24p8i8Cg

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.809451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.809451Z digest=sha256:517a78d37b6ca399fb49d53fde083b941db669b104ea8b6774e861c9461d8a66

Observation 608c4094-fe7d-42db-91c4-eba22db57296 · outbound

This paper cites ISBN 9798400702310.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering ISBN 9798400702310

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.813546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.813546Z digest=sha256:f1762176bf4400d6fbf1914e895f2e42b803d5e48617bedae18a9086311b0394

Observation 27f32768-a236-43d8-87c2-5e2116809be0 · outbound

This paper cites Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.818615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.818615Z digest=sha256:789d7dfc88bba59ef27cf006cc2ed03322964a605cead2860cef36ba5b5ceb03

Observation 2b2ab5f4-48a5-42c5-940d-b15044b12f3a · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.823403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.823403Z digest=sha256:11a09372f895786328cb0a8cae5bfb948b07ad5c6b16c1bde605c55b64ef0d65

Observation cb19529a-d2d8-4a57-9bf5-95359a4f35ba · outbound

This paper cites Kyle O’Brien, Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, Ishan Mishra, Geoffrey Irving, Yarin Gal, and Stella Biderman.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Kyle O’Brien, Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, Ishan Mishra, Geoffrey Irving, Yarin Gal, and Stella Biderman

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.828547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.828547Z digest=sha256:5ca2cd2ea0e6f4cf094bf0f2a567c17abe991cc9c04e4f1d788c04924b32aeeb

Observation cf606279-0b01-4581-b299-aef4377bcd57 · outbound

This paper cites an unresolved cited work.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.833451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.833451Z digest=sha256:785fe1271d6291c56b03bd477e2f6f1dea9e7ba07c97d162e9bbd9802e8e4c17

Observation 608a7c66-ce8d-4b04-83d4-bfae2e8e242b · outbound

This paper cites GPT-4 Technical Report.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering GPT-4 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.838787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.838787Z digest=sha256:087ccae2f073fc70d974202bfc9c54a13f313ae1f982020bad913de0b0005d1d

Observation 7013b613-6c98-4212-b40b-295c0e22d8b8 · outbound

This paper cites ISBN 979-8-89176-195-7.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering ISBN 979-8-89176-195-7

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.843583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.843583Z digest=sha256:39423f09f76a2140815e939be6bfd03f4772e1a20798620473766a292f7e6991

Observation 8193f5d1-1bd8-44cb-a112-509f52876962 · outbound

This paper cites Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.848303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.848303Z digest=sha256:bca57609a265e766606d1ba18709f9f035aec4d554dc4223fc42e865918a47dc

Observation 2fdc5205-36e5-4029-b521-a08d041a72e9 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.852644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.852644Z digest=sha256:e2757e7a099dd6bdb6659695518b683b9b648ab52c532b99639413a3238f5ebd

Observation 635d0d1f-45f5-4ec1-8ad7-6709e7139deb · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Jailbroken: How Does LLM Safety Training Fail?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.857511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.857511Z digest=sha256:1d6210fbe0775396ecd77bef6577eada3c5c6dafd0623fb2f81e70d671b8c6c9

Observation f31a21da-f524-4d98-b1b9-ad927f168a17 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Instruction-Following Evaluation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.862258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.862258Z digest=sha256:4c4d02c02f764a7fea771610f0504344c7b3838a35f9f4786b14e90502978672

Observation b65f4ac7-03a7-42dc-ae07-5bddd6793393 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Representation Engineering: A Top-Down Approach to AI Transparency

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.866706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.866706Z digest=sha256:d45b6cb349d8967882e8c7ae4ddb8f92516ff5428049fa8866dfbb84a213f046

Observation f15d922c-090f-4bc9-b23a-810b20f1fdfb · outbound

This paper cites In contrast, MATH is only loosely aligned, reflecting its narrower domain and strict exact-match scoring.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering In contrast, MATH is only loosely aligned, reflecting its narrower domain and strict exact-match scoring

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.871822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.871822Z digest=sha256:a5ba66011a9605ba6db07429559977691666c000a6d6219aa1bb4e7818809897

Observation c6b62e94-65d5-4550-83c5-7ea3324b2509 · outbound

This paper cites Danny Halawi, Alexander Wei, Eric Wallace, Tony Wang, Nika Haghtalab, and Jacob Steinhardt.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Danny Halawi, Alexander Wei, Eric Wallace, Tony Wang, Nika Haghtalab, and Jacob Steinhardt

Reference 23

Resolution
verified exact
doi, observed 2026-08-03T03:49:47.372290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-03T03:46:14.800549Z digest=sha256:d230121f902906f4c981ad9504666384079e1c60b142d5dccf7f438d57f2a2a7

Observation e417595e-ea0f-4b95-bf7e-fc2ac3fc5515 · outbound

This paper cites Program Synthesis with Large Language Models.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Program Synthesis with Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.775654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.775654Z digest=sha256:73891c332eb7a17014a5853ebfe0e57947a88b9266be46edb47514165efb1720

Observation b5f71dea-bc28-4e6c-b53a-e7ab0e8ac61e · outbound

This paper cites The current year is 2025.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering The current year is 2025

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.877674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.877674Z digest=sha256:ff6c21187770481f0913e14a53a08e7c9a6b4672375954dfb3fa76bdd6b5d7e9

Observation 5c7b69ad-4a16-441c-9139-d0d94dad80b3 · outbound

This paper cites Eternal Sunshine of the Spotless Net: Selec- tive Forgetting in Deep Networks.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Eternal Sunshine of the Spotless Net: Selec- tive Forgetting in Deep Networks

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.796043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.796043Z digest=sha256:8d2d338abba4121f4f30d562d247b3ca5abed51eb5714246750b80c367891c0a

Observation 79061286-4169-4c8f-8f42-9b3413502b80 · outbound

This paper cites URL https://openreview.net/forum?id=urjPCYZt0I.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering URL https://openreview.net/forum?id=urjPCYZt0I

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.785957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.785957Z digest=sha256:acfbd79c07ec154b6a16e180bc55936772fc59b9b3260410985ecabb298976be

Observation e96d75aa-c237-4a1d-a88c-2607b78ab112 · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:46:14.781234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:46:14.781234Z digest=sha256:8392c975d96930f4638512b4b4f3e409628bc9186bcfea150a6e44ed80906b60

Pith citing papers

No inbound Pith citation observations are available.