Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:46:14.886511Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2602.06911.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:46:14.886511Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6b3c9b1a-0ab6-4c15-a931-653918c625e9 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering (2024b) and He et al
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fca2ddee-d6b6-4785-bc9d-149889d497c6 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff9a125-00ba-468c-82d0-26bad94a0713 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Toward Secure Tuning: Mitigating Security Risks from Instruction Fine-Tuning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4568aedc-7d58-427e-81da-1c1653284209 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Luxi He, Mengzhou Xia, and Peter Henderson
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2abbcde9-9b3d-4e54-821c-3b459b676b2f · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering URL https://openreview.net/forum?id= dp24p8i8Cg
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608c4094-fe7d-42db-91c4-eba22db57296 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering ISBN 9798400702310
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27f32768-a236-43d8-87c2-5e2116809be0 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2ab5f4-48a5-42c5-940d-b15044b12f3a · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb19529a-d2d8-4a57-9bf5-95359a4f35ba · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Kyle O’Brien, Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, Ishan Mishra, Geoffrey Irving, Yarin Gal, and Stella Biderman
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf606279-0b01-4581-b299-aef4377bcd57 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608a7c66-ce8d-4b04-83d4-bfae2e8e242b · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering GPT-4 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7013b613-6c98-4212-b40b-295c0e22d8b8 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering ISBN 979-8-89176-195-7
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8193f5d1-1bd8-44cb-a112-509f52876962 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fdc5205-36e5-4029-b521-a08d041a72e9 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 635d0d1f-45f5-4ec1-8ad7-6709e7139deb · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Jailbroken: How Does LLM Safety Training Fail?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f31a21da-f524-4d98-b1b9-ad927f168a17 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Instruction-Following Evaluation for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b65f4ac7-03a7-42dc-ae07-5bddd6793393 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Representation Engineering: A Top-Down Approach to AI Transparency
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f15d922c-090f-4bc9-b23a-810b20f1fdfb · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering In contrast, MATH is only loosely aligned, reflecting its narrower domain and strict exact-match scoring
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6b62e94-65d5-4550-83c5-7ea3324b2509 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Danny Halawi, Alexander Wei, Eric Wallace, Tony Wang, Nika Haghtalab, and Jacob Steinhardt
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e417595e-ea0f-4b95-bf7e-fc2ac3fc5515 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Program Synthesis with Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5f71dea-bc28-4e6c-b53a-e7ab0e8ac61e · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering The current year is 2025
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c7b69ad-4a16-441c-9139-d0d94dad80b3 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Eternal Sunshine of the Spotless Net: Selec- tive Forgetting in Deep Networks
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79061286-4169-4c8f-8f42-9b3413502b80 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering URL https://openreview.net/forum?id=urjPCYZt0I
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e96d75aa-c237-4a1d-a88c-2607b78ab112 · outbound
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.