Pith. sign in

Paper Citation Record · LEDGER

Removing RLHF Protections in GPT-4 via Fine-Tuning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2311.05553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.05553 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:58:32.538591Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:39:40.653733Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d3a1fce6-3ae4-4331-beba-8a22e0d1e594 · inbound

A StrongREJECT for Empty Jailbreaks cites this paper.

A StrongREJECT for Empty Jailbreaks Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:02.892798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T21:28:02.745230Z digest=sha256:96f88db1d7a9f694039a7242a3b58aa7fa70d3aa2e01616ea85faa29a3e38505

Observation d438dfad-e3f6-46b6-b417-7baa05af7aa5 · inbound

LLM Agents can Autonomously Exploit One-day Vulnerabilities cites this paper.

LLM Agents can Autonomously Exploit One-day Vulnerabilities Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:18:27.668214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T04:18:27.597704Z digest=sha256:125582ca7d4243125c8ad9e73ef4c02bc99443278b1098907d85f141ac287466

Observation 08b13389-b693-485c-9a26-527420eaa3d1 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 205

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.171172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:ed64d138ee3f86c4a63ab0ea8c81fa48d4c7e621f3d31a04c9a4f58ad6b6f73f

Observation d76278a6-65bb-4712-b4da-f9c535c79cde · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.473852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:316a3dfa52773c934cbffbedadc5462f68a32d88dd13666547007d1bd20a2085

Observation da4add40-2fa4-4b49-8281-4abcaf742df1 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.504406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:0a5445f535618e3862116aaba7509dc207e4412070ad70ac7e4b7e09130fad37

Observation 382b390f-a3a8-4031-8574-31eb66cb9f27 · inbound

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning cites this paper.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.538591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.538591Z digest=sha256:ef1a9c68df54ebd62df7da91e32678b91d7dee03f6f44c86c53da8b9e0dd8ce6

Observation ab1d7d35-a32d-4fb2-9219-90097f0a8d86 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:57.988075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:57.988075Z digest=sha256:d9513858eba2a819578680ece93c49a01a36c8f5ea85331da765d8b0bd097c5b

Observation e2fffb31-5954-4016-ac38-f43310662a1b · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.407857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.407857Z digest=sha256:6df05d4ab9f7cd46bfd0d3f7f19049002a205ae383ea3e473f7664eeee1c6d03

Observation d59c73aa-0579-466e-b148-13041ea73cf9 · inbound

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner cites this paper.

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T18:12:37.016328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:12:37.016328Z digest=sha256:589deb82d492ca3bfe4e9f87735627b04ab0043c8c5dc827e5da17b8ce8cafb4

Observation cee53ca1-667e-4e37-bb0f-bb69aa9df89f · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.962460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.962460Z digest=sha256:6de448921917b834c54001087a51b04e44f4b4660f5bb936f4411b6a9f01662d

Observation 5d1f1798-edec-42e1-96a8-a70f3a11a1e9 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.306994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.306994Z digest=sha256:9767cc92ea9edd8da5816a4d59fd93dd5c5c2c3ed11893c33b17e83f8191efaf

Observation 60840032-8432-4455-8d31-8136cb96982e · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.530996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.530996Z digest=sha256:30b711e6c794b3bb8359b42c8d5e1981e50d93bf0493b7c3f23d9ebd58edffda

Observation d4d2b640-0807-482c-98e5-f9235fcac7af · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:28.837716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:28.837716Z digest=sha256:9c93ba2a34e66ca64a486af89fca06bfa467a9ab93393a5df44d1a8ba82eedf9

Observation 4f08c312-b145-4507-a47f-db9af96708e9 · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.223558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:2eb14a5b44be60bf1b1fd09e85db6869a5d6f1264e7baa3fe444ea8e1f6a9ff0

Observation c9d07f83-bfc7-44a4-aacb-e4d1a4a6dac3 · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:44:37.742433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T10:41:14.970156Z digest=sha256:e633a404d06bcfe69678e79e9f8d2b885c12f74a101693ac94879d4c688eae84

Observation 192a1bc1-ec29-49b7-bb81-ab010516d622 · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T16:13:13.877328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:13:13.877328Z digest=sha256:d0241d27b608a7b59d7c32ed2492c681750b369719e6330fdc03da09d0399f8f

Observation b7e7202d-0fa1-44cd-a743-f8d94550050b · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.404876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:c95371de99c4f28dd81050fc9aa65250e45dda4a071e32e77929b043bb03c7b2

Observation 023e10d7-9b08-4f5e-b37d-f7027db4232a · inbound

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation cites this paper.

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:08.458732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:02:20.836208Z digest=sha256:b8c49d56b3cb19dfa833a51130310037af503fd7164abc8c5514bb14b6e0be65

Observation 1cf88bd4-e26a-4ec8-ae27-9674b19c91ac · inbound

Open Weight AI Models Require Proportional Evaluation Approaches cites this paper.

Open Weight AI Models Require Proportional Evaluation Approaches Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:39:40.655924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T15:39:45.969260Z digest=sha256:850f68d2384a7bc5ccaccb570d86e1f664b9c03a1adf07a0dbb69bc31340e6f4

Observation 0bc464dc-398e-45a6-a3c0-388d06036cc7 · inbound

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift cites this paper.

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:12.338931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:12.338931Z digest=sha256:513febe6c578965536e5572c4a41850ba217aa08caa6bdbd2fb95e751e3d4a29

Observation 2c0ce0e2-6c99-4603-94ce-9450a6693d79 · inbound

AI Security Priorities: A Field-Wide Agenda cites this paper.

AI Security Priorities: A Field-Wide Agenda Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T10:22:17.273139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:22:17.273139Z digest=sha256:8e705668b4e61c8a342996f594e8caaf8ca718481525f08ff406a0ed204aa14a