Pith. sign in

Paper Citation Record · LEDGER

Removing RLHF Protections in GPT-4 via Fine-Tuning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2311.05553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.05553 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:15:32.929824Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:39:40.653733Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d3a1fce6-3ae4-4331-beba-8a22e0d1e594 · inbound

A StrongREJECT for Empty Jailbreaks cites this paper.

A StrongREJECT for Empty Jailbreaks Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:02.892798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T21:28:02.745230Z digest=sha256:b5157a7afa8d085cddd4fcedf4674db2f1e011d8698c198137eabfdc27c7183b

Observation d438dfad-e3f6-46b6-b417-7baa05af7aa5 · inbound

LLM Agents can Autonomously Exploit One-day Vulnerabilities cites this paper.

LLM Agents can Autonomously Exploit One-day Vulnerabilities Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:18:27.668214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T04:18:27.597704Z digest=sha256:be3f75a823d3d3c7a6cde55ba3e172f6629395ea570333a5e2992a2c1b430f64

Observation 08b13389-b693-485c-9a26-527420eaa3d1 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 205

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.171172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:926fe4f7bcc18e2c949bb9124b35bb8667f3afb17fd00d99e269529e14207549

Observation d76278a6-65bb-4712-b4da-f9c535c79cde · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.473852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:521897db14edc27e5ece50f7a61bc6a5beacafc6307de800c471afceeb2db265

Observation da4add40-2fa4-4b49-8281-4abcaf742df1 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.504406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:d1e7d126b4f3a16b0275cd938eb5de19acac387bad55e54e896d89c9960d2afe

Observation c455825a-9c42-49be-b81d-cdd36d81e952 · inbound

Safety Alignment Depth in Large Language Models: A Markov Chain Perspective cites this paper.

Safety Alignment Depth in Large Language Models: A Markov Chain Perspective Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T18:15:32.929824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:15:32.929824Z digest=sha256:419fcfb29ca861066fcda64d813ebc9c4f8e42737af306227c548947b5ef06fe

Observation 75cfd0a4-16c1-4066-8ef6-5fb86992c319 · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.407456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.407456Z digest=sha256:cd4b4014219c15e79ba509c26e0c472dc99585f905290a5509f5511a9c021e79

Observation 382b390f-a3a8-4031-8574-31eb66cb9f27 · inbound

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning cites this paper.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.538591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.538591Z digest=sha256:b76248a626ab09407449f1111fc84a3c59aceaae1d929fc7370bf4fec6f34551

Observation ab1d7d35-a32d-4fb2-9219-90097f0a8d86 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:57.988075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:57.988075Z digest=sha256:f3ed8cf28d77e4a781173f0e13ab81d5903a3120ce4ae4097e2844813dbbbfc0

Observation e2fffb31-5954-4016-ac38-f43310662a1b · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.407857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.407857Z digest=sha256:6df05d4ab9f7cd46bfd0d3f7f19049002a205ae383ea3e473f7664eeee1c6d03

Observation d59c73aa-0579-466e-b148-13041ea73cf9 · inbound

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner cites this paper.

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T18:12:37.016328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:12:37.016328Z digest=sha256:4988fb05cb9e27c171ba3e83def5e14fcb83eae028f7780e40e505083393ceaf

Observation cee53ca1-667e-4e37-bb0f-bb69aa9df89f · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.962460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.962460Z digest=sha256:6de448921917b834c54001087a51b04e44f4b4660f5bb936f4411b6a9f01662d

Observation 5d1f1798-edec-42e1-96a8-a70f3a11a1e9 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.306994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.306994Z digest=sha256:9767cc92ea9edd8da5816a4d59fd93dd5c5c2c3ed11893c33b17e83f8191efaf

Observation 60840032-8432-4455-8d31-8136cb96982e · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.530996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.530996Z digest=sha256:7fa2037779446ec9534a063af21e71a775f66249e045685df7be1126c477d926

Observation d4d2b640-0807-482c-98e5-f9235fcac7af · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:28.837716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:28.837716Z digest=sha256:8e7501ef3e57012977a4955039fca551687b9d0d53ac5554c33332c74d5f7aeb

Observation 4f08c312-b145-4507-a47f-db9af96708e9 · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.223558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:45f579bf14726dce334dcaa9a4d0eb2d09ff19c735a8b0d02722e8d33fe665dd

Observation c9d07f83-bfc7-44a4-aacb-e4d1a4a6dac3 · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:44:37.742433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T10:41:14.970156Z digest=sha256:34b9bc7f28feddec22a8efa5998f37de0ebdb8fedf4c9761c8deb798fe73b0d4

Observation 192a1bc1-ec29-49b7-bb81-ab010516d622 · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T16:13:13.877328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:13:13.877328Z digest=sha256:d0241d27b608a7b59d7c32ed2492c681750b369719e6330fdc03da09d0399f8f

Observation b7e7202d-0fa1-44cd-a743-f8d94550050b · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.404876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:849eb64ff8060b3089d7871104ff8501a15fabc11271cd3da367f44b0e90f748

Observation 023e10d7-9b08-4f5e-b37d-f7027db4232a · inbound

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation cites this paper.

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:08.458732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:02:20.836208Z digest=sha256:74ab3df45e5a1d024f712285eeefa82e57ab1d3ff4c691b53eeff339a372cc83

Observation 1cf88bd4-e26a-4ec8-ae27-9674b19c91ac · inbound

Open Weight AI Models Require Proportional Evaluation Approaches cites this paper.

Open Weight AI Models Require Proportional Evaluation Approaches Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:39:40.655924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T15:39:45.969260Z digest=sha256:de1a0bef0e07aa47bb86f9af3c673c44dce3e2a093d589cda85d55c415be37cd

Observation 0bc464dc-398e-45a6-a3c0-388d06036cc7 · inbound

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift cites this paper.

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:12.338931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:12.338931Z digest=sha256:ebf8c8194c56d5a81d1a224cc6800cb363e7746bf877f7c611db294efa7387c9

Observation 2c0ce0e2-6c99-4603-94ce-9450a6693d79 · inbound

AI Security Priorities: A Field-Wide Agenda cites this paper.

AI Security Priorities: A Field-Wide Agenda Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T10:22:17.273139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:22:17.273139Z digest=sha256:8e705668b4e61c8a342996f594e8caaf8ca718481525f08ff406a0ed204aa14a