Pith. sign in

Paper Citation Record · LEDGER

Universal Jailbreak Backdoors from Poisoned Human Feedback

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2311.14455.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.14455 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:12.430061Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T02:07:33.632292Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 843f8e52-cca2-4732-98aa-63e2f3df7bd8 · inbound

Model-Editing-Based Jailbreak against Safety-aligned Large Language Models cites this paper.

Model-Editing-Based Jailbreak against Safety-aligned Large Language Models Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:06.025124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:06.025124Z digest=sha256:ae6772ba35461eaead906b9d6e0c96c31ed4190cfaa4d806e845ae3affaf059f

Observation ea4c7205-f1b4-4f54-b765-64f8e0121f7e · inbound

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning cites this paper.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.818464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.818464Z digest=sha256:a0c80c60d50bf336b0b18442caf05df146689e6244c65a6bdd28586e7a793d75

Observation 8904d5fa-aaf9-4f38-b061-701e550acba6 · inbound

Trojan Detection Through Pattern Recognition for Large Language Models cites this paper.

Trojan Detection Through Pattern Recognition for Large Language Models Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T18:07:46.877650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:07:46.877650Z digest=sha256:505bc424cc39f4a8b92107d61acf9a5c77982e2233483679f494da24acd65456

Observation 7e0e3eb9-cf42-431b-8978-7f499194f1a5 · inbound

Complete Chess Games Enable LLM Become A Chess Master cites this paper.

Complete Chess Games Enable LLM Become A Chess Master Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:32.870781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:21:32.870781Z digest=sha256:dc9bd7d6f436aca15d5f655b04f55693b916175af6bc3013b5a4f784100a6411

Observation a702fc4c-b6c0-4577-8df6-dde358651754 · inbound

Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription cites this paper.

Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-09T11:34:09.559897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:34:09.559897Z digest=sha256:44bcb305b381e8e4054731853311cf1f93ff066540c3214f2f07dc24c6da8ac5

Observation 800c2181-654a-4592-b3d5-fa4f73d67cc3 · inbound

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations cites this paper.

A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-09T00:50:00.364674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:50:00.364674Z digest=sha256:14be5cfebced63d1bb1830f3aa8c24f8fd454d8838ff7be194dd6061cf680842

Observation a8abec62-2281-4b3a-b1e6-49930318727c · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:12.430061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:12.430061Z digest=sha256:094fe965db7a46ea8cde723ee843290efea12209056fed1e08ce8696f456d5a8

Observation 352d621b-763a-427f-b92e-f1f81a82f522 · inbound

Inducing Vulnerable Code Generation in LLM Coding Assistants cites this paper.

Inducing Vulnerable Code Generation in LLM Coding Assistants Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:20:36.423021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:20:36.423021Z digest=sha256:6d9be20338c85ef52567cdd9aa5170b2f4938704158ccc1d6fd1daf5841764f7

Observation 731fc1b5-9132-4d42-8c8d-061e18962ee0 · inbound

BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts cites this paper.

BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:51.836074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:38:51.836074Z digest=sha256:3832812d2f19e6c88abfac1324c0d14f6855ae0763445606d77145476cbfab95

Observation 055ca13f-a865-4e8b-930a-72311221dd5f · inbound

ACE: A Security Architecture for LLM-Integrated App Systems cites this paper.

ACE: A Security Architecture for LLM-Integrated App Systems Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:06:54.290344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T18:05:50.363944Z digest=sha256:56d9bc6b3a7e689f215906bc71374b60415ae64649e1ee58a908c36fa5fee226

Observation 33bf3e96-e499-4ef8-9020-f5c5256b1d61 · inbound

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors cites this paper.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:33.389364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:33.389364Z digest=sha256:d0d0fa2ad2aabd40fefddb75f51cf696fc36c325bb49c52150c195c999348f42

Observation b6007439-ac16-4601-bf79-aefa8c707234 · inbound

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF cites this paper.

BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:04.647947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:18:04.647947Z digest=sha256:79afcb854ec101532a0ae2f60953f2cc2fe8b140c7b01a87c42588982521694f

Observation 2f3f7c0d-3860-409d-986d-450d99d89b3c · inbound

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users cites this paper.

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:02:07.883772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T05:58:17.452837Z digest=sha256:b17c32c427e5b174e27fd504f018d7059a09ec82d6a6159343986d12662ed157

Observation a3edde7a-2b09-4096-9e4b-c3d793d59ec3 · inbound

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation cites this paper.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.702267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.702267Z digest=sha256:ed73d659a72b5f0f67fde5f47e02680e8c7dfc38f178c12019d3a0be79bce6c3

Observation 7cd51222-9dc3-4c30-9a49-c03d5c2a14bb · inbound

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? cites this paper.

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:48.881724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:00:48.881724Z digest=sha256:b682f68e1a2bdc440b19eab0b5a6be40c604b81d675de31b212a30e4fdd50424

Observation e34afee9-ddf0-49b7-b96f-20acfe8a15b2 · inbound

A Security Analysis of the OpenClaw AI Agent Framework cites this paper.

A Security Analysis of the OpenClaw AI Agent Framework Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:03.421451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T22:22:01.349069Z digest=sha256:9707422b1496ca5f7edbe5b8a56126419eb0af00c17191febdab4c6ca73586c4

Observation 44cd1d8b-909c-469c-a338-a811b1ccf740 · inbound

A Security Analysis of the OpenClaw AI Agent Framework cites this paper.

A Security Analysis of the OpenClaw AI Agent Framework Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:49:50.175932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T06:45:24.663069Z digest=sha256:e69c9d517e783151f484a1b03c329e8de5c95bd8860e5c4b6bfe356e6a9939d4

Observation 447dbc6e-7ee6-4acc-a18d-34d8fdd1ea07 · inbound

Efficient Preference Poisoning Attack on Offline RLHF cites this paper.

Efficient Preference Poisoning Attack on Offline RLHF Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:50:27.068723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T19:29:25.000361Z digest=sha256:dbcadec039555611677d05cca689178cfa3b6c669b99ed6d2c6d3a1bb50b3422

Observation f351d1fe-56f8-4364-b528-5dd31006fb63 · inbound

BadDLM: Backdooring Diffusion Language Models with Diverse Targets cites this paper.

BadDLM: Backdooring Diffusion Language Models with Diverse Targets Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:11:25.953285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T04:30:13.417357Z digest=sha256:c08cf261c6331b449eebe64e7e87e83f7389277fb7bbca3848ee54d4309fccae

Observation 8cda201b-f576-49ee-961d-c1f32dc7a68a · inbound

Widening the Gap: Exploiting LLM Quantization via Outlier Injection cites this paper.

Widening the Gap: Exploiting LLM Quantization via Outlier Injection Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.628969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T20:58:45.542409Z digest=sha256:094ef6fb58a9586844bf651489ae18a2ba34d7ccfdad2255d2b99c7733a0f9e2

Observation 2489d6a4-7dd4-4708-8615-bc63b45180e8 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.163002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:923e58dc6057b7d07bda035da0ed1f5b72200ca61d4d3645fa81ae80da6a25fe

Observation 378a5fb9-9805-4319-b74a-f759f5dbc1a0 · inbound

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI cites this paper.

From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:08:50.603159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T18:08:24.901025Z digest=sha256:cb5764c46a2badd00d7e026e9e1031df4e05ed00725b7284f1115adbe3b4bd0d

Observation 04ad651b-896f-407f-b21a-8938c3e38c47 · inbound

Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs cites this paper.

Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:07:33.634083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T16:10:18.569412Z digest=sha256:eddd3e5016109b6f87009af8089013fbaa42b26d9bb8e9c0ca754ec0eef67302