Pith. sign in

Paper Citation Record · LEDGER

Weak-to-Strong Jailbreaking on Large Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2401.17256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.17256 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:34:34.361272Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:29:44.257667Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c586368a-150e-4f55-84a3-29d1350e9d56 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Weak-to-Strong Jailbreaking on Large Language Models

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.504149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:08a78076fff599c286cb59fd4ea0ed4134a798b039012e5bcc7bcd212b6195cd

Observation 96364801-bc8b-4de6-b570-6925a7125ba3 · inbound

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment cites this paper.

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment Weak-to-Strong Jailbreaking on Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T11:03:01.295916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:03:01.295916Z digest=sha256:bc25a407b180c3bd91fe2776c80cd947cb3e61b7fc327d746fc2dbd82198476b

Observation 4e0b6696-8957-4b9f-bc5e-9e306e10adeb · inbound

T-REG: Preference Optimization with Token-Level Reward Regularization cites this paper.

T-REG: Preference Optimization with Token-Level Reward Regularization Weak-to-Strong Jailbreaking on Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:56.453781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:15:56.453781Z digest=sha256:877fc02ff9a49d3007b4a92384c27892e842ef708ec99dd52910bb35eb6a59c5

Observation 89604a44-bd2e-49d3-b80a-00209d22dc3f · inbound

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds cites this paper.

LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds Weak-to-Strong Jailbreaking on Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T20:53:38.160847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:53:38.160847Z digest=sha256:40c911f1dca70397e578f08a7e3d8a5ecfc4ce226e59ae0e8c6778604600b24e

Observation 52690a97-a512-4916-8229-e365c0105b4e · inbound

Model-Editing-Based Jailbreak against Safety-aligned Large Language Models cites this paper.

Model-Editing-Based Jailbreak against Safety-aligned Large Language Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T18:11:05.809308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:11:05.809308Z digest=sha256:0cae2d95c0175e4c16cc48d8d9e36f234b4d754350630751a06d4c7ffd4c10c2

Observation e2337ab0-aac2-423f-9d0e-34710e82a46d · inbound

Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings cites this paper.

Crabs: Consuming Resource via Auto-generation for LLM-DoS Attack under Black-box Settings Weak-to-Strong Jailbreaking on Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:10.942570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:48:10.942570Z digest=sha256:d6efbf7c6dceeec250dd9b3f03d2dd3a9a59fc77c54603fe1e6c5d5e73ec093f

Observation aa3158fe-f488-43ef-95d9-92f6ee752848 · inbound

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models cites this paper.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.726743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.726743Z digest=sha256:ac59d99dca8e4f0bca3c669caaecd1c7ab4bfd42780539c6bda7c90602fe4730

Observation 9a343cc1-97f1-43e0-91af-e32e44dc361a · inbound

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense cites this paper.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Weak-to-Strong Jailbreaking on Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.187801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.187801Z digest=sha256:c039959548f0856c51337c5b56313df1fb432f44dae98cfd5b4bc7b525234a0a

Observation 22481f73-652b-4e7a-abb2-ec51598fee79 · inbound

Peering Behind the Shield: Guardrail Identification in Large Language Models cites this paper.

Peering Behind the Shield: Guardrail Identification in Large Language Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:45:21.416651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T03:45:14.234545Z digest=sha256:d977795afdbece71a006502c12a1bc91740670fb3c71d09f59933719927cc1a9

Observation 2ed9ab0e-3aec-4809-979b-34f74362d5ad · inbound

Confidence Elicitation: A New Attack Vector for Large Language Models cites this paper.

Confidence Elicitation: A New Attack Vector for Large Language Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T22:06:39.024687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:06:39.024687Z digest=sha256:87666224bfc67d9292df4d15fe9eeab4247239a95151dd2cb78a7aa024b16d1b

Observation e17858c5-6305-4ec4-a4cc-c9d78526c34f · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Weak-to-Strong Jailbreaking on Large Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:42:34.219547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:7fa4a62779fcafcccdc9eadc083ca3ed601cbcca25e69ea79dc590a14dfbd017

Observation 65d59232-1a7e-4c6b-8b98-3e67cee7e60b · inbound

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences cites this paper.

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences Weak-to-Strong Jailbreaking on Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T10:23:06.305919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:23:06.305919Z digest=sha256:733aedcdf08326597e0d9731986cc7fb25345f5fed8fe7d88f95739a5f718ad1

Observation ea5fdb87-2e78-4dfb-9370-991952db1837 · inbound

A Survey of Attacks on Large Language Models cites this paper.

A Survey of Attacks on Large Language Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:34.361272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:34.361272Z digest=sha256:cace4350a6d00692313e7a11e298a1bda5e61ad7c5a571660c98b44d3d50528a

Observation 36ca2ae1-9574-45a8-a4bf-cc5df38bc296 · inbound

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion cites this paper.

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion Weak-to-Strong Jailbreaking on Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:15.103581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:15.103581Z digest=sha256:b0e5dc08451199bcf2164c844b77c50b937730af4f90c5212f51596a8d1401b5

Observation eb482b29-547d-4846-bd6b-42e9cf4b000f · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Weak-to-Strong Jailbreaking on Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.759736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.759736Z digest=sha256:71a7a5261e8b6d8c2980b3f35a3cb027d01dae0a3d56306f25d7a6406f4a2ecf

Observation c35c1c48-c59d-4924-9014-51bde1ac83e3 · inbound

SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues cites this paper.

SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues Weak-to-Strong Jailbreaking on Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:47.630473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:04:47.630473Z digest=sha256:73f4383d8043c301c96ec7e33702afe58a692cf79069f9e8f6a3986cdc3ad4b3

Observation 595452c4-6357-4760-8f52-30d432d8a57b · inbound

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures cites this paper.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Weak-to-Strong Jailbreaking on Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.606600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.606600Z digest=sha256:42736d8c3454ad5c8d5f4345ec503a9c9a3789eca1ab2627239912c3743dd5d6

Observation 014a108e-00d5-492c-8dad-fb42b13f5e67 · inbound

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem cites this paper.

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem Weak-to-Strong Jailbreaking on Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:42:12.897176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T08:40:56.186349Z digest=sha256:a9880efd6fd5910c977a21e967d8d19623fc99ded21c2a439586e43be88639fa

Observation 31368c05-277e-4bbf-b0e3-6e2c92dff6f8 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:57.504508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:57.504508Z digest=sha256:6750dd32424b84030e6d0daf8741a1cdd46a1f7d9e4cc2abd0023c470f901144

Observation 2687ee65-1b68-46c1-91c9-905b049d5d40 · inbound

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training cites this paper.

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training Weak-to-Strong Jailbreaking on Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.155732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.155732Z digest=sha256:dce4a82771b97666da917720553ffc938187c5d16919f30dd0548ae4c02aee45

Observation f6b8c244-c238-478b-9765-920e0081a3a9 · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Weak-to-Strong Jailbreaking on Large Language Models

Reference 196

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:25.675116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:25.675116Z digest=sha256:9039f73ec771a769334f09a4488bff4f2d6c5f9cbff7da019471260f72224c51

Observation cd0431ae-c85a-46f5-878f-6dcce83602e6 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Weak-to-Strong Jailbreaking on Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.631918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.631918Z digest=sha256:4ef08509a6d6651319f0e35a73dc8ac917c6c354bf181ed672214689c4ba5434

Observation f2f4d9b8-4bd6-4321-bdbd-65f60f741f7e · inbound

CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection cites this paper.

CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection Weak-to-Strong Jailbreaking on Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:18:24.534370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:18:24.534370Z digest=sha256:673e09e2d9c8165c451a0c6dcff8013631b9264af7ab65c254fd544f99067468

Observation 4bfedeb4-c79b-4671-9818-36d308a6b432 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Weak-to-Strong Jailbreaking on Large Language Models

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:58.581856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:58.581856Z digest=sha256:29500cc3745f23a15e0925d5dac60a0663fd9826681e2ef87b4846d93befd5d2

Observation b2e2cef7-bad0-463a-8847-12988eecd8cc · inbound

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure cites this paper.

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure Weak-to-Strong Jailbreaking on Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:20:04.101125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:20:04.101125Z digest=sha256:68675e16d06e87d7fefbb416525d0920bab8143702045fd4e29e883f10f979bb

Observation a4cca6e4-e5df-47f5-9d7f-d15684c99b5a · inbound

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification cites this paper.

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification Weak-to-Strong Jailbreaking on Large Language Models

Reference 8

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T12:05:21.962460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T12:04:13.987667Z digest=sha256:6c513ab7bd33cc0d956246ab99564dd0d82f450bc44d9152a1d738adc17fdf4c

Observation fe3c3664-c480-4d42-b6c2-57b112b72c48 · inbound

Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes cites this paper.

Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes Weak-to-Strong Jailbreaking on Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:27:43.007360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T18:25:12.407203Z digest=sha256:ce4bba362c6c8a9705348d9526749b5e2289231640d455156457219dc7469cf5

Observation e0b49d19-bc3f-44ad-a27c-e72e6c28fd64 · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Weak-to-Strong Jailbreaking on Large Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:59.307397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:60faa55126403e7bb24da89a59ca956a038a1bca78f640bbe5f784604fc2afb0

Observation 030fe290-72b6-4030-bff2-7b26a25262c6 · inbound

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs cites this paper.

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs Weak-to-Strong Jailbreaking on Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:29:44.259446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T09:52:52.603708Z digest=sha256:cfe1c15b5930391bdebfd73db3440f3081c89266eea0cadeabc04374b1dbbd86

Observation 760ed90a-bad4-42fb-85e2-c538bd131da1 · inbound

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs cites this paper.

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs Weak-to-Strong Jailbreaking on Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:35.637196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T07:07:45.177242Z digest=sha256:36fc535e6a08cb77cf4ea90010a16a39c2f1d38aa70a1e93e848b2af193f8848

Observation 52fadcb1-083a-450b-ba00-b19f4d4df369 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Weak-to-Strong Jailbreaking on Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.053927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.053927Z digest=sha256:6e545a7e2fad5cda20d00cc8ef37acb925e33ad5770b7c8ccc87822010b9d02a

Observation 47529e3a-b68b-475c-a2c1-45808eaf61f4 · inbound

Weak-to-Strong On-Policy Distillation cites this paper.

Weak-to-Strong On-Policy Distillation Weak-to-Strong Jailbreaking on Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.761661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.761661Z digest=sha256:857aba76e9d425a6ba7aa10d017dec3719f4b96e8466dfb8f4694e12187727e5