Pith. sign in

Paper Citation Record · LEDGER

Weak-to-Strong Jailbreaking on Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2401.17256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.17256 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:06:39.024687Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:29:44.257667Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c586368a-150e-4f55-84a3-29d1350e9d56 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Weak-to-Strong Jailbreaking on Large Language Models

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.504149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:4895c18cac57b40767f41a9fdf79fa5a3316ac5e6c05df2f05a30115b437456b

Observation 22481f73-652b-4e7a-abb2-ec51598fee79 · inbound

Peering Behind the Shield: Guardrail Identification in Large Language Models cites this paper.

Peering Behind the Shield: Guardrail Identification in Large Language Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:45:21.416651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T03:45:14.234545Z digest=sha256:cb719bd29c75c55033352e1ded82e5026939636de43e0a7492379a8e7adc0434

Observation 2ed9ab0e-3aec-4809-979b-34f74362d5ad · inbound

Confidence Elicitation: A New Attack Vector for Large Language Models cites this paper.

Confidence Elicitation: A New Attack Vector for Large Language Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T22:06:39.024687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:06:39.024687Z digest=sha256:925d716d1d2e2de03544904ae3fcca43ac554786c96f5926c31e1ce4df2b022b

Observation e17858c5-6305-4ec4-a4cc-c9d78526c34f · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Weak-to-Strong Jailbreaking on Large Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:42:34.219547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:bdad00e6fb328d4b04dfeea4d5546f59a94bf008d13c01655f761fc89c29c2e0

Observation 65d59232-1a7e-4c6b-8b98-3e67cee7e60b · inbound

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences cites this paper.

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences Weak-to-Strong Jailbreaking on Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T10:23:06.305919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:23:06.305919Z digest=sha256:b2eea452a699d3f4f0ee8247f2a3502789e410117d8957f364d30b56c456baa5

Observation 36ca2ae1-9574-45a8-a4bf-cc5df38bc296 · inbound

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion cites this paper.

Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion Weak-to-Strong Jailbreaking on Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:15.103581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:15.103581Z digest=sha256:4794964c067fbc5f966dcc90df5e4ec2a28d337ecdcef8ebe00482f01ec9ce3f

Observation eb482b29-547d-4846-bd6b-42e9cf4b000f · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Weak-to-Strong Jailbreaking on Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.759736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.759736Z digest=sha256:92cdc773e19e5462ea595a4a95ce31d1a7b52db4f59c7f0f0607f13e21996f3f

Observation c35c1c48-c59d-4924-9014-51bde1ac83e3 · inbound

SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues cites this paper.

SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues Weak-to-Strong Jailbreaking on Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:47.630473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:04:47.630473Z digest=sha256:51465df90182c632f33eb4fb2e20974ff1d6ec273539895d04716cf9058cc3aa

Observation 595452c4-6357-4760-8f52-30d432d8a57b · inbound

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures cites this paper.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Weak-to-Strong Jailbreaking on Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.606600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.606600Z digest=sha256:af24c79450522727c6969d2387b2c31ca7f9ada9bced66ff5a70ea3189d30567

Observation 014a108e-00d5-492c-8dad-fb42b13f5e67 · inbound

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem cites this paper.

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem Weak-to-Strong Jailbreaking on Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:42:12.897176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T08:40:56.186349Z digest=sha256:9750fe2d3eccbf1220d7fffb2dbc16aeeadd2386f0e544c030825336a09f0b69

Observation 31368c05-277e-4bbf-b0e3-6e2c92dff6f8 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Weak-to-Strong Jailbreaking on Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:57.504508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:57.504508Z digest=sha256:458c856162d6e62f3c2850216d7e5e25b24b0de91cf1bdc4e14b0b974aa973f6

Observation 2687ee65-1b68-46c1-91c9-905b049d5d40 · inbound

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training cites this paper.

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training Weak-to-Strong Jailbreaking on Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:46:01.155732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:46:01.155732Z digest=sha256:cd50e6616585bf6119272c9fcff8ad7c464f28541eed2e07628942aa64c04262

Observation f6b8c244-c238-478b-9765-920e0081a3a9 · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Weak-to-Strong Jailbreaking on Large Language Models

Reference 196

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:25.675116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:25.675116Z digest=sha256:21b33225c0aa4b764e2e45a022fbf54f9a04ced05b92a8ad2fb5d1c88b645334

Observation cd0431ae-c85a-46f5-878f-6dcce83602e6 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Weak-to-Strong Jailbreaking on Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.631918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.631918Z digest=sha256:6bc3dd713f71c2faede727bff978e3d4df781577132f39c5c329d729d338f6cc

Observation 4bfedeb4-c79b-4671-9818-36d308a6b432 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Weak-to-Strong Jailbreaking on Large Language Models

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:58.581856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:58.581856Z digest=sha256:12cd5452697bc671e90c4e7f2e306fc4e95c04bdd2ac478e766c29c05c93790a

Observation b2e2cef7-bad0-463a-8847-12988eecd8cc · inbound

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure cites this paper.

Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure Weak-to-Strong Jailbreaking on Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:20:04.101125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:20:04.101125Z digest=sha256:22272d38bbae8dcaba5dc43bb4fb1a3c5f1c9f639c62be0ba47caae486d99b16

Observation a4cca6e4-e5df-47f5-9d7f-d15684c99b5a · inbound

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification cites this paper.

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification Weak-to-Strong Jailbreaking on Large Language Models

Reference 8

Resolution
malformed identifier
arxiv_id, observed 2026-05-10T12:05:21.962460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T12:04:13.987667Z digest=sha256:428f2cbff702e22d4d8f5001e6fed25d899e861390f7271c52ec5f802f7f4867

Observation fe3c3664-c480-4d42-b6c2-57b112b72c48 · inbound

Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes cites this paper.

Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes Weak-to-Strong Jailbreaking on Large Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:27:43.007360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T18:25:12.407203Z digest=sha256:40cbe6fd69708bdbbc2eab75fe9d3caf81af2742492feb6efba380024b5af4a4

Observation e0b49d19-bc3f-44ad-a27c-e72e6c28fd64 · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Weak-to-Strong Jailbreaking on Large Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:59.307397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:6ce433eed2e58f591eaa69f447326eb01efc8c59b63b1c6f651141f1c35defca

Observation 030fe290-72b6-4030-bff2-7b26a25262c6 · inbound

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs cites this paper.

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs Weak-to-Strong Jailbreaking on Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:29:44.259446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T09:52:52.603708Z digest=sha256:b9e51badcc9d0cb136cc13d06ae0d6364a62b3daf7fa5e3fe6c597019c159a3c

Observation 760ed90a-bad4-42fb-85e2-c538bd131da1 · inbound

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs cites this paper.

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs Weak-to-Strong Jailbreaking on Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:35.637196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T07:07:45.177242Z digest=sha256:c88e3641a387709e1a897ff9d9577b3fc6fcb863ef144b07f0c169027ba4679f

Observation 52fadcb1-083a-450b-ba00-b19f4d4df369 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Weak-to-Strong Jailbreaking on Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:55.053927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:55.053927Z digest=sha256:652a9b0df418d4fd6aac65d40f36b310b6358b249b6a8b5b227551af3d0e28d0

Observation 47529e3a-b68b-475c-a2c1-45808eaf61f4 · inbound

Weak-to-Strong On-Policy Distillation cites this paper.

Weak-to-Strong On-Policy Distillation Weak-to-Strong Jailbreaking on Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:23.761661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:23.761661Z digest=sha256:54ccc982c117d3ad79da6da7bfe366a19a20975e199ffaecdbc1bd81b7ff9bce