Pith. sign in

Paper Citation Record · LEDGER

COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2402.08679.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.08679 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:21.278845Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:11.145122Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7dc8bf30-aa44-492d-ab4e-2afbfcc5b876 · inbound

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment cites this paper.

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:38:39.981082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T00:38:36.992597Z digest=sha256:fe6574fd50bf6f2282e5a525b17157399f3041c972bbeb8a428a2570d7c8f098

Observation 5a7c2b66-937a-4345-9359-9e19c5744a60 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:20:44.649287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:47ae17608547c95412b8c0b84a7485084dc4b49af1a65f7e4b69932885763ad2

Observation 03979045-6da7-4365-8526-870ccb24d34f · inbound

Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics cites this paper.

Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T22:32:12.954917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T22:27:18.533162Z digest=sha256:486009fbeee27841dd5b5b83d90da5f00434bbf5b99e675ad38364da1184ec68

Observation a053897b-589c-4975-b2ea-f27e28840223 · inbound

Adversarial Preference Learning for Robust LLM Alignment cites this paper.

Adversarial Preference Learning for Robust LLM Alignment COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:21.278845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:21.278845Z digest=sha256:17abddb335a001ee84a65ff95d15f66928712000cf3659dd87cb8ae49e134166

Observation 644770f6-f978-49a1-8c07-841eeb5244af · inbound

Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models cites this paper.

Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:31.581115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:31.581115Z digest=sha256:b26c698b18c8275b91d56d825d95f8775734d381cee43a88f3b386956e76c78e

Observation eba21bb8-2c2f-4933-94eb-e0b8a11402c2 · inbound

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks cites this paper.

A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:50:24.189376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:50:24.189376Z digest=sha256:4288b46a49ddd40981c2775e5d28512974434ffb9d41b30184e85a1b6be25a7c

Observation 7cdc94e7-169b-41ab-9e61-63d88ead5892 · inbound

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation cites this paper.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.641819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.641819Z digest=sha256:fa34a1b767d706492972527bb3ff8ddd30eba5455f1c2bf8e7454f677ac94d6c

Observation 5d6e4e7d-f0d4-495d-8c1f-029d5b7109fa · inbound

Adaptive Content Restriction for Large Language Models via Suffix Optimization cites this paper.

Adaptive Content Restriction for Large Language Models via Suffix Optimization COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:58.325167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:58.325167Z digest=sha256:320174fa0489907a9dc59dee27399307a79817d1810456c9517e3fc8e3aa8a42

Observation 512e92fe-7ea0-4a83-b1d9-7f13210766df · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.972622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.972622Z digest=sha256:fec4a3ee774be37edf019162acc97127a494e9692caf548493f87aa35c72f6c1

Observation 7b74580c-e5aa-4de2-af2b-41a82de52755 · inbound

A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives cites this paper.

A Systematic Survey of Model Extraction Attacks and Defenses: State-of-the-Art and Perspectives COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T18:12:37.189913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:12:37.189913Z digest=sha256:36f73eac61e0fbd68df60069654c18bf4e5586e221e4d7f44ad7aed728c358da

Observation 959cbe90-79e5-444a-8150-c924af79b856 · inbound

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring cites this paper.

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T14:51:03.667508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:51:03.667508Z digest=sha256:2717c13921bb7cd25e66c6130047aaad6216784b81470f0f673412b8c6a36e19

Observation a676b136-449c-4440-adec-2f98c854da4e · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.277726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.277726Z digest=sha256:e3850f265c9f0016126f3e8f689a134d6423d3d6b7ccb5cb1517d37a3f25c466

Observation c6ba4a17-9d7e-4a19-8b37-598ca4bf46c7 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.515431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.515431Z digest=sha256:7d60e787392d49b1d60f4b8bf0c2a408c3044f52e8affd6c4f87c9de43e81526

Observation ebe02e7b-4f76-4c39-9cb4-dfe6e23130b8 · inbound

Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks cites this paper.

Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:06.471703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:00:06.471703Z digest=sha256:7c1ead42e6ba2b6a5b1163a72c723815a3d59f6d59b60e0c63333357ab12a15e

Observation 38d38698-e6f3-47cf-ac97-c5cf118cb558 · inbound

Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization cites this paper.

Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T11:11:00.443361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:11:00.443361Z digest=sha256:c7a1a1ef11abbf94c8acb525956cb8098435295f66e0a3cb41f0366f4840f577

Observation dc655edd-f014-4170-8710-fc4be3fffd6c · inbound

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs cites this paper.

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:55:38.107722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T01:54:22.995178Z digest=sha256:2345d3c59a95dc803a972b80e9fd92271927d7a73bf176c55ce45d50d8b48fca

Observation 88091f53-7646-4091-8521-989e081c4e5d · inbound

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models cites this paper.

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:57.972169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:12:34.909229Z digest=sha256:22258042a11a2d5a10bd619bf08ab77751df1b2db21220d4616d3254ada1d4e1

Observation 6deb68e8-9e7b-40f3-829f-7af9ddd7628f · inbound

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models cites this paper.

Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T16:34:03.945251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:34:03.945251Z digest=sha256:0cd1aee6a65e343db9e6cd2ce223e95165d99dbf2543a14f5f0379607970ec72

Observation 40d6ff13-23c9-4655-b541-6420a6a35a67 · inbound

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs cites this paper.

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:26:00.029979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:02:52.006859Z digest=sha256:153294da6839fbe60d3b80e7ef3482ae366c4999291acaa28364e16bed4e580b

Observation 58685fd8-51ec-40fa-9dfe-3bc632cd8a88 · inbound

STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming cites this paper.

STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:08:59.273795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T03:06:21.624159Z digest=sha256:ecdce450bc7c37fa7037a2495292318af92d38d2b64ccbdda47cc3b8f76113e6

Observation 3c824040-4556-4982-a3f0-9912bdecf8c6 · inbound

LASH: Adaptive Semantic Hybridization for Black-Box Jailbreaking of Large Language Models cites this paper.

LASH: Adaptive Semantic Hybridization for Black-Box Jailbreaking of Large Language Models COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:49:35.496621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T04:48:19.926845Z digest=sha256:e7d8f4812a5695a3e9bc33c2287acc7a50b23567f8159b48fc6abdb3e5fd386e

Observation 16e4472a-7f36-48ee-af6e-cbe5c49b9e8f · inbound

Adversarial Reframing: A Framework for Targeted Generation in Language Models cites this paper.

Adversarial Reframing: A Framework for Targeted Generation in Language Models COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:36:20.994701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:35:51.862736Z digest=sha256:711ad4f3c4d9099315d3a4cba93e9d65f0f62e52ab161e33a5d22253d566d603

Observation 462a2daf-22b4-43cf-a9c8-9acf5948e827 · inbound

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation cites this paper.

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:50:11.146715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T20:58:53.119386Z digest=sha256:1d78bb6985865263058d03d42692adeea476a9016961943c644b35de50418f36

Observation f687973f-0d21-459e-9f29-1292a72ed43f · inbound

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation cites this paper.

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T10:16:41.626900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:16:41.626900Z digest=sha256:b97ed0dda471bc2998b4dfeeb8f0aaa1b5da6cb99974de95bacb2ef9ba41196d

Observation 9cbd144b-cd6e-40b6-b065-7c8fe71ab14c · inbound

One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models cites this paper.

One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T21:01:04.881568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:01:04.881568Z digest=sha256:aaad5cd3dc690280f201c916998934c5a844bc22809c03dc6e4a6894a082f2be

Observation e76c0b05-541f-4199-9896-e5336725ddf1 · inbound

Weak-to-Strong On-Policy Distillation cites this paper.

Weak-to-Strong On-Policy Distillation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T00:26:21.232534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:26:21.232534Z digest=sha256:568a93d55e53f20995669c7bdf6d1f8f4828f84e016a2409d35c30818b1346cc

Observation a536787a-3a5f-4751-9dba-ec9f93c6c34d · inbound

One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs cites this paper.

One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T23:07:39.746628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:07:39.746628Z digest=sha256:2015997d24265a2e13138bcaace903dc517d6bc46eb4b7626e7073be47cd78f3