Pith. sign in

Paper Citation Record · LEDGER

Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 91 inbound Pith citation observations for arXiv:2310.06387.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.06387 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 91 of 91 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:10:31.064804Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

21
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 117d23d0-55a9-4849-8078-a4f74db8fd9c · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.426486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:4b3fa8f188a7f69cfb6a932014325a553ad419c4a5c20d1a2cf3a807a14fb9df

Observation 6381b773-e8c9-48c4-92ef-f99986a5ad55 · inbound

DROJ: A Prompt-Driven Attack against Large Language Models cites this paper.

DROJ: A Prompt-Driven Attack against Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.300047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.300047Z digest=sha256:90b063fc93c0446ded15d7be987dab0da9b37232f7a94cd966a9a0653bb97ae1

Observation 9ca19ab0-3dea-4262-820a-f10d6891b276 · inbound

DYNASHIELD: A Black-Box Moving Target Defense for LLMs via Dynamic Decoding Customization cites this paper.

DYNASHIELD: A Black-Box Moving Target Defense for LLMs via Dynamic Decoding Customization Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T18:41:06.153149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:41:06.153149Z digest=sha256:678cd561fc558d692cfadc2358faa4b15b581fa3f70cfcbe4657acda3296dc69

Observation c95fb98b-469c-4d12-a165-66b54dc5fc35 · inbound

Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM cites this paper.

Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T18:53:15.305158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:53:15.305158Z digest=sha256:87db23104e3e5064821da41d1f9c1ce01e71c3bc4d582b3f3d56ab38151ee757

Observation 6fedc0ca-9974-49fb-8a0f-8024c1a93efe · inbound

No Free Lunch for Defending Against Prefilling Attack by In-Context Learning cites this paper.

No Free Lunch for Defending Against Prefilling Attack by In-Context Learning Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:49:38.142015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:49:38.142015Z digest=sha256:6417fc4c37ba17ac5c0367c934d7afb2b13de52e4dc7e2d9d2be97d231c82dab

Observation d24f0f6a-2382-49ba-b029-bfaf1f3e4515 · inbound

Towards Responsible Governing AI Proliferation cites this paper.

Towards Responsible Governing AI Proliferation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:48.601255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:48.601255Z digest=sha256:be889a7770e2b57a13ce5335cff7043a2888e3877f77fb5fd3bb15897d10fd94

Observation 4cfd5c5e-551a-4de1-8066-5a1b12d2d127 · inbound

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models cites this paper.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.713244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.713244Z digest=sha256:84f8f429a3fe594ec609714327aab3052114d2de838f7c5fa78ba0218c998e10

Observation 9cc1274f-7c2e-475b-99c0-af3b7e4c34ac · inbound

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models cites this paper.

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:38.873180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:40:38.873180Z digest=sha256:d77aab7f15197b2d26f0ca5a37441e113072cd3b1c76600591a4dffcc51873aa

Observation 7ef6edf3-7144-427a-8a56-67f480a69451 · inbound

SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation cites this paper.

SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:28:56.732300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:28:56.732300Z digest=sha256:b0265f8017a991ac33890d3b74250584a3d4aafd99b8a7246338335b592e19bb

Observation 962b564d-6ac9-41c1-8db9-b5cc36a03762 · inbound

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models cites this paper.

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:25:54.723257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:25:54.723257Z digest=sha256:2c5c3b7c0af6ad0458412b624c95d9ea71b9dad269183631bd1b18d2c5c362fe

Observation d6ef1e05-0fd2-4d45-bd4d-a8d6458c5f61 · inbound

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs cites this paper.

Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:26.984854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:26.984854Z digest=sha256:6c4c98fcecc86670d9454d541a9c8447db9010862a4851ee9dc79a7f5b468bba

Observation 6d605eed-3151-4592-a9e4-07be3a451d65 · inbound

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense cites this paper.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:59.912758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:59.912758Z digest=sha256:b6b7863fb61b2ee32ca4fc132db19f3a6e5d639e6ac4126d499128dd8471d412

Observation ed581a41-2e19-4c3a-93a0-50d01a279b19 · inbound

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning cites this paper.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.856473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.856473Z digest=sha256:1d0e60fb69f8bb763d027e4766f94c5ae1e372768dd6a28a49685478dce9821a

Observation 19cf3e0d-9de7-48fe-b988-f87b47517819 · inbound

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks cites this paper.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.737120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.737120Z digest=sha256:104e1f201c1bdd0bc731064450b1beba4ca01b713bfd4e4ba8d843bc3b88a6ae

Observation cb412c3f-1d20-49c5-8853-766be59699f9 · inbound

Episodic memory in AI agents poses risks that should be studied and mitigated cites this paper.

Episodic memory in AI agents poses risks that should be studied and mitigated Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T17:58:18.112923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:58:18.112923Z digest=sha256:e940fece228399da3cb9f4a1683e24e57c2d0351a019f0520e542a623156bd38

Observation 0175743d-88fe-4eb6-9851-69008c436cdf · inbound

PromptShield: Deployable Detection for Prompt Injection Attacks cites this paper.

PromptShield: Deployable Detection for Prompt Injection Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:38:10.811952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:38:10.811952Z digest=sha256:d35a3d5689b7faf6267aab22273436959ce937b1807ba0cde482d7f4f627c6b2

Observation 0f26dd78-4c35-4556-8f79-98f7892de47d · inbound

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models cites this paper.

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T00:12:04.817181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:12:04.817181Z digest=sha256:d7dca2dd3dde39efd109d14b1743b7efc9e5eba83bd9b1b9258ea54a445030bf

Observation a928bc4c-1dd2-43bd-862a-186b3966b4c8 · inbound

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling cites this paper.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.908552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.908552Z digest=sha256:315855d71a565bb4c95449c5a3902568bc7f51231f74b6ab40915842a1d25c14

Observation e184b199-3148-4634-bc49-02b558a134dc · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.673565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.673565Z digest=sha256:5244b1c07e97174acac64eb2001a5551a51ca937bbe4b364d43fe8193c23b765

Observation e2a488f1-08c4-43a6-a2a4-94b88f78a080 · inbound

MetaSC: Test-Time Safety Specification Optimization for Language Models cites this paper.

MetaSC: Test-Time Safety Specification Optimization for Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:45.010325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:45.010325Z digest=sha256:1ac30cbba0b5756d9e87b82bd326857e9b183d9cbc947d9f9de7128417d30028

Observation c3d0bbf9-03d8-49cd-8153-bb7f999c5e68 · inbound

DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification cites this paper.

DETAM: Defending LLMs Against Jailbreak Attacks via Targeted Attention Modification Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:10:31.064804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:10:31.064804Z digest=sha256:443cff337a616cf65648ed911e3edfc94a641518131819883e010f70c5108913

Observation ecf5f09d-1fdf-4ef4-912b-9a0129a3c8ff · inbound

Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control cites this paper.

Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:52:07.409816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:52:07.409816Z digest=sha256:89f78fcfd8d533f9b36a9e42d62c8d22fcf447b1b7af2a41b193f8a5c8d2de3e

Observation 1a0912ed-9faf-4177-80bf-8c2623f7a839 · inbound

The Automation Advantage in AI Red Teaming cites this paper.

The Automation Advantage in AI Red Teaming Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T05:46:02.176124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:46:02.176124Z digest=sha256:9532cdec88d0df37a707cfa7610b89de5c58ca1e2e4f747b059f353a4535f373

Observation 5cd12cf8-ef86-4a7c-b00c-6289a1ee42e6 · inbound

Attack and defense techniques in large language models: A survey and new perspectives cites this paper.

Attack and defense techniques in large language models: A survey and new perspectives Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T04:33:06.577405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:33:06.577405Z digest=sha256:afb5a552eeee3e538bbd346f822d10364b57dc0f7bbaa75eee300753d726cc65

Observation f2908be1-0bdb-4385-8425-46098b195243 · inbound

Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement cites this paper.

Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.311556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:45:55.311556Z digest=sha256:7834c879a1351384e7a6d69473d0895f6eeed22de2374b6a69616951394d08c7

Observation 2cb7c7aa-8f91-49a8-bb57-79e5a3691a0e · inbound

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks cites this paper.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.380405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.380405Z digest=sha256:11ce348782c7bf076372942b82316d711ca7a53d9502e7afa2169e534b2f0a68

Observation cba2e2e6-8940-425b-aecc-e698ad06b693 · inbound

Advancing LLM Safe Alignment with Safety Representation Ranking cites this paper.

Advancing LLM Safe Alignment with Safety Representation Ranking Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.326956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.326956Z digest=sha256:466b4d306587f6db8ed785960725ece8f0ee8903d9403056a30a7d3b9004ff84

Observation 9bdd8198-b371-4ea6-b84c-bdc93100f652 · inbound

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval cites this paper.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.745729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.745729Z digest=sha256:a5412fb4cfa6c2bce43ec3347c6d27a750a0d58258092beb3356fd98429e7657

Observation 380184e7-86ed-4d1d-acad-bd528028193c · inbound

Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models cites this paper.

Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:41.028053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:41.028053Z digest=sha256:2777abeba71ae571d76518b91a7f2a92b93c470c9d9b7d9cebf94051f47c6ce7

Observation 6f79f5c5-dadf-42b5-871f-460fbc8e94b7 · inbound

Secure LLM Fine-Tuning via Safety-Aware Probing cites this paper.

Secure LLM Fine-Tuning via Safety-Aware Probing Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:11:35.853681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T13:07:09.402763Z digest=sha256:74c545b126ffc2e9f826e80b7c10bdcd070520a3a8d8679e75a66db961345257

Observation 50dc41cf-1536-4e7b-869a-64ad6d1d0615 · inbound

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation cites this paper.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:30.720380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:30.720380Z digest=sha256:751bb2f642406ee322de3ecdc2636cfa9f09585e59cfd5d2b3b729f71e615d0e

Observation 315c024d-59f5-4a55-8fe7-56370521a7b3 · inbound

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts cites this paper.

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:11.052058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:11.052058Z digest=sha256:3a9ac83fea907f5719e626c6bb650e7f43c829d59d18e56fd41b4f5c60c6f3e2

Observation 73d2de4a-8bd0-44b0-b6e4-7b1dec2610c8 · inbound

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models cites this paper.

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:24.102566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:24.102566Z digest=sha256:3f54b2958d6c4ae665e4850e70552c11d81f8ba9bbc0c4cf1143077cf16e4760

Observation 5faeabe1-ec16-46fd-8fbe-ec8a76ffd52e · inbound

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap cites this paper.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:22.236454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:22.236454Z digest=sha256:0c9e59a6233b645bdec83416868938f9318b9bc957ff415cc09f139618d9c3cf

Observation 75792518-7d97-4fe7-917b-fcd454ccace3 · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.987956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:e49d6d1036a07c8f7e3376d692f38f25d896193ea06ecee2582e82c19914fce9

Observation 60f67dd0-5532-4871-9215-078dcc65b23a · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.551096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.551096Z digest=sha256:201af96f6bf2ee071202c43939eae2dd20ddd16938d4d541067c2fe713b2ac02

Observation abd6921c-f91f-4c0e-b1af-47bf8baa381a · inbound

TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts cites this paper.

TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:58.135736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:58.135736Z digest=sha256:3210f51ff964d95f70c05014a5ffe0d7bb12b15cacfe04af83441e8257c75c68

Observation de60991d-b09f-4bc5-815e-e386c2095b22 · inbound

Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations cites this paper.

Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:57:16.542808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T10:54:55.262180Z digest=sha256:8ece955340986a0a19ec0841bd51e1f13be8d7709a428d2040040280fcbcb045

Observation b255213c-9252-42fc-8015-5501bcc45f5a · inbound

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges cites this paper.

LLMs Caught in the Crossfire: Malware Requests and Jailbreak Challenges Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:10.112904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:35:10.112904Z digest=sha256:ce4bd03055db3ad45490c0989163c903904320b28b6c72959de9324399ce5219

Observation 0bfa0aee-8356-4b43-a9f4-82adba2f8a86 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:27.001406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:27.001406Z digest=sha256:c4c94cbab629fbc9e1295641558960b08c2f3ed3d21bd8250e9d35f13ab86001

Observation 55f36b96-93fa-4d92-8919-e1cc47c74088 · inbound

InfoFlood: Jailbreaking Large Language Models with Information Overload cites this paper.

InfoFlood: Jailbreaking Large Language Models with Information Overload Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T01:02:31.037588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:02:31.037588Z digest=sha256:5df32941ccd58f97db57c4501187f16b8377ec0559dabca2e3f76f09c6976937

Observation d1ac8baf-a28d-404b-aed8-130ee71f471c · inbound

From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem cites this paper.

From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:09.987996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:45:09.987996Z digest=sha256:68e45cff8e5b6a550569cc7950874be73c52538a345a0d272935e650c6072a91

Observation 03d5aa00-3481-44ee-a5fb-50e0a81e5135 · inbound

Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models cites this paper.

Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T19:31:54.439019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:31:54.439019Z digest=sha256:6afd55ac428bbdbaae299bef695c45d479487e643ae0a3e7bff89697e5c9ddeb

Observation acd11a48-ed3f-4c4d-be94-32473105ddd4 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:53.797609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:53.797609Z digest=sha256:e0c5d8d8539f3ef1ddd3db1ed6ad97f35956c9c1a37a1ac6aa42f5845baa2d3b

Observation c8564b34-057e-407c-bf81-6ed1135155b4 · inbound

Linearly Decoding Refused Knowledge in Aligned Language Models cites this paper.

Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:35.014424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:35.014424Z digest=sha256:d3ee8051a1d2e7c1080dea326c1027dc815ac498a3eb8f3bc634c7c127de2af1

Observation c9e12c66-942d-4f13-b196-686c4bf8756d · inbound

Defending Against Prompt Injection With a Few DefensiveTokens cites this paper.

Defending Against Prompt Injection With a Few DefensiveTokens Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:42.838856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:42.838856Z digest=sha256:6886b414fb8fb9535068665194c19a45e1d24f2bc446c9b4896622a6af6537fc

Observation 78c43e03-bdbb-458f-83aa-5413134e33b9 · inbound

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation cites this paper.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.743717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.743717Z digest=sha256:4417fde8b3842c48b61eceb860d774b4d37a3796fb021a008ae13f2ca4741301

Observation bf2e2b86-141e-493b-a58e-0e75f00be4e4 · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:55.100347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:55.100347Z digest=sha256:8bfc365917c059e6168cea93cf0da948c666e994fdcc47cf3bfb0d81deac604d

Observation f0d19bb1-5cea-4a7d-928b-271ecf66ead2 · inbound

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? cites this paper.

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:34.156403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:17:34.156403Z digest=sha256:c5d2f34e05ee0356e275da31e26336f91bc7042b61c6cfc087682f8607747fa0

Observation 218856ef-8017-4313-b7d8-3c4288aa9d6f · inbound

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles cites this paper.

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:43:39.293986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:43:39.293986Z digest=sha256:d5e0ed65bfab78353a177e97c80b357704297fbbc1033831ace0e36f43cc48bb

Observation b53d180c-a509-4bc2-8d06-8efb4e8c110b · inbound

Automatic LLM Red Teaming cites this paper.

Automatic LLM Red Teaming Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 1685

Resolution
unresolved
no resolver link, observed 2026-08-06T00:04:25.504135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:04:25.504135Z digest=sha256:015f5a7cab4815564894524429c98fa52e65ef0c4de2e25e5acd2a6ff4c84119

Observation 11d2042c-6acd-4c2f-8ba3-21a2845c9a02 · inbound

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection cites this paper.

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:24:23.428703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:24:23.428703Z digest=sha256:f67239eb0c5f3f46b33461d48b2f0a02e89e74c5ba6aec91cfe355c9d5954b12

Observation 7bbdb906-2757-4b09-abf5-0001b09d1ebc · inbound

A Survey on Training-free Alignment of Large Language Models cites this paper.

A Survey on Training-free Alignment of Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:48.551115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:18:48.551115Z digest=sha256:cbfeea8fc5cda0dede9d9c3171891f1b71f920fbe97e2170e2a92708892833fa

Observation 3e071d78-c297-4a80-8084-43d8cac74c4f · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.559678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.559678Z digest=sha256:eaa4488743863532be36a735599687eeadcb02a5970ef154a0fb57063ae24c2f

Observation 9443b52e-49d8-4b5e-b54b-912e938cb6bd · inbound

Mitigating Jailbreaks with Intent-Aware LLMs cites this paper.

Mitigating Jailbreaks with Intent-Aware LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.137872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.137872Z digest=sha256:86296e61f573f8c386d70fd7e20c4a017aaf3a7570e46b6ec6aee3d8681fddbe

Observation 2380904f-c806-4730-9838-7803ec4e6629 · inbound

CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection cites this paper.

CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:18:24.500459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:18:24.500459Z digest=sha256:0d3a88d403438ac20ca62e73cd04796f7b41545b53a7d2a74ac5b2343518cde5

Observation 103a5d3e-866c-42b7-a0b0-5c77d922d988 · inbound

SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks cites this paper.

SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T18:05:32.424178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:05:32.424178Z digest=sha256:c9c3ac17ff13cfc666516e5ffa27935416dd4d2b448e0fb9defdd9b003b4e992

Observation fe642ff8-1d94-4870-a75b-bc2df70e9cc0 · inbound

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? cites this paper.

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:52.110346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:00:52.110346Z digest=sha256:63a020da6f72c13adba0ab5ce5b6b7338d60ba150bd08fedbf927e14c3e429b9

Observation 921a1417-8384-499b-97e9-4bdae34b8f08 · inbound

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs cites this paper.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.479691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:1eeebc00870d55236552af2ce567d66b8be58506a644ff1d79c8f4a4f9b8cbb5

Observation 262a77df-7593-46d2-b169-d8fe64bae7d3 · inbound

Baichuan-M2: Scaling Medical Capability with Large Verifier System cites this paper.

Baichuan-M2: Scaling Medical Capability with Large Verifier System Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T11:50:14.742567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:50:14.742567Z digest=sha256:4bd53e6bbbd596bff9a9ed0897b430eb5380bee7e0577fae5d66383c6f70e019

Observation fcf06fb2-f84c-486c-9e38-4c9f4c0ca189 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 264

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:08.207569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:08.207569Z digest=sha256:e913a0ee9bff8e9639dfd67d3b7b5cb69f6bbec713d56f6a69cdb24740e29d32

Observation 881ed3e0-8875-48cf-a485-a295d796bf27 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.470662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.470662Z digest=sha256:9372dfa340fc1be5ac27dfa1968b1556141da2a3fdfee8e31870a08cf33f80c1

Observation bb991391-33fa-4d40-b1c4-1646f578c6a8 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:57.778089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:57.778089Z digest=sha256:8c34672a345889c8c11e579d00ed246bffb3c7959fa6a18ce470dc02dd4770af

Observation 8b9af65f-eef9-4e39-b148-2c7c4b2f4287 · inbound

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models cites this paper.

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:25:54.630924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T05:22:30.509147Z digest=sha256:5b9305c786eb69f4c1366f4248dd308913c1f34876a609fc2593a9ae4ffafa95

Observation 5187e3f8-ddbc-49ef-b93f-23f7013e2313 · inbound

Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels cites this paper.

Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T07:06:38.120540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:06:38.120540Z digest=sha256:0f27a18b2f687a492e5c28d8a8cdc2e3a2aab69c2a292cbf08bf6706725c8171

Observation e241d81d-697b-4658-86dc-2532dc2b8cd0 · inbound

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs cites this paper.

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:55:38.050506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T01:54:22.995178Z digest=sha256:b889daeb5a76f00706e95595ee8b5b12dddcb7ea39a9477bf7937de432758d35

Observation 1c861905-1bde-4f80-b222-a54f7756b3e8 · inbound

GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents cites this paper.

GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:05:26.677753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T07:02:28.660002Z digest=sha256:e7177bcd8f9955978d09c80cabd2a0b0e1dd15b02587920442d94c20faf1cf1e

Observation b6802b60-e140-435f-8723-813ea5ddc511 · inbound

ContextCov: Deriving and Enforcing Executable Constraints from Agent Instruction Files cites this paper.

ContextCov: Deriving and Enforcing Executable Constraints from Agent Instruction Files Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:50:12.375753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T17:49:08.192347Z digest=sha256:2b2a3f159e4203b87e9184624c87a9a5896c70d5d1ebb6b7a353a608465278c0

Observation b6f4ff22-d27a-4aa8-b762-1b9e0680ce30 · inbound

TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense cites this paper.

TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:50.778962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T18:26:19.922383Z digest=sha256:ec35289223e727620e8aaf282498b3a71e8a1ea7b173a31b4d1c851c9e630ae9

Observation 3d58642f-7d0f-43f0-9c67-768094314bbb · inbound

GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking cites this paper.

GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:45:37.340698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:40:59.993298Z digest=sha256:7aa33d70d6e799873ae726c34ccdf35d8832c88992fd540c465d631cc516e37c

Observation 25c6cce4-a38e-4317-b3ab-66456ab89088 · inbound

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs cites this paper.

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:26:00.015367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:02:52.006859Z digest=sha256:72edcbe48ef574d1c1851f89a11f28a95276d1ae58d680c5951a12e935e6deea

Observation 7b10e68e-4c58-4538-949f-52ec49ce67bb · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:5d29794f8f5763c1c157a481cf292c06752f06a97446f819f2fad6adfebd7b9b

Observation f2448a1f-80b8-4bbe-b1be-faeaf618d404 · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.931019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:b534ecc27ecd94b240f786fde960da204a361ea9b9fe6318df8fe9f02e8e7a1a

Observation 75fea871-0b2c-455e-b68a-69d3a924bc36 · inbound

A Systematic Study of Training-Free Methods for Trustworthy Large Language Models cites this paper.

A Systematic Study of Training-Free Methods for Trustworthy Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:22:37.353269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T08:19:42.671690Z digest=sha256:bcf2f1f5a0476fe15737e43ea3419f041d2bce5ac0a76c103809f4cae4eef945

Observation c0804e6e-5974-4d59-9fc7-fcc348e25e30 · inbound

Jailbreaking Large Language Models with Morality Attacks cites this paper.

Jailbreaking Large Language Models with Morality Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:21:27.030426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T06:17:14.242224Z digest=sha256:85584cf86054cae0709a89ed852127c63284260219eb88c00e369c996e6968a0

Observation 654cd810-3283-4979-be35-e00290ea9fe6 · inbound

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models cites this paper.

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:04.411354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T02:21:29.463149Z digest=sha256:4ff0d9c6cd5423205b0ff62a8ecd4503e1b21d9bc141fe9fef12dee7db43258c

Observation c698fb39-f359-4ccf-bc76-66c16fef678f · inbound

Automation-Exploit: A Multi-Agent LLM Framework for Adaptive Offensive Security with Digital Twin-Based Risk-Mitigated Exploitation cites this paper.

Automation-Exploit: A Multi-Agent LLM Framework for Adaptive Offensive Security with Digital Twin-Based Risk-Mitigated Exploitation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:09.476633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T11:45:21.608035Z digest=sha256:8da1327d2436761dbee4eb0d12f7fc9a89fe5ff4023c4b72fc5c9764256416bc

Observation 63e4218e-7582-4fa1-9f5b-30946ea95025 · inbound

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework cites this paper.

A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:09.567702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T07:53:13.746141Z digest=sha256:3b6dd220a078673471f83a51a55abd6ac09a107d0207fc3adaf90988dc7bd31f

Observation 5d7afb9e-d275-447f-9f23-c01aedbff02b · inbound

Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models cites this paper.

Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:01:09.629032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T20:39:39.225898Z digest=sha256:cd0210cc02507e36bc3be6a0a0921a18b969f8a505205eb5429eec428c0c7fc9

Observation 598a05e0-a41c-42a9-858b-69653353566e · inbound

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms cites this paper.

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:43.987363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T01:46:49.586630Z digest=sha256:5733a58bfa42ec788ddb90959e3dc674c808eb90dcc8bdaa3c95a6deca5dafb7

Observation 833c9c08-b805-4c2d-828c-8bd4f095fd6c · inbound

PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures cites this paper.

PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:03:36.763330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T18:03:07.646917Z digest=sha256:3a1aa6481aa22d9e0675eed31d247d8511689becba4c506b8ef302f564f93226

Observation c766b4e8-3c80-4acc-acc4-12cc8b8205e3 · inbound

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models cites this paper.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:53:47.699039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:adf193fe220fc7d019fe61e79fa073f7c9473d204eded75a519c35bd0460b729

Observation 440422c9-2f00-4359-a058-80339a4cc457 · inbound

THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models cites this paper.

THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:19.173126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T15:05:31.392966Z digest=sha256:8647649021863ebd1104e5bd92bb0e031ce95db88262c1ad86f8cc273cd99a35

Observation 0806e1bd-bc88-4890-a644-04a84df5052a · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:59.309811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:1b394ea481b61c43e270eeed8d6241c0d67c3d07a996e2d22b43d0a1a0de1ab1

Observation 7ad5939a-73f8-40b2-993c-5fe2b2a0181f · inbound

A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots cites this paper.

A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:09:22.521177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T20:00:14.036515Z digest=sha256:27035a8d41b2f0a11d2be20638ccd8966c76ae9bcff913684abd6fa9c2419dae

Observation 89595de1-b9ad-4080-a147-1ed5fd276557 · inbound

Investigating The Security of Modern AI and Cloud Infrastructure cites this paper.

Investigating The Security of Modern AI and Cloud Infrastructure Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:42.325046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T11:31:39.910784Z digest=sha256:e76a21b8d4517ad208bdf5d09d034207d2113fc90ea479663a8b89b51af82d4d

Observation e938f8b3-7d40-4f1f-9857-d41e30c67f32 · inbound

ToxiREX: A Dataset on Toxic REasoning in ConteXt cites this paper.

ToxiREX: A Dataset on Toxic REasoning in ConteXt Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 223

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:43:07.121701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-29T04:33:18.794505Z digest=sha256:90727e84c1328356de8071802a6a5a0bd3f15ef25b39b9ee135a1b519a67eaac

Observation ccf75546-7565-452c-b4bb-2286973270ba · inbound

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety cites this paper.

Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 104

Resolution
unresolved
no resolver link, observed 2026-07-12T14:28:50.627444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T14:28:50.627444Z digest=sha256:5e2abca7d93fc51605ae69a912da7c00f6800758bd7df5894914c88789d61abc

Observation 121b3da8-53a6-49bf-a5ae-cf67e5e61b74 · inbound

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions cites this paper.

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-09T10:26:11.097166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T10:22:23.782469Z digest=sha256:4266246b34b608e7938bf84b79089bf64decbd35aaa12a2a35fea9b486068621

Observation dde8f715-429a-4a7b-a486-671e594e1b93 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:01.084617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:01.084617Z digest=sha256:aa04864cb482ac5c77830993f0aece0d0c36b788d9bcc2f8aed0515682f11323

Observation 7a179e9a-bb3b-445a-b683-f7a1f9e23786 · inbound

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs cites this paper.

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T16:06:05.960205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:06:05.960205Z digest=sha256:1d121deb2d7b6f4445318d1d14b25e572031d33706b739a4dab7ab19e67b3a39