Pith. sign in

Paper Citation Record · LEDGER

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning

As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2501.07959.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07959 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:35:52.915101Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 51987e9b-db86-4e42-85d4-7ed965eb8c5d · outbound

This paper cites Llama 3 model card.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama 3 model card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.697345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.697345Z digest=sha256:acf4c01565adc25c428e5e759a8fde4dca46215d1c08a66bb05e4ba6be0d617e

Observation 0e399931-dc77-41bf-a8d0-8557088251a1 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Detecting Language Model Attacks with Perplexity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.701802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.701802Z digest=sha256:4b626c99b52d37fd8cdacb06481de5fced4ab8f9ddd53363e9371564f01cd349

Observation 3af79303-da0c-41c9-affa-cbf70b6b7e2c · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.705776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.705776Z digest=sha256:2433ede96dcf638d72f57d63374c5e264b1ea2878fb528642900a7483c3e929e

Observation 094b98a7-515f-4d18-b11a-c305be6ca13e · outbound

This paper cites Many-shot jailbreaking.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Many-shot jailbreaking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.573329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:35:52.710220Z digest=sha256:ffc83c4fa513a5f64fd04a96aafe41d3cf410f8e5f1841632aaa628b9d3c17e8

Observation 6bb033e1-ce8e-48c9-a83d-809d6234bd9a · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.714194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.714194Z digest=sha256:06872e130e8a11b639d702475f69f696a6616869fb246dd595327db8770cb534

Observation b95c695c-de00-4f27-8d69-e1594b0a73d9 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.719204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.719204Z digest=sha256:32e52a70d2e6b1e25aa6ce64721446ef0059609465eeb3ac844fa4d010137d57

Observation bb0f603f-c2e7-4512-bf45-3b8ccb3bf130 · outbound

This paper cites Language models are few-shot learners.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.724012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.724012Z digest=sha256:843c4287572b3633385f1e70dea08844c8471afda9c88ebef33a828d41fb61d6

Observation 62be13bd-c7b2-43f6-9867-100bd8c422fa · outbound

This paper cites Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.727788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.727788Z digest=sha256:c218e43d78e8df540ec8d266e6e980249e1708ab741c2691882138425b1e5e31

Observation ddfa3eb9-4a42-48a3-abe4-935f61abb4e2 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.733088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.733088Z digest=sha256:a1e3f7ccab2eb4fe5db6ae165fe2be8e5fd507f17d8a1727b761c256a30dcf32

Observation b6413099-72d0-4e21-b4d8-71bf0ae35d1e · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.737657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.737657Z digest=sha256:eaab2fb8c67e3811d74d043460459d31e1e4c846485dc91a3e1ca66272a2d74e

Observation 2138c5fa-c552-429f-a962-5fed6f1913ae · outbound

This paper cites Combating misinformation in the age of llms: Opportunities and challenges.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Combating misinformation in the age of llms: Opportunities and challenges

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.742128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.742128Z digest=sha256:54ee7521852eebe04d191f4ee0d5354dbee967910267e02a31d90f363add7c0b

Observation 3d825ee4-f14e-40cc-a7dd-8430b73b6ff0 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.746178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.746178Z digest=sha256:1f6a0d0c9269f0d84d5c42c65cfde7471cf5e69cd884158d152955b33003918a

Observation 5a5e26a7-1a84-457a-9088-1aecc83d8781 · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Multilingual Jailbreak Challenges in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.750811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.750811Z digest=sha256:0f70afaa7fa49e470116fc5e46de4f17b01ca0c03c035616b6bf601a5d83aef3

Observation 7a878be1-1938-426c-8f49-4cc7ed5d7ebb · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.755719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.755719Z digest=sha256:282362054a948fa74fb7e8e724bcfe5ba6d3dc35f87a42268eaf121a52f46817

Observation fc414b65-0879-4503-8536-00c59249ef86 · outbound

This paper cites The Llama 3 Herd of Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.760004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.760004Z digest=sha256:e9a37b275edf7a7a43752d0b91c490bd088a99fa295cf0826587712bdba5171a

Observation bc2925ad-506c-4846-a1f8-ed75d3afb945 · outbound

This paper cites Red- teaming for generative ai: Silver bullet or security theater? In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 421–437, 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Red- teaming for generative ai: Silver bullet or security theater? In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 421–437, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.543563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:35:52.763569Z digest=sha256:0a82508cb71ed6817656cbf2fcdf6547f6a9e9a30b78c3f8aed211cc3a76e52a

Observation 86a9f30c-9aee-4579-b0b6-5d765cd04cf8 · outbound

This paper cites Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b, 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.531168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:35:52.767401Z digest=sha256:8187c5cb5ea69ae1c360293bc3b7d750fb717ba0a97ae8f47c1d1be5ea140c54

Observation 4aebbcbf-ea7f-4815-80f5-81b755a6aee6 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.770507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.770507Z digest=sha256:754afeb4bbf7746fa8b6bf0eb28a4865d60502c3a5e8af87b614233c9eb60d3a

Observation 1a230f98-a27f-41a7-9c84-bd66a97c4009 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.774210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.774210Z digest=sha256:dafe57635c1592d09c7b219a9710b8b413dc1bea9ea9863ec6a1b681707bf349

Observation 641c1b60-bddc-48c9-aeeb-6d8c3eeda8db · outbound

This paper cites Perplexity—a measure of the difficulty of speech recognition tasks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Perplexity—a measure of the difficulty of speech recognition tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.517940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:35:52.777767Z digest=sha256:8391fe20ff97bcfaa56780bef6c5ae34823120be39d523a956354d4d118fd595

Observation b52040c0-11c3-4c80-9989-2315068e87fa · outbound

This paper cites Mistral 7B.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.781447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.781447Z digest=sha256:871d845bdb58278342452bb1a6723991f535a42e5c2ab0219f9fb49a3c0d7312

Observation a1af48f6-8b91-426a-b038-02fffdf6b462 · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.785478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.785478Z digest=sha256:6f2876baee879ed85bd925b039bbae4327a1f9ab08e12bd5eae70e4c03be2dba

Observation b5580c03-ad9d-4d82-a648-b82d36d9a824 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.788923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.788923Z digest=sha256:2331991757be3c5cf895d4ffe95d4c319ec4c2996b1dbdcc4044fa458372deae

Observation d31ed7c8-b58b-43ea-af7d-f97174656187 · outbound

This paper cites Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.792439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.792439Z digest=sha256:a38ed1a44efc6dc11bdc30bb8e57332ca218aed940ade02e07416e0e01c9feb4

Observation e59b2646-f26c-4178-a01d-b79d06bbc569 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.796233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.796233Z digest=sha256:07d98fb920058cec235a4fcccdf13d3a7960aa34249d9d09464247cf3d3b9454

Observation 774b9cd4-2bb3-44be-9d77-7e489c0def0f · outbound

This paper cites RAIN: Your Language Models Can Align Themselves without Finetuning.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning RAIN: Your Language Models Can Align Themselves without Finetuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.799789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.799789Z digest=sha256:c660b730c02b29e4d0b897c61910f556781c98b998959f76c601f45eda4e3b60

Observation 6d3a996f-7d66-422b-affe-4e8ac9491e96 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.803418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.803418Z digest=sha256:adf141c8023078f624a05ee9a9fe65566c1cbe320c611c4506b5771dc02ae3bd

Observation 5c6e3624-1cf6-460d-a82e-630435d229c8 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.807650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.807650Z digest=sha256:386eb8b6ba0b5e2e75361c650d1f54271e29c7e1800981633c9f33459fb8c9c3

Observation f52b7b31-46da-41ec-815f-3d1ac87b91e2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Training language models to follow instructions with human feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.811437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.811437Z digest=sha256:c015a8eeba90f2c9fc12cc30243e7754f15d1810cda32206430e2b2e7c97cc40

Observation e9d98946-ae4a-4eb0-acfd-02da1a80b42a · outbound

This paper cites LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.814907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.814907Z digest=sha256:8b515c57920396b4e1e555fdd554cf1ec7063c378fa62ea83431db6efaa2355d

Observation ea4c7205-f1b4-4f54-b765-64f8e0121f7e · outbound

This paper cites Universal Jailbreak Backdoors from Poisoned Human Feedback.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.818464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.818464Z digest=sha256:a0c80c60d50bf336b0b18442caf05df146689e6244c65a6bdd28586e7a793d75

Observation 8b8d9d38-736e-4310-821b-d742885177bc · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.822255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.822255Z digest=sha256:bde8f6823960170032855e497a34e908805fce51b28c2a5e5a4be7f7fafc9525

Observation 4146389d-132b-4bd6-8081-1fe955e307fc · outbound

This paper cites CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.825816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.825816Z digest=sha256:242181d209a21d7a28e9f4bfef3b0ae8eaaa61c2453517e331ca51ef59759fef

Observation 44356f20-045e-4da2-9e84-8b14a768a0dc · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.829433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.829433Z digest=sha256:eb56482a5060d7c70169db588fc2d934ee0f9ae953ac8ac8bfb2706fb6f5735f

Observation 33988b2a-c35b-48c8-b3ec-b58bdf2dea45 · outbound

This paper cites do anything now.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning do anything now

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.833240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.833240Z digest=sha256:c99c6f1deb1e2e24d39de771a72df1387cac198ca5eb63ea3fb6d748e6f26a3f

Observation 7c72ef52-306c-49f2-815e-41f77be8cc23 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Qwen2.5: A party of foundation models, September 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.837049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.837049Z digest=sha256:41399d57588f12ff0445ed4ba3a74748f0ba3273ec2ed96dd56a0e43231255a1

Observation ff03216b-92ee-4574-83e4-69f22ea01aba · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.840899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.840899Z digest=sha256:4a5fa47f3c2ac9ef45d07ae011788b5b7ca2236a2f03ad57afb2e43a14730962

Observation eeec823e-f6ed-45c5-9d6c-b5e2ec9ce37e · outbound

This paper cites OpenChat: Advancing Open-source Language Models with Mixed-Quality Data.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning OpenChat: Advancing Open-source Language Models with Mixed-Quality Data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.844524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.844524Z digest=sha256:d4111bc3bc24ad68026515a62cdb6e1cb6078a8274efe41cb85fb8ef2f6c987d

Observation 8a8eaf54-aad6-4d17-9a68-39953e2e431d · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.848329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.848329Z digest=sha256:cd416ceccdefc97b8076728eadd65fea39dd06ed841b6faae4ef05d184385658

Observation 524a50a4-a35c-481a-ac54-2ececa3686b3 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.852850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.852850Z digest=sha256:1b4ee18a053b04e1af8e473ead8b51a1296f4f39a90c0056879cbab091c06791

Observation ed581a41-2e19-4c3a-93a0-50d01a279b19 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.856473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.856473Z digest=sha256:1d0e60fb69f8bb763d027e4766f94c5ae1e372768dd6a28a49685478dce9821a

Observation 1738ef74-4719-47c3-8ad8-c8494ed086b2 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminder.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Defending chatgpt against jailbreak attack via self-reminder

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.477073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:35:52.860142Z digest=sha256:dfc508560e57896f8db4230ebd1d55a7c19d6eceec5d6e9c0cd91cfe8a269f45

Observation 493da253-3d00-4cab-9e27-74a02732220c · outbound

This paper cites Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.863562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.863562Z digest=sha256:9be07854373d215a7ae91feb2d45c3b4ddba31e04e73c0e521fee296f6dc2066

Observation 4383d3df-5b9c-40cb-9135-33ab224320d0 · outbound

This paper cites Qwen2 Technical Report.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.867549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.867549Z digest=sha256:e1b38ca9d820dcf8524e6696128cb99d267abde641686e8e5be917be755adaae

Observation 35f49797-2e6a-4e80-ad93-790aa4ed03df · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.870948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.870948Z digest=sha256:75a9dd6254422f885c7e32a72a5fe7ff0b64439dcf6f728a0e587a51c9860f69

Observation 99b54b90-02ac-41d0-919d-823d6d7cfc14 · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning A survey on large language model (llm) security and privacy: The good, the bad, and the ugly

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.874542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.874542Z digest=sha256:074b9caf647e42f154c35949c30a38a8e520e6df3b75c993e0ab8d89800734e2

Observation e088c2cf-5dc4-4985-8771-c216233db363 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Low-Resource Languages Jailbreak GPT-4

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.877810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.877810Z digest=sha256:c5f4eed6870f4ee29022f98029c23dd0b3ceddbe902b7b3169098dbfe7ee103b

Observation 095ec9a2-ea40-4352-a2c7-27ea35e465f3 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.881332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.881332Z digest=sha256:93faee8bb82e95222a390ec27d87e4eb8b5f1875f9daaa31ca301c7e5ba86285

Observation de2ff848-1cfe-45d1-91dd-fbcea379715e · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.885191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.885191Z digest=sha256:4b4768820752d22d619ff6bd063bba2345f205e08370537cf3d84e43fda03b09

Observation 34f0c01b-75ba-442b-891f-23955059e5e6 · outbound

This paper cites Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.889218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.889218Z digest=sha256:6d28d2bb8ddf05d1106c65aedfd39e5344fd0f693fe35d0b68cfb21d4113752d

Observation 2424d0e2-b828-4159-bed6-f8e2cce7f211 · outbound

This paper cites Diversity Helps Jailbreak Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Diversity Helps Jailbreak Large Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:35:52.997180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:35:52.893040Z digest=sha256:63224fb6f247dbea7092343a5f4470146a6ca083c5f2275e92e1854ec3ea169f

Observation 1f9f5f83-9ceb-4d85-b6b7-63b0cdaa32b1 · outbound

This paper cites Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.898323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.898323Z digest=sha256:748fe6ac68ac8789bc0988dca4899b9a2a32a31a38689d1fa6b9fae7724bd05b

Observation ed8e2339-b36c-4fba-a62a-9510ab9119d9 · outbound

This paper cites Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.902122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.902122Z digest=sha256:4d395dced69193b00b4d22c5da7dd3cc67f62cbcb74377f53e0344d8a1f40770

Observation b9476d38-a2a4-4bdb-b25d-3c64abba7121 · outbound

This paper cites Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.455117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:35:52.906616Z digest=sha256:aa79bc3ffb0e4a5176c54671ff7b89836f8daa7a9d956dadd9e4dc5219f77f23

Observation dea701a4-d2cc-4c57-96ba-73755f0076f1 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.910697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.910697Z digest=sha256:7daced72a0da6217a9b1c8bf8228e8c2ee1a0084bbb139d46ded0fb65cafda7d

Observation 47c3d999-626e-41cc-8893-c678846a02e7 · outbound

This paper cites As shown in Table 10, our method can still achieve remarkable performance on HarmBench [28].

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning As shown in Table 10, our method can still achieve remarkable performance on HarmBench [28]

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.439187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:35:52.915101Z digest=sha256:5231badb84ebfab0da314fd597ac3e98609bbe0fd69f8d1a43a4dd92868e287b

Pith citing papers

No inbound Pith citation observations are available.