Pith. sign in

Paper Citation Record · LEDGER

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning

As of 11 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2501.07959.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07959 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:35:52.915101Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 51987e9b-db86-4e42-85d4-7ed965eb8c5d · outbound

This paper cites Llama 3 model card.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama 3 model card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.697345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.697345Z digest=sha256:2291dabaf4f7af2fc8ed7bcdcdd6703121312374bbc3dae9ccf52ab56835998d

Observation 0e399931-dc77-41bf-a8d0-8557088251a1 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Detecting Language Model Attacks with Perplexity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.701802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.701802Z digest=sha256:ed1e455860b1bca3ea3690e368322b59c2cac4afb4e1c1df276b080a22ff9dc7

Observation 3af79303-da0c-41c9-affa-cbf70b6b7e2c · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.705776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.705776Z digest=sha256:6bf328169ed4e8d24a9458439b54de1c28747f725b022ee673b9f4f24e3a57f3

Observation 094b98a7-515f-4d18-b11a-c305be6ca13e · outbound

This paper cites Many-shot jailbreaking.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Many-shot jailbreaking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.573329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.710220Z digest=sha256:8f469f8f2323b060dabc943e6dc8fb069242682759b0d4a0edb5122918a300bf

Observation 6bb033e1-ce8e-48c9-a83d-809d6234bd9a · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.714194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.714194Z digest=sha256:8bc2317514fd55e72888de961338a54fbed17410aaa0c2d5ed8e08280b7198bd

Observation b95c695c-de00-4f27-8d69-e1594b0a73d9 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.719204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.719204Z digest=sha256:c48cb60860e9afb2a0784ac0d4ad5089cf15bef772bdcca291e5f6329cc19baf

Observation bb0f603f-c2e7-4512-bf45-3b8ccb3bf130 · outbound

This paper cites Language models are few-shot learners.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.724012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.724012Z digest=sha256:faf60f2193f0b762575e04971c07fda1e2aa91c6210cd1f456d1cdd4f6b79edb

Observation 62be13bd-c7b2-43f6-9867-100bd8c422fa · outbound

This paper cites Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.727788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.727788Z digest=sha256:28426e03505ffc050c9848fcc227258fc89bd719e497a66aa16c38e6f9550214

Observation ddfa3eb9-4a42-48a3-abe4-935f61abb4e2 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.733088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.733088Z digest=sha256:82e17d0efaa0d2afcf7cf2a663ad965b0b137d7bc9331c0653693a1f75acbc7c

Observation b6413099-72d0-4e21-b4d8-71bf0ae35d1e · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.737657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.737657Z digest=sha256:dc62daaaffd9ed1a74511b4c1fcbc09f7df35e963041ed3d848f4aa115328ea2

Observation 2138c5fa-c552-429f-a962-5fed6f1913ae · outbound

This paper cites Combating misinformation in the age of llms: Opportunities and challenges.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Combating misinformation in the age of llms: Opportunities and challenges

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.742128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.742128Z digest=sha256:f429af84988bd8d967b52cbdccd3b4711eb069c24187dbb1f5a347fa7012a213

Observation 3d825ee4-f14e-40cc-a7dd-8430b73b6ff0 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.746178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.746178Z digest=sha256:3a64214dfeaefb380083d0f201f78cf7fa6abd06f19fd5461fef0b5db1dff101

Observation 5a5e26a7-1a84-457a-9088-1aecc83d8781 · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Multilingual Jailbreak Challenges in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.750811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.750811Z digest=sha256:e51f3ae209963dab00ab1d061b8abf837385c948065d6db22d93ba0edcdff58a

Observation 7a878be1-1938-426c-8f49-4cc7ed5d7ebb · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.755719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.755719Z digest=sha256:9f022333a070b393cf7a57a0f7ff6ae81b88d1fde8a16165ebc7bbe9cd8ec46e

Observation fc414b65-0879-4503-8536-00c59249ef86 · outbound

This paper cites The Llama 3 Herd of Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.760004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.760004Z digest=sha256:6c9824aebeae6511a84bfe8017b1425250ca91e18b7e24b7b9a6a42c17713a03

Observation bc2925ad-506c-4846-a1f8-ed75d3afb945 · outbound

This paper cites Red- teaming for generative ai: Silver bullet or security theater? In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 421–437, 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Red- teaming for generative ai: Silver bullet or security theater? In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 421–437, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.543563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.763569Z digest=sha256:0497991986d3cad27efbc62f1ae8736f686e5c6d4e9d0a8576419d27d0f9413d

Observation 86a9f30c-9aee-4579-b0b6-5d765cd04cf8 · outbound

This paper cites Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b, 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.531168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.767401Z digest=sha256:94cffd8571bd0f044ea02f64538b987d11c3ef6a51dc1991532cdd0ef35e31b0

Observation 4aebbcbf-ea7f-4815-80f5-81b755a6aee6 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.770507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.770507Z digest=sha256:f9310a24f879fe5cf7d409efa226395888c65b05a58f87e739f02790781c2550

Observation 1a230f98-a27f-41a7-9c84-bd66a97c4009 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.774210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.774210Z digest=sha256:85e30aba8bf80796d6db66c7ceb58260bd886e9f9632eca046ef8b8698d7daff

Observation 641c1b60-bddc-48c9-aeeb-6d8c3eeda8db · outbound

This paper cites Perplexity—a measure of the difficulty of speech recognition tasks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Perplexity—a measure of the difficulty of speech recognition tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.517940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.777767Z digest=sha256:52ceaa20cbc615183ba0de1ff232d32f787b70ffe64e771c4d2b9b4704ab8db5

Observation b52040c0-11c3-4c80-9989-2315068e87fa · outbound

This paper cites Mistral 7B.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.781447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.781447Z digest=sha256:6542ab0d9c69780a93f4f89df84182601db7491ade7ade17343135fed9fa146d

Observation a1af48f6-8b91-426a-b038-02fffdf6b462 · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.785478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.785478Z digest=sha256:0a6b93baed103ee8991db3ee3a6df649d5a94c42a22ae3dfe4983c94aae6a982

Observation b5580c03-ad9d-4d82-a648-b82d36d9a824 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.788923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.788923Z digest=sha256:445d8ed5678b83cba142b9fa987c532457df3b32a4bc1e6804ffdb53b6d45a5c

Observation d31ed7c8-b58b-43ea-af7d-f97174656187 · outbound

This paper cites Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.792439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.792439Z digest=sha256:0b5d8fb41c480fb32aa2e332dd72e468404d27890cf539c157dba1b8b102b618

Observation e59b2646-f26c-4178-a01d-b79d06bbc569 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.796233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.796233Z digest=sha256:f6722e5c29e0376aa5653cc3984d75c866b09bcfaf86b0a5ea3158deafe06b65

Observation 774b9cd4-2bb3-44be-9d77-7e489c0def0f · outbound

This paper cites RAIN: Your Language Models Can Align Themselves without Finetuning.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning RAIN: Your Language Models Can Align Themselves without Finetuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.799789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.799789Z digest=sha256:8dd6e239486cf50f0411ac30cc095c03ddf555866b224ed7d8e9b455af320921

Observation 6d3a996f-7d66-422b-affe-4e8ac9491e96 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.803418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.803418Z digest=sha256:fbdf59e101d4ca9d547c16c713ce9c05b33c680f2c813508e6a91e5a1ce7e4a7

Observation 5c6e3624-1cf6-460d-a82e-630435d229c8 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.807650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.807650Z digest=sha256:0aab5485423e12ec3dadba278787e3c511c33dede5a9b3e506e352b1076924a7

Observation f52b7b31-46da-41ec-815f-3d1ac87b91e2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Training language models to follow instructions with human feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.811437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.811437Z digest=sha256:b1a8361b99df4d7a0f9380df79b94ad67da3eef25aac71ee8ae7512ba9aa53ab

Observation e9d98946-ae4a-4eb0-acfd-02da1a80b42a · outbound

This paper cites LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.814907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.814907Z digest=sha256:97e2cee74fad1cec98f8439534e2612e516cfb77f13c4791e6ec8021ac5cd25f

Observation ea4c7205-f1b4-4f54-b765-64f8e0121f7e · outbound

This paper cites Universal Jailbreak Backdoors from Poisoned Human Feedback.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.818464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.818464Z digest=sha256:d4db39a507a7ae5a77c0cd0422d99cf44b81eb2957fa27c254da6a14d968d398

Observation 8b8d9d38-736e-4310-821b-d742885177bc · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.822255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.822255Z digest=sha256:dbb59ca78f9cca93d3a7852c076f8c79f6d6a102d1436c088c8db7c6111b10eb

Observation 4146389d-132b-4bd6-8081-1fe955e307fc · outbound

This paper cites CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.825816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.825816Z digest=sha256:ff7f366fcba5d878fb3b67d6cc5e58b387c01d6b848455141cf8aa740236ff25

Observation 44356f20-045e-4da2-9e84-8b14a768a0dc · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.829433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.829433Z digest=sha256:9a515dc04b65f9a0ccad07ef8fb1d9c1e75c20f6c1f5af5fa777ed1a642443ce

Observation 33988b2a-c35b-48c8-b3ec-b58bdf2dea45 · outbound

This paper cites do anything now.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning do anything now

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.833240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.833240Z digest=sha256:561b684e93444f82d161cf704cd69538b389cde2cb77d81c650f59978fdb7217

Observation 7c72ef52-306c-49f2-815e-41f77be8cc23 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Qwen2.5: A party of foundation models, September 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.837049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.837049Z digest=sha256:7733ff5ea1b07d7da8f5d47f2da538466cfeee947122ca02184c562d5451e87c

Observation ff03216b-92ee-4574-83e4-69f22ea01aba · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.840899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.840899Z digest=sha256:96b397d9e94ca3b1cb17d5045de5205fa983b2845f1053b5db969beb1ba6a0eb

Observation eeec823e-f6ed-45c5-9d6c-b5e2ec9ce37e · outbound

This paper cites OpenChat: Advancing Open-source Language Models with Mixed-Quality Data.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning OpenChat: Advancing Open-source Language Models with Mixed-Quality Data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.844524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.844524Z digest=sha256:562d7e1f512ca8485776e1840a71afdfe8b1d3b8cee054b0d538655134e106ce

Observation 8a8eaf54-aad6-4d17-9a68-39953e2e431d · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.848329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.848329Z digest=sha256:f76401cfded265af794f8c0f6b5a71f0b7fc93209c547c72bac1acb72805ddb3

Observation 524a50a4-a35c-481a-ac54-2ececa3686b3 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.852850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.852850Z digest=sha256:76732d4e24bbffc8d1b351e1ab69697db8691c1766f18be287ca7d28391f71a5

Observation ed581a41-2e19-4c3a-93a0-50d01a279b19 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.856473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.856473Z digest=sha256:85a866de7f145299d2fadb8b3797ec3b74d457546e04a844d9d5c24b1ffdaf47

Observation 1738ef74-4719-47c3-8ad8-c8494ed086b2 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminder.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Defending chatgpt against jailbreak attack via self-reminder

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.477073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.860142Z digest=sha256:61aaec6306a95ac7d7aed2f28931569792cde503845efedc264c2088237e09e5

Observation 493da253-3d00-4cab-9e27-74a02732220c · outbound

This paper cites Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.863562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.863562Z digest=sha256:db04b94f8dda7dd000fb8993394c7b7148807d97a86e39ea5343b6fefa504865

Observation 4383d3df-5b9c-40cb-9135-33ab224320d0 · outbound

This paper cites Qwen2 Technical Report.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.867549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.867549Z digest=sha256:49239cd7c068cea1972e7578b3a54e080995a97e49a0cf13a2a358afa1e80d51

Observation 35f49797-2e6a-4e80-ad93-790aa4ed03df · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.870948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.870948Z digest=sha256:28bdda95324560be6fb875dfee99978ba1314c36b2ddacabd263f18afe1b28c4

Observation 99b54b90-02ac-41d0-919d-823d6d7cfc14 · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning A survey on large language model (llm) security and privacy: The good, the bad, and the ugly

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.874542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.874542Z digest=sha256:2606197c396dd4f45dd108d94075b3eaf10c2424ceb74cccb0b9e6ca54218743

Observation e088c2cf-5dc4-4985-8771-c216233db363 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Low-Resource Languages Jailbreak GPT-4

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.877810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.877810Z digest=sha256:14f7eca3c29658c7141076e2a8ed78d894e1c951458190d963e964ba4a4660d4

Observation 095ec9a2-ea40-4352-a2c7-27ea35e465f3 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.881332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.881332Z digest=sha256:0856b9d9c91d6ba1fd71ecf5be8935dd41aa19476b58f8042dd4dc5bbd87c1f1

Observation de2ff848-1cfe-45d1-91dd-fbcea379715e · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.885191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.885191Z digest=sha256:8c64c317b950ed7b91ae2c7334f72bd1a57620077dc83245754a3a87b5fa2628

Observation 34f0c01b-75ba-442b-891f-23955059e5e6 · outbound

This paper cites Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.889218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.889218Z digest=sha256:393659a959c8a81bf1c33945b81fd0d510b55754d5c66b4c69a8ef65589ff4bb

Observation 2424d0e2-b828-4159-bed6-f8e2cce7f211 · outbound

This paper cites Diversity Helps Jailbreak Large Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Diversity Helps Jailbreak Large Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:35:52.997180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.893040Z digest=sha256:1f7942e2c80610370a365e72668d0c1dfbb70691c5aae621974d6489dee8bb3a

Observation 1f9f5f83-9ceb-4d85-b6b7-63b0cdaa32b1 · outbound

This paper cites Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.898323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.898323Z digest=sha256:485877ebb5367d972f294ad0e6b9e2b7254fb54b0927eeb4e9d8a5d110fc08da

Observation ed8e2339-b36c-4fba-a62a-9510ab9119d9 · outbound

This paper cites Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.902122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.902122Z digest=sha256:b8aff9ab3465c5c133183e9f7e7e14b4cdfbc13e6c73e3b41fcabd919c8f288a

Observation b9476d38-a2a4-4bdb-b25d-3c64abba7121 · outbound

This paper cites Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.455117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.906616Z digest=sha256:bdea1cba98b79b3a7f9a8f13b1ce8f1446ba594eaad72e320ceac5be83e6e17a

Observation dea701a4-d2cc-4c57-96ba-73755f0076f1 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:52.910697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:52.910697Z digest=sha256:b286f69f6ae0d9d2d21d23867e58467416cec7193b371f62d878d29c5edb6a90

Observation 47c3d999-626e-41cc-8893-c678846a02e7 · outbound

This paper cites As shown in Table 10, our method can still achieve remarkable performance on HarmBench [28].

Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning As shown in Table 10, our method can still achieve remarkable performance on HarmBench [28]

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:35:53.439187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T20:35:52.915101Z digest=sha256:739a54fe4fab7f1b88c49d9fc02e4098ad18f72df71534ec247c26724bc27fde

Pith citing papers

No inbound Pith citation observations are available.