Pith. sign in

Paper Citation Record · LEDGER

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling

As of 21 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2502.01925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01925 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:02:38.960422Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:17:34.283917Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation edc6b069-b5ec-4df6-b96a-ea72e66ed362 · outbound

This paper cites write newline.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.672091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.672091Z digest=sha256:aa7d917c00024da3a2689c4c2d13dd462bb3e6a4416ce5e19b5527403245088e

Observation e9a89191-3d1f-429b-8239-d140d4c4eee0 · outbound

This paper cites GPT-4 Technical Report.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.679096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.679096Z digest=sha256:2392404ec56560f6b5b602c831afb2f636a4f940989d8b33fccc88d70c34a494

Observation 0af1b2ab-2c22-4241-bb72-d7ed61f1de33 · outbound

This paper cites What learning algorithm is in-context learning? investigations with linear models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling What learning algorithm is in-context learning? investigations with linear models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.898135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.685122Z digest=sha256:f94f4587e479e10e5c9f4e249db0a14ed2da7463e6214dd035ea977867536f79

Observation c4799bcd-3113-4929-bbf6-ef61d02c8f81 · outbound

This paper cites Jailbreaking leading safety-aligned LLM s with simple adaptive attacks.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbreaking leading safety-aligned LLM s with simple adaptive attacks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.881193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.690848Z digest=sha256:2376dbac68d7644fffaf9a02a08f76abe44ee71bd3bca1316c788293e7f94f25

Observation 810a0264-d52d-464a-bb96-7e44d721b7d6 · outbound

This paper cites J., et al.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling J., et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.865038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.696196Z digest=sha256:9f7152b99328220a72f9ea86c46e327682c93cecafc3b07f6600d06a39ece782

Observation 0f58597b-951d-43b5-b104-348bb47970ad · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.701627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.701627Z digest=sha256:553cf33632e0555608dd8075f373b7abde5c2f22d84e3939a6c9f125763dca11

Observation 9a4038b2-7caa-44c7-8fa2-9afbbd46c36a · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.707132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.707132Z digest=sha256:b77572d2c3464f53f06e164b7edf211b48ea49b2dcb0f7e91555836140241300

Observation 2076e417-f77e-4b9d-abfe-e7389033a782 · outbound

This paper cites J., and Wong, E.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling J., and Wong, E

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.838867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.712934Z digest=sha256:051340670231952ab0e85c5bc8eaae0f3656e0076d6c9d1e1ce4e468dc155523

Observation 53ca5892-b4e7-441f-9fb9-500c604c1d03 · outbound

This paper cites How many demonstrations do you need for in-context learning? In Findings of the Association for Computational Linguistics: EMNLP, 2023.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling How many demonstrations do you need for in-context learning? In Findings of the Association for Computational Linguistics: EMNLP, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.823673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.717968Z digest=sha256:7cef9a24b497d1abfa27a7b89975f32a4a18831d0c30dc5c80937990f4b0c72e

Observation 1dba753b-4f46-419f-afa2-0d2c5c3e4f2c · outbound

This paper cites What does BERT look at? A n analysis of BERT ’s attention.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling What does BERT look at? A n analysis of BERT ’s attention

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.808177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.723140Z digest=sha256:a66163d81ed2f47767cfbbac424de0bd4e958954e441d39326038f2258b322df

Observation 41cf55af-ddbc-455e-b3bd-97318fcc5b60 · outbound

This paper cites BERT : Pre-training of deep bidirectional transformers for language understanding.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling BERT : Pre-training of deep bidirectional transformers for language understanding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.792042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.728274Z digest=sha256:2486b6b4652966912a947920cf866fd7a03c978fa3c522d309960fe5362dfed8

Observation c596d65e-0b2a-41ed-a8ca-41fe1030c439 · outbound

This paper cites L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.774919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.734484Z digest=sha256:21eca004eddb32ed2269711bd6f60d909caf4d40e954c85c03bbb3b6591798f7

Observation 7548d31d-efd0-4ff1-936e-a3329d3fc9cc · outbound

This paper cites X., Wang, B., Tian, Z., Chen, W., and Wen, J.-R.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling X., Wang, B., Tian, Z., Chen, W., and Wen, J.-R

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.759338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.739622Z digest=sha256:fbadfb8d385ea9b5f8d4ffd6d66325838491e5a27cbc4b26e6b31c1896b81105

Observation 6b652bf3-e455-4335-9306-f8562761a66a · outbound

This paper cites The Llama 3 Herd of Models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.744562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.744562Z digest=sha256:97cfaa1ddebb75f4bea79e3d4b9b576da807f832e4603d8043ac144c70c608ec

Observation 339f643a-586c-477f-869c-4a27c619d4d3 · outbound

This paper cites Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.744028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.749852Z digest=sha256:cdca298c45e090fba0415ba5756e2bfe76ba55f661efeeb10b568d1729025da6

Observation 1a8bc742-36e0-4988-9b6b-e375a885ad6c · outbound

This paper cites and Das, K.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling and Das, K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.728024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.754667Z digest=sha256:238aad63e7174b5e6f2de9ccb1fd4917debcf89c1c27ab60068875db20aa8aed

Observation cf188ecf-51e2-484e-92bd-a5e0721bc083 · outbound

This paper cites ChatGLM : A family of large language models from GLM-130B to GLM-4 all tools, 2024.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling ChatGLM : A family of large language models from GLM-130B to GLM-4 all tools, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.712294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.759412Z digest=sha256:bfa3f4b5939658fe9fb2005585775b15478491e30969b243dc1ca89270b7af4e

Observation 3540208b-63db-4f9d-b6f2-6aa9e12ad817 · outbound

This paper cites Comparing results of 31 algorithms from the black-box optimization benchmarking BBOB-2009.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Comparing results of 31 algorithms from the black-box optimization benchmarking BBOB-2009

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.696549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.764123Z digest=sha256:f8e00edcc53e457105aaca34967f6a8236f82e254bba0636bab4b57e9ee541a2

Observation 78bba061-cf75-4aec-9343-966316fe9e89 · outbound

This paper cites Self-attention attribution: Interpreting information interactions inside transformer.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Self-attention attribution: Interpreting information interactions inside transformer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.680842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.769107Z digest=sha256:f425a41ad926eb38b42b362de027e52c68179806b513b616a1cf645762c1c0a3

Observation 573ae094-0bdb-4b6a-aae0-1f0ac36359fa · outbound

This paper cites WizardLM-13B-Uncensored , 2023.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling WizardLM-13B-Uncensored , 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.665681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.774170Z digest=sha256:0fbe5fe537d38a7e363cce30be7650c210dd96be17ba27763a828f2d45f1648a

Observation 1e508619-bdbc-4886-be77-fb0afa741acc · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Measuring mathematical problem solving with the math dataset

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.650630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.779314Z digest=sha256:887c1ccb79d78523f1d9451df6022eb2df0697b6c23689f9386a71961fc896a1

Observation 2d7f3953-7880-46eb-8df1-0541ac21f691 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.784054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.784054Z digest=sha256:aa0816d7678f8555b55d52bbcaf2c6af95f57b8417ab952c00e841714069a93a

Observation 06f05d7f-5f69-49ce-b5fb-96df45cbc25e · outbound

This paper cites LLM maybe LongLM : Self-extend LLM context window without tuning.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling LLM maybe LongLM : Self-extend LLM context window without tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.634246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.789233Z digest=sha256:02a5c76fc1f7336a992d99da8cf66f445e8a6cf7d8cca960b2ebb9c1361ba7d9

Observation e60d5aa6-b1e3-407d-a9b8-d98f36562f04 · outbound

This paper cites an unresolved cited work.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:02:39.618661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.794558Z digest=sha256:244cf99d75f1dcb835d28bc33cd43cf30a45b60bade419a01dc97d0d089cd336

Observation c034bce8-ca56-49ff-b366-31cf55b5cc2c · outbound

This paper cites Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.603277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.799904Z digest=sha256:8c49213bbbd16815ad2a9b29b5d4d2251175c44b749a053a2387648c92eb7f67

Observation 2ba4cd55-4d1e-48f6-a4c9-2645fcc1d322 · outbound

This paper cites Attention-enhancing backdoor attacks against BERT -based models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Attention-enhancing backdoor attacks against BERT -based models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.587553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.804891Z digest=sha256:2a2e12196b266ee542015b11dd649c766632430fae14f4c201ff27fad144fd2f

Observation a2b9f54d-e71d-42f1-9f8a-89a5fbe7ebb7 · outbound

This paper cites Harmbench: A standardized evaluation framework for automated red teaming and robust refusal.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Harmbench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.571273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.809842Z digest=sha256:b1ec82b6661ddd8495b2d956820bae521dff3dfef2dae61e9e5904e4229bcbee

Observation f88212ac-0e77-44f2-8d9f-80d82fd105a0 · outbound

This paper cites Tree of attacks: Jailbreaking black-box LLMs automatically.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Tree of attacks: Jailbreaking black-box LLMs automatically

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.556180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.814969Z digest=sha256:9281cacaf34a29ba6f5302f573707d818dfb28f59eea4e0cbfe50dbb793abf96

Observation e980d4f3-c52f-46e3-ae06-2e25889085ca · outbound

This paper cites Bayesian Optimization : Open source constrained global optimization tool for Python , 2014.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Bayesian Optimization : Open source constrained global optimization tool for Python , 2014

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.540977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.819627Z digest=sha256:74acbd2b46d8f3b2b827cee89cacfd885463af7372d782639fd26ad92c39b196

Observation 08fe398d-46fd-4119-8b82-7e64b8ea9cb2 · outbound

This paper cites 2 OLMo 2 F urious.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling 2 OLMo 2 F urious

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.525142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.824402Z digest=sha256:e67e7ce6e847459d2e87d473b3a25d95c93bd199c38423f46e0cfca0653cadb8

Observation e3e47036-ad29-44a4-9247-54e72e8d252c · outbound

This paper cites Training language models to follow instructions with human feedback.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Training language models to follow instructions with human feedback

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.507753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.829040Z digest=sha256:1048476456aa97f5ffe816d38f17ca223c2586e51ee93d49b971d6fd0153c3eb

Observation e605e30b-a35e-4755-b477-4def7886fac0 · outbound

This paper cites S., Soltanolkotabi, M., and Thrampoulidis, C.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling S., Soltanolkotabi, M., and Thrampoulidis, C

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.491573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.833730Z digest=sha256:2c1952db2324c83900e88fab5c582357e970bbce20819f0ef7790a0d0af2bc20

Observation 8800b52c-2221-4ea7-9382-7e70755cef03 · outbound

This paper cites S., O'Brien, J., Cai, C.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling S., O'Brien, J., Cai, C

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.475920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.838426Z digest=sha256:116faf5b125699bfa1f5bb40e3af3e6c627d3d5c797cd3a9c57ed3b6c5fa08a3

Observation 6a9f2eb1-747e-4c56-b095-68e9a8541ede · outbound

This paper cites Red Teaming Language Models with Language Models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Red Teaming Language Models with Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.843272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.843272Z digest=sha256:467cf59d08c7a504a6c5ca44b8b9432603fcaf20087754fb322c19856c6c4a92

Observation 8f6256a7-e521-4e52-9c58-7aa146b45b97 · outbound

This paper cites Baitattack: Alleviating intention shift in jailbreak attacks via adaptive bait crafting.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Baitattack: Alleviating intention shift in jailbreak attacks via adaptive bait crafting

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.459923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.848658Z digest=sha256:0f7caa0eeea882d88e9d436a716c3d83425fd1f6559f98a437f7271973cf3a58

Observation 07ef3fc9-f027-4e17-9387-645ed9bd8f2a · outbound

This paper cites and Barez, F.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling and Barez, F

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.443220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.853805Z digest=sha256:48400b425ef656268a244287ed67a5ccaeaf31a6b28a7cf26b540928f94772bc

Observation 19d42300-af06-4934-87e9-aa127ecc40df · outbound

This paper cites Language models are unsupervised multitask learners.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Language models are unsupervised multitask learners

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.858787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.858787Z digest=sha256:ccdb821e4ecdbaf873c29eb37514b6edba2ab1e937e6de5ca465860018bdcca2

Observation b3127a67-72f4-42e9-827e-fbd0049284eb · outbound

This paper cites Tricking LLMs into disobedience: Formalizing, analyzing, and detecting jailbreaks.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Tricking LLMs into disobedience: Formalizing, analyzing, and detecting jailbreaks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.416559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.863659Z digest=sha256:5808e0a3d28d74e98850b791f14ce0b98d0a4cd7617ecaa46bbc48bd3cc87c09

Observation 662cfdae-b896-4968-9704-422ae6d95df7 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.868276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.868276Z digest=sha256:49dc3b3fbd395a14cbb91892903a8d66e6b96ac75a01068ac3574a9b33754cf2

Observation ad5cd759-6940-416e-a69a-fe7683933d1c · outbound

This paper cites P., and De Freitas, N.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling P., and De Freitas, N

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.399652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.874036Z digest=sha256:288037e26e83d741a03f829b7081d08430554e672ab2d4ebde7cb497a732be82

Observation bc2101dc-cbca-443c-855f-1461c93cb6be · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.879189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.879189Z digest=sha256:1e35fe8a83a8887e85cfa7a0793760ec98e249d73621553bea397d27cdf81ebe

Observation 2e9a0925-b676-46d4-a4dc-644c39dc8f0a · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Qwen2.5: A party of foundation models, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.382828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.884299Z digest=sha256:7ee441d045c08fddc07c3d895ff509c3ae8cd4884a8f074930ce38d8e1fff4da

Observation 917cc9bb-3b75-4789-9ed4-6272e02ba0ae · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.889082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.889082Z digest=sha256:2e9c733b4b64c83e69e88f7138d5809a046293d36970c450e560cba7d020367a

Observation 8a1cd8ae-a4d1-42e7-aa71-d651bffd4ff7 · outbound

This paper cites Bayesian optimization is superior to random search for machine learning hyperparameter tuning: Analysis of the black-box optimization challenge 2020.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Bayesian optimization is superior to random search for machine learning hyperparameter tuning: Analysis of the black-box optimization challenge 2020

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.366481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.894328Z digest=sha256:72dbae806dcd53e5537fd4677951ef3a0d816a2bca54376666ead4fe077d7d84

Observation 806772cf-42de-477b-9c5a-e628e8470ee2 · outbound

This paper cites Attention is all you need.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Attention is all you need

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.350820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.898986Z digest=sha256:468693b2cea7a9275d17c9011c499b913d7f7779e0622a3eedc1adbdea458bbd

Observation ab6a9ead-82e5-4605-8df1-46eb9577f14c · outbound

This paper cites Jailbroken: How does LLM safety training fail? In Advances in Neural Information Processing Systems (NeurIPS), 2023 a.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbroken: How does LLM safety training fail? In Advances in Neural Information Processing Systems (NeurIPS), 2023 a

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.335411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.903722Z digest=sha256:ac7ad2bbf48c9e65ee156a6fa04e5afa17b82854525c5a0b851500edae424eff

Observation a928bc4c-1dd2-43bd-862a-186b3966b4c8 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.908552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.908552Z digest=sha256:315855d71a565bb4c95449c5a3902568bc7f51231f74b6ab40915842a1d25c14

Observation b7cbf08a-857b-48a3-99eb-54fc721163eb · outbound

This paper cites Never miss a beat: An efficient recipe for context window extension of large language models with consistent ``middle'' enhancement.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Never miss a beat: An efficient recipe for context window extension of large language models with consistent ``middle'' enhancement

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.319450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.913719Z digest=sha256:a67fb2f272e4339815a7764b86cd07ca84271e4d68e1281bc005592841ca877b

Observation 6a37ad53-bd30-4d3e-a0c9-25f8077beb81 · outbound

This paper cites Distract large language models for automatic jailbreak attack.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Distract large language models for automatic jailbreak attack

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.302514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.918497Z digest=sha256:dcc2ec699b6b08fc695ac827e35c42afccdb308a75ac46c154326082d7857d26

Observation 1ad21543-645e-4493-b87e-507f7de8f4bd · outbound

This paper cites Defending chat GPT against jailbreak attack via self-reminders.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Defending chat GPT against jailbreak attack via self-reminders

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.286459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.923177Z digest=sha256:91ad7a9e9e1e3cb92097200fe8977327a827d7325b3a5a1cfbc426936ee892f8

Observation 0a69fe84-05bd-47b1-b7c5-53d4f54bf1bd · outbound

This paper cites Qwen2 Technical Report.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Qwen2 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.927696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.927696Z digest=sha256:e94715bf317bbedf1e608296650b3d1e526972684cfc010de1b22501731bc1ec

Observation eec647d9-6744-45df-bee5-89ac7b834d2a · outbound

This paper cites Tell your model where to attend: Post-hoc attention steering for LLMs.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Tell your model where to attend: Post-hoc attention steering for LLMs

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.270411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.932874Z digest=sha256:4a95f120f9375e1268d600b9ce3918fe8c9b1c45166a11371f8e4ef4d45a1475

Observation cdb0e819-a258-48f8-b051-968249140517 · outbound

This paper cites In-context principle learning from mistakes.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling In-context principle learning from mistakes

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.253913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.937270Z digest=sha256:d10e0a90eba24c3ad068d76d95cce022552a984821fe0b301d44c48a48673633

Observation 3158314c-1474-4972-8274-163ca8dc1a5a · outbound

This paper cites What makes good examples for visual in-context learning? In Advances in Neural Information Processing Systems (NeurIPS), 2023.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling What makes good examples for visual in-context learning? In Advances in Neural Information Processing Systems (NeurIPS), 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.235536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.941925Z digest=sha256:01106e10b7f8c99344c3adf3e397db45a6fc733ea2c9cc53a313fac444eea3e1

Observation 0bddeb3a-23f4-4153-b5ee-d7a272365c80 · outbound

This paper cites Calibrate before use: Improving few-shot performance of language models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Calibrate before use: Improving few-shot performance of language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.218387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.946524Z digest=sha256:cf027c19c8f089bf7e571b646404efd4849d60a412f8ad0f3fc47abf1035f663

Observation 7a23758f-509b-465d-b6d1-76b017de50a5 · outbound

This paper cites Improved few-shot jailbreaking can circumvent aligned language models and their defenses.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Improved few-shot jailbreaking can circumvent aligned language models and their defenses

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.202097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.951304Z digest=sha256:09a49dbbbbaab9c86e7cba007ad6dc613248e8781943d9bec4641c61bb3738ac

Observation 8e68e883-4acb-444c-84d2-412a40e8c2b5 · outbound

This paper cites P., Di Eugenio, B., and Zhang, Y.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling P., Di Eugenio, B., and Zhang, Y

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.184896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.955740Z digest=sha256:1591aaf96a0c074935c18f2941b24c702ba11dd67c328f1e5b0b27a0cb354420

Observation 7078b9d0-2b3a-4673-864d-786cbfc30fb2 · outbound

This paper cites Z., and Fredrikson, M.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Z., and Fredrikson, M

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.168206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.960422Z digest=sha256:d44cf1c6bc4fb6406f11deb41082e273b7b8d3fa2e78505b8a9ab0e838206fb8

Pith citing papers

Observation cc00ba49-84be-496a-8562-a7fb38e9ce60 · inbound

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration cites this paper.

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:14.961667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T00:50:19.648405Z digest=sha256:1ded112f69a468476f35a57cf3b1e49a61c5bd53d1032a85cdc3b70347428025

Observation ce3b0946-6889-419e-b0b4-9ee3f453c93f · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:22:18.823341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:ed98b1705149b3bf4f171662c717d69dbfb59559309d0fccfa9277a671adbfab