Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:02:38.960422Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2502.01925.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:02:38.960422Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:17:34.283917Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
58 of 58 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation edc6b069-b5ec-4df6-b96a-ea72e66ed362 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a89191-3d1f-429b-8239-d140d4c4eee0 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0af1b2ab-2c22-4241-bb72-d7ed61f1de33 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling What learning algorithm is in-context learning? investigations with linear models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c4799bcd-3113-4929-bbf6-ef61d02c8f81 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbreaking leading safety-aligned LLM s with simple adaptive attacks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 810a0264-d52d-464a-bb96-7e44d721b7d6 · outbound
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0f58597b-951d-43b5-b104-348bb47970ad · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4038b2-7caa-44c7-8fa2-9afbbd46c36a · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2076e417-f77e-4b9d-abfe-e7389033a782 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling J., and Wong, E
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 53ca5892-b4e7-441f-9fb9-500c604c1d03 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling How many demonstrations do you need for in-context learning? In Findings of the Association for Computational Linguistics: EMNLP, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1dba753b-4f46-419f-afa2-0d2c5c3e4f2c · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling What does BERT look at? A n analysis of BERT ’s attention
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 41cf55af-ddbc-455e-b3bd-97318fcc5b60 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling BERT : Pre-training of deep bidirectional transformers for language understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c596d65e-0b2a-41ed-a8ca-41fe1030c439 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7548d31d-efd0-4ff1-936e-a3329d3fc9cc · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling X., Wang, B., Tian, Z., Chen, W., and Wen, J.-R
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6b652bf3-e455-4335-9306-f8562761a66a · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 339f643a-586c-477f-869c-4a27c619d4d3 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1a8bc742-36e0-4988-9b6b-e375a885ad6c · outbound
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cf188ecf-51e2-484e-92bd-a5e0721bc083 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling ChatGLM : A family of large language models from GLM-130B to GLM-4 all tools, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3540208b-63db-4f9d-b6f2-6aa9e12ad817 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Comparing results of 31 algorithms from the black-box optimization benchmarking BBOB-2009
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 78bba061-cf75-4aec-9343-966316fe9e89 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Self-attention attribution: Interpreting information interactions inside transformer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 573ae094-0bdb-4b6a-aae0-1f0ac36359fa · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling WizardLM-13B-Uncensored , 2023
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1e508619-bdbc-4886-be77-fb0afa741acc · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Measuring mathematical problem solving with the math dataset
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2d7f3953-7880-46eb-8df1-0541ac21f691 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f05d7f-5f69-49ce-b5fb-96df45cbc25e · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling LLM maybe LongLM : Self-extend LLM context window without tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e60d5aa6-b1e3-407d-a9b8-d98f36562f04 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c034bce8-ca56-49ff-b366-31cf55b5cc2c · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2ba4cd55-4d1e-48f6-a4c9-2645fcc1d322 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Attention-enhancing backdoor attacks against BERT -based models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a2b9f54d-e71d-42f1-9f8a-89a5fbe7ebb7 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f88212ac-0e77-44f2-8d9f-80d82fd105a0 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Tree of attacks: Jailbreaking black-box LLMs automatically
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e980d4f3-c52f-46e3-ae06-2e25889085ca · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Bayesian Optimization : Open source constrained global optimization tool for Python , 2014
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 08fe398d-46fd-4119-8b82-7e64b8ea9cb2 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling 2 OLMo 2 F urious
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e3e47036-ad29-44a4-9247-54e72e8d252c · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Training language models to follow instructions with human feedback
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e605e30b-a35e-4755-b477-4def7886fac0 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling S., Soltanolkotabi, M., and Thrampoulidis, C
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8800b52c-2221-4ea7-9382-7e70755cef03 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling S., O'Brien, J., Cai, C
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6a9f2eb1-747e-4c56-b095-68e9a8541ede · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Red Teaming Language Models with Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f6256a7-e521-4e52-9c58-7aa146b45b97 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Baitattack: Alleviating intention shift in jailbreak attacks via adaptive bait crafting
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 07ef3fc9-f027-4e17-9387-645ed9bd8f2a · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling and Barez, F
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 19d42300-af06-4934-87e9-aa127ecc40df · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Language models are unsupervised multitask learners
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3127a67-72f4-42e9-827e-fbd0049284eb · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Tricking LLMs into disobedience: Formalizing, analyzing, and detecting jailbreaks
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 662cfdae-b896-4968-9704-422ae6d95df7 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5cd759-6940-416e-a69a-fe7683933d1c · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling P., and De Freitas, N
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bc2101dc-cbca-443c-855f-1461c93cb6be · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9a0925-b676-46d4-a4dc-644c39dc8f0a · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Qwen2.5: A party of foundation models, 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 917cc9bb-3b75-4789-9ed4-6272e02ba0ae · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling LLaMA: Open and Efficient Foundation Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a1cd8ae-a4d1-42e7-aa71-d651bffd4ff7 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Bayesian optimization is superior to random search for machine learning hyperparameter tuning: Analysis of the black-box optimization challenge 2020
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 806772cf-42de-477b-9c5a-e628e8470ee2 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Attention is all you need
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ab6a9ead-82e5-4605-8df1-46eb9577f14c · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbroken: How does LLM safety training fail? In Advances in Neural Information Processing Systems (NeurIPS), 2023 a
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a928bc4c-1dd2-43bd-862a-186b3966b4c8 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7cbf08a-857b-48a3-99eb-54fc721163eb · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Never miss a beat: An efficient recipe for context window extension of large language models with consistent ``middle'' enhancement
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6a37ad53-bd30-4d3e-a0c9-25f8077beb81 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Distract large language models for automatic jailbreak attack
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1ad21543-645e-4493-b87e-507f7de8f4bd · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Defending chat GPT against jailbreak attack via self-reminders
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0a69fe84-05bd-47b1-b7c5-53d4f54bf1bd · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Qwen2 Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec647d9-6744-45df-bee5-89ac7b834d2a · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Tell your model where to attend: Post-hoc attention steering for LLMs
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cdb0e819-a258-48f8-b051-968249140517 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling In-context principle learning from mistakes
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3158314c-1474-4972-8274-163ca8dc1a5a · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling What makes good examples for visual in-context learning? In Advances in Neural Information Processing Systems (NeurIPS), 2023
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0bddeb3a-23f4-4153-b5ee-d7a272365c80 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Calibrate before use: Improving few-shot performance of language models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7a23758f-509b-465d-b6d1-76b017de50a5 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Improved few-shot jailbreaking can circumvent aligned language models and their defenses
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8e68e883-4acb-444c-84d2-412a40e8c2b5 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling P., Di Eugenio, B., and Zhang, Y
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7078b9d0-2b3a-4673-864d-786cbfc30fb2 · outbound
PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Z., and Fredrikson, M
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cc00ba49-84be-496a-8562-a7fb38e9ce60 · inbound
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ce3b0946-6889-419e-b0b4-9ee3f453c93f · inbound
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.