Pith. sign in

Paper Citation Record · LEDGER

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks

As of 20 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 5 inbound Pith citation observations for arXiv:2505.13862.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13862 v3

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:12:28.410564Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:26.916448Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T18:56:45.840997Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation df60a4c9-065c-486b-b7ac-20e5933b61f6 · outbound

This paper cites Language models are few-shot learners.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Language models are few-shot learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:27.952093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:27.952093Z digest=sha256:f71d43bbfa4316a2dc2254cf911f43f6eb8909ebd33fb1bc9f8d43dbe1ab6d7e

Observation 28de3028-cefa-44eb-a708-5c9bdebaed77 · outbound

This paper cites The Llama 3 Herd of Models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:27.959029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:27.959029Z digest=sha256:266a4237187d06d902dcbcc1dd1f7bca2484c0615e93368f3263cc8711dc0c3c

Observation 3f0b7cbf-1b15-42e9-a6eb-4ac51fec35f0 · outbound

This paper cites Qwen2.5 Technical Report.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Qwen2.5 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:27.965475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:27.965475Z digest=sha256:7fba349879aad9554286d11c87ac7bc40feee6efaf60b35bacea1753a1b5949f

Observation f42d85c6-a069-494e-8907-c733c42ba4ff · outbound

This paper cites Gemini - Our most intelligent AI models, built for the agentic era, 2025.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Gemini - Our most intelligent AI models, built for the agentic era, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.939360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:27.972524Z digest=sha256:ad6e5418662d6b01830b77e2f8794ffcde46753d2ad5a55674742cd0f60bcc85

Observation aa19e645-c592-4fb8-972f-6a658926bac0 · outbound

This paper cites Eureka: Human-level reward design via coding large language models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Eureka: Human-level reward design via coding large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.918845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:27.981197Z digest=sha256:ec4ebdb94ff2c8f18da58ba581d2a4ed1a673585babc154f33f8deefd5b1313f

Observation 40ca29c7-6ba4-4278-b073-8077625fdc2a · outbound

This paper cites Creating large language model applications utilizing langchain: A primer on developing llm apps fast.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Creating large language model applications utilizing langchain: A primer on developing llm apps fast

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.896603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:27.987784Z digest=sha256:b57f2516a60a5c4a80e2354452a9480bfda567b0f3f85421851e94c47f6ca57c

Observation 9c09b8a8-f03c-48d4-9b3d-966ea09ffdb0 · outbound

This paper cites Taxonomy of risks posed by language models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Taxonomy of risks posed by language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:27.996184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:27.996184Z digest=sha256:3a9c1d67dcf6ddd30e925a1e1aa59e9449373935c6a23ec2b8212f5b3be3cd7b

Observation ef1706d5-da10-47fd-9963-bdbcc9a7ebc2 · outbound

This paper cites Navigating the risks: A review of safety issues in large language models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Navigating the risks: A review of safety issues in large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.850221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.002672Z digest=sha256:966d350d8aaaff2a4eb83f523e883c0ba39b0cbf2bab502c922548cea7e766a8

Observation b0bf3749-b52a-4709-a08e-d22973a5053b · outbound

This paper cites A survey on large language model (llm) security and privacy: The good, the bad, and the ugly.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks A survey on large language model (llm) security and privacy: The good, the bad, and the ugly

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.010548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.010548Z digest=sha256:aaf4812c8b44faa0ddc1bddbbfaa26a1bb005bd7e9555f6b0d0ec11bbff008d6

Observation 9c117881-0b0c-4e9c-b21b-0b863a6aba2e · outbound

This paper cites Multimodal situational safety.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Multimodal situational safety

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.809951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.016883Z digest=sha256:32f5a41dd2923d4613bbe26cc80743b9ab6c440b559c747f6fea46d8c0638e8d

Observation dba124c1-07b9-45af-b6f2-d440207e5696 · outbound

This paper cites Air-bench 2024: A safety benchmark based on regulation and policies specified risk categories.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Air-bench 2024: A safety benchmark based on regulation and policies specified risk categories

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.784390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.023470Z digest=sha256:72ed552075460523bb28c0c5c8ba674953404b909c99eee04afd936cc84e804f

Observation d19f4f79-0f54-4a4d-af8d-e4fa17fd46f8 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Safe rlhf: Safe reinforcement learning from human feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.030062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.030062Z digest=sha256:d2d60a843d1102f62ac8209c8133c151db91921aad5bd375e192a4a3f77f69e2

Observation 62005d79-be73-4562-bb62-37f732cd1b2b · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.036730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.036730Z digest=sha256:a718055581a64fdf02059ddc820d30e4924f5d6d529f32a71e93048b15893deb

Observation 117632e5-4bce-4502-8524-66f8ac717635 · outbound

This paper cites Zico Kolter, and Matt Fredrikson.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Zico Kolter, and Matt Fredrikson

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.044177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.044177Z digest=sha256:0103892c79e9f636b7e6e2e5b2a32a7037282fcfaa89c881470a2ffc74576435

Observation 6fbf680e-e989-4c37-a0fd-9894f09cb4d1 · outbound

This paper cites Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.050206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.050206Z digest=sha256:16b3a9f5827f4e4636f89ac577be3f42788baf59e9dfc6e3eeda66c62dc16378

Observation 547b3c93-8e26-4c08-98c5-34219cd2557d · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.056212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.056212Z digest=sha256:9d32ac0bfed0513914d410568f616a909bcd4b2309a74775f6c48595aea0a970

Observation ce3c3ccf-2d06-49e5-ac7c-7cf0b2153678 · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.062971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.062971Z digest=sha256:db765b45a6208506c4ce17ce563c4b1e44dc436385c99ee279fc286e3386ed98

Observation 3f91d558-0aab-4c3f-9204-a0809c379448 · outbound

This paper cites Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.070723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.070723Z digest=sha256:e18dcff3a7d7f9ccce87c3724d4ba60b45bc325a712cb845ee26b4071433caa7

Observation ca58e9ef-186d-4163-a9bd-0937d8001d7b · outbound

This paper cites Jailbreak antidote: Runtime safety-utility balance via sparse representation adjustment in large language models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Jailbreak antidote: Runtime safety-utility balance via sparse representation adjustment in large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.732854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.076502Z digest=sha256:d1b4c8f1eb4412c66133165f784a7ff56dc2e967a5378f49c8f02d738ba094d4

Observation 4e7c375d-e4b4-4c32-af61-51fd5b2a7bee · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.081961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.081961Z digest=sha256:d58c7b03b2a55f6d33bcfd8c979f559ba7e3e6a75fadbd87332b3e61b1a0c9f8

Observation e23ca5eb-8e7b-4054-8846-bb347eec362d · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Jailbreaking black box large language models in twenty queries

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.683140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.087538Z digest=sha256:741dc2039986b1ebcb23983a0d39d531c93cfbeef97a8d821928314bb3d3f6de

Observation 92d0f952-ef6e-45ba-93d5-a2caea2fa3a1 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Defending chatgpt against jailbreak attack via self-reminders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.092938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.092938Z digest=sha256:c0c676339e8c96269af5c2f442b89f568b6d6854a29c1d2ddeeac5b7cdb5ef69

Observation 6f6b1eee-5318-46c1-bc88-0e3174ba4240 · outbound

This paper cites Defending llms against jailbreak- ing attacks via backtranslation.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Defending llms against jailbreak- ing attacks via backtranslation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.638761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.098601Z digest=sha256:115a2f885b9b8db41616913b82abd10a8541e442cce1240aa330d0e7b07e8771

Observation 4b7e79ab-22a6-4ca3-8149-57b07f1f491b · outbound

This paper cites Pku-saferlhf: A safety alignment preference dataset for llama family models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Pku-saferlhf: A safety alignment preference dataset for llama family models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.494475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.104148Z digest=sha256:e0f276996dc310320c23e90697d72ee41b4b9480a1731ce8670a6702b82a3dfc

Observation fb551892-74a7-40f4-a74b-9679ca250279 · outbound

This paper cites AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks AISafetyLab: A Comprehensive Framework for AI Safety Evaluation and Improvement

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.109825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.109825Z digest=sha256:ba5dc7cab1adb88c5749a8a3b4d90cca977f266849251cbaff13971d6ee48b6a

Observation cc9e03ed-3170-4be0-ac20-a66274b48ff8 · outbound

This paper cites Bag of tricks: Benchmarking of jailbreak attacks on llms.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Bag of tricks: Benchmarking of jailbreak attacks on llms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.473277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.115989Z digest=sha256:2517d0dfd3091a4f08be180ccf58b4b1bdb02c1f78de8c51744c00664c95db11

Observation 36bdcb7e-004a-4ee3-9615-466a71a03cb8 · outbound

This paper cites Safetybench: Evaluating the safety of large language models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Safetybench: Evaluating the safety of large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.122027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.122027Z digest=sha256:a62e0e94de3e6b0aadfdafe41abb0ab06bf1d4853baef5dbbe7e64421719b86b

Observation 525f3b0b-d83b-410f-82f1-d5b88db63591 · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.128237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.128237Z digest=sha256:a0a4bfcb75dc16611ccb147c620f9a6f025f79c53806557c090fc258f26aaf30

Observation 11714a50-da6d-4253-aeef-d85e5da45f71 · outbound

This paper cites A survey on evaluation of large language models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks A survey on evaluation of large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.134188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.134188Z digest=sha256:5117e9c248f2b1d080b270b76d08bdbabfee27e5dc922f36a3cac4140421e6c1

Observation 4f293da4-244d-4fd7-b960-c705317b1ba7 · outbound

This paper cites Advbench: a framework to evaluate adversarial attacks against fraud detection systems.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Advbench: a framework to evaluate adversarial attacks against fraud detection systems

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.400243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.140900Z digest=sha256:9c8bb0292609998d5d18652a53cabe0d4362442db22bf2f905e233f05404044e

Observation ee761edf-f374-4537-aaf6-1bd7323fdb57 · outbound

This paper cites Jailjudge: A comprehensive jailbreak judge benchmark with multi-agent enhanced explanation evaluation framework, 2024.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Jailjudge: A comprehensive jailbreak judge benchmark with multi-agent enhanced explanation evaluation framework, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.374479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.147108Z digest=sha256:f3768d1993674f67eeded90f4880c0edb6748111f062fef58786cbe689ab4ed1

Observation 9c2a1e22-bd67-42b2-8a68-d268c8b5f526 · outbound

This paper cites EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.154029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.154029Z digest=sha256:fadf4bc6755b125eb88ea95437d7ec08e3aec1997193af7e8e0fbbc47a1baa04

Observation 0d86aa3a-948e-4ef6-ab23-a8f3fcd2b301 · outbound

This paper cites Harmbench: a standardized evaluation frame- work for automated red teaming and robust refusal.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Harmbench: a standardized evaluation frame- work for automated red teaming and robust refusal

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.348580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.160971Z digest=sha256:03de1bec70b84d9f08e42759d36540674f16afd9dfba810b6b029bb460511202

Observation 7b881c55-f619-4f66-83f6-45e58a27cd2f · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Gonzalez, Hao Zhang, and Ion Stoica

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.166852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.166852Z digest=sha256:1c183ceeb27fcde236279fb48a68b01fa1386382f1f8deba40b9e209bf51aefd

Observation e428f175-7e1f-42cc-9850-57c51c630ac7 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Sglang: Efficient execution of structured language model programs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.174978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.174978Z digest=sha256:89d3c1c2e29d111ee13a6156f18a721c6fd9a0d41a6600cae27e38352f7e2480

Observation 108b9efa-00dd-481b-9b4a-6ff18aa6a91d · outbound

This paper cites Ollama: Run large language models locally, 2025.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Ollama: Run large language models locally, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.292654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.184031Z digest=sha256:c11ea50f4b05b0dc7f28becd08c3feb2c115c83f0909976141f3001577ee5598

Observation e508776e-ec90-43ec-8240-6c0c65e10bd4 · outbound

This paper cites Jailbreak chat.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Jailbreak chat

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.271428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.193182Z digest=sha256:ad215a7265d846ba5d33bb47a7e0831cf3340495039f2f5a41620b959945fb7d

Observation 1cbf9e98-4525-4e5b-acc8-d1f2c5e070f8 · outbound

This paper cites Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.199798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.199798Z digest=sha256:57d2607df5b332bfdb0827e921f523bd7f0d4381dc44df7bc39b6b60a704b7da

Observation 8b641bc0-a644-4e23-a30f-4ecfbe748f13 · outbound

This paper cites Improved generation of adversarial examples against safety-aligned LLMs.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Improved generation of adversarial examples against safety-aligned LLMs

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.237976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.207613Z digest=sha256:acd146b12e169319489016e0e405e52fc5d224bacfa0376e64893c19e8ecf028

Observation ee501cbe-1cb2-4f4b-8565-f22cad340991 · outbound

This paper cites Jailbreaking leading safety-aligned llms with simple adaptive attacks.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Jailbreaking leading safety-aligned llms with simple adaptive attacks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.217084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.213385Z digest=sha256:e553695b8d2e99766333083bfc27209a057328bde87ac91df6620af966eca4ed

Observation 9d093622-b3b9-4573-9b87-a1c59e925046 · outbound

This paper cites Does refusal training in llms generalize to the past tense? In Neurips Safe Generative AI Workshop 2024, 2024.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Does refusal training in llms generalize to the past tense? In Neurips Safe Generative AI Workshop 2024, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.192233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.220285Z digest=sha256:9b4f6599efee057a29c6eb91f87ac2f7aa9a037dc72b5a8c87336320aa8269c5

Observation 0b6db335-661d-4f74-983b-e9d420caf9d2 · outbound

This paper cites Rain- bow teaming: Open-ended generation of diverse adversarial prompts.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Rain- bow teaming: Open-ended generation of diverse adversarial prompts

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.170330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.226559Z digest=sha256:a14fabcc8360dfa12a248b18cb94bd56f464d3bcc41eaa7fcb2e647a13472308

Observation be8f0d35-a25a-4e36-895f-1e646b4e2338 · outbound

This paper cites Artprompt: Ascii art-based jailbreak attacks against aligned llms.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Artprompt: Ascii art-based jailbreak attacks against aligned llms

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.237023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.237023Z digest=sha256:0365c5e80225f879b5553c98138eb748e458616a38cf878381850d748ab176d0

Observation 06ebc475-7517-4d1d-93eb-36acfae75382 · outbound

This paper cites Deepinception: Hypnotize large language model to be jailbreaker.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Deepinception: Hypnotize large language model to be jailbreaker

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.245128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.245128Z digest=sha256:85edfeb704877930e8522f669375836a1382382fe9b263a357514f3ab60bcad0

Observation 70f47f65-b2bc-4457-a4d7-dd3004d128e7 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.252253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.252253Z digest=sha256:2daa32525a6c640d6874a4756d41e21ef99b838e6c19519833b3d0bc0d2368ff

Observation 74d3bf7e-424f-4930-ade8-9dc28e009b55 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.259897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.259897Z digest=sha256:789fa3f2f01618ac53c14eabe880531a8ff82a35a7c4170d8350be5f89d79a98

Observation c883163e-e2d2-42b8-a865-f62e33ce04dd · outbound

This paper cites Defending large language models against jailbreaking attacks through goal prioritization.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Defending large language models against jailbreaking attacks through goal prioritization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.117089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.267800Z digest=sha256:64f5b700872dad063acefc8e1486eb9798998973887a0ee3d0d7a12e529c965c

Observation 80a7c5a8-b76a-4556-b50e-fb10ebe7720f · outbound

This paper cites Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.274891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.274891Z digest=sha256:00e0845f4bfda619d8a5625f9edbfed325c46d0f5c55fb9c1fe38de9bc0fd7fd

Observation 5e51d652-fb1b-4145-b576-ded5c0ed92b7 · outbound

This paper cites Llm self defense: By self examination, llms know they are being tricked.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Llm self defense: By self examination, llms know they are being tricked

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.284370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.284370Z digest=sha256:229591ce4bc6bf44ad381f5aabe4360fa8b6eb5b6f378f8a783a9789bd477c3d

Observation fb67fe54-6dba-4199-8a35-df3e5318e44d · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Representation Engineering: A Top-Down Approach to AI Transparency

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.291761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.291761Z digest=sha256:9e8c7633ed743f96437651312a776affb8d9304695a2f937e902659164ab9329

Observation fbd7e63d-23be-480e-840e-cbdd99a7fb7d · outbound

This paper cites Safe lora: The silver lining of reducing safety risks when finetuning large language models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Safe lora: The silver lining of reducing safety risks when finetuning large language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:29.079297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.299921Z digest=sha256:b1acd7fc4770d479dfc039646e9823418ac1fe54958bfee9b383598afbd84964

Observation 7fc3552a-a6a5-4bda-af0b-c20fc3a9c00f · outbound

This paper cites Training language models to follow instructions with human feedback.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Training language models to follow instructions with human feedback

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.306823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.306823Z digest=sha256:0cc375fe9b32b4f7640e9fba8a29168dc386f0904ddaf77a11df21871d62fbb0

Observation 6162b60d-3417-4222-a562-470ec91211a5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Direct preference optimization: Your language model is secretly a reward model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.312870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.312870Z digest=sha256:b5939e3bda60ea5a10ca5d77010507fa0f74e0e86e02167c3565db2e51968289

Observation 44affa0e-f407-49f7-8e20-c38b7e980fae · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.318309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.318309Z digest=sha256:041fbe33564ba3b1f67e2189e9cac158f303d6e42c92f52113aca2137a78b65b

Observation f21669cd-3664-4567-9a4c-51c89bbd2c93 · outbound

This paper cites A Survey on LLM-as-a-Judge.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks A Survey on LLM-as-a-Judge

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.324184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.324184Z digest=sha256:49bb62b9cdd006e8bf6aca1270ccb8666dd685c971c79a12751644c9dee8377d

Observation fa8787ee-838d-4920-b296-e9212cf735a3 · outbound

This paper cites Promptbench: Towards evaluating the robustness of large language models on adversarial prompts.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Promptbench: Towards evaluating the robustness of large language models on adversarial prompts

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.330644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.330644Z digest=sha256:78dee2de1593ee9f289c37379d6eb0ffe4b445cd5d8494999235df2075a4550d

Observation 8f3c5e57-b04f-4365-88e2-b767847464de · outbound

This paper cites Decodingtrust: A comprehensive assessment of trustworthiness in gpt models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Decodingtrust: A comprehensive assessment of trustworthiness in gpt models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.337667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.337667Z digest=sha256:d1503bf037b99c82a0691e4a19948b4bc01e93c515cd3492782836e4e488251a

Observation cac776c4-a048-4bac-8563-ec5817bec7ea · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks TrustLLM: Trustworthiness in Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.343741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.343741Z digest=sha256:35ddf91d169477bacbe550cd33145aa49c0c9b9c96a8f57f66290f7d31d5b4d3

Observation 19a5bcf9-4c1b-4e99-a5ca-c75c08940efd · outbound

This paper cites an unresolved cited work.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:12:29.006340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.350734Z digest=sha256:b5f55874c7f568c9979bd9b3f8db8d34067dae5845b4f7d8351596d9b603663c

Observation 7c3bf811-89c5-4df6-8959-ad73037762fb · outbound

This paper cites an unresolved cited work.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:12:28.984753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.356453Z digest=sha256:f459ee6fc255d62f6ee6bab20ab4114ef6bcd8db043f8cfcd105dfb152310fc3

Observation 7c66d790-b8c7-4409-899b-0c6e8fb680c0 · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.362060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.362060Z digest=sha256:fa045928944b7f11df2e398e5ea791814b2992505ea34450131c2fd04d23edeb

Observation 4b203de0-b6f9-4db1-bdb6-cbd4ac463b61 · outbound

This paper cites Hashimoto.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Hashimoto

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.368265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.368265Z digest=sha256:94ab4369100571d4dddabb6f8be32ab867a4ec7c902334af2b716bf5b1928348

Observation 195e4141-5804-4e75-9cdd-497e0384c223 · outbound

This paper cites Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:28.933483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.374011Z digest=sha256:3f78332f5d90547efd35815f0915257bbec70c8733a727cf0dc19d1fb9e3518c

Observation 2cb7c7aa-8f91-49a8-bb57-79e5a3691a0e · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:28.380405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:28.380405Z digest=sha256:11ce348782c7bf076372942b82316d711ca7a53d9502e7afa2169e534b2f0a68

Observation 3cc13dcd-540d-4f6c-b618-b4ee84c85a78 · outbound

This paper cites Uncovering safety risks of large language models through concept activation vector.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Uncovering safety risks of large language models through concept activation vector

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:28.908294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.387527Z digest=sha256:c3042f7223da2c60b4c56817a5cc1dd0bfca8e380acdd8e81533062256d45373

Observation 16ad5843-ad03-4cde-ae97-7ec4da586707 · outbound

This paper cites Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:12:28.483228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.395788Z digest=sha256:ee2d6b22c40605084a8245060c2f87c76c1a5c912ba6ec4ce821a38cea2a5466

Observation 8e03b60a-4202-4add-b3a5-d5e50da0d3a5 · outbound

This paper cites Robust prompt optimization for defending language mod- els against jailbreaking attacks.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks Robust prompt optimization for defending language mod- els against jailbreaking attacks

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:28.885088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.403981Z digest=sha256:7a70a32a962f7465ec55c3aecfb045e325d47a71ecd9ba83ea14062fc2fc4a1e

Observation b1701f1d-7cba-4dae-8844-f7f053c1c0dc · outbound

This paper cites I’m sorry.

PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks I’m sorry

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:12:28.865944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T20:12:28.410564Z digest=sha256:d6486d3426e2129fa79bd5369bffe774f67746c90f1c2925e6603b89e1255357

Pith citing papers

Observation 1f4f6028-cdad-4fcd-9f77-f3ac8d3676a6 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.916448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.916448Z digest=sha256:fc69342812d466ea892b5a10dde33d532707917f0e2e25a22ac799c53896d0a0

Observation bbdcf00c-f351-4fe0-8dbd-a1677a965279 · inbound

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal cites this paper.

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:56:45.844559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T18:56:13.680353Z digest=sha256:2a55a0ae1b4b7cc623c4b251a7466e137f6590aa90614533ee8174b99db7a996

Observation c0a14eb1-5771-4397-bd1c-340f8a782158 · inbound

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks cites this paper.

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T09:40:46.486872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:40:46.486872Z digest=sha256:f2d8bb8b44bb6259ca932dadb4f7fce461e6af198c55d9aa2d55c033218e8e76

Observation bc2eb868-8ecd-46a3-9d65-9f2663b84adc · inbound

Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models cites this paper.

Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:23:23.520526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T20:22:29.726790Z digest=sha256:e828cf650b7f9eccb85ceddd507d4073453b2191a57aa338316430356226c5fe

Observation 65b758ab-7aa9-48e5-8e67-aff1f084d2d1 · inbound

SoK: Robustness in Large Language Models against Jailbreak Attacks cites this paper.

SoK: Robustness in Large Language Models against Jailbreak Attacks PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:01:08.732256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T16:42:41.137808Z digest=sha256:736bbe75fa570eaec201eeeb29d194923b59233597a476281b4ee6bb869a960b