Pith. sign in

Paper Citation Record · LEDGER

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs

As of 6 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 1 inbound Pith citation observation for arXiv:2508.20325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20325 v3

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T21:34:51.665401Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:05:09.033412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T09:21:00.875633Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact10
  • verified fuzzy21
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch22

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f9d800a-8f84-435a-a63d-65f7e40f39df · outbound

This paper cites Journal of experimental political science 9(1), 104–117 (2022).

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Journal of experimental political science 9(1), 104–117 (2022)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.510498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:e280296b99bc3d763ca328b1dad4e953665a006b1823f8f7c08ecc3b7542a4ad

Observation 17c02676-cb49-44b3-9322-6652d95bacfc · outbound

This paper cites Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Generative Language Models and Automated Influence Operations: Emerging Threats and Potential Mitigations

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.538697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:cdd63e3a2901afeb3b83182d9d039b69dce0f9d1fa1083f9178186010e87ac21

Observation b6e80b81-4dcd-4456-8b7d-eb0beca2dcf9 · outbound

This paper cites Computer Law Review International 20(4), 97–106 (2019) 14.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Computer Law Review International 20(4), 97–106 (2019) 14

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.507524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:40ceda504d56dec57e7ac5ab3e8cd7314f5ee29d63f6a4f2538a4c1fab39aea5

Observation 617d0056-91d9-4b76-b5a3-4e7f565a3e5b · outbound

This paper cites https:// www.aepd.es/sites/default/files/2019-12/ai-ethics-guidelines.pdf.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs https:// www.aepd.es/sites/default/files/2019-12/ai-ethics-guidelines.pdf

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.520473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:8faa15f2c3983affba8aaeda2c54fbdf800d9d8dfba3b092a166ead1f0b5b2b2

Observation 36a52f62-46d2-4e14-b124-ab9ce795e576 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.543549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:78eb6bd525d152e307d1c2c19e86a8b580f09348739b75b83477c2c12c3f8548

Observation c96c01fc-a47a-49e4-a28e-f55a81956a11 · outbound

This paper cites Retrieved August 24, 2018 (2016).

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Retrieved August 24, 2018 (2016)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.501505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:745047f59a4e56d09c0c72627c88915aa025d31e82c217a21b89142279979d1d

Observation fd47df4b-6d1b-40c3-b1fd-7424c6a93547 · outbound

This paper cites https://www.whitehouse.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs https://www.whitehouse

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.513184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:ee7bb029f1b28df16e2d2cace37143d5206b5517b665bfcdeeba00793e4e83c5

Observation dd16909e-5b76-4ffa-84f8-68abfe4b6a1c · outbound

This paper cites https: //www.whitehouse.gov/briefing-room/presidential-actions/2023/10/30/ executive-order-on-the-safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence/.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs https: //www.whitehouse.gov/briefing-room/presidential-actions/2023/10/30/ executive-order-on-the-safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence/

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.504555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:71b2bc8e50e81aebe4d30869c2a4293901660266f6ddfe088d6a8c3f611bad55

Observation 6753390a-b2c7-48c6-90a3-c457f2451d9e · outbound

This paper cites https://www.nist.gov/itl/ai-risk-management-framework.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs https://www.nist.gov/itl/ai-risk-management-framework

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.516573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:359f052e90d7e5608c641ec1fe35e5b0b928bfca3370a9af93510f93adede4c9

Observation 3af16540-817b-4bf9-b7e9-95e2647b4e4c · outbound

This paper cites https:// assets.publishing.service.gov.uk/media/64cb71a547915a00142a91c4/ a-pro-innovation-approach-to-ai-regulation-amended-web-ready.pdf.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs https:// assets.publishing.service.gov.uk/media/64cb71a547915a00142a91c4/ a-pro-innovation-approach-to-ai-regulation-amended-web-ready.pdf

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.524524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:d0e7aa00679778bf68a7d083a76c69a4561fa858cc20e470e6c846062bd3edda

Observation dc8c2dad-3780-4218-b9d1-d4dbaed9545b · outbound

This paper cites https://artificialintelligenceact.eu/ ai-act-explorer/.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs https://artificialintelligenceact.eu/ ai-act-explorer/

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.529510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:97fae5b67766b63f321e882aaedd97bfe164f17664856e9f5f3619f10b3f75fe

Observation d3255833-f4a6-456b-8cee-eea0a074bc9e · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs ChatDev: Communicative Agents for Software Development

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.430664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:ee2234388e2d35919eadfd3237620dbafaa526a96ea3384e498b432202fe203c

Observation 0f8952ac-551d-4905-889c-920b7e5bc657 · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.388845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:44f936bd7641a99313d7d6223f98ac04633eb2f4be59071111af4bfec42c011f

Observation d9da99c8-fa42-40dc-9037-86276faca79b · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.379499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:122126cc20420faf02ec79205d52c074aa4c40688c6f9d4e7a229186a39d5fb3

Observation b54939ce-00ce-420e-87cd-32ade625a2bb · outbound

This paper cites In: Proceedings of the 36th Annual Acm Symposium on User Interface Software and Technology, pp.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs In: Proceedings of the 36th Annual Acm Symposium on User Interface Software and Technology, pp

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:41:53.103068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:d76ae7e8aaaa364494bbbbe63ae89cc180f2840cf102010109e0f4ad13bd7806

Observation 50bc6b1c-d234-463a-89f7-0ea1aeb79db7 · outbound

This paper cites Large Language Models Perform Diagnostic Reasoning.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Large Language Models Perform Diagnostic Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:36:52.384146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:b7f9a2cb0cdadfa9690c2041cad55ef70e3609c9cff39b4f93829dcce655193c

Observation 3cb5697e-cea7-4fca-8936-27bc01d45e1a · outbound

This paper cites MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.374799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:14e7d9b55e04ab4d0d2e0fc0b52df9600c1c2fd295d44e9a4b65c0da7f344f6e

Observation 3f102d19-6d21-413d-af7b-8eade6e78125 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.369906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:3b84a7fd35fab8bbe003d6703205573038eaa01a150ed46eda05eac76a4f249f

Observation 1a6a6f27-68d6-4b5a-8c6d-27446e64585a · outbound

This paper cites Advances in Neural Information Processing Systems 35, 24824–24837 (2022).

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Advances in Neural Information Processing Systems 35, 24824–24837 (2022)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:41:53.106256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:b65b83d20b84d09daf7a6fdfca203db61b400840f3b05797298ff279cfbfebca

Observation ef35f386-b49d-4051-89c1-c0d2990413b2 · outbound

This paper cites Multi-step Jailbreaking Privacy Attacks on ChatGPT.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Multi-step Jailbreaking Privacy Attacks on ChatGPT

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.522996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:fc86e1b4b431a00fcbce6a6fbe046e9241b19d61fc6df2420bd6668d1c0f156d

Observation ba88cf10-a96b-4307-8f2c-e877ff3e2726 · outbound

This paper cites In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:36:52.393660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:a1a427d193f6c8c3ff69f35fb8e3599c504eaeb555a0845bf95751468688d052

Observation 4b7e8fc8-a251-4518-9670-6ab5593c6d32 · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.453943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:eb9c9c24dce9ff0dc020fcc5a85806a5c976b2e67efa34eb56aad458a4d72a26

Observation ef1d3bf9-f50f-4b5e-be1a-336db1a2e088 · outbound

This paper cites Automatically Auditing Large Language Models via Discrete Optimization.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Automatically Auditing Large Language Models via Discrete Optimization

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.469255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:9d7ecaa3ca85f3dbb52ac58118c3d2c86981159bf7b008e2a7c984843de03fc3

Observation c058c589-b1a6-4851-92fd-d14f34da76e6 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.491112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:cad1b4aebedcbd90f8adcde0259ab2d946cee70414137dcb1f1b60d004384a81

Observation ebb79019-4428-4388-acd3-2fcb2098df84 · outbound

This paper cites AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.459033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:817145dc1a6984886c54b6c04218170fdd9c3c9bcdbbb1ba25a79403a46be39b

Observation a5be04eb-7089-4b62-a33c-3f3cd94f841e · outbound

This paper cites MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.533755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:fdd28494da653033d963ce054a5d7c3d37474a581a5b2b101e069f869c408664

Observation 921a1417-8384-499b-97e9-4bdae34b8f08 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.479691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:72e8557eb458e8fa0f6d7208644b00dde5a90523932b4b17c843ca30d25fb55a

Observation f92ea18d-b282-49f0-94a7-78cc624a958d · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.473897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:91151764f308da23438a789bd9a2b94394470b6fa1c09bef8faba0e3aa01c633

Observation 51bae0dc-effd-4724-9804-6e3ad33b4bdb · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:36:52.516771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:ef4d4a8086a717cd4b7c620305a0eb23341510d7b45d9c34a6b424ff6674dd64

Observation de79b5a9-c7fc-4908-8bde-64d2917fcd92 · outbound

This paper cites Query-Based Adversarial Prompt Generation.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Query-Based Adversarial Prompt Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:36:52.402868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:e0177862af975c813d66f10bf95b46433b2781e0133239bc0db35ac0cee9eb40

Observation cb1b8fd8-a396-4c50-8c2c-42e2b2672dbc · outbound

This paper cites CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:36:52.485500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:0c2e5d72e2540f2e12fc49406f4b217e27d2b3a5ec04b5bb0b43c33b0f6419ac

Observation 5c51fca5-3ace-4ed8-aead-76b29948150c · outbound

This paper cites DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:36:52.445155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:189adf6c97005b95bf70f86240ab00ad244805cbea789d886a340cbf8e9329ca

Observation 9f740cc8-976e-4e1b-bdac-fa4eb26f8472 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:36:52.497769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:67b9c16127e979f32bf4fa62a0c5fa37505593b1e8a7bea06ffc72b8df6959fd

Observation 7c05e2e6-6f84-4eb6-bb9b-85156b3002b6 · outbound

This paper cites arXiv preprint arXiv:2402.10601 (2024).

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs arXiv preprint arXiv:2402.10601 (2024)

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:36:52.528600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:e87323d2d8692718b0b9d32cd685e90d18a5e6e5df1e467cb274392e04f175cc

Observation 44863539-cf70-48cd-8278-ea533d6fc824 · outbound

This paper cites Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:36:52.440502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:32417cb62008522069552d5465c8a2072fc799aa0793cc00890d01e0f2692e8a

Observation 373986f5-47b5-4ec7-93f6-d1525fb8a244 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.464182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:eff1d8516382fc2f9f5770523360d8b1bdbf7581ae7bf4267511ac7ca985e69d

Observation 8c116005-d260-42e6-8d32-f524cffab5ee · outbound

This paper cites In: Theory and Applications of Ontology: Computer Applications, pp.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs In: Theory and Applications of Ontology: Computer Applications, pp

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:41:53.109543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:a20759c336ae3fae12d382ce06b34636b17f54e0dd6a1cd1592a1a2c901e0df3

Observation 2ef11fd6-9aae-4990-9fe3-7a91b7b30593 · outbound

This paper cites In: 2014 47th Hawaii International Conference on System Sciences, pp.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs In: 2014 47th Hawaii International Conference on System Sciences, pp

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:41:53.112843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:8794f6407dd11e5a89aa83e0ce6f99286dc8b221b9fbdf2df77070dc58aebe0a

Observation 8be137c6-aaf8-4f99-bbaf-974704d4df07 · outbound

This paper cites IEEE transactions on neural networks and learning systems 33(2), 494–514 (2021).

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs IEEE transactions on neural networks and learning systems 33(2), 494–514 (2021)

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.472555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:2b031b8d54adb0d5a05368c70646a8af49405bd7c99fd281569bbecd6b5c3757

Observation b7468a20-16ed-4857-afa4-b956386523d6 · outbound

This paper cites In: Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs In: Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.480798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:893b17f5b3b43019ad75f71029c66fcc87683d72c9f260af7ecc250c36f5249d

Observation a66af267-0a44-4fc0-aa9b-494d07cbae49 · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2024).

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Advances in Neural Information Processing Systems 36 (2024)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.498417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:d0a7b7b369ba2135474d33b33d357e23295a7fcd5f7e999694edb5adca18be16

Observation abe688da-d2bf-4bcd-98bb-5fbcb186008c · outbound

This paper cites https://lmsys.org/blog/2023-06-29-longchat.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs https://lmsys.org/blog/2023-06-29-longchat

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.476942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:c9bd0e41e0cf5845039c449d00cd250dd017740ce9f8987402280f8a24b5b7f3

Observation 0c06697e-3193-432c-a10b-b2c16b5bac22 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-18T21:36:52.510207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:292b308599fbe5a685144877e2e17d8a876121126c6d118d0fed690c943736b4

Observation a0822686-31bd-454d-b969-6fbc4d60c89f · outbound

This paper cites an unresolved cited work.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-18T21:42:51.489696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:9dc9d3945ceb2c5fa2b4fbc5e05dea3c11cc69a5f9f66925262b145883ab8aff

Observation b0039adb-b194-496e-abf9-8576f5d0ec47 · outbound

This paper cites GPT-4 Technical Report.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs GPT-4 Technical Report

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.419620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:3535c9882895288b377acd473811fc0956d47d4811a3067d1e7f47ee4e6f4968

Observation 8fee8f1f-2c85-4974-8f40-39948f62a05c · outbound

This paper cites GPT-4o System Card.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs GPT-4o System Card

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.449485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:77134686ba2bb490f5477782aa89c069cd83af8e3139b877c73ac9a7f232beba

Observation 163d61af-a4ed-4caa-b19a-cd8c1d5de4f5 · outbound

This paper cites Claude-3.5 Model Card (2024).

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Claude-3.5 Model Card (2024)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:42:51.493727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:6c154ea6dab71673335544a16f240148ac5061e5b72fbbfe9d2eacb81e92655d

Observation 4752f6a6-2bfc-40c6-98c3-dc21c202560e · outbound

This paper cites Language models are unsupervised multitask learners.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Language models are unsupervised multitask learners

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:41:53.130547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:c1df4024c0d7fb606d21ad50dfae4b3add59ec2cef1f559c7be59988bb3a33d2

Observation b543987c-e8c9-430d-aa1e-8975058769ab · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.424978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:7e08c5c32c482acabab6f2daff71d6e7b763f4de6d05c2be5b3d2c4edd6c329f

Observation 7604c940-8b49-4e38-a32e-a5341efb20e0 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.435410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:18d83d3d5d0ed3d7d7999bac8c6e8a8c68b6a815705e73ae15d7d38b645fe3b6

Observation 334cbcc0-699e-4b65-9c26-658840696e42 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:36:52.414153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:c9527265614ac66c4967a31537cf288c1242c01663c8d330829e9e14cfac5c23

Observation 0716171d-da89-43a0-b878-46b3aa7e47b1 · outbound

This paper cites an unresolved cited work.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-18T21:41:53.133327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:7d17a1c5caad62f70530e94605a3bb7a42d6ba3ae5ced16aa7def2799e6923de

Observation 2b7a8093-461b-4ecc-ba4e-569bd2e286d8 · outbound

This paper cites Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:36:52.409545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:8d3db1e04b9fae24e8f1381168c1aa1787e0549711ee0d4399cfac626afbdc10

Observation adda398d-c0ac-4c98-aa65-02f01bfa6911 · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Efficient Estimation of Word Representations in Vector Space

Reference 54

Resolution
malformed identifier
local_arxiv, observed 2026-05-18T21:36:52.398224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:7502bdd7fd2b185e7bb974e5c77022f086914204f830ab26e56b11ad978cee4a

Observation 6b0ef2a5-16a2-4f21-9b3f-45dcef4a85ac · outbound

This paper cites an unresolved cited work.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-05-18T21:42:51.465212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:39f63a5ebe204774a95e3114bd2b29fb4f935b3da3e665f79418019343e8f3ef

Observation 091d8c05-5855-4ce3-8e87-5d9326de851b · outbound

This paper cites an unresolved cited work.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-18T21:42:51.484712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:c60eeba872dfcbc057f799d14a6f6f1d2693d0005a78d9bd81a9a5a45ff4c436

Observation 65b22d6d-10c3-45fd-b5e9-47d298672b08 · outbound

This paper cites an unresolved cited work.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-18T21:41:53.122021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:fe309070a3569f7b3c7bbd180e7d709c1efd55b8f9b206dd215aa36eea28145a

Observation a181bf21-cda8-432b-904f-0edbfca095ec · outbound

This paper cites <Example 3> Domains: Education, Consumer services Scenarios: AI systems may be difficult to interpret, leading to incorrect eval- uations or distrust among users.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs <Example 3> Domains: Education, Consumer services Scenarios: AI systems may be difficult to interpret, leading to incorrect eval- uations or distrust among users

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:41:53.124993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:a6e96ad970891f99f7be3ecf0129def0aa6a1e26b042494aeb4eb59df0e588b6

Observation 0fa1df22-167a-4d4f-8988-24f4c43ac3e8 · outbound

This paper cites an unresolved cited work.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-18T21:41:53.127482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:2e4064ec11680c542275d6ee0c05a3dad69a1373020e85e3587fd3a7d65c28af

Observation a29274c5-f99d-46df-8c77-0dc20a427d0a · outbound

This paper cites an unresolved cited work.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-18T21:41:53.115821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:79545a8829988fac22c2e4a1ba7119849785a2e9aeb610b7634af2e4c5792e51

Observation e7ccb396-daa3-4203-8ba9-2295ed8999fd · outbound

This paper cites Stay in Character!.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Stay in Character!

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T21:41:53.119075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:f36cfc5fe1766940cd8d8ecf5809f0c8f9234756fd416e06867b0b21ffa48b72

Pith citing papers

Observation cc542535-d662-4c85-8865-a815766dbdf2 · inbound

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models cites this paper.

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:42:01.658650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T16:05:09.033412Z digest=sha256:a526a7d186d823743eaeef34708e6e5b5fd2a1791cc562a68b397efb075aa238