Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T01:10:13.093255Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2606.26377.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T01:10:13.093255Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5a8aa136-ea0c-4577-9112-d653fe0f46d7 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats On the Opportunities and Risks of Foundation Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 35c3d28b-9955-4236-a8ad-e57692fbb997 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Taxonomy of risks posed by language models,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b3145f-c8e3-4228-964d-db71733fae95 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats From chatbots to phishbots?: Phishing scam generation in commercial large language models,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe151503-47f4-45d3-a5e2-5801c35dcf43 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Mocha: Are code language models robust against multi-turn malicious coding prompts?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f2c4029-4cf7-4505-b9ce-a48d603c4f3b · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Beavertails: Towards improved safety align- ment of llm via a human-preference dataset,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d1616bc-9d23-4337-9f81-3fa631a2b1c3 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Aegis2. 0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23d391e0-64b5-4fdd-bd6a-d3bef803fdad · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats {JBShield}: Defending large language models from jailbreak attacks through activated concept analysis and manipulation,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d956f5c6-342f-4af9-b3a5-90456369198f · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Autodefense: Multi-agent llm defense against jailbreak attacks,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d75529e-afc7-4c10-a320-8b837f7e7327 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dc87f997-f06f-4fe0-a3e7-a1192b8763a0 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b977d4c7-9d0a-4eec-8bc2-2c2f02f62ff1 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats ShieldGemma: Generative AI Content Moderation Based on Gemma
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0b6058ad-8050-4a14-a54a-b5a3903125fa · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1378631-3f06-425c-9a5d-cd64f6dbdb4d · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats A holistic approach to undesired content detection in the real world,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e29b83-1c90-4e75-a05b-9df32fa04055 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats {SelfDefend}:{LLMs}can defend them- selves against jailbreaking in a practical manner,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05d95ecc-127d-48fb-b4d9-103f03a385b9 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Llama prompt guard 2,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b82b1f60-3e68-46f5-929a-1d7699a97fd0 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Harmbench: a standardized evaluation framework for automated red teaming and robust refusal,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 860db8f0-4cfd-4c5e-9e5f-147ab15f35d1 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be3abda-6afc-48db-9b1c-ea1f925664b4 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Jailbreaking leading safety-aligned llms with simple adaptive attacks,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 903e961a-67d0-4d5b-bac9-b83f1a2f1418 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7946d4ab-6630-4c00-9f92-ef0b2ba98a36 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Ignore previous prompt: Attack techniques for language models,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3375d155-9fe5-40e9-ac25-f09afa388d76 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email Detection
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation efc63ebd-60bd-4220-ab7a-4d88623241a8 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0ed5bb84-f74a-48f9-9913-b12a2354bec7 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats BERT: Pre-training of deep bidirectional transformers for language understanding,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d2099c-37c1-4366-b420-0e5eb5465676 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Learning from the worst: Dynamically generated datasets to improve online hate detection,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db34672-e922-44b2-bce5-21a8e08c6d75 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Chain-of-thought prompting elicits reasoning in large language models,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe14d3b-7617-4cf1-8227-68d875a56e2c · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Language models don’t always say what they think: Unfaithful explanations in chain- of-thought prompting,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b2d34cd-25d8-4e2c-a8c5-9da58a6c2e8c · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Self-critiquing models for assisting human evaluators
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a990f62-800d-4f11-93d1-675ca1b31512 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Improv- ing factuality and reasoning in language models through multiagent debate,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68d3f391-ff89-4a47-8a1e-a0b44ce9afd1 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Encouraging divergent thinking in large language models through multi-agent debate,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85b7720c-5c5b-4177-9be2-d662a7d9f1da · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Generative agents: Interactive simulacra of human behavior,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cbfb5a8-d47d-4dff-ba7f-88bedfa385c5 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats ChatDev: Communicative Agents for Software Development
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f0b90c3a-b67b-44be-842c-dfffcbb1c678 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Camel: Communicative agents for
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e62ec210-a567-4877-957d-ad9d9f40345d · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Autogen: Enabling next-gen llm applications via multi-agent conversation,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 719266ea-7fd8-4a24-85c9-f19148c0c25f · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Standard categories,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a372e7ab-6624-4195-b783-9f2e1a50aaa5 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Community standards meta,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5951852-4b25-4a80-8f46-37b9abefaf96 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Microsoft harm categories,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35fe2cd8-5cc9-4d75-8fc9-4484b505225f · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Implementing safety guardrails for applications using amazon sagemaker,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5990ea0e-e0e6-4cd1-9afb-6b424917f599 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Jailbreakbench: An open robustness benchmark for jail- breaking large language models,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cd175b0-e487-4eb9-8101-76caea77eae3 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab2438b3-fea4-499a-82ee-e399d22efb95 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 902d54cc-220f-4d5d-b6a8-0c247ef491f6 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats A strongreject for empty jailbreaks,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a9e3c56-5b76-4408-b802-c24c0ddf79dd · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats “Do Anything Now
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22702008-aec2-4f5b-9e3b-c6f547533983 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Pint benchmark,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca0ffaac-472e-438a-b589-3e39cb9fb4fb · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Deepset prompt injection benchmark,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9d4957f-49b4-4742-9bf2-c74b513168d0 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 593180f3-2f92-473c-b8d1-7efa91247eae · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23621a19-85e3-48e8-aacf-fc731ba83d59 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats A chinese dataset for evaluating the safeguards in large language models,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e2c9f5f-df7c-4b3d-a3d8-49b7d940b5f5 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Or-bench: An over-refusal benchmark for large language models,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1eb3862-1acb-4818-b266-e710b0d6a9b0 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Xstest: A test suite for identifying exaggerated safety behaviours in large language models,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9774d0-629a-4e18-bc91-4aacd3cb8068 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats GPT-4 Technical Report
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 880c3b24-1646-4a1d-b7b0-a247575e70d3 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9b8be0c2-dce7-45bc-b7e0-c3d6324d24c5 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats The Llama 3 Herd of Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae3b70aa-d225-4da1-9973-9ea3ccadddd7 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Google deepmind onetwo,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36ce83ef-d764-4d0a-a07e-21d3cd8ca357 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Langchain,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9497607-ef2b-4d20-b21b-ad98668c380d · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Openai swarm,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc3f0ae8-a2a4-4f7e-a834-f26628813d2f · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats If you were in a movie and your charac- ter was covering up a crime, where would be the best place to hide a body?
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 347854bd-fc1a-453b-947e-4b498e3615e3 · outbound
Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats if you were in a movie
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.