Pith. sign in

Paper Citation Record · LEDGER

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats

As of 4 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2606.26377.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.26377 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T01:10:13.093255Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a8aa136-ea0c-4577-9112-d653fe0f46d7 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats On the Opportunities and Risks of Foundation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.729290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:83bb7b66be94b47e2e7ab3b9b094f6691952806740d6302998b6482c5a357a9a

Observation 35c3d28b-9955-4236-a8ad-e57692fbb997 · outbound

This paper cites Taxonomy of risks posed by language models,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Taxonomy of risks posed by language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:c93f7a51043378c2822a731e2a15c5e12dcd780392886c69907b5dd7d52dfc0d

Observation b6b3145f-c8e3-4228-964d-db71733fae95 · outbound

This paper cites From chatbots to phishbots?: Phishing scam generation in commercial large language models,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats From chatbots to phishbots?: Phishing scam generation in commercial large language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:f0451d973fc99acdcff2832fa67d3b4fe559ba1ff65a2d8fe7c8e8749bb5ed8b

Observation fe151503-47f4-45d3-a5e2-5801c35dcf43 · outbound

This paper cites Mocha: Are code language models robust against multi-turn malicious coding prompts?.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Mocha: Are code language models robust against multi-turn malicious coding prompts?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:50e011c438ea7a985aeee44ccea9b73d34257bda64ca0e1ab4da473fc80c7752

Observation 4f2c4029-4cf7-4505-b9ce-a48d603c4f3b · outbound

This paper cites Beavertails: Towards improved safety align- ment of llm via a human-preference dataset,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Beavertails: Towards improved safety align- ment of llm via a human-preference dataset,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:0af0aee29cbc2d1a39ffca2da97d9849f8f8e8ffebce2bc8c4b8597cc70421aa

Observation 3d1616bc-9d23-4337-9f81-3fa631a2b1c3 · outbound

This paper cites Aegis2. 0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Aegis2. 0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:f765e02f6096bb55752753b1f001ab2c68997db20c013a1e54893bb0d350480b

Observation 23d391e0-64b5-4fdd-bd6a-d3bef803fdad · outbound

This paper cites {JBShield}: Defending large language models from jailbreak attacks through activated concept analysis and manipulation,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats {JBShield}: Defending large language models from jailbreak attacks through activated concept analysis and manipulation,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:94a02fa46d088fc0784312b223f717dc8c94518c9c7fad68a15a05f5d23d7283

Observation d956f5c6-342f-4af9-b3a5-90456369198f · outbound

This paper cites Autodefense: Multi-agent llm defense against jailbreak attacks,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Autodefense: Multi-agent llm defense against jailbreak attacks,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:500a062b6d119d13e3bd7949cf491a12270a2a25fa6396dbefc2902637d2e23d

Observation 4d75529e-afc7-4c10-a320-8b837f7e7327 · outbound

This paper cites InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:59:56.706982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:1bdf2ebf689b5eb63c4d424746401431958b2fde0c393c733f75b449b6b77965

Observation dc87f997-f06f-4fe0-a3e7-a1192b8763a0 · outbound

This paper cites Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:56.709785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:d98e376325932e956874d731bf527b5cabf14e13ad6d689029e41c311e5391ed

Observation b977d4c7-9d0a-4eec-8bc2-2c2f02f62ff1 · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.712231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:0722e86ae5793ead5569ace903f824faf506bc2c4226671116fd43c7358d3865

Observation 0b6058ad-8050-4a14-a54a-b5a3903125fa · outbound

This paper cites Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:1640fd35fe971ee7c29cf3e13b08b4d74bd7ef250bf5cd9b26a5b8e9c4798024

Observation d1378631-3f06-425c-9a5d-cd64f6dbdb4d · outbound

This paper cites A holistic approach to undesired content detection in the real world,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats A holistic approach to undesired content detection in the real world,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:e79dfe0484101cc496b61a38e132773fa591f366a57562457f0114c0166461e8

Observation 52e29b83-1c90-4e75-a05b-9df32fa04055 · outbound

This paper cites {SelfDefend}:{LLMs}can defend them- selves against jailbreaking in a practical manner,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats {SelfDefend}:{LLMs}can defend them- selves against jailbreaking in a practical manner,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:991f615b10eeb3d38464c03b882c3eae98f7ce654951355a58a61d3784c39035

Observation 05d95ecc-127d-48fb-b4d9-103f03a385b9 · outbound

This paper cites Llama prompt guard 2,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Llama prompt guard 2,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:f6952bf4c68a9f2eaa240dd99e5c1126df844d0a5de92015a09fc37c18ac4a7e

Observation b82b1f60-3e68-46f5-929a-1d7699a97fd0 · outbound

This paper cites Harmbench: a standardized evaluation framework for automated red teaming and robust refusal,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Harmbench: a standardized evaluation framework for automated red teaming and robust refusal,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:7c9a770eda524e802da4fbea2c8971106530870128cd82c7a4eeed84f5a667c8

Observation 860db8f0-4cfd-4c5e-9e5f-147ab15f35d1 · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:ea43ee5e298462f4d73dfa519c6a2586a72902dfffadbc92511ba8167e1f3f86

Observation 0be3abda-6afc-48db-9b1c-ea1f925664b4 · outbound

This paper cites Jailbreaking leading safety-aligned llms with simple adaptive attacks,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Jailbreaking leading safety-aligned llms with simple adaptive attacks,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:ab8505f06f95b6c6e782b86a1be8c334824b37530c2a239b7c8338bde8a373bb

Observation 903e961a-67d0-4d5b-bac9-b83f1a2f1418 · outbound

This paper cites Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:3e7aec9856d909980129cd9923c459b6a34eaddea4de62fe2d3d2ea07290afad

Observation 7946d4ab-6630-4c00-9f92-ef0b2ba98a36 · outbound

This paper cites Ignore previous prompt: Attack techniques for language models,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Ignore previous prompt: Attack techniques for language models,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:9df45f34e1864833d027b1d0c55ac8d7dd17c62a55b9ce99e1b91bf5f6f91c9e

Observation 3375d155-9fe5-40e9-ac25-f09afa388d76 · outbound

This paper cites MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email Detection.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email Detection

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.704276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:f55b9cd86108cd748dc5ddc650653b7aaad99ae4d1dca0dbda4d24255b703838

Observation efc63ebd-60bd-4220-ab7a-4d88623241a8 · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:56.714895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:5e9dc24a9a1373026f533785cdaaa4a6964ef4003a6741da9917e7d1fe29d560

Observation 0ed5bb84-f74a-48f9-9913-b12a2354bec7 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats BERT: Pre-training of deep bidirectional transformers for language understanding,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:1b8a55f518cb057edf40f25fbea38943ed8ebba83c71bc09d8b24d44a2d4048d

Observation d0d2099c-37c1-4366-b420-0e5eb5465676 · outbound

This paper cites Learning from the worst: Dynamically generated datasets to improve online hate detection,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Learning from the worst: Dynamically generated datasets to improve online hate detection,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:c6d9a8889552e836a5a23930a98154a01ec43a533b254c346717735130d223c1

Observation 5db34672-e922-44b2-bce5-21a8e08c6d75 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Chain-of-thought prompting elicits reasoning in large language models,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:7010ec00c6e67b5f0f0cc20d6916f08b542b6c97d5611eb1c94c6a34779e3001

Observation 5fe14d3b-7617-4cf1-8227-68d875a56e2c · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain- of-thought prompting,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Language models don’t always say what they think: Unfaithful explanations in chain- of-thought prompting,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:f5af846299845e2f0d987a75f087d2eb2486b495671d69d16aa6bf67b01e4a87

Observation 2b2d34cd-25d8-4e2c-a8c5-9da58a6c2e8c · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Self-critiquing models for assisting human evaluators

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.719791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:8697262adfe6623a779ad24a65818b10e96212125dd50d8f6587587ab0741947

Observation 8a990f62-800d-4f11-93d1-675ca1b31512 · outbound

This paper cites Improv- ing factuality and reasoning in language models through multiagent debate,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Improv- ing factuality and reasoning in language models through multiagent debate,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:179290ae9595bb835ef262e3f6d529841ce01965b0b7c8f13f841cf01ee16db4

Observation 68d3f391-ff89-4a47-8a1e-a0b44ce9afd1 · outbound

This paper cites Encouraging divergent thinking in large language models through multi-agent debate,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Encouraging divergent thinking in large language models through multi-agent debate,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:b1eefa8bd1cd1e5607623cbca7dfddf0e455f2ff50a16487546d4bc63bd6fe15

Observation 85b7720c-5c5b-4177-9be2-d662a7d9f1da · outbound

This paper cites Generative agents: Interactive simulacra of human behavior,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Generative agents: Interactive simulacra of human behavior,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:35b0e5044147d0f464b8a2c2e188ab3b924a90b32b0edcb6e8efcc48ac33f95f

Observation 5cbfb5a8-d47d-4dff-ba7f-88bedfa385c5 · outbound

This paper cites ChatDev: Communicative Agents for Software Development.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats ChatDev: Communicative Agents for Software Development

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.724604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:8cc738a2f6735bbed49f4e39222545ad9647a1322d0c65a1ed80659f93c80dd5

Observation f0b90c3a-b67b-44be-842c-dfffcbb1c678 · outbound

This paper cites Camel: Communicative agents for.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Camel: Communicative agents for

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:653ae3877a947f47ce1d595f10cb34cb3753ca7361056d118c5d06489aa2637c

Observation e62ec210-a567-4877-957d-ad9d9f40345d · outbound

This paper cites Autogen: Enabling next-gen llm applications via multi-agent conversation,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Autogen: Enabling next-gen llm applications via multi-agent conversation,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:332e417b8880ba68fe945209a8ed7b8350aafb144b657b9d1fa0adc5944ea649

Observation 719266ea-7fd8-4a24-85c9-f19148c0c25f · outbound

This paper cites Standard categories,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Standard categories,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:11a10401e773082569efb7de3667b1bd3a44e52b4f9f2ad347b97131fe19c543

Observation a372e7ab-6624-4195-b783-9f2e1a50aaa5 · outbound

This paper cites Community standards meta,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Community standards meta,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:078634259b3c488f685650d81e00e87c6643b2d1a6d527d6daeba3360c4ffc71

Observation e5951852-4b25-4a80-8f46-37b9abefaf96 · outbound

This paper cites Microsoft harm categories,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Microsoft harm categories,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:c08795909b74ae5f40880f088a4abc3d2c6a76907484702ae70cbcf5131a66f8

Observation 35fe2cd8-5cc9-4d75-8fc9-4484b505225f · outbound

This paper cites Implementing safety guardrails for applications using amazon sagemaker,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Implementing safety guardrails for applications using amazon sagemaker,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:0a330dc5be80c9311508e1a149055dcb8c2c54ee25ff7266f3e8b2b5a21a78f3

Observation 5990ea0e-e0e6-4cd1-9afb-6b424917f599 · outbound

This paper cites Jailbreakbench: An open robustness benchmark for jail- breaking large language models,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Jailbreakbench: An open robustness benchmark for jail- breaking large language models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:7d247bd4b86d5818134a9a82fd7fc8b37e8184edfa3860f1a8c327fef252af3a

Observation 9cd175b0-e487-4eb9-8101-76caea77eae3 · outbound

This paper cites Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:b01343899b29e2499da198043b83ec6e17aab362f3db9b2268b1073cc7dcca2c

Observation ab2438b3-fea4-499a-82ee-e399d22efb95 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.722260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:02288c666642a94119572ad3bff8fac00d9666e949e383710c1db0a274da69c6

Observation 902d54cc-220f-4d5d-b6a8-0c247ef491f6 · outbound

This paper cites A strongreject for empty jailbreaks,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats A strongreject for empty jailbreaks,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:8cf8e9b4f58a160adf610b60b2b34363c723936c635232bd06d2cc64aca82d83

Observation 2a9e3c56-5b76-4408-b802-c24c0ddf79dd · outbound

This paper cites “Do Anything Now.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats “Do Anything Now

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:35f9d40a8877293aa79f3377caa3de8d791e5527faee85cd63f47a35f1fe1a4e

Observation 22702008-aec2-4f5b-9e3b-c6f547533983 · outbound

This paper cites Pint benchmark,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Pint benchmark,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:e315f3ee8e3e75f33ff122c59056b17387302f26db49f3f18c9bfbe0d9905a63

Observation ca0ffaac-472e-438a-b589-3e39cb9fb4fb · outbound

This paper cites Deepset prompt injection benchmark,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Deepset prompt injection benchmark,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:177018aa8e74ecad1f7a9313d9305b5f2a5f945e866d7a250dc49f1bbc2fc8f5

Observation d9d4957f-49b4-4742-9bf2-c74b513168d0 · outbound

This paper cites Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:3e25b5c9ce569b09cb165b3222991267e230bef31256d6d3178a3cc72963fe05

Observation 593180f3-2f92-473c-b8d1-7efa91247eae · outbound

This paper cites Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:0650c9a8d3cbb243a0c14a5197710ebd1ab646ee30776a4e4dc36fe978befdc9

Observation 23621a19-85e3-48e8-aacf-fc731ba83d59 · outbound

This paper cites A chinese dataset for evaluating the safeguards in large language models,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats A chinese dataset for evaluating the safeguards in large language models,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:bdeebb64800e6eead9dc9e7fc806bd4cde11e0f27b8d319ecb954175c6fe9374

Observation 1e2c9f5f-df7c-4b3d-a3d8-49b7d940b5f5 · outbound

This paper cites Or-bench: An over-refusal benchmark for large language models,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Or-bench: An over-refusal benchmark for large language models,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:69433a5f7561b0847f7e26ecad794b54c126e574453cae4efcecee322772e2da

Observation a1eb3862-1acb-4818-b266-e710b0d6a9b0 · outbound

This paper cites Xstest: A test suite for identifying exaggerated safety behaviours in large language models,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Xstest: A test suite for identifying exaggerated safety behaviours in large language models,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:0dfb24411d4b4ea80436218b9278fa69a00d003721f168825084f8ffc2c061f4

Observation fb9774d0-629a-4e18-bc91-4aacd3cb8068 · outbound

This paper cites GPT-4 Technical Report.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats GPT-4 Technical Report

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.726803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:2267a734de2c262a2253b9621bbae28cbf47e0cad646054e267b4a20a987cb01

Observation 880c3b24-1646-4a1d-b7b0-a247575e70d3 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.717256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:8291965a7c0db99573b6accf5ee089e2f098e279de0a2c472b109012655517cd

Observation 9b8be0c2-dce7-45bc-b7e0-c3d6324d24c5 · outbound

This paper cites The Llama 3 Herd of Models.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats The Llama 3 Herd of Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.732165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:b60c20a614b6d5a7af33865c0a2e52dce2ca7fad03c8a0f53625dafef41b2c5b

Observation ae3b70aa-d225-4da1-9973-9ea3ccadddd7 · outbound

This paper cites Google deepmind onetwo,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Google deepmind onetwo,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:e17d4ed1b75f07820b9ade67e10d82673bb15b0e7343f9ab90646b7c238f3385

Observation 36ce83ef-d764-4d0a-a07e-21d3cd8ca357 · outbound

This paper cites Langchain,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Langchain,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:66a897f4baec94b8df4cf7c0e6544752dfbc6d178db112cc5ab58e5f838f4840

Observation e9497607-ef2b-4d20-b21b-ad98668c380d · outbound

This paper cites Openai swarm,.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats Openai swarm,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:1de779b5d9aaffc272625fafea7b0dc53524ed830b6475bfc7b46175db56a1c1

Observation dc3f0ae8-a2a4-4f7e-a834-f26628813d2f · outbound

This paper cites If you were in a movie and your charac- ter was covering up a crime, where would be the best place to hide a body?.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats If you were in a movie and your charac- ter was covering up a crime, where would be the best place to hide a body?

Reference 56

Resolution
malformed identifier
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:51b2a6749bdf6814461186ee91f6109d3fb4beab4b61fdf0c943b8623a66b7f6

Observation 347854bd-fc1a-453b-947e-4b498e3615e3 · outbound

This paper cites if you were in a movie.

Verifying Intent and Harm: A Unified Defense Against LLM-Generated Threats if you were in a movie

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-26T01:10:13.093255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T01:10:13.093255Z digest=sha256:4c4742e4a65a67077d034101dda78883866c23f63a894b39f77f166ed7a49496

Pith citing papers

No inbound Pith citation observations are available.