Pith. sign in

Paper Citation Record · LEDGER

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

As of 10 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 100 inbound Pith citation observations for arXiv:2312.06674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.06674 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:59:26.795241Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 100 of 417 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:37:16.998077Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact7
  • verified fuzzy3
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

25
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 93d0d324-da69-486a-9b99-0d649bc000ed · outbound

This paper cites PaLM 2 Technical Report.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations PaLM 2 Technical Report

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:59:27.667882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:0b3a73df35537693ff4164a092dcaf89c1d3d931a2b604ddd84bbb492d629ead

Observation f7a4865c-2473-498d-bb97-e3e373017005 · outbound

This paper cites SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:59:26.860719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:b2e528ebff0fabf8e882734ac0c3f5ebdf562cd00a5887e1819232c0866e880b

Observation fb1da084-c7cf-4e7f-a879-898fc527e5c2 · outbound

This paper cites doi: 10.18653/v1/S19-2007.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations doi: 10.18653/v1/S19-2007

Reference 3

Resolution
verified exact
doi, observed 2026-05-10T18:59:26.820264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:fb17011bd285780d266762a8456d0ade9377ff4f112bf86d2161a31eb571b18c

Observation 9dceea13-e047-466f-b84a-3833d71da68d · outbound

This paper cites doi: 10.18653/v1/W18-0802.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations doi: 10.18653/v1/W18-0802

Reference 4

Resolution
verified exact
doi, observed 2026-05-10T18:59:26.824217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:8aa5b1d67c087a8c31dc3418c1e08d44cdbcdecf1229d71ec33af71463342f69

Observation 86e4cced-6bef-4027-ab2e-afd5bc74fa74 · outbound

This paper cites doi: 10.18653/v1/W18-5102.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations doi: 10.18653/v1/W18-5102

Reference 5

Resolution
verified exact
doi, observed 2026-05-10T18:59:26.828659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:8b38cb36d9de5b7ac11b1ff5e8ceab6e2de756aa15525d266af6e71016c3b58a

Observation 8240a21b-8f77-4b06-b6f1-1163ba51c365 · outbound

This paper cites doi: 10.18653/v1/2021.acl-long.210.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations doi: 10.18653/v1/2021.acl-long.210

Reference 6

Resolution
verified exact
doi, observed 2026-05-10T18:59:26.812597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:80206e8ddde75d0d0da9a4fc0194aa8bd85d5e316d6bb6fbf63ed54ed8e6232e

Observation f125f07f-329d-45d6-abb3-4d877c9641f1 · outbound

This paper cites Exploring social bias in chatbots using stereotype knowledge.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Exploring social bias in chatbots using stereotype knowledge

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:59:26.857527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:42bb69ac5a5d85bb6f8dc2095b841b1c78b1c9384c279a77a9257d01ba5e706d

Observation d1b098cf-ee0e-484f-9dec-0e4a49f53bf8 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:57:53.077868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:3c24f5621ebec375953a53f0f7e90c5723e7afd2606d9618bc3fc3ef6f3f5b83

Observation 71ab6f0d-57b5-408d-a33c-16d1e58fe3b7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T18:59:26.847187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:fbd7557a8371158254eaf280541206a32ad2b1f2d2e31b1a9be42ed7638513b7

Observation 48e938c7-2e65-4984-a082-18817769b950 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:36:47.138740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:ae4ec42549936430a28d480cee76a6de7973068597c0ec9e09dbad54ef792219

Observation 253d1275-8b8b-4783-9e9b-7211aae604d7 · outbound

This paper cites SemEval-2019 task 6: Identifying and categorizing offensive language in social media (OffensEval).

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations SemEval-2019 task 6: Identifying and categorizing offensive language in social media (OffensEval)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T18:59:26.863640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:1335847abde1a57dd396cee6bb43fa6ff126317d9bab6189e516739503292e58

Observation 0c0656c3-286e-4848-bba1-7e3caa68f16e · outbound

This paper cites Zampieri, S.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Zampieri, S

Reference 12

Resolution
metadata mismatch
doi, observed 2026-05-10T18:59:26.816547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:59:26.795241Z digest=sha256:373c49c2af37f0273630fd5fb9634e0cbec95ffbe40f6fcc7ea4f1b5f58e1ab7

Pith citing papers

Observation 59b493df-1188-4c0d-9cf6-e209f1c23bf8 · inbound

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks cites this paper.

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-14T17:11:00.769268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T17:11:00.639293Z digest=sha256:46417c2a8692c7c1803197a87601fcd238976b6efd5a349a903bbeda2dd115ad

Observation b032a506-52d4-41cf-8c2f-027640a9e8d0 · inbound

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models cites this paper.

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:08:05.602244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T06:08:05.386345Z digest=sha256:9991ec4331307e269cb947e85050d86753376d1495ac8792fd9bc78fb95790a7

Observation 72db41cc-e8d5-43e0-9cc8-5b8bcb80f5e0 · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T16:25:14.849301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:ab73e0f1881baac799a93541bdf65ba496d284766bec48865791765ec2ae1dc0

Observation 7062158a-35b6-4172-a476-abd86862d35f · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:17:39.507526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:1b62985cc4d6cfba462ea55948d3638e4d3fbc5ea2a1a0b698658e4a49e817bf

Observation 23e47189-e4f6-4d4e-b0ae-bf4b1b6709ae · inbound

Bias in Large Language Models: Origin, Evaluation, and Mitigation cites this paper.

Bias in Large Language Models: Origin, Evaluation, and Mitigation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:08:12.380117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T17:08:09.267577Z digest=sha256:8efad890e9d1b870e06066f098dd9334837638d27db6667f376dbc4d1ba8612f

Observation c077c955-8436-4525-b0ca-a0fd153a664d · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:38:45.530796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:e87cc00a888264da4ca538de230b9673244d37c1f23e370171ff2e7a27df935e

Observation 1cb77ed6-af48-4b3f-a363-3dafa0f2183a · inbound

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation cites this paper.

Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:16.998077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:16.998077Z digest=sha256:bd026b0275d533662438c11dae3dddde3ca61090ad5cfdb4903a2cacae497ebd

Observation 9e107d9a-5896-4eee-ace5-17ba50ef1103 · inbound

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models cites this paper.

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T00:12:04.851548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:12:04.851548Z digest=sha256:511f89d145ab1c056fccb26723050170d496c0a3200c8c91082738d873d74a98

Observation 8ddc7d46-66fe-495f-a8fc-2e96225335ce · inbound

o3-mini vs DeepSeek-R1: Which One is Safer? cites this paper.

o3-mini vs DeepSeek-R1: Which One is Safer? Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T23:32:16.870220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:32:16.870220Z digest=sha256:2ac59868648a01c06844cb0e53b60938278f4e7faaf4f180073c24f033bc0a60

Observation 9faece8a-cb64-4cee-8fae-f19b7dd27043 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.723901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.723901Z digest=sha256:31fd34a3d3add9a65d8856393d87d8344550d7571643eb99e1a65061ccf9c25b

Observation 706dce9f-555d-4e57-8484-da58ea72d51c · inbound

Peering Behind the Shield: Guardrail Identification in Large Language Models cites this paper.

Peering Behind the Shield: Guardrail Identification in Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-23T03:45:21.328838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T03:45:14.234545Z digest=sha256:39c9f2317d24f1d80e22eb939b8251e846e61213c68bfd7f888a0a80bd83ec8a

Observation ea8c3026-61a9-488a-8974-84c98a1031f0 · inbound

Adversarial Reasoning at Jailbreaking Time cites this paper.

Adversarial Reasoning at Jailbreaking Time Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T14:49:08.742936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:49:08.742936Z digest=sha256:977edfde247615ec669c883a3acd8eb34e16aa3646a2d6ea29ad59ab2abb6bbd

Observation 98c68720-f744-4ff2-869d-311c1b393377 · inbound

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing cites this paper.

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:14:34.027501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:14:34.027501Z digest=sha256:9edbb13b52794838a6f997b36448b50df3c016f7d3e25f1ce69c29f40f382dfa

Observation 26f6c5a6-02f1-4e9b-b6f2-3e021b82b634 · inbound

Position: Adversarial ML for LLMs Is Not Making Any Progress cites this paper.

Position: Adversarial ML for LLMs Is Not Making Any Progress Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T12:47:21.691851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:47:21.691851Z digest=sha256:5769b76c9078ec79c5409cf571a58e1f67b5cd466cc8b9c13eecbe990d789238

Observation 6975a179-af4f-48af-95dc-e328b3791da9 · inbound

MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents cites this paper.

MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:58.258918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:06:58.258918Z digest=sha256:c6f6a36c8c0c099cf0d0d957352a24015075d276ce0f4f9ad8e8354cdee3a7fb

Observation a4e41a44-f263-41ee-84a9-6447e16b6d3b · inbound

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring cites this paper.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.911523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.911523Z digest=sha256:8875b34bddd970e3e780715c4f74869d5b0ac740be7f877b497c6ef32489de6c

Observation 0ec5b2f6-a62a-446a-b3fa-fcddc1de3ada · inbound

InSTA: Towards Internet-Scale Training For Agents cites this paper.

InSTA: Towards Internet-Scale Training For Agents Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T14:24:50.338817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:24:50.338817Z digest=sha256:7d6bb2e6776cbb13e634c18eeb013140e561ac046409b4bce9cb1413888670ca

Observation a452f9e8-410d-4564-a5ab-a0db470e947b · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.574357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.574357Z digest=sha256:4ddbfe15109d1294153b40b90f6792df9d4a2352f54eb335fae69c285eb51545

Observation e0efeb87-91b4-4b79-8353-b4eda5922f4c · inbound

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences cites this paper.

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T10:23:06.114397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:23:06.114397Z digest=sha256:36ba5e241ad29c5610427ee87898c16940c3698376ebcf27ce85050e7429f70f

Observation 98431140-7e89-4d59-ab0a-9a96b91afff6 · inbound

Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks cites this paper.

Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T04:37:28.388533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:37:28.388533Z digest=sha256:0ed00f9f28ff54bcee0c2d1a3c3c3c3207b0756452aa2149e7e8bab3306a1eb5

Observation 23ca78a9-997c-4099-b4ae-02e4c3ff15ca · inbound

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions cites this paper.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.269258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.269258Z digest=sha256:5eaaaf21ede65567d589a6facf3f4bc98806831c43b2d04f78432db14ee3307c

Observation c20642f4-1de5-4ced-adc1-0fe61786e329 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 155

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:20.013350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:20.013350Z digest=sha256:608cad35943f89a2b19feb10fe7600d8d91e88b5d56601cb9b0ee465530bd8f8

Observation 218d685e-3c0d-44e7-84ef-f5d0fec4d481 · inbound

Responsible Federated LLMs via Safety Filtering and Constitutional AI cites this paper.

Responsible Federated LLMs via Safety Filtering and Constitutional AI Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:22:25.232184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T02:20:07.800524Z digest=sha256:ea19db4d16de8a97e3552853929038754186964717f3507f598c18ac56a95507

Observation 91e8a443-2ab7-4a2f-85d5-5e3a7e316920 · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:32:22.485463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:8343981fccc8aea6cc0741bccd3e4bb795f182da3d80d92b9ead2738e40519ea

Observation b95f0903-988f-4d32-ad2b-3001027e94aa · inbound

AI Failures in the Eyes of the Downstream Developer: A First Look at Concerns, Practices, and Challenges cites this paper.

AI Failures in the Eyes of the Downstream Developer: A First Look at Concerns, Practices, and Challenges Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:15:13.131026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:13:37.570261Z digest=sha256:52ce59d0897cb08661c7394e5520bd9c4dd81f8551605bd907d3ba7d94a49026

Observation eab566b9-d1bb-45e1-ad28-c55fc678c3a2 · inbound

Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics cites this paper.

Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:32:12.925336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T22:27:18.533162Z digest=sha256:3317d7a6873745b79f53b6a84cd916aa773cc190a2bf8d91fbd82a9f46436991

Observation 59b5a60d-48a7-4428-9392-cbb269882575 · inbound

Progent: Securing AI Agents with Privilege Control cites this paper.

Progent: Securing AI Agents with Privilege Control Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:12:08.513942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:09:51.782808Z digest=sha256:7d1bcd106d2fdba2b52271248e0cadb8769eabfb440374451ae2e547da6d363d

Observation 1cccfa9b-1361-422b-be2c-5edeaf6d96aa · inbound

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs cites this paper.

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.943099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.943099Z digest=sha256:4b1e01e50649de7a606ddfffd2da425db368a1a54cb855ae87e6c08914bc0af2

Observation b7270254-4aab-4948-9855-19b980d6df2c · inbound

Advancing LLM Safe Alignment with Safety Representation Ranking cites this paper.

Advancing LLM Safe Alignment with Safety Representation Ranking Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.261712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.261712Z digest=sha256:e03ace1adc6370cbc9e9f5581fcb47450ba0e7c49d314b40b5420d1f7f9cf764

Observation 52d6f016-d02b-4072-953a-d445c93b68cd · inbound

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval cites this paper.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.609882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.609882Z digest=sha256:85f78a11532106c4a3b28619f4d4acb4c8c4d6fa9347499f7f6d1faf589be223

Observation f47dd509-f826-4bfd-91bd-15cb70df1707 · inbound

LLM-Powered AI Agent Systems and Their Applications in Industry cites this paper.

LLM-Powered AI Agent Systems and Their Applications in Industry Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:06:38.122153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T14:05:54.535411Z digest=sha256:2724e8a6ab12d84e333c451026ef5ce1b72de7473d77779a06b57d62ec8dcea9

Observation 3d13f261-fc53-4db5-82a8-4d8512cc1b00 · inbound

Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models cites this paper.

Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:42.407078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:42.407078Z digest=sha256:24bf265d296472d5eb137511d0bc14abf9dd36d917209e9847110c004be74ed3

Observation d68efde7-02c4-47b8-9d6d-7f4ab01dc51b · inbound

EnSToM: Enhancing Dialogue Systems with Entropy-Scaled Steering Vectors for Topic Maintenance cites this paper.

EnSToM: Enhancing Dialogue Systems with Entropy-Scaled Steering Vectors for Topic Maintenance Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:01.507759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:03:01.507759Z digest=sha256:b60a3062bfe6bee30a4f0868e80a83a794cbf6148e084eea1edf107d18b8ef93

Observation 44a28b61-7ba4-4a60-89e4-b959bba1cbc3 · inbound

LLM Access Shield: Domain-Specific LLM Framework for Privacy Policy Compliance cites this paper.

LLM Access Shield: Domain-Specific LLM Framework for Privacy Policy Compliance Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:01.235547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:01.235547Z digest=sha256:7cc4dd0e51b1e0fb608dd77332e69d691584d78467b60c952a247a03327baf25

Observation 6f894b69-d7be-4710-a0a9-f5fe74358a9c · inbound

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming cites this paper.

MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:58.418462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:09:58.418462Z digest=sha256:a433574dad00f61b1e506a0ada1c0badee76ae275224c170b88d6b4b2715d8cc

Observation cf682473-c924-4060-b14b-99535db4bba4 · inbound

Content Moderation in TV Search: Balancing Policy Compliance, Relevance, and User Experience cites this paper.

Content Moderation in TV Search: Balancing Policy Compliance, Relevance, and User Experience Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:08.031018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:08.031018Z digest=sha256:c12231da35a3c4eb5178a964288addc4dbb9e63446e2d3b5988cd6d1cfb05162

Observation e1a5f60f-da92-4416-b5b0-7a6f08b3ef81 · inbound

Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs? cites this paper.

Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs? Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:10.176461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:50:10.176461Z digest=sha256:2aa3c25e08e951ae8e22f5d5da524bf62199bfb6d6ee6063916039801f142bc0

Observation 533de5ff-de66-4a20-bafc-ba2d1bba421f · inbound

Mitigating Deceptive Alignment via Self-Monitoring cites this paper.

Mitigating Deceptive Alignment via Self-Monitoring Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.405155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.405155Z digest=sha256:5f508d050cd6a8f62937d0dc8f24055a1c988a215a9c0a28611ad440d2935532

Observation a94f58ae-8610-4a7f-b05b-0dd5f644d0a1 · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.510563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.510563Z digest=sha256:23c3a6d0941b080254e49beb6c7973f26a91831e66e6dedcc5aa5c2d8a5ce408

Observation cd501efa-f61b-4fbe-bb0f-4f56f228df20 · inbound

VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models cites this paper.

VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:17.802792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:13:17.802792Z digest=sha256:3e4694e8897248ef13394a2db036ac4f3f7305dcce62b0730fdc2bb592008a11

Observation acab8d73-6746-4415-b487-bc4857ac4fc4 · inbound

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment cites this paper.

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:53:03.593930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:52:36.688263Z digest=sha256:5111333f4273bece155a0e063d42777c0e6944bc05c6129bf8fe2ba712913443

Observation 3bf6bcac-375c-47b1-a898-025e275b64d8 · inbound

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities cites this paper.

Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:59.394540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:07:59.394540Z digest=sha256:f9014367a9bfa27d5f5c727cad31dff14cace1b1a8dff3e708635db256b87f4e

Observation 0a34b9a9-a42a-4873-81c9-41aee11e8d93 · inbound

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning cites this paper.

SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:34.771964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:02:34.771964Z digest=sha256:50e12bab8cf23a26f6bf69f9d39466559825b06722ba6dd167eae8c5e04c3a4c

Observation 9949848e-882c-4c87-ae21-d591081d1305 · inbound

Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution cites this paper.

Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:23.191027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:23.191027Z digest=sha256:5595763dc61d6d1b4fc607212cae2280334ae08e235e61db668a34276c31263b

Observation 312bed32-3c6c-4f5a-a6d9-bba5dc19d8e0 · inbound

AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output cites this paper.

AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:13.320214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:29:13.320214Z digest=sha256:38073e04aac0cc6eec42a01d1477edbc27c84a1d00c0a33f8c8ad0cd84b15480

Observation 88b28cec-7931-4bd0-ab13-fd3bf806ad74 · inbound

Should LLM Safety Be More Than Refusing Harmful Instructions? cites this paper.

Should LLM Safety Be More Than Refusing Harmful Instructions? Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:47.280600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:27:47.280600Z digest=sha256:022a8f486d657a0567fc755c4d6f0c899bfc000bb0cb988ef252a52e6200f82f

Observation 88d8b3d7-9188-4a13-9f55-20c4d875c63c · inbound

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems cites this paper.

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:02.982826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:02.982826Z digest=sha256:58ce9017d6cd673bed7363e4a838b7c660e3c7e40e80620e15e177a3f64624d9

Observation 3ea8fead-5ff6-4fb7-9de4-da9df154c32d · inbound

A Red Teaming Roadmap Towards System-Level Safety cites this paper.

A Red Teaming Roadmap Towards System-Level Safety Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:19.827281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:19.827281Z digest=sha256:97fcdd728d355c33aed979534712c3d2121599f1f4c7421b8f8716405f87c9c1

Observation ea34ee52-42ec-4c1e-ae5d-96a5e1f37298 · inbound

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law cites this paper.

From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:55.100445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:55.100445Z digest=sha256:cfacdb182c30d7a9d8c89d17726bc42b98fcdca4d41998ea70d2e972597c5f29

Observation ff8c9001-6c80-4574-b97e-7e4483d26373 · inbound

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models cites this paper.

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:06.331292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:06.331292Z digest=sha256:d9ca4a6094087ae3fca1f2596586dd8b9cc080bb1e86e922e3f41a8dbdded279

Observation 9bb5367b-8781-4928-bfc9-7a99b2465ab7 · inbound

JavelinGuard: Low-Cost Transformer Architectures for LLM Security cites this paper.

JavelinGuard: Low-Cost Transformer Architectures for LLM Security Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:25.411365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:25.411365Z digest=sha256:7baf8e1175c7ed8537c9207bc035ae277d06660cf1c78b4b896e6ed926e1fe22

Observation 424edd04-e787-49a4-8572-4b0d520bdb7d · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.886326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.886326Z digest=sha256:7b222b043f0bbc257afc04207b60d14f39330b8fda9999614b058989ee986498

Observation 45db551d-7a63-4765-a360-9c5e1771aa46 · inbound

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models cites this paper.

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:51.256225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:51.256225Z digest=sha256:cafe4d00a528548c6b359c372693bda0e31f94d1fb67a8e14b7a518ec604567d

Observation 427cf786-071f-45a2-b516-8bde402013dc · inbound

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems cites this paper.

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:32:36.550081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:32:36.550081Z digest=sha256:677fb21e0c2baa653a597ffeba265b58d43d11b971c5734226c36dd2aea69443

Observation 855553a9-f7f4-4bf5-a382-1af3ac415157 · inbound

"I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products cites this paper.

"I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:27:19.609985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:27:19.609985Z digest=sha256:4b5db80b637800c6aae2110a8f8c0b6dee13dc53ae7699c6b1954a9f564f9f64

Observation 1683a507-e3d5-4e7f-b7a0-19a007254198 · inbound

PL-Guard: Benchmarking Language Model Safety for Polish cites this paper.

PL-Guard: Benchmarking Language Model Safety for Polish Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:53.313221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:53.313221Z digest=sha256:09a92fe23629ef4b4aa68e4698d1cf2ec91b1da26827e0c5f24fd6a9640343b5

Observation f9e31c90-214d-4919-8250-b2362123671b · inbound

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem cites this paper.

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:42:12.850642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T08:40:56.186349Z digest=sha256:7afe0f86e2b3e823dddf93e2a008063d0a0543626d25e8093da35d9bd5722a70

Observation 8bb9cdef-5d9f-4a89-bffc-0f18e730a2b8 · inbound

Optimising Language Models for Downstream Tasks: A Post-Training Perspective cites this paper.

Optimising Language Models for Downstream Tasks: A Post-Training Perspective Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T22:44:44.066010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:44:44.066010Z digest=sha256:d903d348493ffee33c72fb39e2feadaebd19751ca4018f0b2754f8ad09c6b462

Observation 47456cfa-c3cc-4b83-a3f7-3921fdbc22c4 · inbound

VERA: Variational Inference Framework for Jailbreaking Large Language Models cites this paper.

VERA: Variational Inference Framework for Jailbreaking Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:07:03.267607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:07:03.267607Z digest=sha256:66b6012aa43e26e27046a913cad09c8a910eb6bbca2e89388ad79343f6fb9af6

Observation 934cfb72-d162-4124-9093-7f5d8f904889 · inbound

On the Surprising Efficacy of LLMs for Penetration-Testing cites this paper.

On the Surprising Efficacy of LLMs for Penetration-Testing Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:10:04.541993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:10:04.541993Z digest=sha256:12a2c1b9746d3060a3831e5940dce25bdbac0e3f761fd872a9e6bbce0fa9db95

Observation 52dcb26a-0105-4bc9-ab25-c05e1c2d69d7 · inbound

AI-guided digital intervention with physiological monitoring reduces intrusive memories after experimental trauma cites this paper.

AI-guided digital intervention with physiological monitoring reduces intrusive memories after experimental trauma Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:05:27.679677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:05:27.679677Z digest=sha256:fd8a7de72faad2ce3e8f7d44398c78583286f1660f5bfb4ea2be301225f8b0e6

Observation abbb5f1d-fe91-4e0b-b262-650f9a19e1a3 · inbound

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation cites this paper.

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:53.281864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:53.281864Z digest=sha256:32075885f44371bfe5ea2b2ce787080d38d82cae974873f9fa5f00f53f4911a7

Observation 7353ff7f-513f-4b87-80c8-4a7f4116fc1d · inbound

GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models cites this paper.

GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:26.838437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:26.838437Z digest=sha256:79d661ac662a778ef107b723c515d3354f2fa66029bae48f9bb479e7cc595c5c

Observation 224d1a9d-36e6-4236-9ac3-18498944b9ae · inbound

ModelCitizens: Representing Community Voices in Online Safety cites this paper.

ModelCitizens: Representing Community Voices in Online Safety Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:00.841084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:31:00.841084Z digest=sha256:bb25f5cd0642c7b9b8d23f359f82b943c8bb1d8db4d27e2f9ede512ce0a035ed

Observation d70c4c11-6f58-4a53-84d6-f1905a37804f · inbound

Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms cites this paper.

Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:22.912261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:12:22.912261Z digest=sha256:0d192824fe980ea462c2d26491a2e27dafee9a4d3510fbdddee180cbc0dbf714

Observation deafb81e-b3c7-46e4-95ae-f81a2e7c1048 · inbound

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems cites this paper.

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:43.104007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:28:43.104007Z digest=sha256:0aa97bb32e657bccb1fae95f69a651169b91233b70fc50ef76b5200e90db9a3e

Observation 7362b693-3a9c-47e2-96a8-c13be50fd027 · inbound

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models cites this paper.

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:10.542110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:10.542110Z digest=sha256:c78d91aab56ed42059b260e12e3af63abcd619bd6faca49ff7cd1f8e2067da01

Observation 0f1391c9-8c4e-42ed-b5c0-bf4ce8d633ac · inbound

Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications cites this paper.

Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:49:41.341531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:49:41.341531Z digest=sha256:8f58797cc4da6fe54e988096e9f231d85ff0088aef7ab52ade191a52c418ad76

Observation a92ffb7e-80dc-400f-bbf3-ac1b6d775d14 · inbound

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models cites this paper.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.391760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.391760Z digest=sha256:97eb3ab9dc82e28dcc61aa7ec35324efd9ec611ac3a05c84ad569efb28a4f1e0

Observation ea9101c5-a663-40e6-86d9-309a01082db3 · inbound

LLMs Encode Harmfulness and Refusal Separately cites this paper.

LLMs Encode Harmfulness and Refusal Separately Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:15.860934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:03:15.860934Z digest=sha256:d65416a0875709b3f8b326cfc46402f406cc56dc8dc6e65d7730641d7279bd2b

Observation 210b3c37-d214-4c21-9fc6-bc16dd1be8f5 · inbound

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers cites this paper.

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:28:50.565875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:28:50.565875Z digest=sha256:3d20ae8e6ccf7a1d0c5f3a618b636e227e77e88ebb849ac65a546d19683e5272

Observation a00e24ce-97d0-4933-b7d5-cd408a718413 · inbound

PrefPalette: Personalized Preference Modeling with Latent Attributes cites this paper.

PrefPalette: Personalized Preference Modeling with Latent Attributes Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-06T16:27:57.229187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:27:57.229187Z digest=sha256:47dfbc10d9efcbda01db9903e8f2a8a8022a33b1c67ef6ed82188ba0cbeca362

Observation 30f7ec07-3d44-4cf9-95d7-a88ae65cc2ed · inbound

WebGuard: Building a Generalizable Guardrail for Web Agents cites this paper.

WebGuard: Building a Generalizable Guardrail for Web Agents Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:12:48.722673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:12:48.722673Z digest=sha256:c189b829789aae4efd058a0b63933274b8152c0f5a502292d43358bc52deb3bf

Observation 2de32363-6a72-45f3-a94f-7fa96500ef52 · inbound

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning cites this paper.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.454863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.454863Z digest=sha256:f50be38f11b2e0a947f81a095e39a96af82638fc6638cb3ca0f65f6e2fbcba4a

Observation 8012d4b5-ccb6-40cd-a89d-16d44398b3f4 · inbound

ADEPTS: A Capability Framework for Human-Centered Agent Design cites this paper.

ADEPTS: A Capability Framework for Human-Centered Agent Design Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:11:47.674139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:11:47.674139Z digest=sha256:17169f01fd52131b873fe21f9e7394753b1540369b77de8e7a665d3a50a1c27a

Observation f1902be7-ee35-4736-88f9-ce89c36aae13 · inbound

TorchAO: PyTorch-Native Training-to-Serving Model Optimization cites this paper.

TorchAO: PyTorch-Native Training-to-Serving Model Optimization Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:16.568903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:16.568903Z digest=sha256:506493fe0202b7be1a0897382ea35739ea1de64cbbe11c347e514a03a4618f91

Observation a6528721-9175-4293-aa54-96cae918c15e · inbound

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report cites this paper.

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:22.834304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:22.834304Z digest=sha256:0251500c80f99a88bd53da4c5a8af1a4228bba0a7db2a848c1deb92e82a37b3d

Observation d3197835-56a4-4ad2-9624-0d1e5b01b33f · inbound

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law cites this paper.

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:21.448244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:21.448244Z digest=sha256:a5b6f72cd22c12cdd62c3d0cdafe629af0ae908927500cd9438bef62816ac14b

Observation 111788cd-b8c9-48a2-b775-0811b8f2e177 · inbound

OneShield -- the Next Generation of LLM Guardrails cites this paper.

OneShield -- the Next Generation of LLM Guardrails Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:04.323451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:17:04.323451Z digest=sha256:daeba97d8eaef7436bd08fb2608d6555a8a846cedf8b58a4ca26e72e25ccdce3

Observation 03bd2d9c-1211-4ebb-80a4-726bfce0f132 · inbound

Agentic Web: Weaving the Next Web with AI Agents cites this paper.

Agentic Web: Weaving the Next Web with AI Agents Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-06T13:05:36.990443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:05:36.990443Z digest=sha256:37e219a451f7dfd240c13715c6f5a5cbef7aa66e797afe5392d22d1bb0cfdd2d

Observation 848fe526-4132-4bf4-9e22-842a315f02bf · inbound

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation cites this paper.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:34.709673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:34.709673Z digest=sha256:8954ab6b952cc1e0a026deb76ecdfa1afa1cd6179a392906773a2484e25fdb9b

Observation f8aa02b0-d615-4334-b707-09bfeeaf163b · inbound

The Problem with Safety Classification is not just the Models cites this paper.

The Problem with Safety Classification is not just the Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:25:21.959665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:25:21.959665Z digest=sha256:6d7595231dd23705ca261595bdc9c69dc47e5e8c0d5eb8031509008b4c19637a

Observation 9299ee83-44bc-4628-8e41-0eac91cd7080 · inbound

Libra: Large Chinese-based Safeguard for AI Content cites this paper.

Libra: Large Chinese-based Safeguard for AI Content Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:17:38.676661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:17:38.676661Z digest=sha256:30979468cb1e5619004e4ada9366f2ae0bab41f717b060948c4e993ac445623a

Observation 3b014c58-1e32-4a5c-a8a7-599a49863fbc · inbound

RedCoder: Automated Multi-Turn Red Teaming for Code LLMs cites this paper.

RedCoder: Automated Multi-Turn Red Teaming for Code LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:44.729289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:00:44.729289Z digest=sha256:43171981a816fbff934557e30648529807d5b5f963b19f03bb802b8a18ddfa9f

Observation 0f1e94ae-9858-419d-8101-cc7efc2b17cb · inbound

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report cites this paper.

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:57:29.271933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:57:29.271933Z digest=sha256:2bbd8f45c21db3aa807d9ac589068af997904fb9a0b491977a45941819fd4d26

Observation 4d7dbf2d-5024-4188-9bf2-190d7779a207 · inbound

Provably Secure Retrieval-Augmented Generation cites this paper.

Provably Secure Retrieval-Augmented Generation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:56:20.264246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:56:20.264246Z digest=sha256:790d595c3f3738b52bcf98046d6294b4e8745782ae5809e843b698bb93696910

Observation 7ded2693-e7ac-44c7-9a55-76436bba48b2 · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 124

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:25.435933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:25.435933Z digest=sha256:cc3649042817ad92e4d1568d903a46fef0345066ca77a5a170c038e1a5f881e1

Observation a5365e4f-0f2d-47a2-8d76-5b266eafd26e · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:04.161107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:04.161107Z digest=sha256:bfa8fcfdb080a74097f6c67ed51c9262fa8f24efd6a8be6bde176854688873dc

Observation f5881f87-bd05-4c12-9048-d554c8f0467d · inbound

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection cites this paper.

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T22:24:23.377564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:24:23.377564Z digest=sha256:27ee1e626ab8fa9122e327e7aae3f0a81dda3e19c42cb5bc607f28ddbdb6a0a6

Observation d71ee36e-7699-4e13-b5b2-fbbdf54fb341 · inbound

Blockchain Network Analysis using Quantum Inspired Graph Neural Networks & Ensemble Models cites this paper.

Blockchain Network Analysis using Quantum Inspired Graph Neural Networks & Ensemble Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T21:22:37.200923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:22:37.200923Z digest=sha256:e6f498242b4cc848df5db896be197236a10c8df79dd50daa463405b7cfa1423b

Observation f8c156e9-ba9d-462d-aeb5-7cb3cb65fc2e · inbound

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges cites this paper.

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T16:40:06.696675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:40:06.696675Z digest=sha256:4c81b668042ced1b881583707fe8b8febd7a4e19bf5e60cb51c207ff52d97f52

Observation 1b002207-66d9-43ba-9015-89a5141a2db7 · inbound

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? cites this paper.

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:48.568286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:00:48.568286Z digest=sha256:2e2a5f6c6a2b8af2e8bdfc1fd70f722f37eba1ad0c22a8214b55d606d934c5e9

Observation 70fe987c-33ce-414e-964a-369b40f947d5 · inbound

Reliable Weak-to-Strong Monitoring of LLM Agents cites this paper.

Reliable Weak-to-Strong Monitoring of LLM Agents Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:50.137911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:50.137911Z digest=sha256:cc563bc6d98ab1bf59aa9f4407b1bd268a6bb63ac6fc0e82e03cec7637053b0b

Observation d2bb5a08-5439-4e4e-9501-cb9dbe177c3a · inbound

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement cites this paper.

IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T15:20:14.192838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:20:14.192838Z digest=sha256:651eabf37d874666eefdfa48b067556b88c9881de4682ac33747bf879e7e85f0

Observation 6ade512d-1ab5-4802-8155-1fa8a76cc5a1 · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:41:50.489469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:5d102bde7b7769da1e3171a77bc11281f7d23d3fedec81b3fe610c9c4d2d66d4

Observation 1e67f2ad-fd86-4f28-b849-5f4af713079d · inbound

Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities cites this paper.

Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T11:19:11.199092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:19:11.199092Z digest=sha256:53ecfdc0170d0ce896e3abaac61c86f8c567fc11ce1f1a2cb5a71dba7a634f2e

Observation fee73ba9-20b5-448a-8f5d-59e87ca16c45 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 270

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:08.233150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:08.233150Z digest=sha256:3ad019b7a0b8415cbeaf5e0251c5ff77651245563dc1d878b769d7df7b9d5a75

Observation 94554423-d12c-46ec-8c8f-7be52ffd7e57 · inbound

Scaling behavior of large language models in emotional safety classification across sizes and tasks cites this paper.

Scaling behavior of large language models in emotional safety classification across sizes and tasks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T11:25:23.020751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:25:23.020751Z digest=sha256:aa48d887739618f899bad27d4f0c6bba2f1314c3bc0503109c01465f2f7291d7

Observation 54805df2-0907-45ad-a454-492f38e700bb · inbound

Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair cites this paper.

Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:29:20.630948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:29:20.630948Z digest=sha256:8296b2bf2e0e2b40c96aa15860c3afbd2f618be620cdf3a14f8aa8f5bde7df5b

Observation a43c34b6-1182-4430-99b8-28e4822e7418 · inbound

Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment cites this paper.

Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T05:09:39.841874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:09:39.841874Z digest=sha256:55e63beda4cdfb695acc7117607c67c0179ec208f094aeb96200f86b3b2fd14b