Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:59:26.795241Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 100 inbound Pith citation observations for arXiv:2312.06674.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T18:59:26.795241Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T04:37:16.998077Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
12 of 12 outbound references displayed
External citation measurements
25
pith, observed 2026-08-05T02:28:24.338817Z
Observation 93d0d324-da69-486a-9b99-0d649bc000ed · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations PaLM 2 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f7a4865c-2473-498d-bb97-e3e373017005 · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations SemEval-2019 task 5: Multilingual detection of hate speech against immigrants and women in Twitter
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb1da084-c7cf-4e7f-a879-898fc527e5c2 · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations doi: 10.18653/v1/S19-2007
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9dceea13-e047-466f-b84a-3833d71da68d · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations doi: 10.18653/v1/W18-0802
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 86e4cced-6bef-4027-ab2e-afd5bc74fa74 · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations doi: 10.18653/v1/W18-5102
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8240a21b-8f77-4b06-b6f1-1163ba51c365 · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations doi: 10.18653/v1/2021.acl-long.210
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f125f07f-329d-45d6-abb3-4d877c9641f1 · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Exploring social bias in chatbots using stereotype knowledge
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d1b098cf-ee0e-484f-9dec-0e4a49f53bf8 · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71ab6f0d-57b5-408d-a33c-16d1e58fe3b7 · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 48e938c7-2e65-4984-a082-18817769b950 · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 253d1275-8b8b-4783-9e9b-7211aae604d7 · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations SemEval-2019 task 6: Identifying and categorizing offensive language in social media (OffensEval)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0c0656c3-286e-4848-bba1-7e3caa68f16e · outbound
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Zampieri, S
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 59b493df-1188-4c0d-9cf6-e209f1c23bf8 · inbound
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b032a506-52d4-41cf-8c2f-027640a9e8d0 · inbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72db41cc-e8d5-43e0-9cc8-5b8bcb80f5e0 · inbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7062158a-35b6-4172-a476-abd86862d35f · inbound
ShieldGemma: Generative AI Content Moderation Based on Gemma Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 23e47189-e4f6-4d4e-b0ae-bf4b1b6709ae · inbound
Bias in Large Language Models: Origin, Evaluation, and Mitigation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c077c955-8436-4525-b0ca-a0fd153a664d · inbound
Cosmos World Foundation Model Platform for Physical AI Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1cb77ed6-af48-4b3f-a363-3dafa0f2183a · inbound
Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e107d9a-5896-4eee-ace5-17ba50ef1103 · inbound
Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ddc7d46-66fe-495f-a8fc-2e96225335ce · inbound
o3-mini vs DeepSeek-R1: Which One is Safer? Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9faece8a-cb64-4cee-8fae-f19b7dd27043 · inbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 706dce9f-555d-4e57-8484-da58ea72d51c · inbound
Peering Behind the Shield: Guardrail Identification in Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ea8c3026-61a9-488a-8974-84c98a1031f0 · inbound
Adversarial Reasoning at Jailbreaking Time Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c68720-f744-4ff2-869d-311c1b393377 · inbound
Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26f6c5a6-02f1-4e9b-b6f2-3e021b82b634 · inbound
Position: Adversarial ML for LLMs Is Not Making Any Progress Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6975a179-af4f-48af-95dc-e328b3791da9 · inbound
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4e41a44-f263-41ee-84a9-6447e16b6d3b · inbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ec5b2f6-a62a-446a-b3fa-fcddc1de3ada · inbound
InSTA: Towards Internet-Scale Training For Agents Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a452f9e8-410d-4564-a5ab-a0db470e947b · inbound
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0efeb87-91b4-4b79-8353-b4eda5922f4c · inbound
Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98431140-7e89-4d59-ab0a-9a96b91afff6 · inbound
Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ca78a9-997c-4099-b4ae-02e4c3ff15ca · inbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c20642f4-1de5-4ced-adc1-0fe61786e329 · inbound
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 155
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 218d685e-3c0d-44e7-84ef-f5d0fec4d481 · inbound
Responsible Federated LLMs via Safety Filtering and Constitutional AI Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91e8a443-2ab7-4a2f-85d5-5e3a7e316920 · inbound
SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b95f0903-988f-4d32-ad2b-3001027e94aa · inbound
AI Failures in the Eyes of the Downstream Developer: A First Look at Concerns, Practices, and Challenges Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eab566b9-d1bb-45e1-ad28-c55fc678c3a2 · inbound
Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 59b5a60d-48a7-4428-9392-cbb269882575 · inbound
Progent: Securing AI Agents with Privilege Control Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1cccfa9b-1361-422b-be2c-5edeaf6d96aa · inbound
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7270254-4aab-4948-9855-19b980d6df2c · inbound
Advancing LLM Safe Alignment with Safety Representation Ranking Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d6f016-d02b-4072-953a-d445c93b68cd · inbound
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f47dd509-f826-4bfd-91bd-15cb70df1707 · inbound
LLM-Powered AI Agent Systems and Their Applications in Industry Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3d13f261-fc53-4db5-82a8-4d8512cc1b00 · inbound
Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68efde7-02c4-47b8-9d6d-7f4ab01dc51b · inbound
EnSToM: Enhancing Dialogue Systems with Entropy-Scaled Steering Vectors for Topic Maintenance Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a28b61-7ba4-4a60-89e4-b959bba1cbc3 · inbound
LLM Access Shield: Domain-Specific LLM Framework for Privacy Policy Compliance Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f894b69-d7be-4710-a0a9-f5fe74358a9c · inbound
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf682473-c924-4060-b14b-99535db4bba4 · inbound
Content Moderation in TV Search: Balancing Policy Compliance, Relevance, and User Experience Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1a5f60f-da92-4416-b5b0-7a6f08b3ef81 · inbound
Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs? Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 533de5ff-de66-4a20-bafc-ba2d1bba421f · inbound
Mitigating Deceptive Alignment via Self-Monitoring Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a94f58ae-8610-4a7f-b05b-0dd5f644d0a1 · inbound
Security Concerns for Large Language Models: A Survey Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd501efa-f61b-4fbe-bb0f-4f56f228df20 · inbound
VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acab8d73-6746-4415-b487-bc4857ac4fc4 · inbound
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3bf6bcac-375c-47b1-a898-025e275b64d8 · inbound
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a34b9a9-a42a-4873-81c9-41aee11e8d93 · inbound
SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9949848e-882c-4c87-ae21-d591081d1305 · inbound
Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312bed32-3c6c-4f5a-a6d9-bba5dc19d8e0 · inbound
AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88b28cec-7931-4bd0-ab13-fd3bf806ad74 · inbound
Should LLM Safety Be More Than Refusing Harmful Instructions? Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d8b3d7-9188-4a13-9f55-20c4d875c63c · inbound
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ea8fead-5ff6-4fb7-9de4-da9df154c32d · inbound
A Red Teaming Roadmap Towards System-Level Safety Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea34ee52-42ec-4c1e-ae5d-96a5e1f37298 · inbound
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff8c9001-6c80-4574-b97e-7e4483d26373 · inbound
Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bb5367b-8781-4928-bfc9-7a99b2465ab7 · inbound
JavelinGuard: Low-Cost Transformer Architectures for LLM Security Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 424edd04-e787-49a4-8572-4b0d520bdb7d · inbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45db551d-7a63-4765-a360-9c5e1771aa46 · inbound
SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 427cf786-071f-45a2-b516-8bde402013dc · inbound
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 855553a9-f7f4-4bf5-a382-1af3ac415157 · inbound
"I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1683a507-e3d5-4e7f-b7a0-19a007254198 · inbound
PL-Guard: Benchmarking Language Model Safety for Polish Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9e31c90-214d-4919-8250-b2362123671b · inbound
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8bb9cdef-5d9f-4a89-bffc-0f18e730a2b8 · inbound
Optimising Language Models for Downstream Tasks: A Post-Training Perspective Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47456cfa-c3cc-4b83-a3f7-3921fdbc22c4 · inbound
VERA: Variational Inference Framework for Jailbreaking Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 934cfb72-d162-4124-9093-7f5d8f904889 · inbound
On the Surprising Efficacy of LLMs for Penetration-Testing Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52dcb26a-0105-4bc9-ab25-c05e1c2d69d7 · inbound
AI-guided digital intervention with physiological monitoring reduces intrusive memories after experimental trauma Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abbb5f1d-fe91-4e0b-b262-650f9a19e1a3 · inbound
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7353ff7f-513f-4b87-80c8-4a7f4116fc1d · inbound
GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 224d1a9d-36e6-4236-9ac3-18498944b9ae · inbound
ModelCitizens: Representing Community Voices in Online Safety Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d70c4c11-6f58-4a53-84d6-f1905a37804f · inbound
Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deafb81e-b3c7-46e4-95ae-f81a2e7c1048 · inbound
SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7362b693-3a9c-47e2-96a8-c13be50fd027 · inbound
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f1391c9-8c4e-42ed-b5c0-bf4ce8d633ac · inbound
Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92ffb7e-80dc-400f-bbf3-ac1b6d775d14 · inbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9101c5-a663-40e6-86d9-309a01082db3 · inbound
LLMs Encode Harmfulness and Refusal Separately Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 210b3c37-d214-4c21-9fc6-bc16dd1be8f5 · inbound
Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00e24ce-97d0-4933-b7d5-cd408a718413 · inbound
PrefPalette: Personalized Preference Modeling with Latent Attributes Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 2003
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f7ec07-3d44-4cf9-95d7-a88ae65cc2ed · inbound
WebGuard: Building a Generalizable Guardrail for Web Agents Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2de32363-6a72-45f3-a94f-7fa96500ef52 · inbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8012d4b5-ccb6-40cd-a89d-16d44398b3f4 · inbound
ADEPTS: A Capability Framework for Human-Centered Agent Design Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1902be7-ee35-4736-88f9-ce89c36aae13 · inbound
TorchAO: PyTorch-Native Training-to-Serving Model Optimization Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6528721-9175-4293-aa54-96cae918c15e · inbound
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3197835-56a4-4ad2-9624-0d1e5b01b33f · inbound
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111788cd-b8c9-48a2-b775-0811b8f2e177 · inbound
OneShield -- the Next Generation of LLM Guardrails Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03bd2d9c-1211-4ebb-80a4-726bfce0f132 · inbound
Agentic Web: Weaving the Next Web with AI Agents Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 848fe526-4132-4bf4-9e22-842a315f02bf · inbound
Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8aa02b0-d615-4334-b707-09bfeeaf163b · inbound
The Problem with Safety Classification is not just the Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9299ee83-44bc-4628-8e41-0eac91cd7080 · inbound
Libra: Large Chinese-based Safeguard for AI Content Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b014c58-1e32-4a5c-a8a7-599a49863fbc · inbound
RedCoder: Automated Multi-Turn Red Teaming for Code LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f1e94ae-9858-419d-8101-cc7efc2b17cb · inbound
Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d7dbf2d-5024-4188-9bf2-190d7779a207 · inbound
Provably Secure Retrieval-Augmented Generation Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ded2693-e7ac-44c7-9a55-76436bba48b2 · inbound
Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 124
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5365e4f-0f2d-47a2-8d76-5b266eafd26e · inbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5881f87-bd05-4c12-9048-d554c8f0467d · inbound
A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d71ee36e-7699-4e13-b5b2-fbbdf54fb341 · inbound
Blockchain Network Analysis using Quantum Inspired Graph Neural Networks & Ensemble Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c156e9-ba9d-462d-aeb5-7cb3cb65fc2e · inbound
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b002207-66d9-43ba-9015-89a5141a2db7 · inbound
On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70fe987c-33ce-414e-964a-369b40f947d5 · inbound
Reliable Weak-to-Strong Monitoring of LLM Agents Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2bb5a08-5439-4e4e-9501-cb9dbe177c3a · inbound
IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ade512d-1ab5-4802-8155-1fa8a76cc5a1 · inbound
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1e67f2ad-fd86-4f28-b849-5f4af713079d · inbound
Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee73ba9-20b5-448a-8f5d-59e87ca16c45 · inbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 270
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94554423-d12c-46ec-8c8f-7be52ffd7e57 · inbound
Scaling behavior of large language models in emotional safety classification across sizes and tasks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54805df2-0907-45ad-a454-492f38e700bb · inbound
Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a43c34b6-1182-4430-99b8-28e4822e7418 · inbound
Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.