Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:37:09.539148Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 26 inbound Pith citation observations for arXiv:2412.07724.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:37:09.539148Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:25.702916Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:38:54.910365Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 49f00bda-38d5-4fb2-852f-87a54dd2da73 · outbound
Granite Guardian Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5284848-538a-4e36-81a3-f71074ec8bd1 · outbound
Granite Guardian Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bf5f906-5679-4e01-a3fa-bcb81c063930 · outbound
Granite Guardian AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efbdd58b-621b-45e2-a1f2-cd38ab627bb0 · outbound
Granite Guardian DialFact: A Benchmark for Fact-Checking in Dialogue
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 07af79c9-0ae4-4823-82cc-961e69cda270 · outbound
Granite Guardian WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f658fc9c-8b94-456b-bb53-7b00a8893398 · outbound
Granite Guardian URL https:// aclanthology.org/2021.emnlp-main.619
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a3554e3d-2752-4850-9d35-7b199b36c4ce · outbound
Granite Guardian Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f195c0cd-ef30-4970-a97f-f2953636dfc7 · outbound
Granite Guardian WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33352a5b-9a2c-4faa-9b01-1f4029128e05 · outbound
Granite Guardian Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bcf2ebcf-ae9a-4093-9b9a-ece9f884b2ea · outbound
Granite Guardian On faithfulness and factuality in abstractive summarization
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation edd39436-c194-4c3c-b121-14aa48eee8e8 · outbound
Granite Guardian AI safety v0.5 proof of concept
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 037aa098-ae99-4d5d-b3c9-a76bcda0f0ca · outbound
Granite Guardian Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4ca22ffd-9db1-4b22-bb01-6120be77fc58 · outbound
Granite Guardian OWASP Top 10 for Large Language Model Applications
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 54bb927b-d886-4088-a5ab-7b017bc0189f · outbound
Granite Guardian URL https://doi.org/10.1177/0146167217741313
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e4b4dca-855f-4d30-93cc-20025340c139 · outbound
Granite Guardian doi: 10.18653/v1/2021.naacl-main.383
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11067dae-6004-4bb7-9067-71c638c9c933 · outbound
Granite Guardian doi: 10.18653/v1/P18-2124
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e53b05-2dec-4d57-98cb-8913088a9da9 · outbound
Granite Guardian Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd032d25-2b61-4c97-be6b-39d93cae2718 · outbound
Granite Guardian Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d93d9de-315b-48b3-8b39-a0a0c3dd3959 · outbound
Granite Guardian Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7f2045f-f494-443d-9cb9-762941d3346d · outbound
Granite Guardian The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 276d121f-1343-4c87-9898-763eb331682c · outbound
Granite Guardian MiniCheck: Efficient fact-checking of LLMs on grounding documents
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 15b762d3-4917-49e8-84d7-d0d536cc8010 · outbound
Granite Guardian URL https://aclanthology.org/2024.emnlp-main.499
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cae8885e-c4ac-4346-8433-95b7ff417fed · outbound
Granite Guardian Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26fd4747-a2ad-4cfa-b1e9-139de5cd2713 · outbound
Granite Guardian SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44c2ce5b-1fb6-4ade-9c72-b71c47d59f32 · outbound
Granite Guardian doi: 10.18653/v1/2020.acl-main.450
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e06c2fb5-6eaf-41db-aba3-fab58ca0fd73 · outbound
Granite Guardian A broad-coverage challenge corpus for sentence understanding through inference
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3c6588ef-f5c3-40df-a61c-20a4a51679a6 · outbound
Granite Guardian Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation accd773b-a004-44c0-8907-ca7d81dfe7b0 · outbound
Granite Guardian ShieldGemma: Generative AI Content Moderation Based on Gemma
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62309228-9ea6-43d7-bd1f-8aa2d0dcd2eb · outbound
Granite Guardian PAWS: Paraphrase adversaries from word scrambling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5444ce29-8025-4b98-94da-57558ac03a26 · outbound
Granite Guardian q2: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a0b9f7c6-bdb4-4e89-9339-3819379e908c · outbound
Granite Guardian Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 847d1bb5-808b-45c5-a09b-3f7cb2554165 · outbound
Granite Guardian Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44584a0d-9b93-4765-b0b3-7aa38f54df1c · outbound
Granite Guardian doi: 10.18653/v1/2020.acl-main.173
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 013d6b17-61ed-4132-a7a7-8be16712e111 · outbound
Granite Guardian Latent Hatred: A Benchmark for Understanding Implicit Hate Speech
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04dd7bd4-0fb0-4d2d-9821-c55b6ac75cd1 · outbound
Granite Guardian A large annotated corpus for learning natural language inference
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 66796fb9-0a43-4da8-8b6a-e717994bb43f · outbound
Granite Guardian Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8f23ad9-21a4-4a6f-a0ac-661aecc2c5dd · outbound
Granite Guardian Evaluations of Machine Learning Privacy Defenses are Misleading
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add17f0c-e5c3-43a4-a8b7-5916a2631bf3 · inbound
An Annotated Reading of 'The Singer of Tales' in the LLM Era Granite Guardian
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db93a84-fa5b-4383-ab6d-be8d78f5a327 · inbound
Dark LLMs: The Growing Threat of Unaligned AI Models Granite Guardian
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd1aae8-08a5-4ac8-9781-3ee1135bf707 · inbound
Concealment of Intent: A Game-Theoretic Analysis Granite Guardian
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19441e0-177a-44fe-95a2-815226a4e0b6 · inbound
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment Granite Guardian
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 99b35f78-3fdc-4095-9493-5e5f5da29c49 · inbound
SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues Granite Guardian
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e45241f8-8ebd-45f5-b39e-5051199e0250 · inbound
JavelinGuard: Low-Cost Transformer Architectures for LLM Security Granite Guardian
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3f62cc-7afd-40bc-917f-eeed3bfd9910 · inbound
Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair Granite Guardian
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4710658-57a5-43b5-890e-a48c0ce0be53 · inbound
Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation Granite Guardian
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c160fdaf-a08b-48ce-bcc9-d4c17e103ddd · inbound
Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs Granite Guardian
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ba3cdc33-703f-49e6-854e-6ebd0fc23edf · inbound
RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration Granite Guardian
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3a927148-053b-4437-b096-232ac64b7a31 · inbound
TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts Granite Guardian
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 85cccfbc-7f5e-424a-9a79-d62836e10b1a · inbound
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills Granite Guardian
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bd324d92-bf14-441b-af29-4be24da6c197 · inbound
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models Granite Guardian
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d133d424-1d0a-4521-afc0-26dd46581487 · inbound
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation Granite Guardian
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf45abf5-dbaf-4118-83cb-fa0eccbd1f4d · inbound
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content Granite Guardian
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2790371d-1cb7-40de-a75b-20befd7f763a · inbound
Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability Granite Guardian
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a7d4e77d-b7cd-467d-b680-a602e2547c0f · inbound
Distilling Safe LLM Systems via Soft Prompts for On Device Settings Granite Guardian
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 847622ec-4900-4320-a49c-379d6d20c338 · inbound
BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems Granite Guardian
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 98c98942-e322-4f0e-91f5-dcf8072acd77 · inbound
Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety Granite Guardian
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3b783caa-c649-4cf4-8e97-3c668c3d182a · inbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Granite Guardian
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c039225c-efed-4ebf-addc-d5b8a6f697d3 · inbound
Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers Granite Guardian
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fcbaf36-3046-4b4e-839f-c65659a7c6fb · inbound
CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization Granite Guardian
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abc6bdd9-fa6d-4439-9573-ffd07e8ee868 · inbound
CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization Granite Guardian
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 854e725e-9c82-4c49-80bd-72bfce434bb3 · inbound
DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Granite Guardian
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98e0530f-2a6f-431a-ae7c-c0252cc6c0c2 · inbound
JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety Granite Guardian
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b28a13af-8bc0-4450-9608-587542258f0c · inbound
Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B Granite Guardian
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.