Pith. sign in

Paper Citation Record · LEDGER

Granite Guardian

As of 18 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 26 inbound Pith citation observations for arXiv:2412.07724.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07724 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:37:09.539148Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:25.702916Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:54.910365Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49f00bda-38d5-4fb2-852f-87a54dd2da73 · outbound

This paper cites Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations.

Granite Guardian Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.363361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.363361Z digest=sha256:5e74811cb0007a409cf8a0f848fa4d3ff00058d144a2cd169c6d38cf694950c5

Observation f5284848-538a-4e36-81a3-f71074ec8bd1 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Granite Guardian Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.375357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.375357Z digest=sha256:b7d7eb884b438944ca1aade4df11371ac19d3afa1204e8191a4db1ddbeeafc75

Observation 1bf5f906-5679-4e01-a3fa-bcb81c063930 · outbound

This paper cites AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts.

Granite Guardian AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.396207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.396207Z digest=sha256:e05abdff0ee0c9a921fafd44a68964d5729a5fcc80168fa6b9bf8b7ef418a214

Observation efbdd58b-621b-45e2-a1f2-cd38ab627bb0 · outbound

This paper cites DialFact: A Benchmark for Fact-Checking in Dialogue.

Granite Guardian DialFact: A Benchmark for Fact-Checking in Dialogue

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:37:09.940600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.401605Z digest=sha256:cd05ae6a47524f44988d228a2554b1e7b7f3e2cae04fb9c85b78e2559b421d5f

Observation 07af79c9-0ae4-4823-82cc-961e69cda270 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

Granite Guardian WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.407592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.407592Z digest=sha256:c8e14b681511a8b59c0b0d181f26ef8db57ec60e984f27f619c58a0fc6524992

Observation f658fc9c-8b94-456b-bb53-7b00a8893398 · outbound

This paper cites URL https:// aclanthology.org/2021.emnlp-main.619.

Granite Guardian URL https:// aclanthology.org/2021.emnlp-main.619

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.193980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.417325Z digest=sha256:06e0a91b32a22574d593a7283c21e705effb728e7767817814f2468abd24b88d

Observation a3554e3d-2752-4850-9d35-7b199b36c4ce · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Granite Guardian Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.422521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.422521Z digest=sha256:ed417b69246df805e9ba3b1a97135560b60d493e54f883860f282d08ceed1858

Observation f195c0cd-ef30-4970-a97f-f2953636dfc7 · outbound

This paper cites WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models.

Granite Guardian WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.431837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.431837Z digest=sha256:f5b6113f226e8fb3857b58f15d0d1b9e2ebf54c42af452f3b1d741d755da340a

Observation 33352a5b-9a2c-4faa-9b01-1f4029128e05 · outbound

This paper cites Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation.

Granite Guardian Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.178899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.436854Z digest=sha256:f6e44717d5c7c4bab8c3adfe3aae832c0685113caa8741095f4174f67351593b

Observation bcf2ebcf-ae9a-4093-9b9a-ece9f884b2ea · outbound

This paper cites On faithfulness and factuality in abstractive summarization.

Granite Guardian On faithfulness and factuality in abstractive summarization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.163624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.441455Z digest=sha256:56dd8b570287c18db189e5fffa3fc0abb9d6d8cbaea612d8ec010c14dd22ddc0

Observation edd39436-c194-4c3c-b121-14aa48eee8e8 · outbound

This paper cites AI safety v0.5 proof of concept.

Granite Guardian AI safety v0.5 proof of concept

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.147864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.451799Z digest=sha256:1c9f3b6375fdf4b7754a69d10eb3a81f35495a930651207f4ee1b4cff6d11b3f

Observation 037aa098-ae99-4d5d-b3c9-a76bcda0f0ca · outbound

This paper cites Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization.

Granite Guardian Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.133218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.455833Z digest=sha256:16726c0c41b7e7d6b5f04a702879864da7f719edb20eb423a76089bed600064b

Observation 4ca22ffd-9db1-4b22-bb01-6120be77fc58 · outbound

This paper cites OWASP Top 10 for Large Language Model Applications.

Granite Guardian OWASP Top 10 for Large Language Model Applications

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.119077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.460271Z digest=sha256:bdb1c4f172b844585d24b5b6a8ab6e5745a089380015d6e384b99476c6c661be

Observation 54bb927b-d886-4088-a5ab-7b017bc0189f · outbound

This paper cites URL https://doi.org/10.1177/0146167217741313.

Granite Guardian URL https://doi.org/10.1177/0146167217741313

Reference 21

Resolution
verified exact
doi, observed 2026-08-11T18:37:09.623380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.464445Z digest=sha256:64916bc1b9d59905b5941d047e35388c9599131cfe141ea50fefa3e6d5828ff8

Observation 8e4b4dca-855f-4d30-93cc-20025340c139 · outbound

This paper cites doi: 10.18653/v1/2021.naacl-main.383.

Granite Guardian doi: 10.18653/v1/2021.naacl-main.383

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.468769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.468769Z digest=sha256:a4a8c2e4e92a923c3b3ed967289efaf04896da9a1b87e3fca720341db45b4bab

Observation 11067dae-6004-4bb7-9067-71c638c9c933 · outbound

This paper cites doi: 10.18653/v1/P18-2124.

Granite Guardian doi: 10.18653/v1/P18-2124

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.473630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.473630Z digest=sha256:7463d37380eb57f79e7572076debe1c8d14562b9a307c10714b739a487e878cb

Observation 62e53b05-2dec-4d57-98cb-8913088a9da9 · outbound

This paper cites Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI.

Granite Guardian Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.478432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.478432Z digest=sha256:125975ad7b6c41faecec39254a7b92113762e866159f0f24aff0ca34f973722a

Observation cd032d25-2b61-4c97-be6b-39d93cae2718 · outbound

This paper cites Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition.

Granite Guardian Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.483349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.483349Z digest=sha256:a79f262ba08eed7e19513c403943c4bf255dd4cfef3a4370a461b0a4a341d60c

Observation 3d93d9de-315b-48b3-8b39-a0a0c3dd3959 · outbound

This paper cites Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition.

Granite Guardian Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.488075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.488075Z digest=sha256:28159c7a081a4a77f7ba44e10783ec5c178f39b0ab32f25e64557b4cca865641

Observation f7f2045f-f494-443d-9cb9-762941d3346d · outbound

This paper cites The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence.

Granite Guardian The AI risk repository: A meta-review, database, and taxonomy of risks from artificial intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.492476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.492476Z digest=sha256:70c44b692812dcd4df4485a1c1e7a276a8e1bd4392e37120ae148ef3cfc3deb4

Observation 276d121f-1343-4c87-9898-763eb331682c · outbound

This paper cites MiniCheck: Efficient fact-checking of LLMs on grounding documents.

Granite Guardian MiniCheck: Efficient fact-checking of LLMs on grounding documents

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.103191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.497435Z digest=sha256:09363f16b08b914ca3befb81e1736ee8c7d554699d5363861b5955641e236f03

Observation 15b762d3-4917-49e8-84d7-d0d536cc8010 · outbound

This paper cites URL https://aclanthology.org/2024.emnlp-main.499.

Granite Guardian URL https://aclanthology.org/2024.emnlp-main.499

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.085497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.501895Z digest=sha256:fd52aa025b903d6cf1b1300c2e7271cfdde06487bebcf94612f3fa76ad958982

Observation cae8885e-c4ac-4346-8433-95b7ff417fed · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Granite Guardian Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.506514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.506514Z digest=sha256:f87b1544e6959b69755f8b772f8a64c7c527bb8f5e34ea0e2488201bab9fe501

Observation 26fd4747-a2ad-4cfa-b1e9-139de5cd2713 · outbound

This paper cites SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models.

Granite Guardian SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.511553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.511553Z digest=sha256:c6cb643b4813a0f1937f454cd0f24cea345d5936c91532534b11222cdb0ecb5a

Observation 44c2ce5b-1fb6-4ade-9c72-b71c47d59f32 · outbound

This paper cites doi: 10.18653/v1/2020.acl-main.450.

Granite Guardian doi: 10.18653/v1/2020.acl-main.450

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.516596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.516596Z digest=sha256:87f672218201a33c81e8e5a1c7b6e286b584dddc7b6264044da4b20b720c4252

Observation e06c2fb5-6eaf-41db-aba3-fab58ca0fd73 · outbound

This paper cites A broad-coverage challenge corpus for sentence understanding through inference.

Granite Guardian A broad-coverage challenge corpus for sentence understanding through inference

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.070194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.521102Z digest=sha256:a349c2f28699d2b85c947996b5d6b148db36f1b8ec5aad15e3da44a207da8e72

Observation 3c6588ef-f5c3-40df-a61c-20a4a51679a6 · outbound

This paper cites an unresolved cited work.

Granite Guardian Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:37:10.053749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.525689Z digest=sha256:57d260fa1b9318df90467bcf5ba51446548ffdd2e38ed4af459a960f8e69985a

Observation accd773b-a004-44c0-8907-ca7d81dfe7b0 · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

Granite Guardian ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.530020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.530020Z digest=sha256:72b93d2f261cbef2020490be078aaa04432c06ad64466aee66333df2d976cc2f

Observation 62309228-9ea6-43d7-bd1f-8aa2d0dcd2eb · outbound

This paper cites PAWS: Paraphrase adversaries from word scrambling.

Granite Guardian PAWS: Paraphrase adversaries from word scrambling

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.037724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.534812Z digest=sha256:8d57a3295755220a48231c0808d772a771181d718997a0d7a35a42a6ab397057

Observation 5444ce29-8025-4b98-94da-57558ac03a26 · outbound

This paper cites q2: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering.

Granite Guardian q2: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.209108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.412463Z digest=sha256:7859328a9a97b35f518219f62bf2d3ba15ce29b3370830264ccfe0bfb6549adc

Observation a0b9f7c6-bdb4-4e89-9339-3819379e908c · outbound

This paper cites Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang.

Granite Guardian Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.224013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.386089Z digest=sha256:b52640d11de9f0cb8398c46e1b75aa3f916ff65df37cb5545d276d6e957052c8

Observation 847d1bb5-808b-45c5-a09b-3f7cb2554165 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Granite Guardian Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.539148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.539148Z digest=sha256:a257ecf9ef22b48924eb1460a066c6707de0713e0d9aca395b69062c6fb100e1

Observation 44584a0d-9b93-4765-b0b3-7aa38f54df1c · outbound

This paper cites doi: 10.18653/v1/2020.acl-main.173.

Granite Guardian doi: 10.18653/v1/2020.acl-main.173

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.446715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.446715Z digest=sha256:c4ef5666650829e27a5c85df514d44ed4290fa23d40a608b7f76e93b49f59324

Observation 013d6b17-61ed-4132-a7a7-8be16712e111 · outbound

This paper cites Latent Hatred: A Benchmark for Understanding Implicit Hate Speech.

Granite Guardian Latent Hatred: A Benchmark for Understanding Implicit Hate Speech

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.391296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.391296Z digest=sha256:24b68b8ab45f25e2b37d43c42121f64753b81ee152e6b6d2e325c2a6035f634b

Observation 04dd7bd4-0fb0-4d2d-9821-c55b6ac75cd1 · outbound

This paper cites A large annotated corpus for learning natural language inference.

Granite Guardian A large annotated corpus for learning natural language inference

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:37:10.238649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:37:09.380927Z digest=sha256:7aba7ac631427c6b45b3eaf525477d0847a0a8c8c46a1323c7154add864d17aa

Observation 66796fb9-0a43-4da8-8b6a-e717994bb43f · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Granite Guardian Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.427318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.427318Z digest=sha256:46f7616509cdbda162f49e2f2064308bff6af9045299732503256e6aabbcd937

Observation f8f23ad9-21a4-4a6f-a0ac-661aecc2c5dd · outbound

This paper cites Evaluations of Machine Learning Privacy Defenses are Misleading.

Granite Guardian Evaluations of Machine Learning Privacy Defenses are Misleading

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.369590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.369590Z digest=sha256:d8c0ed123bb80b69347b4bf36bab33da84204364dbd783f9eda51e969d4a3aef

Pith citing papers

Observation add17f0c-e5c3-43a4-a8b7-5916a2631bf3 · inbound

An Annotated Reading of 'The Singer of Tales' in the LLM Era cites this paper.

An Annotated Reading of 'The Singer of Tales' in the LLM Era Granite Guardian

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T20:08:32.725166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:08:32.725166Z digest=sha256:344a1a82dad3bc334fc1be97159e8e281c3ef5e8e4e5c985dc79466cfbee49c6

Observation 9db93a84-fa5b-4383-ab6d-be8d78f5a327 · inbound

Dark LLMs: The Growing Threat of Unaligned AI Models cites this paper.

Dark LLMs: The Growing Threat of Unaligned AI Models Granite Guardian

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:25.702916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:25.702916Z digest=sha256:39a85f9510ecc48f47d999ab5c9a6a6b62d8ef9cd5b25d185f88e679a4724c9d

Observation ebd1aae8-08a5-4ac8-9781-3ee1135bf707 · inbound

Concealment of Intent: A Game-Theoretic Analysis cites this paper.

Concealment of Intent: A Game-Theoretic Analysis Granite Guardian

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:39.461607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:53:39.461607Z digest=sha256:40e084151c81cc89233a7d0cf1843f8419e04e646315dd0d6b8eaa43c7e39817

Observation b19441e0-177a-44fe-95a2-815226a4e0b6 · inbound

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment cites this paper.

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment Granite Guardian

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:53:03.602566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T11:52:36.688263Z digest=sha256:a60ae6b5ce85f3442d2f81205e76bde8b61de445c746ea9bb609d8dac6cfb608

Observation 99b35f78-3fdc-4095-9493-5e5f5da29c49 · inbound

SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues cites this paper.

SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues Granite Guardian

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:46.927403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:04:46.927403Z digest=sha256:3f8ce26f94196490583dc69ac3aa5b06900683a7e45f8419e3cc4624cd915c40

Observation e45241f8-8ebd-45f5-b39e-5051199e0250 · inbound

JavelinGuard: Low-Cost Transformer Architectures for LLM Security cites this paper.

JavelinGuard: Low-Cost Transformer Architectures for LLM Security Granite Guardian

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:25.455039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:25.455039Z digest=sha256:ac2d8b37dc5602535c07eb1c813b52f68e9d442affb42919bd16373d067ed57d

Observation 3e3f62cc-7afd-40bc-917f-eeed3bfd9910 · inbound

Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair cites this paper.

Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair Granite Guardian

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:29:20.657711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:29:20.657711Z digest=sha256:ac93065d788acca9aadabb6d5b5c3c151a189d1b2ff0e06fa2bac6047c58a963

Observation a4710658-57a5-43b5-890e-a48c0ce0be53 · inbound

Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation cites this paper.

Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation Granite Guardian

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:07:25.573450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T06:07:23.060966Z digest=sha256:f8a06551aa61b8fea74b84acb9c5fe66f713818fb1f55ae80224e5dddb9d2f2f

Observation c160fdaf-a08b-48ce-bcc9-d4c17e103ddd · inbound

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs cites this paper.

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs Granite Guardian

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:46:38.138360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:27:13.339411Z digest=sha256:c0292e1d54712e7079e573083ebcdbacbd09a95b9c0b240a4010e30073259f3e

Observation ba3cdc33-703f-49e6-854e-6ebd0fc23edf · inbound

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration cites this paper.

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration Granite Guardian

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:43:02.000605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T08:38:42.029762Z digest=sha256:e5fea2acf1a7a7134b44390b0489977115b71b90ca101e1b330102ad6488fd51

Observation 3a927148-053b-4437-b096-232ac64b7a31 · inbound

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts cites this paper.

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts Granite Guardian

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:08:25.568187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T09:07:57.713675Z digest=sha256:5c09ebd0c8d5c1467080a4e67f76d49b3a36f95c8b3c5f5340ee51721025e2f9

Observation 85cccfbc-7f5e-424a-9a79-d62836e10b1a · inbound

Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills cites this paper.

Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills Granite Guardian

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:16.339429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T16:22:49.737626Z digest=sha256:bbc6ff9f2adbfa59add5260dbb2c106a30c5863e3876c8a0c17e27a3f10e4517

Observation bd324d92-bf14-441b-af29-4be24da6c197 · inbound

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models cites this paper.

When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models Granite Guardian

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:05:49.322040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:43:12.298529Z digest=sha256:6da18e67c53c3e614df498cd41468dc9750339eee32fdb115768840bd447a873

Observation d133d424-1d0a-4521-afc0-26dd46581487 · inbound

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation cites this paper.

Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation Granite Guardian

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T23:33:25.778345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:33:25.778345Z digest=sha256:4401f5e455a19b9d0ddca08d7a3158dbdcbebc2ed3c3a65182abd9c1020a82f3

Observation bf45abf5-dbaf-4118-83cb-fa0eccbd1f4d · inbound

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content cites this paper.

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content Granite Guardian

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:13:15.985968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T09:11:58.843585Z digest=sha256:239db823e145ea0bc2024fb1b3da51b6d3787a6f20765556d58d9820a5a9dbc2

Observation 2790371d-1cb7-40de-a75b-20befd7f763a · inbound

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability cites this paper.

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability Granite Guardian

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:29.819529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T10:00:30.904247Z digest=sha256:9d59885428a2464b3118b40af6018eb160c0febff28a6c7181f9b0f40f2abafd

Observation a7d4e77d-b7cd-467d-b680-a602e2547c0f · inbound

Distilling Safe LLM Systems via Soft Prompts for On Device Settings cites this paper.

Distilling Safe LLM Systems via Soft Prompts for On Device Settings Granite Guardian

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.234744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T17:15:51.375580Z digest=sha256:c852e8fb4966f3125315c0d80fe05581688693e73430609db50c45b641f803b1

Observation 847622ec-4900-4320-a49c-379d6d20c338 · inbound

BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems cites this paper.

BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems Granite Guardian

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:40.766073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T05:05:08.336367Z digest=sha256:adb120d1c705ff9224cefcb2c1583f67632ba22e916f50cc0051368f15e86f68

Observation 98c98942-e322-4f0e-91f5-dcf8072acd77 · inbound

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety cites this paper.

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety Granite Guardian

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:54.912009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T20:38:16.610310Z digest=sha256:816874ffd825b1907ffc7a081a2c3e7b21c406f4934bcf258075a77b3d29bc70

Observation 3b783caa-c649-4cf4-8e97-3c668c3d182a · inbound

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models cites this paper.

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models Granite Guardian

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T05:21:30.132993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:21:30.132993Z digest=sha256:de6329123f6d0897e520e4eb031df0973d8271333d87b120d101cd6845185a4a

Observation c039225c-efed-4ebf-addc-d5b8a6f697d3 · inbound

Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers cites this paper.

Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers Granite Guardian

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T13:40:18.733148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:40:18.733148Z digest=sha256:f082bdd16e5d224d988e4f4cd0725d843f95478c58514c55d32293f445442fd5

Observation 4fcbaf36-3046-4b4e-839f-c65659a7c6fb · inbound

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization cites this paper.

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization Granite Guardian

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T14:56:08.423736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:56:08.423736Z digest=sha256:35b96bac737eadd25877f96f06597b0864df222ebebd42e6c1bdad82539ca6e6

Observation abc6bdd9-fa6d-4439-9573-ffd07e8ee868 · inbound

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization cites this paper.

CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization Granite Guardian

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:56:58.170447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:56:58.170447Z digest=sha256:f168c72033a871b952f90ffab7bb4913e0f9d288fab488f70fd1c07f8f517b1a

Observation 854e725e-9c82-4c49-80bd-72bfce434bb3 · inbound

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection cites this paper.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Granite Guardian

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.436778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.436778Z digest=sha256:e9fb33e710b490f93b99b4001f6261461597cd47b23ec242068287319325d7de

Observation 98e0530f-2a6f-431a-ae7c-c0252cc6c0c2 · inbound

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety cites this paper.

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety Granite Guardian

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:24:49.018907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:24:49.018907Z digest=sha256:6e2d1beef7a97ea4b9273e2ca39699af6690a8864a1363069db1d1c258f34c80

Observation b28a13af-8bc0-4450-9608-587542258f0c · inbound

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B cites this paper.

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B Granite Guardian

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T14:52:12.527253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:52:12.527253Z digest=sha256:31a0eeb596de269c465728aae1d56f7a2acc66e916410d43aa060d5cf0b0efaa