Pith. sign in

Paper Citation Record · LEDGER

AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 53 inbound Pith citation observations for arXiv:2404.05993.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.05993 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 53 of 53 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:11.637137Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a6f947b2-4071-4765-94e4-06c4f286006f · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:25:14.838539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:f3bcfe6fbff32aef22b149884c4309baea340286bf6306e47909165d04bb044b

Observation a4830859-e41c-4dc6-bde5-aa83c5dc7c11 · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.492624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:45167dd9a3f7b97b620855797943ca234c23d971316b3784ca8062cd0a2a4f0c

Observation f7e03006-52c6-478e-905c-a08392c8e9a7 · inbound

Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations cites this paper.

Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T19:41:19.647374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:41:19.647374Z digest=sha256:c85d78d57cd4e0590cd0003dca9a7eb55ac1e95295fbdc48dec24fb3d3a25a76

Observation 81b63bbe-3def-4bcb-8c5d-e368e5ec89cf · inbound

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models cites this paper.

The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:43.310566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:43.310566Z digest=sha256:22005e0d100bcc66b9c0cf12fce336b94d18be36088845967bb667a7288e4cbd

Observation 1bf5f906-5679-4e01-a3fa-bcb81c063930 · inbound

Granite Guardian cites this paper.

Granite Guardian AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:37:09.396207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:37:09.396207Z digest=sha256:e05abdff0ee0c9a921fafd44a68964d5729a5fcc80168fa6b9bf8b7ef418a214

Observation ff891a96-45b4-400e-b5aa-786d1205030d · inbound

Lightweight Safety Classification Using Pruned Language Models cites this paper.

Lightweight Safety Classification Using Pruned Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:11:27.241267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:11:27.241267Z digest=sha256:788b5eb8f4b477c6953f71eb6a388d521539fa5a9c91149626fbec95e1624ae9

Observation c38e0513-05f2-4a38-8b44-ff848401c888 · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:38:45.444469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:281934d7962fd55fbdc596aa9ca6ed813c2a4488d4dbf9993634f1f9a7b0a96b

Observation 9640defd-d58e-48b3-a75b-e21cd1545ae2 · inbound

Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails cites this paper.

Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:54.707197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:15:54.707197Z digest=sha256:24e8fd0249318405f7e3ac35a9b79a05f1a183d86822a6ad856fe0ac19c60a27

Observation e7d7733f-33bc-43aa-8b2e-322d6ba92696 · inbound

Peering Behind the Shield: Guardrail Identification in Large Language Models cites this paper.

Peering Behind the Shield: Guardrail Identification in Large Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:45:21.395850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T03:45:14.234545Z digest=sha256:09ae9b7a1019af7bc19bbb8b63ac57b6a574dffda89c9b24caffda3d973bf324

Observation d434f31f-5769-4c1f-9726-52e078753405 · inbound

Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation cites this paper.

Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:11.637137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:11.637137Z digest=sha256:9f82b2188a58ec87f568386dcab93a0d7b83186b05c291ceb2212f584134e6a5

Observation cb0c9166-3cda-4a8f-a17b-91d9470880b7 · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 176

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:00.883032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:00.883032Z digest=sha256:364a4db4a1e28031d2d37ba4bee3f43cf5cfd2a37c5ca0d5fc3a2180d3baed77

Observation 76a37a7c-027f-4957-9b31-ce521c91672a · inbound

Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing cites this paper.

Unified Multi-Task Learning & Model Fusion for Efficient Language Model Guardrailing AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T06:00:23.878320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:00:23.878320Z digest=sha256:7b6d7f38c9b817c2e354153109362cc224ff781c4deb909d7e4e5b932f7ac07c

Observation 700d9382-4c44-45a7-b731-f2836b8734c7 · inbound

Understanding and Mitigating Risks of Generative AI in Financial Services cites this paper.

Understanding and Mitigating Risks of Generative AI in Financial Services AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:19:10.418119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:19:10.418119Z digest=sha256:704dfa920e92f4d2fb1cf892d957669686a59bb1b6f83481dfe1e1f249fabce4

Observation c0936b09-a7b3-43ff-8f68-e03761ecaa46 · inbound

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning cites this paper.

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:17.363940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:05:17.363940Z digest=sha256:592b66c45bc441d6f07811aaab725e9599038924a41c75dfea1a525b01c92343

Observation 553144f0-55aa-4cc2-913a-7c5a9c7d1345 · inbound

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs cites this paper.

ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:22.394650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:22.394650Z digest=sha256:c262e17d78c37647231e41f502d73bc37f979d393130eb44bf5a67835252d6c9

Observation 2ee2c05a-9166-4ce2-9426-738d099af967 · inbound

Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon cites this paper.

Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:15.911283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:15.911283Z digest=sha256:0c3da4f78f0652a2a71323107fa60f6a711b94fc5aedc9f008c203e4545f683a

Observation 21dc1195-0146-4bcf-875d-07cbab1af98f · inbound

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment cites this paper.

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:53:03.538854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T11:52:36.688263Z digest=sha256:f61a6d05b715a6c279a3b3a29af80b714c7fe2792575fdb73dc41fa5e6d967dc

Observation 95622f99-7e3d-4c7f-80f9-95034ad77d4b · inbound

A Red Teaming Roadmap Towards System-Level Safety cites this paper.

A Red Teaming Roadmap Towards System-Level Safety AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:19.166188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:19.166188Z digest=sha256:da203a5ffc1afc991ee6c0ab103029da9f6f07879c705c2b916ab7e6d60103fb

Observation 9a02c637-236b-48bc-9429-84998efa3c15 · inbound

JavelinGuard: Low-Cost Transformer Architectures for LLM Security cites this paper.

JavelinGuard: Low-Cost Transformer Architectures for LLM Security AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:25.380511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:25.380511Z digest=sha256:94b40799c8a35ec46a58868c4af002efb0e2d0d138705a1f94e845bc98150067

Observation ce2e4144-2beb-4ad1-a941-531422c683b6 · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.807876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.807876Z digest=sha256:02ce8201b7484a38a122d1b6ccbd64d1b2b2b73064b167fbacad5975dc22cada

Observation d563ce13-f171-458d-8fca-4797387eb0e7 · inbound

PL-Guard: Benchmarking Language Model Safety for Polish cites this paper.

PL-Guard: Benchmarking Language Model Safety for Polish AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:52.956697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:52.956697Z digest=sha256:16907d88a34ec5a1f60de7e121f70833d92e2cd97dcd5ed38cbc1d1c2222215c

Observation 86920b5e-5c8a-4613-b743-109ccee0e90b · inbound

GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models cites this paper.

GAF-Guard: An Agentic Framework for Risk Management and Governance in Large Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:26.615591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:26.615591Z digest=sha256:5040fc5f3c6b6926e9661a624cc5e62f901f06b8dd99f8d689b929b3b87c9e24

Observation 81ad3d4c-d0db-443f-94ca-c1a3350de80a · inbound

Agentic Web: Weaving the Next Web with AI Agents cites this paper.

Agentic Web: Weaving the Next Web with AI Agents AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T13:05:34.201626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:05:34.201626Z digest=sha256:8b42a6191ab36ee9afdf7d6f1f4a8c6cffa11c03e3d53e94a1d26d9660a92416

Observation b54d9ee7-1675-406f-9fd5-ca2eec46de03 · inbound

Libra: Large Chinese-based Safeguard for AI Content cites this paper.

Libra: Large Chinese-based Safeguard for AI Content AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:17:38.660420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:17:38.660420Z digest=sha256:27338940b4d0aaec1b04f71d3b58e15bc823056d3c10c66ecd0f59da3f13ec89

Observation 85c84855-85be-479e-8369-f623731c2b37 · inbound

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models cites this paper.

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T19:55:56.480656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:55:56.480656Z digest=sha256:c258d519ba28e1ed5070f46934e4e6bf71192fe96f0af47dd29fa07744f715ce

Observation c2d13794-9a2b-4399-9469-ee2f79f5efd0 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:40.793268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:40.793268Z digest=sha256:439b4fbca6496863014ae658e23760d49f31a7805b9ad50e7130a0b8545ffe97

Observation 002e1da3-e29d-4e38-b67c-429527881c69 · inbound

Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming cites this paper.

Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T11:43:31.089245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T11:43:31.089245Z digest=sha256:c296a0c3f24a9b6d77295c3b83f0600170b77aa53a3f58f56ac85f86985aaed7

Observation 5819d180-8c54-4462-b410-c1dc3ff3402e · inbound

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs cites this paper.

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:46:39.716726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:27:13.339411Z digest=sha256:0543da8d411b8b0b1d725db4c955b9218ceadcda22841c788cde7622e0cab501

Observation a88c7e97-5595-458c-9bd9-36baa95ad573 · inbound

Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines cites this paper.

Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:20:58.822366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:40:23.128451Z digest=sha256:2404e083fd600a7e6d165fff7beabc32f799fbfcef53bd332ea16568a6cd805a

Observation 5d61a91f-2d50-47b3-a9f9-22c53dd92841 · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.127272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:a1cad1fd96115737251a1eb84d3a05d899f72a46e4190340459c4dd73c45c427

Observation 3fca799b-5bf6-498a-af1d-062c65e400b3 · inbound

Cross-Lingual Jailbreak Detection via Semantic Codebooks cites this paper.

Cross-Lingual Jailbreak Detection via Semantic Codebooks AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:20.039365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T16:22:33.567085Z digest=sha256:745f7528f770b145afc1443dc416f57aae4bf9f370e769f7315b067cab8b2d5d

Observation 0da1d02d-669c-4eaf-b5fb-cd6613cb3ea7 · inbound

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning cites this paper.

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:45:44.124627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T18:08:11.122577Z digest=sha256:ec1778948f0c3dfb8965d36048fb6f0b8e3b0c1ca8df075b28e184e30edd3f4b

Observation fa870234-32f9-4df3-a03e-649212b29865 · inbound

Compositional Jailbreaking: An Empirical Analysis of Mutator Chain Interactions in Aligned LLMs cites this paper.

Compositional Jailbreaking: An Empirical Analysis of Mutator Chain Interactions in Aligned LLMs AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:33:38.095886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T18:31:05.740757Z digest=sha256:1255d19ca3d038de70d3e523a86f82382d353bb255235badb247df80e2c81bed

Observation 8f713945-cf5e-4164-9073-3935e0655052 · inbound

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content cites this paper.

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:13:15.967694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T09:11:58.843585Z digest=sha256:ae2542910b3a714be3bbea1774279c67bbedada7677be154dfcb03ed227c7ace

Observation 208241d1-dd44-482d-82a3-02cdc8d0bedc · inbound

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails cites this paper.

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.502677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T23:03:01.151745Z digest=sha256:c03bb3e1309a33e1380ff475a88b4d0fd4aaa51edbff753909956a44807bf8dc

Observation c7ee5234-4cd9-43d3-90ec-657f22eb2075 · inbound

TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety cites this paper.

TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:42:35.873281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T18:55:30.836052Z digest=sha256:97b5d86346a4276e9d0c4e7d44a874fea17281e8c099f9009c6ca83f2283a36d

Observation 3f84cf0d-8e05-47bc-b092-04ecd60e5129 · inbound

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails cites this paper.

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:57.042339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T02:06:31.341519Z digest=sha256:18f6af66a040c8ab0e520c4fdcc3d5ff013c15d096107eb8a35495f45f7f4ac6

Observation 04344fb7-64f4-4e58-aee5-ded80df070a2 · inbound

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective cites this paper.

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.096909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T20:04:17.744876Z digest=sha256:3366b8f47d1b3a524efe53de233736aba3746222fa8f0aa1ccf0b982770f1b97

Observation 13e6fd3d-7253-4416-89b6-1fee6d31ee95 · inbound

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety cites this paper.

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:59:58.605211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T00:03:27.948225Z digest=sha256:e75a56f82eaead313939d519471e200d934f11608d13f0dd6ed790b456b0cbce

Observation c98e7043-5b28-4831-809f-a3f1306651e1 · inbound

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety cites this paper.

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:03:48.828247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T05:16:19.502837Z digest=sha256:85caf307e0724bc3951f2045adacdbe153740c73754745b474bbbb1b158f7ed2

Observation 6a33beb2-2e8f-466e-889a-a105d40632a1 · inbound

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety cites this paper.

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:52:55.775115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T00:46:03.210076Z digest=sha256:541a9cd8a1cef9e284417e517218d3b3fdf38961ee40b056551cab8d29352ac0

Observation 18c20757-5dd1-4e3f-8368-602d7b340dde · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:35:51.361175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-29T03:35:34.594617Z digest=sha256:e01539bc297ec80cf2ba2c77b27b684ab6a5009f3826c4f0133ae19efd75b4fa

Observation 636c13d0-a4cb-41e9-a88b-a04100941816 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:34:34.450551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-30T09:32:59.824110Z digest=sha256:51870f4bff424024dac011fe36a351d4c160b6d888c3696572440f364220303e

Observation 342689fb-6e32-49e4-9c78-1949b8e33380 · inbound

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing cites this paper.

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:34:18.840560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T06:31:55.631719Z digest=sha256:85010cd704072aff582ca7a2217736a777e9cdc3dccb2806de013fd259fe2787

Observation 5eb61755-d7b1-40cf-92ce-9670233dc6cb · inbound

DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail cites this paper.

DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-08T10:04:51.444702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T10:00:02.677030Z digest=sha256:bb6c572cc42016b91123ea30d9859a282fd3d08a9290c75e68b3874986ba920f

Observation 73d0fa2d-a920-49ff-a1d6-3683f8d3d691 · inbound

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models cites this paper.

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T05:21:30.132993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:21:30.132993Z digest=sha256:5605e265056960da26f88742826418f2c52ab2470ab8d596c7eb6512ff7c1e43

Observation 7171ac90-634b-41f6-ae0b-ab4cbebbfe30 · inbound

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows cites this paper.

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T07:06:18.415677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:06:18.415677Z digest=sha256:adda837cb922f650034a005f25c2a4cc4f80d87f006d7814827019f2b32f6e3e

Observation 0a908996-fc3e-49fb-9ad0-651b6a13bb62 · inbound

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space cites this paper.

When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T23:52:51.607980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:52:51.607980Z digest=sha256:c141b0b088fc2f01079d73a69add8739615e2caa1020e18f50981d80776a7c4f

Observation b4ff3818-7ec1-4014-badf-d67bcb19da28 · inbound

A Dual-Hypothesis Reasoning Framework for LLM Guardrails cites this paper.

A Dual-Hypothesis Reasoning Framework for LLM Guardrails AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T17:41:11.392807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:41:11.392807Z digest=sha256:5caa94d84b307910ee061d6bd219a5f393ae14e1bd77d818ef1bdb29b42b7c29

Observation eef59a7f-3937-4d71-a4f2-47736b47e2bf · inbound

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation cites this paper.

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:36:56.254253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:36:56.254253Z digest=sha256:8a98d6aeb7206148e46de7223f8e87697677319cd6798a54f6506bbc0845a549

Observation 785bc3e8-0cbc-4a7b-9825-9c96f929dccf · inbound

Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed cites this paper.

Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 143

Resolution
unresolved
no resolver link, observed 2026-07-31T09:24:20.717521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T09:24:20.717521Z digest=sha256:c39023fd7b1246c7fc74bdab6ef332f4c6eb9e34aec58a267174780eeb2c2fde

Observation 7990e189-4df8-472c-9ac3-9fdd6c542790 · inbound

Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production cites this paper.

Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:39:30.012200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:39:30.012200Z digest=sha256:4f33b82d7419f1f78750a00cf67db750113f5ea7e086998f46ca70e86283c641

Observation 7a6409d7-b35c-4e40-8c46-af18add6e315 · inbound

HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails cites this paper.

HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:40:07.488669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:40:07.488669Z digest=sha256:d0ded64c329929d792df8dd7801e75f52a227ab31445243c8db10fccec213d1b