Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T16:25:14.744887Z
Paper Citation Record · LEDGER
As of 23 July 2026, this Paper Citation Record lists 63 of 63 outbound references and 45 inbound Pith citation observations for arXiv:2406.18495.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T16:25:14.744887Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T05:21:30.132993Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T06:15:00.866473Z
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a331e6aa-2851-4439-819b-9d90f00f4076 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 25708366-4fa9-43e8-96fe-01986aa66005 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Llama 3 model card
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d872d819-978c-4829-b6fe-cbf6ba537626 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs The claude 3 model family: Opus, sonnet, haiku
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 07a1da5b-b4f0-44e8-b987-1a99bf9035de · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f00169e0-0116-4e04-a219-d8721967e6ba · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Training a helpful and harmless assistant with reinforcement learning from human feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 167aa439-8561-44a3-b238-35fe66429d1d · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Longformer: The Long-Document Transformer
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6e6e61f8-16bf-43d6-bf13-78e229ebf1ac · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f95ec4dc-0b2c-44fc-ad56-92978e0a93ee · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e877ed7f-a5d9-4da8-aba0-e684b1d70f20 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 79fc2aab-b7be-4f85-bcf3-e3a4d415dcf5 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs The measurement of interrater agreement
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation dfaf4668-8bf9-4822-ba09-f582d5242f96 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation ff95bfa7-c8e3-4239-a9cf-30004be2e4c6 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Realtox- icityprompts: Evaluating neural toxic degeneration in language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a6f947b2-4071-4765-94e4-06c4f286006f · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 977d9e24-0485-4c73-821a-479f1350a22c · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Ruddit: Norms of offensiveness for english reddit comments
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 2ec2ec9d-f009-48aa-b07e-d632e06d2172 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs An Overview of Catastrophic AI Risks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 72db41cc-e8d5-43e0-9cc8-5b8bcb80f5e0 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 23daa911-e05d-4c4f-984d-71eb7289922e · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Smith, Iz Beltagy, and Hannaneh Hajishirzi
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9c4b6b5a-df39-439b-a808-9646388d6191 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a426fb12-99e2-4cb2-8c5d-2953105d1293 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3ab805f5-3864-4c30-9366-e43732a29c5a · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 2c9d52a0-e208-4458-bc21-f7cf34f28937 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs A new generation of perspective api: Efficient multilingual character-level trans- formers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 0f840343-e05c-4f98-b0ef-1094684dc243 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4ae480b6-9c61-470c-a90d-6399246e790d · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 0f3d717f-0b7a-4062-97c9-c082bb06011d · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs A holistic approach to undesired content detection in the real world
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation ec39ab30-010f-4947-a3c2-01a92ceac38e · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 2763d7aa-27f6-4a31-bef5-7ff2635145be · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Meta llama guard 2: Model cards and prompt formats
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 8e2fb651-e175-4bc0-aa69-5e71125b37b3 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Chenghaomou/text-dedup: Reference snapshot, September 2023
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation ffb72553-54ff-4228-bac8-3028f477a6b2 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Openai moderation api
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 90f0eb21-5407-40b9-92ec-2261ac57a641 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs SOLID: A Large-Scale Semi-Supervised Dataset for Offensive Language Identification
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 00e99c7d-61c7-475f-b471-5822ce9fa71f · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3ec50fd2-280b-4c61-8255-dc92afc44ef1 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Safety Assessment of Chinese Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 10652d37-b599-46f3-819b-5ee9cd9847f8 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7c4c8529-f4d9-4f04-92f0-30d1272e1bc7 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 6d4c4929-b179-4f63-8bb9-a2c6c108bf7e · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3090015e-366a-49b9-b1e8-71a6128993d6 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Llama 2: Open foundation and fine-tuned chat models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4f5f243a-d99e-431f-9ae2-b485bdeb8495 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c95dd37a-65d2-4051-8125-61a75faa3f98 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Introducing v0.5 of the AI Safety Benchmark from MLCommons
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 17ade557-0482-4cf2-ac20-62bb6333c0e0 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Do-not-answer: A dataset for evaluating safeguards in llms
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5a766751-fc22-4886-b1ca-c12b7d1b5dbe · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Ethical and social risks of harm from Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7df29d8a-16a1-4808-9cec-10210980b32d · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Semeval-2019 task 6: Identifying and categorizing offensive language in social media (offenseval)
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation ef6c1793-6a50-4ffe-9a18-c2cadea13613 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Wildchat: 1m chatgpt interaction logs in the wild
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d093f015-3bf6-4ab7-947f-abaab3128c82 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs P Xing, Joseph E
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 10e3fb7d-8348-42b4-aa94-8bc3a8fa6194 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 78e3adc0-7411-4d4b-96d4-4c86ceeeba18 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e2f01e02-8ced-4cb0-9e6a-2ec177ccb688 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Request
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 65cb4354-6cc8-4bd2-b9cc-70ebf1eda3db · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Assistant
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9aba8f34-f3f0-46ed-86b9-834bf9cfcb23 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Assistant
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b1904c56-0eef-4458-9a7c-6d062f2d789e · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Assistant
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 1e0f9a9a-ef06-4b6a-a57e-7843932117aa · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Assistant
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a9a5ad45-6336-48c8-bfb9-9ea6b2735fcf · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Assistant
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e00e7850-a015-49d3-8d96-71f8cf43af09 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Human" for the human response, and key
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 8d0d9d3d-b4bc-4772-82ee-cce4fc887e31 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Assistant
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9344b7f3-8f45-4490-b473-105aff195f39 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Assistant
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b7bc28d3-3f3b-4c1b-9af9-f8626a35b238 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Assistant
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation fed3b2a0-8f3e-4bf9-9f27-f5456269f7e0 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs examples
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 49056f99-63b9-403e-978c-cd1953575d99 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs examples
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 1b91cf77-0732-4c56-8097-cd778d797921 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Assistant
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation eeb36b6e-1a01-4a5f-9fc6-6d8a05725009 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4f045e81-56a9-4fd7-a01d-df436f867b8c · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 93ed501b-c30d-4845-b2f5-677fa58deee7 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Needs Caution
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 21377d56-d16f-4cca-94b5-e3a109ecf8ae · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 64a0daf6-40a9-49cc-994a-77e715d78efd · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 878f09c6-1f54-4edd-8bba-c0012929ce27 · outbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs As an AI language model, I cannot
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b9b2cbac-3fe4-417f-badf-33106ac3a2bd · inbound
ShieldGemma: Generative AI Content Moderation Based on Gemma WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5ac70946-520d-4327-99c8-769fdabb0ffb · inbound
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f7ff959f-544e-46e4-85dd-884ef1060c9b · inbound
Peering Behind the Shield: Guardrail Identification in Large Language Models WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c078129b-204c-43d6-86db-6731ea5b733f · inbound
Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 1e1842fe-dfa8-4f9e-8677-0b6e6b4112da · inbound
BarrierSteer: LLM Safety via Learning Barrier Steering WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation fef350b6-8a75-4a80-86c0-8e20a620f559 · inbound
Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5f64a7ef-0590-41fc-851d-ad40af7bc9dc · inbound
Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 90064aa9-e27c-4ae5-8893-16ab3d6d027c · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation d053ce7b-03aa-4f0a-9ce8-3fa3f2aa411e · inbound
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33dd3198-1d44-4e19-980c-8d1322a4e63a · inbound
Do LLMs Follow Their Own Rules? A Reflexive Audit of Self-Stated Safety Policies WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a7e76d4f-867b-4dff-8756-6912b13c9654 · inbound
ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation f7ca3975-ea2a-495b-ac2d-0bb702c3d5f1 · inbound
The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b5a7804c-8973-4454-8e28-d576a25b6a11 · inbound
Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation fae64d48-2412-475e-bd67-dae325f422c9 · inbound
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation af06ce84-aa00-4fa1-bdb0-d63225cd369a · inbound
STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation a4900594-7bda-4347-b54a-777e9ee387b2 · inbound
How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 838155ab-a03c-46e2-8236-3ff8eea07cda · inbound
Self-Mined Hardness for Safety Fine-Tuning WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3147ba89-70c5-42c2-ad12-3102ab92cd86 · inbound
Self-Mined Hardness for Safety Fine-Tuning WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 4fa2c9bf-fb4e-48f5-8141-764aaa31cca8 · inbound
GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 2b36bcf9-1cb8-4617-9531-ee4770bc2752 · inbound
GLiGuard: Schema-Conditioned Classification for LLM Safeguard WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 46b30a19-b46b-4e6a-8f68-58a9746808ce · inbound
Context-Aware Spear Phishing: Generative AI-Enabled Attacks Against Individuals via Public Social Media Data WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 342ebf4f-b167-4762-85ff-a5d1164f3824 · inbound
Bayesian Model Merging WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 363a5dc2-faac-4a06-95a5-7f16f798f2d0 · inbound
Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 70017aed-3fc4-4d7e-9574-902e1121a368 · inbound
LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation c7e1a42f-efed-4a59-95b5-94d2cb1247a0 · inbound
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 0145f2ca-6abe-4c53-8273-9be1ea18f396 · inbound
Refusal Evaluation in Coding LLMs and Code Agents: A Systematic Review of Thirteen Malicious-Code Prompt Corpora (2023-2025) WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9324d343-6067-4d1e-9dc2-9980417d560c · inbound
AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 8b3ac29c-b960-4245-90aa-9d04b50c75a5 · inbound
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 63c79d65-3c52-414a-a158-c32329a27542 · inbound
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2636ba3-ff4c-4d7d-92c3-6fd7a591f264 · inbound
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b5511fb2-f112-4eb1-b316-e69c7ecfcd4e · inbound
Gate AI: LLM Security Benchmark Evaluation Methodology and Results WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation b955e541-c739-4445-986c-5af19cf2ae5c · inbound
Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 389737a1-8a89-4d75-8470-71c2176a8d47 · inbound
Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5e086606-93da-4a17-88d3-f3aead05ecda · inbound
OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 9ed39ab4-125f-49fb-8297-fcd1b77fa33e · inbound
FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 2d3f59e9-38b7-4943-9475-8c00d7400c23 · inbound
What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations? WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 5edbb1a7-087c-4376-99c8-5b0a8d6c3b0a · inbound
Efficient Safety Benchmarking via Item Response Theory WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 2fb5785a-726f-4627-a9b7-b630cc353306 · inbound
Discriminatory Compliance: How LLMs Answer Queries from Protected Groups WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 7f489fbb-cfb8-4d29-bcb3-7954ac15773f · inbound
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 07667cd3-ac62-4caf-8395-b15867cd85d0 · inbound
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 48ec0eb3-4e0b-44a7-b56f-c20f65a95c89 · inbound
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 3218c63f-1684-4e82-bbd2-ab1414dd18fb · inbound
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation e1539bfc-f4fc-4e8a-bf63-7eb4381e87d0 · inbound
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 77399629-29a7-41ed-b080-477d92753fab · inbound
HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.
Observation 98e0a8fc-8db3-4ed7-b23d-4e9e6490c449 · inbound
HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.