Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:38:37.185167Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 8 inbound Pith citation observations for arXiv:2412.18693.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:38:37.185167Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:23:48.057420Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T20:54:21.584820Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3a455212-b4e4-4735-a571-38a873193b0f · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Evaluating Large Language Models Trained on Code
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7607460b-5317-4a25-9a2f-eb068e8ca21c · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning persuade the user to incorporate daily exercise for health benefits
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 58e53305-9bb9-4e0e-b728-82380d3faa7a · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Scalable Extraction of Training Data from (Production) Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9c53fb-14f8-4ce6-b98f-412e6b863e3f · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e05c1f27-4a4c-4ccf-ac5c-bc5d8f27c3fd · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 982ef8cc-470e-4f0a-bd50-3f8e1e46ee80 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning FLIRT: Feedback Loop In-context Red Teaming
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 939c8fbd-d949-46ea-936d-caa97c95d135 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Gradient-Based Language Model Red Teaming
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a297fb3-6279-4d64-883d-86c94f074646 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning PAL: Proxy-Guided Black-Box Attack on Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05874896-e918-4dcf-8269-3d8d3d9a551e · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73189aad-9820-43d7-8eda-0d60eb61fa66 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Improving alignment of dialogue agents via targeted human judgements
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f232a7-fddc-4604-947a-518bfb78ccda · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc297b22-cd09-48cc-af46-d3c0fd9b0cb8 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Towards evaluating the robustness of neural networks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b5221044-d131-4726-b209-7bc3e89ce20e · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Gradient-based Adversarial Attacks against Text Transformers
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d15b9a1c-e6d5-4571-a6f5-9c82a6980dcf · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Categorical Reparameterization with Gumbel-Softmax
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5185ca4f-084b-450b-8f0b-90d1ccc70238 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Universal Adversarial Triggers for Attacking and Analyzing NLP
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42180de9-4d5b-4b7d-bed6-6c96075f55ef · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Robust Conversational Agents against Imperceptible Toxicity Triggers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eac0baf8-8037-4944-9c6b-671b07c2279e · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Automatically Auditing Large Language Models via Discrete Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e0147cc-0a79-4dc0-953a-3688a585d55d · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08739c8f-2184-4f18-bb4b-202f215b8875 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233c255d-e8ac-46c5-b10a-390fc426c1ca · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Adversarial Training for High-Stakes Reliability
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7042796-280e-4881-bb20-d605d86b1d02 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1166a25e-124b-40ba-97e6-955b28124f1f · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Dynabench: Rethinking benchmarking in NLP
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2cb2de59-09bb-4cb6-a524-6b0eadcbe133 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning doi: 10.18653/v1/2021.naacl-main.324
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5b8c7e7-81f2-4b82-aa35-284e15c98033 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Explore, Establish, Exploit: Red Teaming Language Models from Scratch
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75bd1147-8474-4be8-a15a-a528a8d57538 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca162ccd-40f3-4202-81b4-f32011e6a94f · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Curiosity-driven Red-teaming for Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c403507-d700-42df-8865-88e97d87de29 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning MART: Improving LLM Safety with Multi-round Automatic Red-Teaming
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c4e0ae-1621-409c-8faf-9b0ea080fe6b · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85991f10-4d48-445b-bea2-3e532e1bb01c · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Measuring and mitigating unintended bias in text classification
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9d3caa98-bb36-45a4-8ed2-e9c3daaa4868 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aafab37d-06c4-4361-a6d1-f223b39fd5ad · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Text and Code Embeddings by Contrastive Pre-Training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8284c02-f3aa-4b32-908e-d09f7d04fed6 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning WebGPT: Browser-assisted question-answering with human feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 814faea4-dd24-4b84-a8da-9fe04d1328ae · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning ReAct: Synergizing Reasoning and Acting in Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4320da0f-87a2-46ce-8a50-41aa3f79efdd · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning safety jailbreak
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation efa3be7c-51dd-4c59-99fb-dc2ea3882446 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning HotFlip: White-Box Adversarial Examples for Text Classification
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93787350-0028-449a-9063-78d9f4c6f6c6 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Towards Deep Learning Models Resistant to Adversarial Attacks
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 807791ba-3657-485e-94cc-043103b873cd · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning The Woman Worked as a Babysitter: On Biases in Language Generation
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08db9907-3aa0-47c1-99be-94dc0272f2a4 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ee5f4db-4ee8-44d9-943b-c338d9c299e7 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Red Teaming Language Models with Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb77f118-3d69-4c1b-b2b7-67f426225bb9 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning GPT-4 Technical Report
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97724b42-e4b7-4616-a72b-0a0ba2a48971 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-Fran¸ cois Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, and Jordan Boyd-Graber
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e54732c8-b4b8-4a18-8a83-1a221d2700c8 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfd78291-d084-4771-97dd-d87b2e086922 · outbound
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1b7a8a-c76d-4627-961e-6017443a229b · inbound
Jailbreaking to Jailbreak Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06ea541-bfea-4834-9e5e-d7489f9e8c11 · inbound
When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a57cad7f-0a66-442d-ba67-07d698debdec · inbound
Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac7ef433-3e74-412b-8fe7-9589df932fb0 · inbound
Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536888bb-9079-4d5f-ab57-719aa7f2a0f4 · inbound
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f65c26c6-c0b8-4993-b6c5-7d009c0c4e55 · inbound
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4861d50b-d0bd-417c-95cb-0513527f72a6 · inbound
Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e8a800fe-adaf-4bc1-b6ff-5bf651c93279 · inbound
GPT-Red: Automated Red Teaming via Self-Play at Scale Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.