Pith. sign in

Paper Citation Record · LEDGER

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

As of 23 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 8 inbound Pith citation observations for arXiv:2412.18693.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18693 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:38:37.185167Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:23:48.057420Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T20:54:21.584820Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3a455212-b4e4-4735-a571-38a873193b0f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.470070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.470070Z digest=sha256:e072f110743a322b44b9195046a873925f41c411c172dde245285318b64ee0b9

Observation 7607460b-5317-4a25-9a2f-eb068e8ca21c · outbound

This paper cites persuade the user to incorporate daily exercise for health benefits.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning persuade the user to incorporate daily exercise for health benefits

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:38:37.620528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T04:38:37.114886Z digest=sha256:3649aa7a05e7c993167367ebd285b7657ec1327b6552c3b5c64cb2897cb4f858

Observation 58e53305-9bb9-4e0e-b728-82380d3faa7a · outbound

This paper cites Scalable Extraction of Training Data from (Production) Language Models.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Scalable Extraction of Training Data from (Production) Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.504371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.504371Z digest=sha256:dcb1c582e7d5f81aadf77b6396ec593f5796e8f60aeaf69f1fde2fa798e1e8bd

Observation 2e9c53fb-14f8-4ce6-b98f-412e6b863e3f · outbound

This paper cites Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Gender Bias in Coreference Resolution: Evaluation and Debiasing Methods

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.510827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.510827Z digest=sha256:24903076b6cb02e8d45faa3b8ca0fb8f13c351a3b49cf3f168a73f4566037862

Observation e05c1f27-4a4c-4ccf-ac5c-bc5d8f27c3fd · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.570688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.570688Z digest=sha256:8afab8aaf7c2f4223089ce33785dd233aa34dad45bf4e76ca5ea3fa0f662ef2f

Observation 982ef8cc-470e-4f0a-bd50-3f8e1e46ee80 · outbound

This paper cites FLIRT: Feedback Loop In-context Red Teaming.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning FLIRT: Feedback Loop In-context Red Teaming

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.687422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.687422Z digest=sha256:1dbfc7e61fe6175ec4a15e96f73ee693e9ecfb108087461787ee4d7b2214217e

Observation 939c8fbd-d949-46ea-936d-caa97c95d135 · outbound

This paper cites Gradient-Based Language Model Red Teaming.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Gradient-Based Language Model Red Teaming

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.691260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.691260Z digest=sha256:76b61461498c044887044c256e04e4cd95e05a200e7d2541aace9ce1e0ddd507

Observation 1a297fb3-6279-4d64-883d-86c94f074646 · outbound

This paper cites PAL: Proxy-Guided Black-Box Attack on Large Language Models.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning PAL: Proxy-Guided Black-Box Attack on Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.695356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.695356Z digest=sha256:9070f2c09bd59dfdf2fc51449c75514729de47b47507a61d98a04f7ca99d12de

Observation 05874896-e918-4dcf-8269-3d8d3d9a551e · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.699423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.699423Z digest=sha256:5e7a9e408afcfa57a2fc130b60917f4f0d09a47fb007d344130db1a8611bb8de

Observation 73189aad-9820-43d7-8eda-0d60eb61fa66 · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Improving alignment of dialogue agents via targeted human judgements

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.703622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.703622Z digest=sha256:ce2fa1b5da60e2cf8dae7dfded468e7e817b20e9212b703d91393bb4caf86760

Observation d9f232a7-fddc-4604-947a-518bfb78ccda · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.707434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.707434Z digest=sha256:86dc9ead0d89a95e762d50ea4a713009dab4364b9e08ebb70a0fb78241dfa2f1

Observation bc297b22-cd09-48cc-af46-d3c0fd9b0cb8 · outbound

This paper cites Towards evaluating the robustness of neural networks.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Towards evaluating the robustness of neural networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:38:37.686863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T04:38:36.711351Z digest=sha256:23ae328b314b2adfcd25e4bf74a38aa664b8edf94cd9a63c12f5cf1a05f16427

Observation b5221044-d131-4726-b209-7bc3e89ce20e · outbound

This paper cites Gradient-based Adversarial Attacks against Text Transformers.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Gradient-based Adversarial Attacks against Text Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.757795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.757795Z digest=sha256:ad522ef89acf7d5c8df3f04292cb7cb5e2435f45a9f7822026be39a3cf545092

Observation d15b9a1c-e6d5-4571-a6f5-9c82a6980dcf · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Categorical Reparameterization with Gumbel-Softmax

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.829242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.829242Z digest=sha256:e17dd4b206bc5bbb173307a093d3cf99499bd9ed973e903bc934baf957a4b07a

Observation 5185ca4f-084b-450b-8f0b-90d1ccc70238 · outbound

This paper cites Universal Adversarial Triggers for Attacking and Analyzing NLP.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Universal Adversarial Triggers for Attacking and Analyzing NLP

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.836773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.836773Z digest=sha256:5882bb57057096604ba339778841449bb37c5d971350c1902121367d6737f70a

Observation 42180de9-4d5b-4b7d-bed6-6c96075f55ef · outbound

This paper cites Robust Conversational Agents against Imperceptible Toxicity Triggers.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Robust Conversational Agents against Imperceptible Toxicity Triggers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.841457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.841457Z digest=sha256:3466db4b8ca4a14427debecf841c5727863fab0cfccf852be7a904cd23aa7b9a

Observation eac0baf8-8037-4944-9c6b-671b07c2279e · outbound

This paper cites Automatically Auditing Large Language Models via Discrete Optimization.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Automatically Auditing Large Language Models via Discrete Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.845185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.845185Z digest=sha256:e0d385e6406b14ab16fdf7223abd94d9d51192c495273a04a29e52ea38e7b541

Observation 2e0147cc-0a79-4dc0-953a-3688a585d55d · outbound

This paper cites Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.849025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.849025Z digest=sha256:5f959bf84e2297da74c2bad833e7c204e87b830d6c9360d05421624bbef0c12b

Observation 08739c8f-2184-4f18-bb4b-202f215b8875 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.853362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.853362Z digest=sha256:3effd7c448a46d92587a13177414d4381b87f59bfae3ed6b78c304700da8fc99

Observation 233c255d-e8ac-46c5-b10a-390fc426c1ca · outbound

This paper cites Adversarial Training for High-Stakes Reliability.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Adversarial Training for High-Stakes Reliability

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.857266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.857266Z digest=sha256:5c3c8f1858d4331cb4ca607959b9fb9420a054eadb3cf9b233ed259005383cb4

Observation d7042796-280e-4881-bb20-d605d86b1d02 · outbound

This paper cites Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.861024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.861024Z digest=sha256:daf170a568e1bd90ea48d7d03b44b438d2d4eca237d64bda0b9c23f566381673

Observation 1166a25e-124b-40ba-97e6-955b28124f1f · outbound

This paper cites Dynabench: Rethinking benchmarking in NLP.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Dynabench: Rethinking benchmarking in NLP

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:38:37.641205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T04:38:36.880492Z digest=sha256:f341e1214a18068261d7ad69b25d205163dd955499ec33666eb456abd47722b8

Observation 2cb2de59-09bb-4cb6-a524-6b0eadcbe133 · outbound

This paper cites doi: 10.18653/v1/2021.naacl-main.324.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning doi: 10.18653/v1/2021.naacl-main.324

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.939010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.939010Z digest=sha256:e6c1ab355639f4c44cd291f018985fc0a449e8a7090b5785d7d66f480ebc11bf

Observation e5b8c7e7-81f2-4b82-aa35-284e15c98033 · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.990324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.990324Z digest=sha256:9d2aef3f2c11592051dff7f66e18d91ef201b9616e1855445b9f59644d266095

Observation 75bd1147-8474-4be8-a15a-a528a8d57538 · outbound

This paper cites Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:37.043217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:37.043217Z digest=sha256:64fafb52147d3f1eb60b8c525643c6374b955b7faac7121d6cc44f3554ece4cf

Observation ca162ccd-40f3-4202-81b4-f32011e6a94f · outbound

This paper cites Curiosity-driven Red-teaming for Large Language Models.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Curiosity-driven Red-teaming for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:37.047366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:37.047366Z digest=sha256:b34162a9161e4c1d12bcdbcf1afd22bf61d86da3168c1b04f53f27da63c1875c

Observation 1c403507-d700-42df-8865-88e97d87de29 · outbound

This paper cites MART: Improving LLM Safety with Multi-round Automatic Red-Teaming.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning MART: Improving LLM Safety with Multi-round Automatic Red-Teaming

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:37.051918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:37.051918Z digest=sha256:63aa1d15fd5aabebeda14220d2881d39940e1178dcd3773d66b85d1308ece22f

Observation 42c4e0ae-1621-409c-8faf-9b0ea080fe6b · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:37.055923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:37.055923Z digest=sha256:2b84f1635469d2e572e9620ebc978f89e35e42c570be7e8c2c163a30ceab65ed

Observation 85991f10-4d48-445b-bea2-3e532e1bb01c · outbound

This paper cites Measuring and mitigating unintended bias in text classification.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Measuring and mitigating unintended bias in text classification

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:38:37.631133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T04:38:37.059454Z digest=sha256:a292a213cb8b152ad97b2d9e2fe0ed510a4cc269b212704b8f11bc4b095a17a2

Observation 9d3caa98-bb36-45a4-8ed2-e9c3daaa4868 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:37.062344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:37.062344Z digest=sha256:f16905a876b2d751ad04d981d083a3899603f721c42b09d687ba30fb75537f1c

Observation aafab37d-06c4-4361-a6d1-f223b39fd5ad · outbound

This paper cites Text and Code Embeddings by Contrastive Pre-Training.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Text and Code Embeddings by Contrastive Pre-Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:37.067103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:37.067103Z digest=sha256:9bc511f02a0bec07671c199fcce94c1095816f79e15f3c0bc8e0843d3d3d7272

Observation a8284c02-f3aa-4b32-908e-d09f7d04fed6 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning WebGPT: Browser-assisted question-answering with human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:37.071018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:37.071018Z digest=sha256:f99ed1c8633ef9058b9e0d31e6c29c02a221746ed5096a12f22ff5b3641b1e5b

Observation 814faea4-dd24-4b84-a8da-9fe04d1328ae · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning ReAct: Synergizing Reasoning and Acting in Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:37.074346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:37.074346Z digest=sha256:58d677dfc81e0d76468e52322fb31b20d93ced4b10a4af1c4107f13391047e2c

Observation 4320da0f-87a2-46ce-8a50-41aa3f79efdd · outbound

This paper cites safety jailbreak.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning safety jailbreak

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:38:37.608339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T04:38:37.185167Z digest=sha256:ffd233142c8aa3a7bc1aa3517c20753b91e13db3fecc04cab7d69607b63deee8

Observation efa3be7c-51dd-4c59-99fb-dc2ea3882446 · outbound

This paper cites HotFlip: White-Box Adversarial Examples for Text Classification.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning HotFlip: White-Box Adversarial Examples for Text Classification

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.833046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.833046Z digest=sha256:ff4c698d3a73d4a11a5809dd5abbec1eef68ac898546bda4600927d2c92b12a5

Observation 93787350-0028-449a-9063-78d9f4c6f6c6 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.715580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.715580Z digest=sha256:ff087fbe1462d77abcce6bb234965c11b73e0ac0749d5e35a3485660c2f6abc4

Observation 807791ba-3657-485e-94cc-043103b873cd · outbound

This paper cites The Woman Worked as a Babysitter: On Biases in Language Generation.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning The Woman Worked as a Babysitter: On Biases in Language Generation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.517644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.517644Z digest=sha256:397f9c7d8e5d6da22765b1ca5979c887635ca972fd1c6902c8f7d04c870a1b5b

Observation 08db9907-3aa0-47c1-99be-94dc0272f2a4 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.522179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.522179Z digest=sha256:f68e46bc08f9a99fce23b028469730fee38b657163cc8057776a8a24d9f80b6a

Observation 6ee5f4db-4ee8-44d9-943b-c338d9c299e7 · outbound

This paper cites Red Teaming Language Models with Language Models.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Red Teaming Language Models with Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.499202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.499202Z digest=sha256:c313df8027ed1d6dd76dfdd5a2dd35da23f2d512eecf3cbc286728def1be8209

Observation cb77f118-3d69-4c1b-b2b7-67f426225bb9 · outbound

This paper cites GPT-4 Technical Report.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning GPT-4 Technical Report

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.479367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.479367Z digest=sha256:1a118cf5c5eea1b878b79563f761e376cf0a3cee11b75f72460f747b1a1040f0

Observation 97724b42-e4b7-4616-a72b-0a0ba2a48971 · outbound

This paper cites Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-Fran¸ cois Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, and Jordan Boyd-Graber.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-Fran¸ cois Bouchard, Chenglei Si, Svetlina Anati, Valen Tagliabue, Anson Liu Kost, Christopher Carnahan, and Jordan Boyd-Graber

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:38:37.842335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T04:38:36.656614Z digest=sha256:48affd60ffe6889414647ec467db6c5d7a9415f56a81f2c5c6c2f2c619c2797b

Observation e54732c8-b4b8-4a18-8a83-1a221d2700c8 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.490989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.490989Z digest=sha256:b79dbd79d7bfc54dc4636e6d3df6afbeb6ad202376c2b40936bc84b57072b9fa

Observation bfd78291-d084-4771-97dd-d87b2e086922 · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T04:38:36.495091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:38:36.495091Z digest=sha256:b76d13c2740f4797355d2fbca2f6adca7180ec51bd42470374cea6b9176c7081

Pith citing papers

Observation 2b1b7a8a-c76d-4627-961e-6017443a229b · inbound

Jailbreaking to Jailbreak cites this paper.

Jailbreaking to Jailbreak Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T17:02:47.195821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:02:47.195821Z digest=sha256:f324201b17ac0bf0d8d8795275218a7142fee058fd78bb4b00dbee3364cbcf6d

Observation a06ea541-bfea-4834-9e5e-d7489f9e8c11 · inbound

When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines cites this paper.

When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:23:48.057420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:23:48.057420Z digest=sha256:bfafdd6dbdad8f80f0dcec6e90292be09e323f3217556408e4e9f62ef035684b

Observation a57cad7f-0a66-442d-ba67-07d698debdec · inbound

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning cites this paper.

Jailbreak-R1: Exploring the Jailbreak Capabilities of LLMs via Reinforcement Learning Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.215841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.215841Z digest=sha256:d3f1236dd1016ef351b7d272f8fb17440c8653260f8db2aa06b42358ac295af1

Observation ac7ef433-3e74-412b-8fe7-9589df932fb0 · inbound

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models cites this paper.

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:06.264820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:06.264820Z digest=sha256:a1fb022c85754dfb6e20fbc57a5d8d2072ef3306bde2791afe7509a256f1459c

Observation 536888bb-9079-4d5f-ab57-719aa7f2a0f4 · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:52:07.721429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:2e58bd93131e873208bead29f58df7974c913589c858a5b62322f45434cde46e

Observation f65c26c6-c0b8-4993-b6c5-7d009c0c4e55 · inbound

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models cites this paper.

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.401969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.401969Z digest=sha256:a135e91c6b582ee41712cda6fa34e2e7085dd87e4c8c535427de19d52e69008c

Observation 4861d50b-d0bd-417c-95cb-0513527f72a6 · inbound

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts cites this paper.

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:54:21.586690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T20:53:58.198974Z digest=sha256:fbfb7bfd66567c81bd457215942f4ff55c89a3ffb6278300b2c23a34dea43c01

Observation e8a800fe-adaf-4bc1-b6ff-5bf651c93279 · inbound

GPT-Red: Automated Red Teaming via Self-Play at Scale cites this paper.

GPT-Red: Automated Red Teaming via Self-Play at Scale Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T01:12:44.118461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:12:44.118461Z digest=sha256:70665df92c861ec2e4043ed60774764cef0855e144b2cd89fe0439e478d13b26