Pith. sign in

Paper Citation Record · LEDGER

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2506.04302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04302 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.444420Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f175e702-865c-41fb-9b1a-e41681f97f73 · outbound

This paper cites Training language models to follow instructions with human feedback.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Training language models to follow instructions with human feedback

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.772227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:19.959731Z digest=sha256:093b5091025ca0115567afda8849c8cac846cac949dcf24b4de1a9d399fdb139

Observation dc690450-026d-45d9-979c-2b506f6cd357 · outbound

This paper cites Unsolved Problems in ML Safety.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Unsolved Problems in ML Safety

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:19.968574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:19.968574Z digest=sha256:3e415fd05ac9749a458c962dee797fdbff43355a12cbfc861e08b0c71cdf376e

Observation 17cc6a00-b7c4-4ae6-8e15-cffb15d9a430 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:19.978526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:19.978526Z digest=sha256:84575eeea82be21130626216b276601b601041f3e5648f9e8dbecbb8a8a2a302

Observation 7a20bf65-7fe3-483d-939c-4c47a640d3ac · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:19.989451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:19.989451Z digest=sha256:155e329e0ab81b1a1fd45893ecc3537f28a5fe54933678c620e75f84c5501e11

Observation 381be7f1-ef81-4ae9-b907-d175e55a1aa0 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.000342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.000342Z digest=sha256:54a0c257a797dbab79bc20190f5756de9416b087cf59e2e11074650b86141949

Observation 0c503a99-6875-4ae6-9011-00c5c4b72415 · outbound

This paper cites Query-Efficient Black-Box Red Teaming via Bayesian Optimization.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Query-Efficient Black-Box Red Teaming via Bayesian Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.011756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.011756Z digest=sha256:a6f728821bfae3974af4303b9865d610d505486da745974de11fcdfa68011a0d

Observation 853fdc58-f2a8-47a3-aa54-12b83aee6b31 · outbound

This paper cites Large Language Models as Optimizers.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Large Language Models as Optimizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.027907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.027907Z digest=sha256:d0d54e10b1436920484d209e8c0bc780662be43bb216c96dc370f9e23de29c27

Observation cac5389b-1754-4149-881a-773667e75f80 · outbound

This paper cites Red Teaming Language Models with Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Red Teaming Language Models with Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.037233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.037233Z digest=sha256:358243013df630e26446f4eb3181014eb7c2a44ad835591873fdb6d4bc60f07f

Observation 382dc7af-da47-42fe-b859-be941488cdc2 · outbound

This paper cites Discovering language model behaviors with model-written evaluations.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Discovering language model behaviors with model-written evaluations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.752229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.043431Z digest=sha256:5089d9f70c8a36620a6f7409fbe91b199beef93f39513b4d351f9e09cc6eda1f

Observation b06c2a86-7fd2-4537-92e4-511cf268330d · outbound

This paper cites Attack Prompt Generation for Red Teaming and Defending Large Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.050296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.050296Z digest=sha256:c6dd340d3456d8e269ec4368a3ad277b445d752b9950beaa3d9c29bf198ec01b

Observation ae34b8d3-fbac-43ec-a207-7dba2ef5b36d · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.057855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.057855Z digest=sha256:d44f6fc467a3d9bdeedb9617ecab1662404c7690c250a35e6e44198980bd23a0

Observation 3e93d054-5718-4965-a335-c1904c8f7421 · outbound

This paper cites Curiosity-driven Red-teaming for Large Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Curiosity-driven Red-teaming for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.071156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.071156Z digest=sha256:1d9426548aeb975885c5f00bffd8c1c85ffd67c35145724e0904a40b10311c72

Observation 87957fe8-5de1-449b-9104-f13c3a4e5268 · outbound

This paper cites DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.732369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.078058Z digest=sha256:22934b79b73fe3bbbbc40b8492e4a4c6bce5f9c0e3fcb16da6c8cea78106cb27

Observation 38a38d83-3fff-443d-a6f5-8bef94cf5d1d · outbound

This paper cites CALM: Curiosity-Driven Auditing for Large Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming CALM: Curiosity-Driven Auditing for Large Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:55:20.946358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.085750Z digest=sha256:34f778bf317d79a4a9675ed18c02f6bca99871681b84afd4c4ffa90a2a28aff0

Observation 9b2881ed-b30b-4c61-9fb6-045bdf8030f9 · outbound

This paper cites CIM: Constrained Intrinsic Motivation for Reinforcement Learning.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming CIM: Constrained Intrinsic Motivation for Reinforcement Learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.716179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.096049Z digest=sha256:495875a8d4eed78c515a37c587ad6b19cc58396c8b16b6dd7b182dfd5352a81d

Observation a9f03fec-4556-4d63-a544-00dd68645c7e · outbound

This paper cites Proximal Policy Optimization Algorithms.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.103060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.103060Z digest=sha256:0d976e0efb58cab71618b990fd990dc947b964e8655b02c4e4e528d300ae01d9

Observation d7c316e8-2e99-4ee5-8f7d-4ed2de119e3a · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.110906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.110906Z digest=sha256:aa13ff5e068617035c2ab403e657c0a0208acd059fa0d3a085e98f129bdde239

Observation 33da88f3-d4b9-4424-ba63-b3cec21e6766 · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Tree of attacks: Jailbreaking black-box llms automatically

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.699718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.118783Z digest=sha256:93da08ac4729049f7459a48b675e0f865b4bd7f14d264d1467186c5c56d7b161

Observation 9a313817-9950-4f83-a6e9-e015f0e9b66c · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.683610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.124594Z digest=sha256:5db0ad6c6dc55abd7003611ed956354e8a50816f7866ae29d7d2408fc6e86410

Observation ad3fbefa-428f-42c0-8b06-d0a1cd315538 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.132349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.132349Z digest=sha256:29729ea1a240509c52f893c4a4e6b1170f380be08fbf3719a0468b844a08399e

Observation 8ea00912-2743-44fc-9d6c-47e46ac6d4a0 · outbound

This paper cites https://github.com/ huggingface/trl.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming https://github.com/ huggingface/trl

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.667522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.138450Z digest=sha256:340ab722904c54676df6e72bdf556d07e928bb83eedec0d8de1fa77134d90fa6

Observation 381dc6ef-ac5c-49ae-b65b-9b98b331c786 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Direct preference optimization: Your language model is secretly a reward model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.650520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.147018Z digest=sha256:148ca57cd9e3afef97ad1f92db87f86ef4d87825463f0dc0be79dca2b821aef2

Observation 908c75a0-f94a-4865-a9f4-8d9bbe7e8e9f · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.161901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.161901Z digest=sha256:c79219f042efa5ff67afdb0d81bd54c4340b75f195b439aea5e98b9add440da0

Observation 08285b22-1570-4643-add2-818f40fc8436 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.167213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.167213Z digest=sha256:85c951cb558627e8f6a08e0e3b8dcf0c493f1758c6ed9d108ac5f3606260c133

Observation 4ed6cbbe-99a6-41cd-ae6a-9d427b1592b3 · outbound

This paper cites Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.175303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.175303Z digest=sha256:33322a20b897afca54dcc4a5583da948748d9304d10b8fcbda626881e5c0a1db

Observation 87e49c1d-7f06-4080-852f-d4786f7fc0d8 · outbound

This paper cites GPT-4o System Card.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.188583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.188583Z digest=sha256:e029b965cd8ab3d8f7a7fff55f6fb441154066a8840edc02a2526af9ad44d71c

Observation 8aeef9cb-e71d-4f57-998c-6696040a3799 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.197684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.197684Z digest=sha256:726d3421ab9bbf92d504bfecc1a2819a088bcf3368dd9c126fd24ac5acc1fe39

Observation b4b230da-84bc-4768-b4c3-268d49f2c8c0 · outbound

This paper cites What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.205546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.205546Z digest=sha256:fc3ae95f6996843048f7d5bd11e62bb2477f25bf0ea90527db3f6b4185bc653f

Observation 0ab1fa87-310f-4571-ad83-a3e2d187c08c · outbound

This paper cites URLB: Unsupervised Reinforcement Learning Benchmark.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming URLB: Unsupervised Reinforcement Learning Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.212381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.212381Z digest=sha256:d63734ea7f39566506d86e53c1159a3279b3619385db37989bbc471477b0dd4a

Observation 40351ced-22cb-4d3f-a5c6-0c5bc8cea678 · outbound

This paper cites Exploration by Random Network Distillation.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Exploration by Random Network Distillation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.217369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.217369Z digest=sha256:be7476a851880f438f9fae7db2e3fb5ec7b1fcf2fccb68e6e9356269a7d3b77d

Observation a180c0b8-31ba-47a6-8997-c72253260825 · outbound

This paper cites Large-Scale Study of Curiosity-Driven Learning.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Large-Scale Study of Curiosity-Driven Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.225697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.225697Z digest=sha256:8b76f3e54cd13dd2a8886ec4da55d79576cdbdcd81921702d07065a24e264ddf

Observation 2a95946f-6344-4dda-9356-a73851590d2b · outbound

This paper cites Made: Exploration via maximizing deviation from explored regions.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Made: Exploration via maximizing deviation from explored regions

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.632220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.231970Z digest=sha256:c2f66713fce1072c7d42a2385925f9d851ab67f366e861ca3aa2391c14e4f170

Observation 87a2b6a5-e034-47bd-b063-0a4a9d478ccf · outbound

This paper cites Aps: Active pretraining with successor features.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Aps: Active pretraining with successor features

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.615141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.242868Z digest=sha256:92c894e53e747d3a446329d23caafa4e90164ec877fcdeea775412d6f18f316f

Observation 9902af5a-2f2c-4bf3-9db7-8616ceb4b872 · outbound

This paper cites Provably efficient maximum entropy exploration.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Provably efficient maximum entropy exploration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.597103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.248052Z digest=sha256:38065ccec0dfc7f1c1ccd3136dff1bfa57863cdb9b2ad2a9eb5e89f39852c979

Observation 46edf65b-7e79-4371-8a30-383151690748 · outbound

This paper cites Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.579769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.254621Z digest=sha256:35ea5fe594ba0635ae9e5f24432f56a1ede7f5eec6f600f79a1732e2391aa77f

Observation 63dc068b-a992-475e-836b-fa81c7bae7ed · outbound

This paper cites Behavior from the void: Unsupervised active pre-training.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Behavior from the void: Unsupervised active pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.563908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.260652Z digest=sha256:dc404b706cea2d4aa8761a5d5427afe6a30424636079867a769ab0bc50b58e62

Observation 9623f91d-a38e-4822-b268-deb1e410a141 · outbound

This paper cites METRA: Scalable Unsupervised RL with Metric-Aware Abstraction.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming METRA: Scalable Unsupervised RL with Metric-Aware Abstraction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.265825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.265825Z digest=sha256:86c701dddc071732210bee0ed69462eeb03b36b3706ba3185824673aab29627d

Observation 5debbec7-ef51-442b-821b-4bede4fed053 · outbound

This paper cites Variational Intrinsic Control.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Variational Intrinsic Control

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.272855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.272855Z digest=sha256:2f0f854c7dbda0901a03ee44b00b8e31cda6436f6cc450f045eba8885adf75da

Observation e03d3241-ec77-42e9-b744-9191bd314146 · outbound

This paper cites Dynamics-Aware Unsupervised Discovery of Skills.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Dynamics-Aware Unsupervised Discovery of Skills

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.279140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.279140Z digest=sha256:67d9dfa17391b1ae76484a33263a80a0ed30c965b492f0c5980506276ebc22cf

Observation 3b9a831c-c371-4a1a-8e9f-7b704e083dd4 · outbound

This paper cites CIC: Contrastive Intrinsic Control for Unsupervised Skill Discovery.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming CIC: Contrastive Intrinsic Control for Unsupervised Skill Discovery

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.286572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.286572Z digest=sha256:6014998852f3eb277b6c18ab5ecbdb4c2caf54fd60a6dc29aadd2a6e8aef3b98

Observation 7edfc9e9-e9a3-4498-861d-cc2b94dbd0e2 · outbound

This paper cites Lipschitz-constrained unsupervised skill discovery.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Lipschitz-constrained unsupervised skill discovery

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.546382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.294060Z digest=sha256:a20676dbef4941020873008dde583b907ffb886e7c9d34f9ddb1a8d7b133e28f

Observation 400927c6-059a-448b-99bc-51f51c43c6dc · outbound

This paper cites Contrastive learning as goal-conditioned reinforcement learning.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Contrastive learning as goal-conditioned reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.529985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.300397Z digest=sha256:314f71c061adb1bf05e9ecac2389930bee1c4976a45705bdb20ba3cbf21d191c

Observation 663a31b7-a700-41ba-ad81-9d45e84ce291 · outbound

This paper cites Behavior contrastive learning for unsupervised skill discovery.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Behavior contrastive learning for unsupervised skill discovery

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.514043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.307494Z digest=sha256:24a86253bbdee126d377e7fd0d7eb477d4dc2b91688450e604bdd7a33b5138ae

Observation 58364df4-b6e3-41b3-b97f-1b27c1a821fb · outbound

This paper cites Constrained Intrinsic Motivation for Reinforcement Learning.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Constrained Intrinsic Motivation for Reinforcement Learning

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:55:20.543202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.314427Z digest=sha256:a390f44ec0e104bf737636ea1d3bff1c5204d8647c4b55b243ce885ba880a109

Observation 6827f958-515b-4993-9e21-ce34312fc62d · outbound

This paper cites Texygen: A benchmarking platform for text generation models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Texygen: A benchmarking platform for text generation models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.495004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.322001Z digest=sha256:3dfdfc476b12d3214451cee567d448afbacff68b113f810eecf60903c7770a56

Observation 54a87242-300d-4a7e-ac4e-925cfcd01723 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.471942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.328339Z digest=sha256:521cd4514e1af0daac776e8ed41d3b234178305e573fe5b581c3e111d30c05a8

Observation 3b7f25ba-2680-4287-a94a-9c13dfc3475a · outbound

This paper cites Limitations.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Limitations

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.445445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.333557Z digest=sha256:1e4d2ea733f9f2d49d2219cf163bee7f38dafb29715246e6da589c14999de3e7

Observation 43689bd1-c7ff-4d64-983c-b29be3ea708a · outbound

This paper cites Thus, we do not provide original theoretical results of the each baseline method.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Thus, we do not provide original theoretical results of the each baseline method

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.421981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.338073Z digest=sha256:e26b4d581596888fc7589b90cd273efed8772ef094191b1d8bbe7df7577127e0

Observation 2e769d34-be7d-4293-9717-f63132a532ba · outbound

This paper cites Our experiment results can be easily reproduced with a simple environment setup, the default config files, and the prepared shell scripts.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Our experiment results can be easily reproduced with a simple environment setup, the default config files, and the prepared shell scripts

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.401126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.344055Z digest=sha256:2908123ff151d30b6f7d61e145383e1628e40c1451b96196b30f0d29a4c18db2

Observation 729bd66c-a64a-44e7-8592-905e34f9ad0e · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.375072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.350750Z digest=sha256:918966bb622ead485365459ec5ea672e3c9cd7395c51d311edd03dcdf271c2b8

Observation 541639fb-fc2c-4f21-b843-2c2018df7c94 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not include experiments

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.353669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.359237Z digest=sha256:eca7bac6cff2d6a0d525e4e568f27d350b6e372aa4b6429a4af6468b67da8763

Observation b11c78d3-4c70-461a-bdad-6b0d36c732ab · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not include experiments

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.331413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.364627Z digest=sha256:938ed79c4b0eba9c61bb4f55212419dc17964e4246c9d6a5ad33f0cecd9d935d

Observation ac8777d8-1fcb-4c87-9979-fd85f64a6bfc · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not include experiments

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.308529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.385226Z digest=sha256:81f68965e0c12e94dcb5e73422b6ad0c06ce42937798278673c82f8e9cdd1e7f

Observation 5f88aa89-5bf3-4ead-bab4-b454b1aa1367 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.289513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.392099Z digest=sha256:e176d83e5c99551381caec58f568c8cd231ffff8b0246b0ac7d059d5759e4441

Observation bc5483a2-b170-4143-a810-6eb120b8953f · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.269397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.398177Z digest=sha256:0e43cb4962856cb21a8bdb4a9a18d3c1204028e636ae8e2c71e81bec6e9b037d

Observation 87174eda-9204-4f40-ad21-68fde8abdc1c · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper poses no such risks

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.251148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.402830Z digest=sha256:f3cbeae83cb7442987f2517fb72255cb919c21a2f8502f0675ae0ef44719ca83

Observation bdc3f549-f526-4d11-8653-dc3203cafff8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not use existing assets

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.234174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.407677Z digest=sha256:47fc03bd45b5d826005b2893c92f304ccaf2131c1c76625ae14bd02d1c5034d8

Observation 3a24e7df-fd5c-44df-ba83-91600cc58698 · outbound

This paper cites Please refer to https: //github.com/x-zheng16/RedRFT.git for the document.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Please refer to https: //github.com/x-zheng16/RedRFT.git for the document

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.217962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.414775Z digest=sha256:b0a55813c346965eba92eacf3faf13c0684594608fb27d219fa752f0c255e76c

Observation d7fb2682-a2e7-4ba4-af7f-21bf833ea638 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.200211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.419930Z digest=sha256:5dc84f4d9527e04adf694c00ad827d1106810ec4045d597caddf68b1e2f0c6d1

Observation 431ab6db-b213-4dd0-aa74-f9bc67e192cd · outbound

This paper cites Guidelines: 25 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: 25 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.180137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.425912Z digest=sha256:796d86b416949c5bd866101332dae7a50a885ba26844cf31558db1b2436bccd0

Observation 8f459fa7-9125-4e1c-9e78-53da5a87cf8b · outbound

This paper cites an unresolved cited work.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:21.162162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:55:20.444420Z digest=sha256:bf8ca97eba380eb88eac5fe47e658e3a89e94e4e69e4f96ea5cbf822b984e57f

Pith citing papers

No inbound Pith citation observations are available.