Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.444420Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2506.04302.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.444420Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f175e702-865c-41fb-9b1a-e41681f97f73 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Training language models to follow instructions with human feedback
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dc690450-026d-45d9-979c-2b506f6cd357 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Unsolved Problems in ML Safety
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17cc6a00-b7c4-4ae6-8e15-cffb15d9a430 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a20bf65-7fe3-483d-939c-4c47a640d3ac · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381be7f1-ef81-4ae9-b907-d175e55a1aa0 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c503a99-6875-4ae6-9011-00c5c4b72415 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Query-Efficient Black-Box Red Teaming via Bayesian Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853fdc58-f2a8-47a3-aa54-12b83aee6b31 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Large Language Models as Optimizers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cac5389b-1754-4149-881a-773667e75f80 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Red Teaming Language Models with Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382dc7af-da47-42fe-b859-be941488cdc2 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Discovering language model behaviors with model-written evaluations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b06c2a86-7fd2-4537-92e4-511cf268330d · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Attack Prompt Generation for Red Teaming and Defending Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae34b8d3-fbac-43ec-a207-7dba2ef5b36d · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Explore, Establish, Exploit: Red Teaming Language Models from Scratch
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e93d054-5718-4965-a335-c1904c8f7421 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Curiosity-driven Red-teaming for Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87957fe8-5de1-449b-9104-f13c3a4e5268 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 38a38d83-3fff-443d-a6f5-8bef94cf5d1d · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming CALM: Curiosity-Driven Auditing for Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b2881ed-b30b-4c61-9fb6-045bdf8030f9 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming CIM: Constrained Intrinsic Motivation for Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a9f03fec-4556-4d63-a544-00dd68645c7e · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Proximal Policy Optimization Algorithms
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c316e8-2e99-4ee5-8f7d-4ed2de119e3a · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33da88f3-d4b9-4424-ba63-b3cec21e6766 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Tree of attacks: Jailbreaking black-box llms automatically
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9a313817-9950-4f83-a6e9-e015f0e9b66c · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad3fbefa-428f-42c0-8b06-d0a1cd315538 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea00912-2743-44fc-9d6c-47e46ac6d4a0 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming https://github.com/ huggingface/trl
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 381dc6ef-ac5c-49ae-b65b-9b98b331c786 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Direct preference optimization: Your language model is secretly a reward model
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 908c75a0-f94a-4865-a9f4-8d9bbe7e8e9f · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08285b22-1570-4643-add2-818f40fc8436 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ed6cbbe-99a6-41cd-ae6a-9d427b1592b3 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87e49c1d-7f06-4080-852f-d4786f7fc0d8 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming GPT-4o System Card
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aeef9cb-e71d-4f57-998c-6696040a3799 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4b230da-84bc-4768-b4c3-268d49f2c8c0 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ab1fa87-310f-4571-ad83-a3e2d187c08c · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming URLB: Unsupervised Reinforcement Learning Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40351ced-22cb-4d3f-a5c6-0c5bc8cea678 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Exploration by Random Network Distillation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a180c0b8-31ba-47a6-8997-c72253260825 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Large-Scale Study of Curiosity-Driven Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a95946f-6344-4dda-9356-a73851590d2b · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Made: Exploration via maximizing deviation from explored regions
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87a2b6a5-e034-47bd-b063-0a4a9d478ccf · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Aps: Active pretraining with successor features
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9902af5a-2f2c-4bf3-9db7-8616ceb4b872 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Provably efficient maximum entropy exploration
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 46edf65b-7e79-4371-8a30-383151690748 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 63dc068b-a992-475e-836b-fa81c7bae7ed · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Behavior from the void: Unsupervised active pre-training
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9623f91d-a38e-4822-b268-deb1e410a141 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5debbec7-ef51-442b-821b-4bede4fed053 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Variational Intrinsic Control
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e03d3241-ec77-42e9-b744-9191bd314146 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Dynamics-Aware Unsupervised Discovery of Skills
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b9a831c-c371-4a1a-8e9f-7b704e083dd4 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming CIC: Contrastive Intrinsic Control for Unsupervised Skill Discovery
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7edfc9e9-e9a3-4498-861d-cc2b94dbd0e2 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Lipschitz-constrained unsupervised skill discovery
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 400927c6-059a-448b-99bc-51f51c43c6dc · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Contrastive learning as goal-conditioned reinforcement learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 663a31b7-a700-41ba-ad81-9d45e84ce291 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Behavior contrastive learning for unsupervised skill discovery
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 58364df4-b6e3-41b3-b97f-1b27c1a821fb · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Constrained Intrinsic Motivation for Reinforcement Learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6827f958-515b-4993-9e21-ce34312fc62d · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Texygen: A benchmarking platform for text generation models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 54a87242-300d-4a7e-ac4e-925cfcd01723 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3b7f25ba-2680-4287-a94a-9c13dfc3475a · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Limitations
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43689bd1-c7ff-4d64-983c-b29be3ea708a · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Thus, we do not provide original theoretical results of the each baseline method
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2e769d34-be7d-4293-9717-f63132a532ba · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Our experiment results can be easily reproduced with a simple environment setup, the default config files, and the prepared shell scripts
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 729bd66c-a64a-44e7-8592-905e34f9ad0e · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that paper does not include experiments requiring code
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 541639fb-fc2c-4f21-b843-2c2018df7c94 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not include experiments
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b11c78d3-4c70-461a-bdad-6b0d36c732ab · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not include experiments
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac8777d8-1fcb-4c87-9979-fd85f64a6bfc · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not include experiments
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f88aa89-5bf3-4ead-bab4-b454b1aa1367 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc5483a2-b170-4143-a810-6eb120b8953f · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that there is no societal impact of the work performed
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87174eda-9204-4f40-ad21-68fde8abdc1c · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper poses no such risks
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bdc3f549-f526-4d11-8653-dc3203cafff8 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not use existing assets
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3a24e7df-fd5c-44df-ba83-91600cc58698 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Please refer to https: //github.com/x-zheng16/RedRFT.git for the document
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d7fb2682-a2e7-4ba4-af7f-21bf833ea638 · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 431ab6db-b213-4dd0-aa74-f9bc67e192cd · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: 25 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f459fa7-9125-4e1c-9e78-53da5a87cf8b · outbound
RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.