Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:58.815824Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2506.08266.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:58.815824Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T14:52:19.102106Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T08:31:16.863245Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 56377b31-6649-45fe-88dc-a8000252c936 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d856b60-19ac-4048-827b-1bafa03f8893 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Constrained Markov decision processes
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6aa3d0b-ca63-4b80-8d0d-c16bd0df323a · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe55287-cb2e-4b07-84f9-2299ce6ab95d · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c12c71ee-0016-4ee0-9047-9f03ee6fe616 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Convex optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59d16e5-e0a7-46dc-a0cb-a7bed72ca393 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rank analysis of incomplete block designs: I
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2e01823-14fb-4cb4-b533-7bcffd74649e · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Deep reinforcement learning from human preferences
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43af2ee8-fb62-4003-a3d9-36fbd6627858 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0462d37-6aa5-49ff-b43f-4347a7448210 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Policy Gradients with Variance Related Risk Criteria
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e182abdc-9e22-47b8-9852-ac3d6f0eb28d · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8c042f1-7573-493f-aac2-c9a443c337f4 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f07684d-8a34-45e2-9aaa-acf16bc1838d · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Fundamentals of optimization theory with applications to machine learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16df702a-b512-4412-854f-3cd7bae3fe74 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3268aad4-55be-4c92-9fa7-318e0e010cf5 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Scaling laws for reward model overoptimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94183662-19cc-44e8-b867-3065a1951642 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adafa1f0-2f03-4377-afbc-cc0c6935a6eb · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Fairness guarantees under demographic shift
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb60a32b-e206-4001-9b95-87797be7d16a · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Improving alignment of dialogue agents via targeted human judgements
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4fca4c5-6a67-46c0-8bf4-05c628320eea · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df3cfe86-df38-4464-b643-b1f6920c81ed · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Ethical Challenges in Data-Driven Dialogue Systems
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b9665f-e183-41e6-9990-aaf7b1ae7336 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Probability inequalities for sums of bounded random variables
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf73396-d333-484c-8515-8a597fe4ae00 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints One-shot safety alignment for large language models via optimal dualization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 085d504c-bea5-441e-8287-5688ca3b64fa · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b85f2f4-dc88-4f8d-a4d0-1933183454bd · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Beavertails: Towards improved safety alignment of LLM via a human-preference dataset, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a4a982a-e92c-4f25-8552-7c6e80c6392b · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints ChatGPT for good? O n opportunities and challenges of large language models for education
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78efc2c2-f73b-463f-bdd9-560b19c3c745 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints GPT -4 passes the bar exam
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9a7b0f8-a552-4b74-bd8c-f14dc96d2620 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Buy 4 reinforce samples, get a baseline for free! In DeepRLStructPred@ICLR, 2019
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31cb7100-2d54-4f7c-bc55-069b11578069 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepa \ n o, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, and Victor Tseng
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 094a5a8d-af98-4b91-83a0-bad6ed9f6a95 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enhancing LLM Safety via Constrained Direct Preference Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e2e0a5-26c9-4697-956f-4b4fb406a4e0 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Offline contextual bandits with high probability fairness guarantees
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc679d97-1a53-4ad8-b491-8692792fbc38 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Krumholz, Jure Leskovec, Eric J
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bce4a920-bc4e-44b7-b7fd-fb5af4712eee · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rule based rewards for language model safety
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce0cb1d9-f734-47cb-b6ed-741e6ad2addd · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Training language models to follow instructions with human feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08ef3aa0-5ad6-4b8b-a146-bd0f8c20961f · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5776140-9477-4cb4-93ed-fd1f3ff2abba · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Qwen2.5 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dba9141-8a58-4692-a9b2-e1d3b097c395 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14082912-204e-42f1-be50-db845eab2700 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 681def66-995e-4f7e-b389-928d39c912c2 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation feb3e995-1e75-40f7-870f-c1a9fbca64fb · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Simultaneous statistical inference
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5fcea960-3005-4e63-bfb7-1b218532842a · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Proximal Policy Optimization Algorithms
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0425a0d8-0177-4019-85af-d75be4267884 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Learning to summarize from human feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 449cef09-448b-4918-8f7b-1c2c972aa1c0 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints The probable error of a mean
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7de895b4-008b-4581-abd2-f98f8bea42bc · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc9254e-270a-4133-9587-da6866bd2b74 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Hashimoto
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d515259b-92f9-4ecb-b751-de47abce81f2 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Preventing undesirable behavior of intelligent machines
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64dc4375-c341-49e2-8d1f-aa55230916cd · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints LaMDA: Language Models for Dialog Applications
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6401466-52e9-49e9-8e0b-1907d3b73769 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce9dc7c0-7761-480a-a53a-8b110137fd31 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Tran, Rei Sato, Takumi Tanabe, and Youhei Akimoto
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46904a93-d126-448b-ac6c-959027d47ce8 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enforcing Delayed-Impact Fairness Guarantees
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b6357af-c971-41de-9d6a-10a85f23d23e · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Ethical and social risks of harm from Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 367a0044-733d-4030-86de-d0ac856dc876 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Simple statistical gradient-following algorithms for connectionist reinforcement learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17688df0-e3b8-42e0-91ee-ffcc0610464f · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Recipes for Safety in Open-domain Chatbots
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 215d5810-a9fe-434f-b1cf-63c3a55c3b00 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Qwen2 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7ecda3-d383-418a-9fd8-e16ffa389a83 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints A large language model for electronic health records
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8037334-eadc-4a1f-8d46-ec3f0acf2601 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 373be2dd-3351-4c21-85a9-93d010d415bf · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e532ecc-f56f-4b18-b185-9ca210e2b8c5 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Secrets of RLHF in Large Language Models Part I: PPO
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ba6c1c8-f416-430f-a4e4-1450569e1f9a · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Adversarial Training for High-Stakes Reliability
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f98a0e96-415d-4e86-9bf3-5e8cacb70470 · outbound
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints write newline
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32637181-00ce-4796-a516-eaa778829233 · inbound
Adaptive Margin RLHF via Preference over Preferences Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1437125e-c3ed-4c8c-85f0-a701b5c4ff4e · inbound
Reinforcement Learning from Human Feedback: A Statistical Perspective Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8345e638-7f56-407a-8140-7d5adc8af28e · inbound
Implicit Safety Alignment from Crowd Preferences Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.