Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

As of 7 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2506.08266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08266 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:21:58.815824Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:52:19.102106Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T08:31:16.863245Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact4
  • verified fuzzy17
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56377b31-6649-45fe-88dc-a8000252c936 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.599591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.599591Z digest=sha256:d2075998e99507cde788cc4c986c5ae211ec3f82bdcc8a6e67d951b1dc577726

Observation 1d856b60-19ac-4048-827b-1bafa03f8893 · outbound

This paper cites Constrained Markov decision processes.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Constrained Markov decision processes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.604175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.604175Z digest=sha256:28c0dd6323e7b2371032e6e87f47cfeb8a1a122a7c310c2acd450e46ac10e78e

Observation d6aa3d0b-ca63-4b80-8d0d-c16bd0df323a · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.608048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.608048Z digest=sha256:6479cb0bf8fb7cace6a8edfb116ba62dc2454098f1d10d0432a2ba1637e59565

Observation fbe55287-cb2e-4b07-84f9-2299ce6ab95d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.612077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.612077Z digest=sha256:afc9142810a4895eb5dcd2decb3a3c57533078f00fc7ffc391cd4a25ab5b446e

Observation c12c71ee-0016-4ee0-9047-9f03ee6fe616 · outbound

This paper cites Convex optimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Convex optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.615943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.615943Z digest=sha256:e29123733fdff6a3486d7d0e0646a0979fc9d8b8554670e1e4b07b190064cba6

Observation d59d16e5-e0a7-46dc-a0cb-a7bed72ca393 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rank analysis of incomplete block designs: I

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.520895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.619931Z digest=sha256:ac4a79352a0680480be06d9dffcef4e6d739a91186b322a91d309de2b7cc0141

Observation f2e01823-14fb-4cb4-b533-7bcffd74649e · outbound

This paper cites Deep reinforcement learning from human preferences.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Deep reinforcement learning from human preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.624040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.624040Z digest=sha256:e2486e56359948c1840c46a17d3cba72be681f9bf17d9f083cd9c9e04d121e31

Observation 43af2ee8-fb62-4003-a3d9-36fbd6627858 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.628042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.628042Z digest=sha256:74a623eba80de4cb90d257769fa9dfe4796b20abbcfa711229ce92605a26ce92

Observation d0462d37-6aa5-49ff-b43f-4347a7448210 · outbound

This paper cites Policy Gradients with Variance Related Risk Criteria.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Policy Gradients with Variance Related Risk Criteria

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:59.225698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.631826Z digest=sha256:b423bfe4ed91f08ceb5b0df679f4d7bcb5b1dbb91393dc4681bcc18e63ac299d

Observation e182abdc-9e22-47b8-9852-ac3d6f0eb28d · outbound

This paper cites Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Anticipating Safety Issues in E2E Conversational AI: Framework and Tooling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.636063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.636063Z digest=sha256:d91fd2640e9cc7efc6ae375c089baaec9f3a5e312db99ca2140f7e46c527d6c5

Observation e8c042f1-7573-493f-aac2-c9a443c337f4 · outbound

This paper cites AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.639987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.639987Z digest=sha256:37873435f72e90d3fca776452bce0d32f018108bc9e7213c2549d0bc610684df

Observation 2f07684d-8a34-45e2-9aaa-acf16bc1838d · outbound

This paper cites Fundamentals of optimization theory with applications to machine learning.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Fundamentals of optimization theory with applications to machine learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.509176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.643947Z digest=sha256:388f579426950f86528364479a023f3f2a71cb133f72b5374924805aef5ecfdd

Observation 16df702a-b512-4412-854f-3cd7bae3fe74 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.647177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.647177Z digest=sha256:2c0fe783caad73b0d97108be20ea8d8a98972cd954341b28721e671d7118cab4

Observation 3268aad4-55be-4c92-9fa7-318e0e010cf5 · outbound

This paper cites Scaling laws for reward model overoptimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Scaling laws for reward model overoptimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.497961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.650410Z digest=sha256:e4e87c1eecb598a2a107cf4656f510c2da6e302a1cce74c3d378145f8a242edd

Observation 94183662-19cc-44e8-b867-3065a1951642 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:21:59.485906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.653546Z digest=sha256:0b9d658e3c27fd60a252c564306c2152dd7690210725387dec530040990577a1

Observation adafa1f0-2f03-4377-afbc-cc0c6935a6eb · outbound

This paper cites Fairness guarantees under demographic shift.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Fairness guarantees under demographic shift

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.475161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.659981Z digest=sha256:8cc8f60ddce34c09d701edda41684e859e86bd11ff5369c696279c59dee19be2

Observation fb60a32b-e206-4001-9b95-87797be7d16a · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Improving alignment of dialogue agents via targeted human judgements

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.668478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.668478Z digest=sha256:f8be40186ee410da750164893813258143c1be40cb02bf5f53b60ea838e7ea70

Observation b4fca4c5-6a67-46c0-8bf4-05c628320eea · outbound

This paper cites The Llama 3 Herd of Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.671965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.671965Z digest=sha256:e89e14ccf4ded4c9a7d5702b41085ddbc4d55527a5f8e280d84e264394495835

Observation df3cfe86-df38-4464-b643-b1f6920c81ed · outbound

This paper cites Ethical Challenges in Data-Driven Dialogue Systems.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Ethical Challenges in Data-Driven Dialogue Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.675367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.675367Z digest=sha256:385b2cc5c4164d1d72d8d972f9d53542601084815a9ac376679a4b9562feb819

Observation d3b9665f-e183-41e6-9990-aaf7b1ae7336 · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Probability inequalities for sums of bounded random variables

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.679396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.679396Z digest=sha256:16a0d040bbadd70ffb1c5f93167c80ddc251533fe0412bfc7a685164c2e0aed6

Observation 5bf73396-d333-484c-8515-8a597fe4ae00 · outbound

This paper cites One-shot safety alignment for large language models via optimal dualization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints One-shot safety alignment for large language models via optimal dualization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.456542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.682720Z digest=sha256:dbe07f93b5d71e47f3421084a9ba28dc8ff2a5cf0595ddbf69307fc7b1968c0a

Observation 085d504c-bea5-441e-8287-5688ca3b64fa · outbound

This paper cites Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.685785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.685785Z digest=sha256:90f156619bb7dc4029d4da000da2a1f4f502b3c977cf982699fb4db054bfc263

Observation 4b85f2f4-dc88-4f8d-a4d0-1933183454bd · outbound

This paper cites Beavertails: Towards improved safety alignment of LLM via a human-preference dataset, 2023.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Beavertails: Towards improved safety alignment of LLM via a human-preference dataset, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.689205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.689205Z digest=sha256:6e03e8fdd871a68a1e5821af5475f5b03ca07e8de2d0b6939c6dd101c98ac7a5

Observation 3a4a982a-e92c-4f25-8552-7c6e80c6392b · outbound

This paper cites ChatGPT for good? O n opportunities and challenges of large language models for education.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints ChatGPT for good? O n opportunities and challenges of large language models for education

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.444817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.692927Z digest=sha256:fc999a2e8a61f1faf223a03a93012a9ae12fcbed21dde981904a8497277faaa5

Observation 78efc2c2-f73b-463f-bdd9-560b19c3c745 · outbound

This paper cites GPT -4 passes the bar exam.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints GPT -4 passes the bar exam

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.432588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.696131Z digest=sha256:23507e0764eda982155306367e4b7e3385c3186c7cf363ddc04e903b581d7618

Observation a9a7b0f8-a552-4b74-bd8c-f14dc96d2620 · outbound

This paper cites Buy 4 reinforce samples, get a baseline for free! In DeepRLStructPred@ICLR, 2019.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Buy 4 reinforce samples, get a baseline for free! In DeepRLStructPred@ICLR, 2019

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.699651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.699651Z digest=sha256:48f6662a0d7c1cec8c6192b1c2ec65ee640604a1e7a84e099faaf3ced8f6c583

Observation 31cb7100-2d54-4f7c-bc55-069b11578069 · outbound

This paper cites Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepa \ n o, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, and Victor Tseng.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Kung, Morgan Cheatham, Arielle Medenilla, Czarina Sillos, Lorie De Leon, Camille Elepa \ n o, Maria Madriaga, Rimel Aggabao, Giezel Diaz-Candido, James Maningo, and Victor Tseng

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.413594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.703396Z digest=sha256:db5933004ea96ee6e21e8782ff34a134a374ef34f5f0e794992793c5d6306736

Observation 094a5a8d-af98-4b91-83a0-bad6ed9f6a95 · outbound

This paper cites Enhancing LLM Safety via Constrained Direct Preference Optimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enhancing LLM Safety via Constrained Direct Preference Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.706600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.706600Z digest=sha256:77ba5bd5c051001196661e57d6c1771e2d7869812937bff2d0be38980d42a08d

Observation 00e2e0a5-26c9-4697-956f-4b4fb406a4e0 · outbound

This paper cites Offline contextual bandits with high probability fairness guarantees.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Offline contextual bandits with high probability fairness guarantees

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.401561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.710273Z digest=sha256:eb53c6f3859057d0a402dcf8dcbd2932c720f666b87d12c057361572fedb4009

Observation fc679d97-1a53-4ad8-b491-8692792fbc38 · outbound

This paper cites Krumholz, Jure Leskovec, Eric J.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Krumholz, Jure Leskovec, Eric J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.389891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.713452Z digest=sha256:1cc913b5b230e09a26b6be8d9230ce0e67fb2ea39d731375540184a055b93072

Observation bce4a920-bc4e-44b7-b7fd-fb5af4712eee · outbound

This paper cites Rule based rewards for language model safety.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rule based rewards for language model safety

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.379311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.716656Z digest=sha256:95d25e973a9a5212b5eb35f28b558aed4e79a68f0377c7b1029e9dbaea67a44a

Observation ce0cb1d9-f734-47cb-b6ed-741e6ad2addd · outbound

This paper cites Training language models to follow instructions with human feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Training language models to follow instructions with human feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.720364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.720364Z digest=sha256:16144abff09c2ba7f3fe95db4303ebaa1cff841f4afc338b3ede7b91067b1998

Observation 08ef3aa0-5ad6-4b8b-a146-bd0f8c20961f · outbound

This paper cites Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:59.026775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.724147Z digest=sha256:365ee8cb842ec516a2b2db9eb60bca4216c74d33b42cbcac8e1ff712ba0c0311

Observation f5776140-9477-4cb4-93ed-fd1f3ff2abba · outbound

This paper cites Qwen2.5 Technical Report.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Qwen2.5 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.727951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.727951Z digest=sha256:7cac2fa19013295e56d0b258731cca956c15324f774425085e12736b0565f608

Observation 6dba9141-8a58-4692-a9b2-e1d3b097c395 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.731769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.731769Z digest=sha256:e9b2fc437202703fcbd2b3c0bce5c183905658b6f2ff01b9dc4421cec068bd29

Observation 14082912-204e-42f1-be50-db845eab2700 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.735489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.735489Z digest=sha256:9fa51fa66c9667243eb26114784858d91cdeb9e15d70e2528cb1043ee643e552

Observation 681def66-995e-4f7e-b389-928d39c912c2 · outbound

This paper cites Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.367354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.739826Z digest=sha256:105fe1175046656bb954452a84a860aa348048988430e265dca2468511645f49

Observation feb3e995-1e75-40f7-870f-c1a9fbca64fb · outbound

This paper cites Simultaneous statistical inference.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Simultaneous statistical inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.356121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.743188Z digest=sha256:8bd0bea60779fe9538d622659dddac9760fb38ee0edd5e4aa95e092703c323a3

Observation 5fcea960-3005-4e63-bfb7-1b218532842a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Proximal Policy Optimization Algorithms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.746310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.746310Z digest=sha256:c2b4bce37ada841f1c710f791b4ac4886a8a32c49512fee72a1c965f5e59e415

Observation 0425a0d8-0177-4019-85af-d75be4267884 · outbound

This paper cites Learning to summarize from human feedback.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Learning to summarize from human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.749826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.749826Z digest=sha256:64d94b0f359289f4b6f149fc064ee879547e02ff74ff6a4e209ce12f30e2ee05

Observation 449cef09-448b-4918-8f7b-1c2c972aa1c0 · outbound

This paper cites The probable error of a mean.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints The probable error of a mean

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.345555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.753669Z digest=sha256:a545c3580169d9191f06f91026d3eb37d64155a2ff6526ce6f3c282c86aa5b5b

Observation 7de895b4-008b-4581-abd2-f98f8bea42bc · outbound

This paper cites Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.756875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.756875Z digest=sha256:5438837148405fa4f0296d1b3e8a607e32b857ece4624b4b8115c5b0cd377a67

Observation 5fc9254e-270a-4133-9587-da6866bd2b74 · outbound

This paper cites Hashimoto.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Hashimoto

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.760319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.760319Z digest=sha256:a5722d5f4ae3b45407a6f151e76a41278fae3139ca24abf38cb4dce84de2c832

Observation d515259b-92f9-4ecb-b751-de47abce81f2 · outbound

This paper cites Preventing undesirable behavior of intelligent machines.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Preventing undesirable behavior of intelligent machines

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.327648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.763641Z digest=sha256:f4b67972d6e556b04e52359d6999ad5ca35db4eb3faf3d3f637977349d373014

Observation 64dc4375-c341-49e2-8d1f-aa55230916cd · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints LaMDA: Language Models for Dialog Applications

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.767276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.767276Z digest=sha256:08730c7b4e73019093a05300bfd2ea7049a411b7dcb971163ddd4c8e403aea2e

Observation c6401466-52e9-49e9-8e0b-1907d3b73769 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.770901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.770901Z digest=sha256:da87745827f3bd7eaa44b25b024851b47f04380aed1061c43c405eb84918299f

Observation ce9dc7c0-7761-480a-a53a-8b110137fd31 · outbound

This paper cites Tran, Rei Sato, Takumi Tanabe, and Youhei Akimoto.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Tran, Rei Sato, Takumi Tanabe, and Youhei Akimoto

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.316321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.774310Z digest=sha256:34d409bb9f20589cd6af1c2628d64421b3910a72bd44ff9a3628db7b97998927

Observation 46904a93-d126-448b-ac6c-959027d47ce8 · outbound

This paper cites Enforcing Delayed-Impact Fairness Guarantees.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Enforcing Delayed-Impact Fairness Guarantees

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:58.932492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.778283Z digest=sha256:e2f3cff68a9f30581125519cfe49344dcc9c162cbd628ada651fae3666194c31

Observation 5b6357af-c971-41de-9d6a-10a85f23d23e · outbound

This paper cites Ethical and social risks of harm from Language Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Ethical and social risks of harm from Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.781959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.781959Z digest=sha256:a8afd0e7ccf0c7f9f5ec76ca1402ef93e6044eec3fb0c4dfef7088e06b1d92ca

Observation 367a0044-733d-4030-86de-d0ac856dc876 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.785508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.785508Z digest=sha256:5665804e5502e6bb9c96d04c0866c2a0d915d2a718ba1d98cac2f0b42da091be

Observation 17688df0-e3b8-42e0-91ee-ffcc0610464f · outbound

This paper cites Recipes for Safety in Open-domain Chatbots.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Recipes for Safety in Open-domain Chatbots

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.788856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.788856Z digest=sha256:fd821bfeaeec57a4d5f1899efe4f45e12e6cb3eea03499abe45f061093dc21b9

Observation 215d5810-a9fe-434f-b1cf-63c3a55c3b00 · outbound

This paper cites Qwen2 Technical Report.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Qwen2 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.792512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.792512Z digest=sha256:18e646c43e79ffa5174cf03b863ebb46d70dc5df1edda3649f03a498d734bb69

Observation 3e7ecda3-d383-418a-9fd8-e16ffa389a83 · outbound

This paper cites A large language model for electronic health records.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints A large language model for electronic health records

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:21:59.298310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.796231Z digest=sha256:e19684b52173fe30c41cf4105494f759a7be7bb1dabb1e8b059ef12d43e71a33

Observation f8037334-eadc-4a1f-8d46-ec3f0acf2601 · outbound

This paper cites Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:21:58.886951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T05:21:58.799911Z digest=sha256:6acee8244ae6c46fd01557d9bea5e5b29dedbd52447f6f42781258a2af81b826

Observation 373be2dd-3351-4c21-85a9-93d010d415bf · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.803923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.803923Z digest=sha256:87d5b7c04e55e58d6e516c65bd38b6a084d38722ce684f118655b829ce0655ac

Observation 4e532ecc-f56f-4b18-b185-9ca210e2b8c5 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Secrets of RLHF in Large Language Models Part I: PPO

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.808116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.808116Z digest=sha256:3af389f87d4a831335c96e71aabeeb9b98c32afaa03fc10d5e25839577e8aa65

Observation 8ba6c1c8-f416-430f-a4e4-1450569e1f9a · outbound

This paper cites Adversarial Training for High-Stakes Reliability.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints Adversarial Training for High-Stakes Reliability

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.812008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.812008Z digest=sha256:62673305ee5a84968fcf7fd7742068e90350ec3e7ade30a1bb407a83a9192563

Observation f98a0e96-415d-4e86-9bf3-5e8cacb70470 · outbound

This paper cites write newline.

Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints write newline

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:58.815824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:21:58.815824Z digest=sha256:71536d7e2806b41f01ee29636b27de22031b592e3a76261d2525b3b14245459e

Pith citing papers

Observation 32637181-00ce-4796-a516-eaa778829233 · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:19.102106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:19.102106Z digest=sha256:718ad20cd42ff69bee11b64624cf7f7e3e3c5da0201002f703c448d5f9e329a4

Observation 1437125e-c3ed-4c8c-85f0-a701b5c4ff4e · inbound

Reinforcement Learning from Human Feedback: A Statistical Perspective cites this paper.

Reinforcement Learning from Human Feedback: A Statistical Perspective Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:13:13.609624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T20:10:43.578904Z digest=sha256:f13efc140ca7317f667aafda89b2a58609ec989546bb6323e7b292c41b728258

Observation 8345e638-7f56-407a-8140-7d5adc8af28e · inbound

Implicit Safety Alignment from Crowd Preferences cites this paper.

Implicit Safety Alignment from Crowd Preferences Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:31:16.865653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T08:28:41.652865Z digest=sha256:28d46fec05a07a81d09727e57d1bc4c78fb75f304fe89170cca850ce20cd3a71