Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2306.14111.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:10:42.886665Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:09:51.330383Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f0ca88e5-076b-4d1f-974d-5f4c209eb44a · inbound
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Is RLHF More Difficult than Standard RL?
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc6e284f-0f30-4db7-b4a7-e9324e6af1ae · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is RLHF More Difficult than Standard RL?
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68c03123-55da-4bc6-a574-bbf61d4e86f5 · inbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Is RLHF More Difficult than Standard RL?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a076c2a1-5fe0-4a7e-ae2c-54027c45ddc9 · inbound
On the optimization dynamics of RLVR: Gradient gap and step size thresholds Is RLHF More Difficult than Standard RL?
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6befe8be-e0f8-434d-84d1-92d2498a0dd2 · inbound
OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Is RLHF More Difficult than Standard RL?
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52f4111c-04ef-40a4-8821-5420410760d2 · inbound
Convex Optimization for Alignment and Preference Learning on a Single GPU Is RLHF More Difficult than Standard RL?
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5604856b-fb3b-4589-a8f2-4086f6b651d6 · inbound
Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Is RLHF More Difficult than Standard RL?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d598499-4ca3-44e8-90fc-58b233a99fd8 · inbound
Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Is RLHF More Difficult than Standard RL?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b36b1692-97d3-438d-b55e-ac2f3ea1231f · inbound
Finding Stationary Points by Comparisons Is RLHF More Difficult than Standard RL?
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.