Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2305.14816.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T19:20:13.002840Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T23:29:02.979336Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation ba8ba420-fd18-4f34-b324-9b965813f310 · inbound
Combinatorial Reinforcement Learning with Preference Feedback Provable Offline Preference-Based Reinforcement Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c1e7082-13de-40a7-a7a9-ffe05b95f5db · inbound
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Provable Offline Preference-Based Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c687fe-e2a0-4cef-b0f2-4cef349d565e · inbound
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Provable Offline Preference-Based Reinforcement Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3074639-539a-49a7-97b8-496c5ef4f467 · inbound
Learning a Pessimistic Reward Model in RLHF Provable Offline Preference-Based Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be121500-6594-443b-9b65-15b395ac386f · inbound
Thompson Sampling in Online RLHF with General Function Approximation Provable Offline Preference-Based Reinforcement Learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa4c54d-4083-4620-95f4-cb0d669efcad · inbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Provable Offline Preference-Based Reinforcement Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73973510-7d79-4868-92e3-2ab89a0ef4cb · inbound
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Provable Offline Preference-Based Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8d5c55-211b-4cc2-88ee-685fe1b9f111 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Provable Offline Preference-Based Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 072c8af7-3370-41b3-a0c1-ca10dfdcd607 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Provable Offline Preference-Based Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation efb1a5a2-635e-408b-865b-236fb3a78783 · inbound
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback Provable Offline Preference-Based Reinforcement Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf716446-1873-4fa2-9096-1d4cc83eddd8 · inbound
OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Provable Offline Preference-Based Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad244cf2-7e9a-4b06-b231-43d295016ddc · inbound
Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback Provable Offline Preference-Based Reinforcement Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8da3cfa-a9c5-4e30-8936-b6b158725598 · inbound
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Provable Offline Preference-Based Reinforcement Learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a6624355-e3cf-4a8d-ae53-9721c458c5cb · inbound
When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? Provable Offline Preference-Based Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4f46e08b-8347-4ca8-8c67-e650cc688904 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Provable Offline Preference-Based Reinforcement Learning
Reference 228
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 454061c9-42bb-493b-b17a-5224417792c2 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Provable Offline Preference-Based Reinforcement Learning
Reference 229
Source-reported events for the cited work
Unavailable: canonical work link unavailable.