Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2111.04850.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:19:09.343297Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:09:51.333293Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d9904036-a23e-4ef2-b25c-ac1113470fd4 · inbound
A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1835f1e-c6e8-4cb5-9847-3189e666cdff · inbound
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c62cbf0-9062-4a4a-99ed-a51ca34dc1f6 · inbound
On the optimization dynamics of RLVR: Gradient gap and step size thresholds Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation defe716c-5c46-480f-a793-7748213f5d3a · inbound
OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 07ba240b-ae6f-4e58-b5e9-813337454882 · inbound
Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a19d89e-b388-4be5-b1cd-d27469ec0e8d · inbound
Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 781d6f84-532b-4e3a-83e9-be6b92614edf · inbound
Finding Stationary Points by Comparisons Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 244e8fb4-b71c-45f9-8612-a7f08767163e · inbound
Preference-Based Reward Learning under Partial Observability with Inexact Dynamics Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad2d9dad-f5ce-435c-8795-433203f26192 · inbound
SPLC: Social Preference Learning for Crowd Robot Navigation Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 09abf6a9-e868-4936-8a26-b281ca942b1e · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 264
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ae8a47-bfda-4f03-ad99-e042ad21549f · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Dueling RL: Reinforcement Learning with Trajectory Preferences
Reference 265
Source-reported events for the cited work
Unavailable: canonical work link unavailable.