Pith. sign in

Paper Citation Record · LEDGER

Dueling RL: Reinforcement Learning with Trajectory Preferences

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2111.04850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.04850 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:19:09.343297Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:51.333293Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d9904036-a23e-4ef2-b25c-ac1113470fd4 · inbound

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO cites this paper.

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:09.343297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:09.343297Z digest=sha256:771dd7bb7c28c5c1e77e03e3b49ecde0d5eb50c0daf56061de1712f72b989374

Observation d1835f1e-c6e8-4cb5-9847-3189e666cdff · inbound

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits cites this paper.

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:42.530610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:10:42.530610Z digest=sha256:df834f665d15eb069b3a2a3e6dd656f872522fa57c1ca3a739340eb3e5c13190

Observation 0c62cbf0-9062-4a4a-99ed-a51ca34dc1f6 · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.354937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:3b35a1002ab8806577e697045e439f3b46b5c597c66a8cbbd2c654bae013f7eb

Observation defe716c-5c46-480f-a793-7748213f5d3a · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:40:21.050286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:938507b6241fef20654049398fa16cf32985473557cb92c173090396e6433b1d

Observation 07ba240b-ae6f-4e58-b5e9-813337454882 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:55.881408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:31:11.200818Z digest=sha256:7c7c6bf5871f06cad3275639b7be89dbdcfeecd001cadc5491449318e3bf11c9

Observation 5a19d89e-b388-4be5-b1cd-d27469ec0e8d · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:e70db6724b44954dcda1caa99098713b1b1770a1cce4b4d6e65c51ffe2352b5a

Observation 781d6f84-532b-4e3a-83e9-be6b92614edf · inbound

Finding Stationary Points by Comparisons cites this paper.

Finding Stationary Points by Comparisons Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:51.334817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T05:26:07.720218Z digest=sha256:54b06b0385ad8c93f02999f6f41b10fe2d7e6a86c5b0ea9e66d9da9d69adcb86

Observation 244e8fb4-b71c-45f9-8612-a7f08767163e · inbound

Preference-Based Reward Learning under Partial Observability with Inexact Dynamics cites this paper.

Preference-Based Reward Learning under Partial Observability with Inexact Dynamics Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.876543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T05:19:16.647402Z digest=sha256:c23a5367703dbf9aeb8cea0793f012605390992ecc2e9fcaa08702c8d86d3ca4

Observation ad2d9dad-f5ce-435c-8795-433203f26192 · inbound

SPLC: Social Preference Learning for Crowd Robot Navigation cites this paper.

SPLC: Social Preference Learning for Crowd Robot Navigation Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T12:18:06.219993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T12:10:03.492401Z digest=sha256:86015366fb33b9ead8d037d1d7854bf684a05b5059468e7ce745e0ca09c07964

Observation 09abf6a9-e868-4936-8a26-b281ca942b1e · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 264

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:6eeaf0b2c7e10e0a1619264e6410a163384eda97ae51102cd70e9e9b2dfb80b1

Observation 83ae8a47-bfda-4f03-ad99-e042ad21549f · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 265

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:03.234561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:03.234561Z digest=sha256:4c300022cb1137630f1c3748b4a02585f96cb7e96edf6d30fb5c5d38efd8cb87