Pith. sign in

Paper Citation Record · LEDGER

Provable Offline Preference-Based Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2305.14816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.14816 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:20:13.002840Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:29:02.979336Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ba8ba420-fd18-4f34-b324-9b965813f310 · inbound

Combinatorial Reinforcement Learning with Preference Feedback cites this paper.

Combinatorial Reinforcement Learning with Preference Feedback Provable Offline Preference-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T19:20:13.002840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T19:20:13.002840Z digest=sha256:aecef83aba077773d1c5aa69809d4a0ed24748f9e87ec90b0a116098cb909de4

Observation 4c1e7082-13de-40a7-a7a9-ffe05b95f5db · inbound

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO cites this paper.

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Provable Offline Preference-Based Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:09.500623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:09.500623Z digest=sha256:049c442cafbdb92b71d99256ccc51cd9d09241431a256cdce17cad4590dbea0a

Observation 16c687fe-e2a0-4cef-b0f2-4cef349d565e · inbound

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits cites this paper.

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Provable Offline Preference-Based Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:43.538251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:10:43.538251Z digest=sha256:cfc893b349f05252afa226a2b53b45635bc1e905a6c84fd4936196805429b7b0

Observation a3074639-539a-49a7-97b8-496c5ef4f467 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Provable Offline Preference-Based Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.936809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.936809Z digest=sha256:da7a6e33abb627f9fb50d16c5dbfcbea30d73a098b538a87593a3eb907498430

Observation be121500-6594-443b-9b65-15b395ac386f · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Provable Offline Preference-Based Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.030374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.030374Z digest=sha256:a0d3a38833e48ddf130e6d3253622c990e6b7201f48c39c80f22d1152aac8652

Observation faa4c54d-4083-4620-95f4-cb0d669efcad · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Provable Offline Preference-Based Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.418783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.418783Z digest=sha256:e9a1b647a9711d4b072e16c3510a6273a7f202f30e653177a7360c530064ab76

Observation 73973510-7d79-4868-92e3-2ab89a0ef4cb · inbound

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model cites this paper.

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Provable Offline Preference-Based Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T14:03:36.586968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:03:36.586968Z digest=sha256:fadbb46c918eeb2bab6e4c6d62965f166002da44a29ef37256d615d2c6086538

Observation 0c8d5c55-211b-4cc2-88ee-685fe1b9f111 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Provable Offline Preference-Based Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.562069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:11ab8209fffe8eccfc98e1c569e14365c66a5dc6de1b504b9f77175c07eb0bd9

Observation 072c8af7-3370-41b3-a0c1-ca10dfdcd607 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Provable Offline Preference-Based Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.495290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:bec7438e19396932a66d62e0b5a28c648de16d9c6eb607e27cd5df89d1b08daf

Observation efb1a5a2-635e-408b-865b-236fb3a78783 · inbound

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback cites this paper.

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback Provable Offline Preference-Based Reinforcement Learning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:29.063328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:03:48.813600Z digest=sha256:c3158ca75f30e1abbacf1393d6ce9871fe693c92c80c6592b31a3c6fdb756655

Observation cf716446-1873-4fa2-9096-1d4cc83eddd8 · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Provable Offline Preference-Based Reinforcement Learning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:40:20.967674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:4276d4290870af38ce354032ffeecef3e53d0d0f7375478c0ade22b72e5f643d

Observation ad244cf2-7e9a-4b06-b231-43d295016ddc · inbound

Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback cites this paper.

Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback Provable Offline Preference-Based Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:06:03.890094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T02:23:20.208976Z digest=sha256:d29c72e85bef519d6f31898f369ed1efe97f6c03f389f4c1db0e6a2925850762

Observation f8da3cfa-a9c5-4e30-8936-b6b158725598 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Provable Offline Preference-Based Reinforcement Learning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:31.090069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:75e93cc19b0e0c2ad9c22d851923efadb9cefbdb47d0b6b6554419082f050bfc

Observation a6624355-e3cf-4a8d-ae53-9721c458c5cb · inbound

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? cites this paper.

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning? Provable Offline Preference-Based Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:29:02.980864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T22:06:24.412259Z digest=sha256:62faf77fa4e065eed4fdb54cd2012d951628691635b13a15248dbef4712ab83c

Observation 4f46e08b-8347-4ca8-8c67-e650cc688904 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Provable Offline Preference-Based Reinforcement Learning

Reference 228

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:bc13c072151dbe4ae61d180695d1f681ae36a1e87bed148d8b10f1cd15299e66

Observation 454061c9-42bb-493b-b17a-5224417792c2 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Provable Offline Preference-Based Reinforcement Learning

Reference 229

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:58.612845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:58.612845Z digest=sha256:71f1b3cdc9490c6f48388da16e31e1e92e56e689d806991ee7fd91d73ef20a8a