Pith. sign in

Paper Citation Record · LEDGER

Beyond Reward: Offline Preference-guided Policy Optimization

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2305.16217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.16217 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:45:35.949355Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T21:40:21.053123Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b77ac3fd-0a00-4d17-acf3-b3352b88eeda · inbound

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe cites this paper.

Comparing Few to Rank Many: Active Human Preference Learning using Randomized Frank-Wolfe Beyond Reward: Offline Preference-guided Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:35.949355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:45:35.949355Z digest=sha256:6a69c373a9fdd3fa7c361fec9a754b6af0de6d94f883662edb4ea8c80eaa8f92

Observation 39893ad9-687e-45ad-949f-f43fe432dbf1 · inbound

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries cites this paper.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Beyond Reward: Offline Preference-guided Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.566244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.566244Z digest=sha256:c368926b0511555791471fe3dff6682bcac7c12526c4a24feb0d60acc129e876

Observation 3ebf8c51-d315-40b4-bddf-d33506d93911 · inbound

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning cites this paper.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Beyond Reward: Offline Preference-guided Policy Optimization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.007912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.007912Z digest=sha256:e5be5a0ac521001513fd18188fa0e12330098793e40ed64fb6f8f29f3c19ae01

Observation e2b62f56-ad88-4238-a446-7e976e98ffd3 · inbound

CTR-Guided Generative Query Suggestion in Conversational Search cites this paper.

CTR-Guided Generative Query Suggestion in Conversational Search Beyond Reward: Offline Preference-guided Policy Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:53.964866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:53.964866Z digest=sha256:b78bef384cdd00a5c124856d4e6f959941dd9b7b2f171dd6d762d757612843f8

Observation 807c4657-a484-4c82-bf69-c6e9493dd44a · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Beyond Reward: Offline Preference-guided Policy Optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:40:21.057067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:eb114d476f60894b4ccdf89f9fa8d9fb6eec6403f9e06c4fa3cb520e91518bf4