Pith. sign in

Paper Citation Record · LEDGER

Beyond Reward: Offline Preference-guided Policy Optimization

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2305.16217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.16217 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:58.566244Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T21:40:21.053123Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39893ad9-687e-45ad-949f-f43fe432dbf1 · inbound

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries cites this paper.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Beyond Reward: Offline Preference-guided Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.566244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.566244Z digest=sha256:7a1f5e93afc3632f9300130d8c9b1156343f572cd62cabc47b1ccee92bc59a6e

Observation 3ebf8c51-d315-40b4-bddf-d33506d93911 · inbound

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning cites this paper.

TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning Beyond Reward: Offline Preference-guided Policy Optimization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:40.007912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:40.007912Z digest=sha256:37102cafa04a233606b846e4ef7fae0ef5bb1ef18866d9ba82edb28d088f1d78

Observation e2b62f56-ad88-4238-a446-7e976e98ffd3 · inbound

CTR-Guided Generative Query Suggestion in Conversational Search cites this paper.

CTR-Guided Generative Query Suggestion in Conversational Search Beyond Reward: Offline Preference-guided Policy Optimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:53.964866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:53.964866Z digest=sha256:4813b226912e61cf68e440d468224709778e19560038a45beeec62b38d0d2739

Observation 807c4657-a484-4c82-bf69-c6e9493dd44a · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Beyond Reward: Offline Preference-guided Policy Optimization

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:40:21.057067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:e54841ca69503bf39bbb0bd74c1a1454889b54d6d0dbe3215682520331d098e9