Pith. sign in

Paper Citation Record · LEDGER

Benchmarks and Algorithms for Offline Preference-Based Reward Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2301.01392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.01392 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:00.273913Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:18:06.214750Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 876b151e-b9d3-4fb6-92b3-4596e38486f6 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.273913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.273913Z digest=sha256:b627ad7118b2ffcf0e8271571e2096de9490a09c1520092d507d85a336c36f94

Observation 3931ac7b-a868-4eec-9af1-bf714c641b6c · inbound

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries cites this paper.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.672841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.672841Z digest=sha256:308ec1e7400a75d9453a1c317f22febfa58f6095dd98327eadfb1f547e6af482

Observation ffb01e19-fb44-4a5f-a819-c5e26406ef07 · inbound

Residual Reward Models for Preference-based Reinforcement Learning cites this paper.

Residual Reward Models for Preference-based Reinforcement Learning Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:39.261586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:39.261586Z digest=sha256:b6b4a6aabd4bae821610557657c1b8f8b98265d16fe9d2cfcfe343ce6e1ef8d2

Observation 1454ab01-6b28-4d9b-b892-9307e16442ce · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:40:21.002736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:3bae6cb910acc4af82ac8bd5b5d24f6e062b06758c58af5e1e555a25858efbf1

Observation 62244405-13a5-4f60-aefa-2ef2d81d2be4 · inbound

SPLC: Social Preference Learning for Crowd Robot Navigation cites this paper.

SPLC: Social Preference Learning for Crowd Robot Navigation Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T12:18:06.216403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T12:10:03.492401Z digest=sha256:4af52e8561ec0e6dc87c787afbf2eadddaf31efa0491b64fefae8c076860ece7

Observation 83171e37-a26e-4543-b3c8-9e6034f87ddb · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 281

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:4bb0ed262fb1381598b093032b030447c3a8a0b14772eb877f78e59ef3c807a3

Observation 3bfe8de6-3dbf-4c42-b71c-7908cd313927 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 282

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:05.320411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:05.320411Z digest=sha256:a230669e4cae7d131dc507464c6a480dcc17e41bf6dce82334c56b65270bef5f