Pith. sign in

Paper Citation Record · LEDGER

Dataset Reset Policy Optimization for RLHF

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2404.08495.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.08495 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:04:30.492218Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.365669Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 19322493-ba98-49d0-bc1a-954f3af1adf5 · inbound

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory cites this paper.

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory Dataset Reset Policy Optimization for RLHF

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T01:04:30.492218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:04:30.492218Z digest=sha256:b0882f463ee374cffc5a550cb3c67cd7a3107998986f4e088800be5705f532dd

Observation 288c2f09-c7eb-49ea-adb0-38ff0434e55b · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Dataset Reset Policy Optimization for RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.467898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.467898Z digest=sha256:cbbfc3ed96167b6c20228097b6bb53d789e0a0f6572b13651ffc9c3a54071e8c

Observation 7207f407-d68d-49bb-8f94-1b724ac82cba · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Dataset Reset Policy Optimization for RLHF

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.778548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:86b8703aa35d54f9b4ce664e38d52c5205d0d46a08b48369afdd4641341264b9

Observation 68f28c37-3a68-481a-bcb0-452025cef58d · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Dataset Reset Policy Optimization for RLHF

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:45:59.539703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:0f3f218cd17c058adfbac01bed1e25ffc24bb3bb74b2bdc4b784fb7335f28a1f

Observation 046b41e7-e346-42bc-9424-2f294768668e · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Dataset Reset Policy Optimization for RLHF

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:30.623863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:e1903e4ada6e7b204326e1395b5aacbd87bd41bce7ea2020e25afc844fc9b1be

Observation 6d94f087-3f58-489c-951e-c3b288a1a96a · inbound

Credit Assignment with Resets in Language Model Reasoning cites this paper.

Credit Assignment with Resets in Language Model Reasoning Dataset Reset Policy Optimization for RLHF

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T21:53:59.235733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T21:50:18.827822Z digest=sha256:ccef0094b07087c1317decb48e0c0c7fee235d38b95fa2f0eac6d0fbcba7c0dd

Observation e3f60e55-fd29-438c-8efd-e73cbdb6f780 · inbound

Reinforcement Learning from Rich Feedback with Distributional DAgger cites this paper.

Reinforcement Learning from Rich Feedback with Distributional DAgger Dataset Reset Policy Optimization for RLHF

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:45.920135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:44:56.667364Z digest=sha256:2fdd11df07af7978d3e02a2eae01768c4c72360389409f7379cd5e52ab14a62f

Observation 453eaf44-1f8d-4e7d-b236-d642194e057c · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Dataset Reset Policy Optimization for RLHF

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.367108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:fd6823a5be57fdf370823c58b661a6589973d3620d53d9a073a04238a84da7ca

Observation a4d2d08b-5f84-4ce7-8f7b-6f20102187b4 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Dataset Reset Policy Optimization for RLHF

Reference 292

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:2ca4cff10a831641cb304b2280b6323a6e843376d939a41976c4c0cb01e1da4d

Observation 707e0ab3-b9e3-4180-95bf-bd4f118689c9 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Dataset Reset Policy Optimization for RLHF

Reference 293

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:06.796208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:06.796208Z digest=sha256:9daed36ad1d30cac229a02e049c57c9893ba61d1192deb8e1ebb7b67d9dda1fa