Pith. sign in

Paper Citation Record · LEDGER

Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2402.07314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.07314 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:15:46.376259Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:06:55.872592Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 800496e3-91e4-41d9-88cc-260b1ff6e4e5 · inbound

Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems cites this paper.

Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:35:31.494836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T00:35:31.375020Z digest=sha256:fe53927ad048711ba4239f1a7f167bfe051636653a88cd0a946b24c3d77d9fa5

Observation 4a78435b-0abb-422c-95b9-2722526d7a4a · inbound

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? cites this paper.

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:46.376259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:46.376259Z digest=sha256:e057d57c518af372ef4fb2c522fa71b9da5ed27e5ffe3ffdbbdab37551f96df9

Observation 62417c3d-95fc-47b8-9e9b-f5ba558409e3 · inbound

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits cites this paper.

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:43.394782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:10:43.394782Z digest=sha256:23b8157591771eb9f3bf72ccf001e48ba8754531a2bace465da410cd2bdc5a6d

Observation 683f5707-3ded-4ea2-9359-4f792525dc08 · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.337652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.337652Z digest=sha256:2a44071bed5f7484f1bcd00e97d482215365b153a3eef511502bc495f6423688

Observation 681a11cf-4fc8-4558-aaed-db5ba133eae9 · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.027663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:5d5d1735c46ceb8df062a4f0f727ff26b1b28768d43981114ef92a2cccd2176f

Observation 15262199-a874-49e4-8ca7-17e4457fbee4 · inbound

Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions cites this paper.

Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.615988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T03:46:35.908522Z digest=sha256:71aaeb6a052044a712fb5b20bc350a73127e44eacd4df32d0f7c951abdfdadc9

Observation c02ab51e-a296-4c6f-ab53-b9a3347391f8 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.873915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:31:11.200818Z digest=sha256:875ebab703221c92355abf4f10df81af46dc6d3748cbfb69d3523c0c3a0662d9

Observation d4255a95-1077-4dd9-8b92-ceef7e6f0dd4 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:bf00fb0c4ce642d78e4734c141d9531081f2bf03b2120de9e4d13eb3a67ccc8d