Pith. sign in

Paper Citation Record · LEDGER

Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2402.10342.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.10342 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:07:09.475512Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:59:06.798546Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 653d82f0-e58a-4f0c-9600-dc62f106ada8 · inbound

Direct Preference Optimization-Enhanced Multi-Guided Diffusion Model for Traffic Scenario Generation cites this paper.

Direct Preference Optimization-Enhanced Multi-Guided Diffusion Model for Traffic Scenario Generation Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T20:07:09.475512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:07:09.475512Z digest=sha256:65c7de41ba523fb665809164675accfe94fc3613103eff6a641d5a4bc6002be7

Observation 79b6d38a-f49e-4578-9de2-766c59c3f55a · inbound

Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates cites this paper.

Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:01:34.174175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T12:58:26.626172Z digest=sha256:cccf5dc626fd5049ffdda73804cec806041bf3f61e5d4bef68547a495f4fdb21

Observation 7526c284-aae6-4f09-a9e1-b9d1e312a23d · inbound

Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback cites this paper.

Policy Gradient Primal-Dual Method for Safe Reinforcement Learning from Human Feedback Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:03.976914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T02:23:20.208976Z digest=sha256:fe43eb3cec9210aa26e82bb0a653ff65854cbe04d4484890bd8977d1d6618d09

Observation 2060ce45-6da3-43bd-9359-594773221f40 · inbound

Distributed Zeroth-Order Policy Gradient for Networked Multi-agent Reinforcement Learning from Human Feedback cites this paper.

Distributed Zeroth-Order Policy Gradient for Networked Multi-agent Reinforcement Learning from Human Feedback Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:22:44.682334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T19:21:53.659764Z digest=sha256:e5d84e173a5a72c744362fd534dfb8fe5db6f79766f542b22c90349613bbeff7

Observation c95c199c-91b2-4ff6-a7dc-76749bedb1da · inbound

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards cites this paper.

Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:59:06.801171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:35:56.749428Z digest=sha256:ed3511630da89de0f92c744bfd26218a5972fa7bc3a82276c68ebb38d95c5c2d