Pith. sign in

Paper Citation Record · LEDGER

Is RLHF More Difficult than Standard RL?

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2306.14111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.14111 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:10:42.886665Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:51.330383Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f0ca88e5-076b-4d1f-974d-5f4c209eb44a · inbound

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits cites this paper.

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Is RLHF More Difficult than Standard RL?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:42.886665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:10:42.886665Z digest=sha256:741501dc55537797da06634c84fed5f97d398ae1600470f33a91b0ef08156918

Observation bc6e284f-0f30-4db7-b4a7-e9324e6af1ae · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is RLHF More Difficult than Standard RL?

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.561160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.561160Z digest=sha256:7287729b85c9d81f99b4ea435d0bf719ce759efabd81efa6cd0c181ef8dbf3d7

Observation 68c03123-55da-4bc6-a574-bbf61d4e86f5 · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Is RLHF More Difficult than Standard RL?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:12.854603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:12.854603Z digest=sha256:2422373b9fd1dd4f22fad7471126c40269c9957975f19cf0b767a56ab8377cc2

Observation a076c2a1-5fe0-4a7e-ae2c-54027c45ddc9 · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Is RLHF More Difficult than Standard RL?

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.373729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:de60f59c68a785df39d1f7e714442a1de9269df150b10a5139c3943a07d279d7

Observation 6befe8be-e0f8-434d-84d1-92d2498a0dd2 · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Is RLHF More Difficult than Standard RL?

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:40:20.961449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:d41b45f952e3417747b99741e066715ef42f310eecac06b0fab8139917ace69d

Observation 52f4111c-04ef-40a4-8821-5420410760d2 · inbound

Convex Optimization for Alignment and Preference Learning on a Single GPU cites this paper.

Convex Optimization for Alignment and Preference Learning on a Single GPU Is RLHF More Difficult than Standard RL?

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:05:23.015101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T05:01:31.560963Z digest=sha256:58329ce5a5575a1ef1f6a4aeb9d2e0484338188d3da4ed2c88d7d14fde020276

Observation 5604856b-fb3b-4589-a8f2-4086f6b651d6 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Is RLHF More Difficult than Standard RL?

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:55.886189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T02:31:11.200818Z digest=sha256:a54d89531cb97f0dc1fc55aec14e45b4c392d36f4621858f7db58f009adfd98c

Observation 6d598499-4ca3-44e8-90fc-58b233a99fd8 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Is RLHF More Difficult than Standard RL?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:d46dfee90e9fb29d4a98020319241c78bc738abedd1d3eb8b4e5a6815709661a

Observation b36b1692-97d3-438d-b55e-ac2f3ea1231f · inbound

Finding Stationary Points by Comparisons cites this paper.

Finding Stationary Points by Comparisons Is RLHF More Difficult than Standard RL?

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:51.331914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:26:07.720218Z digest=sha256:e35ae7ba90a82a986ea94548e16994484340a6da02586837db7f9b80656609dc