Pith. sign in

Paper Citation Record · LEDGER

Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2411.04625.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04625 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:24:30.557757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:06:55.879807Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a5c9c86d-4bf6-4af7-9c22-e8a3cc37626d · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.557757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.557757Z digest=sha256:a6a832edb4c8d8ea20da815f2a548ef2ae0375e660e80892337b4a8522f6a910

Observation c246fe7b-c98b-49c4-ab0f-333531e32621 · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:59.510021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:59.510021Z digest=sha256:6f910847a6fa33ee60a4744a440eb0067e58eaab552ad86cfa9402d2a01a0bd2

Observation e5c294c0-4fdc-43c3-b1bc-f90b10593d33 · inbound

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses cites this paper.

$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:58.743351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T01:13:29.292351Z digest=sha256:a47b355cfa3237350b9168c5cb50e3a588b0f708f214a46761f70741aeef762d

Observation df772fe2-d199-4d0a-a663-157a83450cb4 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.881073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T02:31:11.200818Z digest=sha256:a076d9ee3c28b261beb15914ae1337113e21cd07c0fd0def7d75a151cdc4b168

Observation 5c771a33-d4db-4b2d-a6e8-86f3d462dbcb · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:e32b9abb5cf62a1d7c841baa82f111997bdbac7b52f0bab53245e896166c9721

Observation 26fdd42b-e885-4dd0-a223-ff9cfa071e81 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:bc578e7ba27ba70dc40e5ab7489c4c4885ed7fe34abffe62bc8d6fd2a10e6308

Observation 5a1d4ea9-bd30-4d74-a74e-896919596d63 · inbound

Graph Dimensionality Reduction for Contextual Bandits: Structure-Specific Regret Bounds under Approximate Smoothness and Noisy Eigenspaces cites this paper.

Graph Dimensionality Reduction for Contextual Bandits: Structure-Specific Regret Bounds under Approximate Smoothness and Noisy Eigenspaces Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:55:51.230113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T04:24:32.483334Z digest=sha256:cf26f158637c31f9cbdb58d02efef77b2b7d09e312fc787dfcd040d79e7c4b9f