Pith. sign in

Paper Citation Record · LEDGER

RL with KL penalties is better viewed as Bayesian inference

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2205.11275.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.11275 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:10:17.082630Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:27:29.845324Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 20b0241a-bf09-4e85-a708-b778b1c3e90b · inbound

Scaling Laws for Reward Model Overoptimization cites this paper.

Scaling Laws for Reward Model Overoptimization RL with KL penalties is better viewed as Bayesian inference

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:04:53.301252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T09:04:53.129737Z digest=sha256:d2fbd051df63dff3961699bb65213c4984fc526146f319e0d15996cf48ed3187

Observation 51d2c4b0-c08c-4c05-84c8-1ee3ce9edd31 · inbound

CoDe: Blockwise Control for Denoising Diffusion Models cites this paper.

CoDe: Blockwise Control for Denoising Diffusion Models RL with KL penalties is better viewed as Bayesian inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T17:10:17.082630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:10:17.082630Z digest=sha256:702635c1e9d620e771049f3349ed76e83710b85e760b4a5f586bd8e1f539494e

Observation 5f055b0c-aa79-4eea-ab72-558e90452701 · inbound

Discovering Algorithms with Computational Language Processing cites this paper.

Discovering Algorithms with Computational Language Processing RL with KL penalties is better viewed as Bayesian inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:01.136484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:01.136484Z digest=sha256:674cdd1116b22265b5c1ab755378c4a8df09fb898d9181e40fd4f4e25a3919ba

Observation 93333f7d-8616-41bb-b553-deae042a3f7d · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis RL with KL penalties is better viewed as Bayesian inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:56.065982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:56.065982Z digest=sha256:0b7e781fd94f4cd55a92417a0f0a9f9c10fc392775a6b0bcd784430ac3041be5

Observation 3aa4ea9d-4aa4-4680-b70c-b827d8e4e9b5 · inbound

A Survey of Reinforcement Learning For Economics cites this paper.

A Survey of Reinforcement Learning For Economics RL with KL penalties is better viewed as Bayesian inference

Reference 1960

Resolution
unresolved
no resolver link, observed 2026-08-02T18:34:52.804048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:34:52.804048Z digest=sha256:96c413aed811c6ca6e1a51966532d678b6527043e8c05300f79ed0b72b4c6225

Observation 8ea9c9d4-48f6-481e-a9dd-79f0ddbaec96 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow RL with KL penalties is better viewed as Bayesian inference

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.619008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:5c1f122f1937c86b6691d8dd583490e2006ffedf7861af844f8ec63a3b2b9c3e

Observation 4abcc714-8349-4261-8f2f-1f1c616a831f · inbound

Exponential families from a single KL identity cites this paper.

Exponential families from a single KL identity RL with KL penalties is better viewed as Bayesian inference

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.247915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-07T07:16:28.127796Z digest=sha256:341f72265fd4b3a16003c78ec29799db9df7abd515f40c352931fbc2616703a1

Observation ab0a2075-e8f6-41e5-8d5e-75bfff6214c4 · inbound

Binary Rewards and Reinforcement Learning: Fundamental Challenges cites this paper.

Binary Rewards and Reinforcement Learning: Fundamental Challenges RL with KL penalties is better viewed as Bayesian inference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:08.663028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T16:29:16.359577Z digest=sha256:64aa19dfa342df2601913f364fb874a0aa36394d245203d8e277a6fa1daf324d

Observation c4ed330b-c778-45c0-8bda-697343dad096 · inbound

On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective cites this paper.

On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective RL with KL penalties is better viewed as Bayesian inference

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.798482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:52:51.509014Z digest=sha256:6779d6f206145370ecc148648c525fc5b1f738464743ca285309b56d6b0c42d0

Observation 8ac82a3c-dbd6-4e2f-8485-cfeba45aae1e · inbound

The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives cites this paper.

The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives RL with KL penalties is better viewed as Bayesian inference

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:52:08.903060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T02:49:53.547490Z digest=sha256:c240c6e92dfd4bd097cfd90d8460b24498457c561e538ee5a37247e4819bdc53

Observation db64787e-ea9d-4737-881f-b2258609c6e5 · inbound

A Unifying Lens on Reward Uncertainty in RLHF cites this paper.

A Unifying Lens on Reward Uncertainty in RLHF RL with KL penalties is better viewed as Bayesian inference

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.846581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T17:11:46.549150Z digest=sha256:16b72e4b816417cb5ecb488570ed67b27467f3a7d2919e6b7f179898c2ca76e6

Observation dce67dbc-9dce-420e-8da5-547482ef1730 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay RL with KL penalties is better viewed as Bayesian inference

Reference 247

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:1bbf5dd3dd71c328ddd49f1ddc39ee2f71caad0dbaca15e55a13d38c1121d09f

Observation bf3d459f-d00b-45dc-80be-648e96eb1dc8 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay RL with KL penalties is better viewed as Bayesian inference

Reference 248

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:01.139311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:01.139311Z digest=sha256:888b5a3224e89555e9b2921ee15584241a696b816b0b6c5292a0f62dba4cb083