Pith. sign in

Paper Citation Record · LEDGER

RL with KL penalties is better viewed as Bayesian inference

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2205.11275.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.11275 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:32:52.272951Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:27:29.845324Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 20b0241a-bf09-4e85-a708-b778b1c3e90b · inbound

Scaling Laws for Reward Model Overoptimization cites this paper.

Scaling Laws for Reward Model Overoptimization RL with KL penalties is better viewed as Bayesian inference

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:04:53.301252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T09:04:53.129737Z digest=sha256:692b052ed08e581df3ec611d0fa057d6cd97276eed2d62fa88891775d9a7f997

Observation 55a032d8-32db-48b3-affc-f5fd8119b9bf · inbound

Time-Reversal Provides Unsupervised Feedback to LLMs cites this paper.

Time-Reversal Provides Unsupervised Feedback to LLMs RL with KL penalties is better viewed as Bayesian inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T23:21:06.618515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:21:06.618515Z digest=sha256:9cc804d62a40efe53ac2daa2b91f0569501e3ec36bed6f48de02ef6a01197f22

Observation b6138706-4deb-459f-b813-80733664939a · inbound

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models cites this paper.

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models RL with KL penalties is better viewed as Bayesian inference

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T22:57:01.736400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:57:01.736400Z digest=sha256:b524f221e50cfb656db2ad39da1bcfcd12db92cc53b07a3db6a86ce3c39020d5

Observation 2ec41a30-ebf6-429c-b8b5-57a62ecd6fc5 · inbound

InfAlign: Inference-aware language model alignment cites this paper.

InfAlign: Inference-aware language model alignment RL with KL penalties is better viewed as Bayesian inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:57:54.300472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:57:54.300472Z digest=sha256:7c6260c058005d4b10dfd0ba85ea160e812877e26038d09d2cd3013e47cf43af

Observation 70857bde-aac4-4a66-8db7-21e81c66403f · inbound

A General Framework for Inference-time Scaling and Steering of Diffusion Models cites this paper.

A General Framework for Inference-time Scaling and Steering of Diffusion Models RL with KL penalties is better viewed as Bayesian inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:43.817398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:56:43.817398Z digest=sha256:028ef32b59494f2531466a0cfe129f9818133223ec74fd80cbf99b3d857f46c7

Observation 51d2c4b0-c08c-4c05-84c8-1ee3ce9edd31 · inbound

CoDe: Blockwise Control for Denoising Diffusion Models cites this paper.

CoDe: Blockwise Control for Denoising Diffusion Models RL with KL penalties is better viewed as Bayesian inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T17:10:17.082630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:10:17.082630Z digest=sha256:5f72fc74d3c2b0ef0d8cb1569d2fe902e22263990fed4c231e0f75bbecd04c99

Observation a9c9f8f1-eae8-40c6-9d6d-824d3087514b · inbound

Addressing Concept Mislabeling in Concept Bottleneck Models Through Preference Optimization cites this paper.

Addressing Concept Mislabeling in Concept Bottleneck Models Through Preference Optimization RL with KL penalties is better viewed as Bayesian inference

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-16T10:32:52.272951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:32:52.272951Z digest=sha256:65b068a2bc4175b9d78be9d926f2d61d1f14c7806b8a471cef25f5bc98432739

Observation 876999fe-af79-4977-ba31-ae9939d891db · inbound

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training cites this paper.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training RL with KL penalties is better viewed as Bayesian inference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:27.830463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:27.830463Z digest=sha256:a7e155987525154fecb5e20a134e90556f2afae12e139c759d26b6add003b943

Observation 5f055b0c-aa79-4eea-ab72-558e90452701 · inbound

Discovering Algorithms with Computational Language Processing cites this paper.

Discovering Algorithms with Computational Language Processing RL with KL penalties is better viewed as Bayesian inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:01.136484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:01.136484Z digest=sha256:23d0997f09ae4cefb1b1713db1eace6941cae40f67b52781c798c8b7ccd60f28

Observation 93333f7d-8616-41bb-b553-deae042a3f7d · inbound

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis cites this paper.

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis RL with KL penalties is better viewed as Bayesian inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:28:56.065982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:28:56.065982Z digest=sha256:a9d42b52222f240cd9aa3fbbed20fee10d6d463ead6d2fc9570f198f45081674

Observation 3aa4ea9d-4aa4-4680-b70c-b827d8e4e9b5 · inbound

A Survey of Reinforcement Learning For Economics cites this paper.

A Survey of Reinforcement Learning For Economics RL with KL penalties is better viewed as Bayesian inference

Reference 1960

Resolution
unresolved
no resolver link, observed 2026-08-02T18:34:52.804048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:34:52.804048Z digest=sha256:9cce527889712033fcfe9118a559abb72b7fa67b8efc68eed3c26c4c8a2cec14

Observation 8ea9c9d4-48f6-481e-a9dd-79f0ddbaec96 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow RL with KL penalties is better viewed as Bayesian inference

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.619008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:012cdc25ca47001b281e0dc9f8198febc3c6d7af9de9699c3d904096d4c4216b

Observation 4abcc714-8349-4261-8f2f-1f1c616a831f · inbound

Exponential families from a single KL identity cites this paper.

Exponential families from a single KL identity RL with KL penalties is better viewed as Bayesian inference

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.247915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-07T07:16:28.127796Z digest=sha256:7e1ac14f37c9900d8342671b5e3889bda0cdc5d9b59c299ac83fa09fe220f189

Observation ab0a2075-e8f6-41e5-8d5e-75bfff6214c4 · inbound

Binary Rewards and Reinforcement Learning: Fundamental Challenges cites this paper.

Binary Rewards and Reinforcement Learning: Fundamental Challenges RL with KL penalties is better viewed as Bayesian inference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:08.663028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T16:29:16.359577Z digest=sha256:abf08e9194e8d08f7a4e8e984a0d936f1ec253ed87ee12e1fecbda12c124ff4f

Observation c4ed330b-c778-45c0-8bda-697343dad096 · inbound

On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective cites this paper.

On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective RL with KL penalties is better viewed as Bayesian inference

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:36:26.798482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T00:52:51.509014Z digest=sha256:dbecf411b9269a25a55f768c9cb1aebe3f40aef569cd7250132c584f93e214b5

Observation 8ac82a3c-dbd6-4e2f-8485-cfeba45aae1e · inbound

The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives cites this paper.

The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives RL with KL penalties is better viewed as Bayesian inference

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:52:08.903060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T02:49:53.547490Z digest=sha256:7fb559d82f7312f407e6c3b2c2a1658930a58b26c23d0594b7b2003b2760b8d7

Observation db64787e-ea9d-4737-881f-b2258609c6e5 · inbound

A Unifying Lens on Reward Uncertainty in RLHF cites this paper.

A Unifying Lens on Reward Uncertainty in RLHF RL with KL penalties is better viewed as Bayesian inference

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.846581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T17:11:46.549150Z digest=sha256:30480bfede484513f9d9986b6f37f097df687b9522ccf3a31b840f1b97c46971

Observation dce67dbc-9dce-420e-8da5-547482ef1730 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay RL with KL penalties is better viewed as Bayesian inference

Reference 247

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:be3776d4003ec416b64a5fc21ace65f3e2d9f100eeace57cb26b5b8dc4c67080

Observation bf3d459f-d00b-45dc-80be-648e96eb1dc8 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay RL with KL penalties is better viewed as Bayesian inference

Reference 248

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:01.139311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:01.139311Z digest=sha256:ef73b10cc6e112d356460b53fbf7f7808228e0b40450a09e05f216302c93c64f

Observation 8ba7f66a-6485-4d5e-a4f5-ce4d07460d7d · inbound

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration cites this paper.

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration RL with KL penalties is better viewed as Bayesian inference

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T14:33:57.124460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:33:57.124460Z digest=sha256:b59907a8762eba68547522d23b33cea8208914fd7b522ad1c1253a360a885383