Pith. sign in

Paper Citation Record · LEDGER

Probabilistic Uncertain Reward Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2503.22480.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.22480 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:30:32.243421Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.146385Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 05a0c8d6-973a-41d6-9df9-e5c73c95ca7f · inbound

Variance-aware Reward Modeling with Anchor Guidance cites this paper.

Variance-aware Reward Modeling with Anchor Guidance Probabilistic Uncertain Reward Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:12:17.514376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T05:11:05.546092Z digest=sha256:d159da4aa1386f7c425d2ff07108ef30215e436dfb5d7d95908de6a1b1b72aa0

Observation 211795d3-9128-45df-a8d4-2c412bc6daa8 · inbound

A Unifying Lens on Reward Uncertainty in RLHF cites this paper.

A Unifying Lens on Reward Uncertainty in RLHF Probabilistic Uncertain Reward Model

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:29.835034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:11:46.549150Z digest=sha256:f99c191e7ed1e561418adb3440b6aec966d193828ef0c81a7821bb7e8fd6f9d0

Observation 2f649fe5-fadc-4ca0-b36b-0c601a011908 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Probabilistic Uncertain Reward Model

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.147885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:84bda50e65465112b72894baf1fdf24494a5a56a15f674b2856769cf0d0e9a4b

Observation f7ab9048-6927-4354-97c3-af51283e1d62 · inbound

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training cites this paper.

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training Probabilistic Uncertain Reward Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:32.243421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:30:32.243421Z digest=sha256:2724f3956fc51dd6e7f47d7a590786de5789607e78113af8992e5517e8e9ae30