Pith. sign in

Paper Citation Record · LEDGER

Understanding Learned Reward Functions

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2012.05862.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.05862 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:32:44.291764Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T05:56:01.759090Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e96acce3-4665-415c-9753-e3a6fbe10e02 · inbound

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment cites this paper.

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Understanding Learned Reward Functions

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:46:56.829477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T00:46:56.664582Z digest=sha256:4b0045f6dbed4772f7115fceec0ebe38a72ef6d857c37b3b46e06d7a060ed0c3

Observation 8cb72ddc-32ef-4ce5-8ef5-6d952a873e5c · inbound

Active teacher selection for reward learning cites this paper.

Active teacher selection for reward learning Understanding Learned Reward Functions

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:56:01.762123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T05:54:43.873174Z digest=sha256:90db3771bd22def371e9a49a283bc032959e016f019121ff112b47c7f35e5595

Observation 80e47365-ef4f-4799-a910-e53086418e9a · inbound

Improve the Training Efficiency of DRL for Wireless Communication Resource Allocation: The Role of Generative Diffusion Models cites this paper.

Improve the Training Efficiency of DRL for Wireless Communication Resource Allocation: The Role of Generative Diffusion Models Understanding Learned Reward Functions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:32:44.291764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:32:44.291764Z digest=sha256:753a7f0ca3f4d150f4e81fc77c1c12d9b68de9c72d9a15283382b9a26bfb1e02

Observation c6a1849f-7f15-4fbf-93f7-6ee490ec18ea · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Understanding Learned Reward Functions

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.277433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:511694c5054336eee34be42301b4d41cef5aad273a5ce360d501768ab442db55

Observation 4b3466da-ceb1-46c1-be74-f78d98521990 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Understanding Learned Reward Functions

Reference 174

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:f41562f2485e26a3cbcf8d5406de5a532d73cbb3cbf08a66305479822731c23b

Observation befe4177-5313-4b75-bd0f-69694ed8f6a9 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Understanding Learned Reward Functions

Reference 175

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:52.536159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:52.536159Z digest=sha256:813c5ecb6b2675dc335dd588ccd4f4c2ac776afa027b7ad58226db61d06c3930