Pith. sign in

Paper Citation Record · LEDGER

Reward Models in Deep Reinforcement Learning: A Survey

As of 21 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 6 inbound Pith citation observations for arXiv:2506.15421.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15421 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:38:51.086103Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:08:19.797150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:29:50.941267Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact3
  • verified fuzzy2
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49a2ba99-8f48-418d-a2b3-2d472e72612c · outbound

This paper cites Vision-Language Models as a Source of Rewards.

Reward Models in Deep Reinforcement Learning: A Survey Vision-Language Models as a Source of Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.445678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.445678Z digest=sha256:2e55bed080d3b37dec460cc405521ff733cc1d9fa54a0ad0e198e3dabdebdd7a

Observation 4b81b121-5b55-402d-b1e7-7a72d76bf2d1 · outbound

This paper cites Diversity is All You Need: Learning Skills without a Reward Function.

Reward Models in Deep Reinforcement Learning: A Survey Diversity is All You Need: Learning Skills without a Reward Function

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.461416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.461416Z digest=sha256:c769bf4700e4428aa38a7a639763190d10e3f9e32c1e1e7e9356246050a06f3a

Observation f002d8c3-1e2c-423d-8e3e-a11ad8568f40 · outbound

This paper cites Quantifying Differences in Reward Functions.

Reward Models in Deep Reinforcement Learning: A Survey Quantifying Differences in Reward Functions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.478051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.478051Z digest=sha256:19b2762ad7751058633ddd10a4050ba8b445c779199bcd8333d5a0b8c3e5cfe9

Observation 0bff0728-f114-4b9d-8edc-c33ad3690e8d · outbound

This paper cites Preprocessing Reward Functions for Interpretability.

Reward Models in Deep Reinforcement Learning: A Survey Preprocessing Reward Functions for Interpretability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.544516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.544516Z digest=sha256:a8e954db73c12de0f85cb772b8e7266041a40facf4da02ed2cca613b6b87ff5b

Observation f4a188ef-28ba-45b9-b6d5-9f3237367050 · outbound

This paper cites Regularized Inverse Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Regularized Inverse Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.578945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.578945Z digest=sha256:baff6c608a72c0e71027fe38f7ad25e96e639d1c98e9d56a9348e667d67e59f0

Observation 9270cc22-8d05-44ab-bba5-0b97572f10d2 · outbound

This paper cites A survey of reinforcement learning from human feedback.arXiv preprint arXiv:2312.14925, 10,.

Reward Models in Deep Reinforcement Learning: A Survey A survey of reinforcement learning from human feedback.arXiv preprint arXiv:2312.14925, 10,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.627819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.627819Z digest=sha256:b596b72d8fc0326914f753fbb1f85259caa3b9d2b7098f0b8bb1d5c527723dd0

Observation d6570f42-3b80-4095-b1cc-01223c13b89b · outbound

This paper cites Preference Transformer: Modeling Human Preferences using Transformers for RL.

Reward Models in Deep Reinforcement Learning: A Survey Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.639257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.639257Z digest=sha256:bbe15f0302b96ec92be9aecf841c2590e7f8683099e240d3d249d9c827b692d3

Observation dc08f9b9-f160-4338-a95e-482f090b333c · outbound

This paper cites Empowerment: A universal agent-centric measure of control.

Reward Models in Deep Reinforcement Learning: A Survey Empowerment: A universal agent-centric measure of control

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:38:52.185603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:38:50.650866Z digest=sha256:5cf7e0e5573bbeeaf30a845d06a31d01998f685bce31acb21baec37040c28ba5

Observation fb6b9eb2-7d20-4713-b7cb-c33da20962a7 · outbound

This paper cites Goal-Conditioned Reinforcement Learning: Problems and Solutions.

Reward Models in Deep Reinforcement Learning: A Survey Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.662293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.662293Z digest=sha256:5f5187d8ca26366f85e45fa5c53543ac058799ea0b478de3710e309a4be1499a

Observation 5144f0ef-b109-4920-a17b-a2d97aab60bc · outbound

This paper cites Reward Modeling with Ordinal Feedback: Wisdom of the Crowd.

Reward Models in Deep Reinforcement Learning: A Survey Reward Modeling with Ordinal Feedback: Wisdom of the Crowd

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:38:51.656840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:38:50.668075Z digest=sha256:4f677bb9a314fb4727cf14bdbeb7b08a2c3d36a156560a5ce2e8ee723294e09c

Observation a6529e8c-1975-4c4e-86d9-64d2cd39f9a1 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Reward Models in Deep Reinforcement Learning: A Survey ReFT: Reasoning with Reinforced Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.679115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.679115Z digest=sha256:98476f9cf6b781b1a784aa19ab117198363ee2c367fb83fc1b89c9b414300e5c

Observation 4795901f-7eab-45bf-ab6e-6c5ea4aec581 · outbound

This paper cites Choreographer: Learning and Adapting Skills in Imagination.

Reward Models in Deep Reinforcement Learning: A Survey Choreographer: Learning and Adapting Skills in Imagination

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.698388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.698388Z digest=sha256:bb7c0454de9cbf9dd87893d5fa2cc3f2110dc19c50031c2b0e56f1571afdafb7

Observation 61063ea3-cb44-4b2e-9e8c-769bdd690a8a · outbound

This paper cites Learning to Assist Humans without Inferring Rewards.

Reward Models in Deep Reinforcement Learning: A Survey Learning to Assist Humans without Inferring Rewards

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.792909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.792909Z digest=sha256:c9a2e19ddf0389706933252c84555a615882b95cdbf8a602f22112290ffb4312

Observation 598ec36a-b1d1-4312-86f5-6eeb6144bcfb · outbound

This paper cites METRA: Scalable Unsupervised RL with Metric-Aware Abstraction.

Reward Models in Deep Reinforcement Learning: A Survey METRA: Scalable Unsupervised RL with Metric-Aware Abstraction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.841430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.841430Z digest=sha256:d92ca1189d2a5f5249581ea2b1b603434775d1fc3ac8e16b87a512152aa8a6dd

Observation b9c43ce1-7ec8-434d-99cf-8d878f4a37cb · outbound

This paper cites Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.847290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.847290Z digest=sha256:910302f5150a9cafb7ae333b988f7e163b82797b30a79610f50dbaedffd9b0df

Observation 26e7ff2b-2682-45f3-8248-ac6aaa0779a6 · outbound

This paper cites STARC: A General Framework For Quantifying Differences Between Reward Functions.

Reward Models in Deep Reinforcement Learning: A Survey STARC: A General Framework For Quantifying Differences Between Reward Functions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.852504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.852504Z digest=sha256:e6e32063bfd04f6056862d17e50e5d8e1517a0461542a31f05e46a60af5efca8

Observation fd93ee8b-7900-4168-838c-4bfeb2a8f248 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reward Models in Deep Reinforcement Learning: A Survey Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.858219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.858219Z digest=sha256:dbe0b364ba061e4d26e84cbd4117283488ee03f51d51100479dabe7025ec8c2d

Observation b80eeca1-34af-449e-9d5f-de88fe30a17a · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Reward Models in Deep Reinforcement Learning: A Survey Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.863987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.863987Z digest=sha256:11f2834c3fea0df334d50ab5fc9720ac8d2846e3279ecc26fc1bcaa7ee522b2e

Observation 199389b0-881b-4c00-9e30-6e75e105ec4a · outbound

This paper cites Hindsight PRIORs for Reward Learning from Human Preferences.

Reward Models in Deep Reinforcement Learning: A Survey Hindsight PRIORs for Reward Learning from Human Preferences

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:38:51.407764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:38:50.869505Z digest=sha256:1f732fa32dd740fa068b62026d530257f820c26aaaf083b1aae4858e96d27902

Observation bea44e5a-ebf5-4d1a-a72c-bffd2ba218a8 · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:51.030036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:51.030036Z digest=sha256:3735774ca20d8e38033bc01e254ea67aee90d7374d9b6bf4b05237b27c1d7787

Observation 81ce2191-54d4-409a-8c40-53578b8eaa83 · outbound

This paper cites DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search.

Reward Models in Deep Reinforcement Learning: A Survey DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:51.075704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:51.075704Z digest=sha256:4d17fd3da4ab62f5e8e365b80b47a5c599e85869fa657b7acdf3463f274a08e2

Observation 9dbf7c04-eef8-4c1e-83be-43a1cee1455f · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Reward Models in Deep Reinforcement Learning: A Survey Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:51.080835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:51.080835Z digest=sha256:1d0b591a3b9c4cf01de7d75dd63f5a347caed5ed4fce4873cc1eb22e46a76382

Observation 82014311-ddbe-4e41-8db7-7b212b576b57 · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Reward Models in Deep Reinforcement Learning: A Survey Maximum entropy inverse reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:38:52.167651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:38:51.086103Z digest=sha256:12305908b505b586ca5067aac4b473d51695451032d41c81b89d7c98f9df9365

Observation b88bbc5f-199d-4b32-961a-8cb31b19f2c7 · outbound

This paper cites Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery.

Reward Models in Deep Reinforcement Learning: A Survey Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery

Reference 1950

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.489300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.489300Z digest=sha256:244ed455953888bbebbdc91a4a1472ff2aa13e9c60e5e9abbb9d0353831755b5

Observation e4c7d4f9-2d29-479c-9d14-f6b860535c79 · outbound

This paper cites Exploration by Random Network Distillation.

Reward Models in Deep Reinforcement Learning: A Survey Exploration by Random Network Distillation

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.450710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.450710Z digest=sha256:e5652f80718593a425af4b0a92287c37e918f71e2720e3645e0ec6d398e79442

Observation ebb72092-549a-4dec-84b7-b23e74c8330b · outbound

This paper cites Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.410219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.410219Z digest=sha256:53f99b35198af59ba3f535c38bb4fcc7aafe09a36f44f0d30449cf85f2ca7576

Observation baff5cc1-6d0e-4bf8-adc7-3bcd9acfebef · outbound

This paper cites Models of human preference for learning reward functions.

Reward Models in Deep Reinforcement Learning: A Survey Models of human preference for learning reward functions

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.656340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.656340Z digest=sha256:c66f660f17c2bb9d0909e42f3ef792197416400791a90d83acc9f6347cf319bd

Observation d8b3e62f-02a8-49eb-b66c-a8dc62d1b0ee · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.483634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.483634Z digest=sha256:04df13f23de2bd8e48fe443126d0c95acc89e04ed5ad91c876f01dbeea239344

Observation 875992f6-9b86-4dba-a3c3-852304b4c120 · outbound

This paper cites Learning Robust Rewards with Adversarial Inverse Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Learning Robust Rewards with Adversarial Inverse Reinforcement Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.473011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.473011Z digest=sha256:c1f0c80de541e63cd2760501612b7c3483c0f36707b7a5ae205399157398c1c8

Observation d5b5aa4f-bf1a-44fe-8b45-54f78a626e19 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.456084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.456084Z digest=sha256:df7bee70745fdd509e132a7d89f60aa210ba452b8566264f79221352dfd89c58

Observation 53580f90-e529-44f7-a411-577256bbb502 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Constitutional AI: Harmlessness from AI Feedback

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.439486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.439486Z digest=sha256:1b696932ef89ce4953ab167224fa34347ee23a17ffd9c9bec75be8ca758c8bfc

Observation 64e3677c-cd60-4840-8c23-e9dacf260f04 · outbound

This paper cites Never Give Up: Learning Directed Exploration Strategies.

Reward Models in Deep Reinforcement Learning: A Survey Never Give Up: Learning Directed Exploration Strategies

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.433306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.433306Z digest=sha256:adfcc1c3b2704f52fa5bea50b04b714fc96cd16e189f8bc08684106db8fa1617

Observation 1b420d97-7414-4ab5-a1da-b8a05c7f42bc · outbound

This paper cites A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models.

Reward Models in Deep Reinforcement Learning: A Survey A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.467651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.467651Z digest=sha256:bb11e9af96a1a3c645eb687568a84056280839ced9eb3c2575ff16b8db6b93cb

Observation 1be62d4a-6672-4734-b03a-985e95e42ca8 · outbound

This paper cites Motif: Intrinsic Motivation from Artificial Intelligence Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Motif: Intrinsic Motivation from Artificial Intelligence Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.645727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.645727Z digest=sha256:670f48f2246eb89817ca0f1f1d1084a8abe11e603b65d4a6f5b5b7815da71f9c

Observation 2fe16626-45d2-4e09-a520-a493d26b41fe · outbound

This paper cites LiPO: Listwise Preference Optimization through Learning-to-Rank.

Reward Models in Deep Reinforcement Learning: A Survey LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.674235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.674235Z digest=sha256:9e5f1a2a78bb3fa9f25e9e8dc72cceb8fdeb6a10a2cc52f05b1e7365332df054

Observation 5f0cf442-4de2-42dd-a4a1-b515064d66f3 · outbound

This paper cites Dynamics-Aware Comparison of Learned Reward Functions.

Reward Models in Deep Reinforcement Learning: A Survey Dynamics-Aware Comparison of Learned Reward Functions

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:38:51.286313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T19:38:50.959942Z digest=sha256:6f7fc71f51b47361ca1ab0faab90ae93bcd15e69986a3759c2c9ca1c6a1ce93d

Pith citing papers

Observation 9c30b6bc-215c-4a92-8b91-d9a93b451a39 · inbound

Multi-objective Reinforcement Learning With Augmented States Requires Rewards After Deployment cites this paper.

Multi-objective Reinforcement Learning With Augmented States Requires Rewards After Deployment Reward Models in Deep Reinforcement Learning: A Survey

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:41.945064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T09:38:05.280629Z digest=sha256:f1e2dfac5f13cf07b0419fc1a11dadbe729a822321a6ffd0715f7b9749d765be

Observation e53582ca-beba-4ed5-90ca-36d1efeb6910 · inbound

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning cites this paper.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Reward Models in Deep Reinforcement Learning: A Survey

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.223690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:19e0d1eebece779e6e25a5462c0eee1d0909a969bf60af7de771d083498096c9

Observation 79cea2e5-be26-43e8-9720-51304d27cddc · inbound

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models cites this paper.

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models Reward Models in Deep Reinforcement Learning: A Survey

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:47:53.425028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T19:45:26.458859Z digest=sha256:7927c9c521729527b7836eefa953bdc386da11e66401464ad117e9ebfdc53d60

Observation ee58a84b-58e0-4b21-80ad-9659fcf303f4 · inbound

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models cites this paper.

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models Reward Models in Deep Reinforcement Learning: A Survey

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:06.483419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T06:04:17.574622Z digest=sha256:9762b98a479ab595f1592ddee7181e87bb86133181c9779e262e8038efa982d4

Observation 6c8a3fc2-9855-4232-b2f2-9bf323ba31e0 · inbound

AudioProcessBench: Benchmark for Identifying Process Errors in Audio-Grounded Reasoning cites this paper.

AudioProcessBench: Benchmark for Identifying Process Errors in Audio-Grounded Reasoning Reward Models in Deep Reinforcement Learning: A Survey

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:37:27.332423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T18:08:19.797150Z digest=sha256:24cee4432420afe52e8e06b1b7cacadf5b5a2ea29cba654dfc67ace477482f49

Observation ad523e39-1735-4f45-8802-1b24dff731b1 · inbound

PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation cites this paper.

PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation Reward Models in Deep Reinforcement Learning: A Survey

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:50.942774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T05:14:14.053344Z digest=sha256:0d70a03072e27e1e709fd73f2e02f4d99c1fc09c4884a03bd9d2c11e0d6001de