Pith. sign in

Paper Citation Record · LEDGER

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient

As of 5 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 1 inbound Pith citation observation for arXiv:2604.25872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.25872 v1

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T16:24:43.688967Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 100 outbound references displayed

  • verified exact27
  • verified fuzzy69
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a34d8fc-b13d-43e3-ac7f-067bd024a5c4 · outbound

This paper cites On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.863945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:9136d416f1dead81a8262b5036c0c2af0f27b3e9d6b7b22d4526a4b7f6cfad98

Observation 25482215-e86d-4465-b6c4-70102db495c1 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:22.769824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:207a2c6f7c4137f192190fa0ca605adc58d0f810799b18b0b76b371e8c5b8fd5

Observation bd5e698c-3480-452c-9ca7-4047323ee798 · outbound

This paper cites Understanding the impact of entropy on policy optimization.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Understanding the impact of entropy on policy optimization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.904768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:ef12302a508a2aa5c4008f36bf03f588e723697bd5373e43348a82babc7e511f

Observation 0d54c6f2-c896-41a5-ac7e-88e78269e706 · outbound

This paper cites Concrete Problems in AI Safety.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Concrete Problems in AI Safety

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.510706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:e8e5191b3281ae1953b16c6d465066d715d89eecb5f59cbf77de6eaed4e9154e

Observation 246d52d9-a8e3-4c28-b4bb-2de9932b416f · outbound

This paper cites Potential-based shaping in model-based reinforce- ment learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Potential-based shaping in model-based reinforce- ment learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.984806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6068e7396eec30cc062a10dccfb245a40614adc16c3e3633cdd3d9cbd9ca079e

Observation b373cf31-2bdd-4a9f-9bfa-2a15ebe38687 · outbound

This paper cites InfAlign: Inference-aware language model alignment.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient InfAlign: Inference-aware language model alignment

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.360622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:ede2d45a5a35f5a87255eb55d3b1c0dea207c41d62275dd3b2c970af2eec80ce

Observation 053b26ba-67ef-4828-a9e4-f1c62adc4f01 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Dota 2 with Large Scale Deep Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T22:18:16.018301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:7483c0c5a7af5478ffd7622603fe94b0aaaa8fa66f7ff2c7511bf38e2d4c3e72

Observation 51e7cdf1-4e82-42b6-803e-d770919c1c2b · outbound

This paper cites an unresolved cited work.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-27T02:18:25.100242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:2a7c6e40ac06d8658b8d32c88e999eb2a4e7bd8b7bf732bf76a837cbed99fbb2

Observation b77832eb-85f9-4636-a428-0ba3f8298037 · outbound

This paper cites The accuracy paradox in rlhf: When better reward models don’t yield better language models.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The accuracy paradox in rlhf: When better reward models don’t yield better language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.072863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:a5b6b2a978f2dc3136cc9c035b1f05f813abdb111ff11bab67a3aafeaa93b3a3

Observation b0803899-d69e-443e-b11f-e67df679bff4 · outbound

This paper cites Heuristic-guided reinforcement learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Heuristic-guided reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.859056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:ec6227000896e67129d5679ee31228124607e90af6a929bd45f5f8199f2c20b7

Observation d4899a4c-28d0-444c-a2f5-3d47850b65af · outbound

This paper cites Learning navigation behaviors end-to-end with autorl.IEEE Robotics and Automation Letters, 4(2):2007–2014.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Learning navigation behaviors end-to-end with autorl.IEEE Robotics and Automation Letters, 4(2):2007–2014

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.013493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:65ca1767182cefd2e6741b3236be994217081794b975dc235943e4db4ae2f5b1

Observation 2df8d76b-484d-47ff-a1e4-6730ad9ade43 · outbound

This paper cites More is less: inducing sparsity via overparameteriza- tion.Information and Inference: A Journal of the IMA, 12(3):1437–1460.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient More is less: inducing sparsity via overparameteriza- tion.Information and Inference: A Journal of the IMA, 12(3):1437–1460

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.175394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:4854d718ee36bccaa25feea6eaad91b07a47255a862ebbc6c942a877e7649aeb

Observation 0bafe182-0e35-4c0b-bd22-7149147b4048 · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward model ensembles help mitigate overoptimization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.068336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:445bfea6080562149a273f9ba6af1fa0b57380cd6f5847fb34d254df4d091bd5

Observation 001ce548-e2a0-4a8a-b017-38e03f8407c4 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Ultrafeedback: Boosting language models with high-quality feedback

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.141787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:b08a60043719a24a39693dd7131a26952c445792acdd686c56e99acf68674cff

Observation b6bba76d-fb51-4eb3-8356-1311f0257689 · outbound

This paper cites Maximum expected hitting cost of a markov decision process and informativeness of rewards.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Maximum expected hitting cost of a markov decision process and informativeness of rewards.Advances in Neural Information Processing Systems

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.885611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:c58042f5a09ce474afe29999e1cf0d084f3f2ab3639066b2f40e011902656e1a

Observation 3195c83e-82e7-4a9b-983a-e246baccc8b9 · outbound

This paper cites Exploration-guided reward shaping for reinforcement learning under sparse rewards.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Exploration-guided reward shaping for reinforcement learning under sparse rewards.Advances in Neural Information Processing Systems

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.838961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6c848f87251a7e4fe05d27fcc659c9c6034b2f9a52d18cd088753d7b4ae8546f

Observation 61d4bf89-748f-4f6b-87dc-723058a0a7dd · outbound

This paper cites Continuous vs.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Continuous vs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.868010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:b7268cbe1432ceabb9500f8e58ce42d64451ac89051bfd383a11de7c72df1d7a

Observation 77b8edde-0e2a-42b1-b3da-fffbe35e77f9 · outbound

This paper cites The perils of optimizing learned reward functions: Low training error does not guarantee low regret.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The perils of optimizing learned reward functions: Low training error does not guarantee low regret

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.895553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:efa16b0adf1c70832b8b48c182ee22416fd7440d15861abb9aeb9c2c00947a5e

Observation e2304c86-3d52-4ac1-9fc0-d76092cecfea · outbound

This paper cites Is a good foundation necessary for efficient reinforcement learning? the computational role of the base model in exploration.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Is a good foundation necessary for efficient reinforcement learning? the computational role of the base model in exploration

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.786800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:d65975a4ee90f876fc299fac60d77600f737c8abe22905195c7ea734df6b5996

Observation d14e7f16-b218-42ac-aab0-be81f3b28c37 · outbound

This paper cites How to evaluate reward models for rlhf.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient How to evaluate reward models for rlhf

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.123615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:98a54914e0cd63352376f33d71ce9b1d3a3e52b5f9e3dd28dd431131adcfb1ab

Observation 39e03859-ef9c-4662-8d30-f9a6d8735a40 · outbound

This paper cites Scaling laws for reward model overoptimization.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Scaling laws for reward model overoptimization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.805765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:42b7a6eb78c131293b96f118f58f14c6f8b43bd5df1fc08057fdc6d314589fde

Observation bacdaebb-550e-4a6d-8693-a5602a9b3cf6 · outbound

This paper cites An alternate policy gradient estimator for softmax policies.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient An alternate policy gradient estimator for softmax policies

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.018960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:06e4e60abc65808254f0daa9cc79b4ed106e695a0fb0ed606c330cedc2c1c61d

Observation 00c18eea-a1cd-4f85-95c7-8248d629d8c7 · outbound

This paper cites Quantifying differences in reward functions.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Quantifying differences in reward functions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.828078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:bf58cd9319d28f4da07b03e8391197b2553838c8e53229dd277839a88e03d8d3

Observation 5ff05225-ae09-41cd-b833-66e2010ba483 · outbound

This paper cites The Llama 3 Herd of Models.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The Llama 3 Herd of Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.354997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:0c9c781d9a5d120ceb240b9cd852bd64c27db2aeee70a8210b74bd6f370a3c3e

Observation cd9ba846-5e88-440f-81d5-a62981433e19 · outbound

This paper cites Reward shaping in episodic reinforcement learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward shaping in episodic reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.109371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:7c9c775d4162c8bc7635639644d29208e37ccb039a0312625759c3ec4b706b5a

Observation 0cee9084-8afc-4ecf-b712-e6f53d50705b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.569181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:400de57c18afb0028250cb5c822541cf99cb0d4bf8b14d7b5bfe7c3f26ec57e1

Observation b2b53894-57f5-4717-aae4-3f45bd19d9af · outbound

This paper cites Unpacking reward shaping: Understanding the benefits of reward engineering on sample complexity.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Unpacking reward shaping: Understanding the benefits of reward engineering on sample complexity.Advances in Neural Information Processing Systems

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.795897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:f41fc6e2bfdf4c03e1febb7f674cd3e3da4ea4ce1194f1372f97e09755ff46fd

Observation e8bd2e1e-9a1d-4d8e-bf3e-e7939a5ca7fe · outbound

This paper cites Neural Replicator Dynamics.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Neural Replicator Dynamics

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.370655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:3301fd45b5afab46bd5712a614c3d491e4e05ea31da29c34bce2e05c85ec739a

Observation a9db3d6b-33f0-4820-b168-cb647a65ce88 · outbound

This paper cites Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.843538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:46139055613696142f25788add4b8599bc5a30c3fb1d5b0bc59553eaeebfa5c3

Observation 2e768850-5d4d-4dc9-a3bf-9d122f133052 · outbound

This paper cites On the Emergence of Implicit Curriculum in RLVR Learning Dynamics.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the Emergence of Implicit Curriculum in RLVR Learning Dynamics

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.411893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:f71d1011666411bc87c40647383bf157abf2099a87bc55ee89c42f71bd3d5356

Observation f4292780-bbe9-47f2-8cb6-4447e8fb1132 · outbound

This paper cites Pitfalls of rule- and model-based verifiers–a case study on mathematical reasoning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Pitfalls of rule- and model-based verifiers–a case study on mathematical reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.516586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:5a667e203d9ccf3b096796e455bab2c0b71734f17a1a235f6bba6c53b6d4b7f3

Observation 79356e28-dcfb-46ca-b1ba-7c2e125cd99e · outbound

This paper cites Goodhart’s law in reinforcement learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Goodhart’s law in reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.119332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:b003da069a518f53866b8e874b0e600bcfa36ebb34ab6bb35d890379f290c099

Observation 9ee1562b-2bad-4ee7-994d-05965bcfd311 · outbound

This paper cites Beyond stationarity: Convergence analysis of stochastic softmax policy gradient methods.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Beyond stationarity: Convergence analysis of stochastic softmax policy gradient methods

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.890286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:5b2a5d805a1401466e896bccb7a7869a9608c4461b406db8593129f794314688

Observation 7b1c7b91-3a83-4ce4-819c-d1c6b11ad14f · outbound

This paper cites Buy 4 reinforce samples, get a baseline for free!Deep Reinforcement Learning Meets Structured Prediction ICLR Workhsop.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Buy 4 reinforce samples, get a baseline for free!Deep Reinforcement Learning Meets Structured Prediction ICLR Workhsop

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.800342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:163c49f784d4573fd9288e96e9e310c7e8fd55cfd5cc07906e74942a9b55dbac

Observation c313ba31-c8b0-4040-8b8d-569a560a30e6 · outbound

This paper cites A neural collapse perspective on feature evolution in graph neural networks.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient A neural collapse perspective on feature evolution in graph neural networks.Advances in Neural Information Processing Systems

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.791469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:922b44b89863447d9387974b307691b47e4a1c8f1511085bc155c2ab20522bbf

Observation 9f74ff2c-b947-4f68-b9c1-2b78bd4fd77a · outbound

This paper cites Correlated proxies: A new definition and improved mitigation for reward hacking.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Correlated proxies: A new definition and improved mitigation for reward hacking

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.777821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:527b5b1a33e229b52ecabb238af4612d97afff827d8369342ae17ca96e627c2c

Observation b79c9aee-6548-486f-8dd0-87ec58fd2409 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.452354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:75b8fa0ad518e727de995acd474393f4cc176c5011bd3ddb9002f299d67f40cd

Observation 6b3e9138-9638-4f18-9945-7bf498b74769 · outbound

This paper cites Rewardbench: Evaluating reward models for language modeling.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rewardbench: Evaluating reward models for language modeling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.880935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:8982c55a237a3a8d200bd6349b12eafff0a19089b44d59829f39f01e6e8d584d

Observation a9957a1d-b670-4804-bde9-34413deea7cb · outbound

This paper cites The influence of reward on the speed of reinforcement learning: An analysis of shaping.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The influence of reward on the speed of reinforcement learning: An analysis of shaping

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.810202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:82c695726295b5bdfd34af5e63fc5657f38a93653ef1ac8cb1a77a4b8e515869

Observation c5b2414b-01f3-4776-957a-58931e1a29a7 · outbound

This paper cites Softmax policy gradient methods can take exponential time to converge.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Softmax policy gradient methods can take exponential time to converge

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.160033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:90ab338c0c16446e6ff2697e3407c281c0a01bff5f44f21dedb4f350122bdd79

Observation 3acb3b2f-9de6-41fa-b43c-54d351479c5a · outbound

This paper cites Hashimoto.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Hashimoto

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.170643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:b1865691c54c0d28a45b94263e7a26d3dfeb3e3047c346729c97a4c9d80fa990

Observation 054f8504-b2e6-4b03-a827-9ab741972c4c · outbound

This paper cites Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.689881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:86a6403b44b7bbae327ee05c3c5c981a048eea14a821092c0c3b642827078441

Observation 8e41e6ee-74ea-4de0-96cb-53143922b22e · outbound

This paper cites Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:17:43.404381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:75a37073edf14afde01b85e0109ac1cf65716ae06b243dce31526dccd2606c52

Observation 78f1a5c5-0921-403e-ae92-c241984b705b · outbound

This paper cites Elementary Analysis of Policy Gradient Methods.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Elementary Analysis of Policy Gradient Methods

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.581704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:296e5254ed1f7abdbb07fdede31dcf2de4ebbefaa65226dd3d5a7b93e7fcbf5c

Observation dcfc11c6-1ee1-4546-b85c-2d855129d2c7 · outbound

This paper cites RLTF: Reinforce- ment learning from unit test feedback.Transactions on Machine Learning Research.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient RLTF: Reinforce- ment learning from unit test feedback.Transactions on Machine Learning Research

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.150697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:7fb699ac718c6c2e112f9d1ddf404f59d26d2b9b945125f990fade6b18d44cd9

Observation 6c70cdc9-b1a0-4538-86f7-feb0e02f2e9c · outbound

This paper cites Rm-bench: Benchmarking reward mod- els of language models with subtlety and style.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rm-bench: Benchmarking reward mod- els of language models with subtlety and style

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.050469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:912c7552778a0d4486db6b0c471692a4375143f7d8b8e505005103ba15dafcfd

Observation d310456e-86a5-4dd7-a59d-3dad4320a502 · outbound

This paper cites RewardBench 2: Advancing Reward Model Evaluation.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient RewardBench 2: Advancing Reward Model Evaluation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.467386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:1b4f8f135992b82ee7e186119f23dd090627344344036ed59331bf0e543578c5

Observation da614b5e-a305-461e-a951-b4d3dc04a718 · outbound

This paper cites Reward engineering for reinforcement learning in software tasks.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward engineering for reinforcement learning in software tasks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.440658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:43918974a19a564d5f94ea8612d675b0d961e74ec3f0dc55b14eb6d0449d590c

Observation 0c30e2b3-8a65-4ef6-aad2-cc8f10b05a47 · outbound

This paper cites Reward functions for accelerated learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward functions for accelerated learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.059742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:4e0b6af937d6fe226099fb86341268f21e64fe23b2871aeeba3cf5593a71a792

Observation c95685b4-ae34-43b6-b82a-ae9695cbb204 · outbound

This paper cites Escaping the gravitational pull of softmax.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Escaping the gravitational pull of softmax.Advances in Neural Information Processing Systems

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.999383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:8c138b6004f8b8d6ce0cec746707af7dca85857512404970544b714885660acc

Observation 3d0dd837-3d61-420f-bbb6-f7eb95c65c0d · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the global convergence rates of softmax policy gradient methods

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.146068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:404d196b4e85c6db5c201e5f4d34709ea8a4d6cebb6aca47ada550f622627a0f

Observation 960fb550-3502-4e22-9e08-65740ce9b899 · outbound

This paper cites Leveraging non-uniformity in first-order non-convex optimization.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Leveraging non-uniformity in first-order non-convex optimization

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.782435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:e79c5a9cb8007b1f2a5b77ddd6965fb9a3812d2b0c72d47342ed6b148d225780

Observation 248c8549-954e-416d-8029-11395b13baf0 · outbound

This paper cites Ordering-based conditions for global convergence of policy gradient methods.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Ordering-based conditions for global convergence of policy gradient methods.Advances in Neural Information Processing Systems

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.872459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:1591609a165395ebe8a081169aaea11534c281339c78bd71494e79945a4ccb6d

Observation 33d3b4f0-5ba3-434f-ae45-e150ee40740d · outbound

This paper cites Stochastic gradient succeeds for bandits.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Stochastic gradient succeeds for bandits

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.823342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:ee4df68dee0a41eebaa9507dc20ce1c8f169f98373be07697000d2f87b7cc079

Observation 987f6b8a-5dfd-40a6-9e92-7c999a5a1dde · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Policy invariance under reward transformations: Theory and application to reward shaping

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.032695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:61088922aed28e865941a9716ff006d119c6065357a17309da60a62024155ac6

Observation e28f85ac-e154-4a89-a61c-b8fa25db28e3 · outbound

This paper cites 2 OLMo 2 Furious.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient 2 OLMo 2 Furious

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.431341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:beb69fc1852e9507edd1617e68b399d9844982cb3a6f6d1feba62a8908d30751

Observation a170f578-7634-4aa8-bc62-0b5865648f16 · outbound

This paper cites Olmo 3.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Olmo 3

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:41:19.524345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:864d2c61e7ee8b7688b4ac4e588056e8e5b764cac5b37aa3cd589b04ca32bd54

Observation 519a7f06-fdb2-4a40-8947-523801103c68 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.055485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:9aea9dc4cea56fbf923c732986c4d0bcf5c2483ddcd571f5b3308ef6f334f00e

Observation e0aef6b5-3da2-4eea-b445-f7fab22c60b0 · outbound

This paper cites Reward gaming in conditional text generation.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward gaming in conditional text generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.900348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:11ee04a7eec27aece6ae20471d500a53fae81747971deafa144f4f3d1f95c321

Observation 30751c67-af25-4434-8dad-9db33a339f83 · outbound

This paper cites Automatic differentiation in pytorch.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Automatic differentiation in pytorch

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.082466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:20317f643edadc6e4a74a71ae19959fff4844c01c4de66e87a49c55b1d3cf2ad

Observation 57dbbfe6-b229-48be-a261-b000c3d43132 · outbound

This paper cites Amp: Adversarial motion priors for stylized physics-based character control.ACM Transactions on Graphics (ToG), 40(4):1–20.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Amp: Adversarial motion priors for stylized physics-based character control.ACM Transactions on Graphics (ToG), 40(4):1–20

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.155802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:c085a815ec860787bb24583f78642b4b3dfbd76eb984f7f0800a9dff3f2cf2f3

Observation dd783c71-682e-4090-993b-c31105667067 · outbound

This paper cites Generalizing verifiable instruction following.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Generalizing verifiable instruction following

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.037042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:a6bb842601df1e9503812cc9864025d73a79ac669ccc25c695c933b5ccef57ff

Observation 097d2b6e-9754-4d08-8403-6e69c16330f1 · outbound

This paper cites Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-04T02:07:08.062014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:21d5ef36d869f50bcd126c14584b35fb9debf234f1824ed7bb2d369fd9fb40c5

Observation 1760c012-5233-4b15-9b96-c2cf885b0f2f · outbound

This paper cites Learning to drive a bicycle using reinforcement learning and shaping.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Learning to drive a bicycle using reinforcement learning and shaping

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.834343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6315f33d4c6c6868cccb1003118a291a853712ef2cf2b132ce4c34dba81505db

Observation ad7c4af3-c797-4a68-afb8-233e215bcba2 · outbound

This paper cites Implicit regularization in deep learning may not be explainable by norms.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in deep learning may not be explainable by norms

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.994339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:dfa24ea7496e1934e823242311a830a3b35f92d8d50be4b991bf3aa4930b48b2

Observation 10457d44-eb20-4523-a1c2-b0ab66501224 · outbound

This paper cites Implicit regularization in tensor factorization.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in tensor factorization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.091425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:26de940835ad83651c39a23107acf49182f5126fee0db80457ae3cb8f5f216ae

Observation efc34028-c14f-4e67-a47b-ed586a5b2417 · outbound

This paper cites Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.004683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:ba81b1e711e43e0939f88901f12fad11db4996b2a5f1c5ff470580853c9790fe

Observation 471d2902-ec56-4d25-a6a4-00650a010efd · outbound

This paper cites Susskind, and Etai Littwin.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Susskind, and Etai Littwin

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.028147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:23b4220738087d583019543cf6e8a7064794c8f5ab239d7c5cd612f58f4c9dc1

Observation 5b1a8b75-d679-4728-8d7b-6f986db2992e · outbound

This paper cites What makes a reward model a good teacher? an optimization perspective.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient What makes a reward model a good teacher? an optimization perspective

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.095922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:cdccf7bf4bd05c5c22363915bc0bc7cc14517ee17fded6b3806693da4e0e3765

Observation d585a040-940a-4d8f-be86-bd62a25510d7 · outbound

This paper cites Why is your language model a poor implicit reward model? InInternational Conference on Learning Representations.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Why is your language model a poor implicit reward model? InInternational Conference on Learning Representations

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.854294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:7d6d793e8b61fc569a337cb69a0e9c8496d5927b6cdc7427bb155c561fde0ac4

Observation cee29f65-6e2d-4f9f-aa72-c6d6a227f58c · outbound

This paper cites On the effective number of linear regions in shallow univariate relu networks: Convergence guarantees and implicit bias.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the effective number of linear regions in shallow univariate relu networks: Convergence guarantees and implicit bias.Advances in Neural Information Processing Systems

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.773208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:af2ac8a5983ac5412fac121cb3b1db2fc7703ac7ca398ca5092c933bab76ea2a

Observation 4b9d81ba-92ac-41be-82b4-16e2a588389a · outbound

This paper cites Exact solutions to the nonlinear dynamics of learning in deep linear neural networks.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Exact solutions to the nonlinear dynamics of learning in deep linear neural networks

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.848707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:69a3a22b5ce56e1216d28c286621dea6f48415c943521312f2ca6e3af5684fab

Observation 903a3577-8159-47d5-953b-5ea29d577059 · outbound

This paper cites Ray Interference: a Source of Plateaus in Deep Reinforcement Learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Ray Interference: a Source of Plateaus in Deep Reinforcement Learning

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.487655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6fb6503620f9402b5944890e9ee5d38ad521fedbcd80e28f33393eea729b9a36

Observation 16b258bc-2c08-45a3-bc04-c0cba41e2226 · outbound

This paper cites Proximal Policy Optimization Algorithms.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Proximal Policy Optimization Algorithms

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.387711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:fdf6edcc14fed2687907aa107c8bf61c8c7722b981ffdc6cf8e0b2c3d93e9a62

Observation 4d69d36c-3d2d-4992-a032-4116c8f57c0b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.624857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6ea7ff48b4693ed9064f5cb65e8d6921655a4be1c0d644f961c5dc6731b97bc0

Observation 3c3ce965-d6d4-4970-8910-bca64027c993 · outbound

This paper cites Intrinsically motivated reinforce- ment learning: An evolutionary perspective.IEEE Transactions on Autonomous Mental Development, 2 (2):70–82.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Intrinsically motivated reinforce- ment learning: An evolutionary perspective.IEEE Transactions on Autonomous Mental Development, 2 (2):70–82

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.132804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:dc384cdb79d532a88de2702215a920c24d64f0ac02695ca9bb4bbc17189175e7

Observation f18a4585-de3d-4add-b1dd-e5088a989381 · outbound

This paper cites Defining and characterizing reward gaming.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Defining and characterizing reward gaming.Advances in Neural Information Processing Systems

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.041345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:f782569aa80e1850292a4d9743ed5c14f85dd71b3ed1f7a2616e0d5277b4cc94

Observation 4685644c-8a2c-47f1-a102-03505ab7b8ff · outbound

This paper cites Starc: A general framework for quantifying differences between reward functions.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Starc: A general framework for quantifying differences between reward functions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.045761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:f4f07ca4746ed69a39e1299b945231b5c1e930dc2299374179d5c5e8f364d16d

Observation ff578be5-2523-4960-adf2-91febc47327c · outbound

This paper cites The implicit bias of structured state space models can be poisoned with clean labels.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The implicit bias of structured state space models can be poisoned with clean labels

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.814814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:7dbd998c392b0707b3386858288364734cb72c161f7a76ca84388a6ceb154d4f

Observation 8052fc45-4fd5-4da8-9a7f-298f85c40e55 · outbound

This paper cites Reward design via online gradient ascent.Advances in Neural Information Processing Systems, 23.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward design via online gradient ascent.Advances in Neural Information Processing Systems, 23

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.127911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:abeabf8e4e51da1bee02fd3eb66c603c851f09552b4363dd8ac4545a96039bc2

Observation 5ad401cd-72af-4fef-8bb7-e731b91238ce · outbound

This paper cites Learning to summarize with human feedback.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Learning to summarize with human feedback

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.876597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:9b7fdb3bfe5610b9a362e7b64f106f27f4475533031c9143a1e65615e5b56786

Observation 2ae292cf-cc5a-4eb9-afaf-454659e56a41 · outbound

This paper cites Rl grokking recipe: How does rl unlock and transfer new algorithms in llms? InInternational Conference on Learning Representations.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rl grokking recipe: How does rl unlock and transfer new algorithms in llms? InInternational Conference on Learning Representations

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.105202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:d21aed53671a4a6e9f94b0eb9374add6e3f28708ae4ea723e73e87abb210502c

Observation 33501b71-3632-4fad-8b33-1198ac9fb464 · outbound

This paper cites All roads lead to likelihood: The value of reinforcement learning in fine-tuning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient All roads lead to likelihood: The value of reinforcement learning in fine-tuning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.063971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:9c27f3458852301374639c683bc2ad2af9f5881486f8114fd447df0819e54ae8

Observation 82c9f1d3-92ad-4d9d-ba59-3a1cf0f06f06 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Understanding the performance gap between online and offline alignment algorithms

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.645340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:b7c445cc7c4710f7c688ab3822979ed96816074d19b2ed7627e4833a01b3d656

Observation 530bd498-9511-42bb-ac2e-50bc6c38d329 · outbound

This paper cites InThe Thirty-ninth Annual Conference on Neural Information Process- ing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient InThe Thirty-ninth Annual Conference on Neural Information Process- ing Systems

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:41:19.591630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:06059b957688e60a2e0d342179eaa9be3a9742efea988acfa6da9c5cc72a5b8d

Observation e9020eb9-ad60-4f07-bbc6-55043e5a9723 · outbound

This paper cites Perturbation analysis of neural collapse.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Perturbation analysis of neural collapse

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.137475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:e6398cfaea08387469b6653909c6f9659fcc46ccb10e10dc446e759193cb6462

Observation f6f5f133-7298-4f34-8121-1faf2ea73373 · outbound

This paper cites Implicit regularization in relu networks with the square loss.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in relu networks with the square loss

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.819104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:d430c399f6a245221ab2ac4ebee9d0b2fd34940d068904e682ab8f0dcdad351b

Observation 78c93807-1306-4273-b3fb-328262e3caf1 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.698135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:51d14abc3495e90f3089e4d0d41d09f2e63049f307ca845a8f5006af14b1491a

Observation 97aadcd7-abbf-4da2-a2fc-00e035fbd907 · outbound

This paper cites Rethinking reward model evaluation: Are we barking up the wrong tree? InInternational Conference on Learning Representations.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rethinking reward model evaluation: Are we barking up the wrong tree? InInternational Conference on Learning Representations

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.114951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:b905a909a595f028830d810a404d7b32503143612f6688494fe5e69135ca2c2e

Observation a299a1bd-a796-466a-bc73-4d8fd1093a67 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine learning, 8(3):229–256.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine learning, 8(3):229–256

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.165478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:e61aadbb74334de00568e9b755f1168953980e322a1070f85320f517ecb5d350

Observation ee639010-6b2f-4c5c-81d1-f0e14cb02f1c · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.611570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:3e652179f2e9090f93cb2e607add47b594ffb09265f94e0856af237c8939a314

Observation df978ab5-46a8-4420-8cba-df6b7d5ccc91 · outbound

This paper cites Kernel and rich regimes in overparametrized models.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Kernel and rich regimes in overparametrized models

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.008874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:146bd9fccc458a4d31ae4240c8fad313be76a933dd820701129bd20e1d4db9ad

Observation 91cb503f-62a6-4488-aa6a-22baec79bc27 · outbound

This paper cites Dynamics-aware comparison of learned reward functions.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Dynamics-aware comparison of learned reward functions

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.077423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:a4076eb5a5f033f00ea20c897b9e376f0fd997abf00644ad5704825332c3ef05

Observation cf047b6a-c400-4d90-8d3d-0e3bfbe6ea2d · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.562482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:3aec1919d7467787e8e3b8d5a8ed4e0d9495139cddbede3c8a51a08ba8babcf4

Observation f3fcb075-38e6-4eaf-b9b1-ee1850569394 · outbound

This paper cites Qwen3 Technical Report.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Qwen3 Technical Report

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.605231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:59084624b8983806a31e7c45828d10269de513673ac94841942fa21868e49f45

Observation 59af3c34-be70-43a5-9e92-5b9434de1e4d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.708864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:570813525b210dc190d2b23b68ab7b11d5b9a6b7676684901cd87cabcfdb5d28

Observation f7c9d62b-ed80-410c-8034-fa70ad788a31 · outbound

This paper cites On learning intrinsic rewards for policy gradient methods.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On learning intrinsic rewards for policy gradient methods

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.024062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:82da1cb4e3b2066ebd98295d351dd65ad4cd228a5501ca157603c99359ecb4a7

Observation 75db98d7-8b6b-4d13-8254-9d0ce84bccde · outbound

This paper cites Rmb: Comprehensively benchmarking reward models in llm alignment.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rmb: Comprehensively benchmarking reward models in llm alignment

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.087015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:83f69010037565b64381ec7d7d32fd864fbf1b96b706afdcab80164951cc8c3d

Observation 86b2e17f-577a-4982-8958-8f23b466519c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Fine-Tuning Language Models from Human Preferences

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.598253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:ddba0fc9298f14b6cb14c7e052486cca43fc5cf8d0c78d52ea5e448ab8ffb6c2

Observation ae83b582-bad8-46ba-b797-a75062ed3730 · outbound

This paper cites partial rewards.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient partial rewards

Reference 100

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T23:41:19.586625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:e7e92b8e08d033eaa2c9ab80548b1bb6706fcb50233d8418914699b4c153616c

Pith citing papers

Observation 28af7302-b8f7-4f48-885d-77815d82bb5e · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:45a619e147edd1d8ad2420d7aa0b1d568b5b4c434c86f93c479539c30370b4c8