Pith. sign in

Paper Citation Record · LEDGER

Effective Reward Specification in Deep Reinforcement Learning

As of 12 August 2026, this Paper Citation Record lists 100 of 300 outbound references and 0 inbound Pith citation observations for arXiv:2412.07177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07177 v1

Coverage vector

measured 100 of 300 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:09:54.006504Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 300 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved96
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f670229a-7884-4a60-a7c8-b7924ffe79e2 · outbound

This paper cites write newline.

Effective Reward Specification in Deep Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.173661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.173661Z digest=sha256:30f6b4c1a4324d3d3a46846ec5ffe302e283b45a24627613d6558ad087236edd

Observation 5ce6d39f-ebc7-4668-9d47-185b987f835d · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.182649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.182649Z digest=sha256:1338dbe2f525e3e70f7eac3cb37c296b7f3b39266d2d03409f652bfb9a5fa0cc

Observation bf6479eb-f08c-4c86-8f4c-3976865bd0b1 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.189309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.189309Z digest=sha256:ac95d71fb90908cf317a657f051b04bee33991c0dfb0d8b097870899876d9a28

Observation d9582fd8-9708-4b24-893a-6f74f4a45b4f · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.196826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.196826Z digest=sha256:0622281901d80d5a513f28ff23b3c7fb7f71d868b28ea44b20106e4e0b9f713c

Observation e41861d2-2621-4006-a9d9-57e5d91e27e3 · outbound

This paper cites and Ng, A.

Effective Reward Specification in Deep Reinforcement Learning and Ng, A

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.206559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.206559Z digest=sha256:850ae2781dd5690de0e3850d1debb9d390d4c5a8ead1e9cd59a9b2ef23c9339d

Observation 089b4a0c-6696-494f-853d-b41283a0127d · outbound

This paper cites and Ng, A.

Effective Reward Specification in Deep Reinforcement Learning and Ng, A

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.214761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.214761Z digest=sha256:f1be7039ed9c5ccafc5babba8a94711062cfc3d55a8d5a59ed0ef9781cbfb09c

Observation d9720b40-68cf-48ee-a7c9-0ee8239e1cf2 · outbound

This paper cites K., Littman, M., Precup, D., and Singh, S.

Effective Reward Specification in Deep Reinforcement Learning K., Littman, M., Precup, D., and Singh, S

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.221622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.221622Z digest=sha256:39867626efd8efbba7f56794077e8693b9e4d1b004258ff747df86118665e87a

Observation adfd7db3-5749-44a6-8eba-ed931304f984 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.229386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.229386Z digest=sha256:7592c794bcb92105bc17c3db8c7f1bdcdecdf1651dcb6732c7b005eed753b8da

Observation be446dbb-fbc9-4be0-b634-f49cbfb80154 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.235299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.235299Z digest=sha256:927a04b36f3007bc65bb95c95bc87eca59c2124d998243bf153e9f5d58c5d416

Observation a8a3d86c-0c6a-499e-8b9d-b6123b1aacb4 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.242527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.242527Z digest=sha256:eb9193f19c20df9940dab4e5e45267fbc2e07be73a4a515cbefba5fecbb6a930

Observation 6e22357e-ddb0-4685-bbcf-7166c49ae5ea · outbound

This paper cites Feudal Multi-Agent Hierarchies for Cooperative Reinforcement Learning.

Effective Reward Specification in Deep Reinforcement Learning Feudal Multi-Agent Hierarchies for Cooperative Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.250041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.250041Z digest=sha256:b0c58d293543c2307241131048e93c4c69a5d89e312482c936934dd58851c6b9

Observation 4b3ede61-5ab0-4aaf-a747-cec72cf31152 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Effective Reward Specification in Deep Reinforcement Learning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.255834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.255834Z digest=sha256:8a6955805690a42e31e16e86b56748270c9eac48874526e4a3f79a0338469693

Observation e7324c04-d8f0-4e08-b7ba-719ab1f929be · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.266861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.266861Z digest=sha256:2550df3600c675c255389dae3da1ba0448bc05b3f461c67d8b3dfd4013112640

Observation 4da7c10a-aab0-40f4-896d-aaaea15babd1 · outbound

This paper cites Solving Rubik's Cube with a Robot Hand.

Effective Reward Specification in Deep Reinforcement Learning Solving Rubik's Cube with a Robot Hand

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.274117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.274117Z digest=sha256:9c622692af2a82da16b6e003e474b073d99f5ade51b370c9928c53111a65d82d

Observation add28bfa-6651-4b5e-a342-db45149cb1a9 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.283558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.283558Z digest=sha256:6ff34d25bfcead8d0f9e539cb81578aa05445958920b48275b2340fcfd15bcb7

Observation 91c3085d-16f1-4244-97b7-213c21c114fe · outbound

This paper cites Deep Reinforcement Learning for Navigation in AAA Video Games.

Effective Reward Specification in Deep Reinforcement Learning Deep Reinforcement Learning for Navigation in AAA Video Games

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.289927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.289927Z digest=sha256:f90fe4583e82e937a2472feed4efcd3404708241ce1b2d5d0daea6793252e0b9

Observation a974e33a-8fd6-4da6-96d7-8583e287f451 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.300139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.300139Z digest=sha256:c8d216683fe98348d34d1e01daebd681b1c9ef4d46d7a0bc826ce54f9141316a

Observation 595fc39b-f931-4b4d-abbb-d9a20d1d694f · outbound

This paper cites A Survey of Exploration Methods in Reinforcement Learning.

Effective Reward Specification in Deep Reinforcement Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.307309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.307309Z digest=sha256:b98afde816a4812e7604a90b6803194039daeac81061b276d4e34e50286515af

Observation dd3de39f-95d6-46ae-8a40-2e46c3625b5b · outbound

This paper cites Concrete Problems in AI Safety.

Effective Reward Specification in Deep Reinforcement Learning Concrete Problems in AI Safety

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.314188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.314188Z digest=sha256:3c7b2b178761ad893a81b223ff41a0612ee8ac7c929e715c8548c1f7d0bebf73

Observation 7d1d635b-38b9-48a4-aaa5-f5256f9fa57f · outbound

This paper cites Explaining Reinforcement Learning to Mere Mortals: An Empirical Study.

Effective Reward Specification in Deep Reinforcement Learning Explaining Reinforcement Learning to Mere Mortals: An Empirical Study

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.322858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.322858Z digest=sha256:62d27af4df155088dfbd6c20da9745be2e101575b624aceec576d129e031ab6a

Observation e341f36f-1766-48f1-9580-b6bfd63d85c2 · outbound

This paper cites P., and Zaremba, W.

Effective Reward Specification in Deep Reinforcement Learning P., and Zaremba, W

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.331142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.331142Z digest=sha256:a25cb0742ea797dc5311948bda0f4bd164e924d8bd29ff9b8a6c5aae882adb9c

Observation 5d7f753a-6bbe-4747-acbc-fe25c17ebb63 · outbound

This paper cites M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al.

Effective Reward Specification in Deep Reinforcement Learning M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.339566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.339566Z digest=sha256:b900a0e40dadeb2307af740971125850e8b702e96d759fcaf565f75f193fa4ec

Observation d0a6d152-ee32-46e6-aa7b-c3f586a92ed5 · outbound

This paper cites and Doshi, P.

Effective Reward Specification in Deep Reinforcement Learning and Doshi, P

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.355859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.355859Z digest=sha256:011015b7ab4e4d26341a2e6fb896d40bb8a610645f8b0cc708a37a848f9f23da

Observation 628072b7-53e6-49e4-a569-91408f889bcb · outbound

This paper cites Accurately and Efficiently Interpreting Human-Robot Instructions of Varying Granularities.

Effective Reward Specification in Deep Reinforcement Learning Accurately and Efficiently Interpreting Human-Robot Instructions of Varying Granularities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.364279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.364279Z digest=sha256:214bee6ec56068517886556f9adc3cdcd28aa2a15d9fc6b1279e7cda2dc80333

Observation 53131749-ec93-4df1-bf40-04dda5322d80 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Effective Reward Specification in Deep Reinforcement Learning A General Language Assistant as a Laboratory for Alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.371592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.371592Z digest=sha256:30e7978c69c3d14acc449daff0f04c99a655e7f657344c8f21f977802bb1f1a8

Observation 13785aa4-9151-4ed9-91d5-cfab30392e51 · outbound

This paper cites DynGFN: Towards Bayesian Inference of Gene Regulatory Networks with GFlowNets.

Effective Reward Specification in Deep Reinforcement Learning DynGFN: Towards Bayesian Inference of Gene Regulatory Networks with GFlowNets

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.389663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.389663Z digest=sha256:0104e1b9346d95ed557807e909f3a41668e2743416d7c246eda6184d5096735f

Observation bbb50bcb-4d33-4ddd-8651-a3f00a12f6de · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.398571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.398571Z digest=sha256:6dce65f6bb10444ed1895c29d86db1a828c85dc7057a707b29e66b3897763861

Observation 5fd203ce-372e-46af-ab1d-4e7b8a4e940e · outbound

This paper cites Layer Normalization.

Effective Reward Specification in Deep Reinforcement Learning Layer Normalization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.405296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.405296Z digest=sha256:25ae558cd135ed083ea9c6ed87d1fb205c8b84500e6f80ef9816e1f7b7006d76

Observation de6390a0-20a3-407b-91f7-316ce7f5691b · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.412042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.412042Z digest=sha256:516178aba5d70352d2ccb0011c47ef29ecf428d88c0606f0cdcf8c9c960fd730

Observation 8b51f118-118b-4bce-9f9c-9d0bd1048d01 · outbound

This paper cites Learning to Understand Goal Specifications by Modelling Reward.

Effective Reward Specification in Deep Reinforcement Learning Learning to Understand Goal Specifications by Modelling Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.418720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.418720Z digest=sha256:616137edafd47668adca380d80d4d2f3aaaa70ae30bcc7aee1a265a259a7d937

Observation 1972b0e1-6630-4a0f-bd70-08c3b8ddf945 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Effective Reward Specification in Deep Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.425217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.425217Z digest=sha256:e7c74b61f7bcfd8c5271367b0de5e9de0ff9a2438aca231a3d25087842071a66

Observation f65dbb71-f6f0-4ef5-91f6-5a2a5fc7896c · outbound

This paper cites P., O’malley, M.

Effective Reward Specification in Deep Reinforcement Learning P., O’malley, M

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.436472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.436472Z digest=sha256:bd3d3bc636983617c3b10ef07d02ff8ecd05b5eef2cb8d820b61d0f08ae777f4

Observation 5b4979f7-dcff-4dff-9d1f-43fb575db73d · outbound

This paper cites and Narayanan, S.

Effective Reward Specification in Deep Reinforcement Learning and Narayanan, S

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.443398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.443398Z digest=sha256:4184eea6d25bba05b542e74ee20262b62aefec030aba7f83040df04191c45b01

Observation aa328edd-1c67-4763-b170-851b136f5d51 · outbound

This paper cites L., Waytowich, N.

Effective Reward Specification in Deep Reinforcement Learning L., Waytowich, N

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.449852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.449852Z digest=sha256:452accddeca81676fb741ff4b1b72e970e41f708636cca8b7546d106dd9e8351

Observation c9f9c7dc-196f-4ed8-bad2-3f9bb884c6e4 · outbound

This paper cites Graph augmented Deep Reinforcement Learning in the GameRLand3D environment.

Effective Reward Specification in Deep Reinforcement Learning Graph augmented Deep Reinforcement Learning in the GameRLand3D environment

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:09:58.013775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T19:09:53.455937Z digest=sha256:277f0d81f7a64a79189ce5fbf26f9c33c091e47ecc70204f857b83b6f780d442

Observation 6c883da5-c171-4883-aee1-8d3402a55021 · outbound

This paper cites G., Candido, S., Castro, P.

Effective Reward Specification in Deep Reinforcement Learning G., Candido, S., Castro, P

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.463071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.463071Z digest=sha256:a18d735872b29dc289eaf4e4a1a1faf3e8886a4017b0e799b7847da6cb16d7f3

Observation 889c3a95-3e21-460b-834e-2bfb943f18c6 · outbound

This paper cites A Distributional Perspective on Reinforcement Learning.

Effective Reward Specification in Deep Reinforcement Learning A Distributional Perspective on Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.481057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.481057Z digest=sha256:61a59e641990273c0df0191c695372cef85ffc82f9137ff95a917e9f6782ffb4

Observation 2ccca836-b239-4653-986a-01eb009f15c3 · outbound

This paper cites G., Naddaf, Y., Veness, J., and Bowling, M.

Effective Reward Specification in Deep Reinforcement Learning G., Naddaf, Y., Veness, J., and Bowling, M

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.490833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.490833Z digest=sha256:92808e804e87fe89074ee4638969a47b58424225ee63a92589555d4f98eb2789

Observation 1ee61b4a-6e58-4f43-953c-c0b5a7c81cbf · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.496335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.496335Z digest=sha256:8af15ed213cabcad49befab7068a2a0cb97d5f8b02dc18243610e2de4f095715

Observation d72c8e5e-efd3-4fef-82ba-911ded7c3725 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.502689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.502689Z digest=sha256:ef5bc2d60ab4c4f196b39a53b16289d5bfca72fdf40e55994ad32559950a6b2c

Observation 0c2d210c-b3ce-4eec-b8d6-079a4182122d · outbound

This paper cites J., Tiwari, M., and Bengio, E.

Effective Reward Specification in Deep Reinforcement Learning J., Tiwari, M., and Bengio, E

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.508816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.508816Z digest=sha256:fab5bc3efdd920f09ec973b70b0cddd4a32a451bfcb9494be8140b5e284835e6

Observation 04f0d674-0eb1-49b8-9c6b-a0e03efa5b1b · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.515914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.515914Z digest=sha256:999250ba8e9ff92a948a3276650a57f1e2e1a9025e652316947bd47b33e3ffc1

Observation 0c08540c-3345-488b-8f3e-b0a45aeda7af · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Effective Reward Specification in Deep Reinforcement Learning Dota 2 with Large Scale Deep Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.525635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.525635Z digest=sha256:051f6a0a1885934475a8c982ef0e1fc1e9b1bd7ecc37f3be66c98e5942561541

Observation ddfd02f1-1dd0-47ed-8614-3cccdf70e470 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.533327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.533327Z digest=sha256:b67bb093485b6ab4d0018a33f6af157b1cd56f81b663a1c4e2747c1680be11a3

Observation 46888275-f4a3-4d5b-aac9-e4832d064882 · outbound

This paper cites R., Paolini, G.

Effective Reward Specification in Deep Reinforcement Learning R., Paolini, G

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.540326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.540326Z digest=sha256:4205c9b62a14b2cafec989961176aaccf2c4568a9c093087752a8e6b362fdd0f

Observation 81cc074e-07d3-4d19-a596-10041e5454f3 · outbound

This paper cites Value constrained model-free continuous control.

Effective Reward Specification in Deep Reinforcement Learning Value constrained model-free continuous control

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:09:57.928298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T19:09:53.547799Z digest=sha256:4424dac2c8ba1db253228f8a57c7043953e043167cf75245de4ce7c4ff8a3af7

Observation 9eca8b83-2bda-490f-884b-869f29da442b · outbound

This paper cites B., Shah, J., Niekum, S., Stone, P., and Allievi, A.

Effective Reward Specification in Deep Reinforcement Learning B., Shah, J., Niekum, S., Stone, P., and Allievi, A

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.557042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.557042Z digest=sha256:72b92e5542211754116de01fb0b8aa1f34cf17f957d9ea974798a1a87af8b30d

Observation d16b75b8-bc84-4b85-853c-c0933d4b2282 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.564946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.564946Z digest=sha256:0e9ec2b35153c63e46213d857a88f0d0dde37815aaca5373b9bc7a96389d937e

Observation 4eb1e295-03c8-4341-8886-743a9bd84f65 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.574458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.574458Z digest=sha256:30862d5458323ab7d69022cc5deeb1b3d0edc2730f26316c0f772c01ad8096a6

Observation 992bda25-33e8-4904-a021-b0c9ff3f0139 · outbound

This paper cites D., Abel, D., and Dabney, W.

Effective Reward Specification in Deep Reinforcement Learning D., Abel, D., and Dabney, W

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.581171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.581171Z digest=sha256:d73e1ad326a21f4a7eb33b823fb99f5eefe21440be05e23ad3f80d2cfa25f4fb

Observation 7de053a7-3e23-4bf6-9bf8-26736579e514 · outbound

This paper cites OpenAI Gym.

Effective Reward Specification in Deep Reinforcement Learning OpenAI Gym

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.593141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.593141Z digest=sha256:b083c4fc9e89f6f7faeccfbe9cf66cf19a2c0712761eba5fb6beb45ff9d08e55

Observation 14aa8807-f81d-43a9-be22-00705510bf67 · outbound

This paper cites Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations.

Effective Reward Specification in Deep Reinforcement Learning Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.600198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.600198Z digest=sha256:e82f86422113ed8e028b5e1e22ec4469e87549bad5dc0f7281ab93b2e05c9013

Observation 239d6426-e858-4a33-b371-605d06227559 · outbound

This paper cites H., and Vaucher, A.

Effective Reward Specification in Deep Reinforcement Learning H., and Vaucher, A

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.613259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.613259Z digest=sha256:47df87a97cbc3ccdac0b2e624c92c9c70e9c6ef44719ec49673c4b7de180da20

Observation 19bf22df-ddb1-4b32-8e56-758bfff2917a · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.620653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.620653Z digest=sha256:b46abb1151d33e13e16f91210e5c4157a1896cc925c7fc8a7a2141bc15de0c51

Observation 5b02448d-cd8d-4a4a-a6f3-040bad8b795c · outbound

This paper cites Balancing Constraints and Rewards with Meta-Gradient D4PG.

Effective Reward Specification in Deep Reinforcement Learning Balancing Constraints and Rewards with Meta-Gradient D4PG

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:09:57.833406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T19:09:53.627582Z digest=sha256:0c9cf4543f6e2cb392ca8aaf29f42db602765a94d963cc784a0a10f2c2839da2

Observation 02ff1b52-1fc7-4ddb-8da8-65787e12fc85 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.635035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.635035Z digest=sha256:7631103ae8bb4b8d7e02d04c95cb73dc57428d16db5734d61fa15d67632ca59a

Observation 7ee31b36-4e34-46c4-8b35-7772d11a96eb · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.642271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.642271Z digest=sha256:85f1d79ca8f7143c0c4c33c368c02a20c42d24ef07166d4d3ce0ad96ea7e6fdf

Observation 26c6f9e7-cf8c-48d0-825a-abac7f9c29f7 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.653902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.653902Z digest=sha256:5379bc5f3515e7c85b160fcb6442505c48a507226fdae631155757adcf626319

Observation c0a9aee7-44c5-43f6-b7f5-5097d89ab056 · outbound

This paper cites u rnkranz, J., H \.

Effective Reward Specification in Deep Reinforcement Learning u rnkranz, J., H \

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.661473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.661473Z digest=sha256:69a44b0308b143efbe89cad7c73dc84a68991e512758f634737e1f41e2a45f73

Observation 3b879b8c-1080-45df-8598-7506052ce900 · outbound

This paper cites G., and Singh, S.

Effective Reward Specification in Deep Reinforcement Learning G., and Singh, S

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.667235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.667235Z digest=sha256:5839446bf3deef64769b91e9dc9f05f2356daa9958c59111903412bbb576754e

Observation ea2188bb-ae4b-4ce4-ad80-d572e4f02603 · outbound

This paper cites BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning.

Effective Reward Specification in Deep Reinforcement Learning BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.674314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.674314Z digest=sha256:6d7863d9a0787ec2f1db1ce994db792c30298d0117bce910f4825e1aaee94a1f

Observation 94e48b24-8898-4d09-8247-f5de408aea82 · outbound

This paper cites and Kim, K.-E.

Effective Reward Specification in Deep Reinforcement Learning and Kim, K.-E

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.681347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.681347Z digest=sha256:f03602e3014b0e970de7e3c03bdda41163b010d336940d462b812e2aff2c87b2

Observation 7edc8b98-99dc-4e0f-af39-9b949900490a · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.687973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.687973Z digest=sha256:483858ecf16bde24ee0b412ba837553c862d1b598b47e615c0d36d38eea8fb75

Observation a65f529c-6fd3-4e2f-9358-2d9686dcad04 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.696255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.696255Z digest=sha256:3dfbe5c8e9a82045b7c86ac823cb5409f88aeca8feb77d08446cb33f069449ed

Observation caead2f4-d565-4a16-b601-04a375a763eb · outbound

This paper cites Lyapunov-based Safe Policy Optimization for Continuous Control.

Effective Reward Specification in Deep Reinforcement Learning Lyapunov-based Safe Policy Optimization for Continuous Control

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.703918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.703918Z digest=sha256:b32e5e2823230add0931de6bc5dd1e3d6df4cfae97fa5f694f53d617ac90c715

Observation 44b13125-9c77-41dd-8aa1-5684c4b07b6a · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Effective Reward Specification in Deep Reinforcement Learning F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.712753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.712753Z digest=sha256:b186ce850b4d6fd8fcf6c608599e0c18ab93a355e0e2452f858152c003b0cb85

Observation 2758b84a-e19c-47c0-ba0c-8d85921589f9 · outbound

This paper cites and Amodei, D.

Effective Reward Specification in Deep Reinforcement Learning and Amodei, D

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.724470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.724470Z digest=sha256:e027e10ef71d485e9f4bb2eec401c2152a633af2fd248eab79579ab084cd02ce

Observation 2d750245-79b7-4c5b-83e7-4fb08c060cd9 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.733769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.733769Z digest=sha256:25db31a16c8ed932af6d8ea06b83e02da3421c7791391e1605ea3a4a75ab5ac5

Observation d5d4711a-99d6-4167-a660-35c939c1dfed · outbound

This paper cites and Niekum, S.

Effective Reward Specification in Deep Reinforcement Learning and Niekum, S

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.743079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.743079Z digest=sha256:166a562bf7241471c31d0bccb445154af2f894ddf0179568fbdc8e99a6f5605d

Observation 85619297-2c24-467d-9076-c0fadba86258 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.752217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.752217Z digest=sha256:c8b0cf55c80bc4a3e586171b7517e25e111b8a7f771ea6d2df32b3e83ba09480

Observation 3780a273-9c48-4032-a1fb-c49de9d14ff4 · outbound

This paper cites G., and Silver, D.

Effective Reward Specification in Deep Reinforcement Learning G., and Silver, D

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.759937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.759937Z digest=sha256:a08f281b2a1b67906e65032043e9f1eebebc6171d38b9267bb5e82c0da32d1d7

Observation 16dcd7c3-7925-4efb-9112-b146b1a77c9f · outbound

This paper cites G., and Munos, R.

Effective Reward Specification in Deep Reinforcement Learning G., and Munos, R

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.767064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.767064Z digest=sha256:f517559248402b3791c7f8545d2441a1ee6272a9e73bf949bfdaea841fb6b7bc

Observation f9562dfe-e4ee-4a2e-8b28-cb5b2815f8c0 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Effective Reward Specification in Deep Reinforcement Learning Safe Exploration in Continuous Action Spaces

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.774090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.774090Z digest=sha256:555d10e80f51f6ab5a1cb527d29b0ece1029be0ce3caeb526fcc087f542b8833

Observation fac7ee7e-2ea4-4aea-85ff-8e0a3cfa6ad1 · outbound

This paper cites Robotic Table Tennis: A Case Study into a High Speed Learning System.

Effective Reward Specification in Deep Reinforcement Learning Robotic Table Tennis: A Case Study into a High Speed Learning System

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.789787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.789787Z digest=sha256:430f56455fe43bccec1af0456fa8d5a7b35aafbc54e3e0536df39b5975f1ac6c

Observation 15297a9f-e93b-4ee4-9b89-b5847006731b · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.799929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.799929Z digest=sha256:88139cb2057c2871b3f476a8997e8b7a209cb5ed4ac12379dde35f6acccff4cb

Observation 001f70bd-cbd3-49e0-916e-f6f2aeda4092 · outbound

This paper cites and Dennis, J.

Effective Reward Specification in Deep Reinforcement Learning and Dennis, J

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.810292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.810292Z digest=sha256:fecddba1e4ab8b9a0df4ece183d2737538d604754a06058c55a97d02349760f0

Observation 27da0fd3-5053-4011-a8fc-d1c524a851e9 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.818319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.818319Z digest=sha256:516dff4db4e4d006d1d3d1f7521fd04944f7f2f1c1943ab06555300350db0073

Observation f3e78f71-fb45-4da2-8549-60d605283617 · outbound

This paper cites and Hinton, G.

Effective Reward Specification in Deep Reinforcement Learning and Hinton, G

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.823332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.823332Z digest=sha256:a06354ab9c26dd314023fbb6f0076b4d37695ca5d93b9c89dcb3ecec8c68a782

Observation 6481a611-15a4-4696-8bf7-50a41fd79e94 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.829780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.829780Z digest=sha256:33eb8377a47085864e4afc110c300150d4edd09c257c8024081556779ae7772e

Observation a2c90d26-5996-498a-b768-75edae709482 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.838358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.838358Z digest=sha256:1219ad42c5615ff4ef222bfc9deff2a0f17241bdba8ba000a860d848e90234d2

Observation 840cd487-d96a-4fdf-b43f-492de65b5755 · outbound

This paper cites Off-Policy Actor-Critic.

Effective Reward Specification in Deep Reinforcement Learning Off-Policy Actor-Critic

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.847628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.847628Z digest=sha256:c5036a6d609525fcfef20de14964f5fa09c84d051510cf56c850959e9f4090cf

Observation b7d33566-2f04-4e2a-bc4b-e8f91f0f4e92 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.859480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.859480Z digest=sha256:15d07a5269e2156cbe51200f0f5359dd296a8a63ed810ee20d9dfd557cde2f34

Observation f24761a7-93a9-4391-bcbd-ec185d154813 · outbound

This paper cites Navigation Turing Test (NTT): Learning to Evaluate Human-Like Navigation.

Effective Reward Specification in Deep Reinforcement Learning Navigation Turing Test (NTT): Learning to Evaluate Human-Like Navigation

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:09:57.685950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T19:09:53.867372Z digest=sha256:15e19cc4795a1fe75353bcdb8c21e04267d0f76f7ad133356821cb901212e489

Observation 9ef83963-391c-4eb3-9626-493e391a1fad · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.874206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.874206Z digest=sha256:df7c88154f7689ef12ce5c036088375a0ae42d73589895a976a5054fd6f417f5

Observation 119dc798-9155-4d1a-bdd8-c78ebf838a5c · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.884208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.884208Z digest=sha256:5db5e2605ef86a8d9868f1fc3d08d1af5e72adb43d351d3683064e493b1a6113

Observation da1f14d4-daf7-49d3-8e4f-5fa224e7e0f5 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.893045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.893045Z digest=sha256:515008a9da95fe30e724e4141d1e956997fd1f402082fd2a1da1d3d12796bf0c

Observation d48b8830-6454-4c79-a88c-9a5f82028692 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.901334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.901334Z digest=sha256:0d649876da0315d91d08ed42ce5b7113758e1ddd9d3c0b0dd2788546f297f51f

Observation 4b3be87f-bb72-460c-a6e4-7105fdf137d7 · outbound

This paper cites Adapting Auxiliary Losses Using Gradient Similarity.

Effective Reward Specification in Deep Reinforcement Learning Adapting Auxiliary Losses Using Gradient Similarity

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.907440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.907440Z digest=sha256:5b0396b8c321f089480e7314458af60b9489b7b78acd0a13b267404ab39f209b

Observation 0c18a82c-181b-4524-abf1-92d04fec2c01 · outbound

This paper cites J., Li, J., Paduraru, C., Gowal, S., and Hester, T.

Effective Reward Specification in Deep Reinforcement Learning J., Li, J., Paduraru, C., Gowal, S., and Hester, T

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.914874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.914874Z digest=sha256:3eca12e8b6f52551fdbdbeb2912c3ac2ec3f17a6ccab52232d5057e7ccef9b9d

Observation f2f5c98f-554d-4115-8e66-ef28c1c9709f · outbound

This paper cites Challenges of Real-World Reinforcement Learning.

Effective Reward Specification in Deep Reinforcement Learning Challenges of Real-World Reinforcement Learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.923218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.923218Z digest=sha256:e47660dbc9ce4c52f1102619483f762e8930ce665f443be0065e4b808c6cb05a

Observation c92bf16c-a7b9-4c30-b04e-397ca3f3ed6f · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.932442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.932442Z digest=sha256:4bf825f358fa574f2874b4da60823c00cf6e05a303fb251ad56412da0e9d01cd

Observation 99be24b8-21b7-4ab8-ad38-9033571aafd8 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.938774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.938774Z digest=sha256:9c19ca5751bac8b7a377477a9830d3998f9abdd8e2646b014a20d782aef29d5d

Observation f9c15b42-3b19-488e-b03c-07cbbd1e7e9c · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.944887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.944887Z digest=sha256:e01950b4c138a2325c2658c61756e7782a8975d931e07916bc6bdea41a7f9247

Observation faab6a52-19b8-4e91-bb55-212383f004b9 · outbound

This paper cites and Schuffenhauer, A.

Effective Reward Specification in Deep Reinforcement Learning and Schuffenhauer, A

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.951291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.951291Z digest=sha256:f7b65c80e50eeb86e3b74dd91c7190b52a65e6167c3700c3b85c9f70f2701e98

Observation 8c2c311c-7bfd-4269-bdc8-f213738306d2 · outbound

This paper cites and Gao, J.

Effective Reward Specification in Deep Reinforcement Learning and Gao, J

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.958861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.958861Z digest=sha256:e91ca78d1134197b8c9b44e3c1bc90266c87dc6e9a512f63a2833843e7c1d72b

Observation 9bb42d6f-36d9-44ee-80e2-991f392c92c2 · outbound

This paper cites Hyperbolic Discounting and Learning over Multiple Horizons.

Effective Reward Specification in Deep Reinforcement Learning Hyperbolic Discounting and Learning over Multiple Horizons

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.965045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.965045Z digest=sha256:6a145def80b1a0f85f8cf3b791b926bf70da33c13fc521e99f1e4b3cd39f2899

Observation 2e84dc15-c216-4e34-9b02-16f2c88580e5 · outbound

This paper cites Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation.

Effective Reward Specification in Deep Reinforcement Learning Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.971554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.971554Z digest=sha256:4d941227962b955b8d3498ac0f57214467b075121e013315e77b0cfb824833e1

Observation c3164b18-2f78-43eb-bf76-2278eabe7854 · outbound

This paper cites A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models.

Effective Reward Specification in Deep Reinforcement Learning A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.984622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.984622Z digest=sha256:68c4f0fb881d7da553361fd64fe393872c4843c27b714cacfdb4654e5a4c2edd

Observation d6382932-a9f2-460d-b85a-2396b92285ca · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.995214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.995214Z digest=sha256:1fef6c655fcc4c856ad541b1b077e1491659aa3f30173138e0b12e3c7123c6f6

Observation 5ecfacdc-344f-4fdf-b7a1-031a4d44fca1 · outbound

This paper cites A., de Freitas, N., and Whiteson, S.

Effective Reward Specification in Deep Reinforcement Learning A., de Freitas, N., and Whiteson, S

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:54.006504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:54.006504Z digest=sha256:768d15303976916b1a94bb9989880ab1afd81a3815a257f716b9ac6358c518ab

Pith citing papers

No inbound Pith citation observations are available.