Pith. sign in

Paper Citation Record · LEDGER

Effective Reward Specification in Deep Reinforcement Learning

As of 12 August 2026, this Paper Citation Record lists 100 of 300 outbound references and 0 inbound Pith citation observations for arXiv:2412.07177.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07177 v1

Coverage vector

measured 100 of 300 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:09:54.006504Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 300 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved96
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f670229a-7884-4a60-a7c8-b7924ffe79e2 · outbound

This paper cites write newline.

Effective Reward Specification in Deep Reinforcement Learning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.173661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.173661Z digest=sha256:c6908b622d4ab8e9154d0fa9eff9e8ba9a1e1d0880842fe23eece79926ad6861

Observation 5ce6d39f-ebc7-4668-9d47-185b987f835d · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.182649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.182649Z digest=sha256:b63781ccddb27743d9676ce5f5cce05b867440c0c206f953650d0a18bbd121ba

Observation bf6479eb-f08c-4c86-8f4c-3976865bd0b1 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.189309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.189309Z digest=sha256:f23dc5f1f720c3dafaf93f402e8731620e3c5b6911c4f0d8e9f3d1cfda0e8127

Observation d9582fd8-9708-4b24-893a-6f74f4a45b4f · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.196826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.196826Z digest=sha256:df7b33d605f8e0d0e737adbe866f18917f0bf59456ab65543b839ebfc541a0a8

Observation e41861d2-2621-4006-a9d9-57e5d91e27e3 · outbound

This paper cites and Ng, A.

Effective Reward Specification in Deep Reinforcement Learning and Ng, A

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.206559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.206559Z digest=sha256:2a8115c6830bc6488bc5f61433c99ea1f085bacbfc349f6a731e6848756a1853

Observation 089b4a0c-6696-494f-853d-b41283a0127d · outbound

This paper cites and Ng, A.

Effective Reward Specification in Deep Reinforcement Learning and Ng, A

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.214761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.214761Z digest=sha256:091347c986d960dc480c5ac279896be4f2b3404b9d5880074f59d6eb487e54bc

Observation d9720b40-68cf-48ee-a7c9-0ee8239e1cf2 · outbound

This paper cites K., Littman, M., Precup, D., and Singh, S.

Effective Reward Specification in Deep Reinforcement Learning K., Littman, M., Precup, D., and Singh, S

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.221622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.221622Z digest=sha256:ef4e165c8a6da7049938f63c6b4b806be0afec152fa33ef18da8e0e606d1ec30

Observation adfd7db3-5749-44a6-8eba-ed931304f984 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.229386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.229386Z digest=sha256:4f7e8ca1e1bbf74910c79080a439292f1e525c798fd0c4195d9e9ce76d6c3967

Observation be446dbb-fbc9-4be0-b634-f49cbfb80154 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.235299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.235299Z digest=sha256:8075a8c5a78409b76b5710004282d0f365f0980a91cb7af68a2876b4d13e84c7

Observation a8a3d86c-0c6a-499e-8b9d-b6123b1aacb4 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.242527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.242527Z digest=sha256:f28163f683273c5e95a8d8ed6ebabb33ebccdba139467e2c0449c5b9e32ed7a8

Observation 6e22357e-ddb0-4685-bbcf-7166c49ae5ea · outbound

This paper cites Feudal Multi-Agent Hierarchies for Cooperative Reinforcement Learning.

Effective Reward Specification in Deep Reinforcement Learning Feudal Multi-Agent Hierarchies for Cooperative Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.250041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.250041Z digest=sha256:c2c3dbc0306483e47a1c2ad91eed72f054eecf9743628d5e22d3449bcdda10b0

Observation 4b3ede61-5ab0-4aaf-a747-cec72cf31152 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Effective Reward Specification in Deep Reinforcement Learning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.255834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.255834Z digest=sha256:6418890405cefa1ba2a0c2af0c981dc27e9a97783d5f2d9746afc1d1d60af079

Observation e7324c04-d8f0-4e08-b7ba-719ab1f929be · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.266861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.266861Z digest=sha256:f25c740301522f826d77591e67cb6d860e1b7adfd37e13154f8db9bf8155224b

Observation 4da7c10a-aab0-40f4-896d-aaaea15babd1 · outbound

This paper cites Solving Rubik's Cube with a Robot Hand.

Effective Reward Specification in Deep Reinforcement Learning Solving Rubik's Cube with a Robot Hand

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.274117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.274117Z digest=sha256:bb3eeee423e83d27cd4c43a8e014832206d985c230856c1ce96c76b8a6501fd4

Observation add28bfa-6651-4b5e-a342-db45149cb1a9 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.283558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.283558Z digest=sha256:259a58c6c6c5420fd26bfc4d2a4cb85a4140f06a782e4c67f4a3e6a0717454a0

Observation 91c3085d-16f1-4244-97b7-213c21c114fe · outbound

This paper cites Deep Reinforcement Learning for Navigation in AAA Video Games.

Effective Reward Specification in Deep Reinforcement Learning Deep Reinforcement Learning for Navigation in AAA Video Games

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.289927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.289927Z digest=sha256:083a3b81c78db99a87f5877ca68b9dbe943ca413d56910b93b6a2cb5b344605d

Observation a974e33a-8fd6-4da6-96d7-8583e287f451 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.300139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.300139Z digest=sha256:eb7880cc66c82a30242ad07bb1acf4adf5ecb9383e1af720c9c22ed6b491f613

Observation 595fc39b-f931-4b4d-abbb-d9a20d1d694f · outbound

This paper cites A Survey of Exploration Methods in Reinforcement Learning.

Effective Reward Specification in Deep Reinforcement Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.307309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.307309Z digest=sha256:5a0c9d6f6853fbc7d84b3e8e03d53a5cb5ac47fd4549961a4849328f2a0d3d56

Observation dd3de39f-95d6-46ae-8a40-2e46c3625b5b · outbound

This paper cites Concrete Problems in AI Safety.

Effective Reward Specification in Deep Reinforcement Learning Concrete Problems in AI Safety

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.314188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.314188Z digest=sha256:80306ed7cabcaf06a06ed68e959aa6397fe23111c16a9e320b1aaacb4d712364

Observation 7d1d635b-38b9-48a4-aaa5-f5256f9fa57f · outbound

This paper cites Explaining Reinforcement Learning to Mere Mortals: An Empirical Study.

Effective Reward Specification in Deep Reinforcement Learning Explaining Reinforcement Learning to Mere Mortals: An Empirical Study

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.322858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.322858Z digest=sha256:b354c63f19981e5a0b7aa3e82dba1bbfda278c5b871d7421d755f91292495c71

Observation e341f36f-1766-48f1-9580-b6bfd63d85c2 · outbound

This paper cites P., and Zaremba, W.

Effective Reward Specification in Deep Reinforcement Learning P., and Zaremba, W

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.331142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.331142Z digest=sha256:480d77843eb823bed1c7ff62747954ade8690bc40adbcd8c59020add53f11238

Observation 5d7f753a-6bbe-4747-acbc-fe25c17ebb63 · outbound

This paper cites M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al.

Effective Reward Specification in Deep Reinforcement Learning M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.339566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.339566Z digest=sha256:9c674939c3ff510defce2161bdf629f99ad9072a24ee78f6a572449592ffa39e

Observation d0a6d152-ee32-46e6-aa7b-c3f586a92ed5 · outbound

This paper cites and Doshi, P.

Effective Reward Specification in Deep Reinforcement Learning and Doshi, P

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.355859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.355859Z digest=sha256:d24f4e12d3bd8096b6b84f9214019c56ae49a1210e721f4662020087122ef6f8

Observation 628072b7-53e6-49e4-a569-91408f889bcb · outbound

This paper cites Accurately and Efficiently Interpreting Human-Robot Instructions of Varying Granularities.

Effective Reward Specification in Deep Reinforcement Learning Accurately and Efficiently Interpreting Human-Robot Instructions of Varying Granularities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.364279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.364279Z digest=sha256:4dce1bbebf10696adec18f8083c09eba416603062ca63d1a047e3da841ab10f9

Observation 53131749-ec93-4df1-bf40-04dda5322d80 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Effective Reward Specification in Deep Reinforcement Learning A General Language Assistant as a Laboratory for Alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.371592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.371592Z digest=sha256:d6d3ecd8b97264c5584d32f1c1ea81a66ed1d0afb56ccc938151373846e225cd

Observation 13785aa4-9151-4ed9-91d5-cfab30392e51 · outbound

This paper cites DynGFN: Towards Bayesian Inference of Gene Regulatory Networks with GFlowNets.

Effective Reward Specification in Deep Reinforcement Learning DynGFN: Towards Bayesian Inference of Gene Regulatory Networks with GFlowNets

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.389663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.389663Z digest=sha256:f13572c484d90c473b71d073740a5c4559f414826480b95da4e05104c81710ee

Observation bbb50bcb-4d33-4ddd-8651-a3f00a12f6de · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.398571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.398571Z digest=sha256:74cec71bb44c9b1e8dd2468cff6a5b00c19d160a4f3b2174f350f411fc086420

Observation 5fd203ce-372e-46af-ab1d-4e7b8a4e940e · outbound

This paper cites Layer Normalization.

Effective Reward Specification in Deep Reinforcement Learning Layer Normalization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.405296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.405296Z digest=sha256:80f8af7af6fa794351201d693ce15909186ce8bedc4328afaabccd58e132c5ba

Observation de6390a0-20a3-407b-91f7-316ce7f5691b · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.412042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.412042Z digest=sha256:98bc9cb7274f4355b6967bcc05be894addc7016f1e419f72abdfe6970f34b403

Observation 8b51f118-118b-4bce-9f9c-9d0bd1048d01 · outbound

This paper cites Learning to Understand Goal Specifications by Modelling Reward.

Effective Reward Specification in Deep Reinforcement Learning Learning to Understand Goal Specifications by Modelling Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.418720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.418720Z digest=sha256:8c63f2cd9b438bb141ada5e8985918e6c9024a6ccb90a381e6747944b980fe8c

Observation 1972b0e1-6630-4a0f-bd70-08c3b8ddf945 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Effective Reward Specification in Deep Reinforcement Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.425217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.425217Z digest=sha256:43c19ace5ab5283895100d7fa725c63a833ef4d4ee9e499e736cd2601f771c3e

Observation f65dbb71-f6f0-4ef5-91f6-5a2a5fc7896c · outbound

This paper cites P., O’malley, M.

Effective Reward Specification in Deep Reinforcement Learning P., O’malley, M

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.436472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.436472Z digest=sha256:8a749fd3a9302c3ac5a84eeff89c12a839b8fcccee023ef3fb46ce83f27c2cf4

Observation 5b4979f7-dcff-4dff-9d1f-43fb575db73d · outbound

This paper cites and Narayanan, S.

Effective Reward Specification in Deep Reinforcement Learning and Narayanan, S

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.443398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.443398Z digest=sha256:cf151649ac57d508913009e5beb9f9eba4e192cc4bf2dbf538a42005b6433a05

Observation aa328edd-1c67-4763-b170-851b136f5d51 · outbound

This paper cites L., Waytowich, N.

Effective Reward Specification in Deep Reinforcement Learning L., Waytowich, N

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.449852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.449852Z digest=sha256:1a67ce109d37dfc66ee73ce687fd6e597515d2fbfb46e89a5375ad358521e7e5

Observation c9f9c7dc-196f-4ed8-bad2-3f9bb884c6e4 · outbound

This paper cites Graph augmented Deep Reinforcement Learning in the GameRLand3D environment.

Effective Reward Specification in Deep Reinforcement Learning Graph augmented Deep Reinforcement Learning in the GameRLand3D environment

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:09:58.013775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T19:09:53.455937Z digest=sha256:771dba75e74705a4204862497cf8d2807cbaf6fc2d34dc5f70faf5a643df6dbf

Observation 6c883da5-c171-4883-aee1-8d3402a55021 · outbound

This paper cites G., Candido, S., Castro, P.

Effective Reward Specification in Deep Reinforcement Learning G., Candido, S., Castro, P

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.463071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.463071Z digest=sha256:21a8802fcedf5ad6810b0be02bb2e517b1c4ce7316c4b36c297184e541372f86

Observation 889c3a95-3e21-460b-834e-2bfb943f18c6 · outbound

This paper cites A Distributional Perspective on Reinforcement Learning.

Effective Reward Specification in Deep Reinforcement Learning A Distributional Perspective on Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.481057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.481057Z digest=sha256:1e6962d1b840000ab973d7f86a141bf5a4fbaeccb6c4365b6e43c9adb686d873

Observation 2ccca836-b239-4653-986a-01eb009f15c3 · outbound

This paper cites G., Naddaf, Y., Veness, J., and Bowling, M.

Effective Reward Specification in Deep Reinforcement Learning G., Naddaf, Y., Veness, J., and Bowling, M

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.490833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.490833Z digest=sha256:8050533bc3a4c284826f2053ee74c255099b1fbf8084399c83993dbfc8fde9c2

Observation 1ee61b4a-6e58-4f43-953c-c0b5a7c81cbf · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.496335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.496335Z digest=sha256:6b9bb049f2fbf5ede08f78fd841afc80dd947544ce8bfc63ca5092bdbb1ab5d2

Observation d72c8e5e-efd3-4fef-82ba-911ded7c3725 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.502689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.502689Z digest=sha256:20242366c834d8b7f49372a8c51b9dc966c6c23af0beb6163fb1ce826b253138

Observation 0c2d210c-b3ce-4eec-b8d6-079a4182122d · outbound

This paper cites J., Tiwari, M., and Bengio, E.

Effective Reward Specification in Deep Reinforcement Learning J., Tiwari, M., and Bengio, E

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.508816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.508816Z digest=sha256:628229d1aa5eaace4bc52e71d7b06800ab90f5d885405f2b9a8af9b3cf55afb6

Observation 04f0d674-0eb1-49b8-9c6b-a0e03efa5b1b · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.515914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.515914Z digest=sha256:c3064da681ae18ab9bb5a7d03ea0bf7f59d8a31a09c8090df45ae15c9cf7c8fb

Observation 0c08540c-3345-488b-8f3e-b0a45aeda7af · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

Effective Reward Specification in Deep Reinforcement Learning Dota 2 with Large Scale Deep Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.525635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.525635Z digest=sha256:e18686572529f54ebb135df41e60fc08d84eb227218bb7bd71ba49c491357737

Observation ddfd02f1-1dd0-47ed-8614-3cccdf70e470 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.533327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.533327Z digest=sha256:65c73224cac13d7beb6677f1d00d7bf3542564a232a43a11a8acccee24bc2896

Observation 46888275-f4a3-4d5b-aac9-e4832d064882 · outbound

This paper cites R., Paolini, G.

Effective Reward Specification in Deep Reinforcement Learning R., Paolini, G

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.540326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.540326Z digest=sha256:5c2ec794c9e37cd29d114942a74b21484519fa43de12082107cd418880c40cac

Observation 81cc074e-07d3-4d19-a596-10041e5454f3 · outbound

This paper cites Value constrained model-free continuous control.

Effective Reward Specification in Deep Reinforcement Learning Value constrained model-free continuous control

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:09:57.928298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T19:09:53.547799Z digest=sha256:03f09767c8ee88d49cde1bcc66066546e06941bab1ad04397a6a86dd5d6fe49f

Observation 9eca8b83-2bda-490f-884b-869f29da442b · outbound

This paper cites B., Shah, J., Niekum, S., Stone, P., and Allievi, A.

Effective Reward Specification in Deep Reinforcement Learning B., Shah, J., Niekum, S., Stone, P., and Allievi, A

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.557042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.557042Z digest=sha256:462dabba1f0b12b1a8db3f8acf6e14df0b669a4c4d43477c5d4a0f89b91af367

Observation d16b75b8-bc84-4b85-853c-c0933d4b2282 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.564946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.564946Z digest=sha256:491fd18c63138befa3d4e1f882f732ff376997c611106c4086af6622d1160d7b

Observation 4eb1e295-03c8-4341-8886-743a9bd84f65 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.574458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.574458Z digest=sha256:79f8d98508567e037747fe37910aad7935d7f7757d813d08a9fc0ed4cf10bb99

Observation 992bda25-33e8-4904-a021-b0c9ff3f0139 · outbound

This paper cites D., Abel, D., and Dabney, W.

Effective Reward Specification in Deep Reinforcement Learning D., Abel, D., and Dabney, W

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.581171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.581171Z digest=sha256:61d9797b80ebebe87aa26edcdbe64c38121c361d751fc52a93e4d8629ec304e0

Observation 7de053a7-3e23-4bf6-9bf8-26736579e514 · outbound

This paper cites OpenAI Gym.

Effective Reward Specification in Deep Reinforcement Learning OpenAI Gym

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.593141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.593141Z digest=sha256:bd82752544b71de342962b36ed7a74a9e6faf8ea9a9905722cba8875270205d7

Observation 14aa8807-f81d-43a9-be22-00705510bf67 · outbound

This paper cites Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations.

Effective Reward Specification in Deep Reinforcement Learning Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.600198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.600198Z digest=sha256:16b68054b60a819a270f8efa79fd81783fb2466f6c2d54699e5e69b128aaa863

Observation 239d6426-e858-4a33-b371-605d06227559 · outbound

This paper cites H., and Vaucher, A.

Effective Reward Specification in Deep Reinforcement Learning H., and Vaucher, A

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.613259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.613259Z digest=sha256:1ea1cc8b834fcb67f5bc7580ab6959c3a083a678b95c62aa05cfd98f4ef1cec9

Observation 19bf22df-ddb1-4b32-8e56-758bfff2917a · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.620653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.620653Z digest=sha256:8e0e639585ec5cc45e867061a34b616ce9bc60c1c07b2b43e8459e9f51155519

Observation 5b02448d-cd8d-4a4a-a6f3-040bad8b795c · outbound

This paper cites Balancing Constraints and Rewards with Meta-Gradient D4PG.

Effective Reward Specification in Deep Reinforcement Learning Balancing Constraints and Rewards with Meta-Gradient D4PG

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:09:57.833406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T19:09:53.627582Z digest=sha256:c4ef7594c4a7661a9cace6a803f78e45d156ced32ff6b1ec6d0285a9cf687dee

Observation 02ff1b52-1fc7-4ddb-8da8-65787e12fc85 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.635035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.635035Z digest=sha256:e73c8a0dbc903897fca2cc634f6d656bb938db45a06f988cfbfb9d1b2a12f29a

Observation 7ee31b36-4e34-46c4-8b35-7772d11a96eb · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.642271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.642271Z digest=sha256:e8a2dfef096f569cbee98e37c113fdb0c356c10d29cdede7bf27d741722af357

Observation 26c6f9e7-cf8c-48d0-825a-abac7f9c29f7 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.653902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.653902Z digest=sha256:5a0c5aef7cc36bfc0b2ece2fd57f5fb17c830a0c40116e1600145e8858f5c8c0

Observation c0a9aee7-44c5-43f6-b7f5-5097d89ab056 · outbound

This paper cites u rnkranz, J., H \.

Effective Reward Specification in Deep Reinforcement Learning u rnkranz, J., H \

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.661473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.661473Z digest=sha256:210c630498c297a826a77208b28231a5697cee05eca8b7be9b593dd1e5544704

Observation 3b879b8c-1080-45df-8598-7506052ce900 · outbound

This paper cites G., and Singh, S.

Effective Reward Specification in Deep Reinforcement Learning G., and Singh, S

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.667235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.667235Z digest=sha256:0c82197ffde193d2da9090ad0c257f20d426657f0ef50e5ec172d2e95975de23

Observation ea2188bb-ae4b-4ce4-ad80-d572e4f02603 · outbound

This paper cites BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning.

Effective Reward Specification in Deep Reinforcement Learning BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.674314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.674314Z digest=sha256:36b9cca9fc80f84ea82cb5a698bbae2c097b3cea797309ceaaf96520184ccdb8

Observation 94e48b24-8898-4d09-8247-f5de408aea82 · outbound

This paper cites and Kim, K.-E.

Effective Reward Specification in Deep Reinforcement Learning and Kim, K.-E

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.681347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.681347Z digest=sha256:47e70cf85ead33b6ef5b4d9d00e6d72dde731a82cd164e8d4e0279d41d345df6

Observation 7edc8b98-99dc-4e0f-af39-9b949900490a · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.687973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.687973Z digest=sha256:5f4dcd1ab2a852dbdc03d629ce6a0ada4436f856900fac7e461514cbc0aad712

Observation a65f529c-6fd3-4e2f-9358-2d9686dcad04 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.696255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.696255Z digest=sha256:a9cb7feed518667604a9da08cb3f2ecd1d9d66c44a1627ef9bc09994a15d4ff4

Observation caead2f4-d565-4a16-b601-04a375a763eb · outbound

This paper cites Lyapunov-based Safe Policy Optimization for Continuous Control.

Effective Reward Specification in Deep Reinforcement Learning Lyapunov-based Safe Policy Optimization for Continuous Control

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.703918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.703918Z digest=sha256:52cabf34b0429c20e5c4f6a94466b0e38c0be57a0eaa0f30d92985b26cdfb4e4

Observation 44b13125-9c77-41dd-8aa1-5684c4b07b6a · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Effective Reward Specification in Deep Reinforcement Learning F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.712753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.712753Z digest=sha256:e1452904dfc1b5a2774c2d753171a587d32f4e1f435b68090730008c474437ad

Observation 2758b84a-e19c-47c0-ba0c-8d85921589f9 · outbound

This paper cites and Amodei, D.

Effective Reward Specification in Deep Reinforcement Learning and Amodei, D

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.724470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.724470Z digest=sha256:43ed4f477a4b9de38fc22e2cafab5b04843f20e7f776c49bb47e91cddf1169dd

Observation 2d750245-79b7-4c5b-83e7-4fb08c060cd9 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.733769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.733769Z digest=sha256:6843d6b78b98bb7ea94add18ee58f5788bfc378c6fb9bdf4f0484b242e194ff3

Observation d5d4711a-99d6-4167-a660-35c939c1dfed · outbound

This paper cites and Niekum, S.

Effective Reward Specification in Deep Reinforcement Learning and Niekum, S

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.743079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.743079Z digest=sha256:a3f5b192ca8b65d8bd6cd787a6895105298f7bd38666bb84cd47f3c88e083d1c

Observation 85619297-2c24-467d-9076-c0fadba86258 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.752217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.752217Z digest=sha256:9380adde2fafa44998fca06f3738f1928c9ea6454c1b6af942785bd978235ee9

Observation 3780a273-9c48-4032-a1fb-c49de9d14ff4 · outbound

This paper cites G., and Silver, D.

Effective Reward Specification in Deep Reinforcement Learning G., and Silver, D

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.759937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.759937Z digest=sha256:7ad6fc963f1b80a30bf7f3cfdbb704faec403fe608ca5ad0fc242a8f34389e8b

Observation 16dcd7c3-7925-4efb-9112-b146b1a77c9f · outbound

This paper cites G., and Munos, R.

Effective Reward Specification in Deep Reinforcement Learning G., and Munos, R

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.767064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.767064Z digest=sha256:677f4bde97249edb78ed02798e75c7db1e49c4deda484736b0801b20ddc9c0ab

Observation f9562dfe-e4ee-4a2e-8b28-cb5b2815f8c0 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Effective Reward Specification in Deep Reinforcement Learning Safe Exploration in Continuous Action Spaces

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.774090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.774090Z digest=sha256:60c56a9cb10b4b0ca31b8c1837a393a8878fa6688b7a79417b8e80e7251ee1d9

Observation fac7ee7e-2ea4-4aea-85ff-8e0a3cfa6ad1 · outbound

This paper cites Robotic Table Tennis: A Case Study into a High Speed Learning System.

Effective Reward Specification in Deep Reinforcement Learning Robotic Table Tennis: A Case Study into a High Speed Learning System

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.789787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.789787Z digest=sha256:95b78e8f648022985e8840134d365ae3bf0bf363c512c8d8c43b8f9dcd36a742

Observation 15297a9f-e93b-4ee4-9b89-b5847006731b · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.799929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.799929Z digest=sha256:da08f5149b66ee6189869a8cc65c6b9baff561c54bec949500406bd735a5aaa0

Observation 001f70bd-cbd3-49e0-916e-f6f2aeda4092 · outbound

This paper cites and Dennis, J.

Effective Reward Specification in Deep Reinforcement Learning and Dennis, J

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.810292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.810292Z digest=sha256:15921d86994d52865f6952baa9ddf335ff71b22b0b27485e1d31baf896558d6c

Observation 27da0fd3-5053-4011-a8fc-d1c524a851e9 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.818319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.818319Z digest=sha256:31b9b39f5048cf941a8fae3ad9d8774a328e4ab4949a77eaa5dc5996e54edb17

Observation f3e78f71-fb45-4da2-8549-60d605283617 · outbound

This paper cites and Hinton, G.

Effective Reward Specification in Deep Reinforcement Learning and Hinton, G

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.823332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.823332Z digest=sha256:2c8da6d1332a4ff1858ea0099ec03af1b391429dbb9e86a6d0a6e9dacd28c06f

Observation 6481a611-15a4-4696-8bf7-50a41fd79e94 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.829780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.829780Z digest=sha256:7cf4fe5d84f0df3520e70fd65495a3071498de3cc8feeb7f90e1d2a040b7d069

Observation a2c90d26-5996-498a-b768-75edae709482 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.838358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.838358Z digest=sha256:10eaa49b29e37f8326c92aee694ecbb3566177828d90bbd63668049e942aaa2c

Observation 840cd487-d96a-4fdf-b43f-492de65b5755 · outbound

This paper cites Off-Policy Actor-Critic.

Effective Reward Specification in Deep Reinforcement Learning Off-Policy Actor-Critic

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.847628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.847628Z digest=sha256:822bec9433f46a246ed256fd68d2645aaa77357136318aa2f0a9a133181fa97f

Observation b7d33566-2f04-4e2a-bc4b-e8f91f0f4e92 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.859480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.859480Z digest=sha256:f9d1b0714e049545a41ad797d40383d64616b5a4709ecdfd5b2734223689b24e

Observation f24761a7-93a9-4391-bcbd-ec185d154813 · outbound

This paper cites Navigation Turing Test (NTT): Learning to Evaluate Human-Like Navigation.

Effective Reward Specification in Deep Reinforcement Learning Navigation Turing Test (NTT): Learning to Evaluate Human-Like Navigation

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:09:57.685950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T19:09:53.867372Z digest=sha256:a952ee20bd83ffe704731fd830c76b96939d04235e96dd32c6e7a21a70ccf423

Observation 9ef83963-391c-4eb3-9626-493e391a1fad · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.874206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.874206Z digest=sha256:4e8e724fa4f1551e47d6d4583400385d179c90d53120e3b0267305649b4c63ec

Observation 119dc798-9155-4d1a-bdd8-c78ebf838a5c · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.884208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.884208Z digest=sha256:37e374eefae032922ce4bb1c2bc5103eea3e087e95a9e5bce775364a131211ae

Observation da1f14d4-daf7-49d3-8e4f-5fa224e7e0f5 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.893045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.893045Z digest=sha256:da20fe3a98449199c7d18f585417cf9e091172eb83c00b1fb9b761e38c796e8a

Observation d48b8830-6454-4c79-a88c-9a5f82028692 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.901334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.901334Z digest=sha256:794669e1c5921c790173ba1dffd0f47d16da76d2951b8e773a83ced0ec70719f

Observation 4b3be87f-bb72-460c-a6e4-7105fdf137d7 · outbound

This paper cites Adapting Auxiliary Losses Using Gradient Similarity.

Effective Reward Specification in Deep Reinforcement Learning Adapting Auxiliary Losses Using Gradient Similarity

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.907440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.907440Z digest=sha256:1e59f8c66eb1b23e4de49ff4037985a2f3474b435444045a552f63477446c177

Observation 0c18a82c-181b-4524-abf1-92d04fec2c01 · outbound

This paper cites J., Li, J., Paduraru, C., Gowal, S., and Hester, T.

Effective Reward Specification in Deep Reinforcement Learning J., Li, J., Paduraru, C., Gowal, S., and Hester, T

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.914874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.914874Z digest=sha256:670331ca711c0d323e765d58d997acfcc06cd08f24ac392431b185b15e34c8bf

Observation f2f5c98f-554d-4115-8e66-ef28c1c9709f · outbound

This paper cites Challenges of Real-World Reinforcement Learning.

Effective Reward Specification in Deep Reinforcement Learning Challenges of Real-World Reinforcement Learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.923218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.923218Z digest=sha256:62b2a3f4c4bd24e6fa3fa1e09416b07c6c1a439903b79831d4c5e179afe3487b

Observation c92bf16c-a7b9-4c30-b04e-397ca3f3ed6f · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.932442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.932442Z digest=sha256:2f9a12593465cef2babc5d7057adf1d2045844a456c2c6f1ebc7ba7afe95945f

Observation 99be24b8-21b7-4ab8-ad38-9033571aafd8 · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.938774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.938774Z digest=sha256:9e71c0c93ee87e67506a9b917f4e8326c5fe76803534e1332c4ac7fbe7618c1e

Observation f9c15b42-3b19-488e-b03c-07cbbd1e7e9c · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.944887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.944887Z digest=sha256:749c2529dcbba1493b6730952be350393970228edeb8878373acac7eeef097f9

Observation faab6a52-19b8-4e91-bb55-212383f004b9 · outbound

This paper cites and Schuffenhauer, A.

Effective Reward Specification in Deep Reinforcement Learning and Schuffenhauer, A

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.951291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.951291Z digest=sha256:96881a160fe02c59a285591dcf2e90ed74313a2356a7bb22ae8379a7a198eba7

Observation 8c2c311c-7bfd-4269-bdc8-f213738306d2 · outbound

This paper cites and Gao, J.

Effective Reward Specification in Deep Reinforcement Learning and Gao, J

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.958861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.958861Z digest=sha256:372758103128d67b71c5bf028bbcbc87df567fcd45a57413ae4d744d71d1dda4

Observation 9bb42d6f-36d9-44ee-80e2-991f392c92c2 · outbound

This paper cites Hyperbolic Discounting and Learning over Multiple Horizons.

Effective Reward Specification in Deep Reinforcement Learning Hyperbolic Discounting and Learning over Multiple Horizons

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.965045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.965045Z digest=sha256:771db3514d8f2142d7d122118d9f54387f3054d1df976b53d9ee8c695e95e1fa

Observation 2e84dc15-c216-4e34-9b02-16f2c88580e5 · outbound

This paper cites Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation.

Effective Reward Specification in Deep Reinforcement Learning Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.971554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.971554Z digest=sha256:19caebd5e7ee4931eb98237090b38952bff8d144963f292f7059325fde492241

Observation c3164b18-2f78-43eb-bf76-2278eabe7854 · outbound

This paper cites A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models.

Effective Reward Specification in Deep Reinforcement Learning A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.984622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.984622Z digest=sha256:e7c280f5dfe028885f8f1645aef0144b078c3636dac1c9ee6c025165e2c4e455

Observation d6382932-a9f2-460d-b85a-2396b92285ca · outbound

This paper cites an unresolved cited work.

Effective Reward Specification in Deep Reinforcement Learning Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.995214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.995214Z digest=sha256:515aac1c26d1b73ebd88c796fea3ce5c3b31bbc9fe6c6857b6bf1e5839b8e757

Observation 5ecfacdc-344f-4fdf-b7a1-031a4d44fca1 · outbound

This paper cites A., de Freitas, N., and Whiteson, S.

Effective Reward Specification in Deep Reinforcement Learning A., de Freitas, N., and Whiteson, S

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:54.006504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:54.006504Z digest=sha256:627dbf88163b08e9db06b240b06873c01614f544314add82e6c671f96cd9db01

Pith citing papers

No inbound Pith citation observations are available.