Pith. sign in

Paper Citation Record · LEDGER

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

As of 13 August 2026, this Paper Citation Record lists 100 of 265 outbound references and 9 inbound Pith citation observations for arXiv:2411.18892.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18892 v2

Coverage vector

measured 100 of 265 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:48:42.230024Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:24:03.035487Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:56:56.507096Z

Reference resolution

100 of 265 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved91
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6eb8a5da-10ca-41d7-8182-90656dafc1e0 · outbound

This paper cites an unresolved cited work.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:40.860021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:40.860021Z digest=sha256:92bfbe2f7eae972a72af5408e4430ba76c29bbb4c6e7743067e4e36d3d221d02

Observation 8ebf840b-024f-4dc7-9430-0acbf5a89157 · outbound

This paper cites The theory of dynamic programming,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges The theory of dynamic programming,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:40.922785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:40.922785Z digest=sha256:b2954242de214a91d90d19ead60cbd7af0185f27ca32ce132581c21330c22bd8

Observation 4f61ff19-7df0-40e9-8adc-506dd7dd922a · outbound

This paper cites Introduction to Reinforcement Learning.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Introduction to Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:40.987136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:40.987136Z digest=sha256:3ee20381ff91552d18c372fab561c079ff0389f8b36ef4b944641ef3baf8365f

Observation d2f3bd12-48e7-4497-8a14-11b9db7584c6 · outbound

This paper cites Learning to predict by the methods of tempo ral differences,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Learning to predict by the methods of tempo ral differences,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:40.991742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:40.991742Z digest=sha256:be9c7b276582b1518734a51fa38d4bfde81d8a0456cc872585e0860d608a8de7

Observation c9d16a82-da10-4a31-ba3e-91b6e7b6115c · outbound

This paper cites Learning from delayed rewards,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Learning from delayed rewards,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:40.995985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:40.995985Z digest=sha256:7bd85fabc8e30ec10a20084bc48269669fe66b5a5833d3280d3afdd79608444f

Observation 1835866b-d3c8-43dc-ad65-6501660d4f25 · outbound

This paper cites Rein- forcement learning: A survey,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Rein- forcement learning: A survey,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.000214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.000214Z digest=sha256:392f94bdb347b5d2026a682910366ef2c8bb2653d2f9ca285cb0c0bbabc430a6

Observation 4b199c9b-d96b-4a68-b8fc-c8ada6694e3b · outbound

This paper cites Finite-time a nalysis of the multiarmed bandit problem,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Finite-time a nalysis of the multiarmed bandit problem,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.005245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.005245Z digest=sha256:467ae95169ed72858b7eacc8df335182be1301e3591d92e3309af95f307cd8fe

Observation 92fa5ddc-472a-4e16-b03e-49b3a584e10c · outbound

This paper cites Human-level control through deep rein- forcement learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Human-level control through deep rein- forcement learning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.008774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.008774Z digest=sha256:66a142fcea365e1c21298095b352a727885abc10b82e4cf311a97261ae9d6f1b

Observation 4cf596b4-a013-4e06-9002-5f38eaab5471 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Mastering the game of go with deep neural networks and tree search,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.013806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.013806Z digest=sha256:dd99ae4ee61bf713f039f9a096d0228a5dd8ff787395263ba6608296acd98b12

Observation 051f351e-febc-4267-990d-4f0e1beca322 · outbound

This paper cites Reinforcement learning in game industry—review, prospec ts and challenges,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement learning in game industry—review, prospec ts and challenges,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.018254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.018254Z digest=sha256:9382979facd565d5a8d9db5ea2d1164e32900aced6631e413f3859e281aaa606

Observation f996e11a-f84b-463f-8582-921a1199bb80 · outbound

This paper cites Pgx: Hardware-accelerated parallel game simulators for reinforcement learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Pgx: Hardware-accelerated parallel game simulators for reinforcement learning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.023930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.023930Z digest=sha256:fb038404dcfd4ccba8e46f003fdd416cda95d9359c7e34ec6c41f576156e6aa9

Observation 5a89af55-ed63-4ac8-a6f7-459321266e7d · outbound

This paper cites Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.028168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.028168Z digest=sha256:b0c1f129eef025ca8f0e005570eae93cedf680c04360ccfe42611ca3a6f9904b

Observation 195557f6-2a4a-427b-8a69-8a4830fa2488 · outbound

This paper cites Pursuit-evasion gam e strategy of usv based on deep reinforcement learning in comp lex multi-obstacle environment,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Pursuit-evasion gam e strategy of usv based on deep reinforcement learning in comp lex multi-obstacle environment,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.033382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.033382Z digest=sha256:e49d0a4fe0a8b23aba8f2b7dc1096c230295f2df24274efc382c2810034dc3e7

Observation 93ef6eb5-43f9-4b0d-8256-3714a2072a3e · outbound

This paper cites Residual skill policies: Learning an adaptable skill-bas ed ac- tion space for reinforcement learning for robotics,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Residual skill policies: Learning an adaptable skill-bas ed ac- tion space for reinforcement learning for robotics,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.036964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.036964Z digest=sha256:a776b665328f27720270fd0fe8f3f2265cb406205f85668ea453ccf62d7e23b6

Observation 84bd48e5-7633-4d86-92dc-7376a1de70ca · outbound

This paper cites Reinforcement l earning in robotics: A survey,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement l earning in robotics: A survey,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.040700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.040700Z digest=sha256:189a73e55e9686f400d864e2612a79d33e6cfb591a5863e46e5358b709799297

Observation 48dbe19f-733b-4af5-8ea1-01ad2c0102b6 · outbound

This paper cites Intrinsically motivated multi- goal rein- forcement learning using robotics environment integrated with openai gym,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Intrinsically motivated multi- goal rein- forcement learning using robotics environment integrated with openai gym,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.044625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.044625Z digest=sha256:43acdea5eac8399eb814bcd4266397c38d10156e64488630b8133f9821422fad

Observation 466971ec-fac4-4945-82c0-b5155e8b8af2 · outbound

This paper cites i-sim2real: Reinforcement learning of robotic p olicies in tight human-robot interaction loops,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges i-sim2real: Reinforcement learning of robotic p olicies in tight human-robot interaction loops,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.048641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.048641Z digest=sha256:e1efa4f84f1d3e576850b3200433880e4b2943f9013d3af1da5139323e8f1e11

Observation d9567e94-9df7-474e-96fe-e5267b9dab6a · outbound

This paper cites Multi-agent broad reinforcement learning for intelligent traffic light control,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multi-agent broad reinforcement learning for intelligent traffic light control,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.052502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.052502Z digest=sha256:efd510de37a46f53e4d1dffe70096d8085d787fdc396de62e8d8e55b46f73703

Observation 88715876-e49a-47ed-9eab-5b318b5181e4 · outbound

This paper cites Intelligent vehicle pedestrian light (ivp l): A deep reinforcement learning approach for traffic signal con trol,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Intelligent vehicle pedestrian light (ivp l): A deep reinforcement learning approach for traffic signal con trol,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.056097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.056097Z digest=sha256:26c305a742d2883373ea09f40a88188c26de915d6fd492a2474d4df160033214

Observation 06f5cca0-4302-4472-a14f-3641b671d86c · outbound

This paper cites Swarm learning- based dynamic optimal management for traffic congestion in 6g-driven intelligent transportation system,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Swarm learning- based dynamic optimal management for traffic congestion in 6g-driven intelligent transportation system,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.097258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.097258Z digest=sha256:444c50d3a71cc5af3c12eba425bf4704af3fc36584b49f5fe4e45ed12ae92d7e

Observation 90e2b73d-c105-4c16-8093-91f40cdde1ee · outbound

This paper cites Deep multi-agent reinforcement learning for highway on-ramp merging in mixed traffic,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Deep multi-agent reinforcement learning for highway on-ramp merging in mixed traffic,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.192356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.192356Z digest=sha256:06da72098fdd9b6af83cf45a3da5e8c52580490f536bd898eef5eb1ff4c58114

Observation 4b078d56-7a45-49b8-b13e-609b3bf6d650 · outbound

This paper cites Maximizing group-based vehicle communications and fairn ess: A reinforcement learning approach,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Maximizing group-based vehicle communications and fairn ess: A reinforcement learning approach,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.266637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.266637Z digest=sha256:6cb597b6ec594b5ead3c3da3c6adab6009994aa13f7f1631e910370b2a568034

Observation b33735d3-337e-489c-8670-3661bbe81744 · outbound

This paper cites A real-world application of markov chain monte carlo method for bayesian trajectory control of a robotic manipulator,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A real-world application of markov chain monte carlo method for bayesian trajectory control of a robotic manipulator,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.343650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.343650Z digest=sha256:350eca9ff0b0daa0e53d406d2e11738f1760f33e8748210eb413542cc6acaff1

Observation a92b5b29-4294-4a73-919c-a87f8a547f47 · outbound

This paper cites Obstacle avoidance method based on double dqn for agricultural robot s,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Obstacle avoidance method based on double dqn for agricultural robot s,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.348079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.348079Z digest=sha256:02216db79fa5499ac28a71aa90ae6d2a65fe624d470ca7eb04a63950788fbe15

Observation 29cf150b-2f72-4e26-806a-a3e1716ba69f · outbound

This paper cites Adaptive importance sampling monte carl o simulation of rare transition events,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Adaptive importance sampling monte carl o simulation of rare transition events,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.351609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.351609Z digest=sha256:d5d3a5260a3e1a387179f3250b657a7fb6bb08818870e659012d88d2b1d33729

Observation 24fc822a-55e1-485b-b398-24aae442a5bb · outbound

This paper cites Dou- ble dqn-based coevolution for green distributed heterogen eous hybrid flowshop scheduling with multiple priorities of jobs ,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Dou- ble dqn-based coevolution for green distributed heterogen eous hybrid flowshop scheduling with multiple priorities of jobs ,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.355160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.355160Z digest=sha256:9a18ceb0a4e912422929890e6c6055617b04ec06a0b2b2a0db2572169d1a8dd0

Observation f40a9ad3-784b-4de4-ad6d-dd14c1c237d4 · outbound

This paper cites An intel li- gent train control approach based on the monte carlo reinfor ce- ment learning algorithm,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges An intel li- gent train control approach based on the monte carlo reinfor ce- ment learning algorithm,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.358390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.358390Z digest=sha256:e85250be95723faf5e816d034f5cfb1c4c94f5d4f64501e644af12906e70e847

Observation 9bc4987e-dfbe-464b-ad2e-fcbdd4390a9a · outbound

This paper cites Green computing in heterogeneous internet of things: Optimizing energy alloc ation using sarsa-based reinforcement learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Green computing in heterogeneous internet of things: Optimizing energy alloc ation using sarsa-based reinforcement learning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.361862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.361862Z digest=sha256:690e14b83b82f0452509b19f304d7b9fb17597a8134474f802751b147573620b

Observation 6c3b61f3-78de-4142-bff9-b70591534c47 · outbound

This paper cites Power allocation in multi-ag ent networks via dueling dqn approach,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Power allocation in multi-ag ent networks via dueling dqn approach,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.365078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.365078Z digest=sha256:7dd1af02f10558b3eabd7f75f711a9e1fefcb9af62b585f9ee613c6d1ef19113

Observation 69c26950-5197-4e01-96c9-182994abdf5a · outbound

This paper cites Efficient importa nce sampling for monte carlo simulation of multicast networks,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Efficient importa nce sampling for monte carlo simulation of multicast networks,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.368113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.368113Z digest=sha256:e247fc496971097751d88583260864500624b904a2f1e0b88732008cb605ae99

Observation b6a5f9aa-ff6c-4dc3-b266-e99e3e8d0a99 · outbound

This paper cites Reinforcement-learning-based ener gy storage system operation strategies to manage wind power forecast uncertainty,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement-learning-based ener gy storage system operation strategies to manage wind power forecast uncertainty,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.371619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.371619Z digest=sha256:522ac1bc67d42e407d4929ab710ec275d0901995ec63ea78ca4b06e5c2548cbc

Observation 98a18eb9-d08a-4504-ac1b-0827d24807ee · outbound

This paper cites Learning-based power delay profile estimatio n for 5g nr via advantage actor-critic (a2c),.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Learning-based power delay profile estimatio n for 5g nr via advantage actor-critic (a2c),

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.375262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.375262Z digest=sha256:a8287647f6942003b6a4f21e9e55302074ad8b9a47c8712ddc6225337752857e

Observation f29e2377-b74d-49d2-a97f-f4a5550e59f0 · outbound

This paper cites Per-decision Multi-step Temporal Difference Learning with Control Variates.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Per-decision Multi-step Temporal Difference Learning with Control Variates

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:45.171220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:41.378646Z digest=sha256:43dbb4aca350ef851b739931f605cd41ddd8cbbfe15a5e8306b2270d90c8abf0

Observation e05ce389-f4a6-4f68-916b-9ffad6819c7e · outbound

This paper cites A motion control method for agent based on dyna-q algorithm,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A motion control method for agent based on dyna-q algorithm,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.385374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.385374Z digest=sha256:843ab6ce310d04e83f265b0974c69faaf6610007c190bf60c1cf0b691dcb76b0

Observation 9ee22d97-6343-4a12-930c-85f740d0aa62 · outbound

This paper cites Temporal-difference netwo rks,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Temporal-difference netwo rks,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.389494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.389494Z digest=sha256:698aed578b64b3b8171d644e38c4856068c3f64b3060e2d18626fe93675f3c79

Observation 7f6e22e9-7f97-475e-b4ee-56986fd4a500 · outbound

This paper cites Realizing undelayed n-step td prediction w ith neural networks,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Realizing undelayed n-step td prediction w ith neural networks,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.393338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.393338Z digest=sha256:fb1fc1e156308ad4ab3306bf6aa9d5113a57719048a6de4f35bb2b2bd1ea6c15

Observation 48d5d8f5-651c-4722-ab97-d71b3138e40f · outbound

This paper cites Monte Carlo Tree Search for Asymmetric Trees.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Monte Carlo Tree Search for Asymmetric Trees

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:45.075116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:41.396648Z digest=sha256:4a618adf2053b005d2d15fc59748c0e01fc48b06d92338e023e4accf395f299c

Observation 75a70ff1-bb79-4908-90c1-138dad307666 · outbound

This paper cites An overview of q-learni ng based energy efficient power allocation in wban (q-eepa),.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges An overview of q-learni ng based energy efficient power allocation in wban (q-eepa),

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.400456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.400456Z digest=sha256:7cdd2e521b0e997982206bd5807c1fd4ce2d306569ef237a1d005ba4dbc04b0d

Observation 6642a15e-4047-44a8-9054-4edc75a36e4f · outbound

This paper cites Intelligent control of a quadrotor with proximal pol icy optimization reinforcement learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Intelligent control of a quadrotor with proximal pol icy optimization reinforcement learning,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.404069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.404069Z digest=sha256:b6a4285368cce6fd46cccb749ce401fadaa66b8adb904905a8e731c14bd48f67

Observation ed47a431-7ec0-49de-8da3-42ee5acd76d4 · outbound

This paper cites Automated por t- folio rebalancing using q-learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Automated por t- folio rebalancing using q-learning,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.449478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.449478Z digest=sha256:df1e088a2c51a3a5764c664be99abd293d2d12aa51d0d35de27e51dd1c8780d0

Observation 994b04fb-8929-41c8-ad56-53da4699eb7f · outbound

This paper cites Prioritized sweeping reinfor cement learning based routing for manets,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Prioritized sweeping reinfor cement learning based routing for manets,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.527492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.527492Z digest=sha256:e9f8e5ed09a5c8f0a333b3a1f8ec8db9870684cb5c96bd9304aad20a3c0f1906

Observation 503398ca-04d1-41a5-9d89-033d4aeeedee · outbound

This paper cites Multiagent simulation on hide and seek games using policy gradient trust region policy optimizati on,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multiagent simulation on hide and seek games using policy gradient trust region policy optimizati on,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.612441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.612441Z digest=sha256:0891f317befd61b50cbb200cd4e5e38466bbf933a1b66ca01cbc128e74742651

Observation 5fb7190a-52f6-4b2d-862e-564294d1e8b4 · outbound

This paper cites Maximum likelihood parameter est i- mation of superimposed chirps using monte carlo importance sampling,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Maximum likelihood parameter est i- mation of superimposed chirps using monte carlo importance sampling,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.656691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.656691Z digest=sha256:efbdea06f41ba273567da59a17c066288c032b052cbfc4f2e11f3773a32917f8

Observation c6c4bdc8-0c82-417b-83d7-04a38bf00b3d · outbound

This paper cites Bias Resilient Multi-Step Off-Policy Goal-Conditioned Reinforcement Learning.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Bias Resilient Multi-Step Off-Policy Goal-Conditioned Reinforcement Learning

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:45.014071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:41.660384Z digest=sha256:5dac85e5b6f71fe67348e7e75778b9dbb3608b4f12d6e84c16605c71b0bc452b

Observation b0e4a082-e547-4759-a054-d9d4f5dbed7d · outbound

This paper cites Improving nuclei segmentation in pathological image via reinforcement lear ning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Improving nuclei segmentation in pathological image via reinforcement lear ning,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.664784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.664784Z digest=sha256:6dab6d06e22083a54e984ab3a555f51588b7d864ddc40aa22dbe0f03ffb4d142

Observation 2577fc1a-3ff1-4832-8b2e-13d2b61e7fb8 · outbound

This paper cites Image captionin g via proximal policy optimization,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Image captionin g via proximal policy optimization,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.668401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.668401Z digest=sha256:e3e0191b2fa56a8865c31516812811f89f90c281e85de656fbc04f84af9b8287

Observation 574c8881-dd04-41e7-a64e-405690549b63 · outbound

This paper cites Sarsa (0) reinforcement learning over fully homomorphic encryption,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Sarsa (0) reinforcement learning over fully homomorphic encryption,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.672212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.672212Z digest=sha256:5dfb518b77267573c06d989975fc0bf221a7733f14080b31f020f188b240e10e

Observation 773a32b8-7ddc-49f0-92aa-e44cfd582c7e · outbound

This paper cites A reinforce ment learning approach to the shepherding task using sarsa,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A reinforce ment learning approach to the shepherding task using sarsa,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.675622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.675622Z digest=sha256:93ef370942007deb39ef088d6bc204cb3c2170354fbfb7f615b820e0a1ccc953

Observation 195fc4e1-35b2-4ff8-b385-1f383f6e210b · outbound

This paper cites Multiagent monte carlo tree search,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multiagent monte carlo tree search,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.679706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.679706Z digest=sha256:9a9b8e2bea36c70d550dac4c2d2dceccf12dcad94b6638dc193aa04ffd429b5b

Observation 1c72748f-5bf7-4b56-a22b-50e2a4fc50c3 · outbound

This paper cites Multi-objective p ri- 74 oritized task scheduler using improved asynchronous advan tage actor critic (a3c) algorithm in multi cloud environment,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multi-objective p ri- 74 oritized task scheduler using improved asynchronous advan tage actor critic (a3c) algorithm in multi cloud environment,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.683465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.683465Z digest=sha256:0c54a34d95dab2ef2bf7c24136f0df3b1f5173e0039eaf1b6c6ce3044f45115c

Observation 78f91738-355a-4b03-be01-7c5e34d0b9a9 · outbound

This paper cites Deep Reinforcement Learning: An Overview.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Deep Reinforcement Learning: An Overview

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.687037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.687037Z digest=sha256:31091957884f058f1cf28fd8c0c658fba093927ad1dd705861d2167549d08f8e

Observation a8ed3ef8-c8ce-470e-bbb7-ef97ed42e58a · outbound

This paper cites Deep reinforcement learning: A brief survey,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Deep reinforcement learning: A brief survey,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.691249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.691249Z digest=sha256:72550742cc3e37188abe66d45fff11e5fe4d412de3074d0a7711db6f5bd5546a

Observation 14536c70-9b13-4810-a2ee-5a55fb58b640 · outbound

This paper cites Deep reinforcement learning: A survey,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Deep reinforcement learning: A survey,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.694843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.694843Z digest=sha256:0576784f9aa78f08ed33554c26a0b81201b5b7346020f33cf615e51135f85430

Observation 88fd47cc-1c28-4429-a5e0-ca1cbbe61e3a · outbound

This paper cites Deep reinforcement learning: a survey ,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Deep reinforcement learning: a survey ,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.698285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.698285Z digest=sha256:b7701ce57fc40fbff8d0a015a8c38a217cb591c3fa8c84546c45fe1a721bbb11

Observation 073ea16d-9109-45cc-bf18-c0bdc80f36f6 · outbound

This paper cites Model-based reinforcement learning: A survey,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Model-based reinforcement learning: A survey,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.701787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.701787Z digest=sha256:68840210775533f58e84f2ce223a4a3e4a09456f17c509a4b4199946fc01f246

Observation 615d674f-3909-4196-8aac-aa1c0d9dafc9 · outbound

This paper cites Survey of model-ba sed reinforcement learning: Applications on robotics,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Survey of model-ba sed reinforcement learning: Applications on robotics,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.705417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.705417Z digest=sha256:8ded76b2abe4ce523c62f53349bbb385a89fb191502d393990964320fa881b69

Observation 1e284d7c-455e-49ed-b343-f0535609f3ae · outbound

This paper cites A survey on model-based reinforcement learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A survey on model-based reinforcement learning,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.709403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.709403Z digest=sha256:27af4f45867c8e96a97f45ca30badf0e70471ada8631be655503abf6c7df00f1

Observation 1b09f367-56d2-41e9-bc39-cd938c5c95b3 · outbound

This paper cites Model-Free Reinforcement Learning for Financial Portfolios: A Brief Survey.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Model-Free Reinforcement Learning for Financial Portfolios: A Brief Survey

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:44.941407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:41.713057Z digest=sha256:1348b294fd2e92719d5c29a5b9b281a7734da93d4eb890d4aa8adc4e5eadf43e

Observation 02d77d18-9eb2-4a49-bb7f-41423963c1c3 · outbound

This paper cites Model-free rei nforce- ment learning from expert demonstrations: a survey,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Model-free rei nforce- ment learning from expert demonstrations: a survey,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.737970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.737970Z digest=sha256:3997ad73e9b8cdca58fb93a940f2bc3187bda844e472a0439b581430a7090941

Observation 8616bc73-b434-45d1-972d-ad7af7ef9dd2 · outbound

This paper cites Reward function design in reinforcement learn- ing,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reward function design in reinforcement learn- ing,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.798347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.798347Z digest=sha256:f0066249ef9b95732e44efc499a826a614059f5f0fd7a25c4b50fd128119b1b5

Observation b09e2d1a-8933-47a2-866a-965eaba5e8a0 · outbound

This paper cites Reinforcement learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement learning,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.824322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.824322Z digest=sha256:636d5f39acac65c90cca218fe49701b7a07f6115d5021a199a4f4fcb5d30c803

Observation 91946beb-58f4-4a1d-9efa-8a12dc85ee38 · outbound

This paper cites Chapter 8 markov decision processes,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Chapter 8 markov decision processes,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.828612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.828612Z digest=sha256:86e0078338d0fb111896f37a2c0b7d47692446238bdcc07797f351b20af3c823

Observation 2f79d2da-4759-4430-bc1c-1296c0e33faf · outbound

This paper cites Markov decision processes,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Markov decision processes,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.833838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.833838Z digest=sha256:f59e75719da852b4edfed4a99d9415b4c155bf34d1a87c74b057d46d78c44515

Observation 287ccb41-34db-4b70-accd-c703f70769dd · outbound

This paper cites Reinforcemen t learning to rank with markov decision process,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcemen t learning to rank with markov decision process,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.838055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.838055Z digest=sha256:8a75f0711f1607ed3551faa9a06491ea5188dcb818db4c4fc9098d609ee7f2a1

Observation 8c66f994-cd51-4de3-88ba-7f0d08021261 · outbound

This paper cites Foundations of Reinforcement Learning and Interactive Decision Making.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Foundations of Reinforcement Learning and Interactive Decision Making

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.842568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.842568Z digest=sha256:d94f3c2d7dd593182d7f4b2bace28456cdaae2fee5f755d6cf64803397da9ced

Observation d8ec4552-1bbb-4089-b923-d46fa189eba3 · outbound

This paper cites On-line q-learning usin g connectionist systems,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges On-line q-learning usin g connectionist systems,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.846569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.846569Z digest=sha256:d41bce386372482b82e36ad668da0d8a0ce56aa6439245f9c60502897c7874a1

Observation 8cd5143c-a9c9-4996-ab4e-a08bc127b181 · outbound

This paper cites Reinforcement learni ng,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement learni ng,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.851128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.851128Z digest=sha256:9b7e208f27aded03dee56b4545a549b92d6b0f3b522b43b736d78d5784da45d1

Observation 03e1ae82-08a3-4f28-9ac0-85624f3570b3 · outbound

This paper cites Introduction to reinforcemen t learn- ing,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Introduction to reinforcemen t learn- ing,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.855088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.855088Z digest=sha256:2ab29162c914a97b75226557dd7686a6e17de2ce92d0c53911c9514d8888950a

Observation 6d5975b1-855d-45ee-9326-d351c22dd43b · outbound

This paper cites Learning to reinforcement learn.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Learning to reinforcement learn

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.858763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.858763Z digest=sha256:9ba5b359aec66faf71436679d6a972197c51e09916b5a35edf85af22bf636a86

Observation 533a964c-dcf0-439e-9125-26ae9fd38d77 · outbound

This paper cites Introduction to reinforcement learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Introduction to reinforcement learning,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.862807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.862807Z digest=sha256:87c084be6e2ecc59ca91390036b9fef85b3526e613cf86ae112dd1baa62b410b

Observation fdf290a8-9233-44fd-945c-4d8a0183d2a7 · outbound

This paper cites The role of exploration in learning control,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges The role of exploration in learning control,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.866440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.866440Z digest=sha256:34729a84a4b9cd8f8183b5aebb1c7fa42f98e581efa3dfe40ed5e8dc89d52159

Observation 6ac14105-68af-4b66-9308-3a8dd8f4a08a · outbound

This paper cites Near-optimal performance for r e- inforcement learning in polynomial time,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Near-optimal performance for r e- inforcement learning in polynomial time,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.870304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.870304Z digest=sha256:fa7b9e432fbe9f654d6ed20e5bf0092a62569ed9eff8c58ed60b41a4b097f7c2

Observation 0124cffc-11a6-4412-a15c-07f2ecde35c7 · outbound

This paper cites Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.873857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.873857Z digest=sha256:ba862c3577b59d707c40d9498d6c935590f55658d4b8b5a83b253b354bd014ae

Observation be46aac7-92e4-40db-8b90-9b02cc9b6995 · outbound

This paper cites Control of expl oitation– exploration meta-parameter in reinforcement learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Control of expl oitation– exploration meta-parameter in reinforcement learning,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.877478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.877478Z digest=sha256:fef26c9dffd363b977364615dad78e7395695c3c49c9be011d615fac43205749

Observation ce9dc37a-66f9-41b5-b675-1aed0ebc5814 · outbound

This paper cites Decou- pling exploration and exploitation in reinforcement learn ing,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Decou- pling exploration and exploitation in reinforcement learn ing,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.881135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.881135Z digest=sha256:7d15dcb9182e4db4c5055ff1c52b6ee315fe63924940af162faa9ce6728a4f53

Observation 5ec1c723-0024-4eca-ad5e-4cdcedb912f5 · outbound

This paper cites Exploration versus exploitation in reinforcement learning: a stochastic control approach.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Exploration versus exploitation in reinforcement learning: a stochastic control approach

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.884736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.884736Z digest=sha256:2494bbc42fc8547f76be6a29af72b7f7c7964beffcce1ca1834c06d46918de97

Observation b3f740ee-b55b-4f10-b2f8-3ba22c60b706 · outbound

This paper cites V al ue iteration networks,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges V al ue iteration networks,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.889185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.889185Z digest=sha256:b4d6f0eb506765ac143d1d6e684e31b133c60dd2eba7561b19425b2c30b13dc0

Observation 4f2e7d88-c890-4c14-9edb-342ab28b372c · outbound

This paper cites Mod el- free monte carlo-like policy evaluation,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Mod el- free monte carlo-like policy evaluation,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.893729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.893729Z digest=sha256:27811b139e4a129676f5e7189d3791d4c8bb6fc081fc59079fe223f43ed4550c

Observation dc4df819-930b-4c91-847a-034450280120 · outbound

This paper cites Importance sampling: a re- view,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Importance sampling: a re- view,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.897807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.897807Z digest=sha256:684cc82512c71d0a3ef0734f3203359547b8c85a7846336bf6be513f4691f4e2

Observation 81a9b7fa-e4b7-4151-947c-7fa6383f2bb9 · outbound

This paper cites Renewal monte carlo: Re - newal theory-based reinforcement learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Renewal monte carlo: Re - newal theory-based reinforcement learning,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:41.920599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:41.920599Z digest=sha256:f4a0a5ebf64faf109ea96c86b7914aa1ba960d800e234a951a7fb97058e7f2f1

Observation 2feec419-82d7-4992-8df0-814904fa6ef0 · outbound

This paper cites Monte carlo off-policy reinforcement learning: A rough set approach,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Monte carlo off-policy reinforcement learning: A rough set approach,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.012851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.012851Z digest=sha256:aa457574b390357f510ee4ca8fcee9788a1c8a29acdd9b565e2e83488cd25a31

Observation 0bf113c6-0427-425f-8b7c-f41622586e93 · outbound

This paper cites Monte Carlo Bayesian Reinforcement Learning.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Monte Carlo Bayesian Reinforcement Learning

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:44.825597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:42.056987Z digest=sha256:6118df13d1af758cea6818f9c915e921215d340126fbf8f956177c41c9836e24

Observation f50b9b3d-96c4-4c95-a94d-8c16532d50ec · outbound

This paper cites Monte-carlo bayesian reinforcement learning using a compact factored representation,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Monte-carlo bayesian reinforcement learning using a compact factored representation,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.066081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.066081Z digest=sha256:4917b9d49161a29ad50a94add6b24e7f11e07eae32cf5fb34d4583c9117b703a

Observation 29e1b20c-93a8-4c33-aa7f-797989b70db0 · outbound

This paper cites Reinforcement learning i n card game environments using monte carlo methods and artificial neural networks,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement learning i n card game environments using monte carlo methods and artificial neural networks,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.069762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.069762Z digest=sha256:6fc6ab03fca1ba3b3772b40361d4c1dbb164f29c71082974090037e6d59723c5

Observation 44036db0-53b3-4263-a066-8401df9a750b · outbound

This paper cites Importance sampling in the monte carlo st udy of sequential tests,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Importance sampling in the monte carlo st udy of sequential tests,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.073852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.073852Z digest=sha256:140e15af4f76c6b8b844369c4409226ad6a58ebb90df086556d4b32e8a84ac02

Observation 574b3acd-0552-4076-ab3a-c33092538abf · outbound

This paper cites A new importance sampling monte carlo method for a flow network reliability problem,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A new importance sampling monte carlo method for a flow network reliability problem,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.077777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.077777Z digest=sha256:ad6b4dab70479d08269f1c94388d140c278ea26079d469284db736485523fcfe

Observation 9e3a1262-1c24-4eab-a8a1-8fc013af41c9 · outbound

This paper cites On the Convergence of the Monte Carlo Exploring Starts Algorithm for Reinforcement Learning.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges On the Convergence of the Monte Carlo Exploring Starts Algorithm for Reinforcement Learning

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:44.808777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:42.081598Z digest=sha256:afc6abe443704f3401be6f2fa96497ff42cb37525822f89edb32627f8846e0f1

Observation 8cf93119-5962-40c7-913a-b3936919748e · outbound

This paper cites Approximation spaces in off- policy monte carlo learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Approximation spaces in off- policy monte carlo learning,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.085522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.085522Z digest=sha256:be94de38c1cb279880169b34c424caaef357d834b1747d20c5ccc3a5b2c4fea8

Observation 5116603b-f5da-491e-b92a-740519675d08 · outbound

This paper cites Td (0)-replay: An efficient model-free pl anning with full replay,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Td (0)-replay: An efficient model-free pl anning with full replay,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.089817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.089817Z digest=sha256:f9068c920c4a78db594bb5fb19c7309ec5230dcc7982f289518e199c3997916e

Observation afb1355a-5900-4d82-ad5c-43880153c8e1 · outbound

This paper cites KnightCap: A chess program that learns by combining TD(lambda) with game-tree search.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges KnightCap: A chess program that learns by combining TD(lambda) with game-tree search

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:44.630515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:42.093681Z digest=sha256:bac8095fa525338b10cfa8385f663ba453adc61cad2cfee054b7f297b8e102c5

Observation aa0108e7-e66c-48cd-bd5f-63c420c8514b · outbound

This paper cites The convergence of td ( λ) for general λ,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges The convergence of td ( λ) for general λ,

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.097232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.097232Z digest=sha256:13b5d9887779ffd296e7ed03c2a0d7c4c61c3d9846e5854f09fcfc41aae9b075

Observation d4ebe09c-c006-43e5-9e63-2a1745a7871e · outbound

This paper cites Td ( λ) converges with proba- bility 1,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Td ( λ) converges with proba- bility 1,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.141362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.141362Z digest=sha256:a3f1c7cd8d4f57a7cb4ea3dc2a626469dc5be720241c1e0043cf5974245caafa

Observation f9af6569-6c32-4f93-95e9-4031662c5fb3 · outbound

This paper cites Two novel on-policy reinforcement learning algorithms based on td ( λ)-methods,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Two novel on-policy reinforcement learning algorithms based on td ( λ)-methods,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.202259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.202259Z digest=sha256:f142ea22b57c7b550d9f602a3e9c265e01bb93de0462f7626280fb01aa41e76b

Observation 9ac4437f-29de-4015-8d19-05f37774fd49 · outbound

This paper cites Multi-step reinforcement learning: A unifying algorithm ,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multi-step reinforcement learning: A unifying algorithm ,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.206433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.206433Z digest=sha256:eab01c579e85cff14fa1a7f70d0faa72ef773899b06759ab1de28bf1e9f3240d

Observation 2129c930-62aa-4954-bd23-4b068535513a · outbound

This paper cites A Lyapunov Theory for Finite-Sample Guarantees of Asynchronous Q-Learning and TD-Learning Variants.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A Lyapunov Theory for Finite-Sample Guarantees of Asynchronous Q-Learning and TD-Learning Variants

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.209908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.209908Z digest=sha256:08106a664a31a5ede82dac75176c5f6726b9938369d70a8244f2b86ba3d370de

Observation 2db04f4f-5aa0-4f32-91a2-551b43cb69b0 · outbound

This paper cites Analysis of Off-Policy Multi-Step TD-Learning with Linear Function Approximation.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Analysis of Off-Policy Multi-Step TD-Learning with Linear Function Approximation

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:44.581034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:42.214346Z digest=sha256:1ae93e1d1ce3830c9a14973d97a12e02208818e3b12c5f038ef96ba93e949ae6

Observation 0915664e-5ee0-4d0c-a53c-ff090fd0a9f3 · outbound

This paper cites A unified view of multi-step temporal differ ence learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A unified view of multi-step temporal differ ence learning,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.218975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.218975Z digest=sha256:47a1acf523518a0c4a71d753612a6d3bdb19c8b901695b0b067d375ab0b4e0fc

Observation b00d1fc1-9421-4244-9edf-8a360d4c09ce · outbound

This paper cites Szepesv´ ari, Algorithms for reinforcement learning.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Szepesv´ ari, Algorithms for reinforcement learning

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.222847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.222847Z digest=sha256:90beab5524e0656501a8855078ee43e8082eb1ec50613adb0899ca66ef09b837

Observation 60dab504-c2f3-4847-91d3-e85557085b9a · outbound

This paper cites Greedy multi-step off-policy reinf orce- ment learning,.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Greedy multi-step off-policy reinf orce- ment learning,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:42.226607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:42.226607Z digest=sha256:56bdfa9f66619ef45d80aa314c5437ae3a4cc0bb016467a3e073dd4c7c2be2d6

Observation b6dfa5e0-1fec-4b0e-bfeb-b119dbe2fbeb · outbound

This paper cites Multi-step Off-policy Learning Without Importance Sampling Ratios.

A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multi-step Off-policy Learning Without Importance Sampling Ratios

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:44.427484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:42.230024Z digest=sha256:c7f7d7711fb400fe08d830cb0ace80249c73a5790a8868e15bca9e9bb3b7e834

Pith citing papers

Observation 905e1a6b-bb43-4ddc-b38b-1c52f647d88a · inbound

Generative AI for Autonomous Driving: A Review cites this paper.

Generative AI for Autonomous Driving: A Review A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:03.035487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:03.035487Z digest=sha256:b6d6f946706d20950006d37b24a119f6fe6bf9acbcc91125f9882cee3703b43a

Observation 8649973a-e863-46fd-9b3b-8a86ef7f221c · inbound

From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications cites this paper.

From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:51.717425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:51.717425Z digest=sha256:d00d276b1e75499bbeee993170ed25c1678d51aa3e17a1322ce186e829500ee6

Observation 329293e8-818e-4d65-b6a0-c902d1e8f14c · inbound

Frequency Resource Management in 6G User-Centric CFmMIMO: A Hybrid Reinforcement Learning and Metaheuristic Approach cites this paper.

Frequency Resource Management in 6G User-Centric CFmMIMO: A Hybrid Reinforcement Learning and Metaheuristic Approach A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:05.518697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:05.518697Z digest=sha256:842b5de8cfbd42c2173087463d2292a10d0f06f1869e9c8b3548c81cfa69514a

Observation 160563b1-a527-48d6-a5cd-619c26980934 · inbound

Train-Once Plan-Anywhere Kinodynamic Motion Planning via Diffusion Trees cites this paper.

Train-Once Plan-Anywhere Kinodynamic Motion Planning via Diffusion Trees A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:40.287382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:40.287382Z digest=sha256:299d08cad9499cb6819ff4d91dfef2ba6721ed4707abfdfdc838a981bc173ea9

Observation a5a6f549-c83d-4625-8d52-b403006dc3fc · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

Reference 161

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:25.452439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:02e1987702014a154c8b93c53089eb7b04421a6b3db4612741e994ed514efcad

Observation 87812c59-05dd-480d-9180-d7f767a6160c · inbound

Reinforcement Learning Assisted Quantum Simulation of Many-Body Excited States and Real-Time Dynamics cites this paper.

Reinforcement Learning Assisted Quantum Simulation of Many-Body Excited States and Real-Time Dynamics A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:28:12.125836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T10:25:40.269360Z digest=sha256:41c3e6c89f1719689a0c4de825a0b623a18bd99d933e7464ba4dbcbf3e1ad73e

Observation 7550a96b-62b7-4431-afc4-48048639118a · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.336842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:a774b70929a806db76b3397d0af6a2c9d3f80b4e54642d79d19cf9cb4a97c37a

Observation 190a7ef9-50d6-479d-b85a-6fb7661b0e58 · inbound

From Bootstrapping to Sequence Modeling: A Unified Generative Framework for Personalized Landing-Page Modeling cites this paper.

From Bootstrapping to Sequence Modeling: A Unified Generative Framework for Personalized Landing-Page Modeling A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:55:51.427942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T02:56:52.303152Z digest=sha256:e902d07a8a0a011a1e262f39b9660a1127f1a127590a297eeeb51efceefa31ab

Observation 5fd793d5-d6fc-4f7b-9c5c-c60bfd655e3a · inbound

Coachable agents for interactive gameplay cites this paper.

Coachable agents for interactive gameplay A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:56.508979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-02T12:52:05.010028Z digest=sha256:186e38c611a324279481401aa27c4f6f1a2cd5d90535fbdf02b08c9e19ceb8b4