Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:48:42.230024Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 100 of 265 outbound references and 9 inbound Pith citation observations for arXiv:2411.18892.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:48:42.230024Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:24:03.035487Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T12:56:56.507096Z
100 of 265 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6eb8a5da-10ca-41d7-8182-90656dafc1e0 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ebf840b-024f-4dc7-9430-0acbf5a89157 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges The theory of dynamic programming,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f61ff19-7df0-40e9-8adc-506dd7dd922a · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Introduction to Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2f3bd12-48e7-4497-8a14-11b9db7584c6 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Learning to predict by the methods of tempo ral differences,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d16a82-da10-4a31-ba3e-91b6e7b6115c · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Learning from delayed rewards,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1835866b-d3c8-43dc-ad65-6501660d4f25 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Rein- forcement learning: A survey,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b199c9b-d96b-4a68-b8fc-c8ada6694e3b · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Finite-time a nalysis of the multiarmed bandit problem,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92fa5ddc-472a-4e16-b03e-49b3a584e10c · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Human-level control through deep rein- forcement learning,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf596b4-a013-4e06-9002-5f38eaab5471 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Mastering the game of go with deep neural networks and tree search,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 051f351e-febc-4267-990d-4f0e1beca322 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement learning in game industry—review, prospec ts and challenges,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f996e11a-f84b-463f-8582-921a1199bb80 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Pgx: Hardware-accelerated parallel game simulators for reinforcement learning,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a89af55-ed63-4ac8-a6f7-459321266e7d · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 195557f6-2a4a-427b-8a69-8a4830fa2488 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Pursuit-evasion gam e strategy of usv based on deep reinforcement learning in comp lex multi-obstacle environment,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93ef6eb5-43f9-4b0d-8256-3714a2072a3e · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Residual skill policies: Learning an adaptable skill-bas ed ac- tion space for reinforcement learning for robotics,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84bd48e5-7633-4d86-92dc-7376a1de70ca · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement l earning in robotics: A survey,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48dbe19f-733b-4af5-8ea1-01ad2c0102b6 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Intrinsically motivated multi- goal rein- forcement learning using robotics environment integrated with openai gym,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 466971ec-fac4-4945-82c0-b5155e8b8af2 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges i-sim2real: Reinforcement learning of robotic p olicies in tight human-robot interaction loops,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9567e94-9df7-474e-96fe-e5267b9dab6a · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multi-agent broad reinforcement learning for intelligent traffic light control,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88715876-e49a-47ed-9eab-5b318b5181e4 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Intelligent vehicle pedestrian light (ivp l): A deep reinforcement learning approach for traffic signal con trol,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f5cca0-4302-4472-a14f-3641b671d86c · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Swarm learning- based dynamic optimal management for traffic congestion in 6g-driven intelligent transportation system,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e2b73d-c105-4c16-8093-91f40cdde1ee · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Deep multi-agent reinforcement learning for highway on-ramp merging in mixed traffic,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b078d56-7a45-49b8-b13e-609b3bf6d650 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Maximizing group-based vehicle communications and fairn ess: A reinforcement learning approach,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b33735d3-337e-489c-8670-3661bbe81744 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A real-world application of markov chain monte carlo method for bayesian trajectory control of a robotic manipulator,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92b5b29-4294-4a73-919c-a87f8a547f47 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Obstacle avoidance method based on double dqn for agricultural robot s,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29cf150b-2f72-4e26-806a-a3e1716ba69f · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Adaptive importance sampling monte carl o simulation of rare transition events,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24fc822a-55e1-485b-b398-24aae442a5bb · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Dou- ble dqn-based coevolution for green distributed heterogen eous hybrid flowshop scheduling with multiple priorities of jobs ,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f40a9ad3-784b-4de4-ad6d-dd14c1c237d4 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges An intel li- gent train control approach based on the monte carlo reinfor ce- ment learning algorithm,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc4987e-dfbe-464b-ad2e-fcbdd4390a9a · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Green computing in heterogeneous internet of things: Optimizing energy alloc ation using sarsa-based reinforcement learning,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c3b61f3-78de-4142-bff9-b70591534c47 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Power allocation in multi-ag ent networks via dueling dqn approach,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c26950-5197-4e01-96c9-182994abdf5a · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Efficient importa nce sampling for monte carlo simulation of multicast networks,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a5f9aa-ff6c-4dc3-b266-e99e3e8d0a99 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement-learning-based ener gy storage system operation strategies to manage wind power forecast uncertainty,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a18eb9-d08a-4504-ac1b-0827d24807ee · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Learning-based power delay profile estimatio n for 5g nr via advantage actor-critic (a2c),
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f29e2377-b74d-49d2-a97f-f4a5550e59f0 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Per-decision Multi-step Temporal Difference Learning with Control Variates
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e05ce389-f4a6-4f68-916b-9ffad6819c7e · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A motion control method for agent based on dyna-q algorithm,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ee22d97-6343-4a12-930c-85f740d0aa62 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Temporal-difference netwo rks,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f6e22e9-7f97-475e-b4ee-56986fd4a500 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Realizing undelayed n-step td prediction w ith neural networks,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d5d8f5-651c-4722-ab97-d71b3138e40f · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Monte Carlo Tree Search for Asymmetric Trees
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 75a70ff1-bb79-4908-90c1-138dad307666 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges An overview of q-learni ng based energy efficient power allocation in wban (q-eepa),
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6642a15e-4047-44a8-9054-4edc75a36e4f · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Intelligent control of a quadrotor with proximal pol icy optimization reinforcement learning,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed47a431-7ec0-49de-8da3-42ee5acd76d4 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Automated por t- folio rebalancing using q-learning,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994b04fb-8929-41c8-ad56-53da4699eb7f · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Prioritized sweeping reinfor cement learning based routing for manets,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 503398ca-04d1-41a5-9d89-033d4aeeedee · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multiagent simulation on hide and seek games using policy gradient trust region policy optimizati on,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb7190a-52f6-4b2d-862e-564294d1e8b4 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Maximum likelihood parameter est i- mation of superimposed chirps using monte carlo importance sampling,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c4bdc8-0c82-417b-83d7-04a38bf00b3d · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Bias Resilient Multi-Step Off-Policy Goal-Conditioned Reinforcement Learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b0e4a082-e547-4759-a054-d9d4f5dbed7d · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Improving nuclei segmentation in pathological image via reinforcement lear ning,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2577fc1a-3ff1-4832-8b2e-13d2b61e7fb8 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Image captionin g via proximal policy optimization,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 574c8881-dd04-41e7-a64e-405690549b63 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Sarsa (0) reinforcement learning over fully homomorphic encryption,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 773a32b8-7ddc-49f0-92aa-e44cfd582c7e · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A reinforce ment learning approach to the shepherding task using sarsa,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 195fc4e1-35b2-4ff8-b385-1f383f6e210b · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multiagent monte carlo tree search,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c72748f-5bf7-4b56-a22b-50e2a4fc50c3 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multi-objective p ri- 74 oritized task scheduler using improved asynchronous advan tage actor critic (a3c) algorithm in multi cloud environment,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f91738-355a-4b03-be01-7c5e34d0b9a9 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Deep Reinforcement Learning: An Overview
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8ed3ef8-c8ce-470e-bbb7-ef97ed42e58a · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Deep reinforcement learning: A brief survey,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14536c70-9b13-4810-a2ee-5a55fb58b640 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Deep reinforcement learning: A survey,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88fd47cc-1c28-4429-a5e0-ca1cbbe61e3a · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Deep reinforcement learning: a survey ,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 073ea16d-9109-45cc-bf18-c0bdc80f36f6 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Model-based reinforcement learning: A survey,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615d674f-3909-4196-8aac-aa1c0d9dafc9 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Survey of model-ba sed reinforcement learning: Applications on robotics,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e284d7c-455e-49ed-b343-f0535609f3ae · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A survey on model-based reinforcement learning,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b09f367-56d2-41e9-bc39-cd938c5c95b3 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Model-Free Reinforcement Learning for Financial Portfolios: A Brief Survey
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 02d77d18-9eb2-4a49-bb7f-41423963c1c3 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Model-free rei nforce- ment learning from expert demonstrations: a survey,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8616bc73-b434-45d1-972d-ad7af7ef9dd2 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reward function design in reinforcement learn- ing,
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b09e2d1a-8933-47a2-866a-965eaba5e8a0 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement learning,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91946beb-58f4-4a1d-9efa-8a12dc85ee38 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Chapter 8 markov decision processes,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f79d2da-4759-4430-bc1c-1296c0e33faf · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Markov decision processes,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 287ccb41-34db-4b70-accd-c703f70769dd · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcemen t learning to rank with markov decision process,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c66f994-cd51-4de3-88ba-7f0d08021261 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Foundations of Reinforcement Learning and Interactive Decision Making
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8ec4552-1bbb-4089-b923-d46fa189eba3 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges On-line q-learning usin g connectionist systems,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd5143c-a9c9-4996-ab4e-a08bc127b181 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement learni ng,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03e1ae82-08a3-4f28-9ac0-85624f3570b3 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Introduction to reinforcemen t learn- ing,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5975b1-855d-45ee-9326-d351c22dd43b · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Learning to reinforcement learn
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 533a964c-dcf0-439e-9125-26ae9fd38d77 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Introduction to reinforcement learning,
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdf290a8-9233-44fd-945c-4d8a0183d2a7 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges The role of exploration in learning control,
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac14105-68af-4b66-9308-3a8dd8f4a08a · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Near-optimal performance for r e- inforcement learning in polynomial time,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0124cffc-11a6-4412-a15c-07f2ecde35c7 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be46aac7-92e4-40db-8b90-9b02cc9b6995 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Control of expl oitation– exploration meta-parameter in reinforcement learning,
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce9dc37a-66f9-41b5-b675-1aed0ebc5814 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Decou- pling exploration and exploitation in reinforcement learn ing,
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ec1c723-0024-4eca-ad5e-4cdcedb912f5 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Exploration versus exploitation in reinforcement learning: a stochastic control approach
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3f740ee-b55b-4f10-b2f8-3ba22c60b706 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges V al ue iteration networks,
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f2e7d88-c890-4c14-9edb-342ab28b372c · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Mod el- free monte carlo-like policy evaluation,
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc4df819-930b-4c91-847a-034450280120 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Importance sampling: a re- view,
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a9b7fa-e4b7-4151-947c-7fa6383f2bb9 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Renewal monte carlo: Re - newal theory-based reinforcement learning,
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2feec419-82d7-4992-8df0-814904fa6ef0 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Monte carlo off-policy reinforcement learning: A rough set approach,
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bf113c6-0427-425f-8b7c-f41622586e93 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Monte Carlo Bayesian Reinforcement Learning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f50b9b3d-96c4-4c95-a94d-8c16532d50ec · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Monte-carlo bayesian reinforcement learning using a compact factored representation,
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29e1b20c-93a8-4c33-aa7f-797989b70db0 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Reinforcement learning i n card game environments using monte carlo methods and artificial neural networks,
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44036db0-53b3-4263-a066-8401df9a750b · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Importance sampling in the monte carlo st udy of sequential tests,
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 574b3acd-0552-4076-ab3a-c33092538abf · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A new importance sampling monte carlo method for a flow network reliability problem,
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e3a1262-1c24-4eab-a8a1-8fc013af41c9 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges On the Convergence of the Monte Carlo Exploring Starts Algorithm for Reinforcement Learning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8cf93119-5962-40c7-913a-b3936919748e · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Approximation spaces in off- policy monte carlo learning,
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5116603b-f5da-491e-b92a-740519675d08 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Td (0)-replay: An efficient model-free pl anning with full replay,
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb1355a-5900-4d82-ad5c-43880153c8e1 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges KnightCap: A chess program that learns by combining TD(lambda) with game-tree search
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa0108e7-e66c-48cd-bd5f-63c420c8514b · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges The convergence of td ( λ) for general λ,
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4ebe09c-c006-43e5-9e63-2a1745a7871e · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Td ( λ) converges with proba- bility 1,
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9af6569-6c32-4f93-95e9-4031662c5fb3 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Two novel on-policy reinforcement learning algorithms based on td ( λ)-methods,
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ac4437f-29de-4015-8d19-05f37774fd49 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multi-step reinforcement learning: A unifying algorithm ,
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2129c930-62aa-4954-bd23-4b068535513a · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A Lyapunov Theory for Finite-Sample Guarantees of Asynchronous Q-Learning and TD-Learning Variants
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2db04f4f-5aa0-4f32-91a2-551b43cb69b0 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Analysis of Off-Policy Multi-Step TD-Learning with Linear Function Approximation
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0915664e-5ee0-4d0c-a53c-ff090fd0a9f3 · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges A unified view of multi-step temporal differ ence learning,
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00d1fc1-9421-4244-9edf-8a360d4c09ce · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Szepesv´ ari, Algorithms for reinforcement learning
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60dab504-c2f3-4847-91d3-e85557085b9a · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Greedy multi-step off-policy reinf orce- ment learning,
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6dfa5e0-1fec-4b0e-bfeb-b119dbe2fbeb · outbound
A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges Multi-step Off-policy Learning Without Importance Sampling Ratios
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 905e1a6b-bb43-4ddc-b38b-1c52f647d88a · inbound
Generative AI for Autonomous Driving: A Review A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8649973a-e863-46fd-9b3b-8a86ef7f221c · inbound
From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 329293e8-818e-4d65-b6a0-c902d1e8f14c · inbound
Frequency Resource Management in 6G User-Centric CFmMIMO: A Hybrid Reinforcement Learning and Metaheuristic Approach A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 160563b1-a527-48d6-a5cd-619c26980934 · inbound
Train-Once Plan-Anywhere Kinodynamic Motion Planning via Diffusion Trees A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5a6f549-c83d-4625-8d52-b403006dc3fc · inbound
A Survey of Reinforcement Learning for Large Reasoning Models A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
Reference 161
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 87812c59-05dd-480d-9180-d7f767a6160c · inbound
Reinforcement Learning Assisted Quantum Simulation of Many-Body Excited States and Real-Time Dynamics A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7550a96b-62b7-4431-afc4-48048639118a · inbound
Trust Region On-Policy Distillation A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 190a7ef9-50d6-479d-b85a-6fb7661b0e58 · inbound
From Bootstrapping to Sequence Modeling: A Unified Generative Framework for Personalized Landing-Page Modeling A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5fd793d5-d6fc-4f7b-9c5c-c60bfd655e3a · inbound
Coachable agents for interactive gameplay A Comprehensive Survey of Reinforcement Learning: From Algorithms to Practical Challenges
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.