{"work":{"id":"41a65444-c819-4303-a1f1-b075aa86d40c","openalex_id":"https://openalex.org/W4386291772","doi":"10.1137/22m1480409","arxiv_id":"1509.02971","raw_key":null,"title":"Continuous control with deep reinforcement learning","authors":null,"authors_text":"Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa","year":2015,"venue":"cs.LG","abstract":"We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algorithm robustly solves more than 20 simulated physics tasks, including classic problems such as cartpole swing-up, dexterous manipulation, legged locomotion and car driving. Our algorithm is able to find policies whose performance is competitive with those found by a planning algorithm with full access to the dynamics of the domain and its derivatives. We further demonstrate that for many of the tasks the algorithm can learn policies end-to-end: directly from raw pixel inputs.","external_url":"https://arxiv.org/abs/1509.02971","cited_by_count":1,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"1509.02971","created_at":"2026-05-09T05:55:31.013412+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Continuous control with deep reinforcement learning","render_title":"Continuous control with deep reinforcement learning"},"hub":{"state":{"work_id":"41a65444-c819-4303-a1f1-b075aa86d40c","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":156,"external_cited_by_count":1,"distinct_field_count":18,"first_pith_cited_at":"2015-06-08T11:12:48+00:00","last_pith_cited_at":"2026-07-09T17:09:20+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-24T03:29:20.334898+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":11},{"context_role":"method","n":6}],"polarity_counts":[{"context_polarity":"background","n":11},{"context_polarity":"use_method","n":5},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Continuous control with deep reinforcement learning","claims":[{"claim_text":"We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algorithm robustly solves more than 20 simulated physics tasks, including classic problems such as cartpole swing-up, dexterous manipulation, legged locomotion and car driving. Our algorithm is able to find policies whose performance is competitive with those found by a planning algo","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"LQ =E (x,a1:C ,x′)∼B h ˆQ−Q ψ(x,a 1:C) \u00012i , ˆQ= CX t′=1 γt′−1rt′ +γ CEa′∼πθ h Qψ′(x′,a ′) i . (3) where the input state isx= (z rl,s p), ands p denotes the proprioceptive state information,z rl(s)denotes the RL token extracted for states;x ′ denotes the next input state;a ′ ∼π θ denotes taking a sample from the RL policy. In practice, we follow TD3 [19] andψ ′ are the parameters of the target network. Training the RL Policy.Our actor networkπ θ(·|x, ˜a1:C) produces a Gaussian action distributio","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"Continuous deep Q-learning with model- based acceleration. In Maria Florina Balcan and Kil- ian Q. Weinberger, editors,Proceedings of The 33rd International Conference on Machine Learning, vol- ume 48 ofProceedings of Machine Learning Research, pages 2829-2838, New York, New York, USA, 20-22 Jun 2016. PMLR. URLhttps://proceedings.mlr. press/v48/gu16.html. [45] Richard S Sutton, Andrew G Barto, et al.Reinforcement learning: An introduction, volume 1. MIT press Cam- bridge, 2020. [46] Peter D Lax.","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"reinforcement learning methods. Specifically, the DDPG algo- rithm is an off-policy actor-critic approach that is well-suited for handling continuous action spaces. However, the standard DDPG framework, which relies on fully connected deep neural networks (DNNs), is inadequate for modeling the temporal dynamics present in environments such as user mobility [31], [32]. To overcome this limitation, we enhance the DDPG framework by integrating LSTM networks, enabling the model to effectively captur","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"A high-level comparison of the main methodological paradigms discussed above is shown in Fig. 2. B. Reinforcement Learning for Autonomous Parking Compared to imitation learning, which relies on expert demonstrations, reinforcement learning offers better scalability and adaptability for autonomous parking. Common reinforce- ment learning training methods include DDPG [22], PPO [23], SAC [24], and TD3 [25]. Recently, RL has demonstrated IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS 3 Fig","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"and TD-MPC2 [22] have demonstrated strong performance in vision-based domains by planning in learned latent spaces. However, learning accurate dynamics models and performing repeated planning procedures significantly increases per-step training cost [49], often limiting scalability in simulation settings where wall-clock efficiency is critical. Off-Policy Model-Free RL.Model-free off-policy algorithms such as DDPG [44], TD3 [15], and SAC [20] learn policies and value functions directly from repl","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"conventional model-based control approaches, RL can improve performance from data and experience, which makes it especially appealing for systems with incomplete models or complex dynamics. Recent progress in deep reinforcement learning has further demonstrated the capability of RL to address high- dimensional and nonlinear decision problems that are difficult to solve using classical approaches alone [8, 9]. Despite these advances, the direct use of RL in safety-critical control applications re","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Continuous control with deep reinforcement learning because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":6,"context_role":"method"}]},"error":null,"updated_at":"2026-05-25T15:06:04.426276+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"39be4b4d-c785-4604-b80e-2d6b37ab1e6c","orcid":null,"display_name":"Timothy P. Lillicrap"},{"id":"ef649341-56cb-45e6-ac38-ff23ba6d6d8b","orcid":null,"display_name":"Jonathan J. Hunt"},{"id":"bdc893b0-d3ba-4495-83f9-de9faaafd042","orcid":null,"display_name":"Alexander Pritzel"},{"id":"d6ef4bbc-eb23-4c94-8884-f11d3b7ad177","orcid":null,"display_name":"Nicolas Heess"},{"id":"fbc77fc3-ef25-4646-a26b-2e6d19ea5f31","orcid":null,"display_name":"Tom Erez"},{"id":"4ea13bb8-db03-4c10-8a97-2c2cfa374668","orcid":null,"display_name":"Yuval Tassa"}]},"error":null,"updated_at":"2026-05-25T15:06:04.421746+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T17:49:05.394503+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":18},{"title":"Playing Atari with Deep Reinforcement Learning","work_id":"736a8ddf-e365-4940-ad58-4699fddedb86","shared_citers":10},{"title":"DeepMind Control Suite","work_id":"54294ef0-c651-4d5a-a72b-f85a88329a71","shared_citers":8},{"title":"OpenAI Gym","work_id":"6af98f3f-f074-41ae-a689-7dd7b4b8efde","shared_citers":8},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":7},{"title":"Addressing function approximation error in actor-critic methods","work_id":"129bee39-1830-4ff5-a7c3-f8ecae60f370","shared_citers":6},{"title":"Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor","work_id":"6674e5db-4e1c-49c0-b598-c108a0ecadb6","shared_citers":6},{"title":"Deep reinforcement learning and the deadly triad","work_id":"de214ead-4cb0-4abd-be3d-ae3389f55e9b","shared_citers":5},{"title":"Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems","work_id":"597b6f46-d60f-451f-8f34-7d32876a9014","shared_citers":5},{"title":"Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning","work_id":"ab561983-ab59-4f04-a11e-a467ddde4848","shared_citers":4},{"title":"Behavior Regularized Offline Reinforcement Learning","work_id":"95ad303d-9555-46ad-9a95-3bbea22ed9fb","shared_citers":4},{"title":"Mastering Diverse Domains through World Models","work_id":"6aeb260f-8c7c-4f9c-b98b-067cd7c59acd","shared_citers":4},{"title":"Soft Actor-Critic Algorithms and Applications","work_id":"bb49c9fb-03b2-4226-9edb-50186b8193e4","shared_citers":4},{"title":"//arxiv.org/abs/1811.04551","work_id":"146fc4a4-6db2-43a9-a57f-4dd133c0d315","shared_citers":3},{"title":"arXiv preprint arXiv:1907.02057 , year=","work_id":"14cc71c0-924c-4734-8464-145da2d06666","shared_citers":3},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":3},{"title":"AWAC: Accelerating Online Reinforcement Learning with Offline Datasets","work_id":"f0a11265-1acf-4ffc-a822-08bd04b6bddf","shared_citers":3},{"title":"D4RL: Datasets for Deep Data-Driven Reinforcement Learning","work_id":"47082e4e-a4a5-418b-bf4f-4667355065fc","shared_citers":3},{"title":"E., and Levine, S","work_id":"3570547f-d4c5-4808-a945-f27a73bb7d90","shared_citers":3},{"title":"Espeholt, H","work_id":"9bbfa31b-454d-4174-923d-06e74e0c6cab","shared_citers":3},{"title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","work_id":"1fe8c7c8-aff7-4b94-9096-e549d7e60789","shared_citers":3},{"title":"G., Bellemare, M","work_id":"2b8d4fd3-3479-478a-bc44-f6d493e0616c","shared_citers":3},{"title":"Gymnasium: A Standard Interface for Reinforcement Learning Environments","work_id":"5382dc1c-a327-49b9-afda-4794d5847698","shared_citers":3},{"title":"Kaiser, M","work_id":"edc1a23e-c421-4569-ab9e-83b204eeb0fa","shared_citers":3}],"time_series":[{"n":1,"year":2015},{"n":3,"year":2018},{"n":2,"year":2019},{"n":1,"year":2020},{"n":1,"year":2021},{"n":3,"year":2023},{"n":29,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T17:48:42.912749+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T17:48:46.065456+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Continuous control with deep reinforcement learning","claims":[{"claim_text":"We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algorithm robustly solves more than 20 simulated physics tasks, including classic problems such as cartpole swing-up, dexterous manipulation, legged locomotion and car driving. Our algorithm is able to find policies whose performance is competitive with those found by a planning algo","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"LQ =E (x,a1:C ,x′)∼B h ˆQ−Q ψ(x,a 1:C) \u00012i , ˆQ= CX t′=1 γt′−1rt′ +γ CEa′∼πθ h Qψ′(x′,a ′) i . (3) where the input state isx= (z rl,s p), ands p denotes the proprioceptive state information,z rl(s)denotes the RL token extracted for states;x ′ denotes the next input state;a ′ ∼π θ denotes taking a sample from the RL policy. In practice, we follow TD3 [19] andψ ′ are the parameters of the target network. Training the RL Policy.Our actor networkπ θ(·|x, ˜a1:C) produces a Gaussian action distributio","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"Continuous deep Q-learning with model- based acceleration. In Maria Florina Balcan and Kil- ian Q. Weinberger, editors,Proceedings of The 33rd International Conference on Machine Learning, vol- ume 48 ofProceedings of Machine Learning Research, pages 2829-2838, New York, New York, USA, 20-22 Jun 2016. PMLR. URLhttps://proceedings.mlr. press/v48/gu16.html. [45] Richard S Sutton, Andrew G Barto, et al.Reinforcement learning: An introduction, volume 1. MIT press Cam- bridge, 2020. [46] Peter D Lax.","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"reinforcement learning methods. Specifically, the DDPG algo- rithm is an off-policy actor-critic approach that is well-suited for handling continuous action spaces. However, the standard DDPG framework, which relies on fully connected deep neural networks (DNNs), is inadequate for modeling the temporal dynamics present in environments such as user mobility [31], [32]. To overcome this limitation, we enhance the DDPG framework by integrating LSTM networks, enabling the model to effectively captur","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"A high-level comparison of the main methodological paradigms discussed above is shown in Fig. 2. B. Reinforcement Learning for Autonomous Parking Compared to imitation learning, which relies on expert demonstrations, reinforcement learning offers better scalability and adaptability for autonomous parking. Common reinforce- ment learning training methods include DDPG [22], PPO [23], SAC [24], and TD3 [25]. Recently, RL has demonstrated IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS 3 Fig","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"and TD-MPC2 [22] have demonstrated strong performance in vision-based domains by planning in learned latent spaces. However, learning accurate dynamics models and performing repeated planning procedures significantly increases per-step training cost [49], often limiting scalability in simulation settings where wall-clock efficiency is critical. Off-Policy Model-Free RL.Model-free off-policy algorithms such as DDPG [44], TD3 [15], and SAC [20] learn policies and value functions directly from repl","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"conventional model-based control approaches, RL can improve performance from data and experience, which makes it especially appealing for systems with incomplete models or complex dynamics. Recent progress in deep reinforcement learning has further demonstrated the capability of RL to address high- dimensional and nonlinear decision problems that are difficult to solve using classical approaches alone [8, 9]. Despite these advances, the direct use of RL in safety-critical control applications re","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Continuous control with deep reinforcement learning because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (9 contexts).","role_counts":[{"n":9,"context_role":"background"},{"n":6,"context_role":"method"}]},"error":null,"updated_at":"2026-05-25T15:06:04.432485+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Continuous control with deep reinforcement learning","claims":[{"claim_text":"We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algorithm robustly solves more than 20 simulated physics tasks, including classic problems such as cartpole swing-up, dexterous manipulation, legged locomotion and car driving. Our algorithm is able to find policies whose performance is competitive with those found by a planning algo","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Continuous control with deep reinforcement learning because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T17:48:49.064137+00:00"}},"summary":{"title":"Continuous control with deep reinforcement learning","claims":[{"claim_text":"We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algorithm robustly solves more than 20 simulated physics tasks, including classic problems such as cartpole swing-up, dexterous manipulation, legged locomotion and car driving. Our algorithm is able to find policies whose performance is competitive with those found by a planning algo","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Continuous control with deep reinforcement learning because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":18},{"title":"Playing Atari with Deep Reinforcement Learning","work_id":"736a8ddf-e365-4940-ad58-4699fddedb86","shared_citers":10},{"title":"DeepMind Control Suite","work_id":"54294ef0-c651-4d5a-a72b-f85a88329a71","shared_citers":8},{"title":"OpenAI Gym","work_id":"6af98f3f-f074-41ae-a689-7dd7b4b8efde","shared_citers":8},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":7},{"title":"Addressing function approximation error in actor-critic methods","work_id":"129bee39-1830-4ff5-a7c3-f8ecae60f370","shared_citers":6},{"title":"Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor","work_id":"6674e5db-4e1c-49c0-b598-c108a0ecadb6","shared_citers":6},{"title":"Deep reinforcement learning and the deadly triad","work_id":"de214ead-4cb0-4abd-be3d-ae3389f55e9b","shared_citers":5},{"title":"Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems","work_id":"597b6f46-d60f-451f-8f34-7d32876a9014","shared_citers":5},{"title":"Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning","work_id":"ab561983-ab59-4f04-a11e-a467ddde4848","shared_citers":4},{"title":"Behavior Regularized Offline Reinforcement Learning","work_id":"95ad303d-9555-46ad-9a95-3bbea22ed9fb","shared_citers":4},{"title":"Mastering Diverse Domains through World Models","work_id":"6aeb260f-8c7c-4f9c-b98b-067cd7c59acd","shared_citers":4},{"title":"Soft Actor-Critic Algorithms and Applications","work_id":"bb49c9fb-03b2-4226-9edb-50186b8193e4","shared_citers":4},{"title":"//arxiv.org/abs/1811.04551","work_id":"146fc4a4-6db2-43a9-a57f-4dd133c0d315","shared_citers":3},{"title":"arXiv preprint arXiv:1907.02057 , year=","work_id":"14cc71c0-924c-4734-8464-145da2d06666","shared_citers":3},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":3},{"title":"AWAC: Accelerating Online Reinforcement Learning with Offline Datasets","work_id":"f0a11265-1acf-4ffc-a822-08bd04b6bddf","shared_citers":3},{"title":"D4RL: Datasets for Deep Data-Driven Reinforcement Learning","work_id":"47082e4e-a4a5-418b-bf4f-4667355065fc","shared_citers":3},{"title":"E., and Levine, S","work_id":"3570547f-d4c5-4808-a945-f27a73bb7d90","shared_citers":3},{"title":"Espeholt, H","work_id":"9bbfa31b-454d-4174-923d-06e74e0c6cab","shared_citers":3},{"title":"Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation","work_id":"1fe8c7c8-aff7-4b94-9096-e549d7e60789","shared_citers":3},{"title":"G., Bellemare, M","work_id":"2b8d4fd3-3479-478a-bc44-f6d493e0616c","shared_citers":3},{"title":"Gymnasium: A Standard Interface for Reinforcement Learning Environments","work_id":"5382dc1c-a327-49b9-afda-4794d5847698","shared_citers":3},{"title":"Kaiser, M","work_id":"edc1a23e-c421-4569-ab9e-83b204eeb0fa","shared_citers":3}],"time_series":[{"n":1,"year":2015},{"n":3,"year":2018},{"n":2,"year":2019},{"n":1,"year":2020},{"n":1,"year":2021},{"n":3,"year":2023},{"n":29,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"bdc893b0-d3ba-4495-83f9-de9faaafd042","orcid":null,"display_name":"Alexander Pritzel","source":"manual","import_confidence":0.72},{"id":"ef649341-56cb-45e6-ac38-ff23ba6d6d8b","orcid":null,"display_name":"Jonathan J. Hunt","source":"manual","import_confidence":0.72},{"id":"d6ef4bbc-eb23-4c94-8884-f11d3b7ad177","orcid":null,"display_name":"Nicolas Heess","source":"manual","import_confidence":0.72},{"id":"39be4b4d-c785-4604-b80e-2d6b37ab1e6c","orcid":null,"display_name":"Timothy P. Lillicrap","source":"manual","import_confidence":0.72},{"id":"fbc77fc3-ef25-4646-a26b-2e6d19ea5f31","orcid":null,"display_name":"Tom Erez","source":"manual","import_confidence":0.72},{"id":"4ea13bb8-db03-4c10-8a97-2c2cfa374668","orcid":null,"display_name":"Yuval Tassa","source":"manual","import_confidence":0.72}]}}