Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:39:58.748802Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2502.00288.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:39:58.748802Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 91374a6e-fc4f-4e94-bfff-27c6b7243ffc · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Layer Normalization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef8aae6c-ceae-4080-bf79-253b2d321ea1 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network J., Smith, L., Kostrikov, I., and Levine, S
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation af8f9d16-85ec-47ec-8e05-7cc9d5413d18 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Dota 2 with Large Scale Deep Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f8dd92-cde7-4946-9ec5-23b8f5711018 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Offline rl without off-policy evaluation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a4fae73a-bb69-4f9e-a2e2-7d7a2fae07cf · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0150fed3-b684-4b76-8ff9-00776e071b32 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network A., Salazar, G., Tran, H
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a4c28b55-7c16-4d0e-bbf5-c61969531f0d · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Decision transformer: Reinforcement learning via sequence modeling
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b0d109ad-e96a-4b00-994e-d6727e0c5b54 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Rvs: What is essential for offline RL via supervised learning? In International Conference on Learning Representations, 2022
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fecd2b1-1114-40f1-a0e2-6ef574376969 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Counterfactual multi-agent policy gradients
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db3b8a5c-55dd-40ad-a00f-948b388fe3ef · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cf1abf2-b44b-437b-90f8-0fe4bb516db4 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network and Gu, S
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e974eba-5512-4ef2-9190-856862d12ee9 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Addressing function approximation error in actor-critic methods
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26ad1416-bb4e-4828-8072-f452aba5f1ce · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Reinforcement learning with deep energy-based policies
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e7cc4ae-b972-48a4-be71-8365288c2187 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762c91c2-2744-4237-b726-5ae2ce80c82c · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Modem: Accelerating visual model-based reinforcement learning with demonstrations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bd92b01e-1c3a-4daf-b9b1-0e7201050ce0 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Neural networks: a comprehensive foundation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 12eb4e6e-2e22-44b3-ad2c-b7be9411f06e · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Gaussian Error Linear Units (GELUs)
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a2b416e-2a8f-4220-b613-e2dec9408142 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Deep q-learning from demonstrations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c80aa3c-ffd2-4295-be7a-5301295dc6a5 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network B ayesian design principles for offline-to-online reinforcement learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 493f95b0-9b97-4569-bfea-0995ca08aa92 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network R., and Davison, A
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b47e30cf-10fb-42c2-b51d-1a33e754ab05 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Offline reinforcement learning with implicit q-learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccf644af-9f28-4a5f-b182-1954113cb865 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Conservative q-learning for offline reinforcement learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538adca6-c19f-489d-9158-65b072b067f2 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aaf53e1a-99d2-490a-aa35-c56a737c90fb · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aaf30073-6ddb-40f9-9145-f0178adef183 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network A survey of convolutional neural networks: Analysis, applications, and prospects
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e884384c-00bc-4e93-aeb3-71aef5333199 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Continuous control with deep reinforcement learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 824cb68b-7c81-4977-99f0-f85c9e5c024a · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Discrete Sequential Prediction of Continuous Actions for Deep RL
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4425530-19bc-4373-a556-f48c66a2efe0 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network A., Veness, J., Bellemare, M
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02cadaba-4b3f-4c95-bc9b-b71a628de3d4 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Overcoming exploration in reinforcement learning with demonstrations
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f3d12f85-0098-4afe-b502-e03eb612d060 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 52cea845-30ab-41f4-a614-e4c62b127473 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Learning complex dexterous manipulation with deep reinforcement learning and demonstrations, 2018
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 991c977d-4f6b-4758-b19a-9f829615dd29 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 358af754-ac01-460a-8232-e75feb34d301 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Mastering atari, go, chess and shogi by planning with a learned model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceb203cc-861f-4e3c-80ee-a5fddc944378 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Proximal Policy Optimization Algorithms
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b41d11-1f2b-412f-bcda-c5f7599ee9d2 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network and Abbeel, P
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c10b52f3-7744-4a07-94d5-28a36404e590 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Continuous control with coarse-to-fine reinforcement learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 26766bd2-613b-4f3f-bc8a-b5711008786f · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Solving continuous control via q-learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 04f49af1-4824-4fed-b2dd-03e662db1c07 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Growing Q -networks: S olving continuous control tasks with adaptive control resolution
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5ef7617-735e-416b-94eb-1fa96bc507a2 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Mastering the game of go without human knowledge
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 590cead1-f946-4069-816b-55fde7943841 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network and Agrawal, S
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b2aca0e3-7fa3-4625-9fa3-72622036c302 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Action branching architectures for deep reinforcement learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 76bb7886-7220-4fcb-b150-e0503c4d87ab · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Learning to represent action values as a hypergraph on the action vertices
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c3230589-da4a-48a3-a5e1-f6b7a96708c4 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Deep reinforcement learning with double q-learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 909c3e06-0f3d-4fa0-a2cd-0c5d93d6deb8 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Discriminator-weighted offline imitation learning from suboptimal demonstrations
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f2673bbe-0e28-409c-acef-3c4fb7a9d297 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Hd-cnn: Hierarchical deep convolutional neural networks for large scale visual recognition
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b7835a39-4772-4a97-9a2c-832d712639db · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Mastering visual continuous control: Improved data-augmented reinforcement learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 51b5b9c5-0637-44b1-8770-9a8c6aa53510 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network The surprising effectiveness of ppo in cooperative multi-agent games
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0094e1ab-6b8c-48c7-92db-8cf0d2792c8b · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Policy expansion for bridging offline-to-online reinforcement learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 87a41440-f206-42de-8260-e1bae4ec3c3c · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network Z., Kumar, V., Levine, S., and Finn, C
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c7add5a0-d9ee-428d-8d03-a620df032e05 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network D., Maas, A., Bagnell, J
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4a2aa577-c749-4a7a-bd3d-e5e44f236f26 · outbound
Learning from Suboptimal Data in Continuous Control via Auto-Regressive Soft Q-Network write newline
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.