Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T21:33:40.229376Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2510.02590.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T21:33:40.229376Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T21:42:30.436671Z
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d57b61d9-ac4c-46e7-88e6-00379eac8373 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Dopamine: A Research Framework for Deep Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f473e237-c902-4191-878e-2294cefac0a7 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Soft Actor-Critic Algorithms and Applications
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a15785bd-3cd3-434a-b66b-5b95ea1f79b5 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Hyperspherical Normalization for Scalable Deep Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f4cc08e2-a87f-432b-ab5f-013e68fd6365 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b16a3702-e63b-4c3e-9d86-62f537a9605a · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Playing Atari with Deep Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9573638c-3916-440d-b4bd-c68c2c3a6b68 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b41db84b-b6d6-4ed8-9a55-daf50f968a25 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72b104ca-c66f-4952-b4b4-a77ff7925207 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f64fb178-d907-4734-a8d3-54012a4eafb9 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning DeepMind Control Suite
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cd31e76f-1ba0-4de1-84ba-6676815879c5 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Deep Reinforcement Learning and the Deadly Triad
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e68fcd9a-b01c-4b3e-b8b9-c747bbedeae8 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Under Review
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 43ab906f-195e-4aae-af77-35b800c8c918 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning The second inequality holds because theminoperator is also a non-expansion
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a23c9dd6-70b0-4424-90ca-64f62a555cb1 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning 4, SAC is adapted to use a single Q-function critic, following the approach taken in Simba (Lee et al
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 189f580a-e9d7-49db-a419-a87d2ee50837 · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Under Review
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c1a29729-c6ac-49a7-bafd-4007fc18ad7b · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning C.4 ONLINERLANDCONTINUOUSCONTROL For our continuous-control experiments with online reinforcement learning, we adopt SimbaV1 and SimbaV2
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a3a20ee3-2558-4d09-ab56-ae8a34af91af · outbound
Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Identical values are merged
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 81145edf-9c8f-4fdb-b245-a69fa66192cc · inbound
Stable Deep Reinforcement Learning via Isotropic Gaussian Representations Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.