Pith. sign in

Paper Citation Record · LEDGER

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2510.02590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.02590 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T21:33:40.229376Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:42:30.436671Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact9
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d57b61d9-ac4c-46e7-88e6-00379eac8373 · outbound

This paper cites Dopamine: A Research Framework for Deep Reinforcement Learning.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Dopamine: A Research Framework for Deep Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T21:34:22.294809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:85c083d7e17370ab02ed3aa1548592305fdd577c2597f65526469f1196137de0

Observation f473e237-c902-4191-878e-2294cefac0a7 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:34:22.323552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:c0fa922a165dc307d022a4a09a86de9853a158b0190e08ffae5f78ad5dcd1673

Observation a15785bd-3cd3-434a-b66b-5b95ea1f79b5 · outbound

This paper cites Hyperspherical Normalization for Scalable Deep Reinforcement Learning.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Hyperspherical Normalization for Scalable Deep Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:34:22.302146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:7775b1694eacff8629930576064bf57eebbcb587b9b6df906f3af03b658c1283

Observation f4cc08e2-a87f-432b-ab5f-013e68fd6365 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:34:22.289168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:6e5f561d4f86ee55e575fe12341d0403f7f7f5c30115fd8f307031b3aaab4c9f

Observation b16a3702-e63b-4c3e-9d86-62f537a9605a · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:34:22.290821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:b83d09af64cfddb6026fe0553e115d75fe672053b30a6c2a17887a6a4b24f7cc

Observation 9573638c-3916-440d-b4bd-c68c2c3a6b68 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:34:22.318049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:0fa41cde966a8dfb3a7f3bb943e1ac8b087dd1283ba9c3a89bd1806115214e6d

Observation b41db84b-b6d6-4ed8-9a55-daf50f968a25 · outbound

This paper cites HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:34:22.312178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:5257d5b2da6e038e6dc63a8cf9e1b682429d423c71e26a6fc948ee8375532ef2

Observation 72b104ca-c66f-4952-b4b4-a77ff7925207 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:34:22.300517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:89cde611b27b17f02682577698c6a1752848991c60db1149bb2ba95475cbd221

Observation f64fb178-d907-4734-a8d3-54012a4eafb9 · outbound

This paper cites DeepMind Control Suite.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning DeepMind Control Suite

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:34:22.305967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:62afb4b25f186dde20ad4a443b6b34cd681c0c77446ecd650bcf624908219fe8

Observation cd31e76f-1ba0-4de1-84ba-6676815879c5 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Deep Reinforcement Learning and the Deadly Triad

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:34:22.282341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:8b541ea9754b1da9fb6d64e30fcef7dac277de0c7536c4f19079d51970e2c1e0

Observation e68fcd9a-b01c-4b3e-b8b9-c747bbedeae8 · outbound

This paper cites Under Review.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Under Review

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T21:34:22.376244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:433e8cfb0c89c6e13a4ea35a4f8c834b737ffa808d0dd7fd07ea8abc0b3494c8

Observation 43ab906f-195e-4aae-af77-35b800c8c918 · outbound

This paper cites The second inequality holds because theminoperator is also a non-expansion.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning The second inequality holds because theminoperator is also a non-expansion

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T21:34:22.380690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:924b3525bd89521c8e947d2dd96381af51fc4d647de5db68ec792d213e767d3d

Observation a23c9dd6-70b0-4424-90ca-64f62a555cb1 · outbound

This paper cites 4, SAC is adapted to use a single Q-function critic, following the approach taken in Simba (Lee et al.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning 4, SAC is adapted to use a single Q-function critic, following the approach taken in Simba (Lee et al

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T21:34:22.372047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:1739917e445327f5a8cf7b842d9116a2942fc567cabff076e431b80e837b3149

Observation 189f580a-e9d7-49db-a419-a87d2ee50837 · outbound

This paper cites Under Review.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Under Review

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T21:34:22.392002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:e43b588766bc2ff44a5e9cbe7364aab81b0c58923476607f9995101d2eab43b8

Observation c1a29729-c6ac-49a7-bafd-4007fc18ad7b · outbound

This paper cites C.4 ONLINERLANDCONTINUOUSCONTROL For our continuous-control experiments with online reinforcement learning, we adopt SimbaV1 and SimbaV2.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning C.4 ONLINERLANDCONTINUOUSCONTROL For our continuous-control experiments with online reinforcement learning, we adopt SimbaV1 and SimbaV2

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T21:34:22.384735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:afbb9b8ab1d66511d24b998f028d219924b490decf66caaf62016436d55262b6

Observation a3a20ee3-2558-4d09-ab56-ae8a34af91af · outbound

This paper cites Identical values are merged.

Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning Identical values are merged

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T21:34:22.388263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T21:33:40.229376Z digest=sha256:85e9851e36fa2b9433cacfe720fcbd281334c7c91ae3f031c3b65654accfaf08

Pith citing papers

Observation 81145edf-9c8f-4fdb-b245-a69fa66192cc · inbound

Stable Deep Reinforcement Learning via Isotropic Gaussian Representations cites this paper.

Stable Deep Reinforcement Learning via Isotropic Gaussian Representations Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T21:42:30.436671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:42:30.436671Z digest=sha256:f595bfaabc7933adee57472e252458f3578e8b302c6c085e55195f656a2cabac