Pith. sign in

Paper Citation Record · LEDGER

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition

As of 13 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.09762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09762 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:20:20.073159Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7953fa82-b810-499b-bb1c-efcf13a82c01 · outbound

This paper cites A review of learning-based dynamics models for robotic manipulation |Science Robotics.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition A review of learning-based dynamics models for robotic manipulation |Science Robotics

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.874750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.874750Z digest=sha256:2105bfd8984e395bd4d7da4d25667fc5f7ea1b647f84f92abc6b92d3e1ebabd3

Observation 4f175a95-4231-4d2c-8106-761b11c7705a · outbound

This paper cites Learning Con- tinuous Control Actions for Robotic Grasping with Reinforcement Learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Learning Con- tinuous Control Actions for Robotic Grasping with Reinforcement Learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.194190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.881716Z digest=sha256:e408d2d86c37eda53ca0226b729c708e6e586bb498d81323c6077454ef0e562a

Observation d0cd25a3-7530-46ae-af01-50fc90390d03 · outbound

This paper cites General- Purpose Sim2Real Protocol for Learning Contact-Rich Manipulation With Marker-Based Visuotactile Sensors,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition General- Purpose Sim2Real Protocol for Learning Contact-Rich Manipulation With Marker-Based Visuotactile Sensors,

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-11T11:20:20.852697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.889277Z digest=sha256:f7bd724664cf968102e3eadb2547fa444b9739f1cdc7765ae9d2c28acaeb4a1c

Observation 3eac9bce-f94b-4229-a64e-e3e40b92467d · outbound

This paper cites Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.895124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.895124Z digest=sha256:2b7e4119930b4ef8e18d12565227aa21d97fd854c43294c75364383fb51b9a72

Observation 13405272-286b-4dc9-b277-7b1d07376018 · outbound

This paper cites Srl-vic: A variable stiffness-based safe reinforcement learning for contact-rich robotic tasks,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Srl-vic: A variable stiffness-based safe reinforcement learning for contact-rich robotic tasks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.160147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.901989Z digest=sha256:49277ba25324b06e3d897a116bd2aa2db58c778ffcd94615ce7413a43bed0121

Observation 95e6ccef-cb1e-475c-8382-cf36fc8ff216 · outbound

This paper cites Deep reinforcement learning for robotics: A survey of real-world successes,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Deep reinforcement learning for robotics: A survey of real-world successes,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.907467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.907467Z digest=sha256:c44459b3c5c47378cc81cf388e0200a9ffa516e9bb22889cd51664874e632cc1

Observation ee437a5f-e70d-4f2a-be11-d2916a6e6eee · outbound

This paper cites Improving vision-language-action model with online reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Improving vision-language-action model with online reinforcement learning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.121932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.913331Z digest=sha256:d15829b4c6f3554a14c65856e6bbfc0c82ee7251428e1db39a8b705e38e888d6

Observation ae2eab6f-486a-4215-bef2-fd81a99c225d · outbound

This paper cites Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.919454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.919454Z digest=sha256:260692f1fd0e7539fb67691bc3001651d349f4e1e265e383851799b8517445b1

Observation 2bab35a4-89e1-4c86-b096-dcf22b84ad41 · outbound

This paper cites Rl-100: Performant robotic manipulation with real-world reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Rl-100: Performant robotic manipulation with real-world reinforcement learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.924795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.924795Z digest=sha256:35b4c880e9e3c85f01339cfb8bfa0dccad51d49123fae0582bc8848b89bf975b

Observation 96e95e0a-14ed-41b8-afe7-ff5c1cc4a317 · outbound

This paper cites Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.930023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.930023Z digest=sha256:ad2061e615f70bea92a8ea529b23f6ef8e2190df8b4a134c244f97e13907ab80

Observation dca26e27-0d33-4020-bab7-77c33600fa49 · outbound

This paper cites Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation With Large Language Models,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation With Large Language Models,

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-11T11:20:20.456064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.935620Z digest=sha256:b122dd580cefe36709190b7e0e1a401fed0c810c9d76f2d9d588c71ac2c0a311

Observation d314379b-deee-43a2-aa96-5c385f267da2 · outbound

This paper cites Serl: A software suite for sample- efficient robotic reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Serl: A software suite for sample- efficient robotic reinforcement learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.942026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.942026Z digest=sha256:80db465f0fb6ddcadaa65d4ee68cbc63aa24f420b1198a2cffcd658713563abf

Observation b6fd6e15-207f-490b-b2ad-28e8b8d3db58 · outbound

This paper cites Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.947525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.947525Z digest=sha256:de0206632b91743a54c0924616f029a5f3b449fff3753f67b955a1084e095bf6

Observation 8fe47e55-fee9-49fc-b126-586f27ee244d · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.953071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.953071Z digest=sha256:4d69bb918f6e914afcfd5f09382bb4807b86f3e2f1fce39b2d446924621bd20e

Observation 22e82f08-1792-4285-9c83-703042f22bcf · outbound

This paper cites Reinflow: Fine-tuning flow matching policy with online reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Reinflow: Fine-tuning flow matching policy with online reinforcement learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.073542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.958730Z digest=sha256:a57cbb9a8ff8c0d0fdfc97ed548b22cc10ef3eb3488150a0ddc4da9c8c3cf0b0

Observation 61f3d213-e333-428e-8051-c675de64a402 · outbound

This paper cites Hybrid reward architecture for reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Hybrid reward architecture for reinforcement learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.053010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.968287Z digest=sha256:310029073c0007d58f2b1d9476ee49a822f8275589cccad3db7605597a969339

Observation 0142220e-b7b5-49fc-8fe4-3080dc7f2d20 · outbound

This paper cites Efficient online reinforcement learning with offline data,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Efficient online reinforcement learning with offline data,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.974107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.974107Z digest=sha256:b5adc8ee491f802257cca55cc217cbf3144ed40b44dc1f13af2479c5aa46e730

Observation df969530-88a0-45bf-8ae8-eb2b1f682f6d · outbound

This paper cites Deep Reinforcement Learning in Parameterized Action Space.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Deep Reinforcement Learning in Parameterized Action Space

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.980036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.980036Z digest=sha256:fcdc558eb7496f525734746cbd75c873fc8d660251e405049a1a4b37c4624e7c

Observation f8be1923-6991-4df6-864a-890db2122199 · outbound

This paper cites Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.985496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.985496Z digest=sha256:8311b7ac819018360253f115198eb5efd88930c6249b1593181dfa8b558d12a4

Observation fe4d146d-8b36-413d-9a1c-539d15bd0047 · outbound

This paper cites Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.998476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.998476Z digest=sha256:29c33b9d5e6046ad08e3dccaa7967050746e9e339929244ddfe107dfd851b1fe

Observation 8a078d8e-7ac0-4a00-8fc4-eb2c7f5db353 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Multi-agent actor-critic for mixed cooperative-competitive environments,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.006679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.006679Z digest=sha256:b64c0bf8462d40d647cdbcc6818df9987194f8a46b5ff179a5bf7bad3d6e6d48

Observation 36ffc7a6-71f5-4857-84c9-17b733237f35 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition The surprising effectiveness of ppo in cooperative multi-agent games,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.012254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.012254Z digest=sha256:73af299071e9a1d5290fbdc26271d1a38d2c62fe92d9b46fe990740da8734e30

Observation 55b51ff4-6707-4e05-9489-6b5836aefa71 · outbound

This paper cites Asynchronous actor-critic for multi-agent reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Asynchronous actor-critic for multi-agent reinforcement learning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.985986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:20.021539Z digest=sha256:2d52e8506ba9802cc53311b7cabac18f5b626877852b74174da135f42f0cab64

Observation 5d5e5942-bb6d-414a-812b-bd687254c0a5 · outbound

This paper cites Action decoupled sac reinforcement learning with discrete-continuous hybrid action spaces,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Action decoupled sac reinforcement learning with discrete-continuous hybrid action spaces,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.956233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:20.027981Z digest=sha256:3e74854aac14f30787d727d83b5e9d2afbf9de37ae68295422480094ffbae352

Observation 83659ad8-fe57-4c94-b34d-b03c64bccc8a · outbound

This paper cites Effective multi-agent deep reinforcement learning control with relative entropy regularization,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Effective multi-agent deep reinforcement learning control with relative entropy regularization,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.934523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:20.036886Z digest=sha256:4c4754e3169f22d6bf314dc0670e4d41d6d3853d8dc41afc5aba0b1700463431

Observation e5eeae7a-6621-4cf4-a573-156865f8c01e · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Soft Actor-Critic Algorithms and Applications

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.042024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.042024Z digest=sha256:04298ca319424fe58d5f04e71dcbcfbb9867e1d980d12130b08057d11929f9fd

Observation f2cabc70-02ae-44eb-b8ae-be067ec5d036 · outbound

This paper cites Deep residual learning for image recognition,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Deep residual learning for image recognition,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.048651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.048651Z digest=sha256:2745655f55a27a32c78a2789a90cb7d3538447ac506a670a0cac6093ae7dc85a

Observation 5074cb71-2bf8-4341-9c92-132d71c9fe2e · outbound

This paper cites An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.055396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.055396Z digest=sha256:34cf0b9e1eff640d6b56ea072be165dd63db8b069fb3ddc9e0da738c4efa2ead

Observation 641735f2-9b57-48e6-98be-efa0f519e538 · outbound

This paper cites Soft Actor-Critic for Discrete Action Settings.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Soft Actor-Critic for Discrete Action Settings

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.061953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.061953Z digest=sha256:b53489d32df7b2a99856f2e0b57887d22636b831142791f8f67105951296f56c

Observation c43875eb-ce6b-4761-a192-06de815cc0a2 · outbound

This paper cites A high-force gripper with embedded multimodal sensing for powerful and perception driven grasping,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition A high-force gripper with embedded multimodal sensing for powerful and perception driven grasping,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.894818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:20.067472Z digest=sha256:0f6f15eda9c10e7dfc59249ad5ad475bd9873aadf6b245f8332bfc6358edce4a

Observation 80032bd8-1fab-403b-ad4b-09aa3173a7e7 · outbound

This paper cites Root mean square layer normalization,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Root mean square layer normalization,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.073159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.073159Z digest=sha256:f4f2f9b2216396ae6e93db364d344494c0751853c6e00de800ccab747d67bf3e

Pith citing papers

No inbound Pith citation observations are available.