Pith. sign in

Paper Citation Record · LEDGER

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition

As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.09762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09762 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:20:20.073159Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7953fa82-b810-499b-bb1c-efcf13a82c01 · outbound

This paper cites A review of learning-based dynamics models for robotic manipulation |Science Robotics.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition A review of learning-based dynamics models for robotic manipulation |Science Robotics

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.874750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.874750Z digest=sha256:82e795cb89f67a3a94d17826b46c38f18578c070d7e1af67a93bf9e7f84f31ed

Observation 4f175a95-4231-4d2c-8106-761b11c7705a · outbound

This paper cites Learning Con- tinuous Control Actions for Robotic Grasping with Reinforcement Learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Learning Con- tinuous Control Actions for Robotic Grasping with Reinforcement Learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.194190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.881716Z digest=sha256:ba6ab0e14d8d6bbfaf98b912a02de9afc60b9f9805592b0c6985076e44254691

Observation d0cd25a3-7530-46ae-af01-50fc90390d03 · outbound

This paper cites General- Purpose Sim2Real Protocol for Learning Contact-Rich Manipulation With Marker-Based Visuotactile Sensors,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition General- Purpose Sim2Real Protocol for Learning Contact-Rich Manipulation With Marker-Based Visuotactile Sensors,

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-11T11:20:20.852697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.889277Z digest=sha256:bea1ab766c275949eea1d5f2a645a1ec74ef24fd1081176cc03431ad569e28a8

Observation 3eac9bce-f94b-4229-a64e-e3e40b92467d · outbound

This paper cites Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Towards bridging the gap: Systematic sim-to-real transfer for diverse legged robots

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.895124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.895124Z digest=sha256:d3a1d3336877e98352158552a7ca0c64b26abd8203409bf30e4ceccc88ac55b9

Observation 13405272-286b-4dc9-b277-7b1d07376018 · outbound

This paper cites Srl-vic: A variable stiffness-based safe reinforcement learning for contact-rich robotic tasks,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Srl-vic: A variable stiffness-based safe reinforcement learning for contact-rich robotic tasks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.160147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.901989Z digest=sha256:00f2c25ba531fea4eca8d1726f4158a02473a73133354a506062a2a511a8d97e

Observation 95e6ccef-cb1e-475c-8382-cf36fc8ff216 · outbound

This paper cites Deep reinforcement learning for robotics: A survey of real-world successes,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Deep reinforcement learning for robotics: A survey of real-world successes,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.907467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.907467Z digest=sha256:9b92102a1f083225e49a127c3b6f2ab9f0fdd3682b8c29c4aa1ab8b5b0e71271

Observation ee437a5f-e70d-4f2a-be11-d2916a6e6eee · outbound

This paper cites Improving vision-language-action model with online reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Improving vision-language-action model with online reinforcement learning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.121932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.913331Z digest=sha256:a45b8728361f0f2cee9d6199eac6809f1ae56fcc92f1c5f91b5d199229c437fa

Observation ae2eab6f-486a-4215-bef2-fd81a99c225d · outbound

This paper cites Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Hi-WM: Human-in-the-World-Model for Scalable Robot Post-Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.919454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.919454Z digest=sha256:dbdc944eaee5009dfd454ddbee4b5220dd20deb521e1d2e3f0f37c6ac13ab2eb

Observation 2bab35a4-89e1-4c86-b096-dcf22b84ad41 · outbound

This paper cites Rl-100: Performant robotic manipulation with real-world reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Rl-100: Performant robotic manipulation with real-world reinforcement learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.924795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.924795Z digest=sha256:13d2b49a1e075dac5728804665072c843ceb521ff7e0006178d75509648866d5

Observation 96e95e0a-14ed-41b8-afe7-ff5c1cc4a317 · outbound

This paper cites Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.930023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.930023Z digest=sha256:858a1ef7351654569ebec3502ec5b098f145a8ddac5b2b78e57d87221bbc26c0

Observation dca26e27-0d33-4020-bab7-77c33600fa49 · outbound

This paper cites Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation With Large Language Models,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation With Large Language Models,

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-11T11:20:20.456064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.935620Z digest=sha256:f498d71eff73fb07d68b98a9b5caf04e3043a78d6b3d8849d982185a3687dd40

Observation d314379b-deee-43a2-aa96-5c385f267da2 · outbound

This paper cites Serl: A software suite for sample- efficient robotic reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Serl: A software suite for sample- efficient robotic reinforcement learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.942026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.942026Z digest=sha256:6313ff0976442c9bb4eb3fea34a58d72dacc90099f7c35925728b590c4d96f62

Observation b6fd6e15-207f-490b-b2ad-28e8b8d3db58 · outbound

This paper cites Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.947525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.947525Z digest=sha256:12b0487e8ac8517341322fe10d9a574c8f4060a50df6dd8ccb67adb3ee220fc2

Observation 8fe47e55-fee9-49fc-b126-586f27ee244d · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.953071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.953071Z digest=sha256:1cf000e6608d991cb5555c7786a03b0d7a8d4aaa6c67e4dff7bcaee2c56635b5

Observation 22e82f08-1792-4285-9c83-703042f22bcf · outbound

This paper cites Reinflow: Fine-tuning flow matching policy with online reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Reinflow: Fine-tuning flow matching policy with online reinforcement learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.073542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.958730Z digest=sha256:984d75e774e1d0c0fe42218a32b29bc8c1b7b204e2e4acde88700505e2109f2a

Observation 61f3d213-e333-428e-8051-c675de64a402 · outbound

This paper cites Hybrid reward architecture for reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Hybrid reward architecture for reinforcement learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:21.053010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:19.968287Z digest=sha256:358ee698968b9301509a753ef4d76400b344559c56cda87b7393e406c883d3cf

Observation 0142220e-b7b5-49fc-8fe4-3080dc7f2d20 · outbound

This paper cites Efficient online reinforcement learning with offline data,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Efficient online reinforcement learning with offline data,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.974107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.974107Z digest=sha256:cd61320471d7c89e7ed9af3da87f0ec13346d707ab7e674a5d93b8ff495c887a

Observation df969530-88a0-45bf-8ae8-eb2b1f682f6d · outbound

This paper cites Deep Reinforcement Learning in Parameterized Action Space.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Deep Reinforcement Learning in Parameterized Action Space

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.980036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.980036Z digest=sha256:4caf85d2b56e1c4b9c78bc3f22a6358eadf5744072dfed70de50aca5b9e62301

Observation f8be1923-6991-4df6-864a-890db2122199 · outbound

This paper cites Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.985496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.985496Z digest=sha256:2644d41afffdb38f89092f82c0bad07f457ec63591da3d16dd55a0ba144fefab

Observation fe4d146d-8b36-413d-9a1c-539d15bd0047 · outbound

This paper cites Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:19.998476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:19.998476Z digest=sha256:d2a84360b578d6e44006ebc5ab384f563d74c903e4cc15c7f6bba671becdede2

Observation 8a078d8e-7ac0-4a00-8fc4-eb2c7f5db353 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Multi-agent actor-critic for mixed cooperative-competitive environments,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.006679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.006679Z digest=sha256:6c69f096cf9aa6ce18240248face52de05b37785b16956ea7dcdb06e0213f115

Observation 36ffc7a6-71f5-4857-84c9-17b733237f35 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition The surprising effectiveness of ppo in cooperative multi-agent games,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.012254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.012254Z digest=sha256:75c540b66dee8ed424581920feacf2be0ae0e4e1ba72f4ae88b520010a875685

Observation 55b51ff4-6707-4e05-9489-6b5836aefa71 · outbound

This paper cites Asynchronous actor-critic for multi-agent reinforcement learning,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Asynchronous actor-critic for multi-agent reinforcement learning,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.985986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:20.021539Z digest=sha256:a227c22819d2eb908afafd6dce6246aa7ee7206b6e81be40b7d983d8a9fe94ac

Observation 5d5e5942-bb6d-414a-812b-bd687254c0a5 · outbound

This paper cites Action decoupled sac reinforcement learning with discrete-continuous hybrid action spaces,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Action decoupled sac reinforcement learning with discrete-continuous hybrid action spaces,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.956233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:20.027981Z digest=sha256:430b59cae1405eb620d1c8de6ef34cf017887b379d52f2db05690733f2a2944c

Observation 83659ad8-fe57-4c94-b34d-b03c64bccc8a · outbound

This paper cites Effective multi-agent deep reinforcement learning control with relative entropy regularization,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Effective multi-agent deep reinforcement learning control with relative entropy regularization,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.934523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:20.036886Z digest=sha256:dd036dba9f81466b38d4889c812f824742f4048511c0223d3cfcf750c04c668e

Observation e5eeae7a-6621-4cf4-a573-156865f8c01e · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Soft Actor-Critic Algorithms and Applications

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.042024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.042024Z digest=sha256:ed071f01fdbe4ba15099d9c5558dba6e81eea0b07b9adf3752907137d1c22f02

Observation f2cabc70-02ae-44eb-b8ae-be067ec5d036 · outbound

This paper cites Deep residual learning for image recognition,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Deep residual learning for image recognition,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.048651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.048651Z digest=sha256:7f6679525760be9b63110870fadba374478cd55dcc0de7e4b89f501d5b9df10d

Observation 5074cb71-2bf8-4341-9c92-132d71c9fe2e · outbound

This paper cites An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition An Introduction to Centralized Training for Decentralized Execution in Cooperative Multi-Agent Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.055396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.055396Z digest=sha256:3b2c40f7ef42b563f010c105b96c603face7f34eb2b859404867be6bf60f5ccf

Observation 641735f2-9b57-48e6-98be-efa0f519e538 · outbound

This paper cites Soft Actor-Critic for Discrete Action Settings.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Soft Actor-Critic for Discrete Action Settings

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.061953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.061953Z digest=sha256:7c851edce3b0854ae995a9ed53b1c812c4b161e25f2f661698919e03c0fae3c6

Observation c43875eb-ce6b-4761-a192-06de815cc0a2 · outbound

This paper cites A high-force gripper with embedded multimodal sensing for powerful and perception driven grasping,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition A high-force gripper with embedded multimodal sensing for powerful and perception driven grasping,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:20:20.894818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T11:20:20.067472Z digest=sha256:350523c8eab90fbf96dcf2e1ddd141092a179729f0e2fabf054f7c751e8f7482

Observation 80032bd8-1fab-403b-ad4b-09aa3173a7e7 · outbound

This paper cites Root mean square layer normalization,.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Root mean square layer normalization,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.073159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.073159Z digest=sha256:bfab3266dc79a1090b96fe699775917b4aa5666ecec88f3c8f2013a8d886cf1e

Pith citing papers

No inbound Pith citation observations are available.