Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-01T08:05:47.128354Z
Paper Citation Record · LEDGER
As of 22 July 2026, this Paper Citation Record lists 74 of 74 outbound references and 3 inbound Pith citation observations for arXiv:2605.00416.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-01T08:05:47.128354Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T14:57:49.542416Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T08:59:42.033560Z
74 of 74 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c462da4b-e6a1-491f-822f-04f119b74306 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies RT-1: Robotics Transformer for Real-World Control at Scale
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fabb3f4d-e4b6-409e-a8bd-5b33f29847c7 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rt-2: Vision-language-action models transfer web knowledge to robotic control
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8d8c5cda-965c-4867-b44b-61ee5255f588 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Octo: An Open-Source Generalist Robot Policy
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0f33576f-9359-43f5-9b31-6b0bc32819a6 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies OpenVLA: An Open-Source Vision-Language-Action Model
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b8ccb53a-a27c-48d8-ab89-17ed539b5e6b · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0cf90017-53d0-4d79-9ef2-0467b0c9982f · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies π 0.5: A vision-language-action model with open-world gener- alization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 7d87ea34-b387-4550-b3d8-7ee1509c7d5b · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hg-dagger: Interactive imitation learning with human experts
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 17895e33-1613-4c06-8426-9117b50be000 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Q-learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation da354dc1-2e82-43ce-a670-332ebb0a43bd · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Addressing func- tion approximation error in actor-critic methods
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4de9d740-0529-4809-9502-363d822f9f6c · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Contin- uous control with deep reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 20e703b6-0584-4855-b5fe-6e5dd6f1b146 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Soft actor-critic: Off-policy maximum entropy deep reinforce- ment learning with a stochastic actor
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8655e183-df38-424e-a622-1e62cb0f2392 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rl-100: Performant robotic manipulation with real-world reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 789e75af-6c26-4ef8-807f-f90ae2a241d3 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Gr-rl: Going dexterous and precise for long-horizon robotic manipulation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c4d77d0f-fb4d-4a20-b8e7-6f69cf9e9c29 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 31042599-b993-42b1-9b4f-31a07958c008 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies $\pi^{*}_{0.6}$: a VLA That Learns From Experience
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 47a58aa5-fdc0-46bd-a8af-21bfd2e46dd3 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Serl: A software suite for sample-efficient robotic reinforcement learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b99c4d50-6305-4146-b423-8bda435368b0 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f4b8ac92-607d-4c3c-97a1-5471416adf4d · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 4d23dbfa-1a95-47b7-8b53-3123caaf6da4 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Interactive Post-Training for Vision-Language-Action Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 358de7c6-da13-4340-a79c-dbdfb6df5f8f · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies pi rl: Online rl fine-tuning for flow-based vision-language-action mod- els.arXiv preprint arXiv:2510.25889
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 482c806c-1c8e-4c21-b136-39aab369ef48 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Flow-GRPO: Training Flow Matching Models via Online RL
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 3ad599a0-9b5c-4174-99bc-f1362d657a22 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies arXiv preprint arXiv:2505.22094 , year=
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d50a8dca-f1c4-4c76-8f91-55a8d8b8ae43 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline Reinforcement Learning with Implicit Q-Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 109a64ea-7998-4f52-8b58-88895af6c5ef · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation dad920bd-2f6b-4285-afa8-49afeca480a0 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Q-learning with Adjoint Matching
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 22a6d22b-569d-45e3-bd86-0be02f5443cb · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8446b9d6-3128-4dda-b251-182de97f9328 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies GRAPE: Generalizing Robot Policy via Preference Alignment
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 183d3be4-3e32-49e0-9f9b-a82ef44e7052 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rlinf-vla: A unified and efficient framework for vla+ rl training
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c2c2c579-7846-4540-b5f3-5da937635857 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8f56a2a4-f713-460f-b1f0-517b52dff45c · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c668455f-d698-42e8-acbf-a0b1b81c651f · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Behavior- 1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f67119e0-be7b-4c26-8ace-554d8a076960 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 35779cb1-d0c6-4e75-8dec-b4233fa0f725 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Libero: Benchmarking knowledge transfer for lifelong robot learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e9ea6f10-3405-413f-8245-9d2cdaffc4e1 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Robotwin: Dual-arm robot benchmark with generative digital twins (early version)
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 1a7652de-82a3-4881-90b4-07dd324576d0 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 713d09b6-8adf-48be-9c14-3c955c9ce917 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d31d1aeb-cb82-4e75-a965-7ce03e0c0556 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d0872b62-2baf-42d0-a71d-be727a5aeaf6 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Flow q-learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f2f8e029-05fd-47a5-9928-94a5a2510b04 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation b5d88aa3-9402-4fad-a5f9-d67244d7aa3c · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline- to-online reinforcement learning via balanced replay and pessimistic q-ensemble
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 8bbff4a3-ff21-46d1-9343-e5b0b819ff65 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Reincarnating reinforcement learn- ing: Reusing prior computation to accelerate progress
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation f859429f-2929-4a77-b7bb-0e2ab517c43a · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Effi- cient online reinforcement learning with offline data
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation e8b82892-7e02-4b72-aec3-04041763d612 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 39a0a6bc-4bbd-43fc-b36f-1e8850532b61 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d4d498a3-e9d2-4bfd-b160-d7e76b4381fa · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Steering Your Diffusion Policy with Latent Space Reinforcement Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 66da4269-1d56-4946-8c15-7a1729571d80 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Qt-opt: Scalable deep rein- forcement learning for vision-based robotic manipula- tion
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 19d92e87-c654-4ea4-a22e-77542f76a7d5 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 17c5cc86-2c1e-40aa-b398-c061d477a16f · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Pi-qt-opt: Predictive information improves multi-task robotic reinforcement learning at scale
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 902a7074-8583-4df7-9736-685928302ae9 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Sop: A scalable online post-training system for vision-language-action models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 5448ea69-3db1-4919-9941-7a1bec6e497f · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 633d21ce-7b75-4db9-8bda-139822ac0c4b · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Impala: Scalable dis- tributed deep-rl with importance weighted actor-learner architectures
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 37424740-b8c3-4ff7-899f-7c0f9d55b8e9 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Deep RL at Scale: Sorting Waste in Office Buildings with a Fleet of Mobile Manipulators
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c63998fc-d28f-4cdd-b489-ffdc73bd3dcb · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Flow matching for generative modeling
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c43499ee-d956-4d65-ab8b-c9462145ed13 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies A dis- tributional perspective on reinforcement learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 2c291311-8889-4673-b059-44ecf5f1b6c6 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c74f5f11-bd6d-4522-8274-9e8f19b811fa · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d5fa6e00-a2c0-46c4-81bd-6d9de680ed09 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Energy-weighted flow matching for offline reinforcement learning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 9e25219c-802b-430f-b70e-b129003c16bb · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Gemma 3 technical report
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 7b0b1d20-208b-4754-8474-657d6830b089 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Sigmoid loss for language image pre-training
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0a6dd454-d857-412b-8905-d434933e510d · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation a91f9f70-303e-4390-a230-56c7157b4ba8 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Vision trans- formers for dense prediction
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 38067fb4-b886-4c5c-8678-4df8ad88fafe · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 0b80b7f0-6563-4815-853a-8376770dd978 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Decoupled weight decay regularization
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation bb11bb1a-f882-465e-9f6d-256642ab8e5f · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies In our real- robot experiments, we useK= 201atoms over[−0.1,1.1]
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation fdb24c82-a66f-421d-9ad7-90692f548c5f · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 1492b493-37db-4baf-b0a7-485dec5690ce · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 954c73f3-0eb1-4fad-ab9d-325d18f39c9b · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Demonstrations are successful trajectories, rollouts contain both successes and failures, and play data is treated as unsuccessful exploratory data
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 58af25cc-d75f-4e8c-b548-e1ff60d54e12 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies The policy is optimized with AdamW [63] using a base learning rate of2×10 −5 and a cosine decay schedule
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation c18fac3f-1573-45d4-be3e-c108997614bd · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 396975de-941d-452b-8b0c-deb512f0c24a · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies The model is trained with a flow-matching loss, where the interpolated noisy actiona w is defined in Eq
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 18f88483-d5e8-43b7-b1f3-9a39d2806629 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies The comparison isolates the Robot 1 Robot 2
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 595baa96-96fe-431c-bcb7-16445fcc92d6 · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies 9 vi- sualizes the predicted value distributions for the same episodes shown in Fig
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 09c10d04-97e4-410b-9d6e-efb8489b620d · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies (i) Object-storage uploads commit atomically (read- ers see either the fully-uploaded payload or no object) and are retried until persisted
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation d87dbcdb-4e25-47a0-aec3-82a14d10eb4a · outbound
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Table VI reports both on the same 8-hour, 16-actor run as the End-to-End Reliability subsection above
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 6de04dfe-12cf-4fbe-8d39-258337ca15b7 · inbound
UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 42d737be-edc7-42de-a44c-5013aaaf9047 · inbound
FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.
Observation 30afc46d-77f9-40f2-8f24-ea6c22132644 · inbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.