Pith. sign in

Paper Citation Record · LEDGER

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

As of 22 July 2026, this Paper Citation Record lists 74 of 74 outbound references and 3 inbound Pith citation observations for arXiv:2605.00416.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.00416 v2

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T08:05:47.128354Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-21T06:31:05.380196+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T14:57:49.542416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T08:59:42.033560Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact32
  • verified fuzzy39
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c462da4b-e6a1-491f-822f-04f119b74306 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies RT-1: Robotics Transformer for Real-World Control at Scale

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.301652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d6bf3052754b70ab1b22eab43fac37cc093b05bb79c1b9476a64e3e4da75bd4a

Observation fabb3f4d-e4b6-409e-a8bd-5b33f29847c7 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.019797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:09fdc6e9915cd84a0b60e4b2d084ae7398ce8cb5f977f7fe7bace5feab6607d2

Observation 8d8c5cda-965c-4867-b44b-61ee5255f588 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Octo: An Open-Source Generalist Robot Policy

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.328611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:f98f3ee8515c62562824150b1bccf9116a5327f526e03db76ef7567549c4f658

Observation 0f33576f-9359-43f5-9b31-6b0bc32819a6 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies OpenVLA: An Open-Source Vision-Language-Action Model

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.244008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:b4eebc828887c02b85edbfb6395e85d87ae953bd327fa0915001efd3d7da6fb3

Observation b8ccb53a-a27c-48d8-ab89-17ed539b5e6b · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.376300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:7079667b15fd3862ff4c081371f594edc90842363ca1613329ae87e8640d3e83

Observation 0cf90017-53d0-4d79-9ef2-0467b0c9982f · outbound

This paper cites π 0.5: A vision-language-action model with open-world gener- alization.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies π 0.5: A vision-language-action model with open-world gener- alization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.967936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:f1b15e0fbf1c222224e8bdbba6650563b5d3f07c770c04b52db830f8edfd4b3e

Observation 7d87ea34-b387-4550-b3d8-7ee1509c7d5b · outbound

This paper cites Hg-dagger: Interactive imitation learning with human experts.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hg-dagger: Interactive imitation learning with human experts

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.004270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:77c030b73bfa188312cdb6b45583cb1c42331c5c3131a7c1e9fa794c0e3bc825

Observation 17895e33-1613-4c06-8426-9117b50be000 · outbound

This paper cites Q-learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Q-learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.010141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:0474b0756130d734149d82c1248573d884294550381e0234afc3d0838a574b18

Observation da354dc1-2e82-43ce-a670-332ebb0a43bd · outbound

This paper cites Addressing func- tion approximation error in actor-critic methods.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Addressing func- tion approximation error in actor-critic methods

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.012138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:73d3035bce927c035c566c7619786a92ec9b37f1aeb5c2c05fab3ff086176660

Observation 4de9d740-0529-4809-9502-363d822f9f6c · outbound

This paper cites Contin- uous control with deep reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Contin- uous control with deep reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.974588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:cd97dbb665429d6495c53d9119dd4158dd8b8a13f34901e99645e10e36f2cffc

Observation 20e703b6-0584-4855-b5fe-6e5dd6f1b146 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforce- ment learning with a stochastic actor.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Soft actor-critic: Off-policy maximum entropy deep reinforce- ment learning with a stochastic actor

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.965855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:3110164e4554dc802795308aee113d1a65c4620ee61128a9bc2c1921f5a3adb9

Observation 8655e183-df38-424e-a622-1e62cb0f2392 · outbound

This paper cites Rl-100: Performant robotic manipulation with real-world reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rl-100: Performant robotic manipulation with real-world reinforcement learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.288694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:c894dde9fa84561d8c513e9530886a34cf35d497e8c7041d1865be8d57109540

Observation 789e75af-6c26-4ef8-807f-f90ae2a241d3 · outbound

This paper cites Gr-rl: Going dexterous and precise for long-horizon robotic manipulation.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Gr-rl: Going dexterous and precise for long-horizon robotic manipulation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.372068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:dc5f3ba0958634c0ba3a01f8f0eeaea35bdaeafc087e1d60cab108e6580ba757

Observation c4d77d0f-fb4d-4a20-b8e7-6f69cf9e9c29 · outbound

This paper cites ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.386306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d55ea22d6db3587ae342cc629668a234300d60cae8a4075665f71a63d15c0819

Observation 31042599-b993-42b1-9b4f-31a07958c008 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.263986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d7f894c0033652a8bdaeeb68d1a18878043a55086b9c986b41cfb6327a7c6f6f

Observation 47a58aa5-fdc0-46bd-a8af-21bfd2e46dd3 · outbound

This paper cites Serl: A software suite for sample-efficient robotic reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Serl: A software suite for sample-efficient robotic reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.997970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:196634a4e715e2f2e62a9dc29b3d2a623a99b5d634b786425741e47caf7d97d1

Observation b99c4d50-6305-4146-b423-8bda435368b0 · outbound

This paper cites Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.944097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d4656e1d60f4207f64c9eff74bdd0cd436c6e6d01c278fa98e4d75cf69884e33

Observation f4b8ac92-607d-4c3c-97a1-5471416adf4d · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.305693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:a68e2d68861699ac3796c7f14456cbacf4d1cb1f25fd204e7cf1b3212373495e

Observation 4d23dbfa-1a95-47b7-8b53-3123caaf6da4 · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Interactive Post-Training for Vision-Language-Action Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.281810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:12e0155bff8a756c4ebe91819b03c1e48d3d9801f1ae213ce1c5feb1abb1dd7c

Observation 358de7c6-da13-4340-a79c-dbdfb6df5f8f · outbound

This paper cites pi rl: Online rl fine-tuning for flow-based vision-language-action mod- els.arXiv preprint arXiv:2510.25889.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies pi rl: Online rl fine-tuning for flow-based vision-language-action mod- els.arXiv preprint arXiv:2510.25889

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.254193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:76cf926bf51b9746470effd71e4bbeb50d7993fb138ca537e07424ed42c99395

Observation 482c806c-1c8e-4c21-b136-39aab369ef48 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Flow-GRPO: Training Flow Matching Models via Online RL

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.314277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:89234dfc03c204364ab2c2858fb9483c8e17e1bfc248d045ab3272529c575818

Observation 3ad599a0-9b5c-4174-99bc-f1362d657a22 · outbound

This paper cites arXiv preprint arXiv:2505.22094 , year=.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies arXiv preprint arXiv:2505.22094 , year=

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:15:32.234520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:349013f5c3d12c5c2af2543c731a7b9a989d45d92bfce2cc9139b01fa52ba0da

Observation d50a8dca-f1c4-4c76-8f91-55a8d8b8ae43 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline Reinforcement Learning with Implicit Q-Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.277366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:693d45bba13928b7079c3bc7ba957a8fbf7fc08c10ace6f6c9763214d5790a8c

Observation 109a64ea-7998-4f52-8b58-88895af6c5ef · outbound

This paper cites Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.357386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:816c8c2193c3202ee077471a0054d597a765fcddffd3ad0b74e7aeb48f867c1b

Observation dad920bd-2f6b-4285-afa8-49afeca480a0 · outbound

This paper cites Q-learning with Adjoint Matching.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Q-learning with Adjoint Matching

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.273179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:1837da94037c4a702c4679e0ef7ef1a6820c316348b8eeb574f401d9db0a6e2d

Observation 22a6d22b-569d-45e3-bd86-0be02f5443cb · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.014084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:18f5e5ec743a0741d66f54bfd7fb1e2c6cf4d44630b5407ec8c4086d7d7e152f

Observation 8446b9d6-3128-4dda-b251-182de97f9328 · outbound

This paper cites GRAPE: Generalizing Robot Policy via Preference Alignment.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.297095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:b4f526d85d830702951b00b91df6b493ee1ecdbfe642503b441a4fdced4fb6d4

Observation 183d3be4-3e32-49e0-9f9b-a82ef44e7052 · outbound

This paper cites Rlinf-vla: A unified and efficient framework for vla+ rl training.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rlinf-vla: A unified and efficient framework for vla+ rl training

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.332868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:05afc76280e748e37eeccd733e8239ef4c08c10f48cc4acd189e088174761991

Observation c2c2c579-7846-4540-b5f3-5da937635857 · outbound

This paper cites What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies What can rl bring to vla generalization? an empirical study.arXiv preprint arXiv:2505.19789

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.353236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:20ee56307a0432c089874edfe9c51d5e477e9cb765e6b5150d0fd4f8d9249a1a

Observation 8f56a2a4-f713-460f-b1f0-517b52dff45c · outbound

This paper cites RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies RLDG: Robotic Generalist Policy Distillation via Reinforcement Learning

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:15:32.390805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:4e9216005325a8cf0283f5b775e4ad49ce0d1d15006a288ec6de84f8aaae648f

Observation c668455f-d698-42e8-acbf-a0b1b81c651f · outbound

This paper cites Behavior- 1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Behavior- 1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.989557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:0b3bdf8d9184159c064191091fa054fb9000abd56eecba1e0d8c07883e3d0d98

Observation f67119e0-be7b-4c26-8ace-554d8a076960 · outbound

This paper cites ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.259437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:803107c3f46699080c782107f1bda125db938775fb1e635f2f075a573f91090d

Observation 35779cb1-d0c6-4e75-8dec-b4233fa0f725 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.946914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:cde363339a99977cbec979e70c4cfe966bff98c7dac82d66c4fe024f3abc5cc0

Observation e9ea6f10-3405-413f-8245-9d2cdaffc4e1 · outbound

This paper cites Robotwin: Dual-arm robot benchmark with generative digital twins (early version).

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Robotwin: Dual-arm robot benchmark with generative digital twins (early version)

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.993608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:44a9f1349157b04b575b65b932fe05efd11ed4220e8e8b7d3060580bbe33b7e8

Observation 1a7652de-82a3-4881-90b4-07dd324576d0 · outbound

This paper cites Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Rlinf-user: A unified and extensible system for real-world online policy learning in embodied ai

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.239243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:ea2d3810238bc2341347d6c2420e5cb0737354811e4cdddd17029e6a37813567

Observation 713d09b6-8adf-48be-9c14-3c955c9ce917 · outbound

This paper cites WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.346599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:b506f4f2b7eaa4a370aaf31f9002d07db3578b93bb54577ddde76adc8ad79dd5

Observation d31d1aeb-cb82-4e75-a965-7ce03e0c0556 · outbound

This paper cites SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.395187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:f31dd7afd0a4c18b396962a62748128e5f813e46d6c6bfa3fd9d0f35daa3d43c

Observation d0872b62-2baf-42d0-a71d-be727a5aeaf6 · outbound

This paper cites Flow q-learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Flow q-learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.978628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:3fe00700232d8d498908ee9cd7545bd8e83c48102063cd98377cbcaff99f3ac0

Observation f2f8e029-05fd-47a5-9928-94a5a2510b04 · outbound

This paper cites Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.980895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:330ae2801a5e69db3a08331a12548e917e5d0d6fbd2e7abb122a4da3784ecff9

Observation b5d88aa3-9402-4fad-a5f9-d67244d7aa3c · outbound

This paper cites Offline- to-online reinforcement learning via balanced replay and pessimistic q-ensemble.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline- to-online reinforcement learning via balanced replay and pessimistic q-ensemble

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.986777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:9038fa08dd67a49af2193ae2f5476a4c2d5fb7b297ae8d34cc3e6b6e80821096

Observation 8bbff4a3-ff21-46d1-9343-e5b0b819ff65 · outbound

This paper cites Reincarnating reinforcement learn- ing: Reusing prior computation to accelerate progress.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Reincarnating reinforcement learn- ing: Reusing prior computation to accelerate progress

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.976756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d69f1776ab5ad29be01a836971e7007f3c5f012b1a3e9bc2d88ed78f36129aa9

Observation f859429f-2929-4a77-b7bb-0e2ab517c43a · outbound

This paper cites Effi- cient online reinforcement learning with offline data.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Effi- cient online reinforcement learning with offline data

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.006226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:fbaee9fad9a3a872b7174f626c9f30691799270773c4c8c333699e133d42f3c7

Observation e8b82892-7e02-4b72-aec3-04041763d612 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.337531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:edda728ebf5faf1e77cf258d410a68d602d500687c9b41462963c90029fdfca1

Observation 39a0a6bc-4bbd-43fc-b36f-1e8850532b61 · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.342064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:94e58b66cdc9ad54d68f75e40440f6830dad9c4d0ac683be5bf99a9e128f6d8d

Observation d4d498a3-e9d2-4bfd-b160-d7e76b4381fa · outbound

This paper cites Steering Your Diffusion Policy with Latent Space Reinforcement Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Steering Your Diffusion Policy with Latent Space Reinforcement Learning

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.362409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:c1c85c1d62938eda3995da06552816104b557700c8789361211f49c42989d685

Observation 66da4269-1d56-4946-8c15-7a1729571d80 · outbound

This paper cites Qt-opt: Scalable deep rein- forcement learning for vision-based robotic manipula- tion.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Qt-opt: Scalable deep rein- forcement learning for vision-based robotic manipula- tion

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.963482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:ac281199b206bab8bf3daff2254c07742c234caf92cc56eb0291495f8229ff47

Observation 19d92e87-c654-4ea4-a22e-77542f76a7d5 · outbound

This paper cites MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies MT-Opt: Continuous Multi-Task Robotic Reinforcement Learning at Scale

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.309954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:e427d2bca68fc13d51feec3c6680d1cb05ef7885693501f2bfc5c09c0a92d29b

Observation 17c5cc86-2c1e-40aa-b398-c061d477a16f · outbound

This paper cites Pi-qt-opt: Predictive information improves multi-task robotic reinforcement learning at scale.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Pi-qt-opt: Predictive information improves multi-task robotic reinforcement learning at scale

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.023775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:6b4153fca808a8a6be071aa7c36b742780de5241ea14b4a0a17aaac3ffe20c08

Observation 902a7074-8583-4df7-9736-685928302ae9 · outbound

This paper cites Sop: A scalable online post-training system for vision-language-action models.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Sop: A scalable online post-training system for vision-language-action models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.366990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:5a0fb8a940d5f5ecf25494144fafce0a8b281a6875c3ed11e608128474d0d4e3

Observation 5448ea69-3db1-4919-9941-7a1bec6e497f · outbound

This paper cites RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.324014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:6b974e2e72869a57e12b99d8c66c5eeac198116923d86a6794887e88e5723b3c

Observation 633d21ce-7b75-4db9-8bda-139822ac0c4b · outbound

This paper cites Impala: Scalable dis- tributed deep-rl with importance weighted actor-learner architectures.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Impala: Scalable dis- tributed deep-rl with importance weighted actor-learner architectures

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.000134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:a0abec58aff2864a6ff097abeca989bbcc6e1726c4c1a2b3a368de52bd5235f3

Observation 37424740-b8c3-4ff7-899f-7c0f9d55b8e9 · outbound

This paper cites Deep RL at Scale: Sorting Waste in Office Buildings with a Fleet of Mobile Manipulators.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Deep RL at Scale: Sorting Waste in Office Buildings with a Fleet of Mobile Manipulators

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.381350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:b17e1468a3483d79ae85e23a94b882c6664a8bfd220769e757b34ff1d930a35c

Observation c63998fc-d28f-4cdd-b489-ffdc73bd3dcb · outbound

This paper cites Flow matching for generative modeling.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Flow matching for generative modeling

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.015969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:7bbe4d094fde096dafb62d4c229fb25764ec8365b5f74f07f2eff856d4643066

Observation c43499ee-d956-4d65-ab8b-c9462145ed13 · outbound

This paper cites A dis- tributional perspective on reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies A dis- tributional perspective on reinforcement learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.984857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:8fe8991a2e20451c898fcffcdea210e0cc2ef627219c6d135cdda89b6a1ff321

Observation 2c291311-8889-4673-b059-44ecf5f1b6c6 · outbound

This paper cites Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Offline Q-Learning on Diverse Multi-Task Data Both Scales And Generalizes

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.319056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:e4197efa279a859890795abaafe218a58a02020eb268f3744c5c686b65e8e3fa

Observation c74f5f11-bd6d-4522-8274-9e8f19b811fa · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.268543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:b6df8eaa99aac6805cbadf8a43306716acb0d734a649dbe2879a11e5492b552c

Observation d5fa6e00-a2c0-46c4-81bd-6d9de680ed09 · outbound

This paper cites Energy-weighted flow matching for offline reinforcement learning.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Energy-weighted flow matching for offline reinforcement learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.969944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d2269db2cf6d3fb9aa3dc714707920c3420397ca14e0b8dfa12ddae0c467cf36

Observation 9e25219c-802b-430f-b70e-b129003c16bb · outbound

This paper cites Gemma 3 technical report.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Gemma 3 technical report

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.971846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:523f5b80d3bb60747a8a83adf3533ecb9d84f1937ccfbb33d0eab7f3c17491a0

Observation 7b0b1d20-208b-4754-8474-657d6830b089 · outbound

This paper cites Sigmoid loss for language image pre-training.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Sigmoid loss for language image pre-training

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.949397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:682e67e84740252981eaece60735254b253f423dd64e39475a2194736041f0a6

Observation 0a6dd454-d857-412b-8905-d434933e510d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:15:32.248894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:2860717ac5877d30125fbca3d3fc9026dee86a0a158110efd994794020f831e9

Observation a91f9f70-303e-4390-a230-56c7157b4ba8 · outbound

This paper cites Vision trans- formers for dense prediction.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Vision trans- formers for dense prediction

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.017860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:ebf9c60d6ffc7a342efef4b518108f41c21e148d3ba2aca864860e4e1d42f5b2

Observation 38067fb4-b886-4c5c-8678-4df8ad88fafe · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.008267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:f0aaf32026fad3755a6dd730dd94905448473a2076c43d5219fb38e56f7cde36

Observation 0b80b7f0-6563-4815-853a-8376770dd978 · outbound

This paper cites Decoupled weight decay regularization.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Decoupled weight decay regularization

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.002272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:b0504b9eed91da6de3c8dfaa55653599c2919ba5bb14a40c7b33d29cbb3fef5e

Observation bb11bb1a-f882-465e-9f6d-256642ab8e5f · outbound

This paper cites In our real- robot experiments, we useK= 201atoms over[−0.1,1.1].

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies In our real- robot experiments, we useK= 201atoms over[−0.1,1.1]

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.025765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:31c50da74391c85add9a9c26974116f7251531bb9fc58c260f15dd6f67461b23

Observation fdb24c82-a66f-421d-9ad7-90692f548c5f · outbound

This paper cites an unresolved cited work.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Unresolved cited work

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:37.021848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:e567296a4a0d052b1c8ad3e6cf03b5a4c076467c1e5f99c615d47537badd14d8

Observation 1492b493-37db-4baf-b0a7-485dec5690ce · outbound

This paper cites an unresolved cited work.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Unresolved cited work

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.951626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d867d4f798424967441e3c6f4e4a366c398916c4d41d5e576a44338e0609acea

Observation 954c73f3-0eb1-4fad-ab9d-325d18f39c9b · outbound

This paper cites Demonstrations are successful trajectories, rollouts contain both successes and failures, and play data is treated as unsuccessful exploratory data.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Demonstrations are successful trajectories, rollouts contain both successes and failures, and play data is treated as unsuccessful exploratory data

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.940526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:b6981ac95616c7ce40a09b54a01a8799e592166182b1d4d7724f63e27acd8af4

Observation 58af25cc-d75f-4e8c-b548-e1ff60d54e12 · outbound

This paper cites The policy is optimized with AdamW [63] using a base learning rate of2×10 −5 and a cosine decay schedule.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies The policy is optimized with AdamW [63] using a base learning rate of2×10 −5 and a cosine decay schedule

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.995805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:dbd2226ada3da212a0925ae4c4a4f6d96bda02700f06e54f4531705ee5b9035f

Observation c18fac3f-1573-45d4-be3e-c108997614bd · outbound

This paper cites an unresolved cited work.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Unresolved cited work

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.956522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:274515ccd08e1f63808c7033997d5b9c5ba4c3babe560c5d7dd9b8d65b247bc9

Observation 396975de-941d-452b-8b0c-deb512f0c24a · outbound

This paper cites The model is trained with a flow-matching loss, where the interpolated noisy actiona w is defined in Eq.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies The model is trained with a flow-matching loss, where the interpolated noisy actiona w is defined in Eq

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-07-06T13:42:36.961158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:d1d6c4c74ee36c7160054f110789b497f31b4c31e40fc5c585f446e6463cc80b

Observation 18f88483-d5e8-43b7-b1f3-9a39d2806629 · outbound

This paper cites The comparison isolates the Robot 1 Robot 2.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies The comparison isolates the Robot 1 Robot 2

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.991664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:e75eea4750cee214f152e23a3dcbed4281cde3832d3db164d8e1ee7f67e0f0fc

Observation 595baa96-96fe-431c-bcb7-16445fcc92d6 · outbound

This paper cites 9 vi- sualizes the predicted value distributions for the same episodes shown in Fig.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies 9 vi- sualizes the predicted value distributions for the same episodes shown in Fig

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.954173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:127b143d7a8a5d043f0d69a0725666af51facce974759953fd734cbd876a8d61

Observation 09c10d04-97e4-410b-9d6e-efb8489b620d · outbound

This paper cites (i) Object-storage uploads commit atomically (read- ers see either the fully-uploaded payload or no object) and are retried until persisted.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies (i) Object-storage uploads commit atomically (read- ers see either the fully-uploaded payload or no object) and are retried until persisted

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.982931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:46402d5e9f55a77d373e8bdbc1c5d7e12e04c9f93549354ab968f80d18c0715c

Observation d87dbcdb-4e25-47a0-aec3-82a14d10eb4a · outbound

This paper cites Table VI reports both on the same 8-hour, 16-actor run as the End-to-End Reliability subsection above.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies Table VI reports both on the same 8-hour, 16-actor run as the End-to-End Reliability subsection above

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T13:42:36.958802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:35f275a8d43835f51c7ba75f6e1b4cf4d85e694fa768ead8a1c2a4e71cc2d18e

Pith citing papers

Observation 6de04dfe-12cf-4fbe-8d39-258337ca15b7 · inbound

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning cites this paper.

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:58:02.685056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-27T09:46:59.746745Z digest=sha256:422eb9febea022d1d893387e2b51989e7572aaa82cb7183c40c16ec8654f6490

Observation 42d737be-edc7-42de-a44c-5013aaaf9047 · inbound

FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation cites this paper.

FlowDPG: Deterministic Policy Gradient on Flow Matching Policies for Real-World Manipulation Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:59:42.035183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-21T06:31:05.380196+00:00.

source=pdf_text observed=2026-06-26T10:50:24.868343Z digest=sha256:922e31be9fb50868d05c7b382900766f5a5e78bd198239dd5dcf36bb81513dd9

Observation 30afc46d-77f9-40f2-8f24-ea6c22132644 · inbound

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning cites this paper.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:31cac25aaebf704eb11c89e0e5a7efba6bd4b5b8a4c4d2cbedebc0110627250d