Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T14:57:49.542416Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2607.09866.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T14:57:49.542416Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-30T12:43:44.697996Z
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 50263556-4fdd-426a-a44b-99cbef037a72 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 051a79f4-edb9-404c-9c5c-746357829210 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Conservative q-learning for offline reinforcement learning,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623b2441-4071-4c53-b7c7-3d43ad02936c · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f2ffcc6-3379-4e72-91c1-cb495c81f928 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning RT-1: Robotics Transformer for Real-World Control at Scale,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d50bb420-ea63-497a-bb52-ee7f17ec3121 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Rt-2: Vision-language-action models transfer web knowledge to robotic control,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8445b5e4-8191-43f2-a1fc-1bc1ea8050e9 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b397f97-aa47-4945-8f6e-221215a5d993 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Openvla: An open-source vision-language- action model,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03721b18-45f2-40e6-a1d3-f46270a7da22 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning π0: A Vision-Language-Action Flow Model for General Robot Control,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84616b7e-f9f1-429e-a5e2-0640efb96b00 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a75baee0-3cc4-416e-bc08-7f9b85f4ba01 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Pre-Training for Robots: Offline RL Enables Learning New Tasks in a Handful of Trials,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ef8537-bbc0-4c71-8f3c-a1a2c3f39bf8 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Robogene: Boosting vla pre-training via diversity-driven agentic framework for real-world task generation,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 559fdc21-c558-4510-a35f-b9c28a9b35ad · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6526585b-7244-4585-9e09-75c595bd4552 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d624c465-0148-49aa-87d7-86dc0e5cd2a4 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning SimpleVLA-RL: Scaling VLA training via reinforcement learning,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f35c30-b2bb-4d69-af1c-60065ff9e5c6 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2045532-314d-4c85-a6ca-4600ffe36e44 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Human-assisted robotic policy refinement via action preference optimization,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20bceca7-f507-43cf-a628-f4099186667d · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Geco-srt: Geometry-aware continual adaptation for cross-task sim-to-real transfer,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00a3e352-de63-40a0-83fe-51e1bd0933c4 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Deep reinforcement learning that matters,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e02aa08-7bf9-43d5-a0a0-13023ab3e298 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Deep reinforcement learning at the edge of the statistical precipice,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30afc46d-77f9-40f2-8f24-ea6c22132644 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9673a0d2-f7e2-4726-9280-033af1e9cb7d · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Suf: Stabilized unconstrained fine-tuning for offline-to-online reinforcement learning,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08d22497-249b-4301-9a13-2298ccc79f4f · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c96d70e9-1539-4d9a-aae2-51cb3d61555f · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning $\pi^{*}_{0.6}$: a VLA That Learns From Experience
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bec0c268-5d8b-4cce-ac01-c31138ed87c9 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning What matters in learning from offline human demonstrations for robot manipulation,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b91834-faa6-4088-ba45-e858cee524ea · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Scalable deep reinforcement learning for vision-based robotic manipulation,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6f568c5-bf70-437c-bc2a-3641a8206244 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Gr-rl: Going dexterous and precise for long-horizon robotic manipulation,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e193dcf8-6894-4db5-a786-093a6109a036 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning VIP: Towards universal visual reward and representation via value-implicit pre-training,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f44fa275-83f7-426b-8b2b-95ead89734fd · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Conrft: A reinforced fine-tuning method for vla models via consistency policy,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9070b42-f7ad-4792-a1c6-c43588c3b3d6 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Consistency policy: Accelerated visuomotor policies via consistency distillation,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0da0c2-cabb-49b3-b450-035efd82fe4c · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning from imperfect demonstrations from agents with varying dynamics,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5838b69-9958-44a7-b038-f18b9a758597 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Curating Demonstrations using Online Experience,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b63129a2-b76e-4db5-8709-cb0878d0d636 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Octo: An Open-Source Generalist Robot Policy,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cad58b3-f1b7-4482-996c-1e9d086702ba · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Xr-1: Towards versatile vision- language-action models via learning unified vision-motion representations,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03b5fe24-5f32-4c28-886b-63ce8d60a4c2 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning A survey on offline reinforcement learning: Taxonomy, review, and open problems,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20092fc2-7cd1-4b61-ac37-8abc71913b09 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Offline reinforcement learning with implicit q-learning,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f15ee8c-276a-4511-83e5-30b498f521ce · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Policy expansion for bridging offline-to-online reinforcement learning,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70cea663-cafb-4c94-a5a9-5709e7c4cba8 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca81e43d-d058-45b3-b7a0-2de550c049aa · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 017da096-c3f4-4e4d-b475-bf100d9f0af6 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning to predict by the methods of temporal differences,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b495f60-d9ed-4301-a628-679387c354db · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Analysis of temporal-diffference learning with function approximation,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a43b8e8c-b242-4990-aca5-fcafd790b9cc · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Fast gradient-descent methods for temporal-difference learning with linear function approximation,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1605b3e-a32b-4eaa-9741-10af9dab2970 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Human-level control through deep reinforcement learning,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241cd020-8f64-4623-8456-a3dfb70ce679 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Benchmarking deep reinforcement learning for continuous control,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 624625ae-12fc-4932-a403-aa7e35dca02f · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d6d0eaa-823d-4128-9f18-ba2cce5b14e1 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning A vision-language-action-critic model for robotic real-world reinforcement learning,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba35709e-b3b6-4f60-bdf0-0ea88ec5b6a9 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning RL-VLM-f: Reinforcement learning from vision lan- guage foundation model feedback,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b79ed692-067a-4099-ad10-64436323d7fb · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Rank2reward: Learning shaped reward functions from passive video,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b56a7533-ef9d-4c49-9dee-33ddf266a7cf · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Vision language models are in-context value learners,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04c12ceb-5905-4279-b561-25c82c8f7040 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Universal value function approximators,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89aca5ee-1b6f-4425-82bb-501a60a67579 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning ViVa: A Video-Generative Value Model for Robot Reinforcement Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f395872d-654a-471d-9996-46b1785f3507 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning World Value Models for Robotic Manipulation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72620089-25b1-4214-b74b-ca1710b1e33f · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning PaliGemma: A versatile 3B VLM for transfer
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d1623da-494e-4b9a-924c-b9adf44cce9b · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Sigmoid loss for language image pre-training,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21da14cf-bbe1-453e-b7ff-e99db9d5d115 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Perceiver: General perception with iterative attention,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dcfd6e7-403b-45c6-a4b6-1fdd5f111023 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Stop regressing: Training value functions via classification for scalable deep RL,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803bbc2e-3f95-4b7c-b088-455ba9a92ae3 · outbound
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning A reduction of imitation learning and structured prediction to no-regret online learning,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1a3c85-36e6-4578-8995-10c26abf813b · inbound
$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.