Pith. sign in

Paper Citation Record · LEDGER

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2607.09866.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09866 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T14:57:49.542416Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T12:43:44.697996Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50263556-4fdd-426a-a44b-99cbef037a72 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:34bc8db6f9dda6cb24ad3b1f4776dcfc8136c12b2d0fde25ff99e71eacab3db7

Observation 051a79f4-edb9-404c-9c5c-746357829210 · outbound

This paper cites Conservative q-learning for offline reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Conservative q-learning for offline reinforcement learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:4d997da6a3d47a6d257c920e18f6e84135afdc263e6a1aba4743521f3abc9fd4

Observation 623b2441-4071-4c53-b7c7-3d43ad02936c · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:0adc2006bc8db5b12f5c2ebe7fda444da654baa96eea9a1dfbbb1519d3b9a7bb

Observation 0f2ffcc6-3379-4e72-91c1-cb495c81f928 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning RT-1: Robotics Transformer for Real-World Control at Scale,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:3ce0374fe69d113c456e1f426b5ea0cad158e73a7027677d4c2ec6642cfb7a86

Observation d50bb420-ea63-497a-bb52-ee7f17ec3121 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:fa99f10252d3ce6de0577114208e8ab86c51fd55c728c354c5f8df8ba8260368

Observation 8445b5e4-8191-43f2-a1fc-1bc1ea8050e9 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:980806d265d8a1a00b71ce6ee4030561bb1b718b2e7f4c38c7df71704a9d8f3d

Observation 9b397f97-aa47-4945-8f6e-221215a5d993 · outbound

This paper cites Openvla: An open-source vision-language- action model,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Openvla: An open-source vision-language- action model,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:6825c707b296b4983b0a6e41a3053eb5b90edcde39c966b532e290769a543a09

Observation 03721b18-45f2-40e6-a1d3-f46270a7da22 · outbound

This paper cites π0: A Vision-Language-Action Flow Model for General Robot Control,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning π0: A Vision-Language-Action Flow Model for General Robot Control,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:1ecd991afb075c930794978b4d7e72fc9f375bedceb1514224768c436139c173

Observation 84616b7e-f9f1-429e-a5e2-0640efb96b00 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:f4ed548d9a40506f95475ca92b36bf0b14604fb822e4701f031da3d2edec36e2

Observation a75baee0-3cc4-416e-bc08-7f9b85f4ba01 · outbound

This paper cites Pre-Training for Robots: Offline RL Enables Learning New Tasks in a Handful of Trials,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Pre-Training for Robots: Offline RL Enables Learning New Tasks in a Handful of Trials,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:cd6140e42e3201da920b6e0c019d160a6a422cb53bafc7ad99bb6f92ac1801f4

Observation b3ef8537-bbc0-4c71-8f3c-a1a2c3f39bf8 · outbound

This paper cites Robogene: Boosting vla pre-training via diversity-driven agentic framework for real-world task generation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Robogene: Boosting vla pre-training via diversity-driven agentic framework for real-world task generation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:d62188b14ab8079e75e6a0fdec3a19fb7d66ebac935f9c292e5dcd08f71227c0

Observation 559fdc21-c558-4510-a35f-b9c28a9b35ad · outbound

This paper cites Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:6539314566c6782d7ec736d9151f726c6581c92e17fe74aba668a25b52def984

Observation 6526585b-7244-4585-9e09-75c595bd4552 · outbound

This paper cites Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:da80b293764fd0970e60d7be27f222b4d4ab574142b94eea4fead770925534d4

Observation d624c465-0148-49aa-87d7-86dc0e5cd2a4 · outbound

This paper cites SimpleVLA-RL: Scaling VLA training via reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning SimpleVLA-RL: Scaling VLA training via reinforcement learning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:bcac79a5d977b6cdc5b28ada82059c1cd6e1f38d4b8b68cea61875831688bdf5

Observation 83f35c30-b2bb-4d69-af1c-60065ff9e5c6 · outbound

This paper cites Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:98e36c8b0a149d0d3fb34fdda5f033cdee1cf7c83fb5c7bebc2555f09e7c3f01

Observation f2045532-314d-4c85-a6ca-4600ffe36e44 · outbound

This paper cites Human-assisted robotic policy refinement via action preference optimization,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Human-assisted robotic policy refinement via action preference optimization,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:d492b94d63c225e8ec756b03b2e13710e6f30129f3a1024073167ee63b0b9a39

Observation 20bceca7-f507-43cf-a628-f4099186667d · outbound

This paper cites Geco-srt: Geometry-aware continual adaptation for cross-task sim-to-real transfer,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Geco-srt: Geometry-aware continual adaptation for cross-task sim-to-real transfer,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:a1d897b0cf6bccfdcc630619f5990b8a3d3fbbe9d5bd658df200095f1cf27723

Observation 00a3e352-de63-40a0-83fe-51e1bd0933c4 · outbound

This paper cites Deep reinforcement learning that matters,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Deep reinforcement learning that matters,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:6c267c2d5c3ff74f70778aa8a7237f39e6af9a738a5e467d7179f6383bbc7e2b

Observation 1e02aa08-7bf9-43d5-a0a0-13023ab3e298 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Deep reinforcement learning at the edge of the statistical precipice,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:a56e5ea7b32589eb2bc2d5b844c41ad8e9798060cf3f80e4ddaf5293c1498e46

Observation 30afc46d-77f9-40f2-8f24-ea6c22132644 · outbound

This paper cites Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:5853bb96fdb22580899248a1fb3125b42e659046a4a639a6b7122698b5a69e41

Observation 9673a0d2-f7e2-4726-9280-033af1e9cb7d · outbound

This paper cites Suf: Stabilized unconstrained fine-tuning for offline-to-online reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Suf: Stabilized unconstrained fine-tuning for offline-to-online reinforcement learning,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:d875c157890690fe25eba1d51a4c0673a22980293831d57692120a83c498eea4

Observation 08d22497-249b-4301-9a13-2298ccc79f4f · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:ae800bec1af8c266cade81c934d9dce04b20314b954435d51242520d1a3c8e6d

Observation c96d70e9-1539-4d9a-aae2-51cb3d61555f · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:3ccf6c4480d7cc1c6ef558fa0fbeccf863d162a4540f80e3b4ef329633489208

Observation bec0c268-5d8b-4cce-ac01-c31138ed87c9 · outbound

This paper cites What matters in learning from offline human demonstrations for robot manipulation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning What matters in learning from offline human demonstrations for robot manipulation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:302baf31f0f61d303d7f46fdcbd5363fb968a19625fdafc5724dace565bd10cf

Observation d1b91834-faa6-4088-ba45-e858cee524ea · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Scalable deep reinforcement learning for vision-based robotic manipulation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:0bd078e80e717716996c13e9bd515768598390eee65b90249ee6a7087fdf2a4c

Observation b6f568c5-bf70-437c-bc2a-3641a8206244 · outbound

This paper cites Gr-rl: Going dexterous and precise for long-horizon robotic manipulation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Gr-rl: Going dexterous and precise for long-horizon robotic manipulation,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:2367e28c7f9deae0e67cafa7ca5172847979c427d7ce0803a4dd9e74e1227552

Observation e193dcf8-6894-4db5-a786-093a6109a036 · outbound

This paper cites VIP: Towards universal visual reward and representation via value-implicit pre-training,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning VIP: Towards universal visual reward and representation via value-implicit pre-training,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:70cf991c6d38424bff3299f1047354e01e809960303c762583225e075f8c2028

Observation f44fa275-83f7-426b-8b2b-95ead89734fd · outbound

This paper cites Conrft: A reinforced fine-tuning method for vla models via consistency policy,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Conrft: A reinforced fine-tuning method for vla models via consistency policy,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:fad54f63d758f5325e3d2d734df3c3202d66210bd0a056319afb2c2528a53fef

Observation a9070b42-f7ad-4792-a1c6-c43588c3b3d6 · outbound

This paper cites Consistency policy: Accelerated visuomotor policies via consistency distillation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Consistency policy: Accelerated visuomotor policies via consistency distillation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:dd17a02438a2cd55974df13241a5ddd29428472f75f23a97eba9b88ab88c4c85

Observation fe0da0c2-cabb-49b3-b450-035efd82fe4c · outbound

This paper cites Learning from imperfect demonstrations from agents with varying dynamics,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning from imperfect demonstrations from agents with varying dynamics,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:ed7d8aca3040901162adc78d9e3d93e976028aa821c1c9ef80e2b10aa3710cf4

Observation e5838b69-9958-44a7-b038-f18b9a758597 · outbound

This paper cites Curating Demonstrations using Online Experience,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Curating Demonstrations using Online Experience,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:57a91cf140345d42471a12b8503203663f52a31f4b9d2e1b0eaaf21ac9fa4ca0

Observation b63129a2-b76e-4db5-8709-cb0878d0d636 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Octo: An Open-Source Generalist Robot Policy,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:2ab45aef4d4abff591792dc6b5869238bfbf93c7b6fbd7d65d134b8441460755

Observation 9cad58b3-f1b7-4482-996c-1e9d086702ba · outbound

This paper cites Xr-1: Towards versatile vision- language-action models via learning unified vision-motion representations,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Xr-1: Towards versatile vision- language-action models via learning unified vision-motion representations,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:6fd5be5df2960ba8262cd8e259bb5efa5191eeaae81a8368549b3f75d985c553

Observation 03b5fe24-5f32-4c28-886b-63ce8d60a4c2 · outbound

This paper cites A survey on offline reinforcement learning: Taxonomy, review, and open problems,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning A survey on offline reinforcement learning: Taxonomy, review, and open problems,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:fe0fca8504ce07c339aeba64a1b21f7575df37c159ea5ff4926240432de9d135

Observation 20092fc2-7cd1-4b61-ac37-8abc71913b09 · outbound

This paper cites Offline reinforcement learning with implicit q-learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Offline reinforcement learning with implicit q-learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:943b1dc9e476477ddebb376513d9ae3b9e42ab699fd224768311f450954fce93

Observation 8f15ee8c-276a-4511-83e5-30b498f521ce · outbound

This paper cites Policy expansion for bridging offline-to-online reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Policy expansion for bridging offline-to-online reinforcement learning,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:f826c06da5e665767bed35ce055628e89caf114541e70f63b26413b0eb82619b

Observation 70cea663-cafb-4c94-a5a9-5709e7c4cba8 · outbound

This paper cites Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:b84990ecb2962d01830ee7894c7d7a31e65ebc13294adec92a381dd2c29ce56d

Observation ca81e43d-d058-45b3-b7a0-2de550c049aa · outbound

This paper cites Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:237af6970156dfa56db59da8dbbd208cb206145a44662ecf6051d52712a04d5b

Observation 017da096-c3f4-4e4d-b475-bf100d9f0af6 · outbound

This paper cites Learning to predict by the methods of temporal differences,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning to predict by the methods of temporal differences,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:05433bdb7b5dfe58787e9c3b5735e77d3b71ebd82f0aee162caa259de545d746

Observation 4b495f60-d9ed-4301-a628-679387c354db · outbound

This paper cites Analysis of temporal-diffference learning with function approximation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Analysis of temporal-diffference learning with function approximation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:2b20f8d51e7bda62f99d4f0b2bbadb09e3fdb1c1a18ea0faa9dbc8227591b9cf

Observation a43b8e8c-b242-4990-aca5-fcafd790b9cc · outbound

This paper cites Fast gradient-descent methods for temporal-difference learning with linear function approximation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Fast gradient-descent methods for temporal-difference learning with linear function approximation,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:85d8f3285c2df61ad26b0d7f397a34ed6cbc41d451e92b452110bd65ac46b41e

Observation d1605b3e-a32b-4eaa-9741-10af9dab2970 · outbound

This paper cites Human-level control through deep reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Human-level control through deep reinforcement learning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:af4ad7c1a6be176c39467487a04ae8f8e21801bb4c8fdcf690094ebe26999eef

Observation 241cd020-8f64-4623-8456-a3dfb70ce679 · outbound

This paper cites Benchmarking deep reinforcement learning for continuous control,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Benchmarking deep reinforcement learning for continuous control,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:e15f7f8283ada68066174f52e228b1fc4cc41bc96f08e4bfdebde4eec535d258

Observation 624625ae-12fc-4932-a403-aa7e35dca02f · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:29c6f1096940cb36d5aeae9c661612b9bc12d0f0f52927369f6e09931c71b345

Observation 3d6d0eaa-823d-4128-9f18-ba2cce5b14e1 · outbound

This paper cites A vision-language-action-critic model for robotic real-world reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning A vision-language-action-critic model for robotic real-world reinforcement learning,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:74fa29628b7c1e9543011f776f5f44285829cf9ff3058a4869cd76081413a517

Observation ba35709e-b3b6-4f60-bdf0-0ea88ec5b6a9 · outbound

This paper cites RL-VLM-f: Reinforcement learning from vision lan- guage foundation model feedback,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning RL-VLM-f: Reinforcement learning from vision lan- guage foundation model feedback,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:7d97dd2b71c0bef1857f8843d12b54b99794d4f9a6a24c754e97da0ebeaa719b

Observation b79ed692-067a-4099-ad10-64436323d7fb · outbound

This paper cites Rank2reward: Learning shaped reward functions from passive video,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Rank2reward: Learning shaped reward functions from passive video,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:22644f163e93ac388d067c4655f8ac63fdb698733307a74a6ee7b3a78d63dc2e

Observation b56a7533-ef9d-4c49-9dee-33ddf266a7cf · outbound

This paper cites Vision language models are in-context value learners,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Vision language models are in-context value learners,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:bb8fef2662962407b850068c0b093522569f8dc48a06f31d4773659d68227080

Observation 04c12ceb-5905-4279-b561-25c82c8f7040 · outbound

This paper cites Universal value function approximators,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Universal value function approximators,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:d4e3945fa74dcb416a3a78a87b1888bed9f661eb396f2526b237494407654787

Observation 89aca5ee-1b6f-4425-82bb-501a60a67579 · outbound

This paper cites ViVa: A Video-Generative Value Model for Robot Reinforcement Learning.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:7f7134be0e941f4fd15c3defb2271ddea3a6b5a2d3f962863b8c978803912fc0

Observation f395872d-654a-471d-9996-46b1785f3507 · outbound

This paper cites World Value Models for Robotic Manipulation.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning World Value Models for Robotic Manipulation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:a7ec92c67567962e2e5f11125901480805893195a3e5fcadc64571329789d56a

Observation 72620089-25b1-4214-b74b-ca1710b1e33f · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning PaliGemma: A versatile 3B VLM for transfer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:87948fb6a35276e99856c2c90390e6b7aad6c53d7021a7fe80300f5039f51963

Observation 1d1623da-494e-4b9a-924c-b9adf44cce9b · outbound

This paper cites Sigmoid loss for language image pre-training,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Sigmoid loss for language image pre-training,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:bfa712b3f7a67c85e1e8acd545357fa653c3e8606435d9e1b319b7000446a79d

Observation 21da14cf-bbe1-453e-b7ff-e99db9d5d115 · outbound

This paper cites Perceiver: General perception with iterative attention,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Perceiver: General perception with iterative attention,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:238b02e0de40f036f8db8d77e153260ebf4984e5bdef3eb26ebd5294892a7b23

Observation 5dcfd6e7-403b-45c6-a4b6-1fdd5f111023 · outbound

This paper cites Stop regressing: Training value functions via classification for scalable deep RL,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Stop regressing: Training value functions via classification for scalable deep RL,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:1d0376b35c60a42ce9ea983b571a24fc9806a9116375345b555d14ae40614c7a

Observation 803bbc2e-3f95-4b7c-b088-455ba9a92ae3 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning A reduction of imitation learning and structured prediction to no-regret online learning,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:8ebd0bf6944805bc5707040b0ca2fef5e623bb79edc583214ad8c1439ed50403

Pith citing papers

Observation 6b1a3c85-36e6-4578-8995-10c26abf813b · inbound

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens cites this paper.

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-30T12:43:44.697996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:43:44.697996Z digest=sha256:a0f07cca6e768619cd4ff1944764f212a96f8a60d4892de3614f3780ed58fcac