Pith. sign in

Paper Citation Record · LEDGER

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

As of 22 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2607.09866.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09866 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T14:57:49.542416Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T12:43:44.697996Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50263556-4fdd-426a-a44b-99cbef037a72 · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:f2b8948bea2d815650fd9fd1b83e02783637ca622b31177a9ca986ea172ef7e2

Observation 051a79f4-edb9-404c-9c5c-746357829210 · outbound

This paper cites Conservative q-learning for offline reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Conservative q-learning for offline reinforcement learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:75e3d498b434aeccd0f6757b8866fb28398b004d6666e29eb78a1dd90bda75e5

Observation 623b2441-4071-4c53-b7c7-3d43ad02936c · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:362c8f1bf4105476fd155b2b8556222646a696b17acc792ce750c4b98d27f0f1

Observation 0f2ffcc6-3379-4e72-91c1-cb495c81f928 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning RT-1: Robotics Transformer for Real-World Control at Scale,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:6813a25f0da691b9d0d9c67921c3026e2c1bd60fe6682e337b1b88260a06d754

Observation d50bb420-ea63-497a-bb52-ee7f17ec3121 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Rt-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:7e04989ee9ef2a082ee043ae43fd57d6607231c3d0009a3111cd65af89a0c3a6

Observation 8445b5e4-8191-43f2-a1fc-1bc1ea8050e9 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:26f80d2c9bf49278d550a4c37b68fc74861d62143fd71c2db8020af1bdd927f8

Observation 9b397f97-aa47-4945-8f6e-221215a5d993 · outbound

This paper cites Openvla: An open-source vision-language- action model,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Openvla: An open-source vision-language- action model,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:44737998bbd4c4d06db18bafb0163b04b0251e06e5b660b6d6f466f7a22b7ed1

Observation 03721b18-45f2-40e6-a1d3-f46270a7da22 · outbound

This paper cites π0: A Vision-Language-Action Flow Model for General Robot Control,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning π0: A Vision-Language-Action Flow Model for General Robot Control,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:5e430b8980f970ef0fdb49557cae5c190dea5a147d9f23059b15ecce379c0163

Observation 84616b7e-f9f1-429e-a5e2-0640efb96b00 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:ef4ea733066e108e8d9222694d09e24c2bf5e6324fcbcd85601d571adfa02902

Observation a75baee0-3cc4-416e-bc08-7f9b85f4ba01 · outbound

This paper cites Pre-Training for Robots: Offline RL Enables Learning New Tasks in a Handful of Trials,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Pre-Training for Robots: Offline RL Enables Learning New Tasks in a Handful of Trials,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:223e0dfe9df0cc5c05ea43278ab71576b5fcc9975ab2b0897466d48a1e76189e

Observation b3ef8537-bbc0-4c71-8f3c-a1a2c3f39bf8 · outbound

This paper cites Robogene: Boosting vla pre-training via diversity-driven agentic framework for real-world task generation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Robogene: Boosting vla pre-training via diversity-driven agentic framework for real-world task generation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:4f1a2ea0911ec8e7da932fe738e2075cf130ffbbea5158d5386398a8b7795022

Observation 559fdc21-c558-4510-a35f-b9c28a9b35ad · outbound

This paper cites Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:9df40dd931f6ade8c1f0ba9afc50e4e62e0f002ffd59e0771647d201504b86fa

Observation 6526585b-7244-4585-9e09-75c595bd4552 · outbound

This paper cites Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Robot fine-tuning made easy: Pre-training rewards and policies for autonomous real-world reinforcement learning,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:f5ffbd418b147c31f03f804948fa76437beac36b178bb0533b4078c0d640412f

Observation d624c465-0148-49aa-87d7-86dc0e5cd2a4 · outbound

This paper cites SimpleVLA-RL: Scaling VLA training via reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning SimpleVLA-RL: Scaling VLA training via reinforcement learning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:709c18ebb747e999e88f2c9d412d2c4af85304ea35f20b88de10e6cc848275d6

Observation 83f35c30-b2bb-4d69-af1c-60065ff9e5c6 · outbound

This paper cites Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Phoenix: A motion-based self-reflection framework for fine-grained robotic action correction,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:d0144ede5e805e373b61d3e8e3a50de5fcf0f492a87deda7a66487c2b4120e38

Observation f2045532-314d-4c85-a6ca-4600ffe36e44 · outbound

This paper cites Human-assisted robotic policy refinement via action preference optimization,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Human-assisted robotic policy refinement via action preference optimization,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:07873b310067c59d49e2cedff133ef4aebae5f55e35673cddddbbec1e29878da

Observation 20bceca7-f507-43cf-a628-f4099186667d · outbound

This paper cites Geco-srt: Geometry-aware continual adaptation for cross-task sim-to-real transfer,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Geco-srt: Geometry-aware continual adaptation for cross-task sim-to-real transfer,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:615fcd4e3c22d0a511c4fb9b2df446b7ee6d1248d85a8c4c8d3a74efb4bcb101

Observation 00a3e352-de63-40a0-83fe-51e1bd0933c4 · outbound

This paper cites Deep reinforcement learning that matters,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Deep reinforcement learning that matters,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:e02d6d228f1f8e4d5b41df65a770d14fc46088de093c04d518f8daaa194fce8a

Observation 1e02aa08-7bf9-43d5-a0a0-13023ab3e298 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Deep reinforcement learning at the edge of the statistical precipice,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:83422718d28996bfd1cfcd701d19e6044fa0b054f3c478c51fba3cd628a0a0c4

Observation 30afc46d-77f9-40f2-8f24-ea6c22132644 · outbound

This paper cites Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:d4c6fac16f7db324897e1ca3cbde4a70984c787ed094c36ca54797647bccc10a

Observation 9673a0d2-f7e2-4726-9280-033af1e9cb7d · outbound

This paper cites Suf: Stabilized unconstrained fine-tuning for offline-to-online reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Suf: Stabilized unconstrained fine-tuning for offline-to-online reinforcement learning,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:e21de2eeb494e4320824616fa65d6a1a627c8e91dae6ddb3ba41971616a97e5e

Observation 08d22497-249b-4301-9a13-2298ccc79f4f · outbound

This paper cites Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:ff3dca140a59454aba00bad915cb87110a0f2dad36d21d3fcc67111735561c0a

Observation c96d70e9-1539-4d9a-aae2-51cb3d61555f · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:bdbdca90ee2b311ef8a2b4b07fe1c31823586059a8cff11df0d1676eafc5ae06

Observation bec0c268-5d8b-4cce-ac01-c31138ed87c9 · outbound

This paper cites What matters in learning from offline human demonstrations for robot manipulation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning What matters in learning from offline human demonstrations for robot manipulation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:67031b462d763ae3a7ec8a63eee1cbb54c1c9072f214f55de9fcb8de79694bbc

Observation d1b91834-faa6-4088-ba45-e858cee524ea · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Scalable deep reinforcement learning for vision-based robotic manipulation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:aa77a11ce57d4a193977be8fd292257fe6e8f47afe6065913c7f58f537994f22

Observation b6f568c5-bf70-437c-bc2a-3641a8206244 · outbound

This paper cites Gr-rl: Going dexterous and precise for long-horizon robotic manipulation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Gr-rl: Going dexterous and precise for long-horizon robotic manipulation,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:e6abdcc2ec352e8de12a4c8a9c1ccdc235a40cd63e231d72375e74fe7ad2e141

Observation e193dcf8-6894-4db5-a786-093a6109a036 · outbound

This paper cites VIP: Towards universal visual reward and representation via value-implicit pre-training,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning VIP: Towards universal visual reward and representation via value-implicit pre-training,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:7a75dda96a84870c389b1e831b1fd0d74d34140fdd3f90c46e1a78886649673b

Observation f44fa275-83f7-426b-8b2b-95ead89734fd · outbound

This paper cites Conrft: A reinforced fine-tuning method for vla models via consistency policy,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Conrft: A reinforced fine-tuning method for vla models via consistency policy,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:05287f913818793244cd1f3fc8ace201cca54fd917d6bd42e8837a3e62957fd6

Observation a9070b42-f7ad-4792-a1c6-c43588c3b3d6 · outbound

This paper cites Consistency policy: Accelerated visuomotor policies via consistency distillation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Consistency policy: Accelerated visuomotor policies via consistency distillation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:53bde172973d8d886b7ca9b7be4cef5e8c649cb5854913ea8f3525a582335ac8

Observation fe0da0c2-cabb-49b3-b450-035efd82fe4c · outbound

This paper cites Learning from imperfect demonstrations from agents with varying dynamics,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning from imperfect demonstrations from agents with varying dynamics,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:a59b48da44702e231000e63c4deb337587eaedc3de17a7c3d3c59b128dbd3256

Observation e5838b69-9958-44a7-b038-f18b9a758597 · outbound

This paper cites Curating Demonstrations using Online Experience,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Curating Demonstrations using Online Experience,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:6633b6ca075289bcd17b1771295f5ef79423a0ed0afb5435920e04fd028bba97

Observation b63129a2-b76e-4db5-8709-cb0878d0d636 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Octo: An Open-Source Generalist Robot Policy,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:64f20e36303f27a8c2474253ce01148e2be3ae5b990b7940b940699ddd8bf1ef

Observation 9cad58b3-f1b7-4482-996c-1e9d086702ba · outbound

This paper cites Xr-1: Towards versatile vision- language-action models via learning unified vision-motion representations,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Xr-1: Towards versatile vision- language-action models via learning unified vision-motion representations,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:ce940b4acd896f5ec2ee4aeaa891700c80dbc8b3a09a5c08c2fc252727acb45c

Observation 03b5fe24-5f32-4c28-886b-63ce8d60a4c2 · outbound

This paper cites A survey on offline reinforcement learning: Taxonomy, review, and open problems,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning A survey on offline reinforcement learning: Taxonomy, review, and open problems,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:724e223cec95f26dd8eb4d6c9ca5c803b9b980c6fb0adc29bdf5e0373b8b8686

Observation 20092fc2-7cd1-4b61-ac37-8abc71913b09 · outbound

This paper cites Offline reinforcement learning with implicit q-learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Offline reinforcement learning with implicit q-learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:47e42ccb3cc501dc336a0deeb0421fe39382d7e7fb2a9fb57ea21a93d8c5fe8e

Observation 8f15ee8c-276a-4511-83e5-30b498f521ce · outbound

This paper cites Policy expansion for bridging offline-to-online reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Policy expansion for bridging offline-to-online reinforcement learning,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:eab0d4af321429bc7b9508160eccbaa2ccd49c1300079cb2d4c9c3eb46153ffb

Observation 70cea663-cafb-4c94-a5a9-5709e7c4cba8 · outbound

This paper cites Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Q-transformer: Scalable offline reinforcement learning via autoregressive q-functions,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:179838468fa8f221e753eb5762257472c07e885e9cc66489f46e85346a9cb7e3

Observation ca81e43d-d058-45b3-b7a0-2de550c049aa · outbound

This paper cites Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Finetuning from Offline Reinforcement Learning: Challenges, Trade-offs and Practical Solutions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:72c41ee55f92ae266eacc4bd3a6f33b3ba41290fa0d548a9193774bf17a72f20

Observation 017da096-c3f4-4e4d-b475-bf100d9f0af6 · outbound

This paper cites Learning to predict by the methods of temporal differences,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Learning to predict by the methods of temporal differences,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:b62a392d11e26257b49f243653ae416e0fa135621ae3c71f51e0bb6111093899

Observation 4b495f60-d9ed-4301-a628-679387c354db · outbound

This paper cites Analysis of temporal-diffference learning with function approximation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Analysis of temporal-diffference learning with function approximation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:866cef57bd8c9f3b8ff22857e447cdeabec26073426490e299cd7d33995f362d

Observation a43b8e8c-b242-4990-aca5-fcafd790b9cc · outbound

This paper cites Fast gradient-descent methods for temporal-difference learning with linear function approximation,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Fast gradient-descent methods for temporal-difference learning with linear function approximation,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:19228b38533b94014444016e5d11ca116fc2512706d98a20c70c34b95ac8ac5b

Observation d1605b3e-a32b-4eaa-9741-10af9dab2970 · outbound

This paper cites Human-level control through deep reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Human-level control through deep reinforcement learning,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:ff20a3da7cb192bc728e75fe51279bc505ba548cfe2140e25fd63fce33cc3254

Observation 241cd020-8f64-4623-8456-a3dfb70ce679 · outbound

This paper cites Benchmarking deep reinforcement learning for continuous control,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Benchmarking deep reinforcement learning for continuous control,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:e1af2b426a73cac8b6f12b6fffdc26aa0006e2d6073b68360f2a7ea28c837616

Observation 624625ae-12fc-4932-a403-aa7e35dca02f · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:428102417d1e6918b7430f2c6b29dcfb2ef21624dcdbf4a282c4e31f67c6fbbc

Observation 3d6d0eaa-823d-4128-9f18-ba2cce5b14e1 · outbound

This paper cites A vision-language-action-critic model for robotic real-world reinforcement learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning A vision-language-action-critic model for robotic real-world reinforcement learning,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:6937d9dec76a390ef60a710b1f3c6a85c99378d10bc41dccf4cdae996d681796

Observation ba35709e-b3b6-4f60-bdf0-0ea88ec5b6a9 · outbound

This paper cites RL-VLM-f: Reinforcement learning from vision lan- guage foundation model feedback,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning RL-VLM-f: Reinforcement learning from vision lan- guage foundation model feedback,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:ded034209eda72f40d77052f123160514ac721b3c2591293558244a6943ccdaa

Observation b79ed692-067a-4099-ad10-64436323d7fb · outbound

This paper cites Rank2reward: Learning shaped reward functions from passive video,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Rank2reward: Learning shaped reward functions from passive video,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:b24471022b1837123543f5df9aed28ee44b59cb51a1e60933b79f59f37baa012

Observation b56a7533-ef9d-4c49-9dee-33ddf266a7cf · outbound

This paper cites Vision language models are in-context value learners,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Vision language models are in-context value learners,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:ed59e6f79638eb0d5867215758cf7c53d078f1f11cb849df1cb541d3b83d655d

Observation 04c12ceb-5905-4279-b561-25c82c8f7040 · outbound

This paper cites Universal value function approximators,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Universal value function approximators,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:7ca4d1e34339b435c40c8b28bdd34d41ff0e35122395411af082742e3477d550

Observation 89aca5ee-1b6f-4425-82bb-501a60a67579 · outbound

This paper cites ViVa: A Video-Generative Value Model for Robot Reinforcement Learning.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:ba1e0ccae894d8b88844dcd821b1ad9891c2a3fdb47ef0c586c06288556b7bad

Observation f395872d-654a-471d-9996-46b1785f3507 · outbound

This paper cites World Value Models for Robotic Manipulation.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning World Value Models for Robotic Manipulation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:d73737090542b595325bccdd256eae01703f187650e0f3bf23d9e187d5a3a711

Observation 72620089-25b1-4214-b74b-ca1710b1e33f · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning PaliGemma: A versatile 3B VLM for transfer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:e283e361a03be172163d3fa9a350ca020f86ddf71a8ca78f92f4d426689d3b99

Observation 1d1623da-494e-4b9a-924c-b9adf44cce9b · outbound

This paper cites Sigmoid loss for language image pre-training,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Sigmoid loss for language image pre-training,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:e178240021eeecc14560dbec0b6caa7cf5e4c06ea8f93e6cca4712ff2fef3083

Observation 21da14cf-bbe1-453e-b7ff-e99db9d5d115 · outbound

This paper cites Perceiver: General perception with iterative attention,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Perceiver: General perception with iterative attention,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:7f5271431ffc699528dc89d709a9e9017c3975f11c938497831026d64fbacaa1

Observation 5dcfd6e7-403b-45c6-a4b6-1fdd5f111023 · outbound

This paper cites Stop regressing: Training value functions via classification for scalable deep RL,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning Stop regressing: Training value functions via classification for scalable deep RL,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:a003f8a3734abf46ba53e180b3d9f7ccb90e4f7f39f0859f5fed1e0cdc361653

Observation 803bbc2e-3f95-4b7c-b088-455ba9a92ae3 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning,.

Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning A reduction of imitation learning and structured prediction to no-regret online learning,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T14:57:49.542416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:57:49.542416Z digest=sha256:e8e68e459e7d77ddccfd7906c1f8259655f0be91fe45f0489874b418c69ecb67

Pith citing papers

Observation 6b1a3c85-36e6-4578-8995-10c26abf813b · inbound

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens cites this paper.

$N_0$-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-30T12:43:44.697996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:43:44.697996Z digest=sha256:1169fa7fd238536a96d4f22564c2b86d4b4f229925389434031b8a4287c47ad8