Pith. sign in

Paper Citation Record · LEDGER

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning

As of 5 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2605.05172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05172 v3

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact14
  • verified fuzzy28
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5602730c-eea9-4bb8-8ab6-0d835fcdfff0 · outbound

This paper cites An optimistic perspective on offline reinforcement learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning An optimistic perspective on offline reinforcement learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.550048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:c9767888d60c946d301dc258bc8da90db96c06c52b13efa287fdd5acec55b15a

Observation ec675fdf-0f16-4ab3-9ae9-fe4362c8e060 · outbound

This paper cites Efficient online reinforcement learning with offline data.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Efficient online reinforcement learning with offline data

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.542569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:26b89b2925f84125797df45028ef73e96ca6cba6fc5a9daa62435e489d19ad83

Observation 40fb11ec-a401-4d49-9371-4915c8b53d7d · outbound

This paper cites In: 2025 IEEE International Conference on Robotics and Automation (ICRA), pp.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning In: 2025 IEEE International Conference on Robotics and Automation (ICRA), pp

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:25:06.449172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:aca6aa054f5ff824356a549309df825dc8f894d1453c3e39e397bde55928d58d

Observation afcec783-e7ff-49f7-b16c-ba5c6cc6b472 · outbound

This paper cites Learning under misspecified objective spaces.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Learning under misspecified objective spaces

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.578393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:5515d81c6d32732a4132d625ad5c3cf398bbcef130c47032ce919467a8cfb054

Observation cee03278-b4c2-46c0-bdda-6fe75859e134 · outbound

This paper cites Berkeley UR5 demonstration dataset.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Berkeley UR5 demonstration dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.553573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:0458d4428336e36ea31c9c7f5016c92bf9c6785085aa759c222e66f1494dd03f

Observation fbda2b36-d222-40b5-8850-acf6c272cfd7 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Diffusion policy: Visuomotor policy learning via action diffusion

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.582066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:64ccbec79031c35783dd711969f92945af8cd49c71c05ae74d6491c2d46ecfd9

Observation 090b4944-1677-4cef-9417-cdb2076198fa · outbound

This paper cites Accelerating residual reinforcement learning with uncertainty estimation.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Accelerating residual reinforcement learning with uncertainty estimation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:25:07.016936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:9963bf047a5e102af354962ddcb7c5ba918ea0959d53337943f30d6bf5298711

Observation d828322f-2d29-4f44-a7f2-160889b7446b · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:25:07.008761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:2b44eaddab37e8799566ad63b4a08a08d46489f86b8e6a46c456cad78a57c632

Observation b301ab83-4438-49ac-8af7-1f2ce61b8075 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:25:07.011373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:c8305932fa90b8631a8c86fcc4f1ac6227babee83ed218f35f516b82a3f484ce

Observation 04ade8ca-0c7f-45bc-b8bf-78c093134bd8 · outbound

This paper cites Iq-learn: Inverse soft-q learning for imitation.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Iq-learn: Inverse soft-q learning for imitation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.583858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:6b9d0eea3ba617e82ec8ad51b5e506e2303ced723dba0f9c6ea609ad21ce39ee

Observation 61193c24-ff50-4af3-9647-2e0712bb3ca6 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.589328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:9feb3125d604d730234e6bbdeb81f53997f3c9966d2c35d16281c5b307b8f337

Observation ef12c4ca-ea25-453b-b74f-dac739faf1bc · outbound

This paper cites Imitation bootstrapped reinforcement learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Imitation bootstrapped reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.576621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:76848b9b2bf1fa3ef7f3b7dbc19dc876c556fc1f4281e25a8e9287f5219a4e97

Observation dc07ca71-9f12-4533-ad96-e30468b741cf · outbound

This paper cites Streaming flow policy: Simplifying diffusion/flow- matching policies by treating action trajectories as flow trajectories.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Streaming flow policy: Simplifying diffusion/flow- matching policies by treating action trajectories as flow trajectories

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:25:07.017300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:e31631d15bc5ce5ce6ab16b384435cd0761745fdd26f356e43e0811f929ebc17

Observation 5e3db33b-65f0-477a-8864-9fe2681b38eb · outbound

This paper cites Residual reinforcement learning for robot control.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Residual reinforcement learning for robot control

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.569539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:618c48dc8d7e8aaac8f559d182f21eaa930dbbc730476d3020fdd146ee9f937d

Observation 3479c9a6-c299-4550-8fdd-defdc43c8018 · outbound

This paper cites Kochenderfer.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Kochenderfer

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.571173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:7f59e8d213ba3a5a6ffa4833a63d8ba6add242ba72d885e4319f1d5a3e3bad46

Observation 266fc8f0-5368-481c-8e22-1dc227da4630 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Conservative q-learning for offline reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.574748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:0c7ba60e3d3b7434fdb002ee006e5ff21b9aee30768b0c774323e164edf4ff15

Observation b38b8777-5ff5-4285-9591-7cbe021e792d · outbound

This paper cites The boltzmann policy distribution: Accounting for systematic suboptimality in human models.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning The boltzmann policy distribution: Accounting for systematic suboptimality in human models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.572938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:bf469355983ee004f018525b2ce6bfade42f3bc05b69a8203683395d1ea1f636

Observation e7b46636-abdd-4eff-9a6b-205ad4755113 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:25:07.019630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:755926f1ebda1dfb3fba9212db2f29f3f76fb44dbad9e62d13b53aca891851ee

Observation 1ee4fa5f-567e-4480-84f9-50f2fd84a1bc · outbound

This paper cites Flow Matching for Generative Modeling.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Flow Matching for Generative Modeling

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:25:07.009290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:e16dd5ad80ca405d48c4860f291cf25c0c5d15f213262a3ad512b783fd431924

Observation f085d9e5-5610-4152-938a-fb62dacf2d2d · outbound

This paper cites Individual choice behavior, volume 4.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Individual choice behavior, volume 4

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.557216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:0f3a388724e32feaa9ba48bc1209554f078973abb3124cdededbdeec2ad6c45d

Observation 8d774425-0168-4a58-9545-731ade182813 · outbound

This paper cites Serl: A software suite for sample-efficient robotic reinforcement learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Serl: A software suite for sample-efficient robotic reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.560789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:e7a1e473cfbc0f1d1b4108f00d457f6a6fa0b6227b44853e268797f8f0ee9bb7

Observation c8491618-9b9e-4e45-9760-05d901a53835 · outbound

This paper cites Fmb: a functional manipulation benchmark for generalizable robotic learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Fmb: a functional manipulation benchmark for generalizable robotic learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.559032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:b365f633208b652e414fb106784dbcedfa1cccb07be37f6f64d4b742d16a73ff

Observation ff3fb84b-bc06-4e46-b90a-2969a314c33e · outbound

This paper cites Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.564289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:5a8a465e8c40878262c6b0670e1a43291fc62086a6bce76d1f63f846491ee5ec

Observation 1f0865d0-68a8-4228-885e-25fde3669b47 · outbound

This paper cites Predicting human reaching motion in collaborative tasks using inverse optimal control and iterative re-planning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Predicting human reaching motion in collaborative tasks using inverse optimal control and iterative re-planning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:25:06.444836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:56d726a5a382b85c56eec4064489fabaee735a78a62e8c48a07e4dd214f993c7

Observation d185f49e-a74c-4841-818b-4e1d1deb8645 · outbound

This paper cites Learning to Generalize Across Long-Horizon Tasks from Human Demonstrations.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Learning to Generalize Across Long-Horizon Tasks from Human Demonstrations

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:25:07.013978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:eab6faf4d1b73995b5c9407d04874c89cbd62fb631439f7e8637e7718ec3a59d

Observation 5539257a-b5de-4170-bcb4-15e5ffda870c · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:25:07.019368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:646564ae9d2307d056c01f00ef8b9bd49ddbca6ba479715ac385f7fb2af71c6d

Observation ce945726-8c85-482b-b5c2-109058879548 · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.567832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:8ee0cd504ee3a63d505e62ba04b445d1f441149e2529cdf91a29a71624b4a81f

Observation 4eaf58da-2f1a-4a59-8d03-2afb824aa92b · outbound

This paper cites Learning and Retrieval from Prior Data for Skill-based Imitation Learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Learning and Retrieval from Prior Data for Skill-based Imitation Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:25:07.024819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:7aff714b7f4847cb6df884c938fe96902ccb3339061333f797bc0fa12d4ed8b2

Observation 8bc422b2-a28b-42b8-ba0d-84a302031e91 · outbound

This paper cites Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.587523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:725f91913b6dd37d212b0f330d5f0a258196e9fdbb04c8970e2d9d9b1b3c7a66

Observation 18fb4877-2f9c-401e-a47a-1b5f33f24ea7 · outbound

This paper cites Efficient reductions for imitation learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Efficient reductions for imitation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.551896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:6f7a043bafc14a152f384d5c1ab1a5da562c3aa8b0448747eb821cf81729d006

Observation edba41ed-d8bd-404c-ab17-5a5529852dba · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning A reduction of imitation learning and structured prediction to no-regret online learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.546104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:d765be6e58eaaa9ebe503a69dd9badc329c321dc61a54a0db2d7ab710eadac11

Observation 87a40884-94be-492d-b050-86649de82208 · outbound

This paper cites Quantization-free autoregressive action transformer.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Quantization-free autoregressive action transformer

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:25:07.014602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:8fc34390b2efc23194310f60da32c2b8de7180e762cb11d3f1c5337d73dd32e0

Observation db5fba52-cb65-4be0-aec6-226e5f585cd7 · outbound

This paper cites Residual Policy Learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Residual Policy Learning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:25:07.021746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:627559e9e6f56f7592a2f0961becfbd146b44930feaf7aa81cc9ac884cfdf0ad

Observation fad9eb19-1d8d-4ff4-9d9e-0ec0316b1377 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Octo: An Open-Source Generalist Robot Policy

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:25:06.995578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:cadf620eefe46c7a0e29a9f65172732a77f543a61ce753cf60085868b1ea7fb5

Observation fad6c837-a116-4a1e-9048-48c49ec8ba58 · outbound

This paper cites Image augmentation is all you need: Regularizing deep reinforcement learning from pixels.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Image augmentation is all you need: Regularizing deep reinforcement learning from pixels

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.580296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:f5a62f3316e44101a4b8909927837c3e96f5dbc28ec08e7b5d2158cb73498c69

Observation 361f3720-953c-4b56-907f-ef6db5555a0c · outbound

This paper cites Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:25:07.000747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:bec672a1b64e445051ebb8090b8d47457a1a89c75d619770d195ee354c1afc1a

Observation ffb24849-a9e4-4a2c-888b-6c08219be25b · outbound

This paper cites Deep imitation learning for complex manipulation tasks from virtual reality teleoperation.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Deep imitation learning for complex manipulation tasks from virtual reality teleoperation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.547957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:f5cd5aa6edbb6c04785ad48ab60d13aca8fb73b74317958d1b1d30ccf43548ea

Observation 14de2f99-8035-4933-a6fe-fe78f40067df · outbound

This paper cites Efficient online reinforcement learning fine-tuning need not retain offline data.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Efficient online reinforcement learning fine-tuning need not retain offline data

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.585649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:e450fdd0e0f05bd253801209e313bcee29691f0d0935586348901a52ebe1ed39

Observation 274594d7-c1da-4d05-8b44-8b9a730cc74c · outbound

This paper cites Deep residual learning for image recognition.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Deep residual learning for image recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.591028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:9224ce134d16a90359856976972b7ec3166141decc65a95204a401f77a29a6eb

Observation a3ead827-a68e-4013-91a0-c583278eb055 · outbound

This paper cites Estimating mixture entropy with pairwise distances.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Estimating mixture entropy with pairwise distances

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.555446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:25d39bef10e19f38ba1899ee50ce1eb270f8dac8c0c258049c79fb9080a5d03b

Observation 12b3a106-9801-4736-8660-4e89ea1cb8ba · outbound

This paper cites Fmb: a functional manipulation benchmark for generalizable robotic learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Fmb: a functional manipulation benchmark for generalizable robotic learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.544286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:b08927fe42ee9ce67e06b918063facc30d6407533cb53ad69222f8ff5576b3d1

Observation 44fe016b-b099-4bbc-8b46-a938939b96a6 · outbound

This paper cites Agentlace, framework for distributed agent policy, May 2024.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Agentlace, framework for distributed agent policy, May 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.562573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:f9d129ed549b5fcdd372a92c1a8a892a2f481781bd9798a8fa0d778d104989ac

Observation 4366d565-8ed3-4f19-b8de-df3e23f2dbb7 · outbound

This paper cites Daydreamer: World models for physical robot learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Daydreamer: World models for physical robot learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T11:13:41.566097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:4987ea75e4a08a8a0488ef08001dbbb55896ba8705a0eadb8bd71691de613314

Observation 366ee27e-4144-432c-9abc-a9622a296745 · outbound

This paper cites Residual Policy Learning.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Residual Policy Learning

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T23:25:06.452640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-11T11:50:26.030339Z digest=sha256:e2694c527b2de161052b3143050a1b1686ce32eab4941d14d4a18498c320b632

Pith citing papers

No inbound Pith citation observations are available.