Pith. sign in

Paper Citation Record · LEDGER

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies

As of 11 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2605.03065.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.03065 v4

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T00:02:55.449923Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact36
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ac0b1b8-3d1f-4090-841c-1f42b4993154 · outbound

This paper cites Stochastic Interpolants: A Unifying Framework for Flows and Diffusions.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.676445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:b691010cf49057e8c14e6dc87cee4a4d355ac33c4cc2dd7ff36bf68fed877040

Observation e96d7d8c-10ea-4970-aa9e-90730a5b9759 · outbound

This paper cites Learning long-term dependencies with gradient descent is difficult.IEEE Transactions on Neural Networks, 5(2):157–166.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Learning long-term dependencies with gradient descent is difficult.IEEE Transactions on Neural Networks, 5(2):157–166

Reference 2

Resolution
verified exact
doi, observed 2026-07-01T00:05:08.773784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:8997ae632743bb9234f1a39e5ee69e228f8a0549ebf8a48d0d2fb04ea9f48a32

Observation 76562e94-ed6e-4e04-9278-9e2fe6457c2a · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Training Diffusion Models with Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.726625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:063ed176293b1e27194aabbd6f37cbde8417b3ba20723bf50bc7c33f7d262751

Observation 22b44d03-983e-4d85-bc96-bb9f0b07a5bc · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.678913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:6b78d132f900c74ab831cad0a62e91ade236a9c0180edb48ed13dcb5f166298a

Observation 8d89a0db-29e9-447b-acf6-4dc613ddc45c · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.681731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:c5ae2c69f744a06e82392dbdbd06993b4e4f427070dabfc85221c5d0e0570f8e

Observation 2c6efab4-06a8-4f45-a19c-dc765a6f0c1a · outbound

This paper cites Randomized Ensembled Double Q-Learning: Learning Fast Without a Model.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:05:09.696207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:fc344a0d123af2ffe7e235d409cf6e09fd02b86c4a0b3155d890af8beca18643

Observation f24fe5e6-d9d2-48d8-a029-14f0bc103e8c · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.688829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:d8bad6d47662d71b00ad3eeea64e5a1b28d1f2897c689db21453525e5aee4201

Observation ec093a9a-b7b7-49a3-a3ab-a27d39b7e0b2 · outbound

This paper cites EXPO: Stable Reinforcement Learning with Expressive Policies.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies EXPO: Stable Reinforcement Learning with Expressive Policies

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.712302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:c57c4bf0d7fcdcd56745205646f7e3f815825a0e6acc5f03d9e2dfabcb480603

Observation c400599a-47be-4c03-b5fc-8696aad2b151 · outbound

This paper cites IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.721995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:7bc4019ac9ecda7d0f9017f4864052e71aa23611204b6bc2c16837feeb67fe5b

Observation 718720ac-6e24-4d11-9a34-0e915dc477d5 · outbound

This paper cites Implicit reparameterization gradients.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Implicit reparameterization gradients

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.366090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:e3ebf7b3758f7c27c0693c33ae05c4d97bc913305f486f27d23edf1231a0d29a

Observation 1fa0c259-8845-4328-bd98-786bfbffb5aa · outbound

This paper cites One Step Diffusion via Shortcut Models.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies One Step Diffusion via Shortcut Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.693742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:f4a26d2367b78f941722f24b71dab7eb275b687356e3860fd8a2af65693f3b09

Observation 7f2b5ac5-f29a-402d-86da-8695f7888a87 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.698570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:50abc8426c2e64b0b5bd6c410d178ae91f74bd90714ad19b396b9a4d452e8cc2

Observation 826bc7b6-a777-46cd-8bec-b6b2a8ac091f · outbound

This paper cites Addressing function approximation error in actor-critic methods.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Addressing function approximation error in actor-critic methods

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.369650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:431b62e1d5344d164b367800f4b226008f2437be0a634097cab0734820c530ff

Observation 242da48d-3abb-4790-885d-79a84a51b9fd · outbound

This paper cites Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Relay Policy Learning: Solving Long-Horizon Tasks via Imitation and Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.684233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:1c25cbb7d64ee9b64e086be0b895a71d2c5acdb22f44e6d36e044ab7cdc5c84d

Observation 9b1a8f0b-ebef-4659-bae9-c20cdb464f3a · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.363407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:3dabbe3b2dc7e4a61af26fa4695f64e88eba6ead8eafdbe9e115bac1c13f5276

Observation 1758a274-10c2-421a-84d9-a9a2f6b67ff3 · outbound

This paper cites Denoising diffusion probabilistic models.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Denoising diffusion probabilistic models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.367851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:bedd9e2686d3767c8a20f4b9ed4a1021cfe7cd131f3e0ca1c90140d625828698

Observation e3f67899-95f2-467a-8963-69b9da1675cc · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:05:09.669377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:ed6e28c54513fa5906cb41f4e1e2d3a3ad13c6256e117d219dcf2f7f1162d088

Observation 3533448f-c5f0-4446-b417-5e2ad50e496f · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.718397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:92119af72f233b372b3c22cd10b63e87ef104913fd6d08e07ef37779cbd4f58c

Observation f07f14f8-a433-4c7b-85e9-fe3c158a7592 · outbound

This paper cites Auto-Encoding Variational Bayes.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Auto-Encoding Variational Bayes

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.666809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:9b70e82a146e1b05adfb4c035b1c05a88e966050001b34c9a6fc577e2da34ea0

Observation 20ee65ab-e992-42e3-8906-0af0ad9a33db · outbound

This paper cites Rl-100: Performant robotic manipulation with real-world reinforcement learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Rl-100: Performant robotic manipulation with real-world reinforcement learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.661869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:b80551711279f3c9b95adf260a1eb59d01fffe42f621a7a1bc692ce3006c6a49

Observation ac7c232e-910f-48e8-b43d-f7cc059fd905 · outbound

This paper cites Reinforcement Learning with Action Chunking.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Reinforcement Learning with Action Chunking

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.703276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:78582a1f87cdc2dec2b38b8c0f0919fccf71414d47ed9a1f48e7cd010c2708d3

Observation d1670ae1-8dc0-4b81-9e93-0f62e2754523 · outbound

This paper cites Flow Matching for Generative Modeling.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Flow Matching for Generative Modeling

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.710124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:d5cac5607146e085be130c47eb43d28e967649c41d1dc20409e121a949c316d2

Observation 875e7b2c-640a-4a2e-8ceb-1985281c6e68 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Libero: Benchmarking knowledge transfer for lifelong robot learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.360744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:a03431357afcf0c32a7e6d3a6be6a6d8deedd160b26e960d24a77746f324ec45

Observation 4cd68232-5d9f-45f6-bc62-2fee9011ed3b · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Flow-GRPO: Training Flow Matching Models via Online RL

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.659427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:0e032eb56d29940562a461ac20b219173eca396b0e32a79e2daccd8fb41d946b

Observation 25748936-99cd-4c1e-9a20-cd462e37b2ca · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.664320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:787b3d1922f52c5405c6f4d8a0557d8bdc9b75e70562b448aa1dcae0c7fbb09f

Observation 51d81c37-bc1c-4839-b742-29f641082b59 · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.671629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:5e8fe8bb61a91f036f1c49dd1218a314c3685e5ff80d4dc2ad3c867fc29db799

Observation c6b3ca7a-5ffb-44a1-9f66-0bf9626af0c8 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.674048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:2b5fa9d6a318ed6ec0a57daa68c89ed8449b5db98dc89342acb9fe4993936fb5

Observation d1f51504-4b4f-4ca0-9bee-40f1946ef096 · outbound

This paper cites Flow Matching Policy Gradients.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Flow Matching Policy Gradients

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.701041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:a081bbe7d8228e9ba6cba1b670335ba195fd278619eef04772508ffa60fa506d

Observation 61bda8c9-cd75-41ab-89db-a16147a5a301 · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.686495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:c6f6b7d301ef1763a650310a6a40609f6eecb729f982af01e9ba7d6725441963

Observation f28012cd-88ca-40cc-a7a6-200d7b3d585a · outbound

This paper cites Steering your generalists: Improving robotic foundation models via value guidance.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Steering your generalists: Improving robotic foundation models via value guidance

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.362708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:6cb3eeec8330da5055b93c98eaaa0d12c4a4319e614e06a249c402c3aa32b5af

Observation 82bbdec9-a19b-4bc5-ac61-840847a6ec38 · outbound

This paper cites Self-imitation learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Self-imitation learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.366900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:ede088d7875172c8e6818f2c7908851feedd8fbcd8397adde11861c742baecce

Observation 96f5956d-98bf-4472-abc2-40c0ee4b1942 · outbound

This paper cites Training language models to follow instructions with human feedback.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Training language models to follow instructions with human feedback

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.355488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:cce0980e70d69723a0652c53ff86ba0e62552f966218cb1eb8bd2844e6eb651c

Observation 74ca1875-142c-4ee6-9ada-bfd5271e0c0c · outbound

This paper cites Much ado about noising: Dispelling the myths of gener- ative robotic control.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Much ado about noising: Dispelling the myths of gener- ative robotic control

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.654833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:ab75c850ec5951b8a85b6eab77a490bb0f65154caa123d71b0e59f978e33361d

Observation 05fabe9c-601c-4697-9543-92f0a4db82a2 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.657171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:0b16a38d64d76e95872b61c08f4b6e3f1d5129ac61eafcdcb8b05cf2e73dc4c7

Observation cc0b41cb-7127-422d-a17a-d2688821a6a0 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Film: Visual reasoning with a general conditioning layer

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.357630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:65ccb1af0c3cd7e7ecdff3e62275a9d09938f5abc64be9a8488cffd04c79fa45

Observation 0fb987d5-a3af-463b-96b2-97cc4f8c2ef1 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.691460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:39bd443e065ced0ec307fa5c864cf5311a57bd48d201cc6e0bb85c2d71f040b4

Observation f35b40ff-a3e3-4c8a-bb2d-b5bc8413b31c · outbound

This paper cites Information Theory: From Coding to Learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Information Theory: From Coding to Learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.355896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:70a6ab2bf266af8386d46f5c2c7b125f77a8a1436e61d17866df55b7f222ac7f

Observation 2af59ee1-3ad0-4010-935a-3ed5a7f4f754 · outbound

This paper cites Diffusion Policy Policy Optimization.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Diffusion Policy Policy Optimization

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.649912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:68578f7c9df4cf2339e3ab5a0a751b1fd705acbeb76ecf8a315e22592bdfae38

Observation e22bb612-fe09-4054-9c4b-2cc9240603dc · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies High-resolution image synthesis with latent diffusion models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.364361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:c6847fbd4d8f435a92e7fffc7ef7aca0c0737b13d7c6b2c14da28fec9ad47977

Observation 2f2da0d9-5324-4f28-8621-91f535dd0509 · outbound

This paper cites Trust region policy optimization.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Trust region policy optimization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.358942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:106a318d101c0a02f800105b0657a9fef009bea21b6b84dc8e21921e03b4d5c2

Observation 4e406356-796c-47db-a645-4c6fcdf9c760 · outbound

This paper cites Proximal Policy Optimization Algorithms.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Proximal Policy Optimization Algorithms

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.647772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:67bc0876d1237d772c8076551e665af79a04fc06dd2c022916c50e1c13f5d2e1

Observation 9de63b67-d0db-426d-ae15-87987fce5a98 · outbound

This paper cites e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.652333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:df509de13729022900db2ecc35810e22a4e17ba6f4b9fec4f03d5cb296831c7c

Observation 1c1afc29-c1c9-4d50-826b-43b5ac824d9c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.707869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:b57de2363981d2609064583b08e20424a262c0c3b3ed39f8c71c1078eba70f01

Observation 641428c8-6308-451c-952c-4f35cf1e688b · outbound

This paper cites Denoising Diffusion Implicit Models.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Denoising Diffusion Implicit Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.705464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:d6c5d7a1949fb1c9599195e62717894b21cfbd75356d731d05566eaa7652ec84

Observation 4da2841a-48e1-41ad-906f-22b7e271b714 · outbound

This paper cites Consistency models.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Consistency models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.357242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:512746297595aa7e40e0b2aec29668db0fcfee67df8d9a1fa499de677b9b9bc7

Observation 5d78a724-dce9-4cc4-8e7b-f394573bde90 · outbound

This paper cites Do differentiable simulators give better policy gradients? In International Conference on Machine Learning, pages 20668--20696.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Do differentiable simulators give better policy gradients? In International Conference on Machine Learning, pages 20668--20696

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.347971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:81ce468e16ff307d1e329550eec107d47b5fb4181242c79ceeaa8314f13ccbc6

Observation 2b5dd790-ece3-4e16-a3d3-9ba8769b83a4 · outbound

This paper cites Jump-start reinforcement learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Jump-start reinforcement learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.368553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:6ca946d72459901e023315ad437165fbdca632188608645e8a5fec23045917d7

Observation a3361abc-5bee-4405-a918-7a7159f73ee3 · outbound

This paper cites Steering Your Diffusion Policy with Latent Space Reinforcement Learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Steering Your Diffusion Policy with Latent Space Reinforcement Learning

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.640722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:7822fa26869e4f5089f5f2de7362d7003164f2242573fe65faddd1e66c50a230

Observation 9b275baa-22b5-4e01-9924-3de4269dfa18 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Bridgedata v2: A dataset for robot learning at scale

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.351717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:46953f6fa5d1191d7324978cb883835043071e80271bd1987e983f9ddd115b48

Observation b0014d3f-484a-48b6-8972-c028d1cd7f9e · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.370461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:6bfd7dbbcbb89c3d05e86f6d96f70b79eaad1717cc3be052ebef9313d301ae7d

Observation 64634089-8a90-4fd4-9cc1-cbc55417b948 · outbound

This paper cites Diffusion models for robotic manipulation: A survey.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Diffusion models for robotic manipulation: A survey

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T09:33:36.351866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:0e4b0e36d5d92334202fab495d2e9a234e7b928ef62720af63055aad832ce99c

Observation 672855d0-1ae7-4527-81e8-11987f4ee878 · outbound

This paper cites Multilingual Universal Sentence Encoder for Semantic Retrieval.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Multilingual Universal Sentence Encoder for Semantic Retrieval

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.714639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:2d05fd098f0a299aeee376732a38ada4cd115ab16f6923e4cf4d1979a54ebbea

Observation 4124e0b7-f1b1-489f-b42e-b4610a8c1a21 · outbound

This paper cites Flow policy gradients for robot control.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Flow policy gradients for robot control

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.724331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:9668eaa477e0d30ba8792d90c54378a320801ecef6f46908ffba9d3af22b07fd

Observation 975da54b-5f96-4560-b4d1-21dcb172b1fc · outbound

This paper cites Affordance-based robot manipulation with flow matching.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Affordance-based robot manipulation with flow matching

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.638373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:31860620422fe762fe4ddb31a9467f4b2010827a98ef1448498f14a35e2e4fc1

Observation 122ca670-a47d-4927-8753-a54de95dd54c · outbound

This paper cites arXiv preprint arXiv:2507.09061 , year=.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies arXiv preprint arXiv:2507.09061 , year=

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:05:09.643171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:5698371a4d363ba6c6f47cde6c11828e7e8eca6c87269de623eb2c0b31a527d3

Observation 77c6309b-ae97-4580-a1c6-82345a315b12 · outbound

This paper cites arXiv preprint arXiv:2505.22094 , year=.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies arXiv preprint arXiv:2505.22094 , year=

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:05:09.645589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:98dda685442216c8b1958696f3430e518cb5335e826614c56f6cb35fcd2f4dd0

Observation a10e5ef2-f1fb-4953-b599-84cc3723bbc4 · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:05:09.635889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:c33b614729731400fdf77689a8fc9ca65037178ef70c06d8c61d0405d9304a45

Observation 5b4de9fb-0cae-4a8d-a72f-9485efdf5270 · outbound

This paper cites Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.633737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:fea5ebd03f954ba18a2ce7fc4bddb2fe50d67854eb3939878099356949312b4e

Pith citing papers

No inbound Pith citation observations are available.