Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:01:25.366899Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 22 inbound Pith citation observations for arXiv:2412.13630.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:01:25.366899Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:24:21.470866Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:29:45.616079Z
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 84083f53-db9e-4e1e-87d2-aefc233d3208 · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Such a noisy gradient can easily cause the policy to deviate significantly from the initial weights
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 07fda337-93dc-4ebf-a97f-63d872366498 · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model As the task horizon increases, the agent’s likelihood of discovering sparse rewards through random exploration diminishes
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c23fbfca-699f-4f31-96da-4d0a6c6e8fec · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a6712a2a-97b4-438c-8f69-cd22fad55102 · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Our Cal-QL baseline uses only 25 human demonstrations, ensuring fair comparison with other learning-from-demo baselines that only utilize demonstrations
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fdfa53e9-1a87-4791-a72c-ae7b238f8d5c · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e603f9ed-1b35-4443-b97d-5e6837ac14a3 · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 80f5c9d8-65f6-4c84-8583-5d73d4f75a7f · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 70b70686-dc1a-46bc-8204-659a2df0a51d · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model 24, we experimented with all the aforementioned Q-function architectures in SAC fine-tuning experiments
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4f2054e0-e763-4dd1-a2b3-c72d04eb99c5 · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model These deviations prevent the agent from receiving success signals necessary for guiding learning (see this video for an example)
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation caa91bc3-a834-4fe7-840a-04857f344744 · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model two-layer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 692d356b-0a9c-4913-be21-0068e277ac1f · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model • The PDF of the Gaussian distribution (orange): fGaussian(x) = N (x; µ3, σ2 3)
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f64f459f-cd68-43e2-82cc-c028fe1f6adb · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model • The parameters used in the plot are: w1 = 0.5, w 2 = 0.5, µ 1 = 0.5, µ 2 = 0.5, µ 3 = 3, σ 1 = 1, σ 2 = 1, σ 3 = 1
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cbbb7007-cf81-40fb-923f-353f1240776a · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model However, its online performance is poor, as reported by Ren et al
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 86cde545-1a7c-4cc1-bb26-0c2e06fa85b8 · outbound
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
Reference 2066
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a35557b9-47ce-4cc4-95b6-172f129c4be9 · inbound
SIME: Enhancing Policy Self-Improvement with Modal-level Exploration Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db2435d-d8c6-4259-8446-4e8bf1354ca8 · inbound
Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aceb5ec-e8d9-41e4-afe3-81d0909dd872 · inbound
Touch begins where vision ends: Generalizable policies for contact-rich manipulation Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede22336-157c-43da-ab6b-86a9365cee86 · inbound
Steering Your Diffusion Policy with Latent Space Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bfebf9bf-9bf3-40e5-8603-09c727963f92 · inbound
EXPO: Stable Reinforcement Learning with Expressive Policies Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b73fdf97-579e-41e0-b234-a2c896b84743 · inbound
LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12f079ee-6bef-46af-9d90-a8879cac7c7f · inbound
Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 830346d9-43bd-4595-9c0a-22f5a4151a38 · inbound
From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddd1c633-f0b9-4142-87b2-06ffe2c4c6ea · inbound
ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 68f58fec-9dc3-4c2f-a0c0-32a3bc0c5986 · inbound
Fisher Decorator: Refining Flow Policy via a Local Transport Map Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0d782675-0f10-46d3-97e8-d1ebc68790f4 · inbound
When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 361f3720-953c-4b56-907f-ef6db5555a0c · inbound
When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ef78a9ee-5492-475d-b983-a820646fe8ec · inbound
DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 187
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2f923d3a-0bf9-44d4-a8a3-6771421ce4cc · inbound
Closed-Loop Neural Activation Control in Vision-Language-Action Models Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e8c970f3-c9b0-4aa7-9a49-d86c4558e8ee · inbound
Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 27786b88-2395-4034-a4a5-3b442a5e6a2a · inbound
Flow-based Policy Adaptation without Policy Updates Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c2759d66-e6e2-41af-8b2a-4f227c3a8fd0 · inbound
ReCoVLA: VLM-Guided Reward Compilation for Failure Recovery in Vision-Language-Action Policies Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 76b6f7e1-9290-471f-bdb8-8ec433e0de47 · inbound
MODIP: Efficient Model-Based Optimization for Diffusion Policies Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b86085e0-62ed-4183-b92e-aab854cfc8c4 · inbound
HiL-ResRL: A Model-Agnostic Finetuning Adapter via Human-in-the-loop Residual Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc1f47b7-549c-45f0-8083-09dd79928b23 · inbound
Learning Process Rewards via Success Visitation Matching for Efficient RL Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 40375bd5-b6b6-4e05-b9bb-b5d83c2a812b · inbound
Adapting Generalist Robot Policies with Semantic Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dd34a979-2de1-4501-af9e-250173936c66 · inbound
VINE: Taming Generative Control Policies for Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.