Pith. sign in

Paper Citation Record · LEDGER

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

As of 23 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 22 inbound Pith citation observations for arXiv:2412.13630.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13630 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:01:25.366899Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:24:21.470866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:45.616079Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved5
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84083f53-db9e-4e1e-87d2-aefc233d3208 · outbound

This paper cites Such a noisy gradient can easily cause the policy to deviate significantly from the initial weights.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Such a noisy gradient can easily cause the policy to deviate significantly from the initial weights

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:01:25.556226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.316791Z digest=sha256:b866beda950d0f076dffac7c1620b7ebab90e26409dec1ac958960a1ebc5618a

Observation 07fda337-93dc-4ebf-a97f-63d872366498 · outbound

This paper cites As the task horizon increases, the agent’s likelihood of discovering sparse rewards through random exploration diminishes.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model As the task horizon increases, the agent’s likelihood of discovering sparse rewards through random exploration diminishes

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:01:25.544744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.321009Z digest=sha256:dd71a23d013165b1d4a4c4cb52a1bb063e2fdeef4c07853658107fe222afccca

Observation c23fbfca-699f-4f31-96da-4d0a6c6e8fec · outbound

This paper cites an unresolved cited work.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:01:25.509718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.334021Z digest=sha256:a48ab5ea3672d9044165c3112668d5d75d8c896eeb1e3f32fb6490f78c080e74

Observation a6712a2a-97b4-438c-8f69-cd22fad55102 · outbound

This paper cites Our Cal-QL baseline uses only 25 human demonstrations, ensuring fair comparison with other learning-from-demo baselines that only utilize demonstrations.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Our Cal-QL baseline uses only 25 human demonstrations, ensuring fair comparison with other learning-from-demo baselines that only utilize demonstrations

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:01:25.533220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.325417Z digest=sha256:19732af2e88d6c3429f9389eb0f33645b9da4b7b56736ca14bfe51fc6405404f

Observation fdfa53e9-1a87-4791-a72c-ae7b238f8d5c · outbound

This paper cites an unresolved cited work.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:01:25.520845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.329864Z digest=sha256:51cee7a70bd2fb15bf58991fac1ced227c08f62231e488a2cfc0b77d44e83926

Observation e603f9ed-1b35-4443-b97d-5e6837ac14a3 · outbound

This paper cites an unresolved cited work.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:01:25.498335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.338149Z digest=sha256:ebc9d3defab22e78c62e8f6fd953a62aea3885c7bc006d230e63761c19308837

Observation 80f5c9d8-65f6-4c84-8583-5d73d4f75a7f · outbound

This paper cites an unresolved cited work.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:01:25.486593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.342115Z digest=sha256:46a30e77f2800f163a204cd5a915728ca3059a82de49950fb1bf79b561c4c961

Observation 70b70686-dc1a-46bc-8204-659a2df0a51d · outbound

This paper cites 24, we experimented with all the aforementioned Q-function architectures in SAC fine-tuning experiments.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model 24, we experimented with all the aforementioned Q-function architectures in SAC fine-tuning experiments

Reference 9

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T13:01:25.475035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.346157Z digest=sha256:c483e82585c73e960d16afa1c8b58244842518611b4c941448987cb91c390200

Observation 4f2054e0-e763-4dd1-a2b3-c72d04eb99c5 · outbound

This paper cites These deviations prevent the agent from receiving success signals necessary for guiding learning (see this video for an example).

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model These deviations prevent the agent from receiving success signals necessary for guiding learning (see this video for an example)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:01:25.462819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.350840Z digest=sha256:f1a445fac6327cc4f70a8271492d0e05c999be203fd7cf8b47a0f8c90b66d1f9

Observation caa91bc3-a834-4fe7-840a-04857f344744 · outbound

This paper cites two-layer.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model two-layer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:01:25.450517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.354703Z digest=sha256:5093206f06330a861a39568db202512e2a0210701f04d9d5f1d56fe02c2f5d9e

Observation 692d356b-0a9c-4913-be21-0068e277ac1f · outbound

This paper cites • The PDF of the Gaussian distribution (orange): fGaussian(x) = N (x; µ3, σ2 3).

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model • The PDF of the Gaussian distribution (orange): fGaussian(x) = N (x; µ3, σ2 3)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:01:25.427222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.362974Z digest=sha256:12a8b85f5bf6ff20e64a4a91ab1f9a95af31616e66cfff82f7ee57422eaf9694

Observation f64f459f-cd68-43e2-82cc-c028fe1f6adb · outbound

This paper cites • The parameters used in the plot are: w1 = 0.5, w 2 = 0.5, µ 1 = 0.5, µ 2 = 0.5, µ 3 = 3, σ 1 = 1, σ 2 = 1, σ 3 = 1.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model • The parameters used in the plot are: w1 = 0.5, w 2 = 0.5, µ 1 = 0.5, µ 2 = 0.5, µ 3 = 3, σ 1 = 1, σ 2 = 1, σ 3 = 1

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:01:25.414446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.366899Z digest=sha256:3c7c455feb80bad0d5fdffb1fb81e56dc06cc6bdbec76773726feb322d6dc0ac

Observation cbbb7007-cf81-40fb-923f-353f1240776a · outbound

This paper cites However, its online performance is poor, as reported by Ren et al.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model However, its online performance is poor, as reported by Ren et al

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:01:25.439210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T13:01:25.358805Z digest=sha256:d31b4afb175674846a0783cb93c739f3e07bc92a6cfa7bbdb474a92781ebe517

Observation 86cde545-1a7c-4cc1-bb26-0c2e06fa85b8 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 2066

Resolution
unresolved
no resolver link, observed 2026-08-11T13:01:25.310151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:01:25.310151Z digest=sha256:29140f736e958e267ae908fd722232a05a13b76219d0da6d7a5e51238453e1e8

Pith citing papers

Observation a35557b9-47ce-4cc4-95b6-172f129c4be9 · inbound

SIME: Enhancing Policy Self-Improvement with Modal-level Exploration cites this paper.

SIME: Enhancing Policy Self-Improvement with Modal-level Exploration Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:24:21.470866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:24:21.470866Z digest=sha256:b44b4e197b593e9e4b084a66b84fb52ad304a7d291766655c91a820b2da3ae8e

Observation 1db2435d-d8c6-4259-8446-4e8bf1354ca8 · inbound

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only cites this paper.

Efficient Online RL Fine Tuning with Offline Pre-trained Policy Only Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:46.019748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:46.019748Z digest=sha256:32c63b760ca89df2555e3393d6d6612b8f35584856fc79e19b6acec79b3802c3

Observation 4aceb5ec-e8d9-41e4-afe3-81d0909dd872 · inbound

Touch begins where vision ends: Generalizable policies for contact-rich manipulation cites this paper.

Touch begins where vision ends: Generalizable policies for contact-rich manipulation Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:46.173598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:46.173598Z digest=sha256:eed04a517ae1943a122e0b7013d92e1b3c69968dcf8ac465df39e0ff27751f1c

Observation ede22336-157c-43da-ab6b-86a9365cee86 · inbound

Steering Your Diffusion Policy with Latent Space Reinforcement Learning cites this paper.

Steering Your Diffusion Policy with Latent Space Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:55:46.247634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T21:55:46.183007Z digest=sha256:242ae9d8be47c4bb028d6d8fa0b4e1492edc7f9af69b96ba437259e1f43606d7

Observation bfebf9bf-9bf3-40e5-8603-09c727963f92 · inbound

EXPO: Stable Reinforcement Learning with Expressive Policies cites this paper.

EXPO: Stable Reinforcement Learning with Expressive Policies Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:12:05.263582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T05:09:02.111308Z digest=sha256:48056e14464617a56f86022d04bdb4d93fb1c6c6a6f131c52c2419324ff991e8

Observation b73fdf97-579e-41e0-b234-a2c896b84743 · inbound

LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations cites this paper.

LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-15T17:08:57.846352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:08:57.846352Z digest=sha256:1b7ac3538e8883fd42007185435d0cc1c6cffde4b414b4b79d112534a09ab82b

Observation 12f079ee-6bef-46af-9d90-a8879cac7c7f · inbound

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning cites this paper.

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:10:13.387889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T14:07:10.387869Z digest=sha256:7a8d7d57b118e047545aafa14d09865d5ef008caa305ceeda3c3084755999a29

Observation 830346d9-43bd-4595-9c0a-22f5a4151a38 · inbound

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning cites this paper.

From Prior to Pro: Efficient Skill Mastery via Distribution Contractive RL Finetuning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T23:46:32.301737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:46:32.301737Z digest=sha256:1af6c313421f5c25b80446ba7a9ed3514f2b7c52f48889dced1c2f91161525fa

Observation ddd1c633-f0b9-4142-87b2-06ffe2c4c6ea · inbound

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors cites this paper.

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:39:55.156113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T09:37:02.168898Z digest=sha256:84dcca502b41a367a0a624cbce7fd15019de4f0df5bc0ff276c1d7a8717df4aa

Observation 68f58fec-9dc3-4c2f-a0c0-32a3bc0c5986 · inbound

Fisher Decorator: Refining Flow Policy via a Local Transport Map cites this paper.

Fisher Decorator: Refining Flow Policy via a Local Transport Map Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:51:46.620645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T05:28:12.298066Z digest=sha256:7b5efec6d9909907f78399d5117c03523bf103f94428dcad693d6f15f487a351

Observation 0d782675-0f10-46d3-97e8-d1ebc68790f4 · inbound

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning cites this paper.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:11:05.045910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T16:34:18.603134Z digest=sha256:97483db4d77b9d77f57d34944f60c5ced80a023f4a05928c3737d6ef9fe0c435

Observation 361f3720-953c-4b56-907f-ef6db5555a0c · inbound

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning cites this paper.

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:25:07.000747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T23:24:57.811919Z digest=sha256:2865ace52cf4b5c02563c0d2a50aabe65164cb1a01b62afd1db799c4fb8db29c

Observation ef78a9ee-5492-475d-b983-a820646fe8ec · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 187

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:17.324262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:6ad69085af160fa5dd1fed8f728ce1a92ee11c228364d4b8079e813ac00966dc

Observation 2f923d3a-0bf9-44d4-a8a3-6771421ce4cc · inbound

Closed-Loop Neural Activation Control in Vision-Language-Action Models cites this paper.

Closed-Loop Neural Activation Control in Vision-Language-Action Models Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:36:08.209066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T22:23:32.953232Z digest=sha256:d73c4886dd15d33fca7451f1984e3161782c569cae213db198a90d08bc43d74e

Observation e8c970f3-c9b0-4aa7-9a49-d86c4558e8ee · inbound

Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation cites this paper.

Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:56:34.486703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T09:37:58.434897Z digest=sha256:da9672d9c961f37dcd3d5b741840cef658dbea32fe13127b8927681cc88d3e87

Observation 27786b88-2395-4034-a4a5-3b442a5e6a2a · inbound

Flow-based Policy Adaptation without Policy Updates cites this paper.

Flow-based Policy Adaptation without Policy Updates Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:26:59.178158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T01:18:08.001674Z digest=sha256:ef4f8dca3ac0c98f0c6eb0e6d3e528a812937426f680e9945e585ff0da8b8f99

Observation c2759d66-e6e2-41af-8b2a-4f227c3a8fd0 · inbound

ReCoVLA: VLM-Guided Reward Compilation for Failure Recovery in Vision-Language-Action Policies cites this paper.

ReCoVLA: VLM-Guided Reward Compilation for Failure Recovery in Vision-Language-Action Policies Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.813776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T16:42:30.970472Z digest=sha256:2e7db90a2be19b7439a05a70358ad3e4eaec4b46739ca2eea61111bb32cf3238

Observation 76b6f7e1-9290-471f-bdb8-8ec433e0de47 · inbound

MODIP: Efficient Model-Based Optimization for Diffusion Policies cites this paper.

MODIP: Efficient Model-Based Optimization for Diffusion Policies Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:37:37.581918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T13:44:19.708550Z digest=sha256:5367539448bedfb1da426c2f47dceedb92af492f108217e475f35181622ca30b

Observation b86085e0-62ed-4183-b92e-aab854cfc8c4 · inbound

HiL-ResRL: A Model-Agnostic Finetuning Adapter via Human-in-the-loop Residual Reinforcement Learning cites this paper.

HiL-ResRL: A Model-Agnostic Finetuning Adapter via Human-in-the-loop Residual Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:29:45.617889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T08:45:04.483217Z digest=sha256:27a417fa41a46ac6cd35890938027423eb15ac02a437d6c4f49039b2bb7ff08d

Observation bc1f47b7-549c-45f0-8083-09dd79928b23 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.575789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:766d1d3e0271fa092cb3841638a38bc53f40da412a4c60981fa2753f4f7873da

Observation 40375bd5-b6b6-4e05-b9bb-b5d83c2a812b · inbound

Adapting Generalist Robot Policies with Semantic Reinforcement Learning cites this paper.

Adapting Generalist Robot Policies with Semantic Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.679466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:09:29.625066Z digest=sha256:c85d4561a89bfbb67d0c780d91b419291bb4964cfa052fe5559a012a5aa07819

Observation dd34a979-2de1-4501-af9e-250173936c66 · inbound

VINE: Taming Generative Control Policies for Reinforcement Learning cites this paper.

VINE: Taming Generative Control Policies for Reinforcement Learning Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T12:17:04.321971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:17:04.321971Z digest=sha256:3ff5ed7b68412704223c81cb959b630277b243360364396ccc86130e8be88af6