Pith. sign in

Paper Citation Record · LEDGER

GRAPE: Generalizing Robot Policy via Preference Alignment

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2411.19309.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19309 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:50:57.959258Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cb8b47de-5906-4be4-b685-a637e828abe3 · inbound

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation cites this paper.

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:50:57.959258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:50:57.959258Z digest=sha256:0120abe6b7536eca76af52ef834378377e8c8c601dd04ee592357c76d41b2995

Observation 83a3fcea-f4b4-4d05-886d-741da72770f3 · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:48:48.806390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:3644f98f96a32b2b5b2fdab612ae6c87f776c57dae99b15c62b603da41c9584e

Observation d799988c-429d-4e5f-9812-7004b9454f9e · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:32:22.363054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:cbf398ac5440bf023698853c3c22d6e950a43ae69a527a81cedaf78849cbdb4d

Observation 2fe91573-9106-4706-b7c4-aa0c4edc2f88 · inbound

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems cites this paper.

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:09:24.594651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T15:09:24.367362Z digest=sha256:7aad70a978c8723422b7fb114c6d05ca8f992e86fdf8bc29bbca1f5eb452cc10

Observation a6898c71-0575-4ffd-8d6e-122c1bd2b242 · inbound

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions cites this paper.

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:28:06.959009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T15:28:06.883492Z digest=sha256:e3abcce987ad2b196180fe3bfa824be313b51d90f773f441b59e4032aca6e204

Observation a50457bd-f4cd-4f9a-a775-a7e4dfb9c6c6 · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:55:40.398126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:c5571890c92b5b30d9d02c7ea601a7c4a88d65d2c756813174fc688570ee210a

Observation 4bface85-924a-42ee-a04b-36476e85f515 · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.572064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.572064Z digest=sha256:7f3374a9eb919154ae707437374906c7307ddac101a1918355d7adb90be4b124

Observation 29e67213-daad-427c-b70d-6acbf6f1b574 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 258

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:37.279568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:37.279568Z digest=sha256:26a6fa75107668d02b8bb6a7fbe2501155a17843b29d6f41c83d717c5105b94b

Observation 8878e613-9cc9-4eda-923c-404de6112c0d · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.109526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:cd7abd7719e9b5b2f59b889441a0089512a77265ebf484f79f2b6b231db6314e

Observation 02898c68-a010-4aa3-adf2-0cfe4243bf80 · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.688049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.688049Z digest=sha256:32b2cc313a34f431850007c278e47d16b52bcb97f1cbd223ee19eaaa0e42f52c

Observation fd5ff45e-a0ba-4ba4-952a-a685a8f4638b · inbound

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning cites this paper.

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:02:11.464210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T08:02:11.189795Z digest=sha256:0c96fdaf9a6e78ce4692f5f5232eb2eb839e70a6c7fb80fa923f9984cd9f51d6

Observation 54f664df-f197-40ff-89df-f5497b457f57 · inbound

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models cites this paper.

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:24:57.958256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:24:57.958256Z digest=sha256:bfaf7205d72472f80a85381efb3765945a514fcaceb2ff6b78acca6a70e1ecdf

Observation 40ef77c0-734c-44db-9fd0-8c8f4fb483ea · inbound

Reflection-Based Task Adaptation for Self-Improving VLA cites this paper.

Reflection-Based Task Adaptation for Self-Improving VLA GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:31:02.933528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T07:28:11.187479Z digest=sha256:b72387434eed840dac4a15096ead002a229456bb9fc323e67ba9356ac5dae3d5

Observation a64082d7-3ec8-4d36-bd0f-2c30aa744d6b · inbound

RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation cites this paper.

RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:10:57.868145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T06:10:47.309028Z digest=sha256:aa8a75b863deae30cfe5433b81dec744068bf9dbe1204bdb0a5181142c2701af

Observation 1dd654af-0a47-4061-a394-7a1a3f3bf613 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.387695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:c7f9f702650ff5de3ab87acde89927275a6bc1fb1cbf91bfc702857b065a086c

Observation eb701436-3068-4bac-a2b7-ca52cf2f6edc · inbound

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models cites this paper.

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:10:48.869856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T03:09:09.713822Z digest=sha256:01e34c8da523666432324ef8358402694a80af6b84239380225c6930af1fcb3e

Observation fd542b69-2fa3-45a5-a2df-22e01d8028ab · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:27.515205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:27.515205Z digest=sha256:8a4201a0e29d326b5949af55363d5db7d20de0d60c1baec77051a83d8a8fc381

Observation 62cb33a6-30c1-48ce-b3d1-b892d91f022f · inbound

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation cites this paper.

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.900849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:13:53.818915Z digest=sha256:e48813ae85caf470e45984fe095468a8499c73a6689f5ce10e1eee141ee38a97

Observation 904c8a16-d0a8-4b19-8bae-3e0284a62a0e · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:c9835f8c812756717d8b25dfacca2b2af54705e5f005522587e04f88403cc776

Observation 42866686-c264-452a-a320-63705cc5313b · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:30.009180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T05:47:17.494531Z digest=sha256:d22c056390caeeb7653c41694a58ca471b1c23311b80e6b26b2e29c94e242c92

Observation e7e21ad5-2002-4d86-9cfb-3157990b3a32 · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:17:03.135989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:00:26.352130Z digest=sha256:070d19345d1653127fa5ff18bc0f4d024e17ec1571c876b8d34683717b11fbe5

Observation 5605a0c1-a662-42a1-acba-56c0b3227373 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:10.861354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T19:31:38.069592Z digest=sha256:20a15b7a6cc7836ed0450ec3f8f5b9f1556d2d0e1534a5db2a16894297d021eb

Observation 8446b9d6-3128-4dda-b251-182de97f9328 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.297095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:bdd6b072aa3afba6d0d791ee49908eac13b3d5938267756a822952b1fe1ba477

Observation 14c156f9-c290-4a3b-b2ff-a42e8914e864 · inbound

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation cites this paper.

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:41:19.940732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:40:22.192349Z digest=sha256:b181e5637f4787d67be0792770af9200633bddb73f8132d12be16e8e3da2dcef

Observation 406737e0-8593-4ee5-98f8-7b550cf92ba6 · inbound

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models cites this paper.

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:25.984235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:39:32.715650Z digest=sha256:01b1e3d9927f26ee458d894116c0ff377ffa4ae63f1c74d89ee0c3eb37bd0303

Observation 5e16325d-038d-4ed9-998a-680fd90a8f0c · inbound

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation cites this paper.

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:17.884932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T04:44:03.661688Z digest=sha256:e797203a0839378b4c949c5ae48e866fe2092e4632d6cf8a71078df70798fee9

Observation c38a1916-b160-4388-893a-8d8bf55445c7 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 191

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:17.293277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:06e55242e2ff27479a64a36f404bd77f4927419a0c90ac6a5a0d782d46892603

Observation 5114bd6a-3171-4846-b807-df20d8a2dfa4 · inbound

PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models cites this paper.

PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:03.383440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:12:26.907425Z digest=sha256:51842c30c81dacc84e6d75fcb70fb2e8b0e976e6f625284f87701919d4b6a606

Observation 9a8982b7-5c89-4284-bcae-25bde5ea9884 · inbound

Position: Good Embodied Reward Models Need Bad Behavior Data cites this paper.

Position: Good Embodied Reward Models Need Bad Behavior Data GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.973243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T17:18:17.337336Z digest=sha256:9fb57690ee7727899b93c88c84ecadef6b413b0945f92c5f48ec81b5cb3b8b34

Observation 02a91f91-65be-4ab1-a6f7-ca0aaa2d4ece · inbound

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization cites this paper.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:06:49.389175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:bd1c3d5360ba4137efa2075fe14ca1c8c855f4ff8ff02d1e6c912e851052efa0

Observation 7f8b034d-70c4-4bbb-b528-7fe26894c452 · inbound

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning cites this paper.

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.728427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:46:59.746745Z digest=sha256:cd61cba04532e551de35daef13f3c27d616ace7254a07b0682fe4ba873f41e03

Observation 627851db-f62c-4f08-9b3e-9bb97f8af871 · inbound

Foresight: Iterative Reasoning About Clues that Matter for Navigation cites this paper.

Foresight: Iterative Reasoning About Clues that Matter for Navigation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.175007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:35:05.441401Z digest=sha256:5308466f5281aea40872cd088dfda84d1a0b9fd98bec87bd8cd009715cdfcf12

Observation f7277805-4317-4e28-9e61-fc005ca76c40 · inbound

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack cites this paper.

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T11:31:37.690666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:31:37.690666Z digest=sha256:6f33c880d47d4458e5ddd060369e05f826175dbdd93b742b6dd28aad107cf5c1

Observation 62cfe2bf-a65b-4d10-9fe0-c2e626669ab7 · inbound

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model cites this paper.

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:18:44.353030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T04:13:22.598591Z digest=sha256:51589a1e0c4d92220571ea486e1fcd037cde3c39f792f2fb1ba2f8ded7c2971d

Observation dd067f52-1652-4806-8e97-4d554a3ac16d · inbound

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model cites this paper.

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T04:20:31.958581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T04:13:22.598591Z digest=sha256:d1180bf6602cf7a9dad78ce9d74ee0565b0881478adea63e85b66a63f499dd7b

Observation 3e82a913-f3bb-4323-ad4c-b10b8512e9dc · inbound

HiL-ResRL: A Model-Agnostic Finetuning Adapter via Human-in-the-loop Residual Reinforcement Learning cites this paper.

HiL-ResRL: A Model-Agnostic Finetuning Adapter via Human-in-the-loop Residual Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.655525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:45:04.483217Z digest=sha256:de873881012785a48deaddc1457f68c6df741823225d46dee55526774f4469c3

Observation 37132ba4-8cb9-454d-87ad-233ba7680559 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.474130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:fcd5abe283168b504390d5f30c35c6860019b806ad63f12f79a1091ab8ac31ec

Observation 276ca8c9-6be5-404b-8e94-6c71abf6c499 · inbound

ROAD-VLA: Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models cites this paper.

ROAD-VLA: Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.062167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T20:22:16.508280Z digest=sha256:134470864f1c999a8b6882d4f1681edc247284c5818f103b513c516f3bd9ee21

Observation f7e59c4a-1b10-4957-ab3e-fd326e41c4b5 · inbound

Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models cites this paper.

Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:04:21.163124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:02:15.781538Z digest=sha256:01adfbf8a50c53ccffe25be11372ba6c9b71afa1b278b41030a94891de546df6

Observation 5c19d9a5-74d1-4f82-8f75-e53ecebe2fc2 · inbound

Rethinking Foundation Model Collaboration: Enhancing Specialized Models through Proxy Task Reasoning cites this paper.

Rethinking Foundation Model Collaboration: Enhancing Specialized Models through Proxy Task Reasoning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:45:29.789741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:38:44.170473Z digest=sha256:e47b8e35f9587e774b6775ecbee08ae6fcaf33ed64253225ac6a95ec21298a62

Observation 534908e9-3bd8-4b9e-a92e-584e4d8aa8a2 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:55:41.966122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T04:58:28.971536Z digest=sha256:33431c95c7a366ceae9a7275c7dc13a3f8005bd3f4c6f4f09105697fcc2ff0de

Observation 2d48f57a-1bc6-4b95-bcde-dd18adfd100a · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T16:55:18.028851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:55:18.028851Z digest=sha256:55b9bbf87e73a24171d219ec036c24b59d2fb129422080b6decef1395dfdd7d4

Observation 1b5e5b47-3b10-4bcf-8972-443f065ab236 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T02:32:19.142604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:32:19.142604Z digest=sha256:994e0c7dbc323bc192e5fbb97e018d625176822b99954ad26f386b3066fc17d3

Observation 2f0f7346-005b-42ac-ba88-c2df11dc3a3b · inbound

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models cites this paper.

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-30T14:45:01.915411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T14:45:01.915411Z digest=sha256:09dfcdf28b6f37540e762a8c2bec2cbca38f2b6fd36c8660f18372c959fbf3f7

Observation 8247ecf9-282f-4cea-99d9-e682e690fe3f · inbound

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models cites this paper.

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T10:23:14.040221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:23:14.040221Z digest=sha256:7e0de2bfb72e98f0a2e18c46fe886b82d1eb9c69efa94f8f87a4705fc184478f

Observation a95f753a-b662-479e-bbb8-80748f72cebb · inbound

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy cites this paper.

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T01:32:25.304633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:32:25.304633Z digest=sha256:0963fd6460216db4ae10739e908b1b6b1d2466681d0aea7dd89641fe488d8639

Observation 97963ac0-26fe-4a7e-b568-b59535017128 · inbound

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation cites this paper.

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T19:38:41.768698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:38:41.768698Z digest=sha256:660d31b0ebfb12b092f445e666485c0ceb3377f8f6da42cce5a64eb9bcd25e85