Pith. sign in

Paper Citation Record · LEDGER

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2501.18867.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18867 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:23:17.797242Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:19:38.145306Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 19898cd3-2219-4907-9f3d-17d2f2cee578 · inbound

ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning cites this paper.

ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:17.797242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:17.797242Z digest=sha256:a8ed696f31fdc45714fbe0db2d7355492b6d1472b9883f1271566ebc53e97294

Observation 0bbf6716-25e6-401d-bc0f-69f035bbb1ee · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.525266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:3621bab3b8457dac1f9fb04a8dbc24e6156a2849e76431fd9c2aaeda7c362bfa

Observation 1834e392-4731-4b51-b42b-5e7fecf03bef · inbound

Improving Generalization of Language-Conditioned Robot Manipulation cites this paper.

Improving Generalization of Language-Conditioned Robot Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:00:00.281156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:00:00.281156Z digest=sha256:2a54a0bb7b1c5c1cf597bb698a702c18302c39166108ab55d4b6baf8da96ee6f

Observation 5017bbf7-2e89-4a9d-94ce-046ac81d647c · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:47.885219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:47.885219Z digest=sha256:e71dae64218af587e4af266fb99f2785c084fc82181dbaab190d08152bdcc541

Observation 13d393e9-8b69-4fe8-8a73-c3a3660eb217 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:45.491893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:45.491893Z digest=sha256:2334caa2dd038776cc8fb8ad69dfd5aa70c0c6f8eaf8953fb7cb8d4561ed35ad

Observation 1cea0819-f384-45cc-824d-f912b80e0671 · inbound

Ctrl-World: A Controllable Generative World Model for Robot Manipulation cites this paper.

Ctrl-World: A Controllable Generative World Model for Robot Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-16T01:14:10.412898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T01:14:10.174044Z digest=sha256:f37af62ca95bafcfc6763faf023e8dc73bba3c48ab272391260087208660c1ae

Observation 3d7c705f-3896-4d61-92be-52552f46fe2b · inbound

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models cites this paper.

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:01:20.326259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:01:13.910539Z digest=sha256:afb41a7e0ec6afc43e69b3b0684c461e934c1ecf4951b0c3c113e1a8cd13c250

Observation 82e66e5a-afd9-4728-95eb-7f7b46bde77d · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.889361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.889361Z digest=sha256:c1784e963b6b4ed1cca9408921467f840a40b638f4060a2f5022e49f168e1f3a

Observation fc62d89c-84fb-4888-a78c-81b968be093b · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:08:02.045723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:4f1544284ce54df28532ab30f6cb0f4de8ade181ca229daebbe65ce83529387a

Observation a5932355-278d-421b-aec1-1ead141326b3 · inbound

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation cites this paper.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.167790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:bd2bb19e83337b55b571b7c837f45dafe6f85b7bbe008c2d9ca6e58f8a3bdbf2

Observation c8cd7914-b628-4f76-a057-a5d7f4368dd6 · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.971064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.971064Z digest=sha256:47cb10f6c19b89d6cee6e168e67927bcb396fbf4ca7597c67e13f854a32c26b9

Observation 8c502b30-011a-4097-8701-a4f775628619 · inbound

Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance cites this paper.

Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:28:23.112425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:27:35.706107Z digest=sha256:75909832f5e5735f30a32ed90acebc6c17fa6562c4a94c50ad4d87e145036f07

Observation fe21e7e9-76b4-4ea2-bf6f-c929037bf61d · inbound

DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching cites this paper.

DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:59:34.036536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T22:59:21.473359Z digest=sha256:ecb22bf97606891a445f1adf5155b74b67b6bbd3618d73b9b8f9e5cf3d4b18ca

Observation 215344d3-9b62-4ee6-90c6-e167ac1d2ca6 · inbound

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? cites this paper.

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation? UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:10:49.501816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:10:54.362107Z digest=sha256:5ad7ee3fd6dba804bc892290f3d3ef76070d9346250e4a1a873475b0590cfdbc

Observation 6a9b6c95-3c09-4cdd-bcbb-b2b4d9df7eac · inbound

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning cites this paper.

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:48:48.368301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:06:38.517652Z digest=sha256:95057a3d5517e7721ff63a7af89f62d6b03848cc85062c38886b98b252d6ae4f

Observation bbe15e70-d9c6-4266-b018-17bf1e4020f0 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:53.701218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T03:39:41.090350Z digest=sha256:4bd68960f8823d20d295252c484498fd620c19abc504ea8f52749c11c25ab583

Observation a546ed1d-8b0c-41fb-96f5-371f64fb4702 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:26:29.199579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:53:54.608425Z digest=sha256:cb9ba97e0d283934cf65ec9cf4ce91428779ee0bc44ad264f93b6895b0a41dfe

Observation 1e489c8a-6af9-4c50-bf3b-8b25fb7bb44d · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:19:49.907080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T06:16:48.180290Z digest=sha256:c67d1ed927a2e151e0226cb52d040c4b3640d5b354e04b8967f97dc5e4474f8d

Observation eed7992c-03d2-4756-9207-b9e06610f666 · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:26.881607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:14:54.885244Z digest=sha256:98a3d1d1ea652586b9979080e4c714a9f5e88f3e426844adf2e276026ef9c9cb

Observation 06681ab4-54d2-4ec5-9873-e7dee94d86a3 · inbound

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models cites this paper.

ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:17:59.541966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:14:56.501485Z digest=sha256:4830d59a7999d8cbec9c2df5d46f9ccb3461f9d06cbfc112f60e1d9c1976d56f

Observation 8a60aa44-b1bc-419a-abb0-009de18a24d1 · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:25.755642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:08:43.222818Z digest=sha256:adaa5e013430e35efbb3783ad89fdc6169e751f70419c1f99c7b8ae86d042dcc

Observation bbf4990f-9cc8-4f36-8c36-7d5d8f5c96cf · inbound

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation cites this paper.

UniSteer: Unified Noise Steering for Efficient Human-Guided VLA Adaptation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T14:25:59.846732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:25:59.846732Z digest=sha256:364026480dbf5a88dab3d07a254e6f0fdacb1f5e67987d2d6168f649888a7ab2

Observation 244fea20-afda-4dff-920c-68fe5a1cf106 · inbound

UAM: A Dual-Stream Perspective on Forgetting in VLA Training cites this paper.

UAM: A Dual-Stream Perspective on Forgetting in VLA Training UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:28:55.072771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T19:24:57.339949Z digest=sha256:7ecedc6a3a3b888f6321a6c0de6142a17df9081b9ec95d94899d1d8216bd5e90

Observation a7ccadef-7dfe-4e6e-8a80-ae221c8a6dec · inbound

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control cites this paper.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.582695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:6dc73cbee2fccc15d66c757f965892c04fe9746e69c7bead070953b90fa0c686

Observation a635d8f1-ac24-40cf-a4fe-0205460e94f5 · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:06:08.671326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:05:52.461348Z digest=sha256:e9b8fab1d944a23bfd085b3c31e3ba667537d871da81698c594fb4d26ae853a5

Observation d45834e8-781a-4585-a3e2-d40fe382e7bb · inbound

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model cites this paper.

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:54:58.304885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T16:52:05.565650Z digest=sha256:cd6f0c8f609a6198e26c241a3c661f537b6714ba9243f247862dc5b8dde50140

Observation 0e610237-103c-4e29-85ab-bfb040883cc3 · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.812492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:3b9e5626fe599a358bafc3645adee55079b81e2d6d64eeeedffa07b141d75c35

Observation 0c011b33-cf82-4e6a-acac-4920df51827d · inbound

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding cites this paper.

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:16:59.243402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T01:23:02.576098Z digest=sha256:9e2ecff15263d52a6365c95dac522a9be8937b16d33fadb140dadeb9c02e9ea8

Observation b5e73954-1eb4-493a-b5d1-a114cc96fe0c · inbound

World Pilot: Steering Vision-Language-Action Models with World-Action Priors cites this paper.

World Pilot: Steering Vision-Language-Action Models with World-Action Priors UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.287276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:40:02.137152Z digest=sha256:0fdb5636d4cd8238d82fd3f85b5ad2fef80be377deb88aa0d0fbf6cbc365dc25

Observation 5fe65f78-e8ee-401e-8667-7858116c64e6 · inbound

MV-WAM: Manifold-Aware World Action Model with Value Augmentation cites this paper.

MV-WAM: Manifold-Aware World Action Model with Value Augmentation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:19:38.147129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T14:36:42.051049Z digest=sha256:bdaad6635ed0ef7776f42fc61b3f0376fadb0ab8155634b96d6b234739fafb6d

Observation e6920ae1-14bf-43af-b651-170eb1996765 · inbound

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model cites this paper.

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:04:21.013976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T07:01:48.330369Z digest=sha256:8acadad28d15c775cf54149e8a2fe51762ca4e878ac15b60e10869bba7cfec5c