Pith. sign in

Paper Citation Record · LEDGER

Unified Vision-Language-Action Model

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 55 inbound Pith citation observations for arXiv:2506.19850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19850 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:35:33.987020Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T04:16:48.663138Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 740f90a9-fd72-4524-a61a-0cb9f27d7acf · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI Unified Vision-Language-Action Model

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.573510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:5163f6ea706270039b974318b22fe29516d94b4c0ab157b4483139d127b556c2

Observation 35480d38-3d8d-435c-a75a-6db8a4a4b601 · inbound

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver cites this paper.

ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver Unified Vision-Language-Action Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:32:47.040368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:32:47.040368Z digest=sha256:963889d6c26483abd5815a51f674c87e7aeb3f70ab4dad9435fc7ff1beb4897d

Observation 2d23d0d7-e8a3-41e0-bbc4-277c8bdf65e2 · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey Unified Vision-Language-Action Model

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.143162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:d97311ae8393d2892a6d2fe8e873e2a09ff85c562ea712a4a2904ee21ac007b2

Observation cb223d44-20aa-48b3-af85-73896e5f6ff4 · inbound

LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes cites this paper.

LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes Unified Vision-Language-Action Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:12:57.859388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:12:57.859388Z digest=sha256:f3ca90dad0bc20cf62c2099944dc5e40d1067a8d8105f5c58c5db0f257720c85

Observation 2ec0e7ea-4a8b-4ff6-b753-ca79fcb00ab2 · inbound

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions cites this paper.

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions Unified Vision-Language-Action Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:42:47.075327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T12:42:46.978148Z digest=sha256:c56755d2db1c3809c80f9325a5f2da3f56f77b010d17e66d02d702f3adf43005

Observation 4044bc7b-f83c-48b7-8182-ae80c3942219 · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Unified Vision-Language-Action Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:09:40.041366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:b150395343c34a4cc4915f358744c1d1cea7ef8c1701e282b7deae3775dab31f

Observation 7a4cab14-b8d7-443e-8de5-6a34809b66cf · inbound

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey cites this paper.

Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey Unified Vision-Language-Action Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:13.692307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:13.692307Z digest=sha256:0fbc96b02d481caced97b808238da9eb8aa37672478791fdad00277ff4f14b69

Observation c4569fcf-d0ab-4f00-9262-a77197c27da2 · inbound

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention cites this paper.

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention Unified Vision-Language-Action Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:29:09.967252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T06:28:22.652509Z digest=sha256:aa53fc44fc218227f4a153a2388e63c7cf9af20b151313c786a6449fbe632899

Observation 83fcaaeb-26c8-4395-90d0-f07501b58f5c · inbound

Mixture of Horizons in Action Chunking cites this paper.

Mixture of Horizons in Action Chunking Unified Vision-Language-Action Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:32:29.361062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:32:29.361062Z digest=sha256:ab86ef0b7080d33f04629e300b8068cf88e969c51cef22d42b55fe8480c16aae

Observation 7b6fea4d-07a6-4972-92d1-cbe45c9df9aa · inbound

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models cites this paper.

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:01:20.332076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T23:01:13.910539Z digest=sha256:311b8676355889b4163b6d96c1bc8465fbe3788e0c9fc274ba14cbabbd4d3be5

Observation e3b59dde-b461-40da-b899-03f301ec01a7 · inbound

Causal World Modeling for Robot Control cites this paper.

Causal World Modeling for Robot Control Unified Vision-Language-Action Model

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:53:52.389656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T13:53:52.188890Z digest=sha256:238789060e124e794bbc948472536b11a854a88e855ebfd4b33872cfdbc292c2

Observation 290bd683-3ae9-4992-81a2-33e6ab46285a · inbound

MobileManiBench: Simplifying Model Verification for Mobile Manipulation cites this paper.

MobileManiBench: Simplifying Model Verification for Mobile Manipulation Unified Vision-Language-Action Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T04:23:45.812263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:23:45.812263Z digest=sha256:61012a05b11db55dc3f8f46064069d38967703d3dd71eb1f2678a7186d588be8

Observation 60360840-9f9a-46e7-bb28-410e232a71ad · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs Unified Vision-Language-Action Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:30.135571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:30.135571Z digest=sha256:3710e726d889fb9068d2d182fe827db085a5d2aaf722da1525c724b9f1b9999b

Observation 1394df3e-c321-4484-8a4e-b39177d036cd · inbound

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation cites this paper.

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation Unified Vision-Language-Action Model

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:30:20.954689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T21:22:41.935691Z digest=sha256:05e11ce5cac838ac1754f240fd1719d340933b83666b9309de8465d365802488

Observation d04fda09-55f1-434f-8131-9139b03cbe07 · inbound

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models cites this paper.

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T19:35:38.523434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:35:38.523434Z digest=sha256:273c681c2332f5ead2e776626cdf27a11cb86d3b473a29fbda4d9d3df28c26e4

Observation 133f9c06-ab77-4dcd-8a2e-422ef9694f65 · inbound

History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation cites this paper.

History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation Unified Vision-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-15T13:48:22.524177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:48:22.524177Z digest=sha256:eebe74adda4fa7b5471554634ff449353462fc87f84c319fc838d748e9d3e184

Observation c7bcba76-6aee-4ecf-aa5d-2dfe5acc0bf6 · inbound

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies cites this paper.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Unified Vision-Language-Action Model

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.572350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:134569cfbba64a6296253d4d2f928d12a4a8cffb834002b4cff3bef2428c1dcd

Observation 7a5dfea7-4534-44d5-821d-af2a7f0a68b6 · inbound

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising cites this paper.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Unified Vision-Language-Action Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.384271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T10:40:04.767657Z digest=sha256:4cd9e968d4c52bd3d23f53e2d08b973811bd91fdeac8450d6b16564cc36d451c

Observation 89faec9a-5620-4c9d-be02-6942e9af5b20 · inbound

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising cites this paper.

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Unified Vision-Language-Action Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:12.234668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T03:20:04.435904Z digest=sha256:4579f6fca409b5fdceabba334bf56453dd005727b12b1b73b66924221d2e7d6f

Observation 598b4953-0f3a-4c91-b086-01c75764fc8b · inbound

Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation cites this paper.

Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation Unified Vision-Language-Action Model

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:07.202750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T19:51:05.220627Z digest=sha256:ee4b5ebdc6f9e9974720c027787eda9bc3b7e375266d66ab09bcf3b0023ea202

Observation 777055cf-9ac9-48db-a2c4-5098a9a3d26e · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report Unified Vision-Language-Action Model

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:51:30.868908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T16:06:10.595164Z digest=sha256:73ad2c4c6d81853f25f2589839e31a49f5dbdad6914ad017c06149eebc423604

Observation cd6730fa-b2b6-4168-87df-a0da8ee9ca5a · inbound

RLDX-1 Technical Report cites this paper.

RLDX-1 Technical Report Unified Vision-Language-Action Model

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:34.928333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T19:02:19.756092Z digest=sha256:d7dad1a3d1b5e7c4ab1ce797453861241d99e988d18041692cdd5ed95bb5b113

Observation 3801cb71-65e2-46bc-8bd0-82fb2858dfe2 · inbound

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation cites this paper.

OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation Unified Vision-Language-Action Model

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:11.089901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T08:52:07.545191Z digest=sha256:2f9634a5083a3ebd75c27796d7229d4a345cf3460e479b60efc065a53980b806

Observation a146e105-e257-450b-bd03-d69e71862306 · inbound

Test-Time Training for Visual Foresight Vision-Language-Action Models cites this paper.

Test-Time Training for Visual Foresight Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:51:48.668131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T01:33:30.939425Z digest=sha256:e5bb06ee1ec7485954b4215b9644196f613c03cea3d10fcb086856524919fd2a

Observation 8ae03021-5ce3-4e7f-8c2b-4c6594346514 · inbound

Test-Time Training for Visual Foresight Vision-Language-Action Models cites this paper.

Test-Time Training for Visual Foresight Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:25:09.775680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T00:17:58.955835Z digest=sha256:5d3c4e189bc6c013cee0a94d28a93c3e618b6b61da263bbfd58d594f628580ad

Observation 053bd747-96f1-4063-b084-19f93aeab42e · inbound

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models cites this paper.

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:32:34.405866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T18:31:09.883545Z digest=sha256:a485f77feee36fd76e53ba5084f64f7498512eabc03b77263dda54043ec586d0

Observation 49cd27f8-16ba-4051-82a3-d5fb648e8d3e · inbound

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models cites this paper.

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T14:10:07.016537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:10:07.016537Z digest=sha256:98a3ebcc098b5cf63ae27f386d52e769cca49f7ee2da320f1ab9c3a0b73c51dc

Observation 5bda2089-fd48-483b-9c49-b3b519d8bc53 · inbound

EponaV2: Driving World Model with Comprehensive Future Reasoning cites this paper.

EponaV2: Driving World Model with Comprehensive Future Reasoning Unified Vision-Language-Action Model

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:02.541356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T05:14:28.714494Z digest=sha256:ecd308c5ce0124b5b884c9c7f8a12308eeaf6dfd0b77b7380c0811089345a4c7

Observation 0abff421-cff0-4b99-8fbe-300cd05bce3e · inbound

CLOVER: Closed-Loop Value Estimation and Ranking for End-to-End Autonomous Driving Planning cites this paper.

CLOVER: Closed-Loop Value Estimation and Ranking for End-to-End Autonomous Driving Planning Unified Vision-Language-Action Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:59:01.691669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T20:57:06.455108Z digest=sha256:2dd0ab6c6053d05ee0c87f68b8621f025513f71fa5e0f91efb8261e8be099205

Observation 8d9aed89-8314-4974-815b-6ecdb9d97484 · inbound

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models cites this paper.

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:53:04.974302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T04:48:39.069675Z digest=sha256:5fd17032c62b38c6d36c4d70e7963cae3b2df79b4d597575c7c393da890d387d

Observation ea7158fc-6c61-4ae3-80f8-d051d1b3c218 · inbound

ChainFlow-VLA: Causal Flow Planning with Vision-Language Models cites this paper.

ChainFlow-VLA: Causal Flow Planning with Vision-Language Models Unified Vision-Language-Action Model

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:23.842498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T04:37:56.487048Z digest=sha256:dd6e218aa8ecd19352037385dda7288988dafd7a0cf9c4a25c2c63817a8f2ffa

Observation dbba18dc-66ca-4180-bd02-66cb42fef98a · inbound

Point Tracking Improves World Action Models cites this paper.

Point Tracking Improves World Action Models Unified Vision-Language-Action Model

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:56:37.138179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T03:51:10.921839Z digest=sha256:fd0d0d53c50664522d4c5c5fe048c124d432e30636cae1b04c1e4db2f537b01a

Observation 023294ce-a3c6-4703-bb55-836769d19cf3 · inbound

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation cites this paper.

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation Unified Vision-Language-Action Model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.302091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T21:27:26.682689Z digest=sha256:96c72935a93ed89184b8da7699e462108b70d4bd8701134b08a22e594916702b

Observation c087c14f-7309-451f-ab3e-31e832e2427d · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence Unified Vision-Language-Action Model

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.871097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:0f8f1d2ce08341c3e1accbd2dbd14b06171a5dbabb1a7ed79842dccbfd8f0f74

Observation 01ab5dfb-8e37-4b87-987f-45a99ffe4a0b · inbound

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling cites this paper.

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling Unified Vision-Language-Action Model

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:27.830310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-29T13:33:03.368006Z digest=sha256:6acd203ac1d6e312835dc36e3cc44d790536ffd1a5cc344ec2b8e3de25859f0b

Observation f875f143-b957-43f9-82a7-4df1734eb19e · inbound

OneVLA: A Unified Framework for Embodied Tasks cites this paper.

OneVLA: A Unified Framework for Embodied Tasks Unified Vision-Language-Action Model

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:26:14.048938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T17:05:08.124096Z digest=sha256:f92751e9c200ce6a62f1a75c2b089871cf30ee3e1b94fd92026e717d1c6ab116

Observation 1dd09f7c-acf4-452d-9574-ee945187fe57 · inbound

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs cites this paper.

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs Unified Vision-Language-Action Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:36:24.018891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T14:04:43.270155Z digest=sha256:34643bfd3fd62010c100f2f2678be104ebc8b58099ade65341e20cfae2c9274e

Observation 501e6f11-2725-4b97-b708-156fd4598da2 · inbound

PointAction: 3D Points as Universal Action Representations for Robot Control cites this paper.

PointAction: 3D Points as Universal Action Representations for Robot Control Unified Vision-Language-Action Model

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:16:34.956886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T10:09:48.280446Z digest=sha256:cd152ca6f1c5b0a3a8a45cebb7c0305ee4c5a2f4666cf18ea904ad637b76b394

Observation df4ccc11-83d6-4e41-b137-b58af8f3162d · inbound

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis cites this paper.

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis Unified Vision-Language-Action Model

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.933519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T01:05:26.382826Z digest=sha256:1cc59e7cdc371a2e7f17b21d527861247342a65a115b007d4beba86e34a841b8

Observation 187c603b-3dfc-40e3-becf-c82124ea6d0d · inbound

MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action cites this paper.

MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action Unified Vision-Language-Action Model

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.243423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T01:10:20.066545Z digest=sha256:acb7aae7069e2d893b574c0c8b638be2fc2a2ce3045b9b44efd2d466a9e93c68

Observation c40d48a5-4484-48c9-b08c-0b7937045748 · inbound

Test-Time Trajectory Optimization for Autonomous Driving cites this paper.

Test-Time Trajectory Optimization for Autonomous Driving Unified Vision-Language-Action Model

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:17.899646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T21:53:02.554405Z digest=sha256:22cef4124084af4aa17b73c7125a65eb8328ddc6852bc34575b8560ab4db50a5

Observation 6c1ea0c2-47a8-4187-a1ea-dd764e601de9 · inbound

TBD-VLA: Temporal Block Diffusion Vision Language Action Model cites this paper.

TBD-VLA: Temporal Block Diffusion Vision Language Action Model Unified Vision-Language-Action Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T18:57:16.673533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T21:49:20.600217Z digest=sha256:91dfb895949844eb0cf720a72065e847a97b3a55bc43934211900cd1ed275894

Observation 1e221957-2a6b-4ecd-b5bb-341d38b16780 · inbound

Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments cites this paper.

Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments Unified Vision-Language-Action Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T11:22:19.744983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:22:19.744983Z digest=sha256:8a2b25a03027b076605563b0bd7688f33b968206d741f3f7cc7abb2df119941e

Observation 480057ea-12ef-494b-8d7a-29072ff6bc27 · inbound

ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation cites this paper.

ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation Unified Vision-Language-Action Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:58.672755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T00:51:05.494085Z digest=sha256:8b4acd1b73f7529fa49925ad4ed1ea00d7bf9c8756a8bc4a20306be5f09b43e0

Observation c25c564c-8273-4211-9a7e-a372c55ad769 · inbound

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? cites this paper.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Unified Vision-Language-Action Model

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.478757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:8cbc3e613a8f90e19c781b62c92f0745d24f07a633d1d28b46ba91b36a5acbde

Observation 2ab186b7-f130-47a7-bc97-51d946774f12 · inbound

Inductive Generalization for Robotic Manipulation cites this paper.

Inductive Generalization for Robotic Manipulation Unified Vision-Language-Action Model

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:09:36.791150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T14:52:15.202629Z digest=sha256:ba0f9e6eddc15bf6618a546d25875c941a1f57a334b46393507190db617cad93

Observation a5102fc2-e695-4665-8e37-989022e009ae · inbound

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling cites this paper.

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling Unified Vision-Language-Action Model

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:37.521710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T14:27:40.585782Z digest=sha256:6a22882f0120828df31d4e588c88a38bca73d9d54b3dd4f7a3795fb6d6df5335

Observation 4f5f818d-5c17-4041-a84e-1a7f4a14072d · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Unified Vision-Language-Action Model

Reference 156

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:90159a7dbb4bcf10a9e704ac99be36d346d3cc4c96b66484a98498f8ced41f46

Observation 5647610d-2f1d-43ae-b7ce-0f628088b26e · inbound

Learning 4D Geometric Priors for Inference-Efficient World Action Models cites this paper.

Learning 4D Geometric Priors for Inference-Efficient World Action Models Unified Vision-Language-Action Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-11T14:23:57.266710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T14:23:57.266710Z digest=sha256:05e59b2b261710553cf8a2c9c1314c992491feb951b856dd6387532e6b54b053

Observation d0a44ef8-cb7c-4f95-afad-eb9faee15b42 · inbound

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Unified Vision-Language-Action Model

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.201663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:7121681f4f83f33e223bb98ad40cc79fa9a62d616c0d3a6786de6181c14906a2

Observation 593091f1-a278-4ee3-b6cf-f5cef74062cb · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Unified Vision-Language-Action Model

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:16:48.664420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-10T04:12:29.153764Z digest=sha256:3fb4046c57d6072475fd6fe486a900eeef9f4792a1a3c4e081c3ed9460cd7e93

Observation 53f81131-c0f2-4389-a5d1-33932b1f6dde · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Unified Vision-Language-Action Model

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:43.506974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:43.506974Z digest=sha256:d07c45f7d53e04e8ef9d1a0db09e5807e42d1cc67c57f5b6e518e889b8a2adc3

Observation 9ebb8a0b-fd17-4ff9-8878-e3d4c6ab1d5c · inbound

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model cites this paper.

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model Unified Vision-Language-Action Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:23.625183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:23.625183Z digest=sha256:f00a2557aaf7656bcbc1beb045f31d993eef590d47a296a97a81519794e19279

Observation c0d5ca24-f249-4b16-aee9-1bbd471e85db · inbound

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models cites this paper.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.335648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.335648Z digest=sha256:7641686ffe2ab60c6313c027143d144b88dae1cfbefd2aeb5ee6d1f1419da3d1

Observation ae01dba0-592c-4607-8fed-b77e1323147c · inbound

LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding cites this paper.

LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding Unified Vision-Language-Action Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:35:33.987020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:35:33.987020Z digest=sha256:2276b053093f4b1116e1c003fbe48f8fc72fbf6f8d62d04c20a8fb9ec4a49991