Pith. sign in

Paper Citation Record · LEDGER

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

As of 4 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 5 inbound Pith citation observations for arXiv:2602.20200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.20200 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T13:22:16.242427Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:21:28.231732Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T10:19:47.638604Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact35
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c36938f3-25c1-48c7-bbf1-d04605fb7022 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.189417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:71405e40303c29ecaa7bcbc324bd7a8a635bd4e8d667ae93264a5da12aa0307d

Observation cfb9c089-2ee1-4949-a5ad-ebba25ab8930 · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.205486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:5361b2432ebbb277ac6fb2497537f26afd2d024961f9c6da0be7ae4ac0c747bd

Observation 6889f253-dfc8-4dcd-a448-0c93cc281b4f · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.178582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:e08b9c15afccaad0bf56b1c727dd79205e518f985bf4eca0db1b813073b84804

Observation a3c147d8-c0db-40ed-9a30-ce7d1de15001 · outbound

This paper cites Lion: Empowering multimodal large language model with dual-level visual knowledge.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Lion: Empowering multimodal large language model with dual-level visual knowledge

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.910409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:ca7359c70e47d69afff528fbc7a412a56f77799629d950856444a0b5bdb910ba

Observation c5b74ff7-c0e0-4a79-8cb2-5640795465aa · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.116156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:97e91a9a8a63ab53322d1c46331b4da0a0f16d006a176698655290118bc9bfa8

Observation 03741836-0b9e-43b2-b70f-2cc24e2faa48 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action dif- fusion.The International Journal of Robotics Research, 44 (10-11):1684–1704.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Diffusion policy: Visuomotor policy learning via action dif- fusion.The International Journal of Robotics Research, 44 (10-11):1684–1704

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.913419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:2af580c858a89916d30777e5d2c5e80536179711a613bd3bb6b00f40cf879041

Observation 9ed80f28-0240-40c4-9ae9-0c40f03559de · outbound

This paper cites GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.163597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:9ca12e55427a6664d58f49cb034643b7a00242ec3e9f64918b859ac402a0bb9d

Observation d4bc80f0-0a42-42c3-8262-92a482562e0c · outbound

This paper cites Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.178405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f1b34e19cfe2fc586ccc7828d4627ed3befbffaee4b1f6892984550402ed29d8

Observation e126d505-56f0-48d2-be18-992373d08f99 · outbound

This paper cites Mamba: Linear-time sequence mod- eling with selective state spaces.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Mamba: Linear-time sequence mod- eling with selective state spaces

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.904026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:6cfc6c9bb209b202f9e1b206ae6ec036112f936825e0e1b353ac762e47fdd5ff

Observation 36515dd8-4490-48dc-befb-56f4de6bc702 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.185798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:51bdbefbb7b91f944885ff5682f4189604b45893ca75cda6ec699b79ea75ae60

Observation 8467f861-361f-4b42-b5a7-ebd69db3e302 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation An Embodied Generalist Agent in 3D World

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.080578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:2bdcfb6549f137f66579149c3d35260a0c5ac9c76d2df8ac050f13e0186e8c87

Observation 92826e6a-2113-45d5-840c-7794690bd75e · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.161214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:3d48389ce911c1a3fc232402fb18be7caab3da66d32e2573b611c0894a670463

Observation e1cd81eb-75cf-4725-9244-0eb46bf1966c · outbound

This paper cites Vima: General robot manip- ulation with multimodal prompts.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Vima: General robot manip- ulation with multimodal prompts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.906703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:81f8413496f9893944a7c62f24d18f67c91e50f132f52f6bba892be029f4df36

Observation a7c8f85c-63f3-42e1-b441-8cbdeda8e98a · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.198218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:ef90156e0328bc5cd8ec5aa906a1cbe70736a4caeb9d33be8f09ce2a638cf130

Observation 3e512f60-c9d7-414f-a56a-fa636be755ec · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.135802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:1f9914df0f154f72a867e11584f32826e82aaffeacade96b949697fd3620a7a4

Observation ae2fb953-5677-4694-8c96-a38ec29182c7 · outbound

This paper cites Star: Learning diverse robot skill ab- stractions through rotation-augmented vector quantization.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Star: Learning diverse robot skill ab- stractions through rotation-augmented vector quantization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.916269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:610d810eeb7be256db2f46058a9dabedf8337161194e138aed6a833a36f02d34

Observation 69740d5b-f5af-4cf8-a247-527255dc14e6 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.201565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:e07d769b7dc9cdd89c1393ebbfe8db91b046c5ba70a4d3c429a1dcf1c332d871

Observation 8eec5656-5317-4d5f-8da4-dde412307071 · outbound

This paper cites CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-28T02:04:09.707556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:a741be3591071e66e5212997be12c1bc2a1f3b801525f44dfacbfded7c3cbf74

Observation 1a1a5060-f768-4905-a509-fc1a6fa2f401 · outbound

This paper cites Semanticvla: Semantic-aligned sparsification and enhancement for effi- cient robotic manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Semanticvla: Semantic-aligned sparsification and enhancement for effi- cient robotic manipulation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.897875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:ee3cabe37f0fc32af763419abdbd3631d29aa79f3cdde79433f9c2c3584bb739

Observation 3248ea7c-a93b-453a-b155-b54935d2e0af · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.195439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:ef574e3abd484972cc4859568384fac1412209d5cbc234f469305cf76b720b74

Observation 22555f50-e36e-4cc4-a59f-c14e83dbfb50 · outbound

This paper cites Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.100723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:2b8023c7c4aae268f6f0df2dfb0ded827b51d74339eabff295f3d15fe187d402

Observation 268df47c-f363-4d13-9491-846643828fb9 · outbound

This paper cites Optimus-3: Towards generalist multi- modal minecraft agents with scalable task experts.arXiv preprint arXiv:2506.10357, 2025f.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Optimus-3: Towards generalist multi- modal minecraft agents with scalable task experts.arXiv preprint arXiv:2506.10357, 2025f

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:24:11.108246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:87c9af3101a00a14bb0cf05b72fcf233db75cbb5aa540532294f54e9dff977c8

Observation 5435e479-3ab4-4038-8cea-ce6a44324272 · outbound

This paper cites Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.895060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:0043df73baaaa5acb6bdf09854e6b31a414ce521495cf9a3da18bcbab7dda597

Observation 560d6cc9-c24c-4395-8197-9ffa9b06767c · outbound

This paper cites Flow Matching for Generative Modeling.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Flow Matching for Generative Modeling

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.170701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:7bb962f72bba6164eac18dcfc9816d2453b40d264381805f7099eab83c15c7a5

Observation 12423316-0a45-4a43-9215-ef136d276862 · outbound

This paper cites Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.883955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:532a73e0dc256b68765c5d3067d2579430671d42949a8c2662aa7b7b6a618cf4

Observation c79dc206-a108-4f70-bbe8-062edc62ad09 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Visual instruction tuning.Advances in neural information processing systems, 36

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.880264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:6a445f54e3083c466b9f7ebad6eb8c654fcf5f8b3b180ff266edd9ea2c921abc

Observation bc46999c-8e9b-4670-80bc-cba205012ee1 · outbound

This paper cites Towards generalist robot policies: What mat- ters in building vision-language-action models.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Towards generalist robot policies: What mat- ters in building vision-language-action models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.886754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:d7283c0838d84fbeaf567bd5fa0bc067feb892493f5b40fe45506db18e508a32

Observation 65913c17-6396-4754-8d7c-dd5542ba1599 · outbound

This paper cites Towards generalist robot policies: What mat- ters in building vision-language-action models.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Towards generalist robot policies: What mat- ters in building vision-language-action models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.889683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:eaf21d1423e1a2986c6630b47c8235af2d347ba7fe2d8c7254450586b2a87592

Observation 99a8d1b0-6843-436d-bb04-5566eb15c7ec · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.175324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f6bb6c260ae359ba9f6f4828f8b9ab2fa0882601766b7d56cb3a919273551ca6

Observation 1a2feb2c-5b28-462f-8172-6a7e755df62b · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.201236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:108ec42023b57a2f9cd15c1f9003811e010c38fbafbc7f0509cc76e25bfad976

Observation 7b3933e7-42fd-40c8-acf8-a37fcb7ca00c · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.197847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:b4928d6bce3c72ffcba943b012cd2978da366425e6e5af4d4ff66fee85787efe

Observation 8eeca8bb-fd18-4e3c-b03a-7443b11eab27 · outbound

This paper cites Puma: Layer-pruned language model for efficient unified multimodal retrieval with modality-adaptive learning.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Puma: Layer-pruned language model for efficient unified multimodal retrieval with modality-adaptive learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.874915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:581fa651ae2e11265a6293e3d4ad5a4319c8442e232ed408271d3ea368ccc85f

Observation d7ff144e-b594-459b-a49a-3fd014be0709 · outbound

This paper cites Hierarchical diffusion policy for kinematics-aware multi- task robotic manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Hierarchical diffusion policy for kinematics-aware multi- task robotic manipulation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.872206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:2633b06cb2bbe3633328d91654ab44ca180c547e1dc83ce4b18aca6349cb7138

Observation 8cc85b31-ddc2-41d8-8af8-e4037b105c9c · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.IEEE Robotics and Automation Letters, 7(3): 7327–7334.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.IEEE Robotics and Automation Letters, 7(3): 7327–7334

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.877656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:3a74e3cb94d1b4f82356ac056cf04c1865f2aa2f2563dadcbbb08a2dbe5016d9

Observation b0457362-ecb4-480e-87f0-0b02f5e9d0d2 · outbound

This paper cites Vision-based framework to estimate robot configuration and kinematic constraints.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Vision-based framework to estimate robot configuration and kinematic constraints

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.892303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:876f21c68b5d4746778ff0fd8d6f644b605ddee368897e2a0fe49a402c35e55f

Observation 23e97b66-5306-47be-8230-9a484a30460b · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.167016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:2013e49a42453536b138a38bb30a1ff128b7a1befad33d11151713d5d6ba84ed

Observation 1a2c6248-809f-41f4-a7bd-8e9016fe2289 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.150059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:e669b512df2d91fa7ae7528084d506dcd4d194532e4340dc27be02c821bd6ff4

Observation 86fed208-8d76-4852-be0c-cd5ab6d1b360 · outbound

This paper cites Multi-adversarial discriminative deep domain generalization for face presentation attack detection.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Multi-adversarial discriminative deep domain generalization for face presentation attack detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.860328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:3296404d8c48a2c5e3b5c0d02e1c00a5da4ebb08ca8d4c12b94cb80ca86154b0

Observation d54a68db-5012-4b0a-822d-c90371397bd0 · outbound

This paper cites Detecting and grounding multi-modal media manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Detecting and grounding multi-modal media manipulation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.864451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f5d62e6722dfb73b9f50b61bcb6da80ed8c2fde1b9cdee4be686c83d7f6dffed

Observation d0ec3246-ea55-4ac6-9286-42f13e2cf9f5 · outbound

This paper cites Detecting and grounding multi-modal media manip- ulation and beyond.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Detecting and grounding multi-modal media manip- ulation and beyond.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.861627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:af7a53693d6d36ce0d8cdfbc94f6730fd78e4b2e02e1988b417efb4753210767

Observation 25273529-c0a6-4d7e-b3b8-4655cd7ff9d9 · outbound

This paper cites Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.186177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:29f468d064e2381474026939aa36fcb2f21444cc50f93fdadf7690665e4947cc

Observation dabd7e79-f8e3-4ab5-908a-f41f7b0cb269 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.163984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:25b0e1fccd66034b8f10cd82101e84cd07f907f13246e33a0d3ce8a4cbbc0049

Observation 04b328bb-984c-41b3-aa2e-30daf4076e18 · outbound

This paper cites Accelerating vision-language-action model integrated with action chunking via parallel decoding.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Accelerating vision-language-action model integrated with action chunking via parallel decoding

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.104987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:b2ceede055eea2f29b6f3abb4318c9f0db72431d742aa47e7743707a6c41bc0b

Observation fa147f79-109b-48ed-8d65-e62d9182f7f1 · outbound

This paper cites ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.084402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:74357f6c7ae310962d646c7f23c3d4d4a97347757e7671a0676833ddc620e806

Observation 29591d5d-bcda-4ac1-aa91-922df8fbbd3f · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Octo: An Open-Source Generalist Robot Policy

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.169651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:0a56e1ee83859e38fdda4c20c86d340a2cc0d61ed09d8eb2bcb378949e44cedd

Observation da44e4ee-eb70-4808-9b8a-531ca00c7d85 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:38:25.165602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:9924d966a69b53b063fd68aaa9110fcf1b59701e1b7f24319aac47fb68b61ee2

Observation fcf1428c-b943-469f-8db4-0501335a499a · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.864229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:995c8852d40b7a9d69975fe3f964d7e394d0532d9ced698b5c3ce5ba0e2d5749

Observation c09fb7b5-11e2-4a40-bbab-303e06716379 · outbound

This paper cites Bitvla: 1-bit vision-language-action models for robotics manipulation.arXiv preprint arXiv:2506.07530.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Bitvla: 1-bit vision-language-action models for robotics manipulation.arXiv preprint arXiv:2506.07530

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.148294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:eff6bee302026c1426c181448be275bd02934d1c062b5be7529b2edb9abb2520

Observation 7a373401-68f9-4793-a618-b3787c3f6b97 · outbound

This paper cites Vla-adapter: An effective paradigm 10 for tiny-scale vision-language-action model.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Vla-adapter: An effective paradigm 10 for tiny-scale vision-language-action model

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.172579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:205f81bbb1a79a51882bb3bf963967a1dc73eb15fa6871348d98561c04dc0e44

Observation 081862ef-3547-4637-a3b0-b37c748d350f · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision- language-action models for robotic manipulation.IEEE Robotics and Automation Letters.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Tinyvla: Towards fast, data-efficient vision- language-action models for robotic manipulation.IEEE Robotics and Automation Letters

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.919460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:aa6b07d66d1e30dcebc6de1ac586007f9cf8f60a072ced68a46f8d922dd0e4e6

Observation 1307338c-d422-4ffc-9c28-35f999134ecd · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.189681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:66c0bb30c63f2bbac4c81f2016fe0bf627d36ea2435c4b47df6e4007fdde35b9

Observation a5932355-278d-421b-aec1-1ead141326b3 · outbound

This paper cites UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.167790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:5c8c2d05e7a252536532acd4101791a4a6d91fa7377d9a5f2717813b00cc63c3

Observation e843d8a9-ef1a-4a6e-b895-d146ec04b0f2 · outbound

This paper cites Flowpolicy: Enabling fast and robust 3d flow-based policy via consistency flow match- ing for robot manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Flowpolicy: Enabling fast and robust 3d flow-based policy via consistency flow match- ing for robot manipulation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.901013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:87a2b47e437e4c21a5033154385894b30b9c301a191d21ed31117e636ce92de1

Observation 918a5376-5ae0-43a9-8086-5803079a16b1 · outbound

This paper cites Falcon: Resolv- ing visual redundancy and fragmentation in high-resolution multimodal large language models via visual registers.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Falcon: Resolv- ing visual redundancy and fragmentation in high-resolution multimodal large language models via visual registers

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.858971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:a14a1ce94218810b72a4906d1040cc6501274e40ee479c136658a3631cf7b21d

Observation 11b810bb-8db2-4bb7-a497-8c2121b4a601 · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.193217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:a3fe2ce6a1d4366430db0d6dea47ab844c449f1132c4ab02cc437f6c04b53b04

Observation c64e0994-e7e7-4b8d-94dc-2a69484b31de · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.152024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:b4f8e1fca8b146a6118df2f85ffaebec7fdaf468de92523d0fa86d88b90ebcf6

Observation 68c8f3e0-eeb5-416f-b3a2-f951d32269d2 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.192512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:c5ac43e26590036557a6e67987bd36b8e38cf86c9bdbaa4b82fede935ecd661a

Observation a4fd4e77-e2d6-450e-b6be-13cb9e44ee71 · outbound

This paper cites Hiconagent: History context-aware policy optimiza- tion for gui agents.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Hiconagent: History context-aware policy optimiza- tion for gui agents

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.143272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f092c5515caae25cc9ccc1145cead31ffb80e12c6c6636493a1389394f120e3e

Observation 074b22af-c00e-45dc-af7c-ed6ca4a05a06 · outbound

This paper cites H-gar: A hierarchical interaction framework via goal-driven observation-action refinement for robotic manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation H-gar: A hierarchical interaction framework via goal-driven observation-action refinement for robotic manipulation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.866903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:fe680b567b27a355cff3ef9c1fd3f09a454cd573b29199c8a87d961f2dd46d6a

Observation 1de2ed2c-ed6f-4d6c-8006-06f02068c351 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.869605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:32b788342f60bdc4f71a34d259492e92f8499a0ec0f4e996d4614b8b416f10f9

Pith citing papers

Observation d64ebffc-673f-48f0-b3af-27509d2fe77b · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:04:26.381581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:af7d17c57841fcaf5df0aefeb81281b8ce4db20c25d24bccc7bdf7231a8c0af8

Observation c34fc912-f8fd-41a6-af73-f7b58750e817 · inbound

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark cites this paper.

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:26.381581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T03:50:18.706396Z digest=sha256:a1f85680b890110531942469b093218e44d089fe567d1e56a3ca277ec46f32e9

Observation 3cafd5e3-bee0-4479-8b13-815daa6d21a4 · inbound

$\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models cites this paper.

$\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T08:17:45.666875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T10:50:55.866910Z digest=sha256:fbd5ec2a7217eee070f0cbc61877c026a29161ac905e585ce2fe67929377cae7

Observation e51edb37-ed75-4a39-936a-e3de65548c87 · inbound

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models cites this paper.

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:19:47.640060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T08:57:01.861091Z digest=sha256:57fb11ce0c72e95ee413af3b754774f332f911ad156155349ebeb14f6bc2ed70

Observation e15d5c37-b752-4ab5-9b42-521937032fc4 · inbound

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model cites this paper.

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:28.231732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:28.231732Z digest=sha256:69e8d6428fa2c6341c197a2536c85f9bf7877a68b1ace054d0e31c1ae1f2c688