Pith. sign in

Paper Citation Record · LEDGER

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

As of 14 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 5 inbound Pith citation observations for arXiv:2602.20200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.20200 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T13:22:16.242427Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:21:28.231732Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T10:19:47.638604Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact35
  • verified fuzzy24
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c36938f3-25c1-48c7-bbf1-d04605fb7022 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.189417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:63277b734ec24a10373c832a80cc9a3f21fe0f2979cb397b05527743d3bc2d06

Observation cfb9c089-2ee1-4949-a5ad-ebba25ab8930 · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.205486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:7a0b1080019c9f580d2fe6b43d1ca5e961b59a0c42978b6eb061021989aea9ad

Observation 6889f253-dfc8-4dcd-a448-0c93cc281b4f · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.178582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f3a3a18bc0d7f0b219b3672f5aa427381a7a3bc6b6c612d92536e61fb1cb106b

Observation a3c147d8-c0db-40ed-9a30-ce7d1de15001 · outbound

This paper cites Lion: Empowering multimodal large language model with dual-level visual knowledge.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Lion: Empowering multimodal large language model with dual-level visual knowledge

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.910409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:cf496cc4feafa0f32bc0ac66251af9511a9cffabafb4491a0ce91608fad30efd

Observation c5b74ff7-c0e0-4a79-8cb2-5640795465aa · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.116156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:ba879be31c1f9763a0225a8897ec8d8e55f9bba4007c79e516eddc9c2904faf3

Observation 03741836-0b9e-43b2-b70f-2cc24e2faa48 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action dif- fusion.The International Journal of Robotics Research, 44 (10-11):1684–1704.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Diffusion policy: Visuomotor policy learning via action dif- fusion.The International Journal of Robotics Research, 44 (10-11):1684–1704

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.913419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:4d8d1f29c22e700aeae03b47928fa52fcceda6360ea0b129a52712cdb3d6de33

Observation 9ed80f28-0240-40c4-9ae9-0c40f03559de · outbound

This paper cites GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.163597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:bead584943c17a401b21844925519ed00b501e7bfd95d69f9a2a307704e652d0

Observation d4bc80f0-0a42-42c3-8262-92a482562e0c · outbound

This paper cites Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Interleave-vla: Enhancing robot manipulation with interleaved image-text instructions

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.178405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:10d3947c608c6442aec938cb6a8ba523ebb18a51a3c3373f8222aa39587ef3d0

Observation e126d505-56f0-48d2-be18-992373d08f99 · outbound

This paper cites Mamba: Linear-time sequence mod- eling with selective state spaces.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Mamba: Linear-time sequence mod- eling with selective state spaces

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.904026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:a81430c7a715681618a7d5c104de27684bb95a1e807bd56ee092b6003c206ff0

Observation 36515dd8-4490-48dc-befb-56f4de6bc702 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.185798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:bf605fe5ee3b20f097450c3688e27e28482acf75a63c1474e7dd695a5a9ce2d6

Observation 8467f861-361f-4b42-b5a7-ebd69db3e302 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation An Embodied Generalist Agent in 3D World

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.080578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:65434d70026249afe1b50cda14e23a2ff32a4cfad7f49e20ae4d4122aeb80716

Observation 92826e6a-2113-45d5-840c-7794690bd75e · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.161214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:624e33f92dfebc2971c91c159f0d4e6e4c9b68e1a68bbc1474bbea751ac61f2f

Observation e1cd81eb-75cf-4725-9244-0eb46bf1966c · outbound

This paper cites Vima: General robot manip- ulation with multimodal prompts.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Vima: General robot manip- ulation with multimodal prompts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.906703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:ec3d707d9556aadcbff111304483dcacc51ab23ab05fe513de3771911bcf75cc

Observation a7c8f85c-63f3-42e1-b441-8cbdeda8e98a · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.198218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:746cd68ea358e16ab1b81bcc5c660868a179eb57f239e3406b52824adb8636fe

Observation 3e512f60-c9d7-414f-a56a-fa636be755ec · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.135802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:b5fdaf515b22c60dfabd87a571e6d41f478d74e1d6190b864d41ab45ff386608

Observation ae2fb953-5677-4694-8c96-a38ec29182c7 · outbound

This paper cites Star: Learning diverse robot skill ab- stractions through rotation-augmented vector quantization.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Star: Learning diverse robot skill ab- stractions through rotation-augmented vector quantization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.916269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:08db9eec415c43d0803cbbc3a5b91f075c03919414c9a2c647e02ef5cbf9ce24

Observation 69740d5b-f5af-4cf8-a247-527255dc14e6 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.201565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f158962731cb70988cb25dba10a26f4f041ccbcbf0046fa8a8fd9d002c508f8e

Observation 8eec5656-5317-4d5f-8da4-dde412307071 · outbound

This paper cites CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-28T02:04:09.707556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:8af2d8d5682af2fbc0cc6cbb96ba07a80d08e4ba1321e10771a3867963ceccf3

Observation 1a1a5060-f768-4905-a509-fc1a6fa2f401 · outbound

This paper cites Semanticvla: Semantic-aligned sparsification and enhancement for effi- cient robotic manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Semanticvla: Semantic-aligned sparsification and enhancement for effi- cient robotic manipulation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.897875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:3f2dc3bdd16357bc5ca3b08d1591010a91c3c9effd90df1f9e9e2d93c18f2a46

Observation 3248ea7c-a93b-453a-b155-b54935d2e0af · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.195439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:e761a8995d9451b2aa5e64ed956ac313fad1bdeafd040b32985efe5633dc2272

Observation 22555f50-e36e-4cc4-a59f-c14e83dbfb50 · outbound

This paper cites Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.100723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:bf530570dc2fd0bf7483927c390be5dd3b9d63857a60e73333d72cb6c28fc714

Observation 268df47c-f363-4d13-9491-846643828fb9 · outbound

This paper cites Optimus-3: Towards generalist multi- modal minecraft agents with scalable task experts.arXiv preprint arXiv:2506.10357, 2025f.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Optimus-3: Towards generalist multi- modal minecraft agents with scalable task experts.arXiv preprint arXiv:2506.10357, 2025f

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T13:24:11.108246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f2c6691b833db7636e17c3a93580ed1a982b817161d224a66430ba4629985fbc

Observation 5435e479-3ab4-4038-8cea-ce6a44324272 · outbound

This paper cites Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Optimus-2: Multimodal minecraft agent with goal-observation-action conditioned policy

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.895060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:ea81da0b9b8601c82af6a3c2725d4f88c64173ab71ab7c480e4d1861ebb23d00

Observation 560d6cc9-c24c-4395-8197-9ffa9b06767c · outbound

This paper cites Flow Matching for Generative Modeling.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Flow Matching for Generative Modeling

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.170701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:25c97cf7b8e46736c4ea54129efc5bfdd13bd2643a2f5eee55719689cf5cb88e

Observation 12423316-0a45-4a43-9215-ef136d276862 · outbound

This paper cites Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.883955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:bc0aab126153989a08ccb00e871ec0be66a281ee3854835bdc2caff672e1b0cc

Observation c79dc206-a108-4f70-bbe8-062edc62ad09 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Visual instruction tuning.Advances in neural information processing systems, 36

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.880264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:d934ddbea3dc97cc9c77b009ba1415b3d84c716b696cdb657bbfcefd89e5c0c3

Observation bc46999c-8e9b-4670-80bc-cba205012ee1 · outbound

This paper cites Towards generalist robot policies: What mat- ters in building vision-language-action models.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Towards generalist robot policies: What mat- ters in building vision-language-action models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.886754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f3beb268ae70cf49ffb34c604261a83177d7878495cefcf4efc87e1670d9a7ba

Observation 65913c17-6396-4754-8d7c-dd5542ba1599 · outbound

This paper cites Towards generalist robot policies: What mat- ters in building vision-language-action models.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Towards generalist robot policies: What mat- ters in building vision-language-action models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.889683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:5bf7e5a042968712cb7495762a6468e2f4aa73e1346de892adf773ef8da58c33

Observation 99a8d1b0-6843-436d-bb04-5566eb15c7ec · outbound

This paper cites HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.175324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:8bbfa74917f6fd52a1a1dce346532f6525e2f9c9c0dc64c85e6900ae25e9d51b

Observation 1a2feb2c-5b28-462f-8172-6a7e755df62b · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.201236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:4c94fb4e341ce15c780397fe0c4cfc3c5ac39899ef4fb4a0f92568cff8490d1d

Observation 7b3933e7-42fd-40c8-acf8-a37fcb7ca00c · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.197847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:c2da52a65db3361f057ee175bad41437090e2e1d83c375eba64ec510398b092e

Observation 8eeca8bb-fd18-4e3c-b03a-7443b11eab27 · outbound

This paper cites Puma: Layer-pruned language model for efficient unified multimodal retrieval with modality-adaptive learning.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Puma: Layer-pruned language model for efficient unified multimodal retrieval with modality-adaptive learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.874915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:7da1ddc47d734932025b800ca3d0652674dffded743fc527077f6b76faa099b2

Observation d7ff144e-b594-459b-a49a-3fd014be0709 · outbound

This paper cites Hierarchical diffusion policy for kinematics-aware multi- task robotic manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Hierarchical diffusion policy for kinematics-aware multi- task robotic manipulation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.872206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f1779c54dae29425622df6e58ab0e832754d65c09443fcb70f242d3bee218781

Observation 8cc85b31-ddc2-41d8-8af8-e4037b105c9c · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.IEEE Robotics and Automation Letters, 7(3): 7327–7334.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.IEEE Robotics and Automation Letters, 7(3): 7327–7334

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.877656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:d55df881ec7079a0c3c53bb20ddffc5c15fc5d9938e9f8cbb6eed5ae11695db8

Observation b0457362-ecb4-480e-87f0-0b02f5e9d0d2 · outbound

This paper cites Vision-based framework to estimate robot configuration and kinematic constraints.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Vision-based framework to estimate robot configuration and kinematic constraints

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.892303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:66116eed0c86579f2de09be5f95d7326687553e91260f545f84bb21b82a41b8e

Observation 23e97b66-5306-47be-8230-9a484a30460b · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.167016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:a6a5ccb72b5b36aac6a86581cf1cf1bfd472bffbd26a466cc5cf39900471da5f

Observation 1a2c6248-809f-41f4-a7bd-8e9016fe2289 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.150059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:2586cc8ade28b5f361993fa561a44a0481a3b23578eb6e3b02c08ba0b26b664b

Observation 86fed208-8d76-4852-be0c-cd5ab6d1b360 · outbound

This paper cites Multi-adversarial discriminative deep domain generalization for face presentation attack detection.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Multi-adversarial discriminative deep domain generalization for face presentation attack detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.860328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:ba802cae6ea3ea12d0a8842ae036cd2902c18234280e9568386196bcd114aef6

Observation d54a68db-5012-4b0a-822d-c90371397bd0 · outbound

This paper cites Detecting and grounding multi-modal media manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Detecting and grounding multi-modal media manipulation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.864451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:cffa15bf0241e613f8557c6f520411f6d6f374680b0951e9ec8308522b14da5b

Observation d0ec3246-ea55-4ac6-9286-42f13e2cf9f5 · outbound

This paper cites Detecting and grounding multi-modal media manip- ulation and beyond.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Detecting and grounding multi-modal media manip- ulation and beyond.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.861627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:a09c2a23e9af566e58365992cd2dfce546ddd352b2620ec1fc38bbd593e0b0b1

Observation 25273529-c0a6-4d7e-b3b8-4655cd7ff9d9 · outbound

This paper cites Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.186177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:76fbad689f856e0bef6949cb6671a17f77914c11b68e76e6f7f7ffe9485b2597

Observation dabd7e79-f8e3-4ab5-908a-f41f7b0cb269 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.163984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:c97f07344728be8077d741f9ff508e491c1104871b1ba8b750160e161d785dfe

Observation 04b328bb-984c-41b3-aa2e-30daf4076e18 · outbound

This paper cites Accelerating vision-language-action model integrated with action chunking via parallel decoding.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Accelerating vision-language-action model integrated with action chunking via parallel decoding

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.104987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:7d921b2802b5f91c657fb5fbf34b8b7eb2ef64cade7caa4afc8cc70b20b44591

Observation fa147f79-109b-48ed-8d65-e62d9182f7f1 · outbound

This paper cites ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.084402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:5e508b9c61a1f5c32d8bf69ec6b2ae95b0ab3707954768249582660ffd98d390

Observation 29591d5d-bcda-4ac1-aa91-922df8fbbd3f · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Octo: An Open-Source Generalist Robot Policy

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.169651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:45b381fba3757254f3c66fe9aee56df72aeb35c0b0f548cb21caa3df0e9602df

Observation da44e4ee-eb70-4808-9b8a-531ca00c7d85 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:38:25.165602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:1f7eecf230c1c0b5a604a8b285f37ab3df72aaee9a8f4fb711cc19ed35e40eee

Observation fcf1428c-b943-469f-8db4-0501335a499a · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.864229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f0f581bab40b9f11a903f3ce1a6833d88a2cfd11d27128b9aa0779104ff85d58

Observation c09fb7b5-11e2-4a40-bbab-303e06716379 · outbound

This paper cites Bitvla: 1-bit vision-language-action models for robotics manipulation.arXiv preprint arXiv:2506.07530.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Bitvla: 1-bit vision-language-action models for robotics manipulation.arXiv preprint arXiv:2506.07530

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.148294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:858a224ce6879692ed3e8222d4531fa7725d1554fca80e4cad470f8e23254363

Observation 7a373401-68f9-4793-a618-b3787c3f6b97 · outbound

This paper cites Vla-adapter: An effective paradigm 10 for tiny-scale vision-language-action model.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Vla-adapter: An effective paradigm 10 for tiny-scale vision-language-action model

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.172579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:60fd8c6271b86a4639b5034e9974175d20781a7a81eb2d38aea5b488ef1c263a

Observation 081862ef-3547-4637-a3b0-b37c748d350f · outbound

This paper cites Tinyvla: Towards fast, data-efficient vision- language-action models for robotic manipulation.IEEE Robotics and Automation Letters.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Tinyvla: Towards fast, data-efficient vision- language-action models for robotic manipulation.IEEE Robotics and Automation Letters

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.919460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:404ed384cf23992e009ce448b43844aca2c110754bd69c9b6bffda76b7429011

Observation 1307338c-d422-4ffc-9c28-35f999134ecd · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.189681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:bd22710d341057f6601b1769eab903077e9e17c8539563e977d642c80cb89978

Observation a5932355-278d-421b-aec1-1ead141326b3 · outbound

This paper cites UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.167790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:ea63d772696088a4bac7da1d8c68643589b4e5d51b72812b5042e836fa87b8ae

Observation e843d8a9-ef1a-4a6e-b895-d146ec04b0f2 · outbound

This paper cites Flowpolicy: Enabling fast and robust 3d flow-based policy via consistency flow match- ing for robot manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Flowpolicy: Enabling fast and robust 3d flow-based policy via consistency flow match- ing for robot manipulation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.901013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:f8bd85d39d89600d40c0f3a02502c6f90ebab0bb5fc40b5d6e43f1626f529925

Observation 918a5376-5ae0-43a9-8086-5803079a16b1 · outbound

This paper cites Falcon: Resolv- ing visual redundancy and fragmentation in high-resolution multimodal large language models via visual registers.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Falcon: Resolv- ing visual redundancy and fragmentation in high-resolution multimodal large language models via visual registers

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.858971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:73aa42aac0b63f9d82eb18970b325c24ce073df49e474bfbee372f1ddd330409

Observation 11b810bb-8db2-4bb7-a497-8c2121b4a601 · outbound

This paper cites DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.193217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:914b4d0afd33edf3e34b2d95e72cdb2a21b9593a0f35726536f069723df6aada

Observation c64e0994-e7e7-4b8d-94dc-2a69484b31de · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.152024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:2a4e907a12667677ce064fdd84dd1936c0a65387eab2256c62bc4f73a3b9e3cf

Observation 68c8f3e0-eeb5-416f-b3a2-f951d32269d2 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.192512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:1c67e54d37a54ae0aa8399d6113814e8e6872c53dad506f9e637173273f3f3bc

Observation a4fd4e77-e2d6-450e-b6be-13cb9e44ee71 · outbound

This paper cites Hiconagent: History context-aware policy optimiza- tion for gui agents.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Hiconagent: History context-aware policy optimiza- tion for gui agents

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:24:11.143272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:808ab7b3ff076c755caab401f6bfdd3651407efc91aa76d8f130ce85106db229

Observation 074b22af-c00e-45dc-af7c-ed6ca4a05a06 · outbound

This paper cites H-gar: A hierarchical interaction framework via goal-driven observation-action refinement for robotic manipulation.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation H-gar: A hierarchical interaction framework via goal-driven observation-action refinement for robotic manipulation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.866903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:0718149c74f8843c128052f0f33f221f84145cc722a3d90b0696eebd6ca56b2b

Observation 1de2ed2c-ed6f-4d6c-8006-06f02068c351 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T13:24:11.869605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:5fe0bcb79fbd237fe022eb51585cdc0c0df7fcee07e26b7d7cced27adfea7e6f

Pith citing papers

Observation d64ebffc-673f-48f0-b3af-27509d2fe77b · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:04:26.381581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:b32c1981671eba1afb2069efc94ff341e0b7958ff72033f8db5c2c021633d030

Observation c34fc912-f8fd-41a6-af73-f7b58750e817 · inbound

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark cites this paper.

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:26.381581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T03:50:18.706396Z digest=sha256:3f71ebb74a7c7e0e5cdaeacf197feb465a3df7bedf58fe8c3fe00f8f5d62d8af

Observation 3cafd5e3-bee0-4479-8b13-815daa6d21a4 · inbound

$\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models cites this paper.

$\mu$VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T08:17:45.666875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T10:50:55.866910Z digest=sha256:f494ff649a8551d5f41653f8aeb949cf3944f747a0a40b62c04e4f6f4a2dbf13

Observation e51edb37-ed75-4a39-936a-e3de65548c87 · inbound

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models cites this paper.

UniFS: Unified Fast-to-Slow Hierarchical Architecture for Vision-Language-Action Models Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:19:47.640060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T08:57:01.861091Z digest=sha256:416a796ceb700df9628ad274f9e9f0705f1c19c9be98b75d80ed9795250bf19e

Observation e15d5c37-b752-4ab5-9b42-521937032fc4 · inbound

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model cites this paper.

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:28.231732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:28.231732Z digest=sha256:a83908c87f77846be34bfc165011d8f6af6ed1cd19be4a77eecd853d4a8c2002