Pith. sign in

Paper Citation Record · LEDGER

Vision-Language Foundation Models as Effective Robot Imitators

As of 11 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 92 inbound Pith citation observations for arXiv:2311.01378.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.01378 v3

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T21:44:27.562453Z

measured 117 of 117 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 92 of 92 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:22:28.270135Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T11:41:02.629883Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact14
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cbe03ef-32ab-43bd-960c-956b0c11ab89 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Vision-Language Foundation Models as Effective Robot Imitators Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.586835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:1f2f491d904051302cc4da7099047687c2d30e360e0146557c614339ff631dce

Observation ae782965-892a-4f8c-8340-0a8814ea5859 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Vision-Language Foundation Models as Effective Robot Imitators OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.594224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:78dfc14fc4a8f37a109338202c33a14f2abb1133e12070b2c79e5195b722ab89

Observation b76ed509-3bc4-4db6-8bce-c1d6b97074d4 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

Vision-Language Foundation Models as Effective Robot Imitators GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:44:27.599526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:9a7ab3b22d186772e3a1d9e5d312a4e15f3cbbd7a088ba5ea6d290f890eca654

Observation 21e3ab03-a12d-4689-9be7-9794c12146e1 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Vision-Language Foundation Models as Effective Robot Imitators RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.605460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:7a60380b31552fc0b0be27846a6cd25a6ebfe6c444772befab67d00596337f7b

Observation 16fb9d8a-ec60-4dba-9f3f-d33c6e9f8f4e · outbound

This paper cites Language models are few-shot learners.

Vision-Language Foundation Models as Effective Robot Imitators Language models are few-shot learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:44:27.684676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:821a8875148b2d3a43c672b7700f447e03f95410d9e67c441d262dd8f78ab3ed

Observation cc775273-3999-4bae-a26f-35f8c8eb9903 · outbound

This paper cites Universal Sentence Encoder.

Vision-Language Foundation Models as Effective Robot Imitators Universal Sentence Encoder

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.640258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:fe8d9c9bf3769f4883a7f72374f1928e3e20ac6630e2f80beffea839b6bed102

Observation 013ea9a7-7e86-44fc-b6ae-ba413f602e8e · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Vision-Language Foundation Models as Effective Robot Imitators PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:36:10.299419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:552800646890aa3ee0e34ce0aa0bd9d552ef83eafb0503382841d6bae5333342

Observation c4b870d8-c583-4d48-959f-e5c60d8f7618 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Vision-Language Foundation Models as Effective Robot Imitators PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.650262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:319fe7223fc9349b9686a07b786f25f093b1581125105eebbf0515de7ed74cf1

Observation 0fde5446-ce46-40d3-8080-6026e75ffaf5 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Vision-Language Foundation Models as Effective Robot Imitators PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:44:27.655058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:fca0315fa0a56b6ef283f0976624db1243c9951adcd4eee1855d99214ca8e7b9

Observation a2ed258a-4c73-4e87-a83d-d994d9cfca3e · outbound

This paper cites Language-Driven Representation Learning for Robotics.

Vision-Language Foundation Models as Effective Robot Imitators Language-Driven Representation Learning for Robotics

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:44:27.661155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:51a0991157efc464bbffcae975f1ff2b64d0b6c0be7d5114475571eec54bbf16

Observation 149dc577-de54-4a50-a79e-d9dd3eb604de · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

Vision-Language Foundation Models as Effective Robot Imitators M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.666644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:7abc965b6bdac7b5752b8b5d02ea583cb513485c4c332ecc3a71eb85b7aa32d6

Observation 1f06a7a9-cea8-49fb-ad6c-4f29bc989509 · outbound

This paper cites Robotic indoor scene captioning from streaming video.

Vision-Language Foundation Models as Effective Robot Imitators Robotic indoor scene captioning from streaming video

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:44:27.690836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:afad03dab71cb654b597f7b993126cdbe2f9e42c5e5099e38ca75190d2528016

Observation e5fa4bbc-eb12-4e2d-9344-5bebd5c1a737 · outbound

This paper cites Energy-Based Imitation Learning.

Vision-Language Foundation Models as Effective Robot Imitators Energy-Based Imitation Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.671932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:483ada17b9b3eef03f37577295e96ca0ca865c4995ace19bb70a4234eaf4f24b

Observation b9f768c3-f6c6-4e19-8ced-8946163901b7 · outbound

This paper cites Goal-Conditioned Reinforcement Learning: Problems and Solutions.

Vision-Language Foundation Models as Effective Robot Imitators Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.676683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:459f9d0aa0fed8814f9d7e553c84f0e27f78a5ac3646334ff785da54e457418a

Observation 7323d1d8-b846-4324-822a-759850c23c31 · outbound

This paper cites What matters in language conditioned robotic imitation learning over unstructured data.

Vision-Language Foundation Models as Effective Robot Imitators What matters in language conditioned robotic imitation learning over unstructured data

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:44:27.699498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:6f72a8f6d4c81d724547202e211be8fd0169235516373740c950ab1ec663eeda

Observation 68418dd1-9afa-4db1-a4de-4e00b7713817 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

Vision-Language Foundation Models as Effective Robot Imitators R3M: A Universal Visual Representation for Robot Manipulation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.681200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:bb6162543d76571471a43a6a4a80954658accb543be9c668627d538fa1ca7e89

Observation df1f8824-7038-47c7-af60-9fcdbb989026 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Vision-Language Foundation Models as Effective Robot Imitators Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.610048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:021828afd6b85ed187d3880434c14b363a42f1b10b42e5af3fe39dc47b22ab5d

Observation 612f1bec-e0ef-481d-b834-23a4ac1ca143 · outbound

This paper cites Instruction Tuning with GPT-4.

Vision-Language Foundation Models as Effective Robot Imitators Instruction Tuning with GPT-4

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.614501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:8c7b3a9d45e7147cec13a94b672d60d23bd8c973242349bb26d49e1beb7d2736

Observation f3b40001-ab12-4ff5-b376-dde687899770 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Vision-Language Foundation Models as Effective Robot Imitators Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T21:44:27.619248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:ed30691601c7e62d26040acb6e56b041729b736c8503e5e56e9bb890089178f4

Observation 0f5d8f62-1dab-4419-8d7b-df71be54d9e7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vision-Language Foundation Models as Effective Robot Imitators LLaMA: Open and Efficient Foundation Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:44:27.623676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:80b462de562626ec98563ef60416c1ec13b375bad3c8e435eb423679a0abd3a6

Observation 7ba2036a-28ce-4fb5-9b14-2f9c48d94575 · outbound

This paper cites EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning.

Vision-Language Foundation Models as Effective Robot Imitators EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:44:27.628753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:f76bb8644debd32d5e2dca4ae79526380154f1ebfa0f63942d79003ed13593d9

Observation 78aefb08-78d4-476d-a35f-ae965cd083a0 · outbound

This paper cites X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks.

Vision-Language Foundation Models as Effective Robot Imitators X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.636185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:e8fbe62a523d4299894ad034d534538389f1ce435e09e8f5996698bce62edd50

Observation 78db5617-2162-4a2f-a754-1bfd92eb2af7 · outbound

This paper cites Deep imitation learning for complex manipulation tasks from virtual reality teleoperation.

Vision-Language Foundation Models as Effective Robot Imitators Deep imitation learning for complex manipulation tasks from virtual reality teleoperation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:44:27.693786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:20f46eed7e431481803b1dfc1b47afd2aba5d3c0663352f5d35593f7b40d0f9d

Observation 73934fbc-1691-4f82-9cbc-9d7fc1677cc1 · outbound

This paper cites We loaded the pre-train 14 Preprint Table 5: Comparison of co-trained models and fine-tuned models on the CALVIN benchmark.

Vision-Language Foundation Models as Effective Robot Imitators We loaded the pre-train 14 Preprint Table 5: Comparison of co-trained models and fine-tuned models on the CALVIN benchmark

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:44:27.696647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:dc6e2f282bfd47d1efbdf50428ab1b1e50aefcb4e62087e1e290990ddf5684dc

Observation 2861e5a5-e962-4179-9624-2a09f20465a4 · outbound

This paper cites As for V oltron, we also include a version that fine- tunes the representation layers.

Vision-Language Foundation Models as Effective Robot Imitators As for V oltron, we also include a version that fine- tunes the representation layers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T21:44:27.687783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:44:27.562453Z digest=sha256:5d7aebe2d9ba06665fed1e72c3ac2adbaadf5788925142a2a3367fa755fff621

Pith citing papers

Observation cfc1417d-b307-4bef-8db1-9221044fe4ca · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction Vision-Language Foundation Models as Effective Robot Imitators

Reference 114

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T14:25:59.328882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:009de8969ecc7c2d7c498c0ccbb1469a960582a03615d2e5620d0943bd3da2c0

Observation 1050991e-bc6c-437c-8c5a-6284f60016c6 · inbound

3D Diffuser Actor: Policy Diffusion with 3D Scene Representations cites this paper.

3D Diffuser Actor: Policy Diffusion with 3D Scene Representations Vision-Language Foundation Models as Effective Robot Imitators

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:00:04.959480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T22:00:04.812281Z digest=sha256:3a23fab92cc71f6b639f27a073962890a324e8fb7d104a44bfda277069660a60

Observation b23c365e-ba9c-43f8-a8d5-e0e29b49fb2d · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation Vision-Language Foundation Models as Effective Robot Imitators

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:55:20.466880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:501b54ba32afcda283efbfc4819f85ef3f20c989fed9360bb1e5b170cffe884c

Observation 8f2515bb-e095-434d-b654-653a3a14d2a5 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI Vision-Language Foundation Models as Effective Robot Imitators

Reference 104

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:25:54.448890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:d08fcbf3236bd223260081a93f981effecf4a4050b351273b6d1b8959c8335d8

Observation 8b8992ab-39f2-4bd9-9008-25defdb0c618 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:8d2a6d081964a0a085300c979acc8dc598320cef64c602e8b314ba3fc67f60f5

Observation 826d4c58-8d89-4f86-a768-72cedeec22af · inbound

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation cites this paper.

GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:09:33.761708Z digest=sha256:2b4e94ca4afad118bc2c2bd86a9dcbb50b4904bc86276c7a6789543674a7134c

Observation 0350bb56-cf30-4593-a6f5-ce0ff128479c · inbound

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation cites this paper.

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T07:33:25.188358Z digest=sha256:8c960f406796abbab5dd9bf30759210fa669b0ae7e56eb5ddb60b77c4571aea9

Observation 7486c90b-204f-4f2b-8b9e-e26d128809bb · inbound

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction cites this paper.

CARP: Visuomotor Policy Learning via Coarse-to-Fine Autoregressive Prediction Vision-Language Foundation Models as Effective Robot Imitators

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:28.270135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:28.270135Z digest=sha256:8b1a2a78c3b73f819de8518043c12fa5bcbe1411d8e67eb2a48127d34de82c97

Observation 14ecc731-396f-4840-86c5-745e17fe6df1 · inbound

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons cites this paper.

From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons Vision-Language Foundation Models as Effective Robot Imitators

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T17:55:13.405881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:55:13.405881Z digest=sha256:73a74937a5c72ff8d5b0633d474a06c51adda953fcbff53481c90bf6d34b3611

Observation 3a307198-f89e-4159-b195-7ec5d128335b · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots Vision-Language Foundation Models as Effective Robot Imitators

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.726848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:bc0e8666e3e6f599fb4aa0b0338fe2c7e951d561cd488a92c1952cb705e846f7

Observation 97f2b47b-c758-4e7a-8a9f-0696a40473f4 · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Vision-Language Foundation Models as Effective Robot Imitators

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:3b5da7333a72834fa87655d75651e15600ff0b02d259b7bc9afb90fd884517a4

Observation b620cb81-1c65-4201-ac9c-6e1e5feead9c · inbound

QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning cites this paper.

QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning Vision-Language Foundation Models as Effective Robot Imitators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:22:41.393711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:22:41.393711Z digest=sha256:ee388466b88852d391b3d9f90fc28ae45373f68dc384580f6a9ccae4871b5fe5

Observation a64c9f9a-58c5-41cf-967c-ebc2d33f5a63 · inbound

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance cites this paper.

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance Vision-Language Foundation Models as Effective Robot Imitators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:37.398060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:37.398060Z digest=sha256:885e24b1c31442a598187156018470c75b6de0789c0031df744c165b3480eecd

Observation 9a1560bc-e52e-4e33-bd79-dae0622a990d · inbound

OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints cites this paper.

OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints Vision-Language Foundation Models as Effective Robot Imitators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:50:41.217206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:50:41.217206Z digest=sha256:5e9de54c90bc3e7b3634ce05a7f83c6f4d0a7b4aceaa39c05073cb2e9945ffbd

Observation 2bba8958-450f-46b2-93ec-4f1f13e1f4df · inbound

Visual Language Models as Operator Agents in the Space Domain cites this paper.

Visual Language Models as Operator Agents in the Space Domain Vision-Language Foundation Models as Effective Robot Imitators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:43.834162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:43.834162Z digest=sha256:0852fefa11d0b505d6690b6f40651a459c6c1095594191fcfe9367d86c684a42

Observation 134239e5-fb44-436c-a382-b566955b0ca3 · inbound

RoboReflect: A Robotic Reflective Reasoning Framework for Grasping Ambiguous-Condition Objects cites this paper.

RoboReflect: A Robotic Reflective Reasoning Framework for Grasping Ambiguous-Condition Objects Vision-Language Foundation Models as Effective Robot Imitators

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:11:32.505646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:11:32.505646Z digest=sha256:11e9d228017bc90c92087a114bd058949864d71574c784fedd94282ea0b42376

Observation f5e07be3-73be-4491-8298-d82fb1aca33c · inbound

GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation cites this paper.

GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:47:36.507541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:47:36.507541Z digest=sha256:a80adfc0771ef159ee318aad66a0817f386278ee0ab61e95517e6e272f2e1057

Observation ac41728e-41a9-47c1-bddc-354ccd085d3e · inbound

Improving Vision-Language-Action Model with Online Reinforcement Learning cites this paper.

Improving Vision-Language-Action Model with Online Reinforcement Learning Vision-Language Foundation Models as Effective Robot Imitators

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T11:39:56.058905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:39:56.058905Z digest=sha256:727811e5ae92e611c9808a6c4577c3c8b0ebe84a06244b93e141c3259ec3639e

Observation 55ed24c6-9829-42df-af65-7ad72d392227 · inbound

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent cites this paper.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent Vision-Language Foundation Models as Effective Robot Imitators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.916949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.916949Z digest=sha256:814d3deed46d13f7bb55de09d8649e190ff0b1b6cc3acbe7cbf47275aa4b66a3

Observation a5a5dec7-0536-402e-b65b-24890623dc09 · inbound

RoboBERT: An End-to-end Multimodal Robotic Manipulation Model cites this paper.

RoboBERT: An End-to-end Multimodal Robotic Manipulation Model Vision-Language Foundation Models as Effective Robot Imitators

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:17.882033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:35:17.882033Z digest=sha256:187e401e917f7e868f31e64329aef99d0f94f4abc05d81931c71f11d76055f84

Observation f51fffb1-90e2-4b88-aa30-54da0b47c14d · inbound

Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation cites this paper.

Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T00:05:05.719831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:05:05.719831Z digest=sha256:8abb93f6e3287ef9def99db02e6d7ea6251f27fd8cc768138da0a153176f177c

Observation 5dac971e-7165-498f-bab5-2a201487cda7 · inbound

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning cites this paper.

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning Vision-Language Foundation Models as Effective Robot Imitators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:24.768834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:21:24.768834Z digest=sha256:95bf076289dc757c2ff8d7f0a93cf18f49aee3da474adf0494dffd6ea1712ec8

Observation 70c4fcec-1f1e-4baa-af95-4b606682321e · inbound

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation cites this paper.

GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:13.376171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:13.376171Z digest=sha256:3d337f373a38f73f78292170292d1e3a57ba4dd5d310d87309d3a18b8bc7eeb9

Observation ffd48dfc-9335-4f6a-9b3b-2c2d7b535ee2 · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success Vision-Language Foundation Models as Effective Robot Imitators

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:4133c54219872ef132cc1d0e7c3c274b8f0588653c6a2ab9bdb6a13940c347e3

Observation 80991980-ebd8-45ff-9f24-8ebd1953dba0 · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:2bbc01bf283d948848053eb0ff2bda6cd74d4f9db467e3dc27e5db40fa2a9b05

Observation fd9bba19-1692-4008-bfa8-67a6719c60d6 · inbound

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots cites this paper.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Vision-Language Foundation Models as Effective Robot Imitators

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:08327175b5111874f8c7840b7b1ab80778210efe0352f3d06353453b13859b4c

Observation 00183152-6ce8-4b2f-b379-b4eff41b1f29 · inbound

VLAs are Confined yet Capable of Generalizing to Novel Instructions cites this paper.

VLAs are Confined yet Capable of Generalizing to Novel Instructions Vision-Language Foundation Models as Effective Robot Imitators

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T16:46:47.468917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T16:46:05.993833Z digest=sha256:ca5bed02ab037f47d8379be718ce1b2c2378a9ac48c03c583a06729fd551b930

Observation c52d2050-4d24-444c-86b3-ef4f2f636ada · inbound

FLARE: Robot Learning with Implicit World Modeling cites this paper.

FLARE: Robot Learning with Implicit World Modeling Vision-Language Foundation Models as Effective Robot Imitators

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:59:09.007087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T15:59:08.846629Z digest=sha256:63159fd7d19ec9e65bf7df7ba3c9ae3a7beb6bf92d0b012dbb918e787246e07d

Observation f249808f-6129-4452-aa95-f196ffe61df3 · inbound

ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models cites this paper.

ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:23.808381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:23.808381Z digest=sha256:e4a3b35877f35b03c257b2295550616c77cf35ee7dc3f4ebcdd6ef3a21814310

Observation 1b16272f-66ac-4e50-a783-e906c8945fd4 · inbound

BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization cites this paper.

BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled Optimization Vision-Language Foundation Models as Effective Robot Imitators

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:59.182171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:59.182171Z digest=sha256:ec962fea84ccdaa677cc74afaaa325460325e2d1a85d1ec1dfbe6e2f7fcb0347

Observation a517b294-5f40-42d8-90fe-a0e1c789012c · inbound

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models cites this paper.

Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:01.760587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:43:01.760587Z digest=sha256:7ce8fb90d33783680e9c367a0d23985a945103670205efe2ac84abd1351e6180

Observation 7b8e407a-0420-49b2-a79f-b764eee9cfff · inbound

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge cites this paper.

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge Vision-Language Foundation Models as Effective Robot Imitators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:25:58.804742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:25:58.804742Z digest=sha256:a657ffedcc8fd1580528b2cf7ec714b6eabc9f8f6d0737faeee782193abb4607

Observation c774c3c4-96de-4b34-b385-85d1e4b2264a · inbound

Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning cites this paper.

Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning Vision-Language Foundation Models as Effective Robot Imitators

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:17.665412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:34:17.665412Z digest=sha256:0ae0a35306e3800d02e6c9f7f7df401c2c3c8ae15d31bcdeddb8bfd4db8dcf24

Observation 95a71858-8c5c-4c83-9ddb-8808717a9a11 · inbound

SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models cites this paper.

SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:38.999918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:38.999918Z digest=sha256:9d55528cc75007cce48edeabac8e3fdd7c54ab59e980147e0bc3db8042aefcea

Observation 8ca80984-43d2-4b6f-9dd9-e4f89fb2afb8 · inbound

CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation cites this paper.

CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:20.313536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:20.313536Z digest=sha256:05832f41608443ca6dfd4ab3513d77da96259dca327968aeaaa2e6abdbbb7e4c

Observation b6f21b8b-0c64-4036-950b-5921283d297f · inbound

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models cites this paper.

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:22.448483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:22.448483Z digest=sha256:f0d84e75ac1369809b713862bd4e74c5da29a6ef524a591d6dfc1fa35ab4eaf0

Observation ebc8f8a7-6d31-4b44-b8b3-bbad4f0e1b1b · inbound

GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation cites this paper.

GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:09.575868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:09.575868Z digest=sha256:7d7b1186d6dfc164fcd518e18e2eb3d292d4bc5e962f7be72d904e1638ee055b

Observation b8d93450-37ea-40c6-8ed9-3283ec429749 · inbound

Dynamic Double Space Tower cites this paper.

Dynamic Double Space Tower Vision-Language Foundation Models as Effective Robot Imitators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:12.492546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:12.492546Z digest=sha256:2ebc2d98ed1c1bbf64c42f29ff799ec9d0346219cca48158e1469b3fc7bc4612

Observation 7b4db716-9f2c-4606-9d99-6282f0170e54 · inbound

ROSA: Harnessing Robot States for Vision-Language and Action Alignment cites this paper.

ROSA: Harnessing Robot States for Vision-Language and Action Alignment Vision-Language Foundation Models as Effective Robot Imitators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:31.758456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:31.758456Z digest=sha256:935fcf81865be90783b46ee67737eb524d7ff75d39f367cf5540391ac16a71a0

Observation 0019b4a4-5b4e-4be1-a351-23cdb4da1fd4 · inbound

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models cites this paper.

T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:05.276267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:05.276267Z digest=sha256:91169afd4d9b02146d3b0d45e22e3a9988ba09ebdb6f97f6a59940eb6527c9fa

Observation 16e31991-7784-4d6d-aad9-0f1592599885 · inbound

WorldVLA: Towards Autoregressive Action World Model cites this paper.

WorldVLA: Towards Autoregressive Action World Model Vision-Language Foundation Models as Effective Robot Imitators

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T22:57:07.883617Z digest=sha256:2841a236ac19d67991ee88fd07c55f1df614d379e74576436a41a7d4ee5604d7

Observation ccb96fcb-72d1-4808-86e8-e17c38e1c6c1 · inbound

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation cites this paper.

AC-DiT: Adaptive Coordination Diffusion Transformer for Mobile Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:30.058020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:30.058020Z digest=sha256:26add0d83187da5e4a8ab610ac726949a042440490e21d8aa7897b12c76149b7

Observation 94a21d8c-9bb6-4cc0-b235-b45d901d2b1d · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report Vision-Language Foundation Models as Effective Robot Imitators

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.691874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:261367a202a549db7717f6c5bfc6852f0df7d70e5525239944ebe67b658aa85d

Observation bb2bda5c-3247-4b2b-a1a6-d06443aac02e · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning Vision-Language Foundation Models as Effective Robot Imitators

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:45.149931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:45.149931Z digest=sha256:627709b85f0123d8f04a20dda0020c599ea60bc8025ffa9a27ca634449117110

Observation 35f40323-ed57-40d4-b829-b4e32bb89201 · inbound

Two-flow Feedback Multi-scale Progressive Generative Adversarial Network cites this paper.

Two-flow Feedback Multi-scale Progressive Generative Adversarial Network Vision-Language Foundation Models as Effective Robot Imitators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T17:34:00.544575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:34:00.544575Z digest=sha256:e1ef473ea772d783444b73f5930f2ded0e9c5260d271e360f089de3ea5c4a4df

Observation 77e40e6f-bbe4-4281-b24c-86264bed0b64 · inbound

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training cites this paper.

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training Vision-Language Foundation Models as Effective Robot Imitators

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T12:15:01.429344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:15:01.429344Z digest=sha256:0c38ea8839cb22aa79c615e35a525b77019656c089f6cb2d0545bf5bdac97f59

Observation dce0a682-6ed5-4854-914e-3d1a0ef91947 · inbound

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies cites this paper.

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies Vision-Language Foundation Models as Effective Robot Imitators

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T05:48:48.088373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:48:48.088373Z digest=sha256:9acd31b9089130d2386381b847da7642e873f12127bb46830f753c474b62ae90

Observation 1a5b98ee-3931-467c-b4b6-6d209502f48d · inbound

LLaDA-VLA: Vision Language Diffusion Action Models cites this paper.

LLaDA-VLA: Vision Language Diffusion Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.929977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.929977Z digest=sha256:47fd7257aa6cbcd882ef09dd8fcdc288a1748c3c7e57d2090b0ca5f12353e9b1

Observation b022643b-7b10-4474-b2e0-7bb1e4370340 · inbound

TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models cites this paper.

TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T21:29:09.018713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:29:09.018713Z digest=sha256:a306445fa3808a796c12f2afc2b5117d0e9620a81f27de5b4dec14a171f28e3a

Observation 6b2874ea-6c8b-49ce-8efa-7104d79a790e · inbound

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model cites this paper.

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T07:00:46.400140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:00:46.400140Z digest=sha256:a14daa3e7742e2d03090d0235e42dd9e837e92b4a4e37922a3bf00edc5d7890c

Observation 2672b566-8f58-4b39-a5f8-0b221330dcda · inbound

Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception cites this paper.

Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception Vision-Language Foundation Models as Effective Robot Imitators

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:05:15.679521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T21:03:28.341042Z digest=sha256:053a9bc39599e19cf906230c0cd7f742a92d0a3ef3284fd43809ce0cdd735fbc

Observation 7816679a-32e7-4762-ba51-ee9b28fac194 · inbound

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference cites this paper.

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference Vision-Language Foundation Models as Effective Robot Imitators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T21:13:57.062334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:13:57.062334Z digest=sha256:7853344232130b65b46b3bc96e17c113fb328b4fc74b5e33ceca655f04037f73

Observation b49db4d9-ec83-487b-8134-59b8054e7b2f · inbound

Large Video Planner Enables Generalizable Robot Control cites this paper.

Large Video Planner Enables Generalizable Robot Control Vision-Language Foundation Models as Effective Robot Imitators

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:26:32.048309Z digest=sha256:eea6804203ff5c2eced10bbe767ab673f374edb0a53c41ff31c9621997ddf9e5

Observation 3cd072e4-9009-4ed1-9522-9bc9dcc93d22 · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.815895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.815895Z digest=sha256:389ed0d83ef1c619d5d5e9cb319ffe70290f902abe3a5e25a5923f73e1234cee

Observation 109a6f98-431f-40d0-be5d-ec143c757832 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:3665850372d5f4dad5bbd72bac61e61e3df95c5ad7690303a2c90e913cc0f96c

Observation 2fca8ba7-75e2-492c-90b4-926308054eb2 · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Vision-Language Foundation Models as Effective Robot Imitators

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.024414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.024414Z digest=sha256:7f80675b23dc5df4f03f790abc9400b2769deddc2956c4e1dea4f5dbeeaf1376

Observation 0156e3b7-b707-4466-aee0-8f412012e693 · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs Vision-Language Foundation Models as Effective Robot Imitators

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:27.691000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:27.691000Z digest=sha256:5fab64f121fa4600abf4a4b53abfcfebc7347963b85811ee4510445de9f54e0e

Observation 3248ea7c-a93b-453a-b155-b54935d2e0af · inbound

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation cites this paper.

Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:24:11.195439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T13:22:16.242427Z digest=sha256:a245ad71cb507872288717e7fc7540e08ad19ba198411a23c269c30011fac6e2

Observation 1b5d37b1-f4c6-48f6-beab-1444b650778f · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:7287f0aa1d1eb7a98d9c194d43d5aed375d62a0bb36dc9c36e2ef52e1e15e9eb

Observation acf4a9f8-ca4f-4cc7-a719-92cd552dba1c · inbound

Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks cites this paper.

Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks Vision-Language Foundation Models as Effective Robot Imitators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:12:11.244767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:12:11.244767Z digest=sha256:21ca316a26bf71c631a34b57dfede58898cc203a1d43b2807bcaf59164231871

Observation c9976df8-3e13-467b-b346-626101002b82 · inbound

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment cites this paper.

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment Vision-Language Foundation Models as Effective Robot Imitators

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:27:12.286456Z digest=sha256:9206ca011fdf6ec7b57fd3f3bc097316a88cda6de7bd41423dd2393e39006ef8

Observation f223ed7b-05f6-4852-b60d-e27e6b6246b4 · inbound

R3D: Revisiting 3D Policy Learning cites this paper.

R3D: Revisiting 3D Policy Learning Vision-Language Foundation Models as Effective Robot Imitators

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T10:57:54.273960Z digest=sha256:6dcb8727dda7c657b476ec799489d49683866ad0a80f86650bf11f1b8bb9e524

Observation 2457dd32-3bd9-4ea0-b750-6b37fd5912d1 · inbound

Steadily moving semi-infinite fracture in plane poroelasticity cites this paper.

Steadily moving semi-infinite fracture in plane poroelasticity Vision-Language Foundation Models as Effective Robot Imitators

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T11:41:02.631612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-05T11:39:05.686584Z digest=sha256:f0bf0f6c54c06838f689383bba391c20857ed3a6295591fcbe58b285cb3d1af2

Observation d989c838-42b4-4de4-9e6c-bce3fb67cb03 · inbound

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments cites this paper.

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments Vision-Language Foundation Models as Effective Robot Imitators

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T05:46:36.865150Z digest=sha256:d4200b55d4ee2393168fd2f6a85d78d396192e57f8fe47a712d0ede20db0ee17

Observation 3cae46c5-e933-4d17-ade1-e02b383aab9c · inbound

Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models cites this paper.

Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-09T23:59:32.295839Z digest=sha256:2500c232d3c5064e3d70319fae472e173a8bfecc8a0e5607bb48929d1042ebc2

Observation ed14748f-ba53-4b04-8253-407f2af5c6fd · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy Vision-Language Foundation Models as Effective Robot Imitators

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T03:39:41.090350Z digest=sha256:64a95487674f621b68c7c3a46f5c4194a336775599ff2404d8964e7b6f0e57ce

Observation eda8b2af-9ba1-4373-9ccc-ee5dc093d9a6 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy Vision-Language Foundation Models as Effective Robot Imitators

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T02:53:54.608425Z digest=sha256:6387579d68babceba25c9f113bf8b61df5f7c1f3232289333f79687e21c28811

Observation 99a02836-19fc-4e0d-8fdf-06d50d8168b6 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy Vision-Language Foundation Models as Effective Robot Imitators

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:16:48.180290Z digest=sha256:a3741ef5ba5572f75cc0915ed622162f47662d129a7f8fc7e07f8472b9157339

Observation 63d57f8c-ef74-40da-b75e-15824cdc024f · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Vision-Language Foundation Models as Effective Robot Imitators

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:64df29bb8e45071d567e79ffc85f594cf09c620cf12d2de3df4284768e52f2f6

Observation c4c0af91-7269-4c83-b653-1933350d8d1e · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Vision-Language Foundation Models as Effective Robot Imitators

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:52.451085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:52.451085Z digest=sha256:068c5cc0751dd3cf23d87e8674eb7ce6d81de0db9d56cdd162c33eafaa4883da

Observation 0f162522-895d-4f9f-9e24-27ed585a84e7 · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI Vision-Language Foundation Models as Effective Robot Imitators

Reference 266

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:44:27.700418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:21607a7488c1ea1b73ca39e95ec1d34f7ca89ea945b856f16ca0b849d2ce8b8b

Observation 7c181494-535f-451a-9f0a-6c9783ba3979 · inbound

Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation cites this paper.

Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation Vision-Language Foundation Models as Effective Robot Imitators

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T18:43:38.654786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T18:41:12.533556Z digest=sha256:48083fcec01faf66cb3a2b5e08dce3c215b5eccda8136ca85883380f21069087

Observation 89061257-a158-4bfc-b6c6-86b99145f0db · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Vision-Language Foundation Models as Effective Robot Imitators

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:43:16.959685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:5d190da42c397502a76bfce5af2fda06e98678b70e5aab541b8955238fe4729a

Observation 514e7fa4-60ff-4164-b8a9-058e32ba8694 · inbound

DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation cites this paper.

DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation Vision-Language Foundation Models as Effective Robot Imitators

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:59:36.280986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T04:58:46.675822Z digest=sha256:efbf5082ae75717107d42782b3ea85e1e099228842f2e0fde9287e1cb9607776

Observation 976c9f63-34ef-4bcd-bdcd-c1bb6d6532a9 · inbound

ProgVLA: Progress-Aware Robot Manipulation Skill Learning cites this paper.

ProgVLA: Progress-Aware Robot Manipulation Skill Learning Vision-Language Foundation Models as Effective Robot Imitators

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T11:53:23.407726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T11:51:42.029338Z digest=sha256:dc911bbd3ab9d4b424a50bd02d158a9ed3e15f986a0e5f5951738607ac5e98dc

Observation 32c9a440-5ccd-4b66-8b71-731a5e88279d · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence Vision-Language Foundation Models as Effective Robot Imitators

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.884284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:6b92374b7d2b91668c4de15fe10c0e0209725db4ca0c2233c133673f3e033cb0

Observation 23cdc0b3-ffa8-4502-b64e-1d0fdd5b567b · inbound

PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation cites this paper.

PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T11:53:24.289208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T11:46:14.257386Z digest=sha256:4b46b0664cd3c85ea519d526c7cd771fdb355f3acb4941eca4cfd270111cfe75

Observation 69a65c01-7b88-4177-8bbb-e1f733ce53ad · inbound

PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation cites this paper.

PrimitiveVLA: Learning Reusable Motion Primitives for Efficient and Generalizable Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T12:57:54.300162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:57:54.300162Z digest=sha256:4e8f00b8620e0bbd8ed29e5a9a38c1350632cd5a087bc72abf61ffe087e74064

Observation 41f5cdd1-553a-4629-8013-e5488a2c66e4 · inbound

Wall-OSS-0.5 Technical Report cites this paper.

Wall-OSS-0.5 Technical Report Vision-Language Foundation Models as Effective Robot Imitators

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:26:00.546630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T22:35:00.258436Z digest=sha256:ad35719672977f71a9e27825d12075f850715bee1a442205cb8b718ce3dabc08

Observation b3cd827f-3772-4f8b-b639-290db411c575 · inbound

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling cites this paper.

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling Vision-Language Foundation Models as Effective Robot Imitators

Reference 102

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:33:27.797624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-29T13:33:03.368006Z digest=sha256:18ef29f41eab97fed8fbdf36b5f2d80e918b02b20a41d873e248a6633faafcd0

Observation 1b5bfa74-4664-4ff1-beb6-78e8eea08836 · inbound

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding cites this paper.

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Vision-Language Foundation Models as Effective Robot Imitators

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:16:59.183095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T01:23:02.576098Z digest=sha256:cf737e1d4235ff010a329337671ebf4623b6210bea1cefa59912a69655e74326

Observation 0f594631-394e-41eb-9553-f68c7f431a11 · inbound

Remember what you did?: Learning Behavioral Memories for Partially Observable Object Manipulation cites this paper.

Remember what you did?: Learning Behavioral Memories for Partially Observable Object Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:39:37.140664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T14:21:46.672591Z digest=sha256:2c87f4991d522d3a5ea23834f8eaa04493adc2384b2fe94478339a5174833cbb

Observation b83e0735-4bf9-4e9c-923d-8f30147b996d · inbound

KEMO: Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies cites this paper.

KEMO: Event-Driven Keyframe Memory for Long-Horizon Robot Manipulation with VLA Policies Vision-Language Foundation Models as Effective Robot Imitators

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:59:45.680478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T08:22:29.894757Z digest=sha256:be9a12e6870e51370aa45fc9142ac54b1bb28b175ac8ae79e02e71e1f874091a

Observation b32f87a6-b19b-4c5c-9c5f-99324bec936d · inbound

VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon cites this paper.

VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon Vision-Language Foundation Models as Effective Robot Imitators

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T12:28:07.296016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-03T12:20:11.540651Z digest=sha256:b4e06f85e803fcb4b277f4dd8e0fddb726136aacd33164485d3c80a17e30a392

Observation da9a171c-ec69-422b-b275-2573ef587e77 · inbound

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging cites this paper.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Vision-Language Foundation Models as Effective Robot Imitators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:de15bcb0713a45888f102c0836dfb043f344cf1ddd72248a7bbc4eb54052a184

Observation 178ff601-9fc6-4c5b-8cfa-38572c73e082 · inbound

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation cites this paper.

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T05:40:47.306935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:40:47.306935Z digest=sha256:4db15f9f97293dbadbfbd9d8af5625ced0ac32adc4e48aeb052336f31068e14d

Observation 54fea40b-07c2-4328-bae1-71a7d0e0786e · inbound

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment cites this paper.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Vision-Language Foundation Models as Effective Robot Imitators

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:16:38.875488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:16:38.875488Z digest=sha256:e739567814104cadccb6db053a66bfed3372565fb3065f79c71c6c2a1bc28519

Observation 0e991754-2032-4230-9537-a18f423502a0 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Vision-Language Foundation Models as Effective Robot Imitators

Reference 207

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.334279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.334279Z digest=sha256:032fcc7f093fc500cd3df43b42897c5279a61e8724483bdd300616c1d2db321d

Observation c5fb9ad4-a1da-4803-aea9-f43f45530565 · inbound

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference cites this paper.

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference Vision-Language Foundation Models as Effective Robot Imitators

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T22:36:29.282379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:36:29.282379Z digest=sha256:95726321cabdf042ca965d52de8e00dbd15d2a75d4979e76e5eff8190f5f3e6e

Observation 8da74966-0b46-4aee-a303-caf09a14c868 · inbound

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model cites this paper.

CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model Vision-Language Foundation Models as Effective Robot Imitators

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:25.275883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:25.275883Z digest=sha256:5ec98826e5f9599379c8534f2e28d38fcac4a0aa38e207e571babc9de8038a00

Observation 70d692ef-3573-4fbe-9135-d9fe257aecc2 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Vision-Language Foundation Models as Effective Robot Imitators

Reference 142

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:34.913478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:34.913478Z digest=sha256:8337cbcd570b554547609f7ca0d0dbde87475b521d9ce54dbafd4e112a1a5d52

Observation 901f4227-b934-4ab4-95f0-eac1b048c079 · inbound

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models cites this paper.

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T06:10:32.906456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:10:32.906456Z digest=sha256:dc11c446ca4f12283c5591eb9958a330df80a5f49e6e066479ce4c830a9f26c8