Pith. sign in

Paper Citation Record · LEDGER

What Matters in Building Vision-Language-Action Models for Generalist Robots

As of 6 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 47 inbound Pith citation observations for arXiv:2412.14058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14058 v4

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T21:37:50.617813Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:55:52.186146Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T05:36:01.252344Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact30
  • verified fuzzy20
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f43182c-4218-4708-81e3-a183d7ba0b05 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

What Matters in Building Vision-Language-Action Models for Generalist Robots Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.961534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:f865408eddf43c44d14b5d7aaffb17107975f31042acee60309d790f86ca66b6

Observation 6327640b-6520-4734-9dcb-076e52bd1caa · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

What Matters in Building Vision-Language-Action Models for Generalist Robots Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.668786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:71aee5ceb3dcc7d1f967e269ff6348444be8643da5b6645dbb77ed29a81793da

Observation d2b49e18-aab9-46ff-9f7e-53173b115b85 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

What Matters in Building Vision-Language-Action Models for Generalist Robots PaliGemma: A versatile 3B VLM for transfer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.676778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:7da5ef458e74686a8665faef4906c98d6cf1239dfa6850fc259d4813185aa5df

Observation bcb6b6ea-2692-4ab3-9e0e-7092a78ef483 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

What Matters in Building Vision-Language-Action Models for Generalist Robots $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.683754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:c1b78f158647bd8e2c240ad12a61a92cd45bb788369a6ebbf1b873e1a0fb97d1

Observation 5b4e2902-16db-4ccf-a9b8-c08bcffd4191 · outbound

This paper cites RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation.

What Matters in Building Vision-Language-Action Models for Generalist Robots RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.692026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:98a66321e15bc6d169ea030bc40b68addc5ac120734f280cd043f72b206fa54c

Observation 0205a320-063b-4fa4-a7bf-8f828a9ac790 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

What Matters in Building Vision-Language-Action Models for Generalist Robots RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.700111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:2f12cd7491aa980066c57ce45a3e07b73ff5572cd1f3f3cfcb92f07c1b14ce85

Observation 8f70d6d8-e7a0-4264-823a-6d06bd4319d0 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

What Matters in Building Vision-Language-Action Models for Generalist Robots RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.706165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:a50dbc6f0429a9737389333a73454a025c8e17b6dba3397cb90fbee7552c5066

Observation 9418149c-13e5-4fdc-84d4-6477d619c478 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

What Matters in Building Vision-Language-Action Models for Generalist Robots GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.712601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:30a2ef8e6c84f37c53b856acb33f9ba7cd3ec878d40f67544c0a17984cb55579

Observation 6bede542-7190-4ddd-980e-1361f1aaf294 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

What Matters in Building Vision-Language-Action Models for Generalist Robots Diffusion policy: Visuomotor policy learning via action diffusion

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.867856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:56649f60734f705ac198eb48d8021812bd29cabab857aa55c7f3b7dde2f442c9

Observation 5c1194ee-339c-4ccb-8aac-2e6e19e9d6db · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

What Matters in Building Vision-Language-Action Models for Generalist Robots Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.828876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:08df77f3e17a80dffee46a129fd65970578513b1d64e5405ee5d591a5fdcf3ff

Observation 1be9bf5b-b4c3-48c8-a1f7-7f4725d8ecee · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

What Matters in Building Vision-Language-Action Models for Generalist Robots An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.833067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:0076553bff1c4af53407604d91d311a8716811192532b0a72c167069b5add89a

Observation 2bb8beeb-d451-4af2-9fbb-94ce424e233a · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

What Matters in Building Vision-Language-Action Models for Generalist Robots Model-agnostic meta-learning for fast adaptation of deep networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.903103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:bd7b436d026f89aa794b18b92d9f23e93c13a6d1423fb1e96f5c3c3e44afa1f7

Observation 84705487-e9db-481b-96a7-2f4d9c1ba5d4 · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences.

What Matters in Building Vision-Language-Action Models for Generalist Robots Gpt-3: Its nature, scope, limits, and consequences

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.910214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:83655a18c613e3b5ab96fb854225878ee61bee310e0e0d456ee1ee2277601588

Observation ec2dab5e-509a-47da-9237-afa7bfcc9518 · outbound

This paper cites Long short-term memory.

What Matters in Building Vision-Language-Action Models for Generalist Robots Long short-term memory

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.917436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:7a86256882c1d7f972a4d17dcb19e2607dc930599d5a75ba4208d970c3f9bb1e

Observation 1bff4d0c-4c89-49f2-abc0-fa8959b304d5 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

What Matters in Building Vision-Language-Action Models for Generalist Robots $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.837969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:b5e8bd163c23add5f0fd718349fd0700ba77c6d74b434f84ed3427dcbc0fc62f

Observation e78c6868-9ab3-415c-8280-acc279f32aec · outbound

This paper cites Perceiver: General perception with iterative attention.

What Matters in Building Vision-Language-Action Models for Generalist Robots Perceiver: General perception with iterative attention

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.928095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:c690a732e250a2c419a9379cd0a56aa136888d84da06bbe071738585ee4da31c

Observation b6e94f44-696c-491b-9efe-f8a7f1ee7f22 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

What Matters in Building Vision-Language-Action Models for Generalist Robots Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.934432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:5cfeccc019a33f42d737e4c3a98e279bda16566848b9005ae9e99fc403292ab0

Observation cf67eb99-1617-40ce-a7c2-52496ef7c858 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

What Matters in Building Vision-Language-Action Models for Generalist Robots VIMA: General Robot Manipulation with Multimodal Prompts

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:37:50.842888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:cbaf072e75bc9daebf7e531ef39a0a7a214e8802274e6b639bda536add80ace7

Observation d55420aa-f5ab-40ef-8bfb-b696c5d40ff1 · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

What Matters in Building Vision-Language-Action Models for Generalist Robots 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:00:05.246520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:14c04ac1c6854afbd1804602538a2ede16c0c6db28da4e9caa2c607d6492ff62

Observation 042a928e-46ca-4f70-822b-68c8bad0a779 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

What Matters in Building Vision-Language-Action Models for Generalist Robots OpenVLA: An Open-Source Vision-Language-Action Model

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.861026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:c92567ded2434acd9425aabf4deb15436735c45bbe62413f4dec4c1abd9b8a88

Observation df37c3b8-5d67-4673-8bb0-e408b0af1b93 · outbound

This paper cites GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy.

What Matters in Building Vision-Language-Action Models for Generalist Robots GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:37:50.720922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:de2810a7bfddcad4f1a04f7ec6f654bb4cb80cc0106368ee0cd6f7155ef2efc9

Observation 3a307198-f89e-4159-b195-7ec5d128335b · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

What Matters in Building Vision-Language-Action Models for Generalist Robots Vision-Language Foundation Models as Effective Robot Imitators

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.726848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:382b11864220c8a7954b9c435bddbd921804185faf8188f06ef5be21f257cf8a

Observation 0b081c3b-5c2f-46de-a2e5-e7d32ea62e91 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

What Matters in Building Vision-Language-Action Models for Generalist Robots Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.732573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:09c16f4ef142bd6896113be95815b2eb720f66b89d00ed20127c1eca2420a964

Observation 3ed40d69-2c67-4419-8d4b-30fa59529416 · outbound

This paper cites Flow Matching for Generative Modeling.

What Matters in Building Vision-Language-Action Models for Generalist Robots Flow Matching for Generative Modeling

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.739139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:f5f0b0f5fe03c30c8c381f2086ca2504a0a4eb8a05d413b1723de9c96d2fc5b0

Observation 02893c26-0624-4d53-828b-b8e2cadf64bc · outbound

This paper cites RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation.

What Matters in Building Vision-Language-Action Models for Generalist Robots RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.745251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:ffab749118a6707bb343bba2f50637799976ef923a87de38776e2e7c2bdca99c

Observation 3b7fd351-9604-4a32-b8fd-48b7d726c916 · outbound

This paper cites Visual instruction tuning.

What Matters in Building Vision-Language-Action Models for Generalist Robots Visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.967725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:96b1a05e6bb081a9a67033be2570e760b31abda9dacd3bc86618931903ade7cb

Observation 3886197b-86db-422c-aa03-5c4ff445cc46 · outbound

This paper cites Embodied intelligence: A synergy of morphology, action, perception and learning.

What Matters in Building Vision-Language-Action Models for Generalist Robots Embodied intelligence: A synergy of morphology, action, perception and learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.970968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:af82dffb1e39f95c48ce679e8f642c3df86ce9c28925b39dc2299e9150d00933

Observation 290b5786-1f73-4067-a976-d11cdcb7f124 · outbound

This paper cites RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation.

What Matters in Building Vision-Language-Action Models for Generalist Robots RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:37:50.751231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:f6e5e2a3e64a9074ea70dfd35d5b300fdf8724e2ab8ba6038cc7b37fcd058cc4

Observation 1edfb71d-d626-4eaa-8ad5-5afc19a43340 · outbound

This paper cites Recurrent neural networks.

What Matters in Building Vision-Language-Action Models for Generalist Robots Recurrent neural networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.880492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:70e8b14ba0ddb4ca70c5160ee6887f948d8e40d2b8dde83448d276ab0ca37f7b

Observation aac55734-6788-43db-aa99-6dd726fbc1bb · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.

What Matters in Building Vision-Language-Action Models for Generalist Robots Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.890282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:bd8bc7896c71a8426281374ddc7e81686d565b8f1f15b5f1490dd63d04a0baf8

Observation 877634e0-c398-45ea-bd70-ae617c33e6f8 · outbound

This paper cites Attention bottlenecks for multimodal fusion.

What Matters in Building Vision-Language-Action Models for Generalist Robots Attention bottlenecks for multimodal fusion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.921133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:1930b9305343c495347b6bb145e6c8c633e3907eaa3c7520684cf906244a2849

Observation c37c1c17-84f1-47af-b705-f85104062e88 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

What Matters in Building Vision-Language-Action Models for Generalist Robots R3M: A Universal Visual Representation for Robot Manipulation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.756429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:5be5862b8356639dcd740f149135186613d07e46553525a73ecccef37aaaf4a0

Observation cc30b2aa-16b7-48f4-800c-82aa1ba557b4 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

What Matters in Building Vision-Language-Action Models for Generalist Robots Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.763828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:277fcf62b4f7accce700531b80faff754f937b833b7f3d16fa469ac7e6491edc

Observation 9919b5a0-4dbd-4ad9-9e9f-be9110487b93 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

What Matters in Building Vision-Language-Action Models for Generalist Robots Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.769303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:7cbd018e148e036da772da2b65ff2c9ea2c14875afa485b8c7c290ff4ffdea22

Observation b33b9288-af0f-45a5-a5e7-5f48c27ddbdb · outbound

This paper cites Real-world robot learning with masked visual pre-training.

What Matters in Building Vision-Language-Action Models for Generalist Robots Real-world robot learning with masked visual pre-training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.951449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:5a1c9c8152471a488dc07b55c098e21a942256a9bc691c4317a75f189d3c9fbc

Observation e91eae74-9107-415f-8f31-3356a00f644b · outbound

This paper cites A Generalist Agent.

What Matters in Building Vision-Language-Action Models for Generalist Robots A Generalist Agent

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.774708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:fdc1ab0d046d8adcdabb09fc94a32ccfd7b0075b42d1df81344db77154dde08f

Observation 345cdf66-8b1d-49ad-bb48-0028d7eba8e1 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

What Matters in Building Vision-Language-Action Models for Generalist Robots Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.958151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:423b5018fb103fb7ef70abbca8e13cb0f2ad0dbee8b1687bbaff107219b7d712

Observation 288a73e2-40a5-425a-a837-c0f7d44d0638 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

What Matters in Building Vision-Language-Action Models for Generalist Robots Octo: An Open-Source Generalist Robot Policy

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.779214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:ffe6256d1b3c4e5bcf1421be01def2da8a540db92a69a657387b65fe41aa4ab7

Observation 53be6ba4-0d61-4c3e-81dd-6bd5782e8c53 · outbound

This paper cites Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation.

What Matters in Building Vision-Language-Action Models for Generalist Robots Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.784456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:ab4f6c658885821f56e992a69b9390b86e8f22076175afa09340012063640ef6

Observation de5219c6-f0ea-4710-8d13-ef8de3f2de45 · outbound

This paper cites Uform: Pocket-sized multimodal ai for content understanding and generation.

What Matters in Building Vision-Language-Action Models for Generalist Robots Uform: Pocket-sized multimodal ai for content understanding and generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.974261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:74977363a3445e84ac9ad88e3fb2e7e68491ffe98b2887a5376ada0600f993d4

Observation dca11815-5cef-4c2a-8bf5-6bd5aea0f969 · outbound

This paper cites Attention is all you need.

What Matters in Building Vision-Language-Action Models for Generalist Robots Attention is all you need

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.940877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:e15431711e3519e91257568d2df4e6a30532f11dcbaeb710ee7156e93229ffb0

Observation 450d7ae3-4b84-4d3d-b226-755bdec2009b · outbound

This paper cites Moondream, tiny vision language model.

What Matters in Building Vision-Language-Action Models for Generalist Robots Moondream, tiny vision language model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.944398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:59f8e260942d4ff90ae3733fe4b18b5234faaf5409b35435d3c5a21f5472ebeb

Observation 11ef8a8c-0395-4273-b77e-63458688bd34 · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

What Matters in Building Vision-Language-Action Models for Generalist Robots Bridgedata v2: A dataset for robot learning at scale

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.947933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:a826d8dd905a17ff26c42ed8a930f552e1677d84e948532765fef187a2a06eb9

Observation 3680f861-18c4-4642-a5d4-6e6b6f7a5924 · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

What Matters in Building Vision-Language-Action Models for Generalist Robots Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.954947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:86726f0b86e4fd80100b8a66a9648c5833deb98098573ccd17cbe9b152be7803

Observation 45b2bbee-8bf4-46c8-a56c-125019b862e6 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

What Matters in Building Vision-Language-Action Models for Generalist Robots Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.789922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:3af183130fd29dec05bc11084d1b11b34e0ee6ca2e8cf7d57984a29304a5c304

Observation 2f274156-4873-4ef9-ab0f-4a87e6bc7137 · outbound

This paper cites VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding.

What Matters in Building Vision-Language-Action Models for Generalist Robots VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.795057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:4e41b0afbad4c97ce5ff7c356bed2b2f36b40547382d986fdec45dd68fb4fdcf

Observation affd3e59-1ee0-4735-8917-06ed7dde6ea3 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

What Matters in Building Vision-Language-Action Models for Generalist Robots The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.799854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:ee9639f1018558c89e92a369e3e1f2a73ea865fc5578bfc17ab4ef55582e40ea

Observation 4b663d66-8493-4372-a97a-1ac02bb8f69b · outbound

This paper cites Latent Action Pretraining from Videos.

What Matters in Building Vision-Language-Action Models for Generalist Robots Latent Action Pretraining from Videos

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.804809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:197c8ae4d5bc5e4b03e0222e92b27005082e1d3e74aec2d60db37ac255e50cc8

Observation 6b7cc466-0404-408a-9c6b-c4427ec8cb95 · outbound

This paper cites DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution.

What Matters in Building Vision-Language-Action Models for Generalist Robots DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.810099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:cefcaa1d7a77112f4c16382260958a04aa78d6d4d168f404ece8f655aa367fee

Observation 9f20cad6-497c-44f3-a411-8e205100dd2a · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

What Matters in Building Vision-Language-Action Models for Generalist Robots Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.814685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:73ce6cf8e73d108ce973a76daddea6c15432a95c88b1d1f5d84872415ca34bc4

Observation 26e200d2-dd94-4409-b05e-db71f1772e7a · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

What Matters in Building Vision-Language-Action Models for Generalist Robots Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.819324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:0a3249f9efccaae7fd573c06ba38e597e27dafbc3c7d7904c8aa7eb37b886b07

Observation a0045528-21f9-42a5-af08-787fab6777af · outbound

This paper cites Sim-to-real transfer in deep reinforcement learning for robotics: a survey.

What Matters in Building Vision-Language-Action Models for Generalist Robots Sim-to-real transfer in deep reinforcement learning for robotics: a survey

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:37:50.964769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:141c61ecf39a97f52674e1fdda7b2cb6c28a5d7b355b056a0a8dad1f18720141

Observation 5a3c3806-e811-4803-a045-c40be984e3f8 · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

What Matters in Building Vision-Language-Action Models for Generalist Robots 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.824010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:547a6bab64f1ea1c162a20841e1700860dc78b5a4f67c911a72db42b469a6049

Observation 7b90b889-cb7a-4be7-bbaf-b74925dd9442 · outbound

This paper cites ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge.

What Matters in Building Vision-Language-Action Models for Generalist Robots ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:37:50.657904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:4f657604dbd37824568aa19f8be7d107a569d0d403246cbad7f3628fd9ab3a0a

Pith citing papers

Observation 6f02b858-b7d1-47a3-81c8-de4291621983 · inbound

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model cites this paper.

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T06:12:19.643111Z digest=sha256:32e39d5213c6137e5b4e2d8e1350e231f1a240bc2d5d8643e9e05cf9a2231827

Observation 3e5c5f32-e6a1-4867-b353-6c4259945e00 · inbound

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model cites this paper.

HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T22:00:48.667428Z digest=sha256:a48e6ba3ba474f9da0a3caa6783bc1bb2c9d05f2f861bc63005cbc52f7a5c5ad

Observation b2d9fe60-5df8-4a13-999d-af208700494e · inbound

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions cites this paper.

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T15:28:06.883492Z digest=sha256:e3c552b0f3f51e3ee3ecb90d4d1da95a88de30586c0163674607b15bce3ad079

Observation 14a5a2ee-4f77-46f6-a14f-3c287f392eb1 · inbound

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models cites this paper.

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:40:32.989611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T08:38:02.944230Z digest=sha256:8b82a80859883b6b857c70a578fc23f5955c3f4466fb02bd499602c6e266d90c

Observation e5560d0c-f671-43d8-b56b-f9f4c0d786df · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:ea62343a732918b0763d31dd883f8c946bba334bbd59c0d5cc29ec4c07b34434

Observation aa552093-d4af-4d3e-82ba-3b401294b059 · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:c93d7a6fd0a3849f76c9d461d974eb5f80d6c6a069dfd98f60f9b9fe49e8b558

Observation c6f04bca-d4aa-40e5-bc1a-5929e0aaff62 · inbound

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models cites this paper.

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T21:52:02.893886Z digest=sha256:a26f6eecbf30845b7b54df3aa59a566a3cc2a9a9c17b2aae646c68a0d3abf89c

Observation 941c0693-7845-4c8e-900e-1492f4a0ccc8 · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.186146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.186146Z digest=sha256:167639bb0f9b7707807681f7fbb1d4b76e1082bc27e97f7395e6bb3768c81ad4

Observation 05320bbe-34ca-4a1e-a85c-b9294b827571 · inbound

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation cites this paper.

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T15:26:40.661829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:26:40.661829Z digest=sha256:e890ad38c8fdc7d61e3572ac02b8fcc80747abb8279a468cef75c851bbf86664

Observation a9002ed8-887a-49fa-9d4f-fdc60d2ae041 · inbound

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies cites this paper.

FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T05:48:46.550293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:48:46.550293Z digest=sha256:af77191408ac590c2422e20d23ebe259745318377a81ffd1821e59d9026e2e8f

Observation 6887cc6b-0618-4f86-a323-39a4b08a9c20 · inbound

LLaDA-VLA: Vision Language Diffusion Action Models cites this paper.

LLaDA-VLA: Vision Language Diffusion Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.925406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.925406Z digest=sha256:3e23e13672833a01829069aade466627f079660516ee8f7e9ab1363cd77694b9

Observation ad47c271-72d7-4eb1-b772-2be9468b1762 · inbound

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation cites this paper.

RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T20:10:13.532490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:10:13.532490Z digest=sha256:9ba0a308657cedaed63d44641059cbc2758f133ae7f97960ba557624d14ec39d

Observation 43c0e31a-92c7-4be8-ac5a-6b9597484dea · inbound

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models cites this paper.

FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:49.747548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:54:49.747548Z digest=sha256:7ac5ede371160fc38b719d1f30b6389fe8ebccafef91330cd729ff35ddfd2918

Observation ae57cef9-d09c-4f3a-a437-919db1abb4e4 · inbound

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models cites this paper.

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:43.447966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:34:43.447966Z digest=sha256:d1a18e95750c057f1edcf477947549d616713fae1e4c5b501a02e2c9616de80a

Observation bfca82c3-030d-4500-ba8d-b02ce8fcf60f · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:0e855ef34652927bcaa51ea1821dd132ebdf431122bdd209bf680c8b92830d9b

Observation 0823ed07-a098-43fd-bb65-84a784894e10 · inbound

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models cites this paper.

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T23:01:13.910539Z digest=sha256:139bd23b4f97a3faf8ce81b8f764b7f8c3f2b7a480e2320656c36af5341ca46f

Observation 6a301265-94fd-406f-b8c2-9b42cef5380c · inbound

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models cites this paper.

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T12:30:33.830536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:30:33.830536Z digest=sha256:79ec624d4c86fc2767a5d97178bdecbce2277e0e6c8af68c717e0760195efc37

Observation 17df520a-3ab1-4f7b-8b08-0b58539720cf · inbound

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation cites this paper.

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:13.372878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:13.372878Z digest=sha256:6dbc786442eaa5615a712fdbdfef93fd4d43a9e1dc3f83a378d6e9e3fac93efe

Observation bde4c1cf-9f0f-45c3-a83d-eca0a94d1a50 · inbound

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning cites this paper.

PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:05:30.324309Z digest=sha256:26f90de1309bfb1f7e9081142588ddbf2659f75e79c1504655f4a443184e05f2

Observation 3ada60f4-cda7-444f-9bde-9c004195bb68 · inbound

Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks cites this paper.

Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:12:12.842554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:12:12.842554Z digest=sha256:0b8dc928cd4711f0578bd15df075075b0b9f796b26af3872e515a6b88e17a40c

Observation 87aaae01-72b2-4596-833b-9378e6f387a7 · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:04.939146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:04.939146Z digest=sha256:22a2a3ce0488ea6ffa1aa6cfed932dc3ff337fa571e08bfdea7813ea58cb9fd6

Observation 54a2240a-1abb-4040-8116-622980df07a2 · inbound

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model cites this paper.

Multi-View Video Diffusion Policy: A 3D Spatio-Temporal-Aware Video Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T18:54:07.081457Z digest=sha256:7fcd24ec38f65032913c492ff216aa8258d69fade0c50953177621a7a27c426a

Observation 76a9d1ea-a4be-41d3-9f12-e9bd4cee6c22 · inbound

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model cites this paper.

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:31:23.255452Z digest=sha256:9e408fbb65c0918d6539f4932436b2c2746ea9354f2d08807ef2803714eece90

Observation 27b6b2fc-8d27-4e7d-a3c6-33eed07672a0 · inbound

JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy cites this paper.

JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:05:49.874253Z digest=sha256:5a155eaea150a93a8a0abfa139fa157e786267f992eb1eb712d0829a02aa1f42

Observation 32ca6a19-7b21-40a5-9197-1f9b20191251 · inbound

Bimanual Robot Manipulation via Multi-Agent In-Context Learning cites this paper.

Bimanual Robot Manipulation via Multi-Agent In-Context Learning What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T00:25:24.362191Z digest=sha256:3745d7be38b34911c5fcccb48fcfa5188d5546b6f0e193202059719a1887a4ea

Observation 86bfb119-9e23-4e0e-8518-b0d0aa742974 · inbound

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts cites this paper.

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T09:11:21.715023Z digest=sha256:cd45f6276b74097301cfb5239c1fe1f95635c3faa625028be228b54bf6c3323a

Observation 909513c3-68d7-4d84-a971-79577968e705 · inbound

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts cites this paper.

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:05:07.509291Z digest=sha256:c0c20f1e95cf44596ed84f4890db4b6c00f020e8528a636f84f7e517dbacfb10

Observation 737b6042-96ad-4d5f-bfa9-86e7651e7b7b · inbound

Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning cites this paper.

Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:15:19.347519Z digest=sha256:3058d2cbb36db49d69a59a46877454f5d36e3d2ae4274304238bf86f84c86894

Observation 5a725cd6-4940-4c71-a2f7-9ca3778c0476 · inbound

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation cites this paper.

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T04:44:03.661688Z digest=sha256:5b41c493b1aa9b93ba878b398b7c014f077c59fdabc452158335c6d6a295407e

Observation 8c45d145-87a7-47b6-a5fe-19764a7312ed · inbound

RotVLA: Rotational Latent Action for Vision-Language-Action Model cites this paper.

RotVLA: Rotational Latent Action for Vision-Language-Action Model What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:37:50.975242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T17:48:06.734816Z digest=sha256:7a8d698584707acf930cd3adfafb8b15173b073669a9b512a412bf3b1862ef24

Observation c4008f89-b3fa-4021-b7c3-d0a3999d9938 · inbound

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation cites this paper.

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:35:46.750240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T20:55:56.318157Z digest=sha256:8bdc8509ddb44315b8e6572b26a01551dfae13564a8ebb9b922e2d82d0d1bad1

Observation 9fdcabab-c52d-45b7-8230-af335289e687 · inbound

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation cites this paper.

IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T18:59:49.358909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:59:49.358909Z digest=sha256:37ce4db605a988447c8fd2c3ff52e405d1dcb1a2e9b708bbdbba7bf46b2622f8

Observation 1ba6987e-b015-410e-a7b4-6cb70b030d51 · inbound

PhysBrain 1.0 Technical Report cites this paper.

PhysBrain 1.0 Technical Report What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:37:39.950177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T16:34:44.204055Z digest=sha256:46c87a3cd9566aa52b09f1946290ca2701912ead315e7251616cd99c0ad452b2

Observation 5fa77bc2-1fb2-489a-96f9-3d1cca4507bf · inbound

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation cites this paper.

OASIS: Observation-Action Space Alignment via SE(3) Trajectory Prediction for Robotic Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:33:59.329882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T21:27:26.682689Z digest=sha256:c6a282c36d93d8a61e241a239bd906077187ae422e3f590039deb6cca0e057d1

Observation 33aced77-c7df-4dee-bb7d-8c91acfe6590 · inbound

VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking cites this paper.

VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:03:26.136063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T12:59:17.093731Z digest=sha256:3a791dab562430d8e54ea040e51a42744d302a6aa31eb2934af8dc34dc9caae6

Observation 76a1b0c9-b239-4667-8ca3-67eb07f8b574 · inbound

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling cites this paper.

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 100

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T13:33:27.784888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-29T13:33:03.368006Z digest=sha256:d302f4c7a77e9221aa8e6866b66874f8471ece6110c2a25ec26250da406f1215

Observation ed99c49c-fc2a-4b33-86ad-b7653c49ee90 · inbound

TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models cites this paper.

TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:26:28.502942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T10:09:08.056968Z digest=sha256:e6c695b77d0fb37029997b3eb2070b836f314d5d11fc2bedbb49e3dd36209bbe

Observation dfd867b6-7c44-4bfd-888d-9ad0abaf0381 · inbound

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training cites this paper.

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:36:44.041558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T07:20:40.998337Z digest=sha256:64df70bd338eaee515b467ce5ad237fba517949fe59d32de872610c8ff6d8231

Observation aedf29e1-5680-4cfa-a61e-0f673e20da5b · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T16:47:09.496354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T22:25:17.522240Z digest=sha256:5c54ba0e39aeaaf721b4dfb7f0c4247751aa485306814e65550b2de29f88aa72

Observation 1fe5ca7b-9480-4b7b-ad3c-be6f2500dcbc · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T08:35:34.440088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T07:17:42.045939Z digest=sha256:cd0bd3732d6e0cd87bd1c287e4192dd95892684effc8169e12cf2346bae2edde

Observation 8380a49b-2ac3-4eea-a7bd-13f20dfd7dfe · inbound

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents cites this paper.

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:47:38.576010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T13:36:56.552395Z digest=sha256:1da279fa90d089ff989b3ec50594b3a731e8f07ace130baaf8a9a5305ac824db

Observation 92c15b20-0b1d-4bff-b1a3-7b61a365a99d · inbound

VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models cites this paper.

VeriSpace: Spatially Grounded Action Verification for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T06:17:41.790875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T12:47:35.314486Z digest=sha256:ec77efde074f6b118447b2def73c168da72c95f9b3595fc64c375aab17eb2389

Observation 9891ac9f-4a8e-4538-86fe-69fe07102273 · inbound

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory cites this paper.

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:49:46.505486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T08:27:15.889885Z digest=sha256:5fcea8e9e15ae9834ee96cc465a87c4f008c05d306b4e0db866dd371450fdb50

Observation 0dde5dc2-ca81-465d-b62b-251bf7a85223 · inbound

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:36:01.254012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-09T05:29:36.753268Z digest=sha256:45bcfc9ae32a486514e9099850dbaf3e77dace7390db9600d0beac2821226674

Observation 9dc5e2df-7a63-44a3-9226-42faad756336 · inbound

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories cites this paper.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T00:04:04.138071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:04:04.138071Z digest=sha256:3fa25a77149bb7ef1c920bb5c05927a985aca50eb7b160d8ce8169d64fe11c0f

Observation 2b2bb9c8-09c2-484a-9e7b-208f63fd66f2 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:44.845464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:44.845464Z digest=sha256:01bd42eeaeef23867b0868512e8f912d1d0a5fa5fd144207f6e7732f74931cb1

Observation 8c5727ca-8a7d-4120-8242-c4e4033704a9 · inbound

Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models cites this paper.

Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T14:14:43.293927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:14:43.293927Z digest=sha256:26aac36b989c4c5de21e1753ce35401e7e56282827db30761ad4fc19e423d7ba