Pith. sign in

Paper Citation Record · LEDGER

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

As of 5 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 75 inbound Pith citation observations for arXiv:2502.19417.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.19417 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T22:53:37.120692Z

measured 126 of 126 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 75 of 75 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:55:52.126320Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T04:16:48.655148Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact26
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a958c7d-c056-49bb-94b6-790021a569e6 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models RT-H: Action Hierarchies Using Language

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:53:27.953855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:8bbe86c5f9a3a27509dba704c4fa60d0c657397280762fde784d58ebb0bb4ab3

Observation ec2d3a1c-da81-46e2-b3aa-fa771f7aa19d · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models PaliGemma: A versatile 3B VLM for transfer

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T22:53:37.179006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:6af1ea1fd363c909aa0d3c4ab3617c07659579dbe3fb6b73f789f6afbb1718eb

Observation 2d4f3a82-d24f-407b-9751-e0f38e9fdf9e · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.185946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:3c67899037a09d22bd3995f52dbc8970b6af9835a32788f5202120dfbded109d

Observation 94388843-6dc1-4022-8ce5-4f4fb8f9dd76 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.192807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:9646339500a5a6d06234fc695ec537867ffb47ef39b6cc924e42e1f489c0e413

Observation 95d55974-f9f6-4a4d-918f-621136103741 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.199637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:77d9a7c270ae27d7188b80370e2eee5c601a0cffce7fc82c7eb4e84fd9b958ef

Observation 3816f3b5-fb3e-472e-ac94-41b7b1b5ec31 · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Do as i can, not as i say: Grounding language in robotic affordances

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.397186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:28c4afd6a4751ce77608560861145223fe6ba17998c51ad7704b39c476b9278d

Observation 0735a5b0-3d82-4339-aa7b-ff4c591293fa · outbound

This paper cites Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Automating Robot Failure Recovery Using Vision-Language Models With Optimized Prompts

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.329808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:f3919b04c5b05c942ff75c6f4cfd31a68b566ba49b107b8413f3bc2e8a417bdd

Observation 0823afd3-d37d-46ce-a7af-e0e517a8bbec · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Diffusion policy: Visuomotor policy learning via action diffusion

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.406335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:bebc154a0e0e826521a8a8ad0acc7f5dc1a14c62c41454bfdca0c280c9171233

Observation 651dd099-c50c-4915-83a2-5fefd5a8cb12 · outbound

This paper cites RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.337900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:6f3d3e0786aa1467d0229e01633235debaccd0169f07beb43dc15e881c952f48

Observation 2a3ad65f-e7ed-476b-b96b-4df071b5b615 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T22:53:37.344281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:124a5654e8de38752877798bbae3f904eb9259d86e669885b6d9a6848ce81237

Observation a8b1fd37-20ca-4377-a929-ac09c1984afb · outbound

This paper cites Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.350550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:01111be781af00083faa6ca7f02225b169dad7bf2cce835e6d6ba8f2151b7d22

Observation 4d482576-0e4f-46c0-b471-50f0e5da1b27 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.357899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:4ae72842f13d9e2455a7efa44873f554e10bbfcbda6fc4b4b25d6180f1afa1f4

Observation 6ddefe39-adc7-40b4-9a91-13b71d5082e1 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Language models as zero-shot planners: Extracting actionable knowledge for embodied agents

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.426870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:a4c8a2e1173278c26f64b014ac7be32b4a2239b25f72686094477f218987dee9

Observation c3105b97-d7d5-4263-85bf-d9cd826e7546 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.364907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:3fd8de76f06228189c507311ccb630b8f37955f8c6c0fe492d57915004406bce

Observation e04c1e3b-a2cf-4f0b-bd4d-66d15c54f904 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.435427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:bafc78fc0cee7709bb271727c0fdf011daa12bd772fd63bc6c93a5408b7a7cd4

Observation 40e4c502-03cb-4713-af37-3a070f5a52e3 · outbound

This paper cites Thinking, fast and slow.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Thinking, fast and slow

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.439459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:fe307140b0aabf11dcd6157a4bc42e0faa82ff463df59e4eb64d6bd9b9834fb2

Observation 99c0666b-25f3-436a-987a-25319cdc5d3f · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T22:53:37.371736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:617b0485faa46cfea530487ad1b07f403991dba3e24bfa5082be71e8f3e15fc7

Observation 73eb7159-c35f-4db8-9ae0-51c67852d8a4 · outbound

This paper cites Interactive Task Planning with Language Models.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Interactive Task Planning with Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:53:37.379054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:4081b01ffc890202ae9e81108bb14b49046bd69d5b9cfa81f0c04acc9747e0ad

Observation 6d7b5753-3a88-4cf9-ab63-0a5e52d7584a · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.385552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:b3df7015b39e4b230309c3c997b13ea07bf224f5434909f7a11f05c06c3a04e7

Observation 5b7af09a-3387-4883-a626-7e73b7d4fa0c · outbound

This paper cites HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:53:37.392536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:ae80b7797d2812e0bc76e2c01371948a0c7d009649308152044911c034eb886c

Observation 18c5638a-d462-46c1-80e2-db4b451f5212 · outbound

This paper cites Code as policies: Language model programs for embodied control.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Code as policies: Language model programs for embodied control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.460513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:921bd8984caa5ddfa837cda53749440228c64e173dafe93b07118b26f9b596ce

Observation 95b78ed9-22bd-41c2-92f6-696ba21c9402 · outbound

This paper cites Moka: Open-vocabulary robotic manipulation through mark-based visual prompting.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Moka: Open-vocabulary robotic manipulation through mark-based visual prompting

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.464763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:4cb42d4b7bc649c381dbb16a797def828c918210e003f39f93824e1ad1645ed5

Observation 2d7d0f23-c3b9-433b-ae0b-928593ef3915 · outbound

This paper cites Interactive Robot Learning from Verbal Correction.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Interactive Robot Learning from Verbal Correction

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:53:37.205688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:664ed9bb0c9417d32ff01ae2c31b3fe5d14c4bcbcacdc545fde26a00bd7c7776

Observation 8568d71b-1d54-41f3-8620-7868421ca1eb · outbound

This paper cites OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.212576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:2793291a4b0b66af41a1734a408a0c4008aff719d4a6c176f971c5c5b6f73a59

Observation e33e9e85-65d1-4061-a93a-d09867759e79 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.219528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:5aaa04514065fc4c997d627df61aec0306f3cad39fef5d36ec7f5dc8d0500f76

Observation 9d8d0d14-e00a-45de-b9d8-2643144fb1f6 · outbound

This paper cites and Hutter, F.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models and Hutter, F

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.414665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:da9017d67ca815b698b014b7a2176bd64405d381d58d17c168e7781fd4be152b

Observation ec7e8029-ce24-4e92-b52c-fa52778d9631 · outbound

This paper cites Learning to parse natural language commands to a robot control system.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Learning to parse natural language commands to a robot control system

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.418645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:aca97f1789cdbfe0f8e82debb5441fe25e5f3807b01cb5ea61119d645ccf048a

Observation 75c4a7f8-2f7b-45b6-9803-2473cf569ca7 · outbound

This paper cites Is feedback all you need? leveraging natural language feedback in goal-conditioned rl.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Is feedback all you need? leveraging natural language feedback in goal-conditioned rl

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.422712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:8309ad8f3f9159ae380e97080454c825e3a0c719289a575f3d12a7b99a8d6ecd

Observation 32563067-76a8-4493-b945-b2997a98e215 · outbound

This paper cites Learning neuro-symbolic programs for language guided robot manipulation.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Learning neuro-symbolic programs for language guided robot manipulation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.430904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:82755e6291722b9608e0881cd7f525751c7d1de286d420cdbc28cb820d80f25b

Observation bac584d5-64b5-4b8d-85e2-2279e245f1fa · outbound

This paper cites PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.226372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:787ef06aaf39aac00dd05e01194b6a2343eb64dec716a3d10da9f23b7f3e1d5b

Observation f2a38f43-56c5-4fe6-a3f4-7edf90c441e7 · outbound

This paper cites Y., Sanketi, P., Vuong, Q., Xiao, T., Sadigh, D., Finn, C., and Levine, S.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Y., Sanketi, P., Vuong, Q., Xiao, T., Sadigh, D., Finn, C., and Levine, S

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.447769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:0d4c36e94b0c26c627a5db5b0d443ccc8930f1ae06a783d1a0d85a12e8bbb91e

Observation c6b3ef52-bd69-48ec-ae3c-d33268d8bd1f · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.452474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:07591f14d51f31af1d15be3164af9ea5398168ee5a3a0c44b1198a90b526ed1b

Observation ae395392-7a19-4573-8fa2-838f1232e18a · outbound

This paper cites F., Walter, M.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models F., Walter, M

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.456387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:329bc690b5f870732744b19b79ffcd3c7789a97525e029eed57a42db565aab69

Observation 432d5666-6801-4c84-88bc-4e89fa89b527 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.233298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:1b15058201f006fc8fbe1519c9a12a8cb1f5417ba3866ada9399c32ed7402dfd

Observation 641b521d-332a-45b7-b303-32f9ccc3ec5e · outbound

This paper cites Open-vocabulary Mobile Manipulation in Unseen Dynamic Environments with 3D Semantic Maps.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Open-vocabulary Mobile Manipulation in Unseen Dynamic Environments with 3D Semantic Maps

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.240701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:301a56f95f81632c0ff2d8cf246c7a645d55dc6f5df41b4a14e690a591aab461

Observation 80222e84-084d-4b9a-8ddc-af934b1f55a3 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.410596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:7e1514d36dcffe14189d52196a947595a8888ffc3700f12f61f6e71965ca7c2e

Observation 19bdf996-31dd-4e97-9e13-67ad52c450de · outbound

This paper cites BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.247250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:bf7d90534da48198ac413f09b4e835d9511ee1c5fdacb7d2f6a34d15ddd5ef12

Observation 6fc61aad-38f3-49f8-97aa-45a6a422476c · outbound

This paper cites Yell At Your Robot: Improving On-the-Fly from Language Corrections.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Yell At Your Robot: Improving On-the-Fly from Language Corrections

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.254239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:203254bb2aa0ce56c29769701f01dc83d493a54a7f36695e955576c081eff398

Observation 9ae65d96-6411-495f-8d08-1fa058f87ea2 · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Progprompt: Generating situated robot task plans using large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.401770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:eb90f1c743d65b049024c85fcb38ad3c1a13e18ff2195cf937376e086f39e467

Observation 3ee2b254-045e-4727-b55f-fe12164054e8 · outbound

This paper cites Training Fast Robot Policies with Slow Foundation Models.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Training Fast Robot Policies with Slow Foundation Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-24T00:23:02.532335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:9be14d80e4e612be2d9f0e63dcfdebf1313ecb3596d1975b8e0611583a7d1cf1

Observation ef5cb244-e53e-4bf8-bc3d-385b39448694 · outbound

This paper cites RLVF: Learning from Verbal Feedback without Overgeneralization.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models RLVF: Learning from Verbal Feedback without Overgeneralization

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.267277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:54dcae1161d0902c08934cf30fe1c2118d8fa175caec7654e32385f47d98390f

Observation cb1064fc-f406-4c05-916d-f63a04455ab6 · outbound

This paper cites Language-conditioned imitation learning for robot manipulation tasks.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Language-conditioned imitation learning for robot manipulation tasks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.443628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:d8a018a4c0310379c63271b29f41243766dee0cc14633d04b5d19c078a20b6eb

Observation ca32e68d-fc72-41e1-957f-8f24fc0e225e · outbound

This paper cites Open-World Object Manipulation using Pre-trained Vision-Language Models.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.275383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:33f9f6291c98847b3a549cea9fddf1c22e7308cb2589e532e4d37352ead6d22b

Observation cc439541-8449-4002-a0cd-3466210e0189 · outbound

This paper cites A computational model for the alignment of hierarchical scene representations in human-robot interaction.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models A computational model for the alignment of hierarchical scene representations in human-robot interaction

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T22:53:37.469021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:b3ac6182798fcd69968b95f1de1a3b7c66d8d8a8a5fd673fa9ada9a0edccd539

Observation 0cdda36b-5a7f-4d42-9578-097b375bf7df · outbound

This paper cites LLM3:Large Language Model-based Task and Motion Planning with Motion Failure Reasoning.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models LLM3:Large Language Model-based Task and Motion Planning with Motion Failure Reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.282338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:ff223087aafa87307271de270f20a83a719cf0506b1e6eca5034d9b9c005818e

Observation 59ad44b2-61a6-44d4-886a-06fcdd0365a0 · outbound

This paper cites TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:12:26.208161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:84e2fdcd0692906d6199cab86a8ad7fadc8bad5795826aa9ca9f8ee1aaae2c51

Observation b40cdafd-277a-4546-87cc-bc79b45938e1 · outbound

This paper cites Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Robi Butler: Multimodal Remote Interaction with a Household Robot Assistant

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.296992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:346bbec895180cbfe70250bbd81a4025fde041615a7963e1fc05b8b2c786447b

Observation 4addb3e2-b222-4426-9475-4765e7396165 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:53:37.303492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:a28a9f64ff8b71e19d1e415ff97bbef3a7e0807b93fd88ac2fc23373a9ba2b75

Observation 2c1b937b-7b89-4beb-a411-74e7fa79e76f · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T22:53:37.310049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:4e5ade3acca23d21c75d308d8307c1287054504081b1fd2afabdfc1659b55354

Observation a7db7abc-abe5-4901-8b10-5a40f51ccbb1 · outbound

This paper cites Universal Actions for Enhanced Embodied Foundation Models.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Universal Actions for Enhanced Embodied Foundation Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.317372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:5f72ed71191a2c8fca10d2cbcfd9cef88b6a30640c8eb1520989ae840c81dfad

Observation ce8349bf-d08d-4091-bc0e-9d0ff3388db8 · outbound

This paper cites Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V.

Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.323534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T22:53:37.120692Z digest=sha256:8c0371685786c4a940cc2d65f70031dfbe3836d3ec466be85a9d3f167af3725a

Pith citing papers

Observation 50987af3-8267-4e20-8e85-dcb663ef658f · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-23T01:32:22.535668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:966eb69852be8e9449fcead5ede538d4411a2900725db4b2c9ab6803fbc19317

Observation 2c161a03-7f64-4220-b7b8-029c50aa2b7c · inbound

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning cites this paper.

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:47:10.259876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:47:10.146795Z digest=sha256:407d3ddbad0b40441205917164ca4f7faea83d081c17a900490e6f6a6b0619ae

Observation 4ca952a4-2f8c-4c18-bd64-911acd1d9eb5 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.968391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:5d408c8f209c2acf2a2bf91a6e35e616c54904a649a6530c268dfcd4b19aa88e

Observation 8a692af5-b525-41ab-82c7-06c243af5102 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:08:35.042707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:f26e10d66fbcea7e75a9821bf9c35ce9ac8ffed51d7f766f24da14a3c9d58dca

Observation 66ed0c3a-cdc7-4cd9-9d27-6cfaa1ae04ba · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:42:41.606100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:4bdf78f381bde760e2047e11fbeae89725ab923e2e671d45339f80fe57d45736

Observation ef797cb4-db67-41e1-95c9-abee9238f569 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:01.079183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:5112df0ae497010d7834a51890f6d281c2047d9dd528b4a7b2e228c103362824

Observation b65a24f0-cafa-4ccb-a942-93d8d69415ba · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.126320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.126320Z digest=sha256:8ae1c3debd78607c206d4a8d90f48421d3f910be494cd50cecb02a5cab1ffe63

Observation 30ab6450-eb2a-4fa8-9537-1346c02fd10e · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.765861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.765861Z digest=sha256:737e96bd74f07b639b65e7ae1301d99088ddb7018b41b9151a76fe5d1304890a

Observation a5bc8679-83f1-4d36-b2b5-378c2170ad53 · inbound

Galaxea Open-World Dataset and G0 Dual-System VLA Model cites this paper.

Galaxea Open-World Dataset and G0 Dual-System VLA Model Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T13:31:09.955938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:31:09.955938Z digest=sha256:e7da7663ed3bd2f6228b12fb2b776c708c220bb7db1f07afad25a47ed9d3fb57

Observation 7f3bc4a7-4b41-4881-97f3-ca1276a2d3d2 · inbound

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation cites this paper.

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T05:26:37.189981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:26:37.189981Z digest=sha256:d752c81d98fb3562998ca2eb867e2340c9c459f4de242c57bbceacaf57533e70

Observation 248a9431-08ea-4cb8-9302-b9d334ec5d38 · inbound

Ctrl-World: A Controllable Generative World Model for Robot Manipulation cites this paper.

Ctrl-World: A Controllable Generative World Model for Robot Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-16T01:14:10.531571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T01:14:10.174044Z digest=sha256:27db03eeb2ce1450996a57941f0f85c6ef765801bb28b230a12c7968a944bfad

Observation 5d8032de-8d82-4270-a002-d8d4ebad9ef2 · inbound

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy cites this paper.

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:09:39.677347Z digest=sha256:b4d819a83f8f5254f8544a772b57e1878a9623811e5961ef54421523f065eb46

Observation 070e2f59-2fb9-4760-a14c-4a030752f44b · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:40.382324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:40.382324Z digest=sha256:58f371e4d1f0d8d33d4b11cc89f40719d85ba33a578d8f11173ef766948aad26

Observation 3c453cca-20f2-4bff-873f-c86d8923f044 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:d7b90fbd309b55656114385199e22aa651ab9deb033f9e9536b16fc2ed306df2

Observation a8f7678e-4112-411e-83ce-3985fcda86e9 · inbound

Mixture of Horizons in Action Chunking cites this paper.

Mixture of Horizons in Action Chunking Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:32:28.968924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:32:28.968924Z digest=sha256:059f31621f7357076d47a8e8c7ea33cceb2b9209a9957aeffbac39acea9fb846

Observation a3f6b563-cf25-4616-b239-2a8a523767a1 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 242

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:b65072eecafcb228a0001e7f803cca59c685562f325433024d68a1ae467c5e5c

Observation 4b97762f-3eb2-4c18-a3ac-67b6873a2ef8 · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:26.173751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:26.173751Z digest=sha256:68e2960f8b630ab70def9b04dddd6a3803986b73eaf80a4daf0a3263933cd65d

Observation 365ffa90-3b94-405a-86ff-ebdf9eb7ff8a · inbound

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training cites this paper.

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T13:29:50.876590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:29:50.876590Z digest=sha256:4b95f0c614b72f7f1471d2e53737c51102a071eca93ca43bdc58411869b4d422

Observation aaa2c874-3f2f-45a2-b3a5-7a2ff28b3497 · inbound

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control cites this paper.

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T22:05:39.797848Z digest=sha256:a9910313b3908deb76764b23f56642b60d3883461bd995a8d60b02730b95af75

Observation df08cd3d-ce2b-4ff2-87f9-8429a1be6fdd · inbound

World Action Models are Zero-shot Policies cites this paper.

World Action Models are Zero-shot Policies Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T16:18:15.003371Z digest=sha256:825807551162c050abfcac99fb868bc3634d32e7d2a5a59b2203287eead94198

Observation 8280753a-024f-4345-b561-bdd53e654a5e · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:29.479477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:29.479477Z digest=sha256:7420a6070ae1ee01dcb0d1539bde2a7d9ea6eee43522f2aa884d2f59f55cd692

Observation 845b1f96-7378-41f9-9a69-70a2e694fef6 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:07fd37b74f6a2c9a7b9bbbac9e0418b7be51064b8f0029865e99f06ed45025e3

Observation f8280362-d1e6-47af-8c7f-4634cf54b711 · inbound

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism cites this paper.

OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:50:03.915325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T11:46:12.134869Z digest=sha256:626498bb9b733f9bc3f6753f648a2456ee08a4f6a4de5c22883beefdb33aea48

Observation 4c4d5b42-98b8-475f-84eb-a87a98d54db9 · inbound

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models cites this paper.

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:54:36.125258Z digest=sha256:3f203e7e71a7f63daabc3db77864d9f886137d484e53088cecd3b0370d90bb5e

Observation 8d42439e-13a1-4b76-bead-14d28d5b7f3c · inbound

TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches cites this paper.

TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T19:50:50.950451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:50:50.950451Z digest=sha256:4e2d57527800e0642d43c153ed0f579b91e3a369cb6293b16a756ac8e0c1d540

Observation 914c24af-428e-41b4-9f11-03491d3bdce8 · inbound

ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-Making cites this paper.

ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-Making Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:46:03.360770Z digest=sha256:2252211d05ffea8859fae394fb8ff8288453965a7953a515a4af09b46299f01d

Observation aaae5352-3537-4fdb-ae78-25a238d7d4da · inbound

QuadAgent: A Responsive Agent System for Vision-Language Guided Quadrotor Agile Flight cites this paper.

QuadAgent: A Responsive Agent System for Vision-Language Guided Quadrotor Agile Flight Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:16:36.843013Z digest=sha256:8d12571034f4f17915b919c9a6e04bc00f1487a8262049a6639dc7f60ff02547

Observation 0fdc938a-94b4-4552-97bf-2c04e541adda · inbound

ExpressMM: Expressive Mobile Manipulation Behaviors in Human-Robot Interactions cites this paper.

ExpressMM: Expressive Mobile Manipulation Behaviors in Human-Robot Interactions Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:55:47.280663Z digest=sha256:66868bb483293a210de5640811ef26abf13c45afaec21827878dc768b3866400

Observation d1e66d14-0647-467b-bcf8-ebe8b6400adb · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:3fe7bc40dc5b12c774886f3712ec7c7c7a6d3f481ddd85c5ef0d18fe9b47d203

Observation 4fbe1e40-b620-4f53-b54e-1362db033d34 · inbound

AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement cites this paper.

AnySlot: Goal-Conditioned Vision-Language-Action Policies for Zero-Shot Slot-Level Placement Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:41:28.529494Z digest=sha256:9005c505023124b65bb6412bcee32faf0b892dd28edddff61802425bd55767ba

Observation a3bed610-9656-41fd-a8be-9408cbc8ec84 · inbound

HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System cites this paper.

HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:01:50.429917Z digest=sha256:6501ec7b99a49801d96fecb0e00c709a882f6dcb056ee156b6637617a517d503

Observation 1ac999ca-3d36-4ac7-9d26-60f8d00ab776 · inbound

HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System cites this paper.

HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:51:46.849464Z digest=sha256:3fd15211b206ef15d7d2997ccd2e76e6ad9de0906c9c05550c36c6e0946127f1

Observation cbc1d3c1-e7c9-43dc-8af6-2acac9628246 · inbound

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities cites this paper.

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:42:34.409651Z digest=sha256:f8b34e51661cfa46748cc092bfc3e147629317d23f9a5612920744eef326ad2e

Observation 803aa985-99d7-4029-ade5-a4f87932ccd6 · inbound

Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment cites this paper.

Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T03:04:46.460069Z digest=sha256:7722373d3ef985a405a92cd40d84c729c6200e01761ab9cd48c255cfe9cecae7

Observation 169523e9-0d5a-4408-bb95-d5d2eea44f0b · inbound

SEIF: Self-Evolving Reinforcement Learning for Instruction Following cites this paper.

SEIF: Self-Evolving Reinforcement Learning for Instruction Following Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:51:09.514927Z digest=sha256:79a62ada5bf4265aa18f19e8db61dcbdbb14a12638850eba7c70968063706ef8

Observation 8ca4eef0-e1bd-4d8f-92e3-b750f6589665 · inbound

G-Zero: Self-Play for Open-Ended Generation from Zero Data cites this paper.

G-Zero: Self-Play for Open-Ended Generation from Zero Data Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:a7c85595e191862afb6fa6aec53a2794ab87d88e785170479d3bf0f9cf3c1caa

Observation 62c97075-5d5e-4478-8c9d-5416001a01c2 · inbound

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models cites this paper.

Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:53:37.470313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T18:35:22.595183Z digest=sha256:4c9928ac9b529d142478cf9c8dd0e5df5e675b1d0d6045c3ced70f61e871d8a1

Observation 2bae7c3c-2551-4290-b2f1-fa36352554a7 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:43:16.954611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:718c96bfcb0e87e0417d80f508850f9d1f7ffba0aeac8356042870c9188c8c83

Observation 94237870-196e-4dcc-b8eb-663c8804ea3b · inbound

RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation cites this paper.

RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:38:17.018853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T12:34:11.744980Z digest=sha256:aaab1ca213a76a90c9954303ccb63a07d3537c78faa21f108f0078af39a313e5

Observation a4f803d9-cfcf-4bf1-939a-a1b638b79833 · inbound

Action with Visual Primitives cites this paper.

Action with Visual Primitives Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:31:08.115243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T05:26:55.722814Z digest=sha256:ef03b401ef53808805c0706fc3911f1e91dba0435185e1df52622d127b88e493

Observation e5551921-f9cd-43d0-8029-9bf34b16985f · inbound

GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations cites this paper.

GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T04:54:36.746142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T04:51:21.649138Z digest=sha256:409ea4cdeef0286fca7e24e50590779be00c3935a03e656b36f7051fe138eaa6

Observation 68737d5f-53ef-4b44-a07f-3738098b9817 · inbound

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation cites this paper.

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-22T04:46:04.531661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T04:46:03.020800Z digest=sha256:7c14d8e69dcd768c787075a0879899362c137510f948bc66249cd10ec5b61159

Observation 779e8078-d37b-4970-8c9b-3f1d3416e2f6 · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:33:58.968708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:e52a3b121d72affd67e7028000f0904ac077bee3395f717f0190c674c0f242a8

Observation 77cfea7c-fdc1-4b5c-936a-9945994b8735 · inbound

Wall-OSS-0.5 Technical Report cites this paper.

Wall-OSS-0.5 Technical Report Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:26:00.445957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:35:00.258436Z digest=sha256:ffdef7b0c8694bab92ac3ed552206bace02f2e7ab7f112e40a54503dd835213c

Observation 18b87252-64d6-4bd0-b933-da0bb49959bf · inbound

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs cites this paper.

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:36:24.052646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T14:04:43.270155Z digest=sha256:b92fbde16186e93399ff6816ee928f7b04f01037dd1db59c8504080feef27be3

Observation ec7e2c67-ad2a-4736-af5a-65a820c2e7d4 · inbound

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation cites this paper.

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:46:33.253997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T09:40:04.685274Z digest=sha256:9f8a159871047cab00188c8de5d199a108285d947deaa5cd8e090e95c0accf5f

Observation a08dcb0d-ee6a-48a8-a674-d42812e1e26a · inbound

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding cites this paper.

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:16:59.247183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T01:23:02.576098Z digest=sha256:be77573e0f7301c2b9d4b2c1afcc26f240f8970a1ae80d3220045fd6f23361a0

Observation 6c0a7078-71da-49a4-a094-a0129b447742 · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T16:47:09.455277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T22:25:17.522240Z digest=sha256:62f50b1ddc44af176a60c3d052030483188da1976291950cd2dfb802a7807946

Observation 0e6af7fd-a55f-4de5-9cb1-0f59b107e7ef · inbound

LARA: Latent Action Representation Alignment for Vision-Language-Action Models cites this paper.

LARA: Latent Action Representation Alignment for Vision-Language-Action Models Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T08:35:34.455697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T07:17:42.045939Z digest=sha256:9fc349a5a8577c8c345bf9d17138b75acc72f8b360d97617ee14519b5da4a065

Observation 24baadb7-8a8e-4df9-b1ed-9da6f6426033 · inbound

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation cites this paper.

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:07:18.185116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T21:40:00.330510Z digest=sha256:6442310ae2747c86243ecfe18425e1a93c8c768941535848a4d8ef06eef01f16

Observation f2a8c76a-05f9-48db-b3cb-bbd976b06c7b · inbound

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents cites this paper.

What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:47:38.578601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:36:56.552395Z digest=sha256:6bffdb1cbce929159d7a8453a02e7100761263cc0b2dfb2e8f2ac85a653b73cf

Observation 3a5f5cd2-023a-4867-ba6e-dbea9522f1a8 · inbound

JOIN: Anchor-Grasp-Conditioned Joining via Opposition, Inference, and Navigation for Bimanual Assistive Manipulation cites this paper.

JOIN: Anchor-Grasp-Conditioned Joining via Opposition, Inference, and Navigation for Bimanual Assistive Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T05:57:41.724633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T12:59:50.106598Z digest=sha256:6d568abd29785cbd83f94942b3a76ff82c6da7230a03bf5e695b8fc655163527

Observation b37a4ce4-6dba-415e-8231-3a56c7c229f3 · inbound

DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners? cites this paper.

DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners? Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:08:03.023092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:42:17.813126Z digest=sha256:1d3fbbb75ac24ac4116fe8bfee67f300ee53190c578ba9c1bd19384e4f28b10c

Observation 9f2d1c0b-67d9-4d20-a339-c72d382858a1 · inbound

Learning to Assist: Collaborative VLAs for Implicit Human-Robot Collaboration cites this paper.

Learning to Assist: Collaborative VLAs for Implicit Human-Robot Collaboration Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:37:56.654045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:57:14.119334Z digest=sha256:9f3081116b062da5fdbcfb5d4f12e9c8d448419afa623630416ed3b0a00c3a10

Observation 7c71fc48-fff0-4613-b4d9-bb17fba5cccf · inbound

EA-WM: Event-Aware World Models with Task-Specification Grounding for Long-Horizon Manipulation cites this paper.

EA-WM: Event-Aware World Models with Task-Specification Grounding for Long-Horizon Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:48:35.303669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T06:22:27.987539Z digest=sha256:983622c20de1db78c1e8bc91165b5c341c2a0d599cb1d378493cc947cd94fde3

Observation 5eb0afc2-384a-4880-b56d-4c8d8b945008 · inbound

Geometric Action Model for Robot Policy Learning cites this paper.

Geometric Action Model for Robot Policy Learning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:38:43.878770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T03:59:17.983035Z digest=sha256:2b4c25c0156c954a8ace7ee8ca7d31ea9b5389b79bd412fca2d2c10bb514d26c

Observation 07ad6dc1-1fd3-4a4b-8233-9b1ee89c5903 · inbound

MemoryWAM: Efficient World Action Modeling with Persistent Memory cites this paper.

MemoryWAM: Efficient World Action Modeling with Persistent Memory Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-07-04T04:29:34.822739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:00:43.138441Z digest=sha256:c49d4d66dc21d923c1e32ed9bf9a22e501e3919d314295662d452a0a29cd9614

Observation 86462f29-4896-4aa8-b206-0689bac539ce · inbound

InSight: Self-Guided Skill Acquisition via Steerable VLAs cites this paper.

InSight: Self-Guided Skill Acquisition via Steerable VLAs Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:58.295928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T00:10:51.721485Z digest=sha256:36ee9e00efe386b124f2435d4dbaed232c5aa142f3f3402325af6e0b5e47212f

Observation 3c84ee17-17b2-4b28-8d4a-dd6466e08a7d · inbound

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation cites this paper.

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T15:59:56.659449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T01:10:24.450509Z digest=sha256:9fe0e31aeaac9363ea09277d29847895492ae9d0cf9998e7bb39b388c488eb4c

Observation 1aec08bd-09df-4cc8-bd44-2cb15d2ad425 · inbound

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining cites this paper.

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:59:52.647202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T04:41:13.473022Z digest=sha256:83a130b7bb4a59de1a00f04742cf31ee93f0b0a04864b8096afe195920d301a3

Observation 1889901c-8948-4a31-aea5-a829b69fad20 · inbound

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining cites this paper.

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:53:56.078322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T04:38:02.797595Z digest=sha256:f9aea98f7acb1ce7db6289dd5da6d7344f4df5632758e33665515e73dff53147

Observation 13f5f1f1-e77e-46a0-bf46-4186c37f3ffd · inbound

Adapting Generalist Robot Policies with Semantic Reinforcement Learning cites this paper.

Adapting Generalist Robot Policies with Semantic Reinforcement Learning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:45:42.746779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T05:09:29.625066Z digest=sha256:534d64ddb05cffc9f1c0c81b7c6bf900dc4c4cbc2e07830997f35df8c718faaa

Observation a2fae6f9-0aae-4919-90a1-65721e2ed10f · inbound

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots cites this paper.

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:37:56.102497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T10:37:31.925624Z digest=sha256:237e6efbef74fc54cceccb5b76473978cea29ee12908df7e9fe79bed40463f73

Observation d0f7b84b-c647-4728-b060-0a6ad9b24bba · inbound

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots cites this paper.

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T08:00:25.815355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:00:25.815355Z digest=sha256:d18b6b0524def09d070eff8d5f134ed73ac5ef15d24a12befe24fc69075cc824

Observation 6f0e86ae-3c9b-41fd-8caf-88eccfd4a26b · inbound

PriGo: Test-Time Primitive Guidance to Diffusion and Flow Policies for Adaptive Robotic Manipulation cites this paper.

PriGo: Test-Time Primitive Guidance to Diffusion and Flow Policies for Adaptive Robotic Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-09T20:46:33.633766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T20:38:38.404267Z digest=sha256:343e5ec93f68dba1521c11864183aa6effde2dbfe1a773d54cd135514159f263

Observation f57cc924-132f-4cb7-af06-bdc38b5474cb · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-10T04:16:48.656424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T04:12:29.153764Z digest=sha256:d7dfe718c04a2cfe75078ba91d56912ccf06fc26327c833b9ee5568858fa8891

Observation 7743bd04-6843-40a3-9be4-6d6e366a7489 · inbound

Native Video-Action Pretraining for Generalizable Robot Control cites this paper.

Native Video-Action Pretraining for Generalizable Robot Control Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T07:53:40.859548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:53:40.859548Z digest=sha256:28c1df311ed618f135b235c3322d5e12ca8d5b0fe5cc45b9f85aa66d1c5000f3

Observation 93d8a16a-d707-4932-9130-093735406ff1 · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-14T12:26:27.446079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:26:27.446079Z digest=sha256:f8065a34b3c547b4ed7875bc1cc2e1989dbb013c52ad540a7bd000b7b8ec6852

Observation 4f636bde-a806-42f0-a043-33e6e6a357e0 · inbound

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory cites this paper.

ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T07:21:30.948050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:21:30.948050Z digest=sha256:31e934d1c241e09dea8c08d6239e5b3a0e0258efba819682e3279ca2fcc75e58

Observation 73da70a0-e6e0-4e29-a36c-c75fc076fc8a · inbound

ABot-N1: Toward a General Visual Language Navigation Foundation Model cites this paper.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T12:10:21.115628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:10:21.115628Z digest=sha256:de70b001ead9ed9161205f2f57500fdd4f171217e3b32ed53c81d437e867a5ac

Observation 7cadf22a-03b2-4143-8bf3-e1a1e5b5c8d7 · inbound

ABot-N1: Toward a General Visual Language Navigation Foundation Model cites this paper.

ABot-N1: Toward a General Visual Language Navigation Foundation Model Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T07:19:41.156501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:19:41.156501Z digest=sha256:ae7adf9902d0093eb5cfc2aa032d457fcdc2cad0a13014b899241377e01aaa10

Observation 7180c7de-5691-43d7-af8f-814f712160f4 · inbound

Orbis 2: A Hierarchical World Model for Driving cites this paper.

Orbis 2: A Hierarchical World Model for Driving Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T22:05:04.115994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:05:04.115994Z digest=sha256:d4f46cfbd25c8421ffa1678dc18d0ae6ab1b5f99b4ddc7147b8c30ab30581227

Observation bfeedb94-76f1-423e-9cad-4ba9dee30d3b · inbound

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning cites this paper.

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-01T16:19:09.606067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:19:09.606067Z digest=sha256:032db3fc33b853fd31ba8794b1ad020652be3ce89e8fa3facdb69dc80f14876c

Observation 191f28b7-9290-4e84-ac84-0a4f29b1a1d5 · inbound

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation cites this paper.

RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-01T14:39:52.439334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T14:39:52.439334Z digest=sha256:80284ab463cf3c2f80cfcb7e035b37602d9449804b110fd6b1a5803b04787a7b

Observation 704dceff-fb85-414a-bc98-3db68466de4c · inbound

Addressing the Orchestration Gap in Generalist Robots via Physical Agency cites this paper.

Addressing the Orchestration Gap in Generalist Robots via Physical Agency Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:55:48.334745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:55:48.334745Z digest=sha256:adaae2dd44d7be09ab7bfa3ae780807db274aabecb3a738c97e4838a4b12f5a3