Pith. sign in

Paper Citation Record · LEDGER

3D-VLA: A 3D Vision-Language-Action Generative World Model

As of 23 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 100 inbound Pith citation observations for arXiv:2403.09631.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.09631 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T18:18:27.211034Z

measured 162 of 162 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 100 of 133 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:08:58.446528Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact21
  • verified fuzzy33
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

14
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 97c26ff7-2daf-42b0-9f32-c61a0597ce06 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

3D-VLA: A 3D Vision-Language-Action Generative World Model Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.315022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:a8776315b488221434f0ec1f516854deadf34b60851cb98098fdee3c913f0268

Observation 54c519dd-d490-47a3-84ab-5477256c9f96 · outbound

This paper cites ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth.

3D-VLA: A 3D Vision-Language-Action Generative World Model ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:12:47.371512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:71cc119d916370c13dfa8fcb76512fd826feb65c824f91a7db4d64241733daf2

Observation ecbe30c9-510a-4a0b-91f1-8e058c01ba94 · outbound

This paper cites Zero-shot robotic manipulation with pretrained image-editing diffusion models.

3D-VLA: A 3D Vision-Language-Action Generative World Model Zero-shot robotic manipulation with pretrained image-editing diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.326993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:40014836ef60b04883eec00245d93ed9a366aef2054167a714751f90b99595c3

Observation 26525cfe-1dc5-4f41-9bed-18f1fcf5a0e2 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

3D-VLA: A 3D Vision-Language-Action Generative World Model RT-1: Robotics Transformer for Real-World Control at Scale

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.262017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:c97d488fee3cbd10717d4b751183a8a3a430d57438a052539a19d9342d32eb7b

Observation fb41528e-1b26-46b0-ae0e-888efea28365 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

3D-VLA: A 3D Vision-Language-Action Generative World Model RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.264601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:43c794b4c2fdaf893ee9b1c9a2bdd07175c44dfc2e40afdcd9c78de50e5e5a5c

Observation 0f4e72b5-d1fc-4267-b9e7-f2ecc317f2b0 · outbound

This paper cites an unresolved cited work.

3D-VLA: A 3D Vision-Language-Action Generative World Model Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-13T18:18:27.328615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:a267dad33cb8dc8a3903c1402b4e22de653cfcc20e8b34387d57b9331c258b28

Observation f82a101a-3f98-4ca4-9c88-9ec18aa98a34 · outbound

This paper cites Playfusion: Skill acquisition via diffusion from language-annotated play.

3D-VLA: A 3D Vision-Language-Action Generative World Model Playfusion: Skill acquisition via diffusion from language-annotated play

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.330344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:8c9f88f55d3e8acf807146970727225926170c85ccdd23ec6978766858e736cd

Observation 94c34a7a-f505-4621-a41e-235d45706b52 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning, 2023 b.

3D-VLA: A 3D Vision-Language-Action Generative World Model Ll3da: Visual interactive instruction tuning for omni-3d understanding, reasoning, and planning, 2023 b

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.332108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:2d9506d831ebe27c03df49b6013cc04598effa591bb18ebc0ba1ef1554f8a73b

Observation 4ec0708d-5dad-472e-beaf-36a3579bf0aa · outbound

This paper cites X., Savva, M., Halber, M., Funkhouser, T., and Nießner, M.

3D-VLA: A 3D Vision-Language-Action Generative World Model X., Savva, M., Halber, M., Funkhouser, T., and Nießner, M

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.333726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:7acee92e0be62d0fc6d0ecd70e33405b0b826ae91a42c89ec0e25d206accd2a0

Observation 2acc2055-fb9e-4b7f-8f8e-6bdad4863da6 · outbound

This paper cites M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al.

3D-VLA: A 3D Vision-Language-Action Generative World Model M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.335712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:0fd77304fd021530b3201c67c9820f2a20ce940e639458158e14d690c0a48cc3

Observation e56ab7a1-7359-4dc6-ad3d-3cb9d452bd56 · outbound

This paper cites an unresolved cited work.

3D-VLA: A 3D Vision-Language-Action Generative World Model Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-13T18:18:27.337529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:a8105c424f4c929207cc41eba7ba67b98753161d250fb39d16f2d398d3573029

Observation 9696f091-f581-41ee-ac2f-5be77b0e1e27 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

3D-VLA: A 3D Vision-Language-Action Generative World Model Objaverse: A universe of annotated 3d objects

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.339047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:11e3485feb479f6179a1607f19a2911aae4d9e0c977d6fe902b3b7619e164179

Observation b84cd66f-e925-47e8-99c8-681c07eb4b41 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

3D-VLA: A 3D Vision-Language-Action Generative World Model DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.239045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:517bf75f3a0937306bbe762507904a42a8dc43bf02159c8dcf85ee78eda6af9d

Observation 79465bc0-9128-4873-b79b-f6626afb40a9 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

3D-VLA: A 3D Vision-Language-Action Generative World Model PaLM-E: An Embodied Multimodal Language Model

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T18:18:27.245035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:819e50ef143d1b63a4b7ca362d618e2befa9a4c622caa35dde315ee144e62266

Observation 28a5589f-a5a9-4c7a-870c-92474b4b305c · outbound

This paper cites an unresolved cited work.

3D-VLA: A 3D Vision-Language-Action Generative World Model Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-13T18:18:27.340489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:be11b1d36b277bb1fc799018309079acbc38c5b1a873eb3dc8f7b738d9b954fe

Observation a1e5f78b-b147-426a-a44b-18ff132741de · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

3D-VLA: A 3D Vision-Language-Action Generative World Model Structure and content-guided video synthesis with diffusion models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.342207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:b86b5b5953899228fc1c811f539fadaef9936868b3415885a78bda5d10a7fa9e

Observation 2ebb3c51-960b-47df-8712-0be93cd09567 · outbound

This paper cites RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot.

3D-VLA: A 3D Vision-Language-Action Generative World Model RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.269979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:8ce9f11aa1db1814781441ebc43661ac08efd11f35bc00c24a145eaf0d415957

Observation 9372c22d-f714-43ae-b152-3b0be8d7a963 · outbound

This paper cites Finetuning Offline World Models in the Real World.

3D-VLA: A 3D Vision-Language-Action Generative World Model Finetuning Offline World Models in the Real World

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.283738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:094feef5c393f318cba8788bf971a42d225f8a7a8b68ea316d1700295ec43f25

Observation daef5360-dc86-4303-9daf-02ee05aec76d · outbound

This paper cites Point-bind & point-llm: Aligning point cloud with multi-modality for 3d understanding, generation, and instruction following.

3D-VLA: A 3D Vision-Language-Action Generative World Model Point-bind & point-llm: Aligning point cloud with multi-modality for 3d understanding, generation, and instruction following

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.343881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:40ec9e94f26cfe27c0070b64cb26736c085b081b9a88525e2a0e7ed28c6ee513

Observation 6c53d0dd-a18c-4c88-9f4c-840e83f7097a · outbound

This paper cites 3D-LLM: Injecting the 3D World into Large Language Models.

3D-VLA: A 3D Vision-Language-Action Generative World Model 3D-LLM: Injecting the 3D World into Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.295052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:6a1c4b3329941aa90ba40248bda57021d9a776ce42d92712c0dcd2554b5c2dac

Observation 91595000-032b-4048-ab83-af3e71f081f2 · outbound

This paper cites MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World.

3D-VLA: A 3D Vision-Language-Action Generative World Model MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.304794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:4cda0892ba731cff07ca2c8c46cc8980ce26561cdee1465e75bc3bd584e70379

Observation 6bb457c2-2154-4c7a-be26-4f239908b6ec · outbound

This paper cites and Montani, I.

3D-VLA: A 3D Vision-Language-Action Generative World Model and Montani, I

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.345448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:15b8589bd08449f557a75332d6f6f4eed5437d0dee16b8d1eff4b221f6087470

Observation 713f92ed-29cc-46f3-af21-25a6bee88c14 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

3D-VLA: A 3D Vision-Language-Action Generative World Model LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T18:18:27.241928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:a3b76159a46c0c7135ac2c5c44090c64d0f25393f48a89251896ab70a6681516

Observation 3468c4cc-5af2-4be1-aead-87c78eb6d5ef · outbound

This paper cites Chat-3d v2: Bridging 3d scene and large language models with object identifiers, 2023 a.

3D-VLA: A 3D Vision-Language-Action Generative World Model Chat-3d v2: Bridging 3d scene and large language models with object identifiers, 2023 a

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.347029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:2c28379ed9950ea4418a3b8c68ad6640f5e12a0e7d7cd70a861f3e1648f3f50d

Observation fe44a321-56a2-416f-8c0e-a8cbd7e22ba0 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

3D-VLA: A 3D Vision-Language-Action Generative World Model An Embodied Generalist Agent in 3D World

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:22:18.773673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:8da0300805c1ea170ac07abd53d61653fb925d159f8f121c0b8f6b0aec729973

Observation 413872a8-85d0-4228-b6ab-d2a98e89f293 · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

3D-VLA: A 3D Vision-Language-Action Generative World Model Language Is Not All You Need: Aligning Perception with Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:32:23.026112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:ce20172487c1c92bfc5bda7ea771b9d44c02185ddd3d8b895468c8c8720c67be

Observation 6dd03766-d22d-41dc-b1cf-64b51615841e · outbound

This paper cites R., and Davison, A.

3D-VLA: A 3D Vision-Language-Action Generative World Model R., and Davison, A

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.349135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:9e5738dd28506dd4a744a2d4fa8bccb65e37668c4c4889f0044f21e38801da60

Observation 628055ba-4871-4010-9d8b-5a332ecff3d7 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

3D-VLA: A 3D Vision-Language-Action Generative World Model Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.350802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:38fb66dc67a76f7397cbfd63d6b18408fcb5adb51689ba7526d6a920f2ad2775

Observation a379add8-27ec-4f8e-89cb-b7ef2806394d · outbound

This paper cites Auto-Encoding Variational Bayes.

3D-VLA: A 3D Vision-Language-Action Generative World Model Auto-Encoding Variational Bayes

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.267199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:a5e749582d30fdbd509f72f58f1920ab2ae299b4d118004d6cde31525cf73fb3

Observation 715ce97a-70ca-4f6f-9b17-991f49e24ccb · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

3D-VLA: A 3D Vision-Language-Action Generative World Model Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.352405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:b097df84963eadfa39d277840fe5e75ada1ef2cc1ce1bef456e01bc83e748a50

Observation ab28b99f-416d-4fc3-9526-e3528ecc82bd · outbound

This paper cites CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding.

3D-VLA: A 3D Vision-Language-Action Generative World Model CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.278672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:076de1dd548b40267de87f9a2eb80d2f7eeebd7014367bb7cea1bb5d76fc79b3

Observation 5981af40-678c-4683-9992-f97fd2844b89 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

3D-VLA: A 3D Vision-Language-Action Generative World Model BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.281074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:911cd481276746022dcf3a6e0c595a000bb7b83647b35021526751e154f324a8

Observation 3b80b2e0-cc5a-429b-b8a9-1f735fda131a · outbound

This paper cites 3dmit: 3d multi-modal instruction tuning for scene understanding.

3D-VLA: A 3D Vision-Language-Action Generative World Model 3dmit: 3d multi-modal instruction tuning for scene understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.353994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:4b672059355b00f8b8ec728db8f7d669fc5da99a7d87c40a962eb93bfecda17f

Observation 78a9492e-e41a-4cf4-bb1c-368f1dfb832c · outbound

This paper cites Visual Instruction Tuning.

3D-VLA: A 3D Vision-Language-Action Generative World Model Visual Instruction Tuning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.286127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:8aac581413a4e8b53b216669e8ce1ec78f151c0956dfac0c06e25329f21824e1

Observation 13a4c3ac-e92d-4788-b118-19e30c4b12ca · outbound

This paper cites Hoi4d: A 4d egocentric dataset for category-level human-object interaction.

3D-VLA: A 3D Vision-Language-Action Generative World Model Hoi4d: A 4d egocentric dataset for category-level human-object interaction

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.355723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:3acc24141753ce8f784afe2cf7cf1c4492383806815aa220cff69b470b205e99

Observation 9e6ad27d-fc14-4499-8d7f-3e24c337c29a · outbound

This paper cites UNIFIED - IO : A unified model for vision, language, and multi-modal tasks.

3D-VLA: A 3D Vision-Language-Action Generative World Model UNIFIED - IO : A unified model for vision, language, and multi-modal tasks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.357528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:a4af6c179ac8330c830b2980f41e4a1c8d16d23e1d150fc216671e573971a0b8

Observation 1ce1ef94-0f93-4ec5-99a0-c2f9185fd7b6 · outbound

This paper cites Language Conditioned Imitation Learning over Unstructured Data.

3D-VLA: A 3D Vision-Language-Action Generative World Model Language Conditioned Imitation Learning over Unstructured Data

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.298710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:12d57b9e53909e1edd04d5f39b65e278f2587af61e1ae3b3577bef7f74647c13

Observation 0f286e5d-f3ce-4600-92a8-f1d175c6dac4 · outbound

This paper cites Interactive language: Talking to robots in real time.

3D-VLA: A 3D Vision-Language-Action Generative World Model Interactive language: Talking to robots in real time

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.359197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:69dcbbd5e295be90d0abe8592e63842dc60eb1ae237c8466a629a59c646055ae

Observation 07e492e3-b632-43ce-a334-cc76f3f9dce7 · outbound

This paper cites Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity.

3D-VLA: A 3D Vision-Language-Action Generative World Model Scaling robot supervision to hundreds of hours with roboturk: Robotic manipulation dataset through human reasoning and dexterity

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.361251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:78ab3b9cad7ccac438abe513c86e3fc5b5508bf2c7aaf894e0013eab06d0e3e6

Observation 215447fc-f02c-40bf-a3e2-2f7d50dffe5a · outbound

This paper cites and Chater, Nick , year =.

3D-VLA: A 3D Vision-Language-Action Generative World Model and Chater, Nick , year =

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:27.235951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:3b904b7d5654c15205543e7f794139e73bcf09d7bb68ad48eba35c930835a069

Observation 83e2b69b-3d7c-46e1-807e-2d6ab947a696 · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks.

3D-VLA: A 3D Vision-Language-Action Generative World Model Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.363121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:4f57ef50c74cb0945b4f0eaa7936ad4b408d1dfc40c064aca0f4024f66ce073b

Observation a72ae54b-0011-465a-8eab-dfc30cf77d7a · outbound

This paper cites Grounding language with visual affordances over unstructured data.

3D-VLA: A 3D Vision-Language-Action Generative World Model Grounding language with visual affordances over unstructured data

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.364872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:6eae183520133d00494d219f54a6f81bb1504c662f817085bea336004a9061d7

Observation 702edcdb-dfc9-4d5e-be54-6c7d9ccab732 · outbound

This paper cites Point-E: A System for Generating 3D Point Clouds from Complex Prompts.

3D-VLA: A 3D Vision-Language-Action Generative World Model Point-E: A System for Generating 3D Point Clouds from Complex Prompts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:51:35.400576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:34d09c3a3c6cd422df3e479f78c3c723c40b7bc74f8b7f1a59f089ed103bb987

Observation 13af462d-d4de-422b-95f2-4a6e7a6c5680 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

3D-VLA: A 3D Vision-Language-Action Generative World Model Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.250764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:0635f846a924ad428096756ebefe5930cf0f4f6bbf6cf8e7cda3b39e57f75a51

Observation 1bef82e3-21cf-4661-a399-b3e06604126e · outbound

This paper cites The effects of contextual scenes on the identification of objects.

3D-VLA: A 3D Vision-Language-Action Generative World Model The effects of contextual scenes on the identification of objects

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.366560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:a26df2cb4439270e09c16870114f8e59ea7a806dd6df8b7ec5fb490c55bc323d

Observation 1016d80c-7344-4cfd-8dec-c1e3aa05f3c2 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

3D-VLA: A 3D Vision-Language-Action Generative World Model Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.256407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:cb226185921309d02e8712a5fbe5b0a184347157b85c252a6a517757ff3ec33a

Observation e6884e3c-673a-4d5c-bda9-b7a6d21dbcaf · outbound

This paper cites Seeing and Visualizing: It's Not What You Think.

3D-VLA: A 3D Vision-Language-Action Generative World Model Seeing and Visualizing: It's Not What You Think

Reference 47

Resolution
verified exact
doi, observed 2026-05-13T18:18:27.232263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:cfa359296b03632c416b4daa6603ea18875bab691a0659aab34bf1029d5e322f

Observation 1d1f9053-699f-49f2-88c8-43a5177d1300 · outbound

This paper cites Gpt4point: A unified framework for point-language understanding and generation.

3D-VLA: A 3D Vision-Language-Action Generative World Model Gpt4point: A unified framework for point-language understanding and generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.368322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:1877b1be9c585528b99b596ba03226b1678897fce86d085f15b8685b9f5ae721

Observation 09d79b3b-787a-4f41-89fa-4a843446cb97 · outbound

This paper cites K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J., Undersander, E., Galuba, W., Westbury, A., Chang, A.

3D-VLA: A 3D Vision-Language-Action Generative World Model K., Gokaslan, A., Wijmans, E., Maksymets, O., Clegg, A., Turner, J., Undersander, E., Galuba, W., Westbury, A., Chang, A

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.306947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:0bb749bf986def6c1b740a9214b0eb44b804e4ed5a8bc3728fcf6ae2fa806d47

Observation f35d3ee3-8890-43dd-adcc-e608f402d12f · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks.

3D-VLA: A 3D Vision-Language-Action Generative World Model Grounded sam: Assembling open-world models for diverse visual tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.309039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:e46f17f908780f9ca2bfd5209e3e51a700733aadb411715dff612c71d8efef99

Observation a039cc22-5a82-4cec-a9a6-7697428d927b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

3D-VLA: A 3D Vision-Language-Action Generative World Model High-resolution image synthesis with latent diffusion models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.311213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:91183700492dbfddc5d06f6a3c238a6ba6dab40f35a341ab904bce7a9432141d

Observation 25e1f738-1cce-4d0f-a042-dce8a4a66f31 · outbound

This paper cites Playing with food: Learning food item representations through interactive exploration.

3D-VLA: A 3D Vision-Language-Action Generative World Model Playing with food: Learning food item representations through interactive exploration

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.312967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:c520d4790442e7eb3d2fac68334bf1321c4fa3f1a5c9bee1cdf349d75950c9a2

Observation f18b2b71-bf5c-46c7-bd02-07188b7f11e6 · outbound

This paper cites RoboVQA: Multimodal Long-Horizon Reasoning for Robotics.

3D-VLA: A 3D Vision-Language-Action Generative World Model RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:27.272725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:8cc6404a2d00d07b838780d71c2d14c3dbee3f9ed15e9875d87b85c98387b509

Observation bf507b6f-0f39-4043-8a16-6330bdf8de9b · outbound

This paper cites On Bringing Robots Home.

3D-VLA: A 3D Vision-Language-Action Generative World Model On Bringing Robots Home

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.275641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:1cbaa6f78bdb362cefe7d4204599d8d1326e4b9d12e13ba8e1e0025e0f8bb6e1

Observation b80b00c1-fc0c-404e-b5b0-488b7d0b8e56 · outbound

This paper cites MUTEX : Learning unified policies from multimodal task specifications.

3D-VLA: A 3D Vision-Language-Action Generative World Model MUTEX : Learning unified policies from multimodal task specifications

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.325265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:c5b3dd4a80769b9696ad8abfa089171bb05439d89b4b04e751ebe8876c268ab2

Observation a66d1196-1bd7-4435-aafe-0978d513753d · outbound

This paper cites Lancon-learn: Learning with language to enable generalization in multi-task manipulation.

3D-VLA: A 3D Vision-Language-Action Generative World Model Lancon-learn: Learning with language to enable generalization in multi-task manipulation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.316902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:6d1c5ccbe4cb2cf72f182af9c2bcc99b222889677db6d943284c251d194de4c2

Observation 4664baef-2491-44d6-8a1d-9ab5f4c7e00d · outbound

This paper cites and Deng, J.

3D-VLA: A 3D Vision-Language-Action Generative World Model and Deng, J

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.318514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:b0d72427852c6b7b0dc421bd33ea95660295bbec1520c48fb87ea34d346695e0

Observation 5b8db1bf-1f28-4e29-b7c2-40499e9029e1 · outbound

This paper cites R., Black, K., Zhao, T.

3D-VLA: A 3D Vision-Language-Action Generative World Model R., Black, K., Zhao, T

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.320098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:a7d3a770d9c6a9d493105a8c56c627ae7f89dd6330b35db9c646e32f31250cc6

Observation db948bda-a4a1-4bf3-85e3-84dacd27d500 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

3D-VLA: A 3D Vision-Language-Action Generative World Model NExT-GPT: Any-to-Any Multimodal LLM

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.288689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:5369feb031192f408e05abb2de654725ecd6ebf903444b2acf2149a314d91a64

Observation f6750f1f-752a-4b20-9c80-744927bceddd · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

3D-VLA: A 3D Vision-Language-Action Generative World Model Pointllm: Empowering large language models to understand point clouds

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.321705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:cae2a1a4b293f21ac41bbc4a050392b13262e96e0d6e74fbe280e1e0c7dc8b07

Observation 9859893a-4dce-486d-b210-b2128a11b2df · outbound

This paper cites Uni3d: Exploring unified 3d representation at scale.

3D-VLA: A 3D Vision-Language-Action Generative World Model Uni3d: Exploring unified 3d representation at scale

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T18:18:27.323374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:16e5db77ee7405d88c91344226b9e1807d7a2eb2ff8e9bf281e0f277524dc313

Observation 650d54a1-714d-4bd0-a627-b48d4a576c78 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

3D-VLA: A 3D Vision-Language-Action Generative World Model MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-13T18:18:27.301610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:1f0c2224d5211b75a4f2a72387b47de385878499bda2a839d8cab3684cd047e4

Pith citing papers

Observation a41e3deb-0fa0-4796-8413-d4d595ae3519 · inbound

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset cites this paper.

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T05:51:18.508352Z digest=sha256:6b517c513b9f35578e883f4c53e6768b08c238d19dba35be6a0fe1bb2a691b4e

Observation 2d975812-61d1-4c09-b249-95e8f1228012 · inbound

OpenVLA: An Open-Source Vision-Language-Action Model cites this paper.

OpenVLA: An Open-Source Vision-Language-Action Model 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T14:46:35.942338Z digest=sha256:2822409a4d842065644863fc40daa965ae1fc6854d526ed633dfbdecb04c7622

Observation d9f196ac-444b-4a81-99da-629611871e77 · inbound

Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms cites this paper.

Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-12T19:10:14.723298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:10:14.723298Z digest=sha256:15be393dd5a15665667cf254db9e9fa2f70295a0a70a870f2bf3a7237315476b

Observation 9c8a6d4a-033b-4043-b8b2-31194363a28f · inbound

ShowUI: One Vision-Language-Action Model for GUI Visual Agent cites this paper.

ShowUI: One Vision-Language-Action Model for GUI Visual Agent 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T12:10:40.564363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:10:40.564363Z digest=sha256:a2cfaad7d0c63ed722504882d28f706d16af360ba031b781f3f1996bf9502f5d

Observation 83f99ab1-5a8a-4afa-a952-dbb517c0ee36 · inbound

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences cites this paper.

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:37.029873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:37.029873Z digest=sha256:c335e84db7a4c0200d3f4b1fbb8ee1526394f25ef7616e7b11546b1e53737a10

Observation 5a3c3806-e811-4803-a045-c40be984e3f8 · inbound

What Matters in Building Vision-Language-Action Models for Generalist Robots cites this paper.

What Matters in Building Vision-Language-Action Models for Generalist Robots 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:37:50.824010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T21:37:50.617813Z digest=sha256:47dcd6b48c19cb00c8d57d91651aac3bd1ad37df042b7162d393d8bf56413e61

Observation 4312bfa2-5b9b-490a-921c-217ff4d7ae90 · inbound

Bridging Adaptivity and Safety: Learning Agile Collision-Free Locomotion Across Varied Physics cites this paper.

Bridging Adaptivity and Safety: Learning Agile Collision-Free Locomotion Across Varied Physics 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:16.466796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:16.466796Z digest=sha256:c28ecfd093f972a7e2088943e136f3c5ffa4a3023b64abc4805342f5c2c6b4ec

Observation 2d10f1bf-b5df-40ff-abd3-8cefb998b810 · inbound

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought cites this paper.

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:09:34.870777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T23:09:34.805552Z digest=sha256:d0573afd5feed4db546071c3f8c01f2614d361f55ef5810892cc376fa507f2da

Observation ff6da187-d1c0-4359-9aea-5a10f721b7dd · inbound

RoboReflect: A Robotic Reflective Reasoning Framework for Grasping Ambiguous-Condition Objects cites this paper.

RoboReflect: A Robotic Reflective Reasoning Framework for Grasping Ambiguous-Condition Objects 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:11:32.501631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:11:32.501631Z digest=sha256:0915f926d3db3cf41ef49319bddb1fdba9b6501a2ad80078f1a45326aa29369e

Observation 74032957-6ba3-4cf5-b1ba-8fb2cc4d88bd · inbound

FAST: Efficient Action Tokenization for Vision-Language-Action Models cites this paper.

FAST: Efficient Action Tokenization for Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T08:52:31.686474Z digest=sha256:b41f7f55fddcd9435d587691a6e2f81050fe695a43ac8b4d41e82805ad426ea2

Observation 3999e930-803c-42fe-ae4e-cee7ca3af88e · inbound

GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation cites this paper.

GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T19:47:36.637748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:47:36.637748Z digest=sha256:1bd68757f9461b9d2563347b3f85c6c9429c959fd3626afbb357a2d2f3206777

Observation 5b9472d6-1d08-4b76-b0b4-0eed4be64b16 · inbound

Generative Physical AI in Vision: A Survey cites this paper.

Generative Physical AI in Vision: A Survey 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 244

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:00.951935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:53:00.951935Z digest=sha256:53ece8d26d2b402e91a5e9f967382f9811f59936cd9120c81a725342b5f0b6c6

Observation 73dd6cfd-4145-4c5a-acc3-baefae20e169 · inbound

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent cites this paper.

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T22:13:14.945152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:13:14.945152Z digest=sha256:79e9bd997ce9f0a7031201d5483ce0fc873cf229baac2aef3268ffc019095519

Observation 6e311560-b342-4516-8f30-a641bd5605d6 · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:48:48.999906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:524efa1b865d301b459359513c9121a697b1185e346fad884d41224a38a19d1d

Observation 71c95e1a-7b16-4ade-90ca-b0ae8150b446 · inbound

RoboBERT: An End-to-end Multimodal Robotic Manipulation Model cites this paper.

RoboBERT: An End-to-end Multimodal Robotic Manipulation Model 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:17.903956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:35:17.903956Z digest=sha256:545e3a2abb7bf36c9d0dd8fd75ccc6a7af31f41e296b188808d243b63a71ad49

Observation 169c83f6-dd58-4877-beda-9487749e1334 · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:d8e2b5e2d854ac0743b933429fb636c8ea972ee3db01196bc484429e7b34f467

Observation c656c790-2eb5-4a55-b911-72ad627b1760 · inbound

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots cites this paper.

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:09:10.112304Z digest=sha256:4ca6ac3b8f1ab24f6917350592131b771d2b2fa18827ff146e179037328badef

Observation c310e0ce-5b0d-49ce-8724-05321ed2f72e · inbound

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models cites this paper.

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:21:45.165892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T05:21:44.903048Z digest=sha256:231dba19069b5cba6fbd4c07f8062fd94d33bd93ea187f52d142ff5184ab1192

Observation 2dba29be-db06-4c0a-8c69-5dff826f3490 · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.949393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:7f89cc2ab079ebfb170c0aea8a0f60ec8c253a984dbebfcc24782d3daa640153

Observation 0a805211-7440-4d80-a92c-34c9e0a3afd7 · inbound

Generative AI in Embodied Systems: System-Level Analysis of Performance, Efficiency and Scalability cites this paper.

Generative AI in Embodied Systems: System-Level Analysis of Performance, Efficiency and Scalability 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T10:08:58.446528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:08:58.446528Z digest=sha256:18d7ea54667598cf2c6c0b26e91e5209586b7260aa44f9707ca07b7317bcd831

Observation 159917a2-244a-47dc-afd8-8c86931c3136 · inbound

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding cites this paper.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.463384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.463384Z digest=sha256:d4c1d524e7bd75517289de27c69441b52015780c91128656ce776305600217bb

Observation 39e33cb3-5677-4d04-a67a-8035344bc37b · inbound

TesserAct: Learning 4D Embodied World Models cites this paper.

TesserAct: Learning 4D Embodied World Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T05:18:56.748799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:18:56.748799Z digest=sha256:aa1b8f39ff3b589fe3a70aaddcba0fce57b891ed2e27e80d6440bc6300a4626a

Observation 745f23fe-ccc6-4bb1-bd3a-9a7f00823540 · inbound

Robotic Visual Instruction cites this paper.

Robotic Visual Instruction 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:34.175869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:34.175869Z digest=sha256:d7ea2b42589249cfa14d4113a393df5b2af3b938a3ede04cadd398f4febb5747

Observation 5972453f-af9c-4f0e-8a91-c0ec38829814 · inbound

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data cites this paper.

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:55:52.314853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T20:55:52.109166Z digest=sha256:fab7bd80055c3a33a1654d3a03045982ff88f310b1a0e8f31383d699e62f8d46

Observation a857750e-8bfe-40c8-b84b-249362800f93 · inbound

VLAs are Confined yet Capable of Generalizing to Novel Instructions cites this paper.

VLAs are Confined yet Capable of Generalizing to Novel Instructions 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-22T16:46:47.455578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T16:46:05.993833Z digest=sha256:b2544b4fb3e246202c508ffab4016f690b42060ea4e83ccbcbff8169220da083

Observation e7086f8d-467e-4228-b22c-bc9786ff61fd · inbound

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding cites this paper.

DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:11.023031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:22:11.023031Z digest=sha256:4553fa4471e6bd22d4f94e26dc58aacad2c14da8d65c41582e5c16424a33bbd3

Observation 5095a52b-d1a3-4946-b666-ab2ad5062b46 · inbound

CLTP: Contrastive Language-Tactile Pre-training for 3D Contact Geometry Understanding cites this paper.

CLTP: Contrastive Language-Tactile Pre-training for 3D Contact Geometry Understanding 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:06:24.167765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:06:24.167765Z digest=sha256:8f6f3dd81bcc8bbabf302752aad3b04d5b83a6fd3842762267a9f85ec5f97708

Observation 30055d34-7ab0-4b3e-9974-361af1ee686c · inbound

Training Strategies for Efficient Embodied Reasoning cites this paper.

Training Strategies for Efficient Embodied Reasoning 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:04:36.375870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:04:36.375870Z digest=sha256:79151c47ae09233d2e35340660b61bb9cef7aa78a031a054d59dbea695f1d6be

Observation 8cff7998-77ee-4db3-8159-e59f5b25357d · inbound

VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation cites this paper.

VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:33:01.955606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:33:01.955606Z digest=sha256:91e88414cc0d6cf8691a565a38f20fb2d1f48b6b332b6f87d4dcc015438e3d03

Observation eebb4f1c-5abf-4e3f-9713-304fc97b826c · inbound

DataMIL: Selecting Data for Robot Imitation Learning with Datamodels cites this paper.

DataMIL: Selecting Data for Robot Imitation Learning with Datamodels 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:34:17.026537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:34:17.026537Z digest=sha256:d78c1f32139f410d63d0818da7b1e3b07187fa05eded1d73da3305a2a28a99e8

Observation a3ad2b89-8e9e-4b9e-9570-0e14eb8bb404 · inbound

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation cites this paper.

Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:12.777054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:12.777054Z digest=sha256:457ce734803a62c9daf4259c832b8d08ee622cd7c85adbfbc57fb9ad7bb4cefe

Observation 6d2c2cdd-fcc1-4865-ab40-e0e70f9af25c · inbound

FLARE: Robot Learning with Implicit World Modeling cites this paper.

FLARE: Robot Learning with Implicit World Modeling 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-17T15:59:08.880467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T15:59:08.846629Z digest=sha256:c1d0103cd074ee1ebe7b6523a0c85f8e475e7ece801e2f2a8fb04ef8a9cd603c

Observation 058a3be1-e52f-4b8d-b071-11be389466cb · inbound

ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning cites this paper.

ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:17.804686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:17.804686Z digest=sha256:a60ec796c3d1f99ba343494aeca9ad8f33b4582f2cfdb6e60f2117497bc602a4

Observation ae070d1c-4485-46b8-ab14-35feabada817 · inbound

Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better cites this paper.

Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:25.445765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:25.445765Z digest=sha256:1b3e7907b73f7c24fff62f3c924fb04c003d0b0686681fe6e59a2cd5ff3feeb5

Observation 38e1098c-acb6-48fe-a4e1-8b5bb8bb162a · inbound

DSG-World: Learning a 3D Gaussian World Model from Dual State Videos cites this paper.

DSG-World: Learning a 3D Gaussian World Model from Dual State Videos 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:51.760825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:51.760825Z digest=sha256:56f70407be2c42c4b372846a0f37cc4a4fa6365f670aaf93ac917dca34f36e43

Observation 2f26de07-ff57-449c-b422-21cc47b85286 · inbound

Real-Time Execution of Action Chunking Flow Policies cites this paper.

Real-Time Execution of Action Chunking Flow Policies 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:18:51.738931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T14:18:51.613045Z digest=sha256:7653b0c105421935feb54f8fff932a0acfc3ddfa5037541539ad0b7e66fe4552

Observation 7601f167-343c-4e52-a9cd-4ace54fb07e0 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:29.901136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:29.901136Z digest=sha256:016b42bf16da449ef571de9b7e6ae4314ed56834a6b614b21d123c9b20005d9a

Observation 28d62628-12bb-4d40-bd52-e4a219ccefa7 · inbound

CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding cites this paper.

CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:03:56.657488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:03:56.657488Z digest=sha256:b5b12927b9c2c42ef4c9f72c03bd5bfd85675e438571c269b9a82f6fd1bea98d

Observation 064c76d7-2151-4ef2-b5fb-b740272f4835 · inbound

GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation cites this paper.

GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 66

Resolution
malformed identifier
local_arxiv, observed 2026-05-25T07:50:29.186539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T07:48:33.832017Z digest=sha256:93f1612d4e09380d6d54dd93638be451836b6eec6d4fbff2f4b14e41349aed42

Observation 598224b6-42f0-4c7a-81d7-4efecd310df9 · inbound

DyNaVLM: Zero-Shot Vision-Language Navigation System with Dynamic Viewpoints and Self-Refining Graph Memory cites this paper.

DyNaVLM: Zero-Shot Vision-Language Navigation System with Dynamic Viewpoints and Self-Refining Graph Memory 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:11.784305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:48:11.784305Z digest=sha256:a3af54fee146fe4f242c7d5b87ccda53a86cd01d6afad875122c7acdc7122048

Observation 336692a1-eda4-4520-9bc3-b1f0b07728e7 · inbound

Beyond Syntax: Action Semantics Learning for App Agents cites this paper.

Beyond Syntax: Action Semantics Learning for App Agents 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:32:08.750282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T07:31:33.690689Z digest=sha256:c6ab6f9f952508c36c0825f1c76494784d645117d406b3c9eee55f1408eca307

Observation b7875404-7e74-441d-8324-22f30dab737a · inbound

WorldVLA: Towards Autoregressive Action World Model cites this paper.

WorldVLA: Towards Autoregressive Action World Model 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T22:57:07.883617Z digest=sha256:6b16238cc91d6ce6e328ae28d3f22e67d7b0528b4575f1fead00075d75a80bb5

Observation dc0d5e19-50a6-4363-b5c3-eef93dbdf956 · inbound

A Survey on Vision-Language-Action Models for Autonomous Driving cites this paper.

A Survey on Vision-Language-Action Models for Autonomous Driving 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 161

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:04.898771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:31:04.898771Z digest=sha256:17c19dbffff48dd8956c9cf4cad8ffb7471e2a301f3a59ecca9d8ecb8e563101

Observation 21d4acac-a0e5-447a-8c71-a4e49754acb8 · inbound

A Survey: Learning Embodied Intelligence from Physical Simulators and World Models cites this paper.

A Survey: Learning Embodied Intelligence from Physical Simulators and World Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:08:53.614836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:08:53.614836Z digest=sha256:1843d878fc46c66264a40c1740415d5a04482b4db0573e17fef249fe2a15b980

Observation 50648549-db24-47ed-9cd1-4360a95c363b · inbound

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers cites this paper.

VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:25.998032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:25.998032Z digest=sha256:5ac149f4ce9231b23ed46e047163340a3ad10ba202c1636876b1076ea5db75a0

Observation f6560375-45a2-4982-b1a7-492b871e2c5a · inbound

MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details cites this paper.

MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:19:44.253040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T21:19:44.144633Z digest=sha256:ef22545bf4f73e7d2b80399c39649d4a6a172a5afa1021af89a451d379b79a8e

Observation c0dbc718-15b3-4fef-ba94-993a9eaf9604 · inbound

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation cites this paper.

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T20:02:33.187976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:02:33.187976Z digest=sha256:edd761a565e947e7341c29360e314fbea1d9b15e59c37d06c9d6a33b2181e195

Observation 2c3c555a-d500-47c1-aa25-5d7a32955703 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:42:41.496742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:5b14437a756b73dae99e5f7d3e2db9a962f003f9024d851a2bacb3868e6dc315

Observation cbe25985-f876-4a60-ab34-47c1b18f445c · inbound

Reconstructing 4D Spatial Intelligence: A Survey cites this paper.

Reconstructing 4D Spatial Intelligence: A Survey 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:02:28.459549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:02:28.459549Z digest=sha256:87a2c88a61bedfbfa4cee7f1d3c07c0c2df16f4ee74eefe649debc42ecf2fdbc

Observation fae05599-aee1-4385-b567-8459311861d5 · inbound

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems cites this paper.

Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T12:39:42.254644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:39:42.254644Z digest=sha256:bf0f324ed309c4d0f0e951ca9bdf366014097e574d5f21c5f130fb6d27359d2d

Observation 72328781-969b-4e9c-89dc-4916df0347ee · inbound

H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation cites this paper.

H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T10:48:19.186087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:48:19.186087Z digest=sha256:1c67140f19e46be7ef97612c9015da43a40851f6957d92fede95421ee5ccec4f

Observation 072f4e33-2712-412c-9145-b74cca030574 · inbound

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions cites this paper.

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-05T23:53:54.204000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:53:54.204000Z digest=sha256:193ccdb52b9b8d56533f31b6f11fe73c2ad8bd423c950437828110c754a70293

Observation cc1f00d7-649d-4957-920c-df764718e823 · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.647280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.647280Z digest=sha256:f977707035e84b2254d20b1bd433f764db3d39258bda57d7884e03d997830c92

Observation 6b6ffc71-a3ec-4fc6-a7d2-f446f491920b · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:48.488263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:48.488263Z digest=sha256:5b8553ef9682adf8fda349b0c33fe5dd422552955e9bbebc0327d276d08559cd

Observation 2cbb195e-7fb0-4695-8fa5-abad61c589d2 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:46.276434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:46.276434Z digest=sha256:0f572cdc50f445393e14631632798dfbf3d43af78f7269b4b9f963f2f77c58af

Observation 71c88cdc-4890-4d1d-9fb9-ab781a9d3a60 · inbound

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning cites this paper.

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:56.709921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:56.709921Z digest=sha256:d9a97c522a3dc9c068c438599f32d79b5b7f407312e3eb4eb82873f45cb94b6e

Observation 8bc92a44-f2b0-4e57-97a0-8e62412361e2 · inbound

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy cites this paper.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.824487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.824487Z digest=sha256:b64d9aa491208d2adce98e6bbcbb918073502fae0156da90d3f3b6c22a5270d1

Observation 7b8aa446-dc3f-4c22-bed2-46c7a8f76ea9 · inbound

Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety cites this paper.

Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T11:54:15.963740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:54:15.963740Z digest=sha256:095722a02cc1638b003932e5b09ff769e48347d041cd34657d2332a82dc66dc7

Observation 09dbc511-d9ba-488f-bf02-af7993566b9b · inbound

Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots cites this paper.

Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T16:42:27.860682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:42:27.860682Z digest=sha256:144329778fcc7d1b01e5493d4aa3f67236e2ef605b451a69d6b24ac9845e13a9

Observation 0bbaa614-04ae-4443-bada-3041e9009be2 · inbound

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective cites this paper.

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 228

Resolution
unresolved
no resolver link, observed 2026-08-04T19:39:24.137294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:39:24.137294Z digest=sha256:5a5486b51d5c46bbe11af2fddfc46794a26108634b37236e476465fbd42f0f86

Observation 55e265b9-81e5-4d9d-a14f-a7bbc814150a · inbound

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models cites this paper.

QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:49.182561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:34:49.182561Z digest=sha256:80745bbeaf2a19cfa18ae6bb0c90aa6ed800b834f7febecf139a3a44f9d7960c

Observation fe397185-f5b2-41a5-9be7-3b448b93b77c · inbound

BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections cites this paper.

BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:42:07.535660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T21:40:51.426489Z digest=sha256:486d4ad6e0bef175bc95083029ed8ce1815c00dc54b38030bff240e418a2f1c7

Observation 91e3baa7-afcb-44f5-84ed-f9500e5b0f16 · inbound

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency cites this paper.

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T19:22:04.165316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:22:04.165316Z digest=sha256:5b1053a17566aa8e145612769a7bb8419ed2767db0be9ba203948ade1f4e39b7

Observation bea9803e-44d8-4f44-8903-8701d55a4814 · inbound

VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer cites this paper.

VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T17:38:20.829550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:38:20.829550Z digest=sha256:1899f15c1bc3799ca211cd815159d13f8b4f08871753bb255ee6da06ec69e267

Observation 662f0ae5-f798-498d-9708-7c14a6018916 · inbound

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision cites this paper.

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T14:00:20.759135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:00:20.759135Z digest=sha256:29a725a0c00cd8c8b4d2a2ad04c4accd6f30c593726ca0f78a3a97b9e30cdad5

Observation c0db86ef-5a47-47d0-8270-c69c21a01831 · inbound

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation cites this paper.

GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:10:25.157527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T17:06:34.398973Z digest=sha256:3c48feeef5457caa2f5903936719120ba66371ff741f743f5b2932133a86b998

Observation 96d4c22e-17a2-424f-9628-7540bcd9a511 · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 150

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:08:02.111684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:2ecb9ec7951a7471b386250d83c7ddf6332a00e34f8c90d73edf1720420d9e40

Observation eed17cf9-dad7-460c-a9ee-fb795ffe05e1 · inbound

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models cites this paper.

AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:02:24.967630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T06:01:30.803128Z digest=sha256:e2a788073ca82155e0269e6be04ceb929b1eaabeff4b18bbc952321e4d654d7a

Observation 68c16829-e20e-4cf0-b57a-7cc02c56c00c · inbound

Learning Native Continuation for Action Chunking Flow Policies cites this paper.

Learning Native Continuation for Action Chunking Flow Policies 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:40:08.665758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T12:38:26.522838Z digest=sha256:305e7ed14af1fd5e5383df3c5d1c887c9d9bf1f248c3f8db33061d66674dc201

Observation b2b36cf2-8633-4e4f-916f-cd52925ae079 · inbound

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models cites this paper.

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-15T20:20:17.659399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:18:31.988002Z digest=sha256:620fc0966584b2f2f023a0e4668ed96de0c61da0d08389dbb9785d52ba762a5b

Observation d1a4d5aa-079b-4c3b-bb8a-141773ed4568 · inbound

Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks cites this paper.

Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T21:12:11.164921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:12:11.164921Z digest=sha256:40cd29e76cf91f197accbd9e2a6c1921bceb130e7335ed69112fd96abdc6f1c6

Observation 110a41a6-daaa-4e9d-af88-49196de51f9d · inbound

VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies cites this paper.

VLA Knows Its Limits: Adaptive Execution Horizons for Robot Policies 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:06:10.804392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:06:10.804392Z digest=sha256:65b1068e5fa6887fedc6f2fca1645fd5b33f2ebc6bf4ee6b7f7595d87e33c782

Observation bcdc9a41-8ff9-429c-8b92-e08d87448d9d · inbound

What if? Emulative Simulation with World Models for Situated Reasoning cites this paper.

What if? Emulative Simulation with World Models for Situated Reasoning 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 124

Resolution
unresolved
no resolver link, observed 2026-07-15T13:51:30.008232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:51:30.008232Z digest=sha256:00a678beca9132cb085ce66ded454c388b0dff62c3d95fd79f1aff2c117daa03

Observation 3d9a6d77-5115-48e0-9b8a-a9e383dea878 · inbound

ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-Making cites this paper.

ThermoAct:Thermal-Aware Vision-Language-Action Models for Robotic Perception and Decision-Making 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:48:24.816800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T00:46:03.360770Z digest=sha256:0b02ed75424b316a704f9061d91f994790e8ed42de197ce3c52cdc4d18d0371d

Observation 5ba0b4dc-7668-4075-93d2-5be0d3cfad58 · inbound

Redefining End-of-Life: Intelligent Automation for Electronics Remanufacturing Systems cites this paper.

Redefining End-of-Life: Intelligent Automation for Electronics Remanufacturing Systems 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 183

Resolution
verified exact
local_arxiv, observed 2026-05-13T19:23:09.512252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T19:21:34.849729Z digest=sha256:6319e296d27f57ff002a26a84d34a28b28d4db11189133963f588782b4a9a526

Observation 85c95fa6-d8a6-4794-bd6b-08754ebac289 · inbound

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment cites this paper.

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:27:12.286456Z digest=sha256:41736159e32caffe9c90058273825086be65860edecd038e67baa97e3632ef54

Observation 4922074b-c631-45c4-a225-c19cc523f6e9 · inbound

Action Images: End-to-End Policy Learning via Multiview Video Generation cites this paper.

Action Images: End-to-End Policy Learning via Multiview Video Generation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:51:05.206602Z digest=sha256:c9d786063805967b70637c68c16cc1d087a7fa9c05bda5f69e06ea1c3ba57fd7

Observation b9ec0899-39dd-410d-a5d3-1e1a2e9d8bf1 · inbound

R3D: Revisiting 3D Policy Learning cites this paper.

R3D: Revisiting 3D Policy Learning 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T10:57:54.273960Z digest=sha256:f846616576a29da88e0d82c4c2fd4d8771011f4eb0999eeb2086c8728dab5493

Observation fcbf9ec7-030e-40b3-bb66-c11208c6e426 · inbound

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities cites this paper.

${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:42:34.409651Z digest=sha256:fa0bf05397879eafe714f8fa4ceb50e6f91134d5fe6a85927451bce9811ad976

Observation 705b8e60-503a-4d9c-a00a-3b90167ef080 · inbound

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning cites this paper.

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T05:06:38.517652Z digest=sha256:d6d1da499970341685f88c5d2ab9601cdfeb8060ec3631cd3e578628fcaf618c

Observation 3606e2d6-207c-4c55-a974-c29958e5db91 · inbound

ST-$\pi$: Structured SpatioTemporal VLA for Robotic Manipulation cites this paper.

ST-$\pi$: Structured SpatioTemporal VLA for Robotic Manipulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T04:46:21.784769Z digest=sha256:c2da59331f0318c393099833665a856a7739de1bc8bbecfa66f158703c84a4ee

Observation dcd18759-44dd-4e78-8fd1-56c7538958a7 · inbound

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model cites this paper.

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T11:45:18.081248Z digest=sha256:5ee6f14cb4bb934aff5bf268d547c3c550420011a846ae43b6fef2223dd6b3e8

Observation d9edb622-9300-427f-b3b8-ba68b530cf62 · inbound

Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines cites this paper.

Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T11:30:49.414976Z digest=sha256:c74d8d0538c73cb78234a91b2951d27676dc763a89672bf1ab42036f76353c8c

Observation 07d746d4-9c3f-4554-baa2-9b13bd8a0739 · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T05:47:17.494531Z digest=sha256:12074a137f4c2f310271e66ba921e53c1d84958efb2a2e26f4710a02633932db

Observation 2dd1db0e-1f53-4ca7-b3c5-3ca150d7c0ec · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T03:00:26.352130Z digest=sha256:bbe43a22a77f7c25924c3e22bf1bbb3c81a5b3f9205cfef92e5123aa56ae3338

Observation 890f35fa-c7d8-4d09-8233-f94583bb73b1 · inbound

Affordance Agent Harness: Verification-Gated Skill Orchestration cites this paper.

Affordance Agent Harness: Verification-Gated Skill Orchestration 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T18:40:53.380512Z digest=sha256:2a5b7e8830a56647dee84380b07a447d2e82d45879248fde696fc478ee1b2c18

Observation d61860be-bb22-407c-9489-0dda0ebb58fd · inbound

Affordance Agent Harness: Verification-Gated Skill Orchestration cites this paper.

Affordance Agent Harness: Verification-Gated Skill Orchestration 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T01:55:07.248106Z digest=sha256:ac9bbc1e368da2a2dca94b97943a98600c08ef4bcade36ef611b2cdb325fee0c

Observation b9f682cb-cc59-4279-872f-df65e9dd9ae5 · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:2361acf27953ab14c8b957dcbcf26bdc08dd72f8fa65287b86f0c0401f85632d

Observation 5f242772-84dc-4baa-9c5c-4cfbb9bd33d0 · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T03:39:41.090350Z digest=sha256:804b24671208989c73c684c8763e16609c5cd96d701deeeab55df7de2305c781

Observation 9a95271f-d267-452b-90a4-e6e7f090c6eb · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:53:54.608425Z digest=sha256:ac8423758e8dc170b6a4261fb378aa148082a7730616406e2605f36a05c0f20c

Observation 717ff208-ad5a-4ddf-9604-cf969d7289fb · inbound

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy cites this paper.

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:19:49.789881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T06:16:48.180290Z digest=sha256:b1ea0df76b35022791083667ee6648afad39450feb88fb3d5e28baa568a329bc

Observation d462be3f-9b37-4448-b91f-275895efa924 · inbound

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models cites this paper.

VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T05:09:21.028373Z digest=sha256:46b174db18dd60ffdb8ec92b55e89816eff20fc5f1a26fe5ff836826f40f7000

Observation 821e4050-6e8b-45fb-aaeb-19d9b37ac222 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:b119670f3c26aa2786bcd4b0168b1c36739500e5f5251b400be7334a2faa5cd9

Observation 11ba71ca-a685-4913-9c87-de0bb4496f54 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:52.756668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:52.756668Z digest=sha256:dacc744f0f82548013f204cc32f0db24eda900b834da5154a8ac13e79131f6f5

Observation 31f76f2d-395f-4460-b1b8-0475b3c65bcb · inbound

World Action Models: The Next Frontier in Embodied AI cites this paper.

World Action Models: The Next Frontier in Embodied AI 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 270

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.368966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T05:01:16.802019Z digest=sha256:4eeb23434dfc81799427a0bb8062422d27570d8d5f404962bcb007fcafb33c67

Observation d652e784-ca52-4d06-9e2c-d8711ced4f10 · inbound

Towards Robotic Dexterous Hand Intelligence: A Survey cites this paper.

Towards Robotic Dexterous Hand Intelligence: A Survey 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 112

Resolution
verified exact
local_arxiv, observed 2026-05-15T05:45:05.971120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T05:44:23.945064Z digest=sha256:b094259e792cc606e761404b1dfc44de85519bddfb66480ae17523278337bed4

Observation ce8d7db6-c8d6-40a3-960d-686f48c24b00 · inbound

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model cites this paper.

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:25:46.235324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T21:39:15.662180Z digest=sha256:9a0dc05b67ad51064a6d70bb7afce1ae592aee439369a26575c9b305a8372f0a

Observation cdc1cc04-548e-49d3-950a-9ed717e5c448 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:43:17.308845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:cdb22ef6c12f82fa2991c60e274faba24eef424335d0ed93468595ac32768cec

Observation 0cfabf81-4168-472c-ae58-194d1fa6af2f · inbound

ECG-WM: A Physiology-Informed ECG World Model for Clinical Intervention Simulation cites this paper.

ECG-WM: A Physiology-Informed ECG World Model for Clinical Intervention Simulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:23:16.862717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:21:16.284834Z digest=sha256:87dc660478060b65043d4abcefe8ea8e7ad5f24537bb391c218f9edeaddef1d2

Observation 404c8e18-f088-4e1a-9456-61cb54ff6eb9 · inbound

GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation cites this paper.

GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T05:03:57.849949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T05:02:54.167808Z digest=sha256:91d0f254de0ee7eb759227a877eba2db7f202170bcc74c69f9aaab2fd876efaa