Pith. sign in

Paper Citation Record · LEDGER

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control

As of 6 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2605.21862.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.21862 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T06:19:36.625269Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact31
  • verified fuzzy13
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b4af943-3622-40a0-9045-eaaa66aaa8fc · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.608332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:15481390503464a27caf89efef2b71f044b5fcde356be1951bd65df3d8aefccf

Observation 41aac3fe-8c8b-4d22-a717-537cb41b89e4 · outbound

This paper cites Zero-shot robotic manipulation with pretrained image-editing diffusion models.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Zero-shot robotic manipulation with pretrained image-editing diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.730016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:96b9076b139ee14c4547384e70a33384f59fa96225b31a9617368d72b0818691

Observation 6321eaad-2be1-4f2b-836f-daf98a98081a · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control RT-2: Vision-language-action models transfer web knowledge to robotic control

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.695330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:269d655e33ace1d4f2161b655bc03acdb35e999d279890c0c0f2d3cfe4327e9c

Observation 9c9e8924-05c9-4d17-a779-44fcf326b546 · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.594971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:a21c1fec85bece9c42d0815657d55319268317540ea51695a05bd470fcb22456

Observation f34672aa-8a32-4f77-8f47-a7b48fc17bcf · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Diffusion policy: Visuomotor policy learning via action diffusion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.683633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:b5b9e3de5b0db04f0a80b4fa2dc62b4eabb94a2ad2e265f7388c22f6098c7abd

Observation 9c43ed15-875c-477e-9062-228970b3fb27 · outbound

This paper cites Embodied-SlotSSM.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Embodied-SlotSSM

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.430495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:b3b55d131550efacc13f52d362013ba8ea0d303edc608933aae7bfce5e202967

Observation 787d98c8-0caf-436a-96cb-99c83ea64977 · outbound

This paper cites Tenenbaum, Dale Schuurmans, and Pieter Abbeel.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Tenenbaum, Dale Schuurmans, and Pieter Abbeel

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.699157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:f3d9f037ef5093ebcb77adcdd8c1f1963d4cc1007700a0b003f79762b03079d4

Observation 547833f9-5ab7-49e8-aa61-ffba752bf852 · outbound

This paper cites RVT: Robotic view transformer for 3d object manipulation.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control RVT: Robotic view transformer for 3d object manipulation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.679878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:8ac6dc6643f7a6c68245c841a621e404ecabb027c99551089cc05e7d23c41b3f

Observation 890f44b6-bd64-44d4-98ae-386abc5485fd · outbound

This paper cites RVT-2: Learning Precise Manipulation from Few Demonstrations.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control RVT-2: Learning Precise Manipulation from Few Demonstrations

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.571277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:ffc38b808d047d0d993785134f8823434b19197769474e0d7318de56e8a221f7

Observation 1f7d4da7-e2ce-4d8c-bc68-2d3a8d95bfc4 · outbound

This paper cites Mastering Diverse Domains through World Models.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Mastering Diverse Domains through World Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.496219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:4c26db8da5bddda583c70cb72af73fb52057c53d519adbc1def2081173e4faaf

Observation d62711b2-fdd5-432f-9682-ce01363f15d2 · outbound

This paper cites Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.466796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:12e795efdf14ed618aecac1b15d6b5cbba97a261f3225b543a03dbae464a6b30

Observation 58331888-189e-430c-a99a-e94316e8943f · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.461423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:f205fba024af1de4edf2a1a310efa10ec1dab130029288542117d1f2e66b1ea9

Observation 269d1c92-213f-489c-af37-210e436168df · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control OpenVLA: An Open-Source Vision-Language-Action Model

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.446926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:52885d3c1352a4c4491ba0590c0bc000221dae331ec70cdd5395746a8d2ffa0c

Observation 7be3712a-c6bc-47dd-8b30-9a7d870055d8 · outbound

This paper cites Tenenbaum.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Tenenbaum

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.722533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:7f5a595eb4924fa5a2f92fa17c54c58b96880f567437f936444bf74e003600e6

Observation e237c64d-1304-4124-839e-68cae4e12638 · outbound

This paper cites HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.601738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:c9426af171cc98ec2587c0c8d5f84ccc7de27830a8a4b2d203753dce50170493

Observation 3517652f-03e1-4d0e-b3fe-fa357c5bec38 · outbound

This paper cites Grounding Image Matching in 3D with MASt3R.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Grounding Image Matching in 3D with MASt3R

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.565270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:78b3e75f17381344a8d4b28a53eb3b25e509b5e4d0c6e929f44910b121203bbf

Observation d02ce6f4-6bcd-4fea-99f9-6b1a3edc0caf · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision- language-action model.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Spatial forcing: Implicit spatial representation alignment for vision- language-action model

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.424582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:c025e9a9b7ff921a8d53b1c1fadc21696902b3a4bd35c0d64d6341508b5fd466

Observation 472685a3-1123-4af7-8378-e0579f1e9710 · outbound

This paper cites 3DS-VLA: A 3d spatial-aware vision language action model for robust multi-task manipulation.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control 3DS-VLA: A 3d spatial-aware vision language action model for robust multi-task manipulation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.725971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:af04b0763694c149b1895960ce18317520317f2f992602171503b913870a7dd3

Observation 88162eca-9fb0-4881-8637-1cb534b1082c · outbound

This paper cites QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:41.992691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:46525c9bde00d7768a74c01d99d7cce76a3ff240666fe43c15ab481dc3b1685c

Observation acdc685e-6e97-44f7-a8ff-508261dc1bbb · outbound

This paper cites HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.501618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:ab093785727330facc09f233fb917925b9fbb3c95ba24854e0085a066c1cda61

Observation 0ca4ce47-d39e-4fd4-b2aa-e2506905621f · outbound

This paper cites an unresolved cited work.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-22T06:21:11.691373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:98085d676f07a0803d425989bdd2cd62f836215cfed1c769665612718b68eade

Observation ac988911-b83c-4a8c-808a-b02abb33336a · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.436146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:3f24d25e8ab1ce298823935d97d2f70595c6952dbf001f3a0e03797ab19eb36f

Observation d22eaa47-ad22-4abd-92e8-fe05abb6e1d7 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.703305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:7ca626c32385b55ff74d16e910375edb0fc6f610e2fe8b129fee99243b2402f3

Observation c281c2df-4646-4c05-868b-6901462a7856 · outbound

This paper cites ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.478952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:a6d2f50faf2cbc788f061ebf7cee06a725cc8d2cc7cf4bab220d21c563b57dc8

Observation 8e5ad554-843e-47de-9721-19f2ee3f4dda · outbound

This paper cites RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version).

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version)

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.491006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:aea422952c664a55f3b000f36d0050dd8868db4528edf5ccd9a43abcf5550447

Observation e78f844a-9646-42d4-857a-dcbb3e0cd5d6 · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Octo: An Open-Source Generalist Robot Policy

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.559150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:7bf153c2718c92353a80f03280cc22a1a87c85b5c69c312b1102b977d3ecd09c

Observation 0cecc5cd-f439-41cd-98aa-b64294b6fe75 · outbound

This paper cites Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.547729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:3e1195cf85fdcb31c7cfbb7eba6b73fe0b9a206dacb69e16b97074125bfa4fee

Observation 96f59bfe-79a7-4bbf-bf34-17f734923c11 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.576633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:def875acfc6c43ec0618ba593759f28cabd15d19a0a07e346d5a0ece29597c47

Observation d887257e-67bf-484a-9866-0ce9d80db026 · outbound

This paper cites LingBot-Depth: Masked depth modeling for spatial perception.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control LingBot-Depth: Masked depth modeling for spatial perception

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.707248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:4b1c28ff3ae8718c7b7d3062057cf2b2c7bb4d813bc44e614de4f2af60c30aba

Observation a3ee32f6-99ec-454d-a117-600dc8fd5af0 · outbound

This paper cites MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.524426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:cd3703cd3dea97fb1460510bdc17c4210746101be60d7e069e981fd98253bb62

Observation 12e7e65e-93f3-4bb5-8f2b-5226cc30d3fc · outbound

This paper cites Perceiver-actor: A multi-task transformer for robotic manipulation.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Perceiver-actor: A multi-task transformer for robotic manipulation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.687414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:b0ffbfe59c41a048b09512bb12f70002bfcab71ed29261de206cf89f7dee2c97

Observation 578df43c-2afb-47ab-92c1-72d2621b825b · outbound

This paper cites Masked depth modeling for spatial perception.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Masked depth modeling for spatial perception

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.518794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:648316400996912c76e8a7f9e03b386c388e2e5a9d67e83aede0479a190932b3

Observation 1faea9da-a788-4ede-8cad-2fbd5945db07 · outbound

This paper cites Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.589464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:8cf852a0bf9c9ae3ab2c17883545f91261a55c7cf81715ca2ad39b79aa114f72

Observation 135b38bf-bec1-4d3a-b423-7176233ac7b9 · outbound

This paper cites Continuous 3D Perception Model with Persistent State.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Continuous 3D Perception Model with Persistent State

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T06:21:10.472626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:55c7e3e5283b944718ac997eb58732970fb49ac445eb968c3f62a2a34a0f0c10

Observation b0fee89f-9cce-4e95-a2b9-5bdbb17ca504 · outbound

This paper cites DUSt3R: Geometric 3D vision made easy.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control DUSt3R: Geometric 3D vision made easy

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.710952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:a45b493f12017761d57ec5b7095e6c1771ff8702ab08099cce75025bcbe420ac

Observation 0abca758-46eb-4a5b-879b-d6435cfb7421 · outbound

This paper cites π3: Permutation-equivariant visual geometry learning.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control π3: Permutation-equivariant visual geometry learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.715360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:9f5f2b8c013cdbcfaef88b226701637127783baedcd863d32f285fa6cc8f70f6

Observation 479f12e8-81d6-49ad-afc0-cbe64a856f4c · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Any-point Trajectory Modeling for Policy Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.541114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:42ea17abbbf35105787c4326ce5567348526349706158a8d75b23143b8e1961f

Observation 2cd015e3-e856-4b58-9457-904d05bbcd49 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.535310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:7388df04408b1b8b7447ff4d9e3959b629a80208715df32d0a4dc19b0a7df350

Observation b305d754-4fcd-4bd4-b506-59363a0340e4 · outbound

This paper cites Day- Dreamer: World models for physical robot learning.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control Day- Dreamer: World models for physical robot learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T06:21:11.719057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:6ac926afd478756097f8cc67969b716ab5ea7c3bf4acfec02ceeee87bada9fbd

Observation 65d13b52-8b43-4fbd-ba34-c742cb265758 · outbound

This paper cites A Pragmatic VLA Foundation Model.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control A Pragmatic VLA Foundation Model

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.529535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:c9403e9c2ed3ab03545226340edda3a8b4d397a2fcf4f1c342ee078e38197772

Observation bf068ccd-44a0-49f5-9a66-38f6484c1cff · outbound

This paper cites AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.512927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:518cfa1eae6bfd826aca7f0b4faaf7f665cc96c2c3cb0d39e2b5d6d4a15671e7

Observation f30414af-78cd-4bf0-a9a4-5e3dcc8270ec · outbound

This paper cites 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.484831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:8c4ba57254e921e80a240f3ee0bc2843986e91382825d5189abc76fbd3104c7f

Observation a7ccadef-7dfe-4e6e-8a80-ae221c8a6dec · outbound

This paper cites UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.582695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:e1c226ab522e03f39d95d52b8a0133ffd49f71d84d3982a1522827d314e1934d

Observation fef3418d-5137-47cb-b0f4-afc761ec4358 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.553383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:aaa47edd20b4d77443586c079151e450df1db62e972523915a99f2692b6afcb0

Observation b3fcbc44-cc6d-46b5-937e-7c36c648c063 · outbound

This paper cites FLARE: Robot Learning with Implicit World Modeling.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control FLARE: Robot Learning with Implicit World Modeling

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:21:10.441559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:4701603c60e1ff4f909f1a1ff0a9b4eaa54b21c7f680c33fa76f556aa3a05b45

Observation ef1d76e6-e3de-470d-bd0c-f96496c23003 · outbound

This paper cites VLA-4D: Embedding 4d awareness into vision- language-action models for spatiotemporally coherent robotic manipulation.

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control VLA-4D: Embedding 4d awareness into vision- language-action models for spatiotemporally coherent robotic manipulation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.453584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:19:36.625269Z digest=sha256:068aa55c182baf913a93bd99259f500454ad005072a5168a36ec4148688243aa

Pith citing papers

No inbound Pith citation observations are available.