Pith. sign in

Paper Citation Record · LEDGER

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging

As of 17 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2607.09818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09818 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T15:19:27.489381Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b267005-5ab1-44ed-97b8-954325acfd22 · outbound

This paper cites Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:2698d90d4fa9cfe967bbab16f9f5276db230eb79b64e6a07f80c547005e94c59

Observation edc2f0b6-8ead-423b-8773-7cadbacbc3cd · outbound

This paper cites LLaDA-VLA: Vision Language Diffusion Action Models.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging LLaDA-VLA: Vision Language Diffusion Action Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:4a43043f45cb455ebf67252fb7eeebf81cf6ae543d154e8fa2f70ce27d0a31b0

Observation 859a8458-897c-4e3e-b624-649a9597d353 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging OpenVLA: An Open-Source Vision-Language-Action Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:218f6c62b60bbcff4e2e57aa91717e9acc10b8c1de95b5dc3826fbb23fade5da

Observation 1e9b0bf7-589f-4912-a635-7616040e92fc · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:b9bb7c4ace224930448df8136678e5f63a7ee43711d7c48ba4b50e44b03a6d51

Observation a1a644a8-9ad8-4ef4-9c83-6870ec0aaf79 · outbound

This paper cites SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:a4a885cfbb4640215d87cedcb1f3c67c2d933c4922af9e3c4370378f63204178

Observation 369bdd64-4844-48fb-95e8-985d5619ca7b · outbound

This paper cites Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:2faac4136479625850b564b25d511671e4c15c0ef775ad17b23a89cac12ce99d

Observation 1259e6dc-5590-49af-98f1-7a7f4d5846e3 · outbound

This paper cites Vla-adapter: An effective paradigm for tiny-scale vision-language-action model,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Vla-adapter: An effective paradigm for tiny-scale vision-language-action model,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:a625f87a86c1ea42e25106c5f4ae9f80fa14c4a04517fa7b1037b84fc0c856a0

Observation 706c131b-1f93-4bdd-85d6-3befd6e6e406 · outbound

This paper cites InterMask: 3D Human Interaction Generation via Collaborative Masked Modeling.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging InterMask: 3D Human Interaction Generation via Collaborative Masked Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:62db7a3b38b07fd517e854dd06eb6bb8c0dae015850a7f145f8a2f52835b99f4

Observation 07a52544-01ca-4b0f-9804-8ac9042403f7 · outbound

This paper cites Libero: Benchmarking knowledge transfer for lifelong robot learning,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Libero: Benchmarking knowledge transfer for lifelong robot learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:6c90a8414fecbdf7bbb0421cb0fca077313b2e83f5cecfa20fb7ecea9eda33a1

Observation 79606a1f-34e9-497b-bbbf-e86757255dc6 · outbound

This paper cites Calvin: A benchmark for language-conditioned policy learn- ingforlong-horizonrobotmanipulationtasks,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Calvin: A benchmark for language-conditioned policy learn- ingforlong-horizonrobotmanipulationtasks,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:59e3f24bd5c718d25206a0c53ecbd8c1ae58bf4200b1929b29924492e99935ca

Observation 00ffe3d2-333d-4c64-8e7c-9990893e3afa · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:b6b69867dbd33be99c48213042a30a26a914cfe67b50a04e0a69a9544f099aa1

Observation 2c193481-5430-49ad-9a74-b8a516e399ca · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging RT-1: Robotics Transformer for Real-World Control at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:61a06fd5b33fce6021c53de4f007714f346121a4b9320583bf86fd560bd6916e

Observation 3b82ab7c-12e0-4319-aff2-f7a6811707e9 · outbound

This paper cites Rt-2: Vision- language-action models transfer web knowledge to robotic control,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Rt-2: Vision- language-action models transfer web knowledge to robotic control,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:6483cbafb0675d6ec0b8d3d6a0928ff30443f23dbe66097f9829f7f528ed8515

Observation d20bc52c-5272-486a-8d8c-0dc5ed4240f6 · outbound

This paper cites Structured denoising diffusion models in discrete state-spaces,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Structured denoising diffusion models in discrete state-spaces,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:8d9b09f6734a0849bf9d49523d00e63e23f9aef0587c2ac15ba813bdb1f72fc6

Observation f08fef00-7267-48db-8903-a0cc860cd472 · outbound

This paper cites Argmax flows and multinomial diffusion: Learning categori- cal distributions,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Argmax flows and multinomial diffusion: Learning categori- cal distributions,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:58d5d2efb5cd112bd303e397e1ca06cad9a265fd0953619e1d7c844347eb7bad

Observation 33b84dd6-d4dc-4b8f-a1a8-7120febb38bc · outbound

This paper cites Vector quantized diffusion model for text-to- image synthesis,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Vector quantized diffusion model for text-to- image synthesis,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:adecd01ca7b584b6e839a06fcaed21bc273d068e645cf1ddddf7e3c03315ec74

Observation 74ee7fd2-4233-423c-a90e-60b1b5a7c823 · outbound

This paper cites Diffusion-lm improves controllable text gener- ation,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Diffusion-lm improves controllable text gener- ation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:8cad8a07b7733edcde89b22ef4300c3c37a78a70d89e0b4c53fa5ba2e51e270d

Observation c19a742a-c009-4406-af5c-2505fd47970f · outbound

This paper cites Maskgit: Masked generative image transformer,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Maskgit: Masked generative image transformer,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:63bcc971d0f90c6316667e93b31ba68758562b96c54da2090473095acb2cdfd5

Observation 9ec63cf4-bed7-4917-8670-0bea5fe5b8fc · outbound

This paper cites Large Language Diffusion Models.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Large Language Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:c1e4ec4703a66c30705a31d5bcb55beb11273ce2f16100e99abfc47c4349f144

Observation 9e9f574d-00ba-4c30-8cb5-3eeffc71803b · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging DINOv2: Learning Robust Visual Features without Supervision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:ae5b46cf2667b5cae40f955212a1df7ee506a8da914400581cb689e24e096751

Observation b7eec16c-baa7-4a21-b11e-8dc943a30131 · outbound

This paper cites Sig- moid loss for language image pre-training,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Sig- moid loss for language image pre-training,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:4c85eb6242f390ec27197468da4f21d593dc598421db709a14d81ab6a93a9392

Observation fd2566e3-8b05-4789-bede-072094f40b39 · outbound

This paper cites Qwen2.5-Coder Technical Report.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Qwen2.5-Coder Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:0061bdb65c0b4402e26c5dc88261db88b5de84f44824bf443274d3e0e4477493

Observation 8b27e478-b6d8-4e74-836d-5d0a8d330349 · outbound

This paper cites Flowvla: Thinking in motion with a visual chain of thought,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Flowvla: Thinking in motion with a visual chain of thought,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:3476225ec2d1220606c2f27a539d8604433498f9aae4df78bee9bf6ec0407150

Observation ac4b013c-ed50-4fd0-9e15-e7a256802667 · outbound

This paper cites Cot-vla: Visual chain-of- thought reasoning for vision-language-action models,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Cot-vla: Visual chain-of- thought reasoning for vision-language-action models,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:8c8c547989f47ab9a1ac16091881a05bd123b40c679e6a9bf8f89d0ca30bc5f0

Observation 2d5d561e-cbf3-44ed-9d84-650fbcb13f37 · outbound

This paper cites ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:97fecb13601ba54308015b05557b79660459599cf4b32bb4f235c73fa39a5cf8

Observation ec5eb7c4-13bc-43c4-90c1-7ce0a8b08a25 · outbound

This paper cites UniVLA: Learning to Act Anywhere with Task-centric Latent Actions.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:3e93cbb16b8f3ef62f2053d83f4edf4f50ed0df46f49edf91b03defc03b27d9a

Observation 8daba2e4-ffd5-4a45-a0fa-89c32a412ad7 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:a4806de159a8b23365260bc0ae9ef3465a73e9da2e200ef817a0bad8bf0989e5

Observation 05c4d5c3-5759-4335-9d8b-b72751d51077 · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:da0880da5788d7e1ac1d51a60ba4cf83ce01a6b1e0263d48aac668b05b55d7de

Observation 756c7481-332b-4c0b-a611-265b8f4fbe84 · outbound

This paper cites OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:4f2e2e4e93a6284172172fd5b90d4e56d58a209a81ce538362ec76a295c202b6

Observation 679555cc-f140-41bf-872f-17c2adabb100 · outbound

This paper cites ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:d6a8f0f79f5f86ca0479b1e2281693ffb27b107dbf16a183daf1ca10c4c28cbb

Observation f67087ff-e350-4716-9512-a028039cd180 · outbound

This paper cites FAST: Efficient Action Tokenization for Vision-Language-Action Models.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging FAST: Efficient Action Tokenization for Vision-Language-Action Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:32b38dfcad01e06d03a7b73df6f7471a45747f424f295f576c65eea4d8385273

Observation b2217ddb-d8c4-43e1-ae7b-9aa2aeb76629 · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:db79f68c53649773e2028cd2c7e889f8358a625ab6da4f12cf72cc693660eea3

Observation 65a9fb3d-95c6-47ff-a54b-f1815e42c368 · outbound

This paper cites Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution,.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Deer-vla: Dynamic inference of multimodal large language models for efficient robot execution,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:b274b7d80c6f0fbb6b47920a17743663d4c6bc0b4b3c5687983bea8e7fe049fc

Observation da9a171c-ec69-422b-b275-2573ef587e77 · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Vision-Language Foundation Models as Effective Robot Imitators

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:8630b63e6b7b0c8cc5b962a3ad5466f7c1a6116d64ab772c8ee394a167093c5c

Observation 95876727-8414-4d8d-9cfd-a16fda10e621 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:d49c054e33b534660cc2427950b0952f0378218230757c52f15cc81620a9175c

Observation fae95702-3314-47fe-92ec-32a640721d30 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:eca717bb994b66861a2afc7e8186de0a6d42689ab47f51a0b29784a2c8305914

Observation 3541f0cc-245d-4df8-a6af-7f0f93129804 · outbound

This paper cites VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:c4b94252fb2d61b4d5084e2de8bc4f9b4b9882ad978d90ef59e64009e39a3196

Observation 3fde7b3c-846c-4c76-a20f-43311929f11f · outbound

This paper cites Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:981044c9b6c799a457f48510ef2ad038358a881f28517e4667bef476db9eb63c

Pith citing papers

No inbound Pith citation observations are available.