Pith. sign in

Paper Citation Record · LEDGER

LLaDA-VLA: Vision Language Diffusion Action Models

As of 21 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 17 inbound Pith citation observations for arXiv:2509.06932.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06932 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:55:30.083174Z

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:40:57.389819Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T18:57:16.674514Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4142e92e-1dd3-4b2d-9b39-eb19daa1a4d6 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

LLaDA-VLA: Vision Language Diffusion Action Models Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.795379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.795379Z digest=sha256:d0c0cafdda5f782d5fcac458d0e156d5a6cb6815cec7de7808b18628dba027b0

Observation d6f912d7-fa1a-47df-9997-d36ecd9e562b · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

LLaDA-VLA: Vision Language Diffusion Action Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.800466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.800466Z digest=sha256:68c87da792a29f85d68a0ddbf8642d7e10d2ac357de71bd82a9d34b4be7749a2

Observation 9a350bb7-8093-45da-abd5-ae7485eb0426 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

LLaDA-VLA: Vision Language Diffusion Action Models RT-H: Action Hierarchies Using Language

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.806153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.806153Z digest=sha256:ec30280c52fc30fc4364db8044c8a576ed02f9bd7266e0a35ece5bd346896fa6

Observation 50b658f0-973b-4902-9b3b-55aff3ec7ea7 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

LLaDA-VLA: Vision Language Diffusion Action Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.811254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.811254Z digest=sha256:07e47672e6b7ce5b4ae4c7fb75c33586af967a3ce9e73b217849ce251f51d2b8

Observation 20987eae-33ed-43e8-8b6a-f52c970e3692 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

LLaDA-VLA: Vision Language Diffusion Action Models Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.816460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.816460Z digest=sha256:614fa202b334071fd119626f76bf67d5badb7ca3bd71ea65b0cd1ed24f34d0cd

Observation 656b043a-7311-4424-9a7f-351e32c9f4b6 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

LLaDA-VLA: Vision Language Diffusion Action Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.820716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.820716Z digest=sha256:1be6e00110c0c879361347bcf69c225f8917b3f0e74b528cf425b79fdd8a7766

Observation 909c812e-2adb-4629-b094-ec1a00797727 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

LLaDA-VLA: Vision Language Diffusion Action Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.826083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.826083Z digest=sha256:b24c291aed1008ba4590533414f86999c9aeb59fba1b128db907e4fb394d4073

Observation d000b2f0-3e47-4c7b-bc31-a64b339ac0cb · outbound

This paper cites Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning.

LLaDA-VLA: Vision Language Diffusion Action Models Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.831085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.831085Z digest=sha256:668511594b6eb1c71cf2bcd58ab0cfbc799d51ccc96eca63434df7eab66bc2df

Observation 66c5254f-5912-42c1-8737-569342e25af1 · outbound

This paper cites Llada-medv: Exploring large language diffusion models for biomedical image understanding.arXiv preprint arXiv:2508.01617,.

LLaDA-VLA: Vision Language Diffusion Action Models Llada-medv: Exploring large language diffusion models for biomedical image understanding.arXiv preprint arXiv:2508.01617,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.836178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.836178Z digest=sha256:6361c0d48b9665447fedb98e3835ed8987cf93aaa010a42d36d65850389c02ff

Observation 954b007b-ede4-4b02-9005-42aca895a98e · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

LLaDA-VLA: Vision Language Diffusion Action Models PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.842053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.842053Z digest=sha256:c934a5872919196b3fd0c4b8ac53264c380d6a9d73dd73cd846fd635191c4afc

Observation b0953557-05f1-4f14-9b5d-0872d9268d10 · outbound

This paper cites Fast ecot: Ef- ficient embodied chain-of-thought via thoughts reuse.arXiv preprint arXiv:2506.07639, 2025.

LLaDA-VLA: Vision Language Diffusion Action Models Fast ecot: Ef- ficient embodied chain-of-thought via thoughts reuse.arXiv preprint arXiv:2506.07639, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.847048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.847048Z digest=sha256:194ee6ca66aac9644b368240d4594888fc71fe445a5e6e2c0572a6e3903b3288

Observation d3d937d0-5e6b-4569-b6bd-b77581ce8987 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

LLaDA-VLA: Vision Language Diffusion Action Models Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.851292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.851292Z digest=sha256:2a88c077140b5fa1a231f3ff611f63903d968f00d74124cba48a6169896a58d9

Observation 0bb9d8e2-3a41-4661-b693-729f8af382ec · outbound

This paper cites Scaling Diffusion Language Models via Adaptation from Autoregressive Models.

LLaDA-VLA: Vision Language Diffusion Action Models Scaling Diffusion Language Models via Adaptation from Autoregressive Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.856347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.856347Z digest=sha256:5800abd3d830ac129300a484eb890eae7ff0595ef9ddc2c3310c2007d9208ae5

Observation 586c3571-f432-46e4-9273-cfc6db6d6fb7 · outbound

This paper cites DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models.

LLaDA-VLA: Vision Language Diffusion Action Models DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.860611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.860611Z digest=sha256:29a401c06c4be3ee19ac093b7d004f0e4fe9ad4f97283b005905dab74a148833

Observation 5dd79bc2-2a4c-4bfe-8968-c395034aee31 · outbound

This paper cites RVT-2: Learning Precise Manipulation from Few Demonstrations.

LLaDA-VLA: Vision Language Diffusion Action Models RVT-2: Learning Precise Manipulation from Few Demonstrations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.864667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.864667Z digest=sha256:bc07ccf4cfba31ebb4255fe96c0251c2322d224438367451c9debf934c1f342a

Observation 1f313bcc-e224-45ff-8c89-d765a66d81ee · outbound

This paper cites Rvt: Robotic view transformer for 3d object manipulation.

LLaDA-VLA: Vision Language Diffusion Action Models Rvt: Robotic view transformer for 3d object manipulation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:31.017539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:29.869969Z digest=sha256:d6c0cfcef14f8968b851db14246269f6733fde23fd557be01c00a90839b83fb7

Observation 70913e5f-3e36-4174-93ca-164a576ae6ca · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

LLaDA-VLA: Vision Language Diffusion Action Models Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.874931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.874931Z digest=sha256:1eb8f155775e24360758931519627187370595ee05ecda8d4ce1b42ebf4f551b

Observation 06a8ee21-23d6-4c63-86b9-ada4b365d3b9 · outbound

This paper cites Coarse-to-fine q-attention: Efficient learn- ing for visual robotic manipulation via discretisation.

LLaDA-VLA: Vision Language Diffusion Action Models Coarse-to-fine q-attention: Efficient learn- ing for visual robotic manipulation via discretisation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.993708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:29.879059Z digest=sha256:87c2f39a21d3f16a0b4afb4fac3457eff808a99e2d268bd243d4f89aa54ff571

Observation 980d009c-05d2-40eb-bf9f-fefe05f94e80 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning.

LLaDA-VLA: Vision Language Diffusion Action Models Bc-z: Zero-shot task generalization with robotic imitation learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.979290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:29.882988Z digest=sha256:1d3701e94cb97251c250d609bd8052e17d8fbe5c1562adb68681aee0fdca6076

Observation af9b2363-d8fc-4031-942b-21afc472065a · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

LLaDA-VLA: Vision Language Diffusion Action Models VIMA: General Robot Manipulation with Multimodal Prompts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.888945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.888945Z digest=sha256:3c0ced26d69437eb8888207fc4ccd9bdba6776eefa6009a1e6c2c400c637602d

Observation d9bdfeeb-eae7-49cb-860a-dad72780a7d4 · outbound

This paper cites Prismatic vlms: Investigating the design space of visually-conditioned language models.

LLaDA-VLA: Vision Language Diffusion Action Models Prismatic vlms: Investigating the design space of visually-conditioned language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.964722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:29.893458Z digest=sha256:4a024871c825abf6362b2b7d37303edc9d7e664d19085a1a75739d05ac3f4a9f

Observation 946783db-520c-4f30-9663-d3b14a69564b · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

LLaDA-VLA: Vision Language Diffusion Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.897295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.897295Z digest=sha256:4d480fce2d79c1e23d4847d39abb2a6ae2c191b56aee09948f972c3687ab6d03

Observation 49a66a05-45a4-48d8-8c9a-f22a3697f4a9 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

LLaDA-VLA: Vision Language Diffusion Action Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.902536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.902536Z digest=sha256:2afb6ff676f3b7af5bc561d0014fcb0326440a0b392827be8b6ad081ea33e03c

Observation 15359c25-1e11-4a3a-a8f0-7f3d52c6286f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LLaDA-VLA: Vision Language Diffusion Action Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.906742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.906742Z digest=sha256:e3cc2673303c6832b90019d1b46e949eb5e3c760aa78263113084a9a23805db6

Observation 61f7c985-98fe-43a7-b0bc-9ab6bd81d5d7 · outbound

This paper cites CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation.

LLaDA-VLA: Vision Language Diffusion Action Models CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.910947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.910947Z digest=sha256:4183d4f85d16b34359c091a05cd7cb8ac3a9f183063c11f374c79486bf63f1db

Observation 56f009d3-a24e-48b5-a52f-14a884af83c9 · outbound

This paper cites LaViDa: A Large Diffusion Language Model for Multimodal Understanding.

LLaDA-VLA: Vision Language Diffusion Action Models LaViDa: A Large Diffusion Language Model for Multimodal Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.915607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.915607Z digest=sha256:3994edaae61073aa0af33ee0e367e1fda7183e6300cdfac38089e0974d0331e7

Observation 7fc5df73-d66f-4695-8f61-a270c864fd79 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

LLaDA-VLA: Vision Language Diffusion Action Models Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.920247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.920247Z digest=sha256:9c709738fc74e9a3f6b83714eda0f8626d0943a3c2a62f8d61bd9e920240334a

Observation 6887cc6b-0618-4f86-a323-39a4b08a9c20 · outbound

This paper cites What Matters in Building Vision-Language-Action Models for Generalist Robots.

LLaDA-VLA: Vision Language Diffusion Action Models What Matters in Building Vision-Language-Action Models for Generalist Robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.925406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.925406Z digest=sha256:a778fcae3f86e91332abc2af1c61f92b00a93e77f65eacc9f23d277e84bec193

Observation 1a5b98ee-3931-467c-b4b6-6d209502f48d · outbound

This paper cites Vision-Language Foundation Models as Effective Robot Imitators.

LLaDA-VLA: Vision Language Diffusion Action Models Vision-Language Foundation Models as Effective Robot Imitators

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.929977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.929977Z digest=sha256:cb735b74ba9c28a25143ba42825fe38998be3ffc518c661f3fa2bb880ccbf928

Observation 6e0c8d87-bb01-4e80-bf68-19b7a6e696be · outbound

This paper cites Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies.

LLaDA-VLA: Vision Language Diffusion Action Models Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.938569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.938569Z digest=sha256:01eea5e422bd5af742b7fc018d2a3cabbd3850431df717ca358b3504276e577c

Observation 5d7f7e4c-f1e9-4c8e-9d55-74b90e094a75 · outbound

This paper cites Improved baselines with visual instruction tuning.

LLaDA-VLA: Vision Language Diffusion Action Models Improved baselines with visual instruction tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.950127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:29.943261Z digest=sha256:8ceb801a43102763dd8078a818bf03f196254bbf115f1f377eab89fc321cc998

Observation ff9079db-8d75-4718-afbc-436168d2ba48 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

LLaDA-VLA: Vision Language Diffusion Action Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.947483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.947483Z digest=sha256:292fda910587058f31283aaa5a1353f998c895b0a43ef9ea8b9d281d4b460dc1

Observation 65e53d20-3824-404a-bb40-255f876bf00a · outbound

This paper cites Longllada: Unlocking long context capabilities in diffusion llms.arXiv preprint arXiv:2506.14429, 2025.

LLaDA-VLA: Vision Language Diffusion Action Models Longllada: Unlocking long context capabilities in diffusion llms.arXiv preprint arXiv:2506.14429, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.951918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.951918Z digest=sha256:9df739e2e58da0cd024f11795d6dadc2976be80dfa23818d5d4245d1bdc7e916

Observation b60e2aac-c39e-4358-873e-532117bedf33 · outbound

This paper cites dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching.

LLaDA-VLA: Vision Language Diffusion Action Models dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.956571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.956571Z digest=sha256:16f77de55359bc5bc4b15ad1b347c5217ab9cc1e3856803c67963a72f98f3098

Observation d6db1e3d-f177-4498-b7b4-303f588ee69c · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

LLaDA-VLA: Vision Language Diffusion Action Models Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.961227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.961227Z digest=sha256:41f7980e28c34877d529c225e25607fd018a59e740e2a05284d959f555701f73

Observation 73c5d6c0-1179-48aa-bec2-020b03d38623 · outbound

This paper cites TESS: Text-to-Text Self-Conditioned Simplex Diffusion.

LLaDA-VLA: Vision Language Diffusion Action Models TESS: Text-to-Text Self-Conditioned Simplex Diffusion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.966504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.966504Z digest=sha256:9516d1631611194f604a4897354404f0b137d3c626c79688aa56f44b0f08eb47

Observation 1be22e55-c035-406e-9265-841f9c47b9c9 · outbound

This paper cites Octo: An open- source generalist robot policy.

LLaDA-VLA: Vision Language Diffusion Action Models Octo: An open- source generalist robot policy

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.925638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:29.972204Z digest=sha256:fbb48267f21be7a34c0801d86f8ab8c43cd201c264f802ba5a2c81f279777852

Observation a3d3fb8f-f93f-4361-a4ad-bdf459571e02 · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022.

LLaDA-VLA: Vision Language Diffusion Action Models Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.910871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:29.976637Z digest=sha256:96b960c43f7761993220aa1d92452b1733ee7ee37af438f6ac0bb176ef30d8d7

Observation 2fe57455-9b4f-4db2-b7f4-889f5eb850c1 · outbound

This paper cites Large Language Diffusion Models.

LLaDA-VLA: Vision Language Diffusion Action Models Large Language Diffusion Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.981892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.981892Z digest=sha256:dcfc561600c87fdae73767e18e45c409f49cf7a339a76d60f832a8b2ee45bd04

Observation c04f1ee4-3dc9-45e2-afbe-3ef261826171 · outbound

This paper cites Llarva: Vision-action instruction tuning enhances robot learning.

LLaDA-VLA: Vision Language Diffusion Action Models Llarva: Vision-action instruction tuning enhances robot learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.895605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:29.986337Z digest=sha256:aa3a0815cb73b465aa978357ed8be12d18911b420fb1ac84afae4a900354e449

Observation ccffb727-f5e1-46de-8b1e-1d179e75304b · outbound

This paper cites Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data.

LLaDA-VLA: Vision Language Diffusion Action Models Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.991033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.991033Z digest=sha256:3311304c97ea37714f995b529d15dddb03ff0b8cb18bfefadaf2e088fd7e992f

Observation 36ab77d8-b68f-419d-8f80-e576c4ff9d51 · outbound

This paper cites Scalable diffusion models with transformers.

LLaDA-VLA: Vision Language Diffusion Action Models Scalable diffusion models with transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:29.995928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:29.995928Z digest=sha256:4904ed4e1d1852abf3fcfa13e863775f9b1b0e03a6f14d8cc3a87e7b408e151a

Observation 582c7308-4847-4f6b-9c7a-99376c836ab2 · outbound

This paper cites Robot learning with sen- sorimotor pre-training.

LLaDA-VLA: Vision Language Diffusion Action Models Robot learning with sen- sorimotor pre-training

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.869928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:30.000327Z digest=sha256:2567954458ec5d1e85254a3417d0d2db16b9740ee7be9940c68da76372cf880d

Observation 68264fe8-9500-4ae4-bdb2-50bc0070022d · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

LLaDA-VLA: Vision Language Diffusion Action Models High-resolution image synthesis with latent diffusion models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.008701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.008701Z digest=sha256:d6deb1b3c18ca7f7f4aa729780949b4c41f1a026ff11780646b23768bb0c505f

Observation f3c6487c-b3d1-4aa9-9df0-68751ad22874 · outbound

This paper cites Simple and effective masked dif- fusion language models.Advances in Neural Information Processing Systems, 37:130136–130184, 2024.

LLaDA-VLA: Vision Language Diffusion Action Models Simple and effective masked dif- fusion language models.Advances in Neural Information Processing Systems, 37:130136–130184, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.846072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:30.013198Z digest=sha256:fb2b4c00ee6710b7a3b583b5bd7e025f7e6ff2574308834ebd07b32505ad4592

Observation 244008f8-0252-400b-9e77-ce6a53cd11ff · outbound

This paper cites Simplified and generalized masked diffu- sion for discrete data.Advances in neural information pro- cessing systems, 37:103131–103167, 2024.

LLaDA-VLA: Vision Language Diffusion Action Models Simplified and generalized masked diffu- sion for discrete data.Advances in neural information pro- cessing systems, 37:103131–103167, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.831101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:30.019309Z digest=sha256:22d5ede51e12dbe28ba90c4bece77cdd19a33e8861532a753ef9f27bf78ba70b

Observation 82e8cb2f-0cb5-404d-9a1f-a75baf94c705 · outbound

This paper cites Perceiver- actor: A multi-task transformer for robotic manipulation.

LLaDA-VLA: Vision Language Diffusion Action Models Perceiver- actor: A multi-task transformer for robotic manipulation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.816526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:30.023735Z digest=sha256:e878ee2926ff3d56cfd34c99a8865fa4b549f96119059de05dfe1ccc7bb0a9b1

Observation 32838acf-7e27-4e86-8596-0da07954ce89 · outbound

This paper cites Denoising Diffusion Implicit Models.

LLaDA-VLA: Vision Language Diffusion Action Models Denoising Diffusion Implicit Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.028421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.028421Z digest=sha256:a9184386325a717209c3f55089d81c9e18a8e6083d2882777eaa735b93e3adc4

Observation b02b9883-2625-4d38-ba88-13ecf12f90ac · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

LLaDA-VLA: Vision Language Diffusion Action Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.032889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.032889Z digest=sha256:40d6d49ab7efd7086ed757205625f406b86356e10792ecb0bf0b5cf24fc53a81

Observation 21e50e02-f01a-4006-91ff-5e59b2eb6a3d · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models.

LLaDA-VLA: Vision Language Diffusion Action Models Open x-embodiment: Robotic learning datasets and rt-x models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:55:30.799714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T22:55:30.039381Z digest=sha256:8af6bbb6ea8daea78d805eb728b026cc32cb70ac0b4cd5183d23375db672aa3d

Observation f42e7f92-73d4-4fcd-8413-493de7130406 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

LLaDA-VLA: Vision Language Diffusion Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.043529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.043529Z digest=sha256:bf4e6a0693ad703d424bade890fcdfea527f634d6f5a2d7ba16224bb87b495ba

Observation e237fcef-4e3d-4e62-8ca6-2cb1b17b89ff · outbound

This paper cites MMaDA: Multimodal Large Diffusion Language Models.

LLaDA-VLA: Vision Language Diffusion Action Models MMaDA: Multimodal Large Diffusion Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.049252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.049252Z digest=sha256:e3d3ee374af54be4360b97d6368359bc8b98d4dfb258377a9862f5431c447aa7

Observation 5386cecc-3486-4735-a8bd-edc25ccff9eb · outbound

This paper cites Dream 7B: Diffusion Large Language Models.

LLaDA-VLA: Vision Language Diffusion Action Models Dream 7B: Diffusion Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.054877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.054877Z digest=sha256:4e6daa52ae0ab39e08a44c2286a3028b60029c280e7e3143a9253e131eb21a52

Observation d0b276d1-3ab8-4f95-884a-a1c333866093 · outbound

This paper cites DINOISER: Diffused Conditional Sequence Learning by Manipulating Noises.

LLaDA-VLA: Vision Language Diffusion Action Models DINOISER: Diffused Conditional Sequence Learning by Manipulating Noises

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.059429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.059429Z digest=sha256:a4d32c3a993e9deda40212209307328d7ee8a8e623c2d4e14be0a63895c48340

Observation 072f861c-8656-4bc4-935d-2ad3398040cc · outbound

This paper cites LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning.

LLaDA-VLA: Vision Language Diffusion Action Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.063762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.063762Z digest=sha256:be00a78866da2ccecf0bbd18e5b53cba06254845a37a6b54775af538f3f8d497

Observation 0d40a83e-56b0-41a9-876b-a924ce0b0b26 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

LLaDA-VLA: Vision Language Diffusion Action Models Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.068480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.068480Z digest=sha256:efa1b6b1109a138f469b75ae2b040a2109572dac33a935fa817a18572c693ed4

Observation fb3247b9-c507-442d-99a3-adf9be5b9ed0 · outbound

This paper cites Diffa: Large language diffusion models can lis- ten and understand.arXiv preprint arXiv:2507.18452, 2025.

LLaDA-VLA: Vision Language Diffusion Action Models Diffa: Large language diffusion models can lis- ten and understand.arXiv preprint arXiv:2507.18452, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.074334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.074334Z digest=sha256:3e2698a0ec009fc8e5d35ee351cd06476b15b7a5a6b82582db51a637ca16ebd5

Observation ee066cac-abb4-4331-9d3d-4ca7f09bc4e9 · outbound

This paper cites ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model.

LLaDA-VLA: Vision Language Diffusion Action Models ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.078793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.078793Z digest=sha256:141717f3ce456b01e67c6aefd4c858d57f638ffed7a1f0f1d8e5c141511f0ec5

Observation 6cb97916-0353-40f0-b242-8e89eafffa7b · outbound

This paper cites LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models.

LLaDA-VLA: Vision Language Diffusion Action Models LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.083174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.083174Z digest=sha256:94bc4220bf57686be680d2cbe058c420a301af44cb553636c2ffd9de8f781aab

Pith citing papers

Observation 44b1ebeb-e89f-47ee-8e4d-539be1f87c66 · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models LLaDA-VLA: Vision Language Diffusion Action Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:30:18.361832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:23694e087fe13eea5ce97e2edd583e345520d9c130401d2f7f0998864e875b2e

Observation 8e4532ff-525b-4139-a226-2159f597f3c9 · inbound

Learning Native Continuation for Action Chunking Flow Policies cites this paper.

Learning Native Continuation for Action Chunking Flow Policies LLaDA-VLA: Vision Language Diffusion Action Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:40:08.711965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T12:38:26.522838Z digest=sha256:f4922851893ef02fefdfa3c6632d12de09d0fc994b88cf7ac985dabc94fb8ec5

Observation 330c185c-d790-4c7b-9afb-5e05ed1e8b2b · inbound

DMax: Aggressive Parallel Decoding for dLLMs cites this paper.

DMax: Aggressive Parallel Decoding for dLLMs LLaDA-VLA: Vision Language Diffusion Action Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:45:55.785103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:58:17.880199Z digest=sha256:7ace240961dc91d803cbbd7dda06f4620bb3846ea48c1c68ca7c94b343052b2c

Observation 6c60f7ac-ceee-4c7e-a488-72c09becbfa4 · inbound

DMax: Aggressive Parallel Decoding for dLLMs cites this paper.

DMax: Aggressive Parallel Decoding for dLLMs LLaDA-VLA: Vision Language Diffusion Action Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:47:40.177584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T16:46:56.743268Z digest=sha256:19eb98eb679a62229e5ae7a7913d1fbbb432ff1a409d09defffdb73615d2c33c

Observation 23710f52-11e2-4168-bf37-59b1a4154203 · inbound

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model cites this paper.

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model LLaDA-VLA: Vision Language Diffusion Action Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:09.570974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T11:45:18.081248Z digest=sha256:1657d3e641dc803602024e92cc66c665171258236c94b6aaa497b1cb762199f2

Observation dc0ba0de-1cfe-4be6-8aec-411d9070a164 · inbound

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors cites this paper.

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors LLaDA-VLA: Vision Language Diffusion Action Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:46:11.943493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T02:30:30.319810Z digest=sha256:ad904714105207005d28e4e59b9f9929e087b72eb6c125456166a93349c917ab

Observation 1daa436e-d978-42e4-9a26-9d4a7e1cb41b · inbound

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors cites this paper.

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors LLaDA-VLA: Vision Language Diffusion Action Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:35:33.394169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T08:32:26.160039Z digest=sha256:c12bb19afb69a2f41d0bf3280732d17f1858a7257a396ec88299c1c0e3ffdb54

Observation 041378e5-c6ad-4bf7-b1c3-d12d8b8e1160 · inbound

ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving cites this paper.

ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving LLaDA-VLA: Vision Language Diffusion Action Models

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:26:08.254357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T16:06:47.108382Z digest=sha256:40fb8f9ad93fc17de4831280f8b8937e951cb3169b3d59f2dc9a419d2d962712

Observation 0bbce6bd-145a-4d92-bf6d-0fc52d895c1c · inbound

ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving cites this paper.

ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving LLaDA-VLA: Vision Language Diffusion Action Models

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:52:05.456030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T01:48:36.105389Z digest=sha256:6cb8bcac098a2d5189d0651d3a692dbb79bd3db98f8198343b53b8955dcf0dfa

Observation 6d02ba59-5586-42d7-a183-625286c931f5 · inbound

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation cites this paper.

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation LLaDA-VLA: Vision Language Diffusion Action Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:25:52.956411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T02:25:23.710842Z digest=sha256:d40e4b2a40bb764b15b33f57376e2c81c280ae7204fd28dceb82781f88cef54c

Observation 9815f35e-d1e3-4d66-a9a3-0a8281aec955 · inbound

BlockVLA: Accelerating Autoregressive VLA via Block Diffusion Finetuning cites this paper.

BlockVLA: Accelerating Autoregressive VLA via Block Diffusion Finetuning LLaDA-VLA: Vision Language Diffusion Action Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:52:32.946824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T17:51:48.905620Z digest=sha256:2a3afecba37e2ac4c5bf023c46b09f24661394ca65708b4b3aa6f2da4fec846b

Observation d0f6ce09-66d8-47c1-922a-148f392f0383 · inbound

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR cites this paper.

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR LLaDA-VLA: Vision Language Diffusion Action Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:18:07.249303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:14:31.613251Z digest=sha256:43dc9efa463487106ef4abbe1ff5d9861d75c6b10f271f07eb3c096175be4ea4

Observation f21a925f-5d06-4557-9dc5-f13b11143ef7 · inbound

TBD-VLA: Temporal Block Diffusion Vision Language Action Model cites this paper.

TBD-VLA: Temporal Block Diffusion Vision Language Action Model LLaDA-VLA: Vision Language Diffusion Action Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T18:57:16.676386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T21:49:20.600217Z digest=sha256:8a7f0b7bfd71d9b10d89e15488023ce40042eb78d4f516608429f2236880935c

Observation d3914466-a61f-4704-821b-c51cb2ed65c4 · inbound

SEAM: Smooth Execution of Action-Chunked Motion for Vision-Language-Action Policies cites this paper.

SEAM: Smooth Execution of Action-Chunked Motion for Vision-Language-Action Policies LLaDA-VLA: Vision Language Diffusion Action Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T16:29:25.372250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:29:25.372250Z digest=sha256:50580ce71de3ccff4e0186a2cf09ab368b19117207db963e8543512eede72e39

Observation edc2f0b6-8ead-423b-8773-7cadbacbc3cd · inbound

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging cites this paper.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging LLaDA-VLA: Vision Language Diffusion Action Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:56571cf27f7191f5ad93bd393ad1c7809ebd50aa6421d2a2ab64e4a856e176e5

Observation 07535b9a-00bd-4064-8cbd-ddd0c529edfd · inbound

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation cites this paper.

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation LLaDA-VLA: Vision Language Diffusion Action Models

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:46.836377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:46.836377Z digest=sha256:b91ca360be5924b9d15001d2993359e574f754421a23c29cd26ef398811d9eb3

Observation 1e4549db-746d-45e5-9b47-f66c91c1b066 · inbound

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA cites this paper.

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA LLaDA-VLA: Vision Language Diffusion Action Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:57.389819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:40:57.389819Z digest=sha256:badc071142e046294af8ba394ffecf96e9752de94cf097dd8565fc1f0ece4f93