Pith. sign in

Paper Citation Record · LEDGER

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2503.19757.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19757 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:10:27.174266Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T14:25:47.243057Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45883420-8b1b-49c5-b289-cdd0952cddb7 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.245817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:90e5e78eff593d1a290b7e9b52eafd270a57db9db7292f74c3576127b59ecf2e

Observation ae01a1b4-d01c-473e-8aa0-0758d5d666d8 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.945513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.945513Z digest=sha256:c832dd40122194c1b71b261f05fdf0ddde019a1dde57739652b74b2809ab5e18

Observation 81c2056a-eaf7-420e-95d3-022c4fd77128 · inbound

Block-wise Adaptive Caching for Accelerating Diffusion Policy cites this paper.

Block-wise Adaptive Caching for Accelerating Diffusion Policy Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:13.968090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T09:36:09.790248Z digest=sha256:420b505507e142923842ca9481b905b3124c4137804e38d597d67d8ff3630760

Observation 90667bea-1a84-4668-95d3-4a67e3bc3843 · inbound

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models cites this paper.

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:10:27.174266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:10:27.174266Z digest=sha256:2c760632173327175586fa03c0d2971682f8ecb9fbfac08c54c3f60b969c32a8

Observation 529a2236-55b9-43e9-bbd7-244fb6f8dbf0 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.584985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:8fb12a760e336399d8c589b83328e99c5d6a566e2ab1706fb7711a4ff863b6ec

Observation 8779decc-64d5-427b-a144-f8ce6795cb08 · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.629382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.629382Z digest=sha256:17ad9b30831f83b5f07d1bf632e1cc4495bbbbc7dce9cd8731368b63f1c96357

Observation cd0ec015-300b-4053-aadc-25259e758238 · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:43.793652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:43.793652Z digest=sha256:188802076123ca1ffc048340aaf978e44db0ea0931bcf68fd5e32e19802cb97e

Observation 77f6fd8f-3d92-47dc-981a-a7cbc7c3b8b9 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:40.247428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:40.247428Z digest=sha256:030a5f2b5f5f5599472aea256155c26a6d60be5ef4e3a4cc50aff67f5c96f6e9

Observation 3834d317-7c77-4a80-85f6-af2a9466ee0e · inbound

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy cites this paper.

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:20:58.769098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:20:58.769098Z digest=sha256:7fcba7b3388518b28b4aea5c000d83f875cc46ae76f853537005f0b0bbef2c09

Observation 8cbbdc43-4b9c-4957-801a-f8822d4b9d6c · inbound

Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies cites this paper.

Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:54:04.480137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:54:04.480137Z digest=sha256:0c7d6aa26438c1f8efabdbeae2e7d70da8a49f99d3380ac3526de51f101a2ab3

Observation 9c4cbf47-44f7-46b8-8c6b-00ce61b3bda3 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.628932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.628932Z digest=sha256:31b7b209d6ef858bc9e05edb09f0b5b5e8f310caf10432946d8eff2a7b520899

Observation 1d7a26cb-bc20-4393-9bc1-12fc57289878 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.274529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.274529Z digest=sha256:2d22702224a6fa0c65a23cd10fdb9df534112e13e36eaf9f2467df20d056d465

Observation ce01b314-0f73-42dd-81e9-bb2039a2668e · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.552915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.552915Z digest=sha256:c9bf4ae85ff075448a08dff60f2920a67d7fe3054910f45a10c386c93c29f947

Observation a3f25050-fd4b-489d-9bea-fcd5afbff5ed · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:30:18.245994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:b258e1cc4eeb679dac65eedde2d9a39fa22d671410cd797aad5b2568b8b87630

Observation 867a5f79-f679-49e5-9187-90e85b92fdf8 · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:50:12.928751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:af3f3570aed49c024815a2b7c45ff75bfa9c748728818a3bb24600117478f278

Observation c4026ade-74dc-4276-8eaf-1bf3f6496f0c · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.338230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.338230Z digest=sha256:b94ad21f616c0dae278d90c840036108faa87aa60c1ea020ba86527be5c367c5

Observation 7f3888f9-e553-430c-aa38-42c8124f6990 · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.038317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.038317Z digest=sha256:4909d21760d10a7b37baa52a4501b417041db5fb748784e71fe423a92399b19d

Observation f42c87c8-2ccc-48bc-9c1d-ecf8e8e44b1c · inbound

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception cites this paper.

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:06.529823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T22:00:06.298685Z digest=sha256:ef9775c70ea5df00409eb493a395e3cf1c3d8da7e6c9265ef075dc6d7d5e7ad2

Observation ecf86412-32d0-4d38-a605-6e5ac4a440c0 · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.263709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:f37e07b02ac54b5425fe912f370be230bfb1d6abc553995dc50495b755cf2da3

Observation a2ccbabb-8d6a-4879-9fec-c936219a8c24 · inbound

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation cites this paper.

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:25:53.215480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T02:25:23.710842Z digest=sha256:197152213de942dd4e5fb28855992f23974dd92996c3bdbc77e03b13118e0146

Observation d53dc7d5-3131-4113-b37b-d60799f5d88c · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.139142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:56a012e0c37a7be93fd87917dfbd3d49d77a1e1d96c04e707c27e2df234f792d

Observation 9c652f31-f114-43ee-b75b-03749e6410f7 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:51.313611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:51.313611Z digest=sha256:f7c4a2fa715faefda2297e14d303cce1117bc5165b4c98d59d2f64c15c957b7d

Observation 87a5084d-c753-4d62-ab24-8d4b5e328164 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:16.973759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:bbf7cebcea6d3ce6bde1c2fff81f9c776e1ae48659a6a2ed81438ffd96f2cb9d