Pith. sign in

Paper Citation Record · LEDGER

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2503.19757.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19757 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:10.945513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T14:25:47.243057Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45883420-8b1b-49c5-b289-cdd0952cddb7 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.245817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:4017ac5cd68dae16c67ba172add8e302a4da3bf8238aa3169d6447605ff38bc5

Observation ae01a1b4-d01c-473e-8aa0-0758d5d666d8 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.945513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.945513Z digest=sha256:bf172191dc02fd10fd117e45e583ae3232dfc3c1d17448b2b03ee32d1acb14d4

Observation 81c2056a-eaf7-420e-95d3-022c4fd77128 · inbound

Block-wise Adaptive Caching for Accelerating Diffusion Policy cites this paper.

Block-wise Adaptive Caching for Accelerating Diffusion Policy Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:13.968090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T09:36:09.790248Z digest=sha256:64fde3b05ef51fd9d7aec5b79711d1e08adb27b529d47cd2b3379350f51c6992

Observation 529a2236-55b9-43e9-bbd7-244fb6f8dbf0 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.584985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:f3f98f4defc87e201f75aa2cf22994a353ed599bbd5b7e1590aa542c64d16515

Observation 8779decc-64d5-427b-a144-f8ce6795cb08 · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.629382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.629382Z digest=sha256:5538d6bb614ebc6e29ce3afe187ffd641a1c9c631674e8e4880b2d81db1d16a7

Observation cd0ec015-300b-4053-aadc-25259e758238 · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:43.793652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:43.793652Z digest=sha256:0a56f005bce8d8079a5315362458662b2662b61ec282529cd11d290424f64e4f

Observation 77f6fd8f-3d92-47dc-981a-a7cbc7c3b8b9 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:40.247428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:40.247428Z digest=sha256:724ee461f815673c165d635708ef093be8367b360cc6fd7065eb487b22929a88

Observation 9c4cbf47-44f7-46b8-8c6b-00ce61b3bda3 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.628932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.628932Z digest=sha256:bb0c90568df3a73624cc24824f58cc04db3e1286ef49248709c6e43945dec5ce

Observation 1d7a26cb-bc20-4393-9bc1-12fc57289878 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.274529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.274529Z digest=sha256:fd59f47ec429552656070a697049f332c5afd7dda45f0f4e8c2ef5d5e8f2220c

Observation ce01b314-0f73-42dd-81e9-bb2039a2668e · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.552915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.552915Z digest=sha256:1a1ed42b514ba1ad3a840e1d417c9e3e5051b3894cc4335b435224f179ab2256

Observation a3f25050-fd4b-489d-9bea-fcd5afbff5ed · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:30:18.245994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:b0b7a58c4f3079061fa6ebd77cbf4a2524590bac881c3c9573ba95db8fcf9438

Observation 867a5f79-f679-49e5-9187-90e85b92fdf8 · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:50:12.928751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:b0dc2f3ac92b6bf944167d4546f2948edd4ca207c97deb995cc329f3d9c264ba

Observation c4026ade-74dc-4276-8eaf-1bf3f6496f0c · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.338230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.338230Z digest=sha256:4ddd4f6e58370ad8bcb4b5aec4c7559e5a0e0a9aad04886f00e04444ef3d8bb5

Observation 7f3888f9-e553-430c-aa38-42c8124f6990 · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.038317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.038317Z digest=sha256:3e346448699b69882e2ec3fdb3712c7bc140f1720d2fc712e8830b0a53226373

Observation f42c87c8-2ccc-48bc-9c1d-ecf8e8e44b1c · inbound

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception cites this paper.

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:06.529823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T22:00:06.298685Z digest=sha256:bb2e5970eccaf83dff211810fbbc994538c464e9192624ea51097a8d65e0be56

Observation ecf86412-32d0-4d38-a605-6e5ac4a440c0 · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.263709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:ffcffcd6062551f22d8dbc66ee594899af8d08fea47cef301e6faf9d3574b3ca

Observation a2ccbabb-8d6a-4879-9fec-c936219a8c24 · inbound

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation cites this paper.

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:25:53.215480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:25:23.710842Z digest=sha256:183f1901913f7b39298eeeddf373aef034ed179003e19af0d7319ffbe436688b

Observation d53dc7d5-3131-4113-b37b-d60799f5d88c · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.139142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:d442e34af8adf586a5050169262d055ce2128cb92a0d4922f37f6170a1a0643e

Observation 9c652f31-f114-43ee-b75b-03749e6410f7 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:51.313611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:51.313611Z digest=sha256:ef0de4928d1f0d690993f0d250d9aec920ae546c725c29b56380be6f10c034cc

Observation 87a5084d-c753-4d62-ab24-8d4b5e328164 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:16.973759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:3a364916d63541859bf771060954e54019dc65ca3fdb3f06548e99be863be068