Pith. sign in

Paper Citation Record · LEDGER

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2503.19757.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19757 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:10.945513Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T14:25:47.243057Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45883420-8b1b-49c5-b289-cdd0952cddb7 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.245817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:9b3d6ee2833bd08675f92c1d7f8725b1341c5fc07b53ed8ee89ba26daaa2b684

Observation ae01a1b4-d01c-473e-8aa0-0758d5d666d8 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:10.945513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:10.945513Z digest=sha256:c7f64f31aab8d187d511d374d57273eb64474ff459fb66f4b2055f6fa2a5860e

Observation 81c2056a-eaf7-420e-95d3-022c4fd77128 · inbound

Block-wise Adaptive Caching for Accelerating Diffusion Policy cites this paper.

Block-wise Adaptive Caching for Accelerating Diffusion Policy Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:13.968090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T09:36:09.790248Z digest=sha256:b639202093a13920edad4f4a8659f9eed6d1c8903a49039b4585ae94ecc1beb8

Observation 529a2236-55b9-43e9-bbd7-244fb6f8dbf0 · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.584985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:52701cea12f1e84efb73ca952d5727e5cb9d495d4b40ebad3a39fd7549305fdc

Observation 8779decc-64d5-427b-a144-f8ce6795cb08 · inbound

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models cites this paper.

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T21:16:50.629382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:16:50.629382Z digest=sha256:f5b4aed488d9b1f044884d17bcf61b0898de54d0abe9e0bf84f4cf6a67a9de22

Observation cd0ec015-300b-4053-aadc-25259e758238 · inbound

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach cites this paper.

Source Component Shift Adaptation via Offline Decomposition and Online Mixing Approach Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:35:43.793652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:35:43.793652Z digest=sha256:5564eaf10bbd035a2acd92545a54cf07fdeb4b97af933751f70ab64816151eb9

Observation 77f6fd8f-3d92-47dc-981a-a7cbc7c3b8b9 · inbound

Leveraging OS-Level Primitives for Robotic Action Management cites this paper.

Leveraging OS-Level Primitives for Robotic Action Management Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:38:40.247428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:38:40.247428Z digest=sha256:51b65601d7c6d028981a767d8faef6f27155f0d2757f52e2e1d513b0d8be96fa

Observation 9c4cbf47-44f7-46b8-8c6b-00ce61b3bda3 · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.628932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.628932Z digest=sha256:e777c887d7cf3a6baecb08cc64452f15c33e744fbae04e7ebd1cb3479b7e72ad

Observation 1d7a26cb-bc20-4393-9bc1-12fc57289878 · inbound

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance cites this paper.

Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T12:02:28.274529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:02:28.274529Z digest=sha256:ff7f793031f5e5c9e8370d9fc699d89897b12e532b4d7b8d1b3f88b79515de67

Observation ce01b314-0f73-42dd-81e9-bb2039a2668e · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.552915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.552915Z digest=sha256:c08f3f87c91ddf3a41f89524e4e0e5f86c093f59d7c26bf3a9d74a8435022427

Observation a3f25050-fd4b-489d-9bea-fcd5afbff5ed · inbound

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models cites this paper.

AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:30:18.245994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T21:28:18.630934Z digest=sha256:31e9ee86c2a060c9d96b9187fcae3fed58e37d3205d8f56cf8ae2d00bdb93e5a

Observation 867a5f79-f679-49e5-9187-90e85b92fdf8 · inbound

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning cites this paper.

Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T14:50:12.928751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T14:50:12.804707Z digest=sha256:802bbf7981cafbb8127c15f5f9eaf74d15e90027b4549d4ad6ab421e1cbf7715

Observation c4026ade-74dc-4276-8eaf-1bf3f6496f0c · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:04.338230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:04.338230Z digest=sha256:80912cbfbfb2123c49f7bde69062af94fca32bd52d1d5f4f7e3522002122cba6

Observation 7f3888f9-e553-430c-aa38-42c8124f6990 · inbound

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation cites this paper.

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T20:18:05.038317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:18:05.038317Z digest=sha256:121bef3c09de7ea6bc18e871a16b06a16ce835b7494ee87a5141cfd3acd95508

Observation f42c87c8-2ccc-48bc-9c1d-ecf8e8e44b1c · inbound

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception cites this paper.

FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:06.529823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T22:00:06.298685Z digest=sha256:eddb7ebc0b3f9807a9346611bf055183721511878ce274c35239d130db65516b

Observation ecf86412-32d0-4d38-a605-6e5ac4a440c0 · inbound

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation cites this paper.

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:06:06.263709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T16:40:19.057979Z digest=sha256:e8f1949ebfbd57f94fc277d5a592f4f1d1ba46bb22198177aac62b7d749a09b8

Observation a2ccbabb-8d6a-4879-9fec-c936219a8c24 · inbound

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation cites this paper.

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:25:53.215480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:25:23.710842Z digest=sha256:79afb1849ad52db69da8f37bb21d10d564542863685fc3a3eaac26ee44543cfc

Observation d53dc7d5-3131-4113-b37b-d60799f5d88c · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.139142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:05:44.188530Z digest=sha256:45e0a55554335b51063d564040572ae1acdad49b22e0fde14cc42b7466a0ac5d

Observation 9c652f31-f114-43ee-b75b-03749e6410f7 · inbound

Nautilus: From One Prompt to Plug-and-Play Robot Learning cites this paper.

Nautilus: From One Prompt to Plug-and-Play Robot Learning Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T14:19:51.313611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:19:51.313611Z digest=sha256:ff31ff2454d3ea271d9a2bfa47f04b71a7258883188c23589f47ac9e30e262dc

Observation 87a5084d-c753-4d62-ab24-8d4b5e328164 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:16.973759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:748bca673be356a40a52dd77bb50edbb7dd8eedac206f52b6ffcf674c5d341fe