Pith. sign in

Paper Citation Record · LEDGER

Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2407.05996.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.05996 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:55:52.058916Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:50.161878Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8d349d7c-ca9e-4355-a268-b399094cd15d · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:25:54.447400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:9612286df941bcf37aa1cfa29469480bd560e99ee8c8738b9c5b759b2c1a9958

Observation 2ad859e1-dc7e-4e85-928d-49d597bcbacf · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:38:11.308512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:fd39cd567a7823ae73987687874254ab80750c8916e6ec70dc06b7e57bc38b3e

Observation f8e1ec0b-6798-440c-a082-68b7a88a3788 · inbound

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success cites this paper.

Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:32.782607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T04:35:31.914360Z digest=sha256:1e0613fdf767d051462bfb888571f186d8e23f943dc27a78cb19948be04fa9ba

Observation d5626ecc-b478-4453-9d54-eb8a36d32af6 · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.258877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:399fe2c0d4cca50cf83a185a827ecea273d4fb35395b04a3058f23ce2f735640

Observation ca04556d-919d-4dd7-b074-3d4ea17824ee · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T08:04:12.547881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:c6405b0865b4341b895a1032eb114542218ff46b5ee62e8f417a946ab2448922

Observation 76e7b3a8-92e5-4153-b977-6f651ae96d02 · inbound

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges cites this paper.

Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:52.058916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:52.058916Z digest=sha256:17073689c8450e24c24ef1ac6df31c8fa56991c3325d9c018c1479585578f5d1

Observation fd48f158-af31-4e1f-9d87-0fe1f2dc01fc · inbound

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation cites this paper.

MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:43:24.525433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T20:43:24.417901Z digest=sha256:c66b02bbbb288dbde73d72323382ad9d564a22edaece81520db89cec4d5d5f30

Observation ffee3e7b-3f88-4282-9ee9-d2b050e62ed1 · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.635769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.635769Z digest=sha256:38dc5f5df6901f8d37d1a65b69f265de35f64dea9f0ede8e57be7376178d626a

Observation 7951d0f7-f695-444d-90ee-be9ed1e47dac · inbound

Reflection-Based Task Adaptation for Self-Improving VLA cites this paper.

Reflection-Based Task Adaptation for Self-Improving VLA Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:31:02.986035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T07:28:11.187479Z digest=sha256:24c240401fe5a238db85249160f6a5954cf4f27fdc8e9f71b071541b39f9e794

Observation ee7f89a2-ab41-4a9a-b748-95bfdec19392 · inbound

RynnVLA-002: A Unified Vision-Language-Action and World Model cites this paper.

RynnVLA-002: A Unified Vision-Language-Action and World Model Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:52.285243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:59:52.285243Z digest=sha256:dde3e884f0d1a07b7a6532d700086ad951af985f3edfe3ede42a20ce537970d1

Observation fc590509-8576-4970-bb43-2010eedc415b · inbound

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation cites this paper.

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:08:02.002359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T15:05:21.907878Z digest=sha256:bf4f659829ffb9a1b4220eaffb09426e29973c070090faa88ff5264c6e7ecd0d

Observation 15fbe715-5e6c-4768-9716-468af86a2bc5 · inbound

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies cites this paper.

Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T03:56:05.396182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T03:56:05.396182Z digest=sha256:f6baf61ba8717b7a917779c99911a38f30689073148a6bb6ffea30a289df4be1

Observation 54761429-a8f5-4208-9360-2821397c809c · inbound

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation cites this paper.

Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:14.643503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:14.643503Z digest=sha256:96ca6d711977fef11b2e55c6816b37e073d09b7bed2e26a6a23a555f981b5e0c

Observation 3bbb0ec4-6d9c-496f-9a46-98a8775494e1 · inbound

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory cites this paper.

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T18:49:51.103217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:49:51.103217Z digest=sha256:fd17534a8ee7fc42ebaddd08cae77c7363603ac781a9032c72923c2795880d24

Observation 254afa48-c502-4b35-ae74-3827c9e0af7f · inbound

VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning cites this paper.

VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T23:00:31.128534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:00:31.128534Z digest=sha256:6d785f9880d0d21ba5ec5b2662f6b42595e4e7ad84a3dd20e92ce779b598d19a

Observation cdeb824a-dafd-4038-a29b-6688a7a3bf9c · inbound

Emergent Neural Automaton Policies: Learning Symbolic Structure from Visuomotor Trajectories cites this paper.

Emergent Neural Automaton Policies: Learning Symbolic Structure from Visuomotor Trajectories Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:08:21.377605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:08:16.186576Z digest=sha256:8aac9d782cabf4ecdfcad160eb505c766407aa5f65fdb518cb48017a301e76dc

Observation 9508c971-13d5-46a9-ac4b-f393b047b750 · inbound

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model cites this paper.

A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:51.465622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:31:23.255452Z digest=sha256:69113ff6b5785f56a5af8d38fbaa44660af7ef2033499c6069436ce5e49a62d7

Observation 931e186a-74ee-431a-a788-864391bcccf0 · inbound

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies cites this paper.

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:17.564383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:33:32.677434Z digest=sha256:581922403d79b2e81a949a981a9556311dcfb4d0b8a7b2e5b1c371adaec468e9

Observation 9ea12bec-3613-4844-bf3b-0cae5d759fef · inbound

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training cites this paper.

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T01:03:18.366526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T01:03:08.765350Z digest=sha256:7a719471fc1eff6b6f3abc8f59c79090978e466dce39322eccab455475450dbc

Observation 685b7205-5126-4cb6-b2cb-1f1abe86f475 · inbound

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation cites this paper.

AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:46:04.600695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T04:46:03.020800Z digest=sha256:90236626134f36495b0fd72eb8496f2e5424e696c6752f2188922a018a9da5f9

Observation ad7ffa25-6642-44fc-b6ec-fe6c10308560 · inbound

What Are We Actually Benchmarking in Robot Manipulation? cites this paper.

What Are We Actually Benchmarking in Robot Manipulation? Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:35.333611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T09:31:39.345777Z digest=sha256:e0d5a264e64baa797c037750ad278101344c306532170f14bf8f264e0311d88e

Observation 558b8666-e791-44fd-980b-315c4bfa1be3 · inbound

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling cites this paper.

UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:37.506820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T14:27:40.585782Z digest=sha256:aa18deea9cd484c3a7d4de9b7f37af01e798e3fc467351d95a9cd9219a119639

Observation d84f9e18-49a6-4d60-812c-133f99a2d089 · inbound

Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure cites this paper.

Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.163333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:32:52.472573Z digest=sha256:850c203fb980effe55986c74de053acfeccd31fa43912ce51ae36f27fe0c5d70

Observation 369bdd64-4844-48fb-95e8-985d5619ca7b · inbound

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging cites this paper.

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T15:19:27.489381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:19:27.489381Z digest=sha256:2c0d4b9154009c234f341204b927e97453f452c458358dbc6b8fdb71cfc21d36

Observation 1ae464e7-d62a-4c6c-83a4-c5bc459552ef · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 213

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.162534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.162534Z digest=sha256:6cee68df96b64911177396c413b1e50d13e22ad57409a27fe5fef546c050901f