Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2407.05996.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:55:52.058916Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:09:50.161878Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 8d349d7c-ca9e-4355-a268-b399094cd15d · inbound
A Survey on Vision-Language-Action Models for Embodied AI Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ad859e1-dc7e-4e85-928d-49d597bcbacf · inbound
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8e1ec0b-6798-440c-a082-68b7a88a3788 · inbound
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5626ecc-b478-4453-9d54-eb8a36d32af6 · inbound
Interactive Post-Training for Vision-Language-Action Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca04556d-919d-4dd7-b074-3d4ea17824ee · inbound
GR-3 Technical Report Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76e7b3a8-92e5-4153-b977-6f651ae96d02 · inbound
Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd48f158-af31-4e1f-9d87-0fe1f2dc01fc · inbound
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffee3e7b-3f88-4282-9ee9-d2b050e62ed1 · inbound
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7951d0f7-f695-444d-90ee-be9ed1e47dac · inbound
Reflection-Based Task Adaptation for Self-Improving VLA Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee7f89a2-ab41-4a9a-b748-95bfdec19392 · inbound
RynnVLA-002: A Unified Vision-Language-Action and World Model Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc590509-8576-4970-bb43-2010eedc415b · inbound
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15fbe715-5e6c-4768-9716-468af86a2bc5 · inbound
Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54761429-a8f5-4208-9360-2821397c809c · inbound
Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bbb0ec4-6d9c-496f-9a46-98a8775494e1 · inbound
VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 254afa48-c502-4b35-ae74-3827c9e0af7f · inbound
VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdeb824a-dafd-4038-a29b-6688a7a3bf9c · inbound
Emergent Neural Automaton Policies: Learning Symbolic Structure from Visuomotor Trajectories Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9508c971-13d5-46a9-ac4b-f393b047b750 · inbound
A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 931e186a-74ee-431a-a788-864391bcccf0 · inbound
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ea12bec-3613-4844-bf3b-0cae5d759fef · inbound
AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 685b7205-5126-4cb6-b2cb-1f1abe86f475 · inbound
AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad7ffa25-6642-44fc-b6ec-fe6c10308560 · inbound
What Are We Actually Benchmarking in Robot Manipulation? Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 558b8666-e791-44fd-980b-315c4bfa1be3 · inbound
UniviewVLA: A Unified Multiview Vision-Language-Action Model with World Modeling Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d84f9e18-49a6-4d60-812c-133f99a2d089 · inbound
Inference-Time Robot Behavior Steering through Physically-Aware Reconfiguration of Task-Structure Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 369bdd64-4844-48fb-95e8-985d5619ca7b · inbound
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae464e7-d62a-4c6c-83a4-c5bc459552ef · inbound
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Reference 213
Source-reported events for the cited work
Unavailable: canonical work link unavailable.