Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:50.880926Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2505.10105.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:50.880926Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T15:44:39.322417Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T09:21:25.450763Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 57b5e0e2-1027-4191-adfa-00eff836bd15 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation MultiMAE: Multi-modal multi-task masked autoencoders
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ece4fe15-9e85-4e1d-9833-e52f89d25242 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Masked autoencoders enable efficient knowledge distillers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e016c546-9b4c-47c8-8370-2516f9b26d97 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 498deb50-496c-4526-a4d7-c07d9fc99fe8 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation On the Opportunities and Risks of Foundation Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6319172f-4ba7-4fad-a6b2-112cc70454da · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Lerobot: State-of-the-art machine learning for real-world robotics in pytorch
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25a01810-a5d5-4ccd-91eb-aca2ba372522 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Emerging properties in self-supervised vision transformers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 268688a2-981d-4806-a70d-286785696a85 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Improved Baselines with Momentum Contrastive Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8ff5c4-666a-4085-89ce-a94c9eca8312 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation An Empirical Study of Training Self-Supervised Vision Transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07532837-838a-4119-8316-6571befb2725 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Diffusion policy: Visuomotor policy learning via action diffusion
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86486ae6-d7f6-4ad9-8c0a-09f4c6b7b29c · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Cleandiffuser: An easy-to-use modularized library for diffusion models in decision making
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a3001712-f0cc-41c3-a1da-c362115fd0a6 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation An image is worth 16x16 words: Transformers for image recognition at scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9c056c97-4264-4fd1-a041-0fe1963541cb · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Rh20t: A robotic dataset for learning diverse skills in one-shot
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dddb9c34-8eaa-4509-a55a-8ad5679b8170 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Masked autoencoders as spatiotemporal learners
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3354df95-b62c-4fab-9a66-b7766e8e590a · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Momentum Contrast for Unsupervised Visual Representation Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9501eef1-cf9f-421a-930b-89b57dc0eb6a · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Girshick
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b683ed56-6684-4aa5-98d0-02ce9945c509 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Ponder: Point cloud pre-training via neural rendering
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 660114b0-4013-491a-a17a-eee8d448fcb9 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation 3d diffuser actor: Policy diffusion with 3d scene representations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2d1d65e2-bbe0-42af-92c6-9e6b74d2e768 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ae154270-be45-4d42-96ef-745ea932c6e0 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Open- VLA: An open-source vision-language-action model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d67d355-a2f0-4bbe-97be-6c75e9a16c15 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58c52df4-6bba-49cc-bd39-6548393d23fe · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59498a5-1193-4543-b8fa-bcae5fdf9dc2 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Generative models in decision making: A survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0dc70a9-6b70-4367-b89b-eb4e38cc685f · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation LIBERO: Benchmarking knowledge transfer for lifelong robot learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1adf7a-e1a3-402c-8d3c-5d6361f9b229 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation RDT-1b: a diffusion foundation model for bimanual manipulation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0b4415-5a07-4d1b-b334-ded8fe97dc39 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Where are we in the search for an artificial visual cortex for embodied intelligence? In Thirty-seventh Conference on Neural Information Processing Systems, NIPS, 2023
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation be8590e8-45ff-4976-b936-3905bb3da8fc · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation R3m: A universal visual representation for robot manipulation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 375ec7e7-5383-4e99-98b3-77090d225bfb · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Octo: An open-source generalist robot policy
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83e1f69a-7a4c-4b9e-a157-2b1852c63231 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd80b5b9-6425-4e08-bd78-521503b346c1 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Masked autoencoders for point cloud self-supervised learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2c8d2d92-be91-4959-82d1-c4a73c2c200f · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Scalable Diffusion Models with Transformers
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143439fd-7eca-42e6-9415-db5e0072b4d5 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Pointnext: Revisiting pointnet++ with improved training and scaling strategies
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 08dc8288-85d7-42a1-8935-18843d3d1e83 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e080f40b-f515-491c-ae4f-0a9f5d1d5d6b · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Learning Transferable Visual Models From Natural Language Supervision
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38350926-7c6c-4bfe-9c5d-cf96dc319e44 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fae65f8f-caf9-4942-9721-cceff07a2907 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Training data-efficient image transformers & distillation through attention
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9fa8a917-7c56-4750-ba55-bbe10182b566 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Zhao, Ken Goldberg, Ryan Hoque, Lawrence Yunliang Chen, Simeon Adebola, Gaurav S
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation afd04968-e46f-4933-9f5b-32cdabe80d1d · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Bridgedata v2: A dataset for robot learning at scale
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ecc42fc2-b8f6-4461-beb9-ad580a6eee69 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Videomae v2: Scaling video masked autoencoders with dual masking
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a9768d38-f625-4d11-affe-8b11a2e56b8e · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 96e0dd93-8314-4b72-8923-68d6c5891859 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4f320ca9-36da-429a-af6c-1b54ca539e08 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Depth anything: Unleashing the power of large-scale unlabeled data
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 13211660-7f98-442c-83ff-82e35a8f25bf · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Depth Anything V2
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 196b199d-7cb6-4631-9c88-298a747c27c4 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 272e37c4-4fb7-4e07-96d9-6d9bc9706969 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56c1656b-ef10-4ec8-a767-e3d4775ca4f4 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Sigmoid loss for language image pre-training
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b6f81b37-908c-4a2b-96df-44a8091e8f72 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation 3d-VLA: A 3d vision-language-action generative world model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 14bb0d37-9ed5-486d-bf0a-dc99246d4476 · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edcd3297-b5cc-4bb9-84ed-d4e0534d5acb · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Point cloud matters: Rethinking the impact of different observation spaces on robot learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3e64e5f7-6038-4b89-bbf9-b27dbdfd476c · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation screwdriver
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 24c5840d-feee-424a-891a-b7c675010abf · outbound
EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3547738d-5d86-41df-a708-4632676f9416 · inbound
Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 471cf710-e970-4c45-9092-2bc0ff23515a · inbound
STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.