Pith. sign in

Paper Citation Record · LEDGER

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation

As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2505.10105.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10105 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:50.880926Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:44:39.322417Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:21:25.450763Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57b5e0e2-1027-4191-adfa-00eff836bd15 · outbound

This paper cites MultiMAE: Multi-modal multi-task masked autoencoders.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation MultiMAE: Multi-modal multi-task masked autoencoders

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.434500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.708785Z digest=sha256:5801751bab394dac3c758ba16a95c73f3aaeb694331b52a9205da1078cccdde0

Observation ece4fe15-9e85-4e1d-9833-e52f89d25242 · outbound

This paper cites Masked autoencoders enable efficient knowledge distillers.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Masked autoencoders enable efficient knowledge distillers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.424665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.712835Z digest=sha256:d66f24edcbaa256e1d1326614f283b1bb785067ce9f7ccad87cb62aa0d804dc2

Observation e016c546-9b4c-47c8-8370-2516f9b26d97 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.715970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.715970Z digest=sha256:8b8af5a4654605c0c5645182925bc479fb510447984579e4d18cb05147e04938

Observation 498deb50-496c-4526-a4d7-c07d9fc99fe8 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.720623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.720623Z digest=sha256:0be621170075fca504f241a7a6484e663e9a1a622d33dd8385b5347cc9a4532d

Observation 6319172f-4ba7-4fad-a6b2-112cc70454da · outbound

This paper cites Lerobot: State-of-the-art machine learning for real-world robotics in pytorch.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Lerobot: State-of-the-art machine learning for real-world robotics in pytorch

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.725034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.725034Z digest=sha256:2cec9660fd46812762d89f21f0a5edba3750bedf1193f7fb55fb5bae81557ab1

Observation 25a01810-a5d5-4ccd-91eb-aca2ba372522 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Emerging properties in self-supervised vision transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.409558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.728435Z digest=sha256:87c8ea42b6fe113b1a31c243899c994d9279151594c80176eafa8b0272e2ddac

Observation 268688a2-981d-4806-a70d-286785696a85 · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Improved Baselines with Momentum Contrastive Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.732017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.732017Z digest=sha256:df48758d6387d0b15b1edfe5e4aa4df486d1d06e2046158cb8447efbdca2609d

Observation de8ff5c4-666a-4085-89ce-a94c9eca8312 · outbound

This paper cites An Empirical Study of Training Self-Supervised Vision Transformers.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation An Empirical Study of Training Self-Supervised Vision Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.735924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.735924Z digest=sha256:23bc328668ab9967ede060bf2405f2c6f2913dc886acaf18a7f7027ffa70f270

Observation 07532837-838a-4119-8316-6571befb2725 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Diffusion policy: Visuomotor policy learning via action diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.739304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.739304Z digest=sha256:ab1b0bee8f998e21bf4685b8b99f23bed72f6e35b5ae84106c2867575a7be262

Observation 86486ae6-d7f6-4ad9-8c0a-09f4c6b7b29c · outbound

This paper cites Cleandiffuser: An easy-to-use modularized library for diffusion models in decision making.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Cleandiffuser: An easy-to-use modularized library for diffusion models in decision making

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.393950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.742645Z digest=sha256:bd4f1ceb7040da5b5ab09b12354e31cd36f05e98f8164fc6c63a028c83b7a6d0

Observation a3001712-f0cc-41c3-a1da-c362115fd0a6 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.384260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.745986Z digest=sha256:baea240c7b5a08cc2b2b9bbcca5fa1250800ab279e1f2082aed44f43d2df1d93

Observation 9c056c97-4264-4fd1-a041-0fe1963541cb · outbound

This paper cites Rh20t: A robotic dataset for learning diverse skills in one-shot.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Rh20t: A robotic dataset for learning diverse skills in one-shot

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.373376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.749364Z digest=sha256:b63c6a15629554cc65c16577e2eb4c237d78c03d60e39ccf5093d8e173317932

Observation dddb9c34-8eaa-4509-a55a-8ad5679b8170 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Masked autoencoders as spatiotemporal learners

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.363119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.753204Z digest=sha256:b5157ff2ea2476ea10a81d6f308bc20f34e02cc517bf5d6921643bc2f4ccc926

Observation 3354df95-b62c-4fab-9a66-b7766e8e590a · outbound

This paper cites Momentum Contrast for Unsupervised Visual Representation Learning.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Momentum Contrast for Unsupervised Visual Representation Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.757581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.757581Z digest=sha256:77194438981a1d1ceb4e46d8f5ffa795115ca90cf3c60de88227e01a100f3bdb

Observation 9501eef1-cf9f-421a-930b-89b57dc0eb6a · outbound

This paper cites Girshick.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Girshick

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.352235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.761089Z digest=sha256:f43f72965fee61f1df653a7e16a0a8eafe447d4233d6ff692f9d8774e22301d7

Observation b683ed56-6684-4aa5-98d0-02ce9945c509 · outbound

This paper cites Ponder: Point cloud pre-training via neural rendering.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Ponder: Point cloud pre-training via neural rendering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.341625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.764217Z digest=sha256:ec7cf6be83a27ce7bafc42abb49aa0100e57ae1815a7c6184849c657c11aab2b

Observation 660114b0-4013-491a-a17a-eee8d448fcb9 · outbound

This paper cites 3d diffuser actor: Policy diffusion with 3d scene representations.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation 3d diffuser actor: Policy diffusion with 3d scene representations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.331124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.767529Z digest=sha256:e184aeb393e8cfbac0fcdc34b3106b9b20b7338e3ca3a79f1c6158c2b1b2a842

Observation 2d1d65e2-bbe0-42af-92c6-9e6b74d2e768 · outbound

This paper cites an unresolved cited work.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:19:51.320153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.770679Z digest=sha256:796909a3a04eba181f831d61b1fc2c166dda369d225c0a7118dc50ac8ecc8d10

Observation ae154270-be45-4d42-96ef-745ea932c6e0 · outbound

This paper cites Open- VLA: An open-source vision-language-action model.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Open- VLA: An open-source vision-language-action model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.774115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.774115Z digest=sha256:070cc07faacc2c540a7e7fec9307fe472e29a11d81781b9e5a44ec3323d8b2b0

Observation 5d67d355-a2f0-4bbe-97be-6c75e9a16c15 · outbound

This paper cites Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.777731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.777731Z digest=sha256:f19dd5a93b62683e9876faa2ed6bb8a8fe81a5d8c2ba1c12f5a0949de55980de

Observation 58c52df4-6bba-49cc-bd39-6548393d23fe · outbound

This paper cites PointVLA: Injecting the 3D World into Vision-Language-Action Models.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation PointVLA: Injecting the 3D World into Vision-Language-Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.782316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.782316Z digest=sha256:397739e1146b9a14b558b4c0999400d7954ae7710a9f4cedf01f455c5a5e1b9f

Observation d59498a5-1193-4543-b8fa-bcae5fdf9dc2 · outbound

This paper cites Generative models in decision making: A survey.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Generative models in decision making: A survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.786523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.786523Z digest=sha256:379afc6ed8538613e433e0781ce2297ee67373cb7aea4013dc60d2dc17f3b79f

Observation a0dc70a9-6b70-4367-b89b-eb4e38cc685f · outbound

This paper cites LIBERO: Benchmarking knowledge transfer for lifelong robot learning.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation LIBERO: Benchmarking knowledge transfer for lifelong robot learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.789781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.789781Z digest=sha256:fd4151cc4dff32da69b05649c27348291c1ebc0281813035dbb16ea5f237ae3e

Observation bd1adf7a-e1a3-402c-8d3c-5d6361f9b229 · outbound

This paper cites RDT-1b: a diffusion foundation model for bimanual manipulation.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation RDT-1b: a diffusion foundation model for bimanual manipulation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.792854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.792854Z digest=sha256:ce718cf484eee32542d26f2884d9bf6e748ffda49bd7cbaf12d425627177487d

Observation 0f0b4415-5a07-4d1b-b334-ded8fe97dc39 · outbound

This paper cites Where are we in the search for an artificial visual cortex for embodied intelligence? In Thirty-seventh Conference on Neural Information Processing Systems, NIPS, 2023.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Where are we in the search for an artificial visual cortex for embodied intelligence? In Thirty-seventh Conference on Neural Information Processing Systems, NIPS, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.293018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.796718Z digest=sha256:b1b7a46d1721c4ad4bc6596dc7bcf19e9c64916c097a78c702a311d81428823a

Observation be8590e8-45ff-4976-b936-3905bb3da8fc · outbound

This paper cites R3m: A universal visual representation for robot manipulation.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation R3m: A universal visual representation for robot manipulation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.281335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.799819Z digest=sha256:44c6fe18bd1c68fa306d94efb9a59670451c59799404ca12bf98f28eec1d80be

Observation 375ec7e7-5383-4e99-98b3-77090d225bfb · outbound

This paper cites Octo: An open-source generalist robot policy.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Octo: An open-source generalist robot policy

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.802727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.802727Z digest=sha256:87fe9bce9430c261a3f76ac7cda8aae39cde96ff8975d9b084b582fffb0d349a

Observation 83e1f69a-7a4c-4b9e-a157-2b1852c63231 · outbound

This paper cites an unresolved cited work.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.806109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.806109Z digest=sha256:1cc38c39df590c0c80ccf1cecdc435ae64d044b8fa5afead108eb9b2e2ebc83e

Observation fd80b5b9-6425-4e08-bd78-521503b346c1 · outbound

This paper cites Masked autoencoders for point cloud self-supervised learning.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Masked autoencoders for point cloud self-supervised learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.259733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.808932Z digest=sha256:4eec73af403149fdbac096394ec6cd54f4edb2c410036d9d6febe96b1fbf346a

Observation 2c8d2d92-be91-4959-82d1-c4a73c2c200f · outbound

This paper cites Scalable Diffusion Models with Transformers.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Scalable Diffusion Models with Transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.812313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.812313Z digest=sha256:fbfefa76ec1a35ac5649ee9e6bcda7ededc07dcbd507cc8c63693fa87aa3529b

Observation 143439fd-7eca-42e6-9415-db5e0072b4d5 · outbound

This paper cites Pointnext: Revisiting pointnet++ with improved training and scaling strategies.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Pointnext: Revisiting pointnet++ with improved training and scaling strategies

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.249206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.815734Z digest=sha256:ec5544ba12c3ca50af8245467a68908c411675ac54ade168d94bf3c9132e7189

Observation 08dc8288-85d7-42a1-8935-18843d3d1e83 · outbound

This paper cites SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.818835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.818835Z digest=sha256:d8e557a75ca5ddd2cd266b695ba04c5722e8d73183dc9f14c086b2748630711c

Observation e080f40b-f515-491c-ae4f-0a9f5d1d5d6b · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Learning Transferable Visual Models From Natural Language Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.822145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.822145Z digest=sha256:c9a24a29c0fb941438653288a471a62a1d96e428f049391a8ec6d25708f1c0dd

Observation 38350926-7c6c-4bfe-9c5d-cf96dc319e44 · outbound

This paper cites VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.237420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.825600Z digest=sha256:2b2c7eae7956b7a1ba0152cabb71b8112c77ba2ee2a6351b652bf975958c34d3

Observation fae65f8f-caf9-4942-9721-cceff07a2907 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Training data-efficient image transformers & distillation through attention

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.227157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.828578Z digest=sha256:eb0e2bff34a593958efc2aca764b12415c5a1771ef65ed3950be7bb22f564324

Observation 9fa8a917-7c56-4750-ba55-bbe10182b566 · outbound

This paper cites Zhao, Ken Goldberg, Ryan Hoque, Lawrence Yunliang Chen, Simeon Adebola, Gaurav S.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Zhao, Ken Goldberg, Ryan Hoque, Lawrence Yunliang Chen, Simeon Adebola, Gaurav S

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.216627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.832074Z digest=sha256:32a11cead84f9207841f80053652fad440b112ad168384021a38b4f53c04c026

Observation afd04968-e46f-4933-9f5b-32cdabe80d1d · outbound

This paper cites Bridgedata v2: A dataset for robot learning at scale.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Bridgedata v2: A dataset for robot learning at scale

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.205072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.835511Z digest=sha256:defb72da905b647f0a550b22f83180ea0b6414c133ffc8c1941a9573539c6e38

Observation ecc42fc2-b8f6-4461-beb9-ad580a6eee69 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Videomae v2: Scaling video masked autoencoders with dual masking

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.194048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.838646Z digest=sha256:1d6b34151f64387ef63c08f88ff4dbf08c71181ce58a158a3e1333c4241216c0

Observation a9768d38-f625-4d11-affe-8b11a2e56b8e · outbound

This paper cites Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.182883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.841830Z digest=sha256:506543dc2cfc41e74864ab313aa02334f1db3136b4f430ea646f743d5a07eb3d

Observation 96e0dd93-8314-4b72-8923-68d6c5891859 · outbound

This paper cites an unresolved cited work.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:19:51.171603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.845679Z digest=sha256:98c9bd1f638ff40aa7ae6f5a5085f3ad7b35a7cd6f2d947f86c4fee84c45d389

Observation 4f320ca9-36da-429a-af6c-1b54ca539e08 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Depth anything: Unleashing the power of large-scale unlabeled data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.161216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.849234Z digest=sha256:b5d251713afbea4946b7c23955fb5f7c6b8840e5dc470baa680a462a6bc9c780

Observation 13211660-7f98-442c-83ff-82e35a8f25bf · outbound

This paper cites Depth Anything V2.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Depth Anything V2

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.852543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.852543Z digest=sha256:2e48ece43e5f591a72affda0c37f5493dcb8ff9edb50a369d30412514178ea43

Observation 196b199d-7cb6-4631-9c88-298a747c27c4 · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.149945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.856204Z digest=sha256:0fd0a91a2ee8e45bbe5f00ce8ce02cedb656d681a0526eed04a90d1a54b9d614

Observation 272e37c4-4fb7-4e07-96d9-6d9bc9706969 · outbound

This paper cites 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.859468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.859468Z digest=sha256:19f2ed70a9e829bed03b1658d83e59709fb298b4a416599b3c0f6ed4318fe9e0

Observation 56c1656b-ef10-4ec8-a767-e3d4775ca4f4 · outbound

This paper cites Sigmoid loss for language image pre-training.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Sigmoid loss for language image pre-training

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.130053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.863313Z digest=sha256:2a2724b93c9fd49621deefb1c7b3109f2c3744939539de5ecea509d96b7b40ee

Observation b6f81b37-908c-4a2b-96df-44a8091e8f72 · outbound

This paper cites 3d-VLA: A 3d vision-language-action generative world model.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation 3d-VLA: A 3d vision-language-action generative world model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.118203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.866531Z digest=sha256:58e58da032444865c8ec0ed54f5e990ceb4ee795e035b65abdc0baf0ef904ce2

Observation 14bb0d37-9ed5-486d-bf0a-dc99246d4476 · outbound

This paper cites PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:50.869787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:50.869787Z digest=sha256:bbb4a5c37ff7859b6f12efc0e3a4d3bda30cb228fa5bcfa23ef61bdf8ee2bff1

Observation edcd3297-b5cc-4bb9-84ed-d4e0534d5acb · outbound

This paper cites Point cloud matters: Rethinking the impact of different observation spaces on robot learning.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Point cloud matters: Rethinking the impact of different observation spaces on robot learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.107391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.873425Z digest=sha256:22c830ac8d1f598505ffe3544735f1d825420e804c07721c0d5f555c126df723

Observation 3e64e5f7-6038-4b89-bbf9-b27dbdfd476c · outbound

This paper cites screwdriver.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation screwdriver

Reference 49

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T21:19:51.095692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.876610Z digest=sha256:db6dd05306903e725d84832a1677bcd5fca71314c83fc8e6a21655aaa4c96ef4

Observation 24c5840d-feee-424a-891a-b7c675010abf · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:51.084967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:50.880926Z digest=sha256:d9348e0e54a238f4c063e5eb97e723088074de4a826894e83d4298e1b398243f

Pith citing papers

Observation 3547738d-5d86-41df-a708-4632676f9416 · inbound

Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images cites this paper.

Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:03.491695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:44:39.322417Z digest=sha256:3f21844a50ac8070e40d8079bd397c6557cb32fce082c928bc990b0c975355da

Observation 471cf710-e970-4c45-9092-2bc0ff23515a · inbound

STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation cites this paper.

STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:21:25.453192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T11:26:53.331753Z digest=sha256:8a5ae876630de3e2a65ece2e59706650eb801d56424dc0b6da994af4b34d217b