Pith. sign in

REVIEW 3 cited by

SPA: 3D Spatial-Awareness Enables Effective Embodied Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08208 v3 pith:NG37UDY4 submitted 2024-10-10 cs.CV cs.AIcs.LGcs.RO

classification cs.CVcs.AIcs.LGcs.RO
keywords embodiedrepresentationlearningspatialawarenessmodelresultsscenarios
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we introduce SPA, a novel representation learning framework that emphasizes the importance of 3D spatial awareness in embodied AI. Our approach leverages differentiable neural rendering on multi-view images to endow a vanilla Vision Transformer (ViT) with intrinsic spatial understanding. We present the most comprehensive evaluation of embodied representation learning to date, covering 268 tasks across 8 simulators with diverse policies in both single-task and language-conditioned multi-task scenarios. The results are compelling: SPA consistently outperforms more than 10 state-of-the-art representation methods, including those specifically designed for embodied AI, vision-centric tasks, and multi-modal applications, while using less training data. Furthermore, we conduct a series of real-world experiments to confirm its effectiveness in practical scenarios. These results highlight the critical role of 3D spatial awareness for embodied representation learning. Our strongest model takes more than 6000 GPU hours to train and we are committed to open-sourcing all code and model weights to foster future research in embodied representation learning. Project Page: https://haoyizhu.github.io/spa/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Lift3D-VLA integrates 3D point cloud encoding and temporal action modeling into Vision-Language-Action models, achieving higher success rates on simulated and real-world robotic manipulation tasks.

  2. VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A convolutional residual VQ-VAE action tokenizer trained on over 100x more data than prior work improves OpenVLA success rates and inference speed on several manipulation tasks.

  3. Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation

    cs.RO 2025-02 conditional novelty 5.0 of 10

    Adding an information bottleneck regularizer that penalizes I(X,Z) between fused input features and the latent representation improves average success rates in behavior cloning benchmarks, though gains depend on a per...

Pith tools