Pith. sign in

REVIEW 1 cited by

Locate 3D: Real-World Object Localization via Self-Supervised Learning in 3D

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.14151 v1 pith:PMILS75C submitted 2025-04-19 cs.CV cs.AIcs.RO

classification cs.CVcs.AIcs.RO
keywords locatelearningself-supervisedcapabilitiesd-jepadatasetgeneralizationgrounding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present LOCATE 3D, a model for localizing objects in 3D scenes from referring expressions like "the small coffee table between the sofa and the lamp." LOCATE 3D sets a new state-of-the-art on standard referential grounding benchmarks and showcases robust generalization capabilities. Notably, LOCATE 3D operates directly on sensor observation streams (posed RGB-D frames), enabling real-world deployment on robots and AR devices. Key to our approach is 3D-JEPA, a novel self-supervised learning (SSL) algorithm applicable to sensor point clouds. It takes as input a 3D pointcloud featurized using 2D foundation models (CLIP, DINO). Subsequently, masked prediction in latent space is employed as a pretext task to aid the self-supervised learning of contextualized pointcloud features. Once trained, the 3D-JEPA encoder is finetuned alongside a language-conditioned decoder to jointly predict 3D masks and bounding boxes. Additionally, we introduce LOCATE 3D DATASET, a new dataset for 3D referential grounding, spanning multiple capture setups with over 130K annotations. This enables a systematic study of generalization capabilities as well as a stronger model.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. G$^2$TAM: Geometry Grounded Track Anything Model

    cs.CV 2026-07 accept novelty 6.5 of 10

    Spatially aligned geometric features serve as implicit memory so one model reconstructs scenes and produces promptable, cross-view consistent instance masks from unordered RGB only.

Pith tools