Pith. sign in

REVIEW 6 cited by

Cubify Anything: Scaling Indoor 3D Object Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.04458 v1 pith:4PMZOKJY submitted 2024-12-05 cs.CV

classification cs.CV
keywords ca-1mcubifycutrdetectionobjectobjectswhileanything
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We consider indoor 3D object detection with respect to a single RGB(-D) frame acquired from a commodity handheld device. We seek to significantly advance the status quo with respect to both data and modeling. First, we establish that existing datasets have significant limitations to scale, accuracy, and diversity of objects. As a result, we introduce the Cubify-Anything 1M (CA-1M) dataset, which exhaustively labels over 400K 3D objects on over 1K highly accurate laser-scanned scenes with near-perfect registration to over 3.5K handheld, egocentric captures. Next, we establish Cubify Transformer (CuTR), a fully Transformer 3D object detection baseline which rather than operating in 3D on point or voxel-based representations, predicts 3D boxes directly from 2D features derived from RGB(-D) inputs. While this approach lacks any 3D inductive biases, we show that paired with CA-1M, CuTR outperforms point-based methods - accurately recalling over 62% of objects in 3D, and is significantly more capable at handling noise and uncertainty present in commodity LiDAR-derived depth maps while also providing promising RGB only performance without architecture changes. Furthermore, by pre-training on CA-1M, CuTR can outperform point-based methods on a more diverse variant of SUN RGB-D - supporting the notion that while inductive biases in 3D are useful at the smaller sizes of existing datasets, they fail to scale to the data-rich regime of CA-1M. Overall, this dataset and baseline model provide strong evidence that we are moving towards models which can effectively Cubify Anything.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

    cs.AI 2026-07 conditional novelty 6.0 of 10

    SpatialCLI shows that a VLM can learn to use localization, segmentation, depth, and pose tools and then internalize the tool outputs into direct reasoning, improving both tool-enabled and tool-free spatial task performance.

  2. Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A 3D-aware VLM, RoboTracer, generates metric-grounded spatial traces for robot manipulation using scale supervision and metric-sensitive reinforcement rewards.

  3. RoboBrain 2.0 Technical Report

    cs.RO 2025-07 conditional novelty 6.0 of 10

    RoboBrain 2.0, a 7B/32B embodied vision-language model built on Qwen2.5-VL, reports state-of-the-art or near-top scores on several spatial and temporal reasoning benchmarks for robotics.

  4. BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection via Real-Time Multi-View Box Fusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    BoxFusion fuses per-frame 3D bounding box proposals from Cubify Anything and CLIP semantics into open-vocabulary 3D detections, reporting state-of-the-art AP among online methods without dense reconstruction.

  5. Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and Mapping

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Rooms from Motion produces a global 3D object map and camera poses from unposed photos by matching 3D bounding boxes across images instead of 2D keypoints.

  6. FlyMeThrough: Human-AI Collaborative 3D Indoor Mapping with Commodity Drones

    cs.HC 2025-08 conditional novelty 5.0 of 10

    A commodity-drone, RGB-only pipeline with human-AI annotation produces 3D indoor maps with localized points of interest, evaluated in 11 of 12 scanned buildings.

Pith tools