Pith. sign in

REVIEW 21 cited by

MegaPose: 6D Pose Estimation of Novel Objects via Render & Compare

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.06870 v1 pith:ERMVCOZY submitted 2022-12-13 cs.CV cs.RO

classification cs.CVcs.RO
keywords objectsnovelposeobjectapproachsyntheticdatasetestimation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce MegaPose, a method to estimate the 6D pose of novel objects, that is, objects unseen during training. At inference time, the method only assumes knowledge of (i) a region of interest displaying the object in the image and (ii) a CAD model of the observed object. The contributions of this work are threefold. First, we present a 6D pose refiner based on a render&compare strategy which can be applied to novel objects. The shape and coordinate system of the novel object are provided as inputs to the network by rendering multiple synthetic views of the object's CAD model. Second, we introduce a novel approach for coarse pose estimation which leverages a network trained to classify whether the pose error between a synthetic rendering and an observed image of the same object can be corrected by the refiner. Third, we introduce a large-scale synthetic dataset of photorealistic images of thousands of objects with diverse visual and shape properties and show that this diversity is crucial to obtain good generalization performance on novel objects. We train our approach on this large synthetic dataset and apply it without retraining to hundreds of novel objects in real images from several pose estimation benchmarks. Our approach achieves state-of-the-art performance on the ModelNet and YCB-Video datasets. An extensive evaluation on the 7 core datasets of the BOP challenge demonstrates that our approach achieves performance competitive with existing approaches that require access to the target objects during training. Code, dataset and trained models are available on the project page: https://megapose6d.github.io/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation

    cs.RO 2026-04 unverdicted novelty 7.0 of 10

    RoboWM-Bench evaluates video world models by converting their manipulation video predictions into executable actions validated in simulation, showing that visual plausibility does not guarantee physical executability.

  2. RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation

    cs.RO 2026-04 unverdicted novelty 7.0 of 10

    RoboWM-Bench evaluates video world models by converting their outputs into executable robot actions and running them on manipulation tasks, showing that physical inconsistencies remain common.

  3. Event6D: Event-based Novel Object 6D Pose Tracking

    cs.CV 2026-03 conditional novelty 7.0 of 10

    EventTrack6D tracks 6D poses of unseen objects from event cameras by reconstructing dense intensity and depth cues between frames, generalizing from synthetic training to real data at high speed.

  4. Emergence of a Shared Canonical Object Frame from In-the-Wild Videos

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    A coarse canonical mesh bottleneck plus multi-view consistency lets a shared object frame emerge from self-supervised training on in-the-wild videos without canonical labels or category conditioning.

  5. MF-UAVPose6D: A Model-Free Monocular 6-DoF Pose Estimation Framework for Fixed-Wing UAVs

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    MF-UAVPose6D estimates 6-DoF poses of fixed-wing UAVs from monocular RGB images without CAD models using heatmap center localization, Perspective-Aware Module, Dynamic Topological Sampling, and decoupled decoding on a...

  6. Pose Anything Anywhere:Model-free Object Poses from Arbitrary References

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    PANY is a multi-view transformer framework for model-free 6D object pose estimation from arbitrary sparse references that reports SOTA gains of +12% on YCB-V and +20% on LM-O.

  7. Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    HA-HOI produces physically plausible 4D HOI animations from monocular videos by anchoring object reconstruction to human motion and refining the result in a physics-based humanoid-object simulator.

  8. Semantic Prior Guided One-View 6D Pose Estimation for Novel Objects

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    OneViewAll achieves 92.5% ADD-0.1 accuracy on LINEMOD for novel object 6D pose estimation using only one real reference view by integrating category, symmetry, and patch-level semantic priors in a projection-equivaria...

  9. Semantic Prior Guided One-View 6D Pose Estimation for Novel Objects

    cs.CV 2026-05 conditional novelty 6.0 of 10

    OneViewAll reports 92.5% ADD-0.1 pose accuracy on LINEMOD from a single real reference RGB-D view, using projection-based refinement with mirror-fusion symmetry priors rather than CAD rendering.

  10. MAPRPose: Mask-Aware Proposal and Amodal Refinement for Multi-Object 6D Pose Estimation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    MAPRPose achieves state-of-the-art 76.5% Average Recall on the BOP benchmark for 6D pose estimation, outperforming FoundationPose by 3.1% AR while delivering a 43x speedup in multi-object inference.

  11. From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data

    cs.RO 2026-04 accept novelty 6.0 of 10

    Video-to-robot control methods cluster into three interface families, and the field’s main bottleneck is grounding video-derived predictions into dependable closed-loop robot behavior.

  12. SAM 3D: 3Dfy Anything in Images

    cs.CV 2025-11 unverdicted novelty 6.0 of 10

    SAM 3D reconstructs 3D objects from single images with geometry, texture, and pose using human-model annotated data at scale and synthetic-to-real training, achieving 5:1 human preference wins.

  13. One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Given one RGB-D photo of an unseen object, an AI-generated 3D mesh, aligned jointly in metric scale and pose, yields state-of-the-art one-shot 6D pose estimation on YCBInEOAT, TOYL, and LM-O.

  14. RadGS-Reg: Registering Spine CT with Biplanar X-rays via Joint 3D Radiative Gaussians Reconstruction and 3D/3D Registration

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A joint 3D Gaussian-splatting-based reconstruction and registration network registers spine CT to two X-rays with 1.14 mm mean error in 0.82 seconds on a small in-house set.

  15. PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A training-free geometry-only pipeline matches RGB images against rendered depth/normal maps to estimate 6D poses of unseen, textureless, and slightly defective objects from one image.

  16. From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data

    cs.RO 2026-04 accept novelty 5.0 of 10

    A survey introduces an interface-centric taxonomy for video-to-control methods in robotic manipulation and identifies the robotics integration layer as the central open challenge.

  17. DynamicPose: Real-time and Robust 6D Object Pose Tracking for Fast-Moving Cameras and Objects

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    DynamicPose maintains real-time 6D object pose during fast camera and object motion by combining VIO-based ROI compensation, depth-informed 2D tracking, and VIO-guided Kalman-filter pose prediction in a closed loop.

  18. IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    The claimed IDC-Net framework is absent; the body text is an unrelated instance-segmentation paper.

  19. Geometry-aware 4D Video Generation for Robot Manipulation

    cs.CV 2025-07 unverdicted novelty 5.0 of 10

    A geometry-aware 4D video generation model trained with cross-view pointmap alignment to produce spatio-temporally consistent future videos from novel viewpoints for robot manipulation.

  20. MAPRPose: Mask-Aware Proposal and Amodal Refinement for Multi-Object 6D Pose Estimation

    cs.CV 2026-04 unverdicted novelty 4.0 of 10

    MAPRPose reports 76.5% Average Recall on the BOP benchmark for multi-object 6D pose estimation, beating FoundationPose by 3.1% while running 43 times faster through mask-aware proposals and amodal refinement.

  21. World Action Models: A Survey

    cs.RO 2026-06 unverdicted novelty 3.0 of 10

    A survey that clarifies boundaries and organizes World Action Models by generation requirements and predictive substrates, identifying a trend toward generating less of the future.

Pith tools