REVIEW 10 cited by
Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images
read the original abstract
Most image-based 3D object reconstructors assume that objects are fully visible, ignoring occlusions that commonly occur in real-world scenarios. In this paper, we introduce Amodal3R, a conditional 3D generative model designed to reconstruct 3D objects from partial observations. We start from a "foundation" 3D generative model and extend it to recover plausible 3D geometry and appearance from occluded objects. We introduce a mask-weighted multi-head cross-attention mechanism followed by an occlusion-aware attention layer that explicitly leverages occlusion priors to guide the reconstruction process. We demonstrate that, by training solely on synthetic data, Amodal3R learns to recover full 3D objects even in the presence of occlusions in real scenes. It substantially outperforms existing methods that independently perform 2D amodal completion followed by 3D reconstruction, thereby establishing a new benchmark for occlusion-aware 3D reconstruction.
Forward citations
Cited by 10 Pith papers
-
EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning
An end-to-end 3D editing framework achieves high-fidelity local edits from coarse bounding boxes and 2D image prompts using region-aware loss reweighting and a large-scale parts-derived training dataset.
-
3DMorph: Single-Image-Guided Local 3D Shape Editing and Morphing
3DMorph transfers local modifications from a single edited 2D image to the corresponding regions of a 3D mesh without training and supports shape morphing between original and edited versions.
-
EGM: Efficient Visual Grounding Language Models
EGM enables 8B VLMs to reach 91.4 IoU on RefCOCO at 737 ms latency, outperforming a 235B model at 4320 ms, by substituting volume of mid-quality tokens for model scale.
-
R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation
R3D2 trains a lightweight diffusion model on synthetic placements of 3DGS-generated assets to produce photorealistic insertions with consistent illumination into autonomous driving scenes.
-
Restore3D: Breathing Life into Broken Objects with Shape and Texture Restoration
Restore3D restores shape and texture of broken 3D objects via multi-view image refinement with a Mask Self-Perceiver and coarse-to-fine mesh reconstruction, outperforming baselines on synthetic and real benchmarks.
-
Occlusion-Robust Multi-Object Decoupling for Physics-Based Robotic Interaction
A pipeline combining SAM2 segmentation, 3D Gaussian Splatting, and joint Score Distillation Sampling with 2D/3D diffusion priors reconstructs decoupled multi-object geometries from occluded sparse views for MPM simulation.
-
Occlusion-Robust Multi-Object Decoupling for Physics-Based Robotic Interaction
A new pipeline for occlusion-robust multi-object 3D reconstruction from sparse views supports physics-based robotic interaction.
-
CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction
The paper's stated CA-World counterfactual claim is absent from the body, which instead describes the SAM3D-Phys pipeline for multi-object interactive reconstruction and simulation.
-
CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction
SAM3D-Phys recovers complete simulatable object geometries from incomplete real-world scene reconstructions by combining SAM3D generative priors with physics-constrained spatial optimization and mask-guided appearance...
-
Principles and Practice of Deep Representation Learning: or a Mathematical Theory of Memory
The book presents principles from optimization and information theory to explain deep network architectures and enable new interpretable models.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.