REVIEW 2 cited by
StructDiffusion: Language-Guided Creation of Physically-Valid Structures using Unseen Objects
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Robots operating in human environments must be able to rearrange objects into semantically-meaningful configurations, even if these objects are previously unseen. In this work, we focus on the problem of building physically-valid structures without step-by-step instructions. We propose StructDiffusion, which combines a diffusion model and an object-centric transformer to construct structures given partial-view point clouds and high-level language goals, such as "set the table". Our method can perform multiple challenging language-conditioned multi-step 3D planning tasks using one model. StructDiffusion even improves the success rate of assembling physically-valid structures out of unseen objects by on average 16% over an existing multi-modal transformer model trained on specific structures. We show experiments on held-out objects in both simulation and on real-world rearrangement tasks. Importantly, we show how integrating both a diffusion model and a collision-discriminator model allows for improved generalization over other methods when rearranging previously-unseen objects. For videos and additional results, see our website: https://structdiffusion.github.io/.
Forward citations
Cited by 2 Pith papers
-
Cascaded Diffusion Models for Neural Motion Planning
A cascaded diffusion planner with a coarse global model, a local refiner, and a one-shot collision-patching step improves success rates by roughly 3 to 5 percentage points over prior learned planners in simulated navi...
-
Goal State Generation for Robotic Manipulation Based on Linguistically Guided Hybrid Gaussian Diffusion
A language-conditioned hybrid Gaussian diffusion network generates mug-hanging poses in simulation, then uses a gravity-based overlap removal step to produce collision-free target states.
Discussion (0). Continue with ORCID to comment.