REVIEW 7 cited by
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Synthesizing interactive 3D scenes from text is essential for gaming, virtual reality, and embodied AI. However, existing methods face several challenges. Learning-based approaches depend on small-scale indoor datasets, limiting the scene diversity and layout complexity. While large language models (LLMs) can leverage diverse text-domain knowledge, they struggle with spatial realism, often producing unnatural object placements that fail to respect common sense. Our key insight is that vision perception can bridge this gap by providing realistic spatial guidance that LLMs lack. To this end, we introduce Scenethesis, a training-free agentic framework that integrates LLM-based scene planning with vision-guided layout refinement. Given a text prompt, Scenethesis first employs an LLM to draft a coarse layout. A vision module then refines it by generating an image guidance and extracting scene structure to capture inter-object relations. Next, an optimization module iteratively enforces accurate pose alignment and physical plausibility, preventing artifacts like object penetration and instability. Finally, a judge module verifies spatial coherence. Comprehensive experiments show that Scenethesis generates diverse, realistic, and physically plausible 3D interactive scenes, making it valuable for virtual content creation, simulation environments, and embodied AI research.
Forward citations
Cited by 7 Pith papers
-
InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360{\deg} Image
InSpace generates complete structure-aware 3D indoor scenes (layout plus textured assets) from a single equirectangular 360° image via three-stage flow matching with view- and asset-selective attention.
-
TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
A training-free pipeline generates instance-level, physically interactive 3D tabletop scenes from text or one image, with a differentiable rotation optimizer and top-view spatial alignment for collision-free layouts.
-
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
VLM4D benchmarks spatiotemporal reasoning in VLMs and finds large gaps versus humans, with proposed methods showing partial improvement.
-
"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth
The authors construct the first Chinese youth-toxicity dataset, show that youth and adult perceptions of toxic language diverge, and report that adding contextual meta information improves detection accuracy.
-
Video Perception Models for 3D Scene Synthesis
VIPScene synthesizes 3D scenes by generating a video with Cosmos, reconstructing it with Fast3R, extracting objects with Grounded-SAM and MASt3R, and assembling them from Objaverse assets.
-
RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills
RobotSmith autonomously designs, 3D-prints, and uses task-specific tools for robotic manipulation, raising task success from 2.8% (no tool) to 50% in simulation.
-
Text2Villa: Hierarchical Generation of 3D Indoor Environments with Physics-Aware Analysis-by-Synthesis
Text2Villa generates multi-story villa-scale 3D indoor scenes from text by fine-tuning a layout model and solving a physics-aware closed-loop placement optimization.
Discussion (0). Sign in to comment.