REVIEW 6 cited by
Meta 3D Gen
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce Meta 3D Gen (3DGen), a new state-of-the-art, fast pipeline for text-to-3D asset generation. 3DGen offers 3D asset creation with high prompt fidelity and high-quality 3D shapes and textures in under a minute. It supports physically-based rendering (PBR), necessary for 3D asset relighting in real-world applications. Additionally, 3DGen supports generative retexturing of previously generated (or artist-created) 3D shapes using additional textual inputs provided by the user. 3DGen integrates key technical components, Meta 3D AssetGen and Meta 3D TextureGen, that we developed for text-to-3D and text-to-texture generation, respectively. By combining their strengths, 3DGen represents 3D objects simultaneously in three ways: in view space, in volumetric space, and in UV (or texture) space. The integration of these two techniques achieves a win rate of 68% with respect to the single-stage model. We compare 3DGen to numerous industry baselines, and show that it outperforms them in terms of prompt fidelity and visual quality for complex textual prompts, while being significantly faster.
Forward citations
Cited by 6 Pith papers
-
VesselTok: Tokenizing Vessel-like 3D Biomedical Graph Representations for Reconstruction and Generation
VesselTok learns compact continuous tokens of large tubular biomedical graphs from centerline points plus a fixed pseudo-radius, enabling reconstruction, generation, and link prediction across anatomies.
-
CustomX: Unified Character, Action, and Scene Customization in Video World Models
AniX generates controllable videos of a user-supplied character performing typed actions inside a user-supplied 3D scene by fine-tuning a pre-trained video generator on small locomotion datasets.
-
Meshtron: High-Fidelity, Artist-Like 3D Mesh Generation at Scale
Meshtron autoregressively generates 3D meshes with up to 64K faces at 1024-level coordinate resolution, a large scale increase over prior work, using an hourglass transformer and sliding-window inference.
-
Wavelet Latent Diffusion (Wala): Billion-Parameter 3D Generative Model with Compact Wavelet Encodings
Wavelet Latent Diffusion (WaLa) shrinks 3D shapes to 6,912-variable latent codes and trains billion-parameter diffusion models that generate 256^3 geometry in 2-4 seconds, claiming state-of-the-art results.
-
Video-Guided Foley Sound Generation with Multimodal Controls
A video-guided diffusion model generates synchronized foley sound from text, audio, and video controls, using joint training on noisy internet videos and professional sound-effect libraries to reach 48kHz output.
-
Material Anything: Generating Materials for Any 3D Object via Diffusion
Material Anything is a unified diffusion pipeline that generates PBR material maps (albedo, roughness, metallic, bump) for arbitrary 3D meshes using confidence masks to handle varying texture and lighting conditions.
Discussion (0). Continue with ORCID to comment.