REVIEW 8 cited by
MeshXL: Neural Coordinate Field for Generative 3D Foundation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, given its unstructured graph representation, the direct generation of high-fidelity 3D meshes is challenging. Fortunately, with a pre-defined ordering strategy, 3D meshes can be represented as sequences, and the generation process can be seamlessly treated as an auto-regressive problem. In this paper, we validate the Neural Coordinate Field (NeurCF), an explicit coordinate representation with implicit neural embeddings, is a simple-yet-effective representation for large-scale sequential mesh modeling. After that, we present MeshXL, a family of generative pre-trained auto-regressive models, which addresses the process of 3D mesh generation with modern large language model approaches. Extensive experiments show that MeshXL is able to generate high-quality 3D meshes, and can also serve as foundation models for various down-stream applications.
Forward citations
Cited by 8 Pith papers
-
Sat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent Diffusion
Sat2City generates explicit 3D city geometry and appearance from a height-map condition using cascaded latent diffusion on sparse voxel grids, beating prior methods on a new synthetic city dataset.
-
Efficient Part-level 3D Object Generation via Dual Volume Packing
From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.
-
Nautilus: Locality-aware Autoencoder for Scalable Mesh Generation
Nautilus introduces a locality-preserving tokenization for meshes that lets autoregressive transformers generate high-fidelity meshes with up to 5,000 faces, outperforming prior methods.
-
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction
TAR3D uses a triplane VQ-VAE to turn 3D shapes into discrete codebook tokens and a GPT-style transformer to generate those tokens autoregressively from text or image prompts.
-
3D Shape Tokenization via Latent Flow Matching
Shape Tokens, a compact continuous 3D latent learned by fitting each shape's surface density with flow matching, match specialized baselines across reconstruction, CLIP, generation, and ray intersection tasks.
-
Meshtron: High-Fidelity, Artist-Like 3D Mesh Generation at Scale
Meshtron autoregressively generates 3D meshes with up to 64K faces at 1024-level coordinate resolution, a large scale increase over prior work, using an hourglass transformer and sliding-window inference.
-
LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models
Fine-tuning LLaMA-3.1-8B on an OBJ-as-text dataset lets one chat model both answer questions and generate simple 3D meshes, with no vocabulary expansion.
-
LTM3D: Bridging Token Spaces for Conditional 3D Generation with Auto-Regressive Diffusion Framework
A conditional 3D generation framework that combines masked autoencoding and diffusion in token space, with prefix learning and reconstruction-guided sampling, reports state-of-the-art results on ShapeNet and Objaverse.
Discussion (0). Continue with ORCID to comment.