Pith. sign in

REVIEW 8 cited by

MeshXL: Neural Coordinate Field for Generative 3D Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.20853 v2 pith:W535MT63 submitted 2024-05-31 cs.CV

classification cs.CV
keywords representationcoordinategenerationmeshmeshesmeshxlmodelsneural
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The polygon mesh representation of 3D data exhibits great flexibility, fast rendering speed, and storage efficiency, which is widely preferred in various applications. However, given its unstructured graph representation, the direct generation of high-fidelity 3D meshes is challenging. Fortunately, with a pre-defined ordering strategy, 3D meshes can be represented as sequences, and the generation process can be seamlessly treated as an auto-regressive problem. In this paper, we validate the Neural Coordinate Field (NeurCF), an explicit coordinate representation with implicit neural embeddings, is a simple-yet-effective representation for large-scale sequential mesh modeling. After that, we present MeshXL, a family of generative pre-trained auto-regressive models, which addresses the process of 3D mesh generation with modern large language model approaches. Extensive experiments show that MeshXL is able to generate high-quality 3D meshes, and can also serve as foundation models for various down-stream applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Sat2City generates explicit 3D city geometry and appearance from a height-map condition using cascaded latent diffusion on sparse voxel grids, beating prior methods on a new synthetic city dataset.

  2. Efficient Part-level 3D Object Generation via Dual Volume Packing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    From a single image, a 3D latent diffusion model generates all parts of an object at once by packing the part structure into two non-overlapping volumes.

  3. Nautilus: Locality-aware Autoencoder for Scalable Mesh Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Nautilus introduces a locality-preserving tokenization for meshes that lets autoregressive transformers generate high-fidelity meshes with up to 5,000 faces, outperforming prior methods.

  4. TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

    cs.CV 2024-12 conditional novelty 6.0 of 10

    TAR3D uses a triplane VQ-VAE to turn 3D shapes into discrete codebook tokens and a GPT-style transformer to generate those tokens autoregressively from text or image prompts.

  5. 3D Shape Tokenization via Latent Flow Matching

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Shape Tokens, a compact continuous 3D latent learned by fitting each shape's surface density with flow matching, match specialized baselines across reconstruction, CLIP, generation, and ray intersection tasks.

  6. Meshtron: High-Fidelity, Artist-Like 3D Mesh Generation at Scale

    cs.GR 2024-12 conditional novelty 6.0 of 10

    Meshtron autoregressively generates 3D meshes with up to 64K faces at 1024-level coordinate resolution, a large scale increase over prior work, using an hourglass transformer and sliding-window inference.

  7. LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Fine-tuning LLaMA-3.1-8B on an OBJ-as-text dataset lets one chat model both answer questions and generate simple 3D meshes, with no vocabulary expansion.

  8. LTM3D: Bridging Token Spaces for Conditional 3D Generation with Auto-Regressive Diffusion Framework

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A conditional 3D generation framework that combines masked autoencoding and diffusion in token space, with prefix learning and reconstruction-guided sampling, reports state-of-the-art results on ShapeNet and Objaverse.

Pith tools