Pith. sign in

REVIEW 27 cited by

LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.09595 v1 pith:M4ATGVBZ submitted 2024-11-14 cs.LG cs.AIcs.CLcs.CV

LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

classification cs.LG cs.AIcs.CLcs.CV
keywords llmsmeshtextgenerationmeshesllama-meshmodelseffectively
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This work explores expanding the capabilities of large language models (LLMs) pretrained on text to generate 3D meshes within a unified model. This offers key advantages of (1) leveraging spatial knowledge already embedded in LLMs, derived from textual sources like 3D tutorials, and (2) enabling conversational 3D generation and mesh understanding. A primary challenge is effectively tokenizing 3D mesh data into discrete tokens that LLMs can process seamlessly. To address this, we introduce LLaMA-Mesh, a novel approach that represents the vertex coordinates and face definitions of 3D meshes as plain text, allowing direct integration with LLMs without expanding the vocabulary. We construct a supervised fine-tuning (SFT) dataset enabling pretrained LLMs to (1) generate 3D meshes from text prompts, (2) produce interleaved text and 3D mesh outputs as required, and (3) understand and interpret 3D meshes. Our work is the first to demonstrate that LLMs can be fine-tuned to acquire complex spatial knowledge for 3D mesh generation in a text-based format, effectively unifying the 3D and text modalities. LLaMA-Mesh achieves mesh generation quality on par with models trained from scratch while maintaining strong text generation performance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. One Video, One World: Turning Monocular Video into Physical 4D Scenes

    cs.CV 2026-06 unverdicted novelty 8.0

    OVOW reconstructs instance-level, simulation-ready 4D mesh scenes from monocular video via a four-stage training-free pipeline and introduces a new benchmark for structured Video-to-4D evaluation.

  2. Engine-Native Editable 3D World Reconstruction with Objects and Lighting

    cs.CV 2026-07 conditional novelty 7.0

    A UE5-derived dataset (Lumera-2K) and VLM pipeline parse single-image scenes into oriented object boxes and parametric light tuples, establishing a measurable benchmark for editable, light-aware 3D reconstruction.

  3. Sculpting NeRF Geometry: Human-Preference Fine-Tuning of a 3D-Aware Face GAN

    cs.CV 2026-06 unverdicted novelty 7.0

    Fine-tunes EG3D using a human-preference reward on NeRF density to improve face geometry, achieving 74.4% user preference in pairwise tests with FID rising from 4.09 to 6.66.

  4. QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning

    cs.GR 2026-05 unverdicted novelty 7.0

    QuadLink generates anisotropic quad-dominant meshes from point clouds via anchor prediction, centroid-conditioned linking, and quad-first assembly, supporting hybrid n-gon topology.

  5. LottieGPT: Tokenizing Vector Animation for Autoregressive Generation

    cs.CV 2026-04 unverdicted novelty 7.0

    LottieGPT tokenizes Lottie animations into compact sequences and fine-tunes Qwen-VL to autoregressively generate coherent vector animations from natural language or visual prompts, outperforming prior SVG models.

  6. MeshTailor: Cutting Seams via Generative Mesh Traversal

    cs.GR 2026-03 unverdicted novelty 7.0

    MeshTailor is a mesh-native generative model that uses ChainingSeams serialization and a dual-stream transformer with pointer layers to trace coherent seams vertex-by-vertex on 3D surfaces.

  7. PartDiffuser: Part-wise 3D Mesh Generation via Discrete Diffusion

    cs.CV 2025-11 unverdicted novelty 7.0

    PartDiffuser is a semi-autoregressive discrete diffusion framework that generates high-fidelity 3D meshes from point clouds by combining inter-part autoregression with intra-part parallel diffusion using a part-aware ...

  8. GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement

    cs.GR 2026-07 conditional novelty 6.5

    GReFEM shows MLLMs zero-shot isolate load-activated geometric features for volumetric mesh refinement with higher precision than matched-budget geometric heuristics.

  9. Nexus: Native Mesh Generation with Diffusion

    cs.CV 2026-07 conditional novelty 6.0

    Nexus replaces autoregressive mesh serialization with two coupled diffusion models — octree vertex generation and a latent topology generator — claiming stronger geometry and perceptual quality on Objaverse and Toys4K.

  10. LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow

    cs.GR 2026-07 conditional novelty 6.0

    Factorizing mesh generation into a vertex flow then a vertex-conditioned topology flow improves fidelity and enables part-wise high-res synthesis and topology-adaptive editing.

  11. ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

    cs.CV 2026-07 unverdicted novelty 6.0

    ELSA3D introduces elastic semantic anchoring via sparse anchor tokens and a scale-aware octree tokenizer to unify 3D generation and captioning at reduced computational cost.

  12. Variance Reduction for Expectations with Diffusion Teachers

    cs.LG 2026-05 unverdicted novelty 6.0

    CARV amortizes upstream diffusion teacher costs over noise resamples with timestep importance sampling and stratified-inverse-CDF sampling, delivering 2-3x effective compute gains in text-to-3D experiments and order-o...

  13. PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects

    cs.CV 2026-05 unverdicted novelty 6.0

    PhysX-Omni unifies simulation-ready 3D asset generation across rigid, deformable, and articulated objects via a new geometry representation, the PhysXVerse dataset, and the PhysX-Bench evaluation suite.

  14. QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning

    cs.GR 2026-05 unverdicted novelty 6.0

    QuadLink generates anisotropic quad-dominant meshes from point clouds via a hybrid centroid-conditioned vertex linking model and a Tri-to-Quad data conversion operator.

  15. TOPOS: High-Fidelity and Efficient Industry-Grade 3D Head Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    TOPOS creates high-fidelity 3D heads with fixed industry topology from single images via a specialized VAE with Perceiver Resampler and a rectified flow transformer.

  16. Beyond Spatial Compression: Interface-Centric Generative States for Open-World 3D Structure

    cs.LG 2026-05 unverdicted novelty 6.0

    C2LT-3D factorizes 3D tokenization into canonical local geometry, partition-conditioned context, and relational seam variables to make latent states operational for assembly-level validation and repair in open-world m...

  17. UniRecGen: Unifying Multi-View 3D Reconstruction and Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    UniRecGen unifies reconstruction and generation via shared canonical space and disentangled cooperative learning to produce complete, consistent 3D models from sparse views.

  18. Twisted Fiber Bundle Codes over Group Algebras

    quant-ph 2026-04 unverdicted novelty 6.0

    Singular chain-compatible fiber twists over group algebras can increase CSS encoded dimension k at fixed blocklength n while examples keep distance d unchanged.

  19. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

    cs.CV 2026-03 unverdicted novelty 6.0

    HVG-3D uses a 3D-aware diffusion architecture with ControlNet to synthesize high-fidelity hand-object interaction videos from 3D control signals, achieving state-of-the-art spatial fidelity and temporal coherence on t...

  20. LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow

    cs.GR 2026-07 conditional novelty 5.0

    LATO.2 factorizes mesh generation into a vertex-generation flow and a vertex-conditioned connectivity flow, beating joint-latent and autoregressive baselines on geometric fidelity and connectivity quality.

  21. Sat2City v2: Native 3D City Asset Generation from a Single Satellite Image

    cs.CV 2026-06 unverdicted novelty 5.0

    Sat2City v2 adapts a pretrained native 3D latent model to generate controllable textured 3D city assets from satellite images via geometry flow fine-tuning and anchored texturing on a collected real dataset.

  22. Variance Reduction for Expectations with Diffusion Teachers

    cs.LG 2026-05 unverdicted novelty 5.0

    CARV introduces a hierarchical Monte Carlo estimator with amortized reuse, importance sampling, and stratification that yields 2-3x effective compute gains on diffusion-teacher pipelines while cutting gradient varianc...

  23. EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers

    cs.CV 2026-05 unverdicted novelty 5.0

    EVA01 introduces a Mixture-of-Transformers model that natively adds 3D mesh understanding, generation, and multi-turn editing to MLLMs by decoupling understanding and generation experts with shared global self-attention.

  24. SynVA: A Modular Toolkit for Vessel Generation and Aneurysm Editing

    cs.CV 2026-05 unverdicted novelty 5.0

    SynVA toolkit generates realistic vascular meshes and anatomically plausible aneurysms, releasing 50,000 labeled samples for medical vision tasks.

  25. CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models

    cs.CV 2026-01 unverdicted novelty 5.0

    CG-MLLM is a multimodal LLM using a Mixture-of-Transformer architecture with separate TokenAR and BlockAR components integrated with a pre-trained vision-language backbone and 3D VAE to enable 3D captioning and high-f...

  26. PartDiffuser: Part-wise 3D Mesh Generation via Discrete Diffusion

    cs.CV 2025-11 conditional novelty 5.0

    A part-wise semi-autoregressive discrete diffusion model for point-cloud-to-mesh generation that separates global structure from local detail, beating prior SOTA on Objaverse.

  27. QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning

    cs.GR 2026-05 unverdicted novelty 4.0

    QuadLink generates anisotropic quad-dominant meshes from point clouds via autoregressive anchor prediction and centroid-conditioned linking, with a Tri-to-Quad data converter and quad-first assembly.