REVIEW 5 major objections 5 minor 97 references
MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation
T0 review · 5 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A dataset of 40 million captions for 8.9 million 3D assets, with five description levels, is claimed to be the largest 3D caption dataset and to beat prior captions by 72–73% in preference.
desk verdict A genuinely large, well-engineered captioning resource whose headline quality numbers are inflated by a length confound, and whose artifacts are not yet released. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the five-stage MARVEL annotation pipeline. It renders each asset from four fixed viewpoints (azimuth 0°, 90°, 180°, 270° at 30° elevation), filters noisy user metadata with Mistral-Nemo, asks InternVL2-40B to write a dense description covering components, geometry, materials, colors, and context, then uses Qwen2.5-72B to compress that description into five hierarchical levels, ending with Qwen2.5-14B ethical filtering. This chain converts a 3D asset into 44.5 million captions. The downstream text-to-3D system is a two-stage chain: LoRA fine-tuned Stable Diffusion 3.5 produces reconstruction-friendly images from these captions, and Stable Fast 3D lifts the image to a textured mesh.
What would settle it
Take 1,000 randomly sampled objects from Objaverse-XL, generate MARVEL-style captions from the four fixed views, and also from a denser orbit rendering (e.g., 12 views or a turntable sweep); if human experts or GPT-4 judge the denser-view captions clearly more accurate, or if objects known to be thin or heavily occluded fail systematically under the four-view protocol, the assumption that four views suffice would be refuted.
Extended reading notes
Core claim
The paper's central claim is that high-quality 3D captions can be produced automatically at a scale that was previously impractical, and that this scale is what unblocks high-fidelity text-to-3D. It reports that MARVEL-40M+ contains 44,510,515 captions for 8,902,103 objects aggregated from seven datasets, making it the largest 3D caption dataset to date. Evaluated on 5,000 samples by GPT-4 and 1,000 samples by five human reviewers, its Level-4 captions are preferred over CAP3D, 3D-Topia, and Kabra captions 72.41% and 73.40% of the time respectively, and its Level-1 captions reach 84.70% (GPT-4) and 82.80% (human) caption accuracy. On the generation side, MARVEL-FX3D, which fine-tunes Stable Diffusion 3.5 on the dataset and passes the resulting image through Stable Fast 3D, reports the highest prompt fidelity (7.71/10) and overall preference (6.94/10) among compared methods while producing textured meshes in about 15 seconds.
Load-bearing premise
The load-bearing premise is that four fixed rendered views (front, back, left, right) give the vision-language model enough information to describe all 8.9M objects correctly, including thin, occluded, or viewpoint-ambiguous ones.
Editorial extensions
If this is right
- A single dataset can now supply training captions for 8.9M objects from seven sources, so text-to-3D models no longer need to reconcile mismatched short captions from different dataset formats.
- Because each asset has five description levels, downstream systems can select the level matched to their cost and fidelity needs, from 150–200-word reconstructions to 10–20-word prototyping tags.
- The annotation pipeline is built from open-source models and costs roughly $2,700–$3,000 to annotate the 800K-sample Objaverse subset, making dataset expansion far cheaper than human annotation.
- Fine-tuning Stable Diffusion 3.5 on MARVEL captions improves prompt fidelity and overall preference over training on CAP3D captions or no fine-tuning, and does so while keeping generation time near 15 seconds.
Reading between the lines
- An implication the paper leaves implicit is that the metadata-filtering step could be generalized to datasets without user metadata by using classifier or taxonomy labels, in which case the value of metadata injection should be reported as a per-domain statistic rather than only qualitatively.
- The five-level hierarchy is likely to be useful beyond generation, for example as a controlled way to probe how much caption detail a retrieval or editing system needs; the paper does not test those downstream tasks.
- The reported win rates are relative preferences, so a natural next benchmark is to measure absolute caption usefulness, such as improvement in reconstruction accuracy or text-to-3D retrieval recall, on a fixed test set.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MARVEL-40M+, a large-scale multi-level caption dataset covering roughly 8.9 million 3D assets and 44.5 million captions aggregated from seven 3D datasets. Captions are produced by a five-stage pipeline: four-view rendering, human-metadata filtering, dense description generation with InternVL2-40B, multi-level elaboration with Qwen2.5-72B, and ethical filtering. The paper also introduces MARVEL-FX3D, a two-stage text-to-3D pipeline that fine-tunes Stable Diffusion 3.5 on the new captions and uses Stable Fast 3D for mesh generation. The central claims are that MARVEL-40M+ outperforms existing 3D caption datasets in annotation quality and linguistic diversity, with reported GPT-4 and human win rates of 72.41% and 73.40%, and that MARVEL-FX3D improves prompt fidelity and overall preference over existing text-to-3D methods.
Significance. If validated, this is a potentially valuable resource: it is substantially larger than existing 3D caption datasets, uses only open-source models, provides a cost analysis of the annotation pipeline, and offers a five-level annotation structure with qualitative examples across diverse domains. The inclusion of failure cases and limitations is a strength. However, the headline quality claims currently rest on evaluation protocols that do not control for caption length or report statistical reliability, so the magnitude of the claimed advantage is not yet established.
major comments (5)
- [Section 4.1, Table 2] The statement that Level-4 annotations were chosen because their average length is 'similar' to baseline datasets is not supported by the table: MARVEL Level 4 averages 44 words, versus 16 for Cap3D, 29 for 3D-Topia, and 5 for Kabra. The win-rate task asks GPT-4 and human judges to select the caption that best matches the rendered 3D images; longer captions can mention more objects, attributes, colors, and contextual details, so even a partially correct longer caption has more opportunities to appear correct. The headline win rates of 72.41% and 73.40% are therefore confounded by length and may partly measure verbosity rather than fidelity. The paper should provide a length-controlled or length-matched evaluation, such as truncated captions, per-attribute precision/recall, or normalized informativeness, together with confidence intervals.
- [Section 4.1, Table 3] The caption-accuracy comparison uses MARVEL Level 1 captions averaging 170 words against baselines averaging 5-29 words. Asking judges whether 'all' mentioned attributes are correct is a reasonable way to penalize verbosity, but the paper does not report how judges handle attributes that are plausible but not visually verifiable from four renders, nor does it report inter-annotator agreement or confidence intervals for the 250-sample human evaluation. Without these details, the claimed 84.70% GPT-4 and 82.80% human accuracy rates are not statistically grounded, and the comparison across very different caption lengths remains difficult to interpret.
- [Section 4.2, Table 4] The claim that MARVEL-FX3D outperforms state-of-the-art text-to-3D methods rests on 50 prompts scored by five users. Several reported differences are within one standard deviation (for example, visual quality 6.58 vs 6.47 and geometric consistency 7.20 vs 7.25), and no significance tests or inter-rater agreement measures are reported. In addition, the baselines are adapted with different training budgets and implementations (LucidDreamer for 3k steps, DreamFusion and HiFA via threestudio for 10k and 24k steps), so it is unclear whether the comparison is apples-to-apples. The paper should report paired significance tests, confidence intervals, and standardized baseline configurations.
- [Section 4.3A] The human-metadata contribution is a stated core contribution, including the abstract's claim that metadata 'reduce VLM hallucinations,' but the ablation is qualitative only. Figure 5 and the supplementary examples show selected successes, but no quantitative measurement such as hallucination rates, error counts, or a randomized ablation over a representative sample is provided. The claim that metadata injection reduces hallucinations is therefore not yet established at dataset scale.
- [Section 3.1 and Section 5] The four-view rendering protocol (azimuths 0, 90, 180, 270 degrees at elevation 30) is assumed to provide sufficient visual information for InternVL2-40B to caption all 8.9 million assets. Section 5 acknowledges failure modes such as misidentification of thin objects, side-view confusion, numerical imprecision, and directional errors, but the paper never quantifies how frequently these failures occur or how they are distributed across the seven source datasets. Since the dataset-wide quality claim depends on caption correctness across a highly heterogeneous corpus, a failure-rate estimate or category-level breakdown is needed before the 'high-quality annotation' claim can be assessed.
minor comments (5)
- [Section 4.1] The text states that the trend extends to 'average word length,' but Table 2 does not contain an average word length column; the table reports average caption length, MTLD, unigram and bigram counts, and GPT-4/human win rates.
- [Section 4.1] The MTLD comparison sentence contains citation mismatches: 'higher than Cap3D [28]' and 'higher than 3D-Topia[53]' appear to swap the reference numbers, since [28] is 3D-Topia and [53] is Cap3D in the bibliography.
- [Supplementary Section 9.5] The supplementary text refers to 'Table 5' when discussing the inter-level semantic retention ablation, but the corresponding result appears as Table 6 in the main paper; this cross-reference should be corrected.
- [Table 3] The table caption contains a typo: 'desite' should be 'despite.'
- [General] For a dataset paper, the manuscript should state clearly where and how the dataset and the annotation pipeline code will be released, including licenses and any restrictions inherited from the source datasets; the current project page URL alone is not sufficient.
Circularity Check
Minor self-referential inter-level retention check; main annotation and TT3D claims are externally benchmarked.
-
other
[Section 3.1 (Multi-Level Visual Elaboration) / Section 4.3B / Table 6 / Section 9.5]
"Qwen2 [85] then processes these descriptions into five hierarchical levels, progressively compressing different aspects of the 3D assets. ... Results in Table 6 show strong semantic retention from Levels 1-4, demonstrating effective compression while preserving meaning."
Levels 2-5 are generated by prompting Qwen2.5 to compress the same Level 1 description, and the paper reports sentence-BERT cosine similarity between adjacent levels as evidence of 'effective compression.' Because each lower level is, by construction, a compressed paraphrase of the same source text, high embedding similarity between parent and child is largely built into the generation procedure; the measurement mostly confirms that the LLM followed the compression instruction.
full rationale
The central claims of MARVEL-40M+ are evaluated against external judges and external baselines, so they are not circular. Annotation-quality win rates (72.41% GPT-4, 73.40% human) are obtained by asking GPT-4 and human raters to compare captions against four multi-view renders; these judges are not the models that generated the captions, and the baselines are prior datasets. Caption accuracy (Table 3) is likewise assessed by external GPT-4 and human reviewers checking captions against rendered views, which penalizes MARVEL's longer Level-1 captions more than short baselines rather than forcing a win. MARVEL-FX3D's prompt-fidelity and overall-preference scores come from human evaluation against Shap-E, DreamFusion, HiFA, and LucidDreamer, again an external comparison. The only near-circular element is the inter-level semantic-retention study (Section 4.3B, Table 6): since Levels 2-5 are generated by compressing Level 1, cosine similarity between adjacent levels partly reflects the construction of the hierarchy rather than an independent quality measure. This does not support the paper's headline claims and is a secondary ablation, so it only slightly raises the circularity score. The paper's own Section 5 candidly lists VLM limitations (numerical precision, thin-object misidentification, ambiguous views), and the verified length mismatch in Table 2 (44 words vs 16/29/5) is a statistical-validity concern, not a circular-derivation concern. No load-bearing self-citations were found; the two overlapping-author references [35,36] are contextual rather than foundational.
Assumptions & free parameters
free parameters (3)
- Level word-count ranges =
L1 150-200, L2 100-150, L3 50-100, L4 ~30, L5 10-20
- Rendering camera settings =
azimuth 0/90/180/270 degrees, elevation 30, distance 1.5
- LoRA rank and alpha =
rank=4, alpha=4
assumptions (6)
- domain assumption Pretrained InternVL2-40B can accurately describe 3D assets from four rendered views.
- domain assumption Filtered human metadata from source datasets is accurate and improves annotation quality.
- domain assumption GPT-4 and human evaluators can reliably judge caption-image alignment and caption accuracy.
- domain assumption Stable Fast 3D is a reliable pretrained image-to-3D converter.
- domain assumption Sentence-BERT cosine similarity measures semantic retention across caption levels.
- domain assumption MTLD and n-gram statistics reflect annotation quality.
Cite this review
Pith. "Pith review of MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation." pith.science (2026). https://pith.science/paper/ILTQ4PMS
@misc{pith2026241117945,
author = {Pith},
title = {Pith review of: MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation},
year = {2026},
howpublished = {\url{https://pith.science/paper/ILTQ4PMS}},
note = {Machine review of arXiv:2411.17945}
}
read the original abstract
Generating high-fidelity 3D content from text prompts remains a significant challenge in computer vision due to the limited size, diversity, and annotation depth of the existing datasets. To address this, we introduce MARVEL-40M+, an extensive dataset with 40 million text annotations for over 8.9 million 3D assets aggregated from seven major 3D datasets. Our contribution is a novel multi-stage annotation pipeline that integrates open-source pretrained multi-view VLMs and LLMs to automatically produce multi-level descriptions, ranging from detailed (150-200 words) to concise semantic tags (10-20 words). This structure supports both fine-grained 3D reconstruction and rapid prototyping. Furthermore, we incorporate human metadata from source datasets into our annotation pipeline to add domain-specific information in our annotation and reduce VLM hallucinations. Additionally, we develop MARVEL-FX3D, a two-stage text-to-3D pipeline. We fine-tune Stable Diffusion with our annotations and use a pretrained image-to-3D network to generate 3D textured meshes within 15s. Extensive evaluations show that MARVEL-40M+ significantly outperforms existing datasets in annotation quality and linguistic diversity, achieving win rates of 72.41% by GPT-4 and 73.40% by human evaluators. Project page is available at https://sankalpsinha-cmos.github.io/MARVEL/.
Figures
Figures from the paper (21 more)
Reference graph
Works this paper leans on
-
[1]
https: //www.blender.org
Blender - a 3d modelling and rendering software. https: //www.blender.org. 3
-
[2]
https:// huggingface.co/spaces/opencompass/open_ vlm_leaderboard
Openvlm leaderboard - a hugging face space. https:// huggingface.co/spaces/opencompass/open_ vlm_leaderboard. 3, 4
-
[3]
https : / / huggingface
Stable diffusion 3.5 large - huggingface. https : / / huggingface . co / stabilityai / stable - diffusion-3.5-large. 2, 3, 4, 5, 7, 8, 16
-
[4]
Chatgpt vs
Mohammed Aldeen, Joshua Luo, Ashley Lian, Venus Zheng, Allen Hong, Preethika Yetukuri, and Long Cheng. Chatgpt vs. human annotators: A comprehensive analysis of chatgpt for text annotation. In2023 International Conference on Ma- chine Learning and Applications (ICMLA) , pages 602–609. IEEE, 2023. 3
2023
-
[5]
DeepFloyd IF: a novel state- of-the-art open-source text-to-image model with a high de- gree of photorealism and language understanding
DeepFloyd Lab at StabilityAI. DeepFloyd IF: a novel state- of-the-art open-source text-to-image model with a high de- gree of photorealism and language understanding. https: //www.deepfloyd.ai/deepfloyd- if , 2023. Re- trieved on 2023-11-08. 3, 5
2023
-
[6]
Modeling for text compression
Timothy Bell, Ian H Witten, and John G Cleary. Modeling for text compression. ACM Computing Surveys (CSUR), 21 (4):557–591, 1989. 8
1989
-
[7]
Sf3d: Stable fast 3d mesh reconstruction with uv- unwrapping and illumination disentanglement, 2024
Mark Boss, Zixuan Huang, Aaryaman Vasishta, and Varun Jampani. Sf3d: Stable fast 3d mesh reconstruction with uv- unwrapping and illumination disentanglement, 2024. 2, 3, 4, 5, 6, 7
2024
-
[8]
An es- timate of an upper bound for the entropy of english
Peter F Brown, Stephen A Della Pietra, Vincent J Della Pietra, Jennifer C Lai, and Robert L Mercer. An es- timate of an upper bound for the entropy of english. Compu- tational Linguistics, 18(1):31–40, 1992. 6
1992
Show all 97 references
-
[9]
Frankland, Thomas L
Declan Campbell, Sunayana Rane, Tyler Giallanza, Nicol`o De Sabbata, Kia Ghods, Amogh Joshi, Alexander Ku, Steven M. Frankland, Thomas L. Griffiths, Jonathan D. Co- hen, and Taylor W. Webb. Understanding the limits of vision language models through the lens of the binding problem,
-
[10]
Chang, Thomas A
Angel X. Chang, Thomas A. Funkhouser, Leonidas J. Guibas, Pat Hanrahan, Qi-Xing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, L. Yi, and Fisher Yu. Shapenet: An information-rich 3d model repository. ArXiv, abs/1512.03012, 2015. 2, 3, 5, 15, 24
2015 arXiv
-
[11]
Pali- x: On scaling up a multilingual vision and language model,
Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Jialin Wu, Carlos Riquelme Ruiz, Sebastian Goodman, Xiao Wang, Yi Tay, Siamak Shak- eri, Mostafa Dehghani, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang...
-
[12]
Single-view 3d scene reconstruc- tion with high-fidelity shape and texture
Yixin Chen, Junfeng Ni, Nan Jiang, Yaowei Zhang, Yixin Zhu, and Siyuan Huang. Single-view 3d scene reconstruc- tion with high-fidelity shape and texture. In 2024 Interna- tional Conference on 3D Vision (3DV) , pages 1456–1467. IEEE, 2024. 2
2024
-
[13]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv prepri...
2023 arXiv
-
[14]
Text-to-3d using gaussian splatting, 2024
Zilong Chen, Feng Wang, Yikai Wang, and Huaping Liu. Text-to-3d using gaussian splatting, 2024. 3
2024
-
[15]
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhang- wei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821, 2024. 2, 3, 4, 5, 8, 14
2024 arXiv
-
[16]
Abo: Dataset and benchmarks for real-world 3d object understand- ing
Jasmine Collins, Shubham Goel, Kenan Deng, Achlesh- war Luthra, Leon Xu, Erhan Gundogdu, Xi Zhang, Tomas F Yago Vicente, Thomas Dideriksen, Himanshu Arora, Matthieu Guillaumin, and Jitendra Malik. Abo: Dataset and benchmarks for real-world 3d object understand- ing. In IEEE/CV...
2022
-
[17]
Objaverse-XL: A universe of 10m+ 3d objects
Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram V oleti, Samir Yitzhak Gadre, Eli VanderBilt, Anirud- dha Kembhavi, Carl V ondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Obja...
2023
-
[18]
Objaverse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...
2023
-
[19]
Reflecting reality: Enabling diffusion models to produce faithful mirror reflections
Ankit Dhiman, Manan Shah, Rishubh Parihar, Yash Bhalgat, Lokesh R Boregowda, and R Venkatesh Babu. Reflecting reality: Enabling diffusion models to produce faithful mirror reflections. arXiv preprint arXiv:2409.14677, 2024. 2 9
2024 arXiv
-
[20]
McHugh, and Vincent Vanhoucke
Laura Downs, Anthony Francis, Nate Koenig, Brandon Kin- man, Ryan Hickman, Krista Reymann, Thomas B. McHugh, and Vincent Vanhoucke. Google scanned objects: A high- quality dataset of 3d scanned household items, 2022. 2, 3, 5, 13, 15, 27
2022
-
[21]
Scaling rectified flow trans- formers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yan- nik Marek, and Robin Rombach. Scaling rectified flow tr...
2024
-
[22]
Compgs: Unleashing 2d compositional- ity for compositional text-to-3d via dynamically optimizing 3d gaussians, 2024
Chongjian Ge, Chenfeng Xu, Yuanfeng Ji, Chensheng Peng, Masayoshi Tomizuka, Ping Luo, Mingyu Ding, Varun Jam- pani, and Wei Zhan. Compgs: Unleashing 2d compositional- ity for compositional text-to-3d via dynamically optimizing 3d gaussians, 2024. 2
2024
-
[23]
threestudio: A unified framework for 3d content generation
Yuan-Chen Guo, Ying-Tian Liu, Ruizhi Shao, Christian Laforte, Vikram V oleti, Guan Luo, Chia-Hao Chen, Zi- Xin Zou, Chen Wang, Yan-Pei Cao, and Song-Hai Zhang. threestudio: A unified framework for 3d content generation. https://github.com/threestudio- project/ threestudio, 2023. 16
2023
-
[24]
Flex3d: Feed-forward 3d generation with flexible reconstruction model and input view curation
Junlin Han, Jianyuan Wang, Andrea Vedaldi, Philip Torr, and Filippos Kokkinos. Flex3d: Feed-forward 3d generation with flexible reconstruction model and input view curation. arXiv preprint arXiv:2410.00890, 2024. 3
2024 arXiv
-
[25]
Classifier-free diffusion guidance, 2022
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance, 2022. 16
2022
-
[26]
What’s left can’t be right – the remaining positional incompetence of contrastive vision-language models, 2023
Nils Hoehing, Ellen Rushe, and Anthony Ventresque. What’s left can’t be right – the remaining positional incompetence of contrastive vision-language models, 2023. 8
2023
-
[27]
Viewdiff: 3d-consistent image generation with text-to-image models
Lukas H ¨ollein, Aljaˇz Boˇziˇc, Norman M¨uller, David Novotny, Hung-Yu Tseng, Christian Richardt, Michael Zollh ¨ofer, and Matthias Nießner. Viewdiff: 3d-consistent image generation with text-to-image models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...
-
[28]
3dtopia: Large text-to-3d gener- ation model with hybrid diffusion priors, 2024
Fangzhou Hong, Jiaxiang Tang, Ziang Cao, Min Shi, Tong Wu, Zhaoxi Chen, Shuai Yang, Tengfei Wang, Liang Pan, Dahua Lin, and Ziwei Liu. 3dtopia: Large text-to-3d gener- ation model with hybrid diffusion priors, 2024. 2, 3, 4, 5, 6, 7, 14, 17, 18, 19, 20, 21
2024
-
[29]
Lrm: Large reconstruction model for single image to 3d
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan. Lrm: Large reconstruction model for single image to 3d. arXiv preprint arXiv:2311.04400, 2023. 2, 3
2023 arXiv
-
[30]
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021. 5, 7
2021 arXiv
-
[31]
Make-a-shape: a ten-million-scale 3D shape model
Ka-Hei Hui, Aditya Sanghi, Arianna Rampini, Kamal Rahimi Malekshan, Zhengzhe Liu, Hooman Shayani, and Chi-Wing Fu. Make-a-shape: a ten-million-scale 3D shape model. In Proceedings of the 41st International Conference on Machine Learning, pages 20660–20681. PMLR, 2024. 2
2024
-
[32]
A survey on text-to-3d contents generation in the wild
Chenhan Jiang. A survey on text-to-3d contents generation in the wild. arXiv preprint arXiv:2405.09431, 2024. 2, 3
2024 arXiv
-
[33]
Shap-e: Generating condi- tional 3d implicit functions, 2023
Heewoo Jun and Alex Nichol. Shap-e: Generating condi- tional 3d implicit functions, 2023. 3, 6, 7, 16, 30, 31
2023
-
[34]
Rishabh Kabra, Loic Matthey, Alexander Lerchner, and Niloy J. Mitra. Leveraging vlm-based pipelines to annotate 3d objects. In Proceedings of the 41st International Confer- ence on Machine Learning. PMLR, 2024. 2, 3, 4, 5, 6, 7, 14, 17, 18, 19, 20, 21
2024
-
[35]
Cad-signet: Cad language inference from point clouds using layer-wise sketch instance guided attention
Mohammad Sadil Khan, Elona Dupont, Sk Aziz Ali, Kseniya Cherenkova, Anis Kacem, and Djamila Aouada. Cad-signet: Cad language inference from point clouds using layer-wise sketch instance guided attention. In In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP...
2024
-
[36]
Text2cad: Generating sequential cad models from beginner-to-expert level text prompts
Mohammad Sadil Khan, Sankalp Sinha, Talha Uddin Sheikh, Didier Stricker, Sk Aziz Ali, and Muhammad Ze- shan Afzal. Text2cad: Generating sequential cad models from beginner-to-expert level text prompts. In Advances in Neural Information Processing Systems, 2024. 2
2024
-
[37]
Geometry-aware score dis- tillation via 3d consistent noising and gradient consistency modeling, 2024
Min-Seop Kwak, Donghoon Ahn, Ines Hyeonsu Kim, Jin- Hwa Kim, and Seungryong Kim. Geometry-aware score dis- tillation via 3d consistent noising and gradient consistency modeling, 2024. 3
2024
-
[38]
Generative ai meets 3d: A survey on text-to-3d in aigc era
Chenghao Li, Chaoning Zhang, Atish Waghwase, Lik-Hang Lee, Francois Rameau, Yang Yang, Sung-Ho Bae, and Choong Seon Hong. Generative ai meets 3d: A survey on text-to-3d in aigc era. arXiv preprint arXiv:2305.06131 ,
-
[39]
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML,
-
[40]
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML,
-
[41]
Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model,
Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan Sunkavalli, Greg Shakhnarovich, and Sai Bi. Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model,
-
[42]
Dual3d: Efficient and consistent text-to-3d generation with dual-mode multi-view latent diffusion, 2024
Xinyang Li, Zhangyu Lai, Linning Xu, Jianfei Guo, Liu- juan Cao, Shengchuan Zhang, Bo Dai, and Rongrong Ji. Dual3d: Efficient and consistent text-to-3d generation with dual-mode multi-view latent diffusion, 2024. 3, 5
2024
-
[43]
Luciddreamer: Towards high- fidelity text-to-3d generation via interval score matching
Yixun Liang, Xin Yang, Jiantao Lin, Haodong Li, Xiaogang Xu, and Yingcong Chen. Luciddreamer: Towards high- fidelity text-to-3d generation via interval score matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 6517–6526,
-
[44]
Magic3d: High-resolution text-to-3d content creation
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), 2023. 2, 3
2023
-
[45]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In NeurIPS, 2023. 2, 3, 4 10
2023
-
[46]
Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024. 2, 3, 4
2024
-
[47]
Openshape: Scaling up 3d shape representation towards open-world understanding
Minghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu, Xu- anlin Li, Shizhong Han, Hong Cai, Fatih Porikli, and Hao Su. Openshape: Scaling up 3d shape representation towards open-world understanding. Advances in neural information processing systems, 36, 2024. 3
2024
-
[48]
One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion
Minghua Liu, Chao Xu, Haian Jin, Linghao Chen, Mukund Varma T, Zexiang Xu, and Hao Su. One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimiza- tion. Advances in Neural Information Processing Systems , 36, 2024. 2
2024
-
[49]
Zero-1-to-3: Zero-shot one image to 3d object, 2023
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3d object, 2023. 2
2023
-
[50]
Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age
Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age. arXiv preprint arXiv:2309.03453, 2023. 2
2023 arXiv
-
[51]
Att3d: Amortized text-to-3d object synthesis
Jonathan Lorraine, Kevin Xie, Xiaohui Zeng, Chen-Hsuan Lin, Towaki Takikawa, Nicholas Sharp, Tsung-Yi Lin, Ming- Yu Liu, Sanja Fidler, and James Lucas. Att3d: Amortized text-to-3d object synthesis. The International Conference on Computer Vision (ICCV), 2023. 3
2023
-
[52]
Neural shape compiler: A unified framework for transforming between text, point cloud, and program
Tiange Luo, Honglak Lee, and Justin Johnson. Neural shape compiler: A unified framework for transforming between text, point cloud, and program. 2022. 3
2022
-
[53]
Scalable 3d captioning with pretrained models
Tiange Luo, Chris Rockwell, Honglak Lee, and Justin John- son. Scalable 3d captioning with pretrained models. In Advances in Neural Information Processing Systems , pages 75307–75337. Curran Associates, Inc., 2023. 2, 3, 4, 5, 6, 7, 14, 17, 18, 19, 20, 21
2023
-
[54]
View selec- tion for 3d captioning via diffusion ranking
Tiange Luo, Justin Johnson, and Honglak Lee. View selec- tion for 3d captioning via diffusion ranking. arXiv preprint arXiv:2404.07984, 2024. 2, 3
2024
-
[55]
X-dreamer: Creating high-quality 3d content by bridging the domain gap between text-to-2d and text-to-3d generation, 2024
Yiwei Ma, Yijun Fan, Jiayi Ji, Haowei Wang, Xiaoshuai Sun, Guannan Jiang, Annan Shu, and Rongrong Ji. X-dreamer: Creating high-quality 3d content by bridging the domain gap between text-to-2d and text-to-3d generation, 2024. 5
2024
-
[56]
McCarthy and Scott Jarvis
Philip M. McCarthy and Scott Jarvis. Mtld, vocd-d, and hd- d: A validation study of sophisticated approaches to lexical diversity assessment. Behavior Research Methods, 42:381– 392, 2010. 6, 16
2010
-
[57]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In Computer Vision – ECCV 2020, 2020. 3
2020
-
[58]
Point-e: A system for generat- ing 3d point clouds from complex prompts
Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generat- ing 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 2, 3, 5
2022 arXiv
-
[59]
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion mod- els, 2022a eprint=2112.10741, archivePrefix=arXiv, prima- ryClass=...
-
[60]
Gpt-4 technical report
Josh OpenAI, Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shya- mal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 2, 3, 4, 6, 7, 8, 14, 15, 16
2023 arXiv
-
[61]
Seeing past words: Testing the cross-modal capa- bilities of pretrained v&l models on counting tasks, 2021
Letitia Parcalabescu, Albert Gatt, Anette Frank, and Iacer Calixto. Seeing past words: Testing the cross-modal capa- bilities of pretrained v&l models on counting tasks, 2021. 8
2021
-
[62]
Barron, and Ben Milden- hall
Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv,
-
[63]
Highly accurate dichotomous im- age segmentation
Xuebin Qin, Hang Dai, Xiaobin Hu, Deng-Ping Fan, Ling Shao, and Luc Van Gool. Highly accurate dichotomous im- age segmentation. In ECCV, 2022. 5
2022
-
[64]
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021. 3
2021
-
[65]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International confer- ence on machine learning, pages 8821–8831. Pmlr, 2021. 2
2021
-
[66]
Sentence-bert: Sentence embeddings using siamese bert-networks
N Reimers. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084 ,
1908 arXiv
-
[67]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 5
2022
-
[68]
Omniview-tuning: Boosting viewpoint invariance of vision-language pre-training models,
Shouwei Ruan, Yinpeng Dong, Hanqing Liu, Yao Huang, Hang Su, and Xingxing Wei. Omniview-tuning: Boosting viewpoint invariance of vision-language pre-training models,
-
[69]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information...
2022
-
[70]
MVDream: Multi-view diffusion for 3d gen- eration
Yichun Shi, Peng Wang, Jianglong Ye, Long Mai, Kejie Li, and Xiao Yang. MVDream: Multi-view diffusion for 3d gen- eration. In The Twelfth International Conference on Learn- ing Representations, 2024. 2, 3
2024
-
[71]
Meta 3d assetgen: Text-to-mesh genera- tion with high-quality geometry, texture, and pbr materials
Yawar Siddiqui, Tom Monnier, Filippos Kokkinos, Mahen- dra Kariya, Yanir Kleiman, Emilien Garreau, Oran Gafni, Natalia Neverova, Andrea Vedaldi, Roman Shapovalov, and David Novotny. Meta 3d assetgen: Text-to-mesh genera- tion with high-quality geometry, texture, and pbr materi...
2024
-
[72]
Llm pruning and distillation in practice: The minitron ap- proach, 2024
Sharath Turuvekere Sreenivas, Saurav Muralidharan, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, and Pavlo Molchanov. Llm pruning and distillation in practice: The minitron ap- proach, 2024. 4, 5 11
2024
-
[73]
Stefan Stojanov, Anh Thai, and James M. Rehg. Using shape to categorize: Low-shot learning with an explicit shape bias
-
[74]
Pix3d: Dataset and methods for single-image 3d shape modeling
Xingyuan Sun, Jiajun Wu, Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Tianfan Xue, Joshua B Tenenbaum, and William T Freeman. Pix3d: Dataset and methods for single-image 3d shape modeling. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 2, 3, 5, 1...
2018
-
[75]
Let me speak freely? a study on the impact of format restrictions on performance of large language models
Zhi Rui Tam, Cheng-Kuang Wu, Yi-Lin Tsai, Chieh-Yen Lin, Hung-yi Lee, and Yun-Nung Chen. Let me speak freely? a study on the impact of format restrictions on performance of large language models. arXiv preprint arXiv:2408.02442,
-
[76]
Textmesh: Gen- eration of realistic 3d meshes from text prompts
Christina Tsalicoglou, Fabian Manhardt, Alessio Tonioni, Michael Niemeyer, and Federico Tombari. Textmesh: Gen- eration of realistic 3d meshes from text prompts. In Interna- tional conference on 3D vision (3DV), 2024. 2, 3
2024
-
[77]
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distilla- tion
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distilla- tion. In Advances in Neural Information Processing Systems (NeurIPS), 2023. 2, 3
2023
-
[78]
A prompt pattern catalog to enhance prompt engineering with chatgpt
Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Car- los Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer- Smith, and Douglas C Schmidt. A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382, 2023. 5
2023 arXiv
-
[79]
Ritual: Random image transformations as a universal anti-hallucination lever in lvlms, 2024
Sangmin Woo, Jaehyuk Jang, Donguk Kim, Yubin Choi, and Changick Kim. Ritual: Random image transformations as a universal anti-hallucination lever in lvlms, 2024. 4
2024
-
[80]
Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation
Tong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang, Liang Pan Jiawei Ren, Wayne Wu, Lei Yang, Jiaqi Wang, Chen Qian, Dahua Lin, and Ziwei Liu. Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation. In IEEE/CVF Conference on Computer V...
2023
-
[81]
Gpt- 4v(ision) is a human-aligned evaluator for text-to-3d genera- tion, 2024
Tong Wu, Guandao Yang, Zhibing Li, Kai Zhang, Ziwei Liu, Leonidas Guibas, Dahua Lin, and Gordon Wetzstein. Gpt- 4v(ision) is a human-aligned evaluator for text-to-3d genera- tion, 2024. 6
2024
-
[82]
Latte3d: Large-scale amortized text-to-enhanced3d synthe- sis
Kevin Xie, Jonathan Lorraine, Tianshi Cao, Jun Gao, James Lucas, Antonio Torralba, Sanja Fidler, and Xiaohui Zeng. Latte3d: Large-scale amortized text-to-enhanced3d synthe- sis. The 18th European Conference on Computer Vision (ECCV), 2024. 3
2024
-
[83]
Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models
Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models. arXiv preprint arXiv:2404.07191 ,
-
[84]
Perspective transformer nets: Learning single- view 3d object reconstruction without 3d supervision
Xinchen Yan, Jimei Yang, Ersin Yumer, Yijie Guo, and Honglak Lee. Perspective transformer nets: Learning single- view 3d object reconstruction without 3d supervision. Ad- vances in neural information processing systems , 29, 2016. 2
2016
-
[85]
Qwen2 technical report, 2024
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jian- wei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jin...
2024
-
[86]
Consistnet: Enforcing 3d consistency for multi- view images diffusion
Jiayu Yang, Ziang Cheng, Yunfei Duan, Pan Ji, and Hong- dong Li. Consistnet: Enforcing 3d consistency for multi- view images diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 7079–7088, 2024. 2
2024
-
[87]
Dreamcomposer: Controllable 3d object generation via multi-view conditions
Yunhan Yang, Yukun Huang, Xiaoyang Wu, Yuan-Chen Guo, Song-Hai Zhang, Hengshuang Zhao, Tong He, and Xi- hui Liu. Dreamcomposer: Controllable 3d object generation via multi-view conditions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , p...
2024
-
[88]
Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models, 2024
Taoran Yi, Jiemin Fang, Junjie Wang, Guanjun Wu, Lingxi Xie, Xiaopeng Zhang, Wenyu Liu, Qi Tian, and Xinggang Wang. Gaussiandreamer: Fast generation from text to 3d gaussians by bridging 2d and 3d diffusion models, 2024. 3
2024
-
[89]
Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets
Longwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu, Anqi Pang, Haoran Jiang, Wei Yang, Lan Xu, and Jingyi Yu. Clay: A controllable large-scale generative model for creat- ing high-quality 3d assets. ACM Transactions on Graphics (TOG), 43(4):1–20, 2024. 2, 3, 5
2024
-
[90]
Gala3d: Towards text-to-3d complex scene generation via layout-guided generative gaussian splatting
Xiaoyu Zhou, Xingjian Ran, Yajiao Xiong, Jinlin He, Zhi- wei Lin, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang. Gala3d: Towards text-to-3d complex scene generation via layout-guided generative gaussian splatting. arXiv preprint arXiv:2402.07207, 2024. 2
2024 arXiv
-
[91]
Hifa: High-fidelity text-to-3d generation with advanced diffusion guidance, 2023
Junzhe Zhu and Peiye Zhuang. Hifa: High-fidelity text-to-3d generation with advanced diffusion guidance, 2023. 2, 3, 6, 7, 16, 30, 31
2023
-
[92]
Gtr: Improving large 3d reconstruction models through geometry and texture refinement
Peiye Zhuang, Songfang Han, Chaoyang Wang, Aliak- sandr Siarohin, Jiaxu Zou, Michael Vasilkovsky, Vladislav Shakhrai, Sergey Korolev, Sergey Tulyakov, and Hsin- Ying Lee. Gtr: Improving large 3d reconstruction models through geometry and texture refinement. arXiv preprint arXi...
2024 arXiv
-
[95]
Dataset Preparation Objaverse: Objaverse† [18] contains 798, 759 3D assets, with metadata (e.g., name, tags, description) available for ∼93% samples after filtering
Additional Details on Captioning Process 8.1. Dataset Preparation Objaverse: Objaverse† [18] contains 798, 759 3D assets, with metadata (e.g., name, tags, description) available for ∼93% samples after filtering. From ObjaverseXL [17], we rendered 8, 031, 637 assets, of which ∼...
-
[96]
La Cava Window
Additional details on MARVEL annotations 9.1. More Results on Effects of Human Metadata Figure 7 showcases examples where human-provided meta- data from source datasets reduce VLM hallucination and enhances annotations with domain-specific information. To generate captions usi...
-
[97]
More Implementation Details As discussed in the main paper, MARVEL-FX3D is a two- stage pipeline
Additional results of MARVEL-FX3D 10.1. More Implementation Details As discussed in the main paper, MARVEL-FX3D is a two- stage pipeline. In the first stage, Stable Diffusion 3.5 [3, 21] is fine-tuned. During each epoch, one annotation is sampled from five levels and paired wi...
-
[98]
UNITED STATES OF AMERICA
Discussion on Application of MARVEL The MARVEL-40M+ dataset, with its scale and diversity, serves as a powerful resource for text-to-3D tasks such as reconstruction, multi-view consistency, and compositional scene generation. A notable real-world use case, illustrated in Figur...
1948
-
[2022]
2, 3, 6, 7, 16, 30, 31
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.