{"id":"afde3114-742b-4415-8b6f-15e722c0e254","arxiv_id":"2411.17945","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MARVEL-40M+ provides multi-level captions for over 8.9 million 3D assets and a two-stage text-to-3D pipeline that generates textured meshes in 15 seconds.","lead":"This paper introduces a large dataset of 40 million text captions for 8.9 million 3D objects, generated by a multi-stage pipeline of vision-language and large-language models, plus a fast text-to-3D generation model. A generalist reader might care because high-quality captioned 3D data is a bottleneck for generating 3D content from text.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported caption-quality win rates are confounded by length: MARVEL Level-4 captions (44 words) are compared against Cap3D (16), 3D-Topia (29), and Kabra (5) without normalization, so the 72.41%/73.40% win rates may partly measure verbosity rather than fidelity.","rationale":"The paper's central claim is that MARVEL-40M+ is the largest and best-annotated 3D caption dataset, with win rates of 72.41% (GPT-4) and 73.40% (human). For that claim to hold, the evaluation must show that MARVEL captions are better, not merely more verbose. The reader's weakest assumption focuses on whether four fixed views at elevation 30 supply enough visual information for InternVL2-40B to caption 8.9M diverse assets. I agree that this is an important unvalidated assumption, and the paper's Section 5 admits related failures (thin objects misidentified, numerical/directional errors). However, the more directly load-bearing weakness is the caption-length confound in the headline quality numbers. Table 2 compares 44-word Level-4 captions against 16-word, 29-word, and 5-word baselines while claiming 'fair comparison due to their similar average length'; that claim is quantitatively false. The same issue carries into Table 3, where 170-word Level-1 captions are scored on all-attributes-correct, penalizing long captions more but also giving them many more chances to match the image. Because both GPT-4 and human judges see the images and the captions but not a length-controlled setting, the reported preference rates cannot be attributed to annotation quality alone. This does not invalidate the dataset's potential utility, and the qualitative examples and downstream fine-tuning gains are real evidence. But the headline quantitative superiority is not yet established. The concrete test — length-matched pairwise evaluation — is inexpensive, requires no model retraining, and would settle whether the win rates survive controlling for verbosity. I therefore keep the reader's CONDITIONAL verdict: accept only after release of artifacts and this validation, since the central claim is credible but not fully proven.","tokens_in":44675,"tokens_out":2719,"duration_ms":29764,"concrete_test":"Take a stratified random sample of 200 Objaverse assets used in the image-text alignment evaluation. For each asset, truncate or compress the MARVEL Level-4 caption to match the word count of each baseline (16, 29, and 5 words), using a deterministic extractive method or a fixed first-N-words rule, and keep baselines at their original length. Rerun the GPT-4 pairwise preference and an independent human rating on these length-matched pairs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quality claim, from the abstract, is that MARVEL-40M+ 'significantly outperforms existing datasets in annotation quality,' supported by win rates of 72.41% (GPT-4) and 73.40% (human) in Table 2. The evaluation asks judges to view four multi-view renders and pick the caption that best matches the 3D model. Under these conditions, longer captions have a systematic advantage: they can mention more object properties, colors, parts, and context, so even a partially correct longer caption will contain more verifiable phrases than a short correct one. MARVEL Level-4 captions average 44 words, versus 16 for Cap3D, 29 for 3D-Topia, and 5 for Kabra. The text states Level 4 was chosen because its length is 'similar' to baselines, but 44 is 2.75x Cap3D and 8.8x Kabra — not similar. The same confound appears in the caption-accuracy comparison (Table 3): MARVEL Level-1 captions (170 words) are judged on whether all attributes are correct, while Kabra (5 words) has far fewer attributes to be wrong. The claim that MARVEL is superior in annotation quality for its intended use is therefore load-bearing on an evaluation design that does not isolate quality from length or detail. This is a correctness risk for the headline numbers, not merely a presentation issue, because the same numbers are used to justify the dataset's value and the downstream fine-tuning comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MARVEL-40M+, a large-scale multi-level caption dataset covering roughly 8.9 million 3D assets and 44.5 million captions aggregated from seven 3D datasets. Captions are produced by a five-stage pipeline: four-view rendering, human-metadata filtering, dense description generation with InternVL2-40B, multi-level elaboration with Qwen2.5-72B, and ethical filtering. The paper also introduces MARVEL-FX3D, a two-stage text-to-3D pipeline that fine-tunes Stable Diffusion 3.5 on the new captions and uses Stable Fast 3D for mesh generation. The central claims are that MARVEL-40M+ outperforms existing 3D caption datasets in annotation quality and linguistic diversity, with reported GPT-4 and human win rates of 72.41% and 73.40%, and that MARVEL-FX3D improves prompt fidelity and overall preference over existing text-to-3D methods.","tokens_in":44992,"tokens_out":5148,"duration_ms":49331,"significance":"If validated, this is a potentially valuable resource: it is substantially larger than existing 3D caption datasets, uses only open-source models, provides a cost analysis of the annotation pipeline, and offers a five-level annotation structure with qualitative examples across diverse domains. The inclusion of failure cases and limitations is a strength. However, the headline quality claims currently rest on evaluation protocols that do not control for caption length or report statistical reliability, so the magnitude of the claimed advantage is not yet established.","major_comments":[{"comment":"The statement that Level-4 annotations were chosen because their average length is 'similar' to baseline datasets is not supported by the table: MARVEL Level 4 averages 44 words, versus 16 for Cap3D, 29 for 3D-Topia, and 5 for Kabra. The win-rate task asks GPT-4 and human judges to select the caption that best matches the rendered 3D images; longer captions can mention more objects, attributes, colors, and contextual details, so even a partially correct longer caption has more opportunities to appear correct. The headline win rates of 72.41% and 73.40% are therefore confounded by length and may partly measure verbosity rather than fidelity. The paper should provide a length-controlled or length-matched evaluation, such as truncated captions, per-attribute precision/recall, or normalized informativeness, together with confidence intervals.","section":"Section 4.1, Table 2"},{"comment":"The caption-accuracy comparison uses MARVEL Level 1 captions averaging 170 words against baselines averaging 5-29 words. Asking judges whether 'all' mentioned attributes are correct is a reasonable way to penalize verbosity, but the paper does not report how judges handle attributes that are plausible but not visually verifiable from four renders, nor does it report inter-annotator agreement or confidence intervals for the 250-sample human evaluation. Without these details, the claimed 84.70% GPT-4 and 82.80% human accuracy rates are not statistically grounded, and the comparison across very different caption lengths remains difficult to interpret.","section":"Section 4.1, Table 3"},{"comment":"The claim that MARVEL-FX3D outperforms state-of-the-art text-to-3D methods rests on 50 prompts scored by five users. Several reported differences are within one standard deviation (for example, visual quality 6.58 vs 6.47 and geometric consistency 7.20 vs 7.25), and no significance tests or inter-rater agreement measures are reported. In addition, the baselines are adapted with different training budgets and implementations (LucidDreamer for 3k steps, DreamFusion and HiFA via threestudio for 10k and 24k steps), so it is unclear whether the comparison is apples-to-apples. The paper should report paired significance tests, confidence intervals, and standardized baseline configurations.","section":"Section 4.2, Table 4"},{"comment":"The human-metadata contribution is a stated core contribution, including the abstract's claim that metadata 'reduce VLM hallucinations,' but the ablation is qualitative only. Figure 5 and the supplementary examples show selected successes, but no quantitative measurement such as hallucination rates, error counts, or a randomized ablation over a representative sample is provided. The claim that metadata injection reduces hallucinations is therefore not yet established at dataset scale.","section":"Section 4.3A"},{"comment":"The four-view rendering protocol (azimuths 0, 90, 180, 270 degrees at elevation 30) is assumed to provide sufficient visual information for InternVL2-40B to caption all 8.9 million assets. Section 5 acknowledges failure modes such as misidentification of thin objects, side-view confusion, numerical imprecision, and directional errors, but the paper never quantifies how frequently these failures occur or how they are distributed across the seven source datasets. Since the dataset-wide quality claim depends on caption correctness across a highly heterogeneous corpus, a failure-rate estimate or category-level breakdown is needed before the 'high-quality annotation' claim can be assessed.","section":"Section 3.1 and Section 5"}],"minor_comments":[{"comment":"The text states that the trend extends to 'average word length,' but Table 2 does not contain an average word length column; the table reports average caption length, MTLD, unigram and bigram counts, and GPT-4/human win rates.","section":"Section 4.1"},{"comment":"The MTLD comparison sentence contains citation mismatches: 'higher than Cap3D [28]' and 'higher than 3D-Topia[53]' appear to swap the reference numbers, since [28] is 3D-Topia and [53] is Cap3D in the bibliography.","section":"Section 4.1"},{"comment":"The supplementary text refers to 'Table 5' when discussing the inter-level semantic retention ablation, but the corresponding result appears as Table 6 in the main paper; this cross-reference should be corrected.","section":"Supplementary Section 9.5"},{"comment":"The table caption contains a typo: 'desite' should be 'despite.'","section":"Table 3"},{"comment":"For a dataset paper, the manuscript should state clearly where and how the dataset and the annotation pipeline code will be released, including licenses and any restrictions inherited from the source datasets; the current project page URL alone is not sufficient.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The dataset and pipeline are likely to be useful to the text-to-3D community, and the open-source, cost-aware design is a genuine strength. However, the headline annotation-quality and text-to-3D superiority claims are not yet supported because the main evaluations are confounded by caption length and lack statistical reliability. I would encourage the editor to request a revised version that includes length-controlled evaluations, inter-annotator agreement, significance tests, and at least a coarse quantification of known failure modes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real contribution. The scale alone matters—40M+ captions across 8.9M assets with a five-level hierarchy and filtered human metadata injection is new, and the open-source pipeline (InternVL2 + Qwen2.5) is a sensible, cost-conscious way to get there. The cost/throughput numbers in the supplement are transparent and useful. The paper also does the honest thing in Section 5: it names VLM counting failures, thin-object misidentification, directional errors, and flat outputs from SF3D. That gives me confidence the authors know where the weak points are.\n\nThe soft spot is the evaluation that carries the abstract's central claim. The image-text alignment study compares MARVEL Level-4 captions (44 words) against Cap3D (16), 3D-Topia (29), and Kabra (5), and the paper calls these lengths \"similar.\" They are not. Judges who see four renders and pick the best-matching caption will systematically prefer longer captions because longer captions contain more verifiable claims. The caption-accuracy comparison has the same problem at a larger scale: MARVEL Level-1 captions average 170 words, versus 5 for Kabra, and the task is whether all mentioned attributes are correct. A 5-word caption has very few attributes to be wrong about. The reported 72.41% and 73.40% win rates are therefore partly measuring verbosity, not fidelity. This is a load-bearing flaw for the paper's central claim, not a cosmetic one.\n\nThe human evaluation is also thin: five reviewers, 200 samples each for alignment, 250 samples total for accuracy, no inter-annotator agreement, no confidence intervals. For a dataset paper that is trying to displace CAP3D and Kabra, that is not enough. And the dataset and code are not released, despite the project page; for a resource paper, that is a central omission.\n\nThe FX3D downstream comparison is plausible but not convincing: 50 test prompts, five users, no statistical tests, and most of the heavy lifting is done by pretrained SF3D. The fine-tuning comparison against Cap3D is more informative, though it inherits the same evaluation-scale problems.\n\nWho is this for? People building text-to-3D systems who want a large, hierarchically annotated training corpus, and people working on 3D captioning evaluation. The resource idea is credible and the engineering is sound, but the empirical case needs a serious rework: length-controlled or per-attribute evaluation, statistical tests, and released artifacts.\n\nRecommendation: send it to peer review—it deserves referee time—but push hard for conditional acceptance. The length confound must be addressed and the artifacts must be made available before the quality claims can be trusted.","headline":"A genuinely large, well-engineered captioning resource whose headline quality numbers are inflated by a length confound, and whose artifacts are not yet released.","tokens_in":45592,"tokens_out":1833,"would_cite":false,"duration_ms":21119,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A dataset of 40 million captions for 8.9 million 3D assets, with five description levels, is claimed to be the largest 3D caption dataset and to beat prior captions by 72–73% in preference.","keywords":["text-to-3D generation","3D captioning dataset","multi-view vision-language models","multi-level annotations","human metadata filtering","Stable Diffusion fine-tuning","image-to-3D reconstruction"],"falsifier":"Take 1,000 randomly sampled objects from Objaverse-XL, generate MARVEL-style captions from the four fixed views, and also from a denser orbit rendering (e.g., 12 views or a turntable sweep); if human experts or GPT-4 judge the denser-view captions clearly more accurate, or if objects known to be thin or heavily occluded fail systematically under the four-view protocol, the assumption that four views suffice would be refuted.","tokens_in":44472,"feed_emoji":"🧊","tokens_out":6620,"duration_ms":53489,"temperature":0.7,"pith_summary":"The paper argues that text-to-3D generation is held back less by generation methods than by the scarcity of rich, accurate text captions for 3D assets. To fix this, it introduces MARVEL-40M+, a dataset of over 40 million captions covering 8.9 million 3D objects gathered from seven existing datasets. The captions are produced by an automated pipeline that renders four views of each object, asks a multi-view vision-language model for a dense description, and then compresses that description into five levels, from 150–200-word detailed descriptions down to 10–20-word tags. Filtered human metadata from the source datasets is injected to add domain-specific names and reduce hallucination. The paper reports that its captions beat prior datasets in 72.41% of GPT-4 comparisons and 73.40% of human comparisons, and that a two-stage text-to-3D model trained on the dataset generates textured meshes in about 15 seconds.","feed_headline":"New 40M-caption dataset beats prior 3D captions 73% of the time","feed_subtitle":"Automated open-source pipeline writes detailed-to-tag descriptions and turns text into textured meshes in 15 seconds.","key_machinery":"The load-bearing mechanism is the five-stage MARVEL annotation pipeline. It renders each asset from four fixed viewpoints (azimuth 0°, 90°, 180°, 270° at 30° elevation), filters noisy user metadata with Mistral-Nemo, asks InternVL2-40B to write a dense description covering components, geometry, materials, colors, and context, then uses Qwen2.5-72B to compress that description into five hierarchical levels, ending with Qwen2.5-14B ethical filtering. This chain converts a 3D asset into 44.5 million captions. The downstream text-to-3D system is a two-stage chain: LoRA fine-tuned Stable Diffusion 3.5 produces reconstruction-friendly images from these captions, and Stable Fast 3D lifts the image to a textured mesh.","core_discovery":"The paper's central claim is that high-quality 3D captions can be produced automatically at a scale that was previously impractical, and that this scale is what unblocks high-fidelity text-to-3D. It reports that MARVEL-40M+ contains 44,510,515 captions for 8,902,103 objects aggregated from seven datasets, making it the largest 3D caption dataset to date. Evaluated on 5,000 samples by GPT-4 and 1,000 samples by five human reviewers, its Level-4 captions are preferred over CAP3D, 3D-Topia, and Kabra captions 72.41% and 73.40% of the time respectively, and its Level-1 captions reach 84.70% (GPT-4) and 82.80% (human) caption accuracy. On the generation side, MARVEL-FX3D, which fine-tunes Stable Diffusion 3.5 on the dataset and passes the resulting image through Stable Fast 3D, reports the highest prompt fidelity (7.71/10) and overall preference (6.94/10) among compared methods while producing textured meshes in about 15 seconds.","pith_inferences":["An implication the paper leaves implicit is that the metadata-filtering step could be generalized to datasets without user metadata by using classifier or taxonomy labels, in which case the value of metadata injection should be reported as a per-domain statistic rather than only qualitatively.","The five-level hierarchy is likely to be useful beyond generation, for example as a controlled way to probe how much caption detail a retrieval or editing system needs; the paper does not test those downstream tasks.","The reported win rates are relative preferences, so a natural next benchmark is to measure absolute caption usefulness, such as improvement in reconstruction accuracy or text-to-3D retrieval recall, on a fixed test set."],"forward_implications":["A single dataset can now supply training captions for 8.9M objects from seven sources, so text-to-3D models no longer need to reconcile mismatched short captions from different dataset formats.","Because each asset has five description levels, downstream systems can select the level matched to their cost and fidelity needs, from 150–200-word reconstructions to 10–20-word prototyping tags.","The annotation pipeline is built from open-source models and costs roughly $2,700–$3,000 to annotate the 800K-sample Objaverse subset, making dataset expansion far cheaper than human annotation.","Fine-tuning Stable Diffusion 3.5 on MARVEL captions improves prompt fidelity and overall preference over training on CAP3D captions or no fine-tuning, and does so while keeping generation time near 15 seconds."],"supporting_citations":[{"why":"InternVL2 multi-view VLM generates the dense descriptions that are the pipeline's first content stage.","marker":"[13, 15]"},{"why":"Qwen2.5 performs the multi-level elaboration that produces the five annotation levels.","marker":"[85]"},{"why":"Mistral-Nemo filters noisy user metadata before it is fed to the captioning model.","marker":"[72]"},{"why":"Objaverse and Objaverse-XL supply the bulk of the 8.9M assets and the human metadata used to ground captions.","marker":"[17, 18]"},{"why":"CAP3D is the main baseline whose annotation quality and cost are compared against MARVEL.","marker":"[53]"},{"why":"Stable Fast 3D is the pretrained image-to-3D module that turns generated images into textured meshes in MARVEL-FX3D.","marker":"[7]"},{"why":"Stable Diffusion 3.5 is the text-to-image model fine-tuned on MARVEL captions in the first stage of MARVEL-FX3D.","marker":"[21]"},{"why":"Provides the MTLD lexical-diversity metric used to claim MARVEL captions are more diverse than baselines.","marker":"[56]"},{"why":"GPT-4 serves as the automated evaluator whose 72.41% win rate supports the dataset-quality claim.","marker":"[60]"}],"fun_headline_variants":["Largest 3D caption dataset beats prior captions 73%","40M captions auto-written for 8.9M 3D assets","Multi-level captions: 40M for 8.9M 3D assets","Text-to-3D in 15s with 40M-caption dataset","Auto-captions win 73% human preference in 3D"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that four fixed rendered views (front, back, left, right) give the vision-language model enough information to describe all 8.9M objects correctly, including thin, occluded, or viewpoint-ambiguous ones.","fun_headline_variants_meta":{"raw":{"variants":["Largest 3D caption dataset beats prior captions 73%","40M captions auto-written for 8.9M 3D assets","Multi-level captions: 40M for 8.9M 3D assets","Text-to-3D in 15s with 40M-caption dataset","Auto-captions win 73% human preference in 3D"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001646,"raw_usage":{"total_tokens":6600,"prompt_tokens":1067,"completion_tokens":5533,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":5429}},"tokens_in":683,"tokens_out":5533,"duration_ms":32142,"temperature":1.0,"reasoning_tokens":5429,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:40:17.697208+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take 1,000 randomly sampled objects from Objaverse-XL, generate MARVEL-style captions from the four fixed views, and also from a denser orbit rendering (e.g., 12 views or a turntable sweep); if human experts or GPT-4 judge the denser-view captions clearly more accurate, or if objects known to be thin or heavily occluded fail systematically under the four-view protocol, the assumption that four views suffice would be refuted.","supporting_citations":[{"cited_title":"Qwen2 technical report, 2024","cited_arxiv_id":null,"evidence_quote":"Qwen2.5 performs the multi-level elaboration that produces the five annotation levels."},{"cited_title":"Llm pruning and distillation in practice: The minitron ap- proach, 2024","cited_arxiv_id":null,"evidence_quote":"Mistral-Nemo filters noisy user metadata before it is fed to the captioning model."},{"cited_title":"Scalable 3d captioning with pretrained models","cited_arxiv_id":null,"evidence_quote":"CAP3D is the main baseline whose annotation quality and cost are compared against MARVEL."},{"cited_title":"McCarthy and Scott Jarvis","cited_arxiv_id":null,"evidence_quote":"Provides the MTLD lexical-diversity metric used to claim MARVEL captions are more diverse than baselines."}],"review_version":1}