{"id":"7ac7dab7-7ab7-4386-9497-916298e188dc","arxiv_id":"2508.20470","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A video diffusion backbone fine-tuned on 4M densely captioned 360-degree renderings generates spatially consistent multi-view images for 3D assets from image plus detailed text input.","lead":"The authors built a 4-million-object dataset of rotating 3D renders with detailed per-angle captions, then fine-tuned a video-generation model to produce multi-view images from an input image and a long text prompt. The result is a 3D asset generator that also shows signs of working on whole scenes, not just single objects.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Evaluation tables are internally inconsistent: Droplet3D's SSIM is lower than both baselines and its own backbone although the text claims SSIM confirms higher quality; without correction/error bars the central superiority and video-prior claims are unsupported.","rationale":"The paper's most valuable contribution—the 4M rendered multi-view video dataset with dense captions—is plausible and independently useful, and the model's qualitative demonstrations are suggestive. However, the central claim of quantitative superiority rests on Tables 2–4, and those tables contain a hard internal contradiction: the model's SSIM is lower than both baselines and lower than its own backbone after fine-tuning, despite prose claiming SSIM confirms higher quality. This is not merely a missing error bar; the numbers as reported undermine the conclusion they are cited to support. A reader cannot tell whether the metrics were computed under comparable conditions. The scene-level claim is qualitative and does not repair this. Since this issue is addressable by rerunning with a standardized protocol and significance testing, the appropriate verdict remains CONDITIONAL (no change from the reader's verdict).","tokens_in":29308,"tokens_out":11368,"duration_ms":133278,"concrete_test":"Using the released weights, re-run the Table 2/3 evaluation on the same 200 GSO samples with 3 or more seeds and a single fixed protocol: render all methods to the same canonical 85-view orbit/resolution and use one third-party script to compute per-sample PSNR/SSIM/LPIPS/MSE/CLIP-S with 95% confidence intervals and paired tests. If, under identical alignment, Droplet3D's SSIM remains below LGM/MVControl while PSNR/LPIPS improve, the paper's 'SSIM confirms higher quality' statement is false and the superiority claim is unproven; if SSIM flips and differences are significant, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The quantitative support for the central claim is Tables 2–4 (§5.2.1). They are not self-consistent. In Table 2, Droplet3D has SSIM 0.76 vs 0.84 (LGM) and 0.88 (MVControl), yet §5.2.1 states 'PSNR and SSIM confirm that the content generated by Droplet3D is of higher quality and closer to the ground truth.' In Table 3, fine-tuning from DropletVideo improves PSNR (20.51→28.36), LPIPS (0.12→0.03), MSE, and CLIP-S but lowers SSIM (0.87→0.76) while the prose claims 'improved generation consistency.' No error bars, significance tests, or description of view alignment/protocol for each method are provided. If metrics were computed under different camera trajectories, reference views, or output formats, the claimed superiority over baselines and the attribution to video priors lose support. The qualitative scene-level results (§5.3.5) are not quantified, so they cannot substitute for the broken quantitative case.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a video-driven paradigm for 3D generation. It introduces Droplet3D-4M, a dataset of about 4M objects from Objaverse-XL, each rendered as an 85-frame 360-degree orbital video with dense multi-view-level captions averaging 260 words. It then fine-tunes the authors' DropletVideo backbone to produce Droplet3D, a model that takes an image and a dense text prompt as input and generates 85 surrounding views, which are lifted to textured meshes or 3D Gaussian splatting. The central claim is that commonsense priors from video models improve spatial consistency and semantic fidelity in 3D generation, and that this enables even scene-level generation despite no scene-level training data. The dataset, code, and weights are open-sourced.","tokens_in":29636,"tokens_out":3763,"duration_ms":44688,"significance":"If the quantitative claims were supported, this would be a significant contribution: a 4M-scale rendered-video dataset with detailed multi-view captions, a video-to-3D fine-tuning recipe, and a demonstration that video backbones can transfer spatial and semantic priors to 3D generation are all valuable to the community. The authors should be credited for releasing resources and for the substantial dataset-construction effort, including the GRPO-based captioning pipeline. However, the paper's core evidence is currently weakened by internal metric inconsistencies and by the lack of statistical rigor in the comparisons. The claimed superiority over baselines and the attribution of that superiority to video priors are not yet established.","major_comments":[{"comment":"There is a direct internal inconsistency between the reported numbers and the prose. In Table 2, Droplet3D has SSIM 0.76, while LGM has 0.84 and MVControl has 0.88, yet the text states that “PSNR and SSIM confirm that the content generated by Droplet3D is of higher quality and closer to the ground truth.” In Table 3, fine-tuning from DropletVideo lowers SSIM from 0.87 to 0.76 while the surrounding text claims “improved generation consistency.” SSIM is presented as a quality/reconstruction metric, so a drop of this size must either be explained (e.g., a different evaluation protocol, reference views, or camera alignment) or the claims must be revised. As written, the table contradicts the central quantitative claim.","section":"§5.2.1, Table 2 and Table 3"},{"comment":"All quantitative results are single-run point estimates on a hand-selected 200-sample subset of GSO, with no error bars, no significance tests, and no description of how camera viewpoints or reference views are aligned across methods. The text states the subset was selected to cover all categories and to be “confirmed uniform,” but this does not replace repeated sampling or statistical comparison. Since PSNR/SSIM/LPIPS can vary substantially with viewpoint alignment and reconstruction protocol, the claimed superiority over LGM and MVControl is not supported without this information. Please report means with standard deviations over multiple runs or bootstrapped subsets, and state the exact evaluation protocol for each baseline.","section":"§5.2.1, Tables 2–4"},{"comment":"The central attribution to “commonsense priors from videos” is not cleanly isolated. Droplet3D differs from DropletVideo not only in fine-tuning data but also in using the 85-view rendered videos, the dense multi-view captions, a longer text token length (400 vs. 226), and a canonical-view alignment module. Table 3 compares Droplet3D-5B with DropletVideo-5B, but this conflates the effect of the video backbone with the effect of continued training on Droplet3D-4M. Table 4 evaluates zero-shot video-generation ability and does not establish that a video backbone is what makes fine-tuning successful. A control that fine-tunes a third-party video backbone on the same Droplet3D-4M data, or that trains Droplet3D from a non-video initialization on the same data, is needed to support the paper's stated claim that video priors significantly facilitate 3D creation.","section":"§4.1, §5.2.2, and abstract"}],"minor_comments":[{"comment":"Typographical issues: “funtions” should be “functions”; “multiview-match” and “Multi-view pattern” formatting is inconsistent. The reward formulas in Eqs. (1)–(4) are not clearly tied to the five claimed dimensions (Subject, material, Functional, Details, OCR), which makes the caption-quality reward difficult to reproduce.","section":"§3.3.2"},{"comment":"“anonical viewpoint” should be “canonical viewpoint”; the same typo appears in the module name. Also, the 200-example dataset for view alignment is described very briefly; please clarify how the four orthogonal ground-truth views were selected and whether the same images were used in evaluation.","section":"§4.2.2"},{"comment":"The scene-level results are qualitative only, and the claim that this capability is “entirely inherited” from DropletVideo is stronger than the evidence supports. Without quantitative consistency metrics or a controlled comparison, the text should be tempered or supplemented with measurements.","section":"§5.3.5"},{"comment":"The paper would benefit from a limitations subsection. For example, the dataset is derived entirely from Objaverse-XL and filtered by aesthetic/quality thresholds; the implications of this distribution for downstream generalization are not discussed. Also, several references appear only as arXiv identifiers and some related concurrent video-to-3D works are not discussed in detail.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The evaluation is heavily self-referential: the backbone, the dataset, and the text-alignment modules all come from the same team, and the quantitative comparison to third-party methods is currently unreliable. I would encourage the editor to request an independent or at least a more carefully controlled evaluation before considering acceptance. The underlying dataset and open-sourced resources are promising, and the authors should be able to address the metric inconsistency and the attribution issue with additional experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe worthwhile thing here is Droplet3D-4M: 4M objects with 85-view orbital renders and dense per-view captions, released along with code and weights. That is a concrete, reusable asset for anyone working on multi-view generation or 3D-language tasks. The model side—fine-tuning a video diffusion backbone on this data for image-plus-text conditioned orbit generation—is a reasonable engineering contribution, and the canonical-view alignment and text rewriting modules are sensible. The qualitative controllable-editing results (QR codes, backpacks, scene-level lifts) are suggestive and could seed real follow-up work.\n\nThe problem is the quantitative evaluation. Table 2 reports Droplet3D with SSIM 0.76 against LGM's 0.84 and MVControl's 0.88, yet the text says SSIM confirms higher quality. Table 3 shows fine-tuning on Droplet3D-4M raising PSNR, LPIPS, MSE, and CLIP-S while dropping SSIM from 0.87 to 0.76, and the prose again claims improved consistency. That is not a typo-level issue; it undercuts the central claim that the method beats baselines and that the video-prior attribution is supported. There are no error bars, no significance tests, and the 200-sample GSO subset selection is only loosely described. The scene-level generalization claim is entirely qualitative. These are addressable, but as is, the superiority claims do not hold up.\n\nI'd still send this to a serious referee. The dataset release alone warrants scrutiny and useful feedback, and the metric inconsistency is likely fixable by re-running with proper protocol and reporting variance. But any acceptance should hinge on a corrected evaluation and a toned-down attribution.","headline":"The dataset is the real contribution; the evaluation tables don't back the paper's own claims, so the superiority story needs a rewrite.","tokens_in":30155,"tokens_out":2049,"would_cite":true,"duration_ms":23287,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuned video model turns one image into a full 3D orbit.","keywords":["3D generation","video diffusion models","multi-view consistency","commonsense priors","dense multi-view captions","orbital rendering","Gaussian splatting","image-to-3D"],"falsifier":"Re-run the comparison on a random, larger sample of GSO with multiple seeds and report confidence intervals; if Droplet3D's PSNR and CLIP gains over LGM and MVControl do not reproduce outside the selected subset, the central claim fails. Additionally, fine-tune a non-video image-based multi-view diffusion model on the same Droplet3D-4M clips and check whether the gains persist without the video backbone's temporal prior.","tokens_in":29252,"feed_emoji":"🎥","tokens_out":3365,"duration_ms":38579,"temperature":0.7,"pith_summary":"This paper tries to establish that the commonsense priors stored in large-scale video data—spatial consistency across views and broad semantic knowledge—can be transferred into 3D generation. The authors build Droplet3D-4M, a dataset of 4 million 3D objects rendered as 85-frame, full-360-degree orbital videos with dense multi-view text captions averaging 260 words. They then fine-tune a pretrained video diffusion model, DropletVideo, on this dataset to create Droplet3D, which takes one image plus a long text prompt and generates a spatially consistent orbital view sequence. Those views can be lifted into textured meshes and Gaussian splats, and the model reportedly extends to scene-level generation even though its training set contains no scenes. If correct, this offers a path around the native-3D data bottleneck by borrowing supervision from abundant video.","feed_headline":"Fine-tuned video model turns one image into a full 3D orbit","feed_subtitle":"Training on 4M rendered orbital clips with dense captions yields 85 consistent views, meshes, splats, and even whole scenes.","key_machinery":"The load-bearing mechanism is the pairing of Droplet3D-4M with a video-diffusion backbone. Droplet3D-4M converts native 3D meshes into a video-native format: 85 frames rendered along a circular camera path with less than 5 degrees between adjacent views, accompanied by two-paragraph captions that first describe the object globally and then describe viewpoint-specific appearance changes. This format lets Droplet3D inherit spatial-consistency priors from DropletVideo while the dense text supervision preserves the model's semantic knowledge. The architecture adds a 3D causal VAE for spatio-temporal latent encoding and a modality-expert transformer for fusing text and video features, plus an inp","core_discovery":"The central claim is that video-derived commonsense priors significantly facilitate 3D creation. Concretely, the authors claim that fine-tuning DropletVideo—a video diffusion model with integral spatio-temporal consistency—on Droplet3D-4M yields a generator that, from an input image and dense text, produces 85 spatially consistent multi-view frames covering a full 360-degree orbit. They further claim that the dense multi-view captions preserve the backbone's semantic understanding, enabling controlled edits (e.g., swapping a character's backpack for a QR code or a crystal orb) and generalization to stylized images such as sketches and comics, as well as scene-level lifting into 3D Gaussian s","pith_inferences":["A testable extension is ablating the second, viewpoint-aware caption paragraph: if removing it degrades cross-view consistency, the paper's attribution of spatial consistency to video priors would be tangled with the effect of dense text supervision.","The dataset's aesthetic and quality filters (scores above 4.0) bias Droplet3D-4M toward clean, well-lit renderings; models trained on it may transfer less well to noisy or in-the-wild imagery, and the scene-level claim suggests this bias was not fatal but deserves direct measurement.","Comparing Droplet3D against an image-based multi-view diffusion model fine-tuned on the same rendered clips would isolate whether the video backbone's temporal prior is the causal ingredient or whether the dataset alone drives the gains.","The canonical-view alignment module was trained on only 200 manually curated examples; scaling that set is a natural next step for robust handling of arbitrary real-world viewpoints."],"forward_implications":["If the claim holds, 3D generators can be built by fine-tuning existing video models rather than collecting native 3D data at the scale of image or text datasets.","Dense multi-view captions provide a supervisory signal that lets the model keep semantic concepts absent from 3D corpora, such as QR codes and stylized accessories.","The same 85-view orbit supports both textured-mesh and Gaussian-splatting reconstruction, so a single generator can feed multiple downstream 3D representations.","Scene-level 3D generation from a single image becomes possible even when the training data contains only object-level orbital videos, implying the scene ability is inherited from video pretraining.","Arbitrary input viewpoints can be aligned to canonical views, loosening the input constraints that typically limit image-to-3D systems."],"supporting_citations":[{"why":"Supplies the DropletVideo-5B video diffusion backbone whose weights initialize Droplet3D and whose spatio-temporal consistency is claimed as the source of inherited 3D priors.","marker":"[74]"},{"why":"Objaverse-XL is the raw source of the 6.3 million 3D models from which the filtered 4 million Droplet3D-4M samples are drawn.","marker":"[7]"},{"why":"GSO is the evaluation benchmark from which 200 samples are selected for the main quantitative comparison against LGM and MVControl.","marker":"[11]"},{"why":"LGM is one of the two image-plus-text 3D baselines that Droplet3D is compared against in the TI-to-3D evaluation.","marker":"[56]"},{"why":"MVControl is the other image-plus-text 3D baseline used in the quantitative and qualitative comparison.","marker":"[33]"},{"why":"Cap3D provides the existing short-caption multi-view annotation approach whose coarser object-level text is contrasted with Droplet3D-4M's 260-word multi-view captions.","marker":"[40]"},{"why":"Hunyuan3D-2 is used downstream to convert Droplet3D's generated multi-view images into textured meshes.","marker":"[77]"},{"why":"The optimization-based 3D Gaussian splatting algorithm is used to reconstruct splats from Droplet3D's generated views, including scene-level results.","marker":"[28]"},{"why":"V3D is cited as prior work recognizing that video diffusion models can be fine-tuned as 3D generators, positioning Droplet3D's approach.","marker":"[4]"},{"why":"SV3D is cited as another video-diffusion-based multi-view generation method that Droplet3D extends with dense text supervision.","marker":"[58]"}],"fun_headline_variants":["Video priors enable one image to become a 3D orbit","Droplet3D: From a single image to 85 consistent views","Videos provide commonsense for 3D generation from one image","One image to full 3D orbit: video-derived priors","Droplet3D: Video teaches 3D model to render 360° from one shot"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The central superiority claim rests on metrics computed once on a hand-selected 200-sample subset of the GSO dataset, and one of the reported numbers (SSIM) actually favors a baseline, so if that evaluation is not representative, the claim that video priors outperform existing methods loses its support.","fun_headline_variants_meta":{"raw":{"variants":["Video priors enable one image to become a 3D orbit","Droplet3D: From a single image to 85 consistent views","Videos provide commonsense for 3D generation from one image","One image to full 3D orbit: video-derived priors","Droplet3D: Video teaches 3D model to render 360° from one shot"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000533,"raw_usage":{"total_tokens":2434,"prompt_tokens":811,"completion_tokens":1623,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":1523}},"tokens_in":555,"tokens_out":1623,"duration_ms":13586,"temperature":1.0,"reasoning_tokens":1523,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:04:39.866434+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the comparison on a random, larger sample of GSO with multiple seeds and report confidence intervals; if Droplet3D's PSNR and CLIP gains over LGM and MVControl do not reproduce outside the selected subset, the central claim fails. Additionally, fine-tune a non-video image-based multi-view diffusion model on the same Droplet3D-4M clips and check whether the gains persist without the video backbone's temporal prior.","supporting_citations":[{"cited_title":"DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation","cited_arxiv_id":"2503.06053","evidence_quote":"Supplies the DropletVideo-5B video diffusion backbone whose weights initialize Droplet3D and whose spatio-temporal consistency is claimed as the source of inherited 3D priors."},{"cited_title":"Google scanned objects: A high-quality dataset of 3d scanned household items","cited_arxiv_id":null,"evidence_quote":"GSO is the evaluation benchmark from which 200 samples are selected for the main quantitative comparison against LGM and MVControl."},{"cited_title":"Lgm: Large multi-view gaussian model for high-resolution 3d content creation","cited_arxiv_id":null,"evidence_quote":"LGM is one of the two image-plus-text 3D baselines that Droplet3D is compared against in the TI-to-3D evaluation."},{"cited_title":"Scalable 3d captioning with pretrained models","cited_arxiv_id":null,"evidence_quote":"Cap3D provides the existing short-caption multi-view annotation approach whose coarser object-level text is contrasted with Droplet3D-4M's 260-word multi-view captions."},{"cited_title":"Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion","cited_arxiv_id":null,"evidence_quote":"SV3D is cited as another video-diffusion-based multi-view generation method that Droplet3D extends with dense text supervision."}],"review_version":1}