{"id":"17787c30-02c6-4b4b-bf4a-615ae95f559c","arxiv_id":"2412.01430","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MVImgNet2.0 expands MVImgNet to 520k objects and 515 categories with higher-quality annotations, and experiments show it improves 3D reconstruction models.","lead":"This paper introduces MVImgNet2.0, a multi-view image dataset of about 300,000 real-world objects in 347 categories, expanding the authors' earlier MVImgNet to roughly 520,000 objects in 515 categories. It adds 360-degree captures and higher-quality masks, camera poses, and point clouds, and shows these improvements help train 3D reconstruction models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 4's Chamfer-Distance evidence for higher-quality point clouds is confounded: the test target is produced by the same MV2 annotation pipeline used as training supervision.","rationale":"The reader's weakest assumption concerns the overlap between the 20 held-out test categories and MVImgNet1.0 categories. I do not think this is the most damaging issue. If the held-out categories overlap MV1 categories, the MV1-only baseline has seen those categories, giving it an advantage; MV2-only still wins in Table 3 by roughly 0.5dB, which is conservative for the paper's claim. The combined-versus-MV1 comparison is also fair in terms of category familiarity because both models have seen the same test categories. Overlap would at most weaken the 'category-agnostic' phrasing, not the core value claim. The real soft spot is the non-independent point-cloud target in Table 4. The CD metric is the only quantitative evidence for the headline point-cloud-quality improvement, and the training target and evaluation target are generated by the same pipeline. This is a concrete, testable concern: independent 3D ground truth would settle it. The paper also lacks error bars and significance tests, as the reader notes, but that is secondary. Since the concern can be resolved with an independent 3D evaluation and does not invalidate the dataset's scale and category contributions, conditional acceptance remains the right verdict, with the added condition of independent point-cloud validation or softer wording of the point-cloud-quality claim.","tokens_in":22440,"tokens_out":16883,"duration_ms":153159,"concrete_test":"Select 30-50 objects from the held-out test categories. Acquire independent 3D shape ground truth by laser scanning or by an external high-fidelity reconstruction protocol (e.g., COLMAP MVS plus screened Poisson, or a structured-light scan), not by the MV2 annotation pipeline. Recompute Table 4's CD for TriplaneGaussian trained with MV1-Anno and MV2-Anno point-cloud supervision against this independent target, using identical train/test splits and seeds. If the MV2-Anno-trained model still has significantly lower CD, the point-cloud-quality claim survives; if the advantage shrinks or reverses, the Table 4 CD evidence is an artifact of train/test target alignment, and the paper should be revised to report absolute accuracy or qualify the point-cloud claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 4.1 states that test-set shape ground truths are 'high-quality dense point cloud reconstructions with manual cleaning by annotators.' These are produced by the MV2 pipeline (PixSfM + Neural-Angelo, per Section 3.2). Table 4 trains TriplaneGaussian with MV1-Anno vs MV2-Anno point-cloud supervision and reports Chamfer Distance against these MV2-pipeline targets. This is not an independent ground truth: a model trained to imitate Neural-Angelo's output will match Neural-Angelo-style test targets (smoothness, hallucinated regions, completeness bias) and achieve lower CD even if its absolute geometry is no better. Manual cleaning does not remove the shared systematic bias, because cleaners edit the same reconstruction rather than independently measure the object. The CD drop from 0.89 to 0.36 is therefore partly an alignment artifact. Because 'higher-quality dense point clouds' is a headline new feature (abstract item iv, Section 3.2) and Table 4 is its main quantitative support, this confound is load-bearing for the claim that higher-quality annotations boost large reconstruction models. The rendering metrics in Table 4 are less affected, but the shape-quality conclusion is not established. The reader's category-overlap concern is less damaging: if the held-out categories overlapped MV1 categories, the MV1-trained baseline would have an advantage, making MV2's win conservative; the larger issue is the non-independent CD target.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MVImgNet2.0, a multi-view image dataset that extends MVImgNet from roughly 220k objects in 238 classes to about 520k objects in 515 classes, with newly collected videos mostly covering 360-degree views. It also updates the annotation pipeline: masks via a detection-segmentation-tracking approach, camera poses via PixSfM, and dense point clouds via Instant-Angelo/Neural-Angelo. The authors claim these new features yield higher-quality annotations and demonstrate that training large reconstruction models (LGM, LRM, TriplaneGaussian) on the new data improves rendering quality and Chamfer distance compared with training on Objaverse or MVImgNet1.0 data. Per-scene reconstruction with INGP and 3DGS also improves when the MVImgNet2.0 camera poses are used. The dataset, reconstructed point clouds, and annotation code are to be released publicly.","tokens_in":22768,"tokens_out":5770,"duration_ms":50715,"significance":"If the claims hold, MVImgNet2.0 is a substantial community resource: no existing real-world multi-view object dataset combines this scale, category coverage, and 360-degree coverage, and the consistent improvements across three large reconstruction models in Tables 3 and 4 are encouraging. The controlled ablations in Figure 9, isolating data scale, category range, and view range, are a useful addition, and the mask-quality evaluation in the supplementary material includes external datasets (ECSSD, DAVIS) plus a manually annotated subset, which is more than many dataset papers provide. However, the headline shape-quality claim rests on a Chamfer-distance evaluation whose test target is generated by the same reconstruction pipeline used to create the training supervision; that confound must be addressed before the higher-quality point-cloud claim is accepted.","major_comments":[{"comment":"The Chamfer-Distance comparison in Table 4 is not an independent evaluation of shape quality. Section 4.1 states that the test shape ground truths are \"high-quality dense point cloud reconstructions with manual cleaning by annotators,\" and Section 3.2 describes these as outputs of the MVImgNet2.0 pipeline (PixSfM poses plus Instant-Angelo/Neural-Angelo dense reconstruction). Therefore, the row trained with MV2-Anno point-cloud supervision is trained to imitate the same reconstruction procedure that generated the test target. The reported drop in CD from 0.89 to 0.36, and even the 0.40 for the Objaverse-trained model, can partly reflect alignment of systematic biases such as smoothness, completeness, or hallucinated regions rather than improved absolute geometry. Manual cleaning edits the same reconstructions and does not supply an independent measurement. Because feature (iv) of the paper is one of the four headline claims and Table 4 is its main quantitative support, the shape-quality conclusion is not established by the current experiment. Please add an external shape ground truth, for example a small set of objects measured by laser scanning or high-fidelity RGB-D scanning, and report CD against that target; alternatively, rephrase the claim as consistency with the MVImgNet2.0 annotation style and support it with rendering metrics alone.","section":"Table 4; Sections 3.2 and 4.1"},{"comment":"The paper reports only single mean values over the 50 per-scene objects and the 1k test samples, with no error bars, confidence intervals, or multi-seed variance. In Table 3, the reported gains of about 0.3-0.5 dB in PSNR are small enough that without variance information it is difficult to rule out run-to-run or test-set sampling variation, even though the direction of the effect is consistent across models. Please report at least bootstrap confidence intervals over the test set and, ideally, standard deviations over a few independent training runs, for the key comparisons in Tables 3 and 4 and Figure 9.","section":"Tables 2-4; Section 4.1"}],"minor_comments":[{"comment":"The text says the 20 held-out test categories are \"unseen in training\" but does not state whether they overlap with the 238 MVImgNet1.0 categories or are among the 277 new categories. If they overlap, the MV1-Data baseline has an advantage, which would make the reported MV2 wins conservative; if they are disjoint, the comparison is more directly interpretable. Please state the overlap explicitly.","section":"Section 4.1"},{"comment":"The last column header reads \"LIPIS\"; this should be \"LPIPS\".","section":"Table 2"},{"comment":"The dataset name is written \"ECCSD\"; the correct name is ECSSD.","section":"Supplementary Table R.1"},{"comment":"The text says MVImgNet2.0 includes 300k videos and 347 object classes, while the abstract emphasizes the combined 520k objects and 515 categories; the relationship between the newly added portion and the combined total should be stated in one place so that the two numbers do not appear inconsistent.","section":"Section 3.3"},{"comment":"The caption says the figure analyzes three factors, but the figure contains four panels, with panel (d) varying the ratio of MV2-Data at constant total scale; please update the caption to describe all four panels.","section":"Figure 9"}],"recommendation":"major_revision","confidential_remarks":"The dataset is clearly a valuable resource, and the main experimental direction is sound, but the core evidence for the higher-quality point-cloud claim is confounded by the shared annotation pipeline between training supervision and test target. The fix, adding an external shape benchmark or downgrading the claim, is feasible within the scope of a revision. I do not see grounds for reject, and the other weaknesses are addressable with additional reporting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a substantial dataset contribution—300k new objects, 277 new categories, mostly 360-degree captures—and the mask and pose improvements are backed by reasonable evidence. The point-cloud quality claim is the soft spot. Table 4's Chamfer-Distance target is generated by the same MV2 annotation pipeline (PixSfM + Instant-Angelo) used as training supervision, so a model trained on MV2 point clouds has an unfair advantage in matching those targets. Manual cleaning does not remove the shared systematic bias. The CD drop from 0.89 to 0.36 is therefore partly an alignment artifact, not proof that MV2 point clouds are absolutely more accurate. The rendering metrics in the same table are less affected, so the broader claim that MV2 data helps large reconstruction models still stands; it is specifically the \"higher-quality point clouds improve shape reconstruction\" conclusion that is not established by this experiment. A cleaner evaluation would compare against an independent reconstruction (e.g., laser scans or a different photogrammetry pipeline) or at least caveat the result.\n\nThe other issues are minor. There are no error bars or significance tests on any of the tables; with 50 per-scene objects and 1k test samples, reporting variance would be easy and should be added. The held-out test categories are said to be 'unseen in training' but not explicitly stated to be disjoint from MVImgNet1.0's 238 categories; if they overlap, that biases the comparison against MV2, so it is a clarification rather than a fatal flaw.\n\nWhat the paper does well: the dataset itself is the contribution. Mask segmentation improvements are validated on external datasets (ECSSD, DAVIS) plus a manual subset, which is the right kind of evidence. Per-scene reconstruction with PixSfM poses shows a clear improvement, especially for 3DGS (5.8 dB). Training LGM, LRM, and TriplaneGaussian on MV2 data yields consistent rendering gains, so the practical value of the data is believable. The limitations section is honest—excluded categories, annotation refinement, and simple camera trajectories are all acknowledged.\n\nWho is this for? Anyone working on 3D reconstruction, novel view synthesis, or multi-view learning; the released dataset alone justifies the paper. It deserves a serious referee, not a desk reject. I'd send it to review with a request to add error bars, clarify the test category overlap, and either fix or substantially temper the CD-based point-cloud claim.","headline":"A genuinely large and useful dataset expansion whose scale and mask/pose improvements are real, but the headline point-cloud-quality result is weakened by an evaluation target produced by the same pipeline being compared.","tokens_in":23267,"tokens_out":3979,"would_cite":true,"duration_ms":32234,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MVImgNet2.0 expands multi-view real-object imagery to 520k objects in 515 categories and shows that the larger scale, 360-degree coverage, and higher-accuracy masks, camera poses, and point clouds improve large 3D reconstruction models.","keywords":["3D object dataset","multi-view images","3D object reconstruction","image-based modeling","large-scale dataset","point cloud","camera pose estimation","object segmentation"],"falsifier":"Check whether the 20 held-out test categories are disjoint from the 238 categories in MVImgNet1.0; if any overlap exists, re-run Tables 3 and 4 using test categories provably absent from both training sets. As an additional check, train LGM on equal-size, category-matched subsets of MV1-Data and MV2-Data; if the PSNR gap between them disappears or reverses, the claim that the new data's higher quality drives the improvement would not hold.","tokens_in":22293,"feed_emoji":"📦","tokens_out":11932,"duration_ms":85204,"temperature":0.7,"pith_summary":"MVImgNet2.0 expands the existing MVImgNet dataset from about 220,000 to roughly 520,000 real-world objects across 515 categories, adding 300,000 new crowdsourced object videos, most captured in full 360-degree orbits, and re-annotating everything with more accurate foreground masks, camera poses, and dense point clouds. The paper's central claim is that this larger, higher-quality real-image dataset supplies a stronger 3D prior for category-agnostic object reconstruction, and it supports the claim by training three large reconstruction models—LGM, LRM, and TriplaneGaussian—on the new data, the old data, and synthetic data. The new data alone improves reconstruction quality over the old data despite being evaluated as a standalone training set, and combining both datasets improves it further; per-scene reconstructions also improve when the new camera poses are used. The work matters because 3D deep learning lacks a real-world dataset at the scale of 2D benchmarks like ImageNet, and a resource of this size and quality could let reconstruction models learn from the physical world instead of synthetic CAD models.","feed_headline":"520k real objects in 515 categories sharpen 3D reconstruction","feed_subtitle":"Adds 360-degree views and better masks, poses, and point clouds that boost large 3D reconstruction models.","key_machinery":"The load-bearing machinery is the dataset itself together with its annotation pipeline. Raw data consists of crowdsourced object-centric videos shot with phone cameras orbiting the object, most covering a full 360-degree view. Annotations are produced by three upgraded components: Pixel-Perfect Structure-from-Motion (PixSfM) for camera poses, which improves keypoint localization and bundle adjustment using dense feature metric refinement; a detection-segmentation-tracking pipeline that combines Grounding-DINO (open-set detection), SAM (segmentation), and DeAOT (video object tracking) to generate foreground masks; and Neural-Angelo/Instant-Angelo neural surface reconstruction with multi-resolution hash grids for dense point clouds. In the experiments, the large reconstruction models LGM, LRM, and TriplaneGaussian act as the measuring instruments: retraining them on MV2-Data versus MV1-Data versus synthetic Objaverse isolates the contribution of the new dataset, and the LGM-tiny ablations systematically vary data scale, category count, and 360-degree view ratio to attribute the observed gains.","core_discovery":"MVImgNet2.0 is a dataset of roughly 520,000 real-life objects in 515 categories, formed by combining the original 220,000-object MVImgNet with 300,000 newly collected object videos that cover 347 classes, 277 of them new. Compared with its predecessor, most of the new videos circle the object through a full 360-degree view; foreground masks come from a detection-segmentation-tracking pipeline built on Grounding-DINO, SAM, and DeAOT rather than from CarveKit; camera poses come from Pixel-Perfect SfM, which refines keypoints and bundle adjustment with dense features; and dense point clouds come from a Neural-Angelo-based reconstruction with multi-resolution hash encoding. The paper's central discovery is that training large reconstruction models on this data improves both reconstruction quality and generalization: LGM and LRM trained on the new data alone outperform the same models trained on the old data, and training on both datasets improves them further; TriplaneGaussian trained with the 360-degree views and the new point-cloud supervision achieves markedly higher PSNR and lower Chamfer distance; and controlled ablations show that data scale, category breadth, 360-degree view ratio, and MV2-style annotation quality each separately raise reconstruction PSNR.","pith_inferences":["The paper never states whether the 20 held-out test categories overlap with the original 238 MVImgNet categories; if they do, the MV2-Data advantage in Tables 3 and 4 could partly reflect category leakage rather than annotation quality. This is an editorial caution, not a claim made in the paper.","The large per-scene pose effect (about 5.8 dB for 3DGS) hints that camera-pose error, more than mask accuracy or point-cloud density, is the binding constraint for high-frequency radiance-field training; the paper's design does not fully separate these three annotation improvements.","Because the earlier MVImgNet already proved useful for view-consistent understanding, multi-view diffusion, and video generation, MVImgNet2.0's larger scale and broader categories will likely benefit those tasks too, but the paper only tests reconstruction, so those transfers remain unverified.","A testable extension would be to train reconstruction models on mixed Objaverse + MVImgNet2.0 data and measure whether the real-to-synthetic domain gap narrows relative to training on either alone; the paper's comparisons treat the two sources separately rather than jointly."],"forward_implications":["If the claims hold, researchers gain a real-image multi-view dataset at about half ImageNet's scale, allowing 3D reconstruction and generation models to be trained or fine-tuned with less dependence on synthetic CAD data.","The improved camera-pose annotations alone produce large per-scene gains—about 1.1 dB for Instant-NGP and 5.8 dB for 3D Gaussian Splatting—so downstream users should get better novel-view synthesis and radiance-field fitting from the same videos.","The 360-degree coverage and denser point-cloud supervision should make reconstructed shapes more complete than the 180-degree MVImgNet views, which leave backsides and occluded regions poorly constrained.","The scaling curves in the ablation study show continued improvement as added objects grow to 140k, suggesting that further collection along the same pipeline would keep yielding gains rather than saturating.","The combined MVImgNet1.0+2.0 corpus, with 520k objects across 515 categories, becomes a more credible real-world counterpart to large synthetic datasets like Objaverse for training category-agnostic reconstruction models."],"supporting_citations":[{"why":"Defines the predecessor MVImgNet dataset and its 220k objects/238 categories, which MVImgNet2.0 expands and improves upon.","marker":"[Yu et al. 2023]"},{"why":"Provides Pixel-Perfect SfM, the structure-from-motion method used to estimate camera poses with lower error.","marker":"[Lindenberger et al. 2021]"},{"why":"Supplies Grounding-DINO, the open-set detector that generates bounding-box candidates in the mask annotation pipeline.","marker":"[Liu et al. 2023b]"},{"why":"Supplies SAM, the segmentation model that turns detection boxes into foreground object masks.","marker":"[Kirillov et al. 2023]"},{"why":"Provides DeAOT, the video object tracker that propagates selected masks across frames for consistency.","marker":"[Yang and Yang 2022]"},{"why":"Supplies Neural-Angelo, the neural surface reconstruction method with multi-resolution hash encoding used for dense point clouds.","marker":"[Li et al. 2023b]"},{"why":"Provides Instant-Angelo, the open-source implementation of Neural-Angelo that the authors use to produce point clouds quickly.","marker":"[Ye 2023]"},{"why":"Defines LRM, one of the large single-view reconstruction models trained on the compared datasets.","marker":"[Hong et al. 2023]"},{"why":"Defines LGM, one of the large multi-view reconstruction models used for the main training-data comparisons and ablations.","marker":"[Tang et al. 2024]"},{"why":"Defines TriplaneGaussian, the single-view reconstruction model whose point-cloud supervision lets the authors test 360-degree views and new dense reconstructions.","marker":"[Zou et al. 2023]"}],"fun_headline_variants":["520k objects, 515 categories, 360° views sharpen 3D reconstruction","MVImgNet2.0: 520k objects, 515 classes, 360° views for better 3D","Full 360° multi-view data for 520k objects boosts 3D reconstruction","MVImgNet2.0 expands to 520k objects, adds 360° views and refined point clouds","MVImgNet2.0: 520k objects, 360° views, refined point clouds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The experiments that attribute the performance gains to MVImgNet2.0's higher quality assume that the 20 held-out test categories are equally unfamiliar to models trained on the old MVImgNet data and on the new data, so that neither training set has already memorized the test categories; if those test categories overlap with the original 238 categories, the reported advantage of MV2-Data over MV1-Data would be biased in favor of the new data.","fun_headline_variants_meta":{"raw":{"variants":["520k objects, 515 categories, 360° views sharpen 3D reconstruction","MVImgNet2.0: 520k objects, 515 classes, 360° views for better 3D","Full 360° multi-view data for 520k objects boosts 3D reconstruction","MVImgNet2.0 expands to 520k objects, adds 360° views and refined point clouds","MVImgNet2.0: 520k objects, 360° views, refined point clouds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001839,"raw_usage":{"total_tokens":7304,"prompt_tokens":1099,"completion_tokens":6205,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":715,"completion_tokens_details":{"reasoning_tokens":6078}},"tokens_in":715,"tokens_out":6205,"duration_ms":40361,"temperature":1.0,"reasoning_tokens":6078,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:22:54.989157+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check whether the 20 held-out test categories are disjoint from the 238 categories in MVImgNet1.0; if any overlap exists, re-run Tables 3 and 4 using test categories provably absent from both training sets. As an additional check, train LGM on equal-size, category-matched subsets of MV1-Data and MV2-Data; if the PSNR gap between them disappears or reverses, the claim that the new data's higher quality drives the improvement would not hold.","supporting_citations":[],"review_version":1}