{"id":"e458e14c-6bc9-40c7-a134-325d50fcbee8","arxiv_id":"2508.13043","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"IntelliCap combines spatial coverage meshes with LLM-ranked spherical proxies to guide photo capture, improving 3DGS and NeRF rendering quality in indoor scenes.","lead":"This paper describes an augmented reality system that tells a person taking photos for 3D scene reconstruction where to point the camera next, using overlays for uncovered areas and spheres around objects that need extra viewing angles. In a 12-person user study, the guided captures produced better final renderings than no guidance or spatial coverage only.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central causal claim is unisolated: the Ours condition bundles sphere proxies, LLM ranking, and significantly more images (320 vs 232, p=0.02), and the auxiliary closeness-to-GT result double-counts the evaluation viewpoints.","rationale":"I read the central claim as: combining spatial-coverage visualization with LLM-prioritized spherical proxies guides operators to collect views that render more accurately than unguided (NV) or spatial-only (SC) capture. The evaluation has real strengths: repeated-measures design, official 3DGS and Nerfacto implementations, COLMAP pose estimation, and examiner-collected ground truth randomly selected from a larger stock, a genuine improvement over the trajectory-biased selection criticized in [35]. The load-bearing weakness is internal validity. The Ours condition bundles three changes at once: sphere proxies, LLM-based object selection, and a significantly larger number of captured images (320 vs 232, p=0.02, §4.3). Because rendering quality is a function of training-view quantity and coverage, the Table 2 margins are compatible with a pure count effect; the Ours-vs-NV comparison is better matched in count, but it cannot separate the sphere component from the spatial component because Ours adds both. The paper never includes a condition with spheres but without LLM ranking, so the mechanism advertised in the abstract, LLM category scoring from §3.2 with an undisclosed threshold, is not isolated. I also weigh the self-asserted limitations: 'We do not fully utilize the LLM scores' and the closed-vocabulary caveat both point to the same gap. The closeness-to-ground-truth result in §4.3 is double-edged in the way the paper itself criticizes in standard evaluations: since the same ground-truth viewpoints are the evaluation set, and these renderers interpolate best near training views, Ours's smaller training-to-ground-truth distance both explains and may inflate its quality scores. This is not an accusation; it is a question of whether the reported superiority reflects the LLM prioritization or simpler quantity and coverage effects. The reader's weakest_assumption identified largely the same gap (unablated LLM premise and capture-count confound); I sharpen it to an attribution problem and add the evaluation-set entanglement. Because the concern reinforces rather than overturns the reader's assessment, the verdict stays CONDITIONAL: the direction is publishable and the evaluation design is thoughtful, but the central causal claim needs the matched-count re-analysis and an ablation separating LLM ranking from mere sphere presence before it can be accepted as established.","tokens_in":16526,"tokens_out":27213,"duration_ms":257599,"concrete_test":"Matched-count re-analysis of the existing data: for each participant, subsample the Ours and NV capture sets down to the number of images captured in that participant's SC session, then re-run COLMAP and the official 3DGS and Nerfacto pipelines with identical hyperparameters, computing PSNR/SSIM/LPIPS at the same examiner ground-truth viewpoints with paired confidence intervals. If the Ours-vs-SC margins in Table 2 collapse or lose significance at matched counts, the headline advantage is a capture-count effect rather than an effect of LLM-prioritized guidance. If the margins survive, the count confound is refuted; the remaining question is whether the LLM ranking (versus spheres on all or random objects) is responsible, which the paper currently does not test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, that LLM-prioritized spherical angular coverage on top of spatial coverage yields photo sets that render better than conventional strategies, requires that the Table 2 advantage be caused by the LLM-prioritized spheres. The experiment does not show this. In §4.3, participants in Ours captured 320±169 images versus 231±131 in SC (p=0.02, r=0.75), so the two conditions differ in treatment and in input quantity at once. Since 3DGS and Nerfacto quality improves with the number and coverage of training views, the Ours-vs-SC margins (3DGS PSNR 18.891 vs 17.086; LPIPS 0.201 vs 0.253) are consistent with a capture-count effect alone. The higher count may be a designed consequence of Ours's completion signal, but that still leaves the distinctive mechanism, the LLM scores of §3.2 averaged over five axes with an undisclosed threshold and a non-deterministic LLM (MS Copilot), unisolated: no condition presents spheres without LLM ranking, so the contribution of the ranking, versus the mere presence of angular-sampling targets, is untested. The paper's Limitations section concedes 'We do not fully utilize the LLM scores' and that the closed vocabulary is a limitation. The auxiliary evidence that Ours training views lie closest to the examiner's ground-truth viewpoints (0.239±0.047 m vs 0.320/0.337 m; 12.4° vs 18.5°/17.1°, §4.3) is double-edged: the same ground-truth views define the Table 2 quality metrics, and renderers generalize best near training views, so closeness to ground truth partly explains, and may inflate, the reported advantage. This does not prove the claim false; it shows the evidence is compatible with simpler explanations. The matched-count re-analysis in the concrete test would decide which explanation holds.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents IntelliCap, an AR-based guidance system for human image capture for novel view synthesis. The system combines a spatial-coverage mesh visualization (pink-white stripes over unreconstructed regions) with spherical proxies placed around objects whose semantic categories have been scored by an LLM (Detectron2 categories, five appearance axes, averaged score). The authors report a within-subjects user study (N=12, three scenes) comparing no guidance (NV), spatial coverage only (SC), and the full system (Ours), with 3DGS and Nerfacto reconstructions evaluated by PSNR/SSIM/LPIPS against ground-truth views collected by an examiner. Table 2 reports that Ours achieves the best means on all six metric/model combinations (e.g., 3DGS PSNR 18.891 vs 17.086 for SC and 16.338 for NV), and the paper claims superior performance in real scenes compared to conventional view sampling strategies.","tokens_in":16907,"tokens_out":6053,"duration_ms":59263,"significance":"The problem addressed—helping non-expert humans capture view sets for 3DGS/NeRF in open scenes—is timely and under-studied, and the use of an external examiner-collected ground-truth set is a genuine methodological improvement over evaluating on participant-captured views. The paper also uses two renderers and three standard metrics, reports usability statistics from a repeated-measures design, and is unusually candid about its limitations, including the under-utilization of the LLM scores and the closed vocabulary. If the performance claim is causally attributable to the proposed guidance, this would be a useful contribution to human-in-the-loop acquisition. However, the current experimental design does not isolate the mechanism responsible for the observed quality gains, and the central quantitative claim is therefore not yet established at the level the abstract asserts.","major_comments":[{"comment":"The central causal claim is unisolated because the Ours condition bundles several components: the sphere proxies, the LLM ranking, and a completion signal that led participants to capture significantly more images (SC: 231.67±131.31; Ours: 320.17±169.03; p=0.02, r=0.75). Since 3DGS and Nerfacto quality generally improve with the number and spatial coverage of training views, the Table 2 margins (e.g., 3DGS PSNR 18.891 vs 17.086; LPIPS 0.201 vs 0.253) are consistent with a pure capture-count effect. The paper needs either an ablation that presents spheres without LLM ranking, or an analysis that controls for image count (e.g., matching subsets, or a regression with capture count as a covariate), before the abstract claim that the proposed strategy yields superior performance can be supported.","section":"§4.3 and Table 2"},{"comment":"The LLM-scoring mechanism that decides which objects receive spherical proxies is not specified reproducibly. The text reports a prompt template, five appearance axes, and that category labels are 'pre-scored by an LLM (e.g., MS Copilot)', but it does not state the threshold used to tag objects, the exact prompt, the LLM version or parameters, how non-determinism was handled, or whether the scores are cached. The Limitations section further concedes 'We do not fully utilize the LLM scores.' Without this information and without an ablation or sensitivity analysis of the threshold, the paper cannot substantiate the claim that LLM-prioritized angular coverage (rather than the mere presence of angular-sampling targets) is responsible for the results.","section":"§3.2"},{"comment":"The auxiliary closeness-to-ground-truth analysis is double-edged and may itself explain the Table 2 improvements. The same ground-truth viewpoints are used both to compute the PSNR/SSIM/LPIPS scores and to measure how close each condition's training views are to those viewpoints; Ours has the smallest distance and angular difference (0.239 m and 12.4° vs 0.320/0.337 m and 18.5°/17.1°). Because radiance-field renderers generalize best near training views, this proximity can inflate the reported quality of Ours independently of any benefit of the guidance. The evaluation should use held-out viewpoints not present in the training set, or should statistically control for training-view-to-test-view distance.","section":"§4.3 and Table 2"},{"comment":"No inferential statistics are reported for the rendering-quality scores. Table 2 gives only means and standard deviations across what appear to be participant/scene replications, yet the paper's headline claim is that Ours 'outperforms the baselines in all image quality metrics.' With N=12 and three conditions, the differences (e.g., Nerfacto SSIM 0.616 vs 0.574) may or may not be reliable; the authors should report paired comparisons per participant (or per scene) with appropriate corrections, or a mixed model with participant and scene as random effects, to support the claim.","section":"§4.2 Table 2"}],"minor_comments":[{"comment":"The text says NV and Ours collected a similar number of images, but no descriptive statistics for NV are given; only SC and Ours means are reported. Please provide the full table of counts and the pairwise test between NV and Ours.","section":"§4.3"},{"comment":"The phrase 'pre-scored by an LLM (e.g., MS Copilot)' suggests the scores are computed offline, but the system description elsewhere implies online detection; clarify the offline/online split and whether the scores are deterministic across runs.","section":"§3.2"},{"comment":"Several parameters (depth tolerance t=5.0 cm, sphere-display FoV thresholds of 20% and 100%, radius scaling k, merge radius cap) are introduced in prose; a table of all parameters with default values and units would improve reproducibility.","section":"§3.3"},{"comment":"The ground-truth set size (20 images per scene) and the criterion 'randomly selected from the full stock' are not justified; please state the total stock size and the selection procedure.","section":"§4.1"},{"comment":"The post-task rankings (Q1–Q6) are summarized without a clear definition of the ranking scale and how ties were handled; please specify the scale and analysis procedure.","section":"§4.3 / Figure 9"},{"comment":"Some symbols and Greek letters appear corrupted in the provided PDF (e.g., 'ﬂX', 'c 2', 'p 𝑚 0:01'); please ensure the camera-ready version renders all formulas correctly.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be a preprint of an already-accepted ISMAR paper; if this is also intended as a journal submission, the authors should clarify the additional contribution over the conference version. The core issue in my review is experimental isolation rather than novelty or presentation. Editors may also wish to confirm that the overlap with the ISMAR version is disclosed appropriately."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take for your file. The paper is worth a serious referee, but the central claim needs to be read with care.\n\nWhat's actually new: IntelliCap combines a spatial-coverage mesh (camera sees missing areas as pink stripes) with spherical proxies around objects that get denser angular sampling, and uses an LLM to score object categories (specularity, transparency, etc.) to decide which objects get spheres. That combination for open multi-object scenes is not in the prior work they cite, which is single-object or forward-facing. The user study is also a genuine attempt to evaluate a moving-target task: they collect 12 participants across three conditions, then render 3DGS and Nerfacto and compare against an independent examiner's ground-truth views. Using a second person's capture to build the test set is a real improvement over the usual 'split the same trajectory' evaluation.\n\nThe results are consistent: Ours beats SC and NV on all three metrics for both renderers. And the paper is honest about limitations, including that it doesn't fully use the LLM scores.\n\nNow the soft spots, in proportion.\n\nFirst, the Ours condition bundles the spheres, the LLM ranking, and significantly more images (320±169 vs 231±131, p=0.02). So the Table 2 gap over SC is compatible with a 'more training views' effect. The paper notes NV and Ours had similar image counts, which argues against count alone explaining Ours vs NV, but it doesn't isolate the LLM ranking from the mere presence of angular targets. There's no condition with spheres but without LLM scoring. The LLM threshold and the number of subsurfaces N are undisclosed, and the LLM is non-deterministic.\n\nSecond, the auxiliary 'closeness to ground truth' evidence is double-edged. The same GT views define the Table 2 metrics, and renderers generalize best near training views. So Ours training views being closest to GT (0.239m vs 0.320/0.337m) is consistent with, and may partly explain, the quality advantage—not independent confirmation.\n\nThird, no inferential statistics on the quality metrics. With n=12, that's underpowered for detecting anything but large effects.\n\nThese don't sink the paper. The system likely does help people capture better data. But the distinctive mechanism—LLM-prioritized angular coverage—is not actually demonstrated. A matched-count comparison or an ablation with spheres but fixed (or randomized) object ranking would settle it.\n\nRecommendation: send to peer review. A good reviewer should ask for that ablation, for confidence intervals on Table 2, and for the undisclosed parameters. If the authors can supply a sphere-without-LLM condition, this becomes a much stronger paper.","headline":"The system is a real step forward for AR-guided capture in open scenes, but the paper never isolates its headline LLM ranking—Ours differs from SC in both guidance and image count, and the closest-to-GT result double-counts the test views.","tokens_in":17494,"tokens_out":3531,"would_cite":true,"duration_ms":35525,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Guided AR capture beats unguided scanning for 3D view synthesis","keywords":["view sampling","3D Gaussian splatting","neural radiance fields","augmented reality guidance","large language models","view-dependent effects","semantic segmentation","novel view synthesis"],"falsifier":"Run the same scenes and participants but replace the LLM priority list with a random or fixed ordering while matching the number of captured images to the guided condition; if PSNR, SSIM, and LPIPS stay statistically indistinguishable, the LLM ranking is not the active ingredient. Alternatively, compute per-object view-synthesis error and check whether objects above the priority threshold show disproportionately large error reductions relative to objects below it.","tokens_in":16326,"feed_emoji":"📷","tokens_out":7399,"duration_ms":73471,"temperature":0.7,"pith_summary":"The paper claims that the weak point in high-quality 3D scene reconstruction is not the rendering algorithm but how a human chooses camera viewpoints, and that a situated augmented-reality guide can fix that. Its IntelliCap system shows two complementary hints live during capture: striped overlays on unvisited scene geometry for spatial coverage, and 3D spheres around objects that likely need dense angular coverage for view-dependent appearance. Object categories come from semantic segmentation and are scored by a large language model on five appearance axes, with only high-scoring objects receiving spheres. In a user study with 12 participants across three real scenes, images collected with this combined guidance yielded the best PSNR, SSIM, and LPIPS for both 3D Gaussian splatting and Nerfacto, compared with no guidance and spatial-only guidance. The practical consequence, if the claim holds, is that non-expert camera operators can collect view-synthesis data with more consistent quality than current best practice.","feed_headline":"Guided AR capture beats unguided scanning for 3D view synthesis","feed_subtitle":"Spatial-plus-angular capture guidance beats unguided and spatial-only scanning in real scenes.","key_machinery":"The load-bearing mechanism is the 3D sphere proxy, a sphere anchored to a detected object whose surface is subdivided into view-sample patches. Each patch corresponds to one viewing direction, and a patch becomes transparent when the user has captured that direction, so the sphere acts as a progress meter for angular coverage. The spheres are generated only for objects whose category exceeds an LLM-computed priority threshold, where priority is the average of scores on geometric complexity, texture complexity, size, specularity, and transparency. A second visualization, the pink-white striped overlay on the live mesh, tracks spatial coverage. The two are combined by merging nearby spheres, hiding spheres behind occluders, and suppressing spheres outside a comfortable capture distance.","core_discovery":"The paper's central claim is that combining spatial coverage guidance with LLM-ranked angular coverage guidance produces view samples that render better than conventional capture strategies in real open scenes. For 3D Gaussian splatting, the guided condition reached a mean PSNR of 18.891 dB versus 17.086 dB for spatial-only and 16.338 dB for no guidance, with the same ordering on SSIM and LPIPS, and the Nerfacto results show the same ordering. The authors attribute the gain to the angular component: without it, users tend to stay at one height and miss reflective and transparent objects, whereas the spherical proxies push them to capture those objects from many directions. They also introduce an evaluation design in which ground-truth views are captured independently by an examiner, arguing that splitting a single operator's trajectory unfairly favors that trajectory.","pith_inferences":["I would test the LLM ranking directly: hold the number of captured photos constant and compare the guided condition against a condition with random or fixed object priorities; if rendering quality does not drop, the active ingredient is extra quantity or coverage rather than semantic ranking.","Because the LLM scores are assigned per category rather than per instance, they could be precomputed and published as a reusable capture-priority table for common object categories, then validated against per-object view-dependent error.","The same dual-proxy design could be adapted to outdoor or dynamic scenes by suppressing moving categories such as people and cars and using reachability constraints for distant spheres, which the paper lists as future work."],"forward_implications":["Capture guidance for 3D Gaussian splatting and NeRF can shift from taking as many photos as possible to targeted, on-the-fly instructions that tell a non-expert when an object's angular coverage is complete.","The same object-priority scoring could be ported to automated next-best-view planners, because it separates the question of which objects matter from the question of where the camera should stand.","Evaluating view-sampling strategies fairly requires independent ground-truth capture by a second person; the paper reports that standard same-trajectory splits inflate PSNR, SSIM, and LPIPS scores substantially, by about 37%, 15%, and 49% respectively.","The closed-vocabulary limitation means the method's benefit is currently bounded by the set of object categories the segmentation model recognizes; expanding to open-vocabulary recognition is a stated future direction."],"supporting_citations":[{"why":"Defines the 3D Gaussian splatting scene representation used as a primary evaluation target.","marker":"[12]"},{"why":"Defines the Nerfacto neural radiance field variant used as the second evaluation target.","marker":"[31]"},{"why":"Supplies the semantic segmentation and category labels that trigger the LLM scoring.","marker":"[34]"},{"why":"Introduces sphere-based angular coverage visualization for single objects, which the method extends to scene-scale scanning.","marker":"[3]"},{"why":"Provides the mixed-reality sphere proxy paradigm and its coverage-by-capture interaction.","marker":"[19]"},{"why":"Represents the prescriptive-sampling baseline for forward-facing scenes that this method aims to generalize beyond.","marker":"[18]"},{"why":"Defines the shiny-object dataset design whose specular and transparent materials motivate the angular coverage evaluation.","marker":"[33]"},{"why":"Provides the structure-from-motion pipeline used to estimate camera poses for both training and ground-truth views.","marker":"[27]"}],"fun_headline_variants":["Spherical proxy guidance adds angular coverage to beat unguided capture","IntelliCap's object-ranked proxies improve view sampling quality","Angular guidance from vision-language ranking beats unguided scanning","Guided capture with angular coverage outperforms spatial-only scanning","Vision-guided spherical proxies yield better renders in real scenes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise, stated in Section 3.2, is that a large language model scoring a fixed set of object categories on five appearance axes correctly predicts which objects need denser angular sampling; if those scores are wrong, the spheres steer users to the wrong objects and the measured gains could mostly reflect the larger number of images taken in the guided condition.","fun_headline_variants_meta":{"raw":{"variants":["Spherical proxy guidance adds angular coverage to beat unguided capture","IntelliCap's object-ranked proxies improve view sampling quality","Angular guidance from vision-language ranking beats unguided scanning","Guided capture with angular coverage outperforms spatial-only scanning","Vision-guided spherical proxies yield better renders in real scenes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00089,"raw_usage":{"total_tokens":3811,"prompt_tokens":889,"completion_tokens":2922,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":2841}},"tokens_in":505,"tokens_out":2922,"duration_ms":23286,"temperature":1.0,"reasoning_tokens":2841,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:15:37.525098+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same scenes and participants but replace the LLM priority list with a random or fixed ordering while matching the number of captured images to the guided condition; if PSNR, SSIM, and LPIPS stay statistically indistinguishable, the LLM ranking is not the active ingredient. Alternatively, compute per-object view-synthesis error and check whether objects above the priority threshold show disproportionately large error reductions relative to objects below it.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the 3D Gaussian splatting scene representation used as a primary evaluation target."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Nerfacto neural radiance field variant used as the second evaluation target."},{"cited_title":"S \\\"u nderhauf, J","cited_arxiv_id":null,"evidence_quote":"Supplies the semantic segmentation and category labels that trigger the LLM scoring."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents the prescriptive-sampling baseline for forward-facing scenes that this method aims to generalize beyond."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the shiny-object dataset design whose specular and transparent materials motivate the angular coverage evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the structure-from-motion pipeline used to estimate camera poses for both training and ground-truth views."}],"review_version":2}