{"id":"be5ac3b1-8fb1-4cc3-a5a1-66ba3a05f610","arxiv_id":"2606.19733","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"QueryGaussian decouples semantic understanding from 3D geometry via 2D model lifting and temporal fusion, matching prior accuracy while cutting memory over 70% and speeding inference 180x on city-scale Gaussian scenes.","lead":"QueryGaussian is a training-free method for finding specific 3D objects in large scenes using everyday language prompts. It avoids the memory explosion of prior approaches by querying instances on demand instead of pre-embedding semantics everywhere.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Whether concurrent max-weight mask lifting plus multi-stage adaptive clustering preserves cross-view consistency without any training or tuning on city-scale data","rationale":"The reader's weakest assumption correctly isolates the single technical hinge on which the training-free scalability argument turns. Because the supplied review is abstract-only and the full manuscript was not examined, no stronger or different objection can be substantiated at this stage.","tokens_in":1757,"tokens_out":306,"duration_ms":11372,"concrete_test":"On a city-scale test scene (≥10M Gaussians) with held-out ground-truth instance labels, run the full QueryGaussian pipeline once with the published association and clustering hyperparameters; compare top-1 retrieval mAP against the strongest reported baseline. If mAP falls >5 points below the baseline, the no-tuning consistency claim does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The efficiency and scalability claims rest on the instance-level query path (pre-trained 2D models + concurrent maximum-weight association + temporal fusion with multi-stage adaptive density clustering) delivering reliable 3D instance masks without scene-specific tuning. Projection ambiguities, view-dependent appearance changes, and density variations are inherent in city-scale Gaussian scenes; the abstract gives no indication that the association and clustering steps are provably robust to these without additional regularization or adaptation. If this step fails to maintain semantic-visual consistency, the claimed 70% memory reduction and 180x speedup cannot be realized while matching SOTA accuracy.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes QueryGaussian, a training-free open-vocabulary 3D instance retrieval framework for large-scale Gaussian scenes. It decouples semantics from geometry by using pre-trained 2D vision models to process natural language prompts, lifting 2D segmentation masks into 3D via concurrent maximum-weight association, and applying a temporal fusion module with multi-stage adaptive density clustering to address projection ambiguities. The central claims are that this achieves accuracy parity with state-of-the-art methods while reducing GPU memory by over 70% and accelerating inference by 180x, enabling city-scale retrieval on scenes with tens of millions of Gaussians using consumer hardware.","tokens_in":1909,"tokens_out":454,"duration_ms":12245,"significance":"If the accuracy and efficiency claims hold under rigorous evaluation, the work would be significant for enabling scalable 3D instance retrieval without scene-specific training or linear memory scaling. The training-free design leveraging existing 2D models and the explicit focus on city-scale feasibility are clear strengths that could impact multimedia analysis and 3D scene understanding applications.","major_comments":[{"comment":"Abstract: the central efficiency claims (70% memory reduction, 180x speedup, city-scale operation) rest on the instance-level query path preserving semantic-visual consistency, yet the abstract provides no information on evaluation datasets, baselines, error bars, or data exclusions, preventing verification of whether the concurrent maximum-weight association and multi-stage adaptive density clustering actually deliver the claimed accuracy parity.","section":"Abstract"},{"comment":"Method description (lifting and fusion): the assumption that concurrent max-weight mask lifting plus temporal fusion with multi-stage adaptive density clustering maintains cross-view consistency without any training or scene-specific tuning is load-bearing for the scalability claim, but no analysis of robustness to projection ambiguities, view-dependent appearance, or density variations in tens-of-millions-Gaussian scenes is referenced.","section":"Method"}],"minor_comments":[{"comment":"The abstract would benefit from explicit mention of the datasets and quantitative metrics used to support the accuracy parity claim.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address each major comment below and propose targeted revisions to enhance the manuscript.","responses":[{"response":"We agree that the abstract would benefit from additional context to support verifiability of the claims. We will revise the abstract to reference the main evaluation datasets (including city-scale scenes with tens of millions of Gaussians), the primary baselines, and indicate that accuracy results with error bars appear in the experiments. This change directly addresses the concern while preserving conciseness.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central efficiency claims (70% memory reduction, 180x speedup, city-scale operation) rest on the instance-level query path preserving semantic-visual consistency, yet the abstract provides no information on evaluation datasets, baselines, error bars, or data exclusions, preventing verification of whether the concurrent maximum-weight association and multi-stage adaptive density clustering actually deliver the claimed accuracy parity."},{"response":"The experiments section validates performance on large scenes via ablations of the fusion module. To strengthen the presentation of the load-bearing assumption, we will add a dedicated robustness analysis subsection (drawing on existing results and qualitative examples) that explicitly discusses handling of projection ambiguities, view-dependent effects, and density variations. This addition will reference the design elements without requiring new experiments.","revision_made":"yes","referee_comment":"[Method] Method description (lifting and fusion): the assumption that concurrent max-weight mask lifting plus temporal fusion with multi-stage adaptive density clustering maintains cross-view consistency without any training or scene-specific tuning is load-bearing for the scalability claim, but no analysis of robustness to projection ambiguities, view-dependent appearance, or density variations in tens-of-millions-Gaussian scenes is referenced."}],"tokens_in":1430,"tokens_out":386,"duration_ms":18748,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this work avoids the linear memory cost of embedding semantics into every 3D Gaussian by querying instances directly. It lifts 2D segmentation masks from pre-trained models using concurrent maximum-weight association and adds a temporal fusion module with multi-stage adaptive density clustering to handle cross-view issues. The result is presented as training-free and able to run city-scale retrieval on consumer hardware while matching prior accuracy.\n\nThis approach is a direct response to the scaling problem in scene-level methods, and the decoupling of semantics from geometry is a sensible move that could make open-vocabulary search practical for very large environments. The use of existing 2D models without any fine-tuning keeps the method lightweight and easy to reproduce in principle.\n\nThe soft spots sit in the evaluation. The abstract states 70% memory reduction and 180x speedup with accuracy parity, yet the paper must show the exact datasets, baselines, run counts, and whether city-scale tests used real tens-of-millions-of-Gaussians scenes or smaller proxies. Projection ambiguities and density variations are common in such data, so the lifting and clustering steps need clear evidence that they maintain consistency without scene-specific tuning; if the ablations or failure cases are missing, that weakens the central claim. Minor gaps in reporting error bars or data exclusions would also need fixing.\n\nThis paper is for researchers working on scalable 3D scene understanding, robotics mapping, or AR retrieval who already use Gaussian representations. A reader focused on efficiency trade-offs would get concrete ideas from the query path even before the numbers are fully verified.\n\nIt deserves peer review so the implementation details and results can be checked against the claims.","headline":"QueryGaussian shifts to instance-level querying with 2D mask lifting and temporal clustering to cut memory and speed up retrieval on large Gaussian scenes, but the big efficiency numbers rest on unshown experimental details.","tokens_in":2412,"tokens_out":422,"would_cite":false,"duration_ms":24431,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"QueryGaussian retrieves open-vocabulary 3D instances from city-scale scenes by lifting 2D masks into 3D without training or scene-wide embeddings.","keywords":["3D instance retrieval","open-vocabulary","Gaussian splatting","training-free","scalable 3D search","semantic lifting","temporal fusion","city-scale scenes"],"falsifier":"Apply the method to a city-scale scene containing at least ten million Gaussians, run a set of text-prompt retrievals on consumer hardware, and check whether it finishes without out-of-memory errors and matches ground-truth instance labels at rates comparable to prior methods.","tokens_in":2672,"feed_emoji":"🔎","tokens_out":699,"duration_ms":16124,"temperature":0.7,"pith_summary":"Existing approaches embed semantic features into every 3D primitive, so memory and compute costs grow directly with scene size and cause failures on large environments. QueryGaussian instead uses pre-trained 2D models to interpret text prompts and projects segmentation masks into 3D through concurrent maximum-weight association. A temporal fusion module with multi-stage adaptive density clustering resolves projection ambiguities across views. The result matches prior accuracy while cutting GPU memory by over 70 percent and speeding inference by 180 times, allowing retrieval on scenes with tens of millions of Gaussians on ordinary hardware.","feed_headline":"Text prompts retrieve 3D instances from city-scale scenes at 180x speed","feed_subtitle":"Training-free method projects 2D masks into 3D instead of embedding every point, slashing memory use by over 70 percent.","key_machinery":"Instance-level query mechanism that lifts 2D segmentation masks into 3D via concurrent maximum-weight association plus temporal fusion with multi-stage adaptive density clustering.","core_discovery":"QueryGaussian is a training-free framework for open-vocabulary 3D instance retrieval that decouples semantic understanding from geometric representation by lifting 2D segmentation masks into 3D via concurrent maximum-weight association and a temporal fusion module with multi-stage adaptive density clustering, thereby avoiding the linear scaling of memory and compute that occurs when semantic features are distilled into every primitive.","pith_inferences":["The same decoupling of semantics from geometry could be tested on other 3D representations such as point clouds or meshes to check whether the efficiency gain generalizes.","The temporal fusion module might extend naturally to video sequences, allowing retrieval in dynamic rather than static scenes.","Because the method avoids storing semantic features per primitive, it opens the possibility of on-the-fly retrieval during interactive navigation of very large environments."],"forward_implications":["GPU memory usage drops by more than 70 percent relative to scene-embedding baselines.","Inference accelerates by a factor of 180 while accuracy remains comparable to state-of-the-art methods.","Retrieval becomes feasible on city-scale scenes holding tens of millions of Gaussians using only consumer-grade hardware.","No per-scene training or tuning is required because the method relies on off-the-shelf 2D vision models."],"fun_headline_variants":["Training-free 3D instance retrieval scales to city scenes via text prompts","QueryGaussian uses 2D masks lifted to 3D for open-vocabulary retrieval","Decoupled approach avoids linear memory scaling in large 3D scenes","2D vision models enable efficient 3D instance search without training"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Lifting 2D segmentation masks into consistent 3D instances through maximum-weight association and adaptive density clustering will preserve semantic-visual alignment across views without any training or scene-specific tuning.","fun_headline_variants_meta":{"raw":{"variants":["Training-free 3D instance retrieval scales to city scenes via text prompts","QueryGaussian uses 2D masks lifted to 3D for open-vocabulary retrieval","Decoupled approach avoids linear memory scaling in large 3D scenes","2D vision models enable efficient 3D instance search without training"]},"model":"grok-4.3","cost_usd":0.007242,"raw_usage":{"total_tokens":3349,"prompt_tokens":689,"num_sources_used":0,"completion_tokens":78,"cost_in_usd_ticks":72424500,"prompt_tokens_details":{"text_tokens":689,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2582,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":689,"tokens_out":78,"duration_ms":19639,"temperature":1.0,"reasoning_tokens":2582,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T17:50:19.058572+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Apply the method to a city-scale scene containing at least ten million Gaussians, run a set of text-prompt retrievals on consumer hardware, and check whether it finishes without out-of-memory errors and matches ground-truth instance labels at rates comparable to prior methods.","supporting_citations":[],"review_version":1}