{"id":"43d935f0-aed8-4a21-a664-ad350aa8d3a1","arxiv_id":"2607.05347","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":7,"one_line_summary":"WildSplat decouples geometry from appearance in a single feedforward pass to produce appearance-conditioned 3D Gaussian reconstructions from unposed in-the-wild images.","lead":"WildSplat is a feedforward 3D Gaussian Splatting system that reconstructs scenes from unposed, in-the-wild photos with varying illumination, conditioning output appearance on a reference image. A smart generalist might read it because it removes the need for per-scene optimization and pre-computed poses when handling internet photo collections, which could streamline VR/AR content pipelines.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Feedforward baselines (AnySplat, WorldMirror) receive no reference image, so the half-image reference protocol gives WildSplat appearance information that competitors cannot use, making the headline performance gaps not directly comparable.","rationale":"The reader correctly identified the small test set and half-image protocol as concerns, but the more load-bearing issue is the information asymmetry between WildSplat and its feedforward baselines. WildSplat receives target-appearance conditioning via the reference image; AnySplat and WorldMirror do not. This means the performance gaps in Tables 1-2 conflate 'better appearance disentanglement' with 'has access to target appearance that competitors lack.' The reader's concern about generalization (14 scenes, no error bars) is valid but secondary — even with more scenes, the comparison would remain unfair if the protocol is unchanged. The ablation discrepancy (Table 3 vs Table 2 for the same 4-view Phototourism setting: 18.52 vs 19.57) also needs explanation. The methodological contribution (dual-branch decoupling with appearance injection) is genuinely novel and the qualitative results are compelling, but the quantitative superiority claim over feedforward baselines is not well-supported under the current protocol. I maintain CONDITIONAL: the verdict should require (1) a fair comparison where baselines receive equivalent appearance information, (2) clarification of the ablation/main-table discrepancy, and (3) broader evaluation. No code is released, making independent verification impossible.","tokens_in":14750,"tokens_out":1833,"duration_ms":74524,"concrete_test":"Re-evaluate AnySplat and WorldMirror with a post-hoc appearance alignment step: use the left-half reference image to apply a color affine transform (e.g., least-squares color mapping) to their rendered right-half outputs before computing metrics. If the gap between WildSplat and these baselines shrinks to under 1 PSNR, the headline claim of 'surpassing feedforward methods' is substantially weakened. Separately, evaluate WildSplat using a reference image from a different viewpoint (not the same GT image) to measure how much of the gain comes from same-viewpoint appearance leakage.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that WildSplat 'surpasses existing feedforward methods' (Tables 1-2). However, the evaluation protocol creates an information asymmetry: WildSplat receives the left half of each ground-truth image as an appearance reference (§4.1, 'we use the left half of the ground-truth image as the reference'), while AnySplat and WorldMirror are feedforward methods with no reference-image input mechanism. The paper itself states these baselines 'cannot decouple geometry and appearance' and produce 'inconsistent appearances' — but this is expected when one method gets target-appearance conditioning and the others do not. The +4.19 PSNR gap over WorldMirror on Phototourism 4-view could be substantially explained by this asymmetry rather than by superior geometry or disentanglement. For optimization-based methods like WildGaussian, the paper notes they 'can optimize an appearance embedding using a reference image,' making those comparisons more fair, but the feedforward comparisons — which are the more novel and interesting baselines — are not apples-to-apples. Additionally, the half-image protocol means the reference comes from the same viewpoint and same image as the target, so the model sees the exact target appearance (just spatially offset), which is an easier conditioning task than cross-view appearance transfer. The ablation numbers (Table 3, 4-view: 18.52) also differ from the main table (Table 2, 4-view: 19.57) for ostensibly the same setting, which is unexplained.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"WildSplat proposes a feedforward 3D Gaussian Splatting framework for unposed in-the-wild images with varying appearances. The method uses a dual-branch architecture: a geometry branch (built on VGGT) that extracts appearance-invariant 3D structure and predicts poses, and an appearance branch that injects target appearance cues from a reference image via global pre-modulation and cross-attention. A multi-reference training paradigm and geometry-guided view sampling strategy are also introduced. The method is evaluated on MegaScenes (10 scenes) and Phototourism (4 scenes) under 4/8/12-view settings, with additional results on DL3DV and pose estimation. The paper also demonstrates appearance editing and interpolation capabilities.","tokens_in":14992,"tokens_out":2223,"duration_ms":87148,"significance":"The paper addresses a genuine gap: existing feedforward 3DGS methods assume photometric consistency and cannot handle in-the-wild appearance variations. The dual-branch decoupling of geometry and appearance, combined with the appearance injector module, is a reasonable architectural contribution. The multi-reference training paradigm is a sensible strategy for encouraging disentanglement. The appearance editing and interpolation results (Figs. 5-6) are a nice demonstration of the decoupled representation. The computational efficiency analysis (Supplementary Fig. B) shows minimal overhead over AnySplat, which is a practical strength. The method also provides pose estimation results (Table 4) competitive with VGGT.","major_comments":[{"comment":"§4.1, Evaluation Protocol / Tables 1-2: The evaluation creates an information asymmetry that is load-bearing for the central claim of 'surpassing existing feedforward methods.' WildSplat receives the left half of each ground-truth target image as an appearance reference, while feedforward baselines (AnySplat, WorldMirror) have no reference-image input mechanism. This means WildSplat has access to the target appearance (same viewpoint, same image, spatially offset) while baselines do not. The +4.19 PSNR gap over WorldMirror on Phototourism 4-view (Table 2) could be substantially explained by this asymmetry rather than by superior geometry or disentanglement. The paper should either (a) provide feedforward baselines with some form of appearance conditioning (e.g., post-hoc color affine alignment to the reference) to make the comparison more equitable, (b) report appearance-agnostic metrics","section":null},{"comment":"Table 3 vs Table 2, Phototourism 4-view: The ablation 'Full' model reports 18.52 PSNR (Table 3), while the main results table reports 19.57 PSNR (Table 2) for ostensibly the same setting (Phototourism, 4 views). This ~1 PSNR discrepancy is unexplained and raises questions about which configuration corresponds to the reported main results. The authors should clarify whether these are different checkpoints, different scene subsets, or different evaluation protocols.","section":null},{"comment":"§4.3, Table 3: The ablation study is conducted only on Phototourism (4 test scenes). Given that the paper evaluates on 14 scenes total across two datasets, restricting the ablation to 4 scenes makes it difficult to assess whether the component contributions (especially multi-reference supervision, which shows the largest drop) generalize across diverse scene types. Extending the ablation to include MegaScenes scenes would strengthen the claims about each component's contribution.","section":null}],"minor_comments":[{"comment":"§3.2, Eq. (5): The notation switches between T_i^{(k)} (content features from geometry branch) and F_content (defined in Eq. 2 as a set). It would help to clarify whether F_content and {T_i^{(k)}} refer to the same quantities or different aggregations.","section":null},{"comment":"§3.3, Algorithm 1: The algorithm references matrices S and R but their precise construction from SfM points is only briefly described in prose. A more formal definition (e.g., how the overlap matrix is computed) would improve reproducibility.","section":null},{"comment":"Table 4: The pose evaluation uses 20 scenes from MegaScenes with 24 views each, but the main NVS evaluation uses 10 MegaScenes scenes. It is unclear whether these are the same or different scene sets.","section":null},{"comment":"Supplementary A.2: The loss weights are listed as λ_i = {10, 1.0, 0.05, 0.05} in the main text (§3.3) but as λ1=10.0, λ2=10.0, λ3=1.0, λ4=0.05 in the supplementary. These appear inconsistent and should be reconciled.","section":null},{"comment":"Fig. 2: The diagram is somewhat dense. Labeling the data flow between the geometry and appearance branches more explicitly (e.g., indicating where F_content is shared) would aid readability.","section":null},{"comment":"§4.1: The paper mentions using an 'off-the-shelf video relighting model [18]' (TC-Light) to generate synthetic multi-illumination training data. The sensitivity of the final results to the quality of this synthetic data is not discussed. A brief comment on this would be valuable.","section":null}],"recommendation":"major_revision","confidential_remarks":"The information asymmetry concern is the primary issue. The paper's framing as 'appearance-conditioned novel view synthesis' is a legitimate problem formulation, but the headline comparisons against feedforward baselines that lack appearance conditioning are not apples-to-apples. The authors need to either acknowledge this limitation clearly and reframe the comparison, or provide a more equitable baseline. The ablation discrepancy (18.52 vs 19.57) also needs a straightforward explanation. If these are addressed satisfactorily, the paper could be a solid contribution."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive review. The referee raises three major points: (1) an information asymmetry in the evaluation protocol where WildSplat receives a reference image that baselines do not, (2) an unexplained ~1 PSNR discrepancy between Tables 2 and 3 for the same setting, and (3) the ablation being restricted to only 4 Phototourism scenes. We address each below. We agree with points (2) and (3) and will revise accordingly. On point (1), we agree the concern is legitimate and will add an ablated baseline with post-hoc color alignment, while also explaining why the core architectural comparison remains meaningful.","responses":[{"response":"We agree that the evaluation protocol creates an asymmetry that should be addressed more carefully, and we appreciate the referee raising this point. We will incorporate the suggested remedy in the revised manuscript. Specifically, we will add a baseline variant in which AnySplat and WorldMirror outputs are post-hoc color-aligned to the reference image half via a per-image affine color transformation (following the appearance modeling approach used by WildGaussians and NeRF-W). This will provide a more equitable comparison where baselines also have access to target appearance cues. We will report these results alongside the existing tables. That said, we wish to clarify why we believe the core comparison remains informative even without this alignment. The fundamental problem WildSplat addresses is that existing feedforward methods produce entangled geometry-appearance representations: when input views have inconsistent lighting, the rendered novel views exhibit mixed, inconsistent appearances that cannot be resolved by post-hoc correction alone. A global color affine can shift the overall tone but cannot fix spatially inconsistent appearance (e.g., one region reflecting one input's lighting and another region reflecting a different input's lighting). This is visible in Figures 3-4, where AnySplat and WorldMirror produce buildings with inconsistent colors across different regions — a problem that a global affine transform cannot resolve. WildSplat's appearance injector, by contrast, predicts per-Gaussian color attributes conditioned on the reference, which is a fundamentally different capability. The post-hoc alignment baseline will help quantify how much of the gap is due to global tone matching versus genuine appearance disentanglement, and we expect it to show that a","revision_made":"no","referee_comment":"§4.1, Evaluation Protocol / Tables 1-2: Information asymmetry — WildSplat receives the left half of the ground-truth target image as an appearance reference while feedforward baselines (AnySplat, WorldMirror) have no reference-image input mechanism. The +4.19 PSNR gap over WorldMirror could be substantially explained by this asymmetry. The paper should either (a) provide baselines with appearance conditioning (e.g., post-hoc color affine alignment to the reference), or (b) report appearance-agnostic metrics."}],"tokens_in":14622,"tokens_out":593,"duration_ms":137271,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The core idea — a feedforward 3DGS pipeline that handles unposed, photometrically inconsistent images and produces appearance-conditioned novel views in a single forward pass — is genuinely new. Nobody else has done this. The dual-branch architecture (geometry branch on VGGT-like backbone, appearance branch with cross-attention injection from a DINOv2-encoded reference) is clean and well-motivated. The multi-reference training paradigm, where a single geometry is rendered under multiple appearance conditions per iteration, is a sensible way to force disentanglement. The geometry-guided view sampling using SfM overlap matrices is a practical contribution for in-the-wild data. Ablations support each component, and the supplementary DL3DV results show the method also helps on consistent-appearance scenes. Efficiency numbers are reasonable — minimal overhead over AnySplat. Credit where earned: the system design is sound and the problem is real. Now the soft spots. The stress-test concern about the evaluation protocol is the main issue, and it lands. WildSplat receives a reference image (left half of the GT) that specifies target appearance; AnySplat and WorldMirror receive no such reference because they have no mechanism for it. So the +4.19 PSNR gap over WorldMirror on Phototourism 4-view is not apples-to-apples — WildSplat gets appearance information the baselines cannot use. The paper frames this as baselines producing “inconsistent appearances,” but that is expected when one method gets target conditioning and the others do not. The optimization-based comparisons (WildGaussian) are fairer since those methods also use a reference. The half-image protocol is also easier than cross-view appearance transfer: the reference comes from the same image as the target, just spatially offset, so the model sees the exact target appearance. The reader's concern about 14 test scenes is valid but secondary — the protocol issue matters more than sample size. The ablation table (Table 3, 4-view: 18.52) vs. main table (Table 2, 4-view: 19.57) discrepancy is unexplained and should be addressed. No code or data released. This paper is for researchers in feedforward 3D reconstruction and in-the-wild NVS. The methodological contribution is real and defensible regardless of the exact magnitude of reported gains. It deserves a serious referee who should push for: (1) a fair comparison protocol where baselines also receive appearance guidance or the asymmetry is explicitly quantified, (2) more test scenes, (3) code release, and (4) an explanation of the ablation/main table discrepancy.","headline":"First feedforward 3DGS for unposed in-the-wild images with appearance conditioning; evaluation protocol creates an information asymmetry that inflates the headline gaps over feedforward baselines.","tokens_in":15801,"tokens_out":619,"would_cite":true,"duration_ms":106912,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"One-pass 3D scene reconstruction from mismatched internet photos","keywords":["novel view synthesis","3D Gaussian Splatting","feedforward reconstruction","in-the-wild images","appearance conditioning","geometry-appearance decoupling","cross-attention","unposed images"],"falsifier":"If the geometry branch's features are not truly illumination-invariant — for instance, if feeding the same scene under drastically different lighting conditions (e.g., day vs night) produces measurably different 3D Gaussian positions or poses — then the decoupling claim fails and the appearance conditioning is merely masking entanglement rather than resolving it.","tokens_in":14988,"feed_emoji":"🖼️","tokens_out":1205,"duration_ms":88339,"temperature":0.7,"pith_summary":"WildSplat is the first feedforward 3D Gaussian Splatting framework that takes a set of unposed, in-the-wild images with varying illumination and appearance, and produces appearance-conditioned novel views in a single forward pass. The central mechanism is a dual-branch architecture that explicitly decouples geometry from appearance: a geometry branch extracts illumination-invariant 3D structure and jointly predicts camera poses, while an appearance branch injects target appearance cues from a reference image into the content features via a globally pre-modulated cross-attention mechanism. A joint multi-reference training strategy renders the same geometry under multiple appearance conditions per training iteration to prevent feature entanglement and stabilize optimization. The paper claims state-of-the-art results on Phototourism and MegaScenes, outperforming both optimization-based and feedforward methods in novel view synthesis and appearance editing from sparse inputs.","feed_headline":"One-pass 3D scene reconstruction from mismatched internet photos","feed_subtitle":"A dual-branch network separates geometry from appearance, letting a single forward pass rebuild scenes from photos taken under different照明 —","key_machinery":"Dual-branch architecture with (1) a geometry branch built on a VGGT-like transformer backbone using DINOv2 encoding that predicts appearance-invariant 3D Gaussian attributes (positions, rotations, scales, opacities) and camera poses, and (2) an appearance branch with an Appearance Injector module that uses AdaLN-Zero global pre-modulation (scale, shift, gating from the reference image's CLS token) followed by cross-attention (content features as queries, reference appearance tokens as keys/values) and self-attention to predict conditioned spherical harmonics color coefficients. A joint multi-reference training strategy renders the same geometry under M randomly sampled reference appearances,","core_discovery":"The key discovery is that explicit geometry-appearance decoupling via a dual-branch architecture with global pre-modulated cross-attention enables a single feedforward pass to reconstruct 3D scenes from photometrically inconsistent images, a setting where prior feedforward methods fail because they entangle geometry and appearance in a shared representation. The geometry branch produces structure invariant to lighting changes, and the appearance branch conditions color attributes on a reference image independently. The multi-reference training paradigm, which supervises multiple appearance renderings of the same geometry within a single iteration, proves the most critical component: removing","pith_inferences":["If the geometry branch truly produces illumination-invariant features, then feeding it images with aggressive color jittering should not degrade pose estimation or geometry quality — this is partially supported by the training-time color jittering, but a systematic evaluation of geometry quality under controlled illumination changes would be a strong test.","The reported failure on localized spatially varying illumination (e.g., projected light patterns) suggests the global pre-modulation captures overall color tone but not spatially varying lighting, implying that a local or spatially-conditioned appearance model would be needed for scenes with cast shadows or spotlights.","The competitive computational overhead (1.3s vs 1.1s for AnySplat at 16 views) suggests the dual-branch design adds minimal cost, but the memory scaling with input resolution noted in limitations could become a practical barrier for consumer-grade deployment on high-resolution imagery.","The small test set (14 scenes total) and the specific left-half/right-half reference evaluation protocol mean the reported gains, while consistent across both datasets, may not reflect performance on scenes with fundamentally different appearance variation patterns (e.g., seasonal changes, night-to-day transitions)."],"forward_implications":["Feedforward 3D reconstruction systems can now operate directly on internet photo collections without requiring precomputed camera poses or photometric consistency, removing two major bottlenecks for practical scene reconstruction at scale.","The geometry-appearance decoupling enables appearance editing as a free byproduct: the same reconstructed geometry can be re-rendered under different reference appearances without re-running the geometry branch, supporting interactive relighting and style transfer applications.","The multi-reference training paradigm, which forces geometry to remain invariant while appearance varies, suggests a general principle for training disentangled representations that could extend beyond 3D reconstruction to other multi-modal or multi-condition settings.","The approach could be extended to dynamic scenes or temporally varying appearances (e.g., construction sites photographed over months) where both geometry and appearance change, by incorporating temporal conditioning into the appearance branch."],"fun_headline_variants":["Separating geometry from appearance rescues 3D reconstruction from mismatched photos","Dual-branch feedforward splatting handles unposed in-the-wild images with varying lighting","Single-pass 3D Gaussian splatting from photos taken under inconsistent illumination","Decoupling geometry and appearance enables in-the-wild 3D scene reconstruction in one pass","Feedforward 3D Gaussian splatting survives lighting variation via dual-branch architecture"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The evaluation rests on only 14 test scenes total (4 from Phototourism, 10 from MegaScenes) with a specific protocol that splits each ground-truth image in half, and the ablation study is conducted on only the Phototourism subset, making it hard to assess whether the reported gains and component contributions generalize across more diverse scene types and appearance variation patterns.","fun_headline_variants_meta":{"raw":{"variants":["Separating geometry from appearance rescues 3D reconstruction from mismatched photos","Dual-branch feedforward splatting handles unposed in-the-wild images with varying lighting","Single-pass 3D Gaussian splatting from photos taken under inconsistent illumination","Decoupling geometry and appearance enables in-the-wild 3D scene reconstruction in one pass","Feedforward 3D Gaussian splatting survives lighting variation via dual-branch architecture"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":631,"prompt_tokens":521,"completion_tokens":110,"prompt_tokens_details":null},"tokens_in":521,"tokens_out":110,"duration_ms":8737,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-07T15:54:31.122914+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the geometry branch's features are not truly illumination-invariant — for instance, if feeding the same scene under drastically different lighting conditions (e.g., day vs night) produces measurably different 3D Gaussian positions or poses — then the decoupling claim fails and the appearance conditioning is merely masking entanglement rather than resolving it.","supporting_citations":[],"review_version":1}