{"id":"aaeb02ea-cb8d-45ee-879a-fdf0d8d4a10b","arxiv_id":"2603.14965","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Feature-space Gaussian Splat Feature Adapter (GS-Adapter) grounds camera-controlled video diffusion in 3D Gaussians, improving geometric consistency and controllability over SEVA and CameraCtrl without retraining geometry models.","lead":"GeoNVS improves novel-view synthesis by injecting 3D Gaussian geometry into video diffusion models in feature space, not as noisy rendered images. The GS-Adapter lifts diffusion features into Gaussians, re-renders them for target cameras, and adaptively fuses them to cut pose error and geometric drift.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged prior-reliability and fairness caveats.","rationale":"The reader’s weakest assumption (reliability of sparse-view 3D-GS after soft lifting, especially far from observations) is exactly the soft spot the paper itself flags and that adaptive fusion is meant to mitigate. Tables 4–7 and the gate visualization (Fig. 12 / Sec. C.3) already show that residual adaptive fusion outperforms both naïve fusion and input-level RGB injection when priors are imperfect, and that PSNRU gains appear even in non-co-visible regions. Best-prior-per-dataset pairing and lack of released artifacts remain legitimate CONDITIONAL reasons, but they do not introduce a new load-bearing technical failure. Therefore the verdict stays CONDITIONAL with high confidence; no adjustment is warranted.","tokens_in":31278,"tokens_out":559,"duration_ms":65942,"concrete_test":"Re-run the small-viewpoint average of Tab. 1 and the Terr/CD columns of Tab. 3b/4 using one fixed geometry prior (e.g., VGGT+InstantSplat) for every combined baseline (GeoNVS, Difix3D, GenFusion, input-level injection) instead of best-prior-per-dataset; if GeoNVS’s relative gains over SEVA and over input-level injection remain >8% PSNR and >1.5× Terr/CD reduction, the fairness concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that soft-lifted 3D-GS features fused adaptively in feature space (Eqs. 4–9, residual + Tanh gate) improve geometric consistency and camera controllability over pure diffusion and over input-level RGB injection—is supported by multi-backbone results (SEVA, CameraCtrl), zero-shot prior swaps (Tab. 5–6), ablations of L_feat / GS-PE / multi-scale / naïve vs adaptive fusion (Tab. 7), and geometry metrics (Terr, Rerr, CD, PSNRU on non-co-visible regions). The soft-assignment uplifting (Eq. 4) and residual adaptive fusion are designed precisely so that unreliable geometry can be down-weighted rather than forced. The paper already documents the residual risk (degradation far from inputs, thin structures, occlusions; Sec. 5, Fig. 26) and the reader correctly treats best-prior-per-dataset pairing and missing code as CONDITIONAL caveats rather than fatal flaws. No additional internal inconsistency or hidden assumption that would overturn the empirical claim was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"GeoNVS couples feed-forward 3D Gaussian geometry priors with camera-controlled video diffusion models for sparse-view novel view synthesis. The core module, GS-Adapter, (1) soft-lifts reference-view diffusion features into 3D Gaussians via rendering-weight-weighted averaging (Eq. 4), (2) rasterizes geometry-constrained novel-view features and refines them with Gaussian positional encoding plus a lightweight ResNet (Eqs. 5–6), and (3) adaptively fuses the refined features with novel-view diffusion features via residual cross-attention and a Tanh gate (Eqs. 8–9). The design is claimed to avoid view-dependent color noise that plagues input-level RGB injection, to improve geometric consistency and camera controllability, and to support zero-shot plug-and-play with multiple geometry models (MVSplat, DepthSplat, VGGT, Pi3) and two diffusion backbones (SEVA, CameraCtrl). Training uses LoRA on the frozen diffusion model plus a cosine feature-alignment loss on reference views. Evaluation spans 9 scenes / 18 settings (small- and large-overlap set NVS, long-trajectory NVS, pose-free ViPE), reporting PSNR/SSIM/LPIPS gains and large reductions in translation error and Chamfer Distance relative to SEVA, CameraCtrl, and prior geometry-injection baselines.","tokens_in":31513,"tokens_out":1143,"duration_ms":10805,"significance":"If the empirical claims hold under independent reimplementation, the paper offers a practical and modular advance for generative NVS: feature-space geometry modulation that is more robust than input-level RGB injection, with demonstrated transfer across geometry priors and two diffusion backbones, plus measurable gains on non-co-visible regions (PSNRU) and camera controllability. The soft-assignment lifting, residual adaptive fusion, multi-scale aggregation, and explicit ablations (Tabs. 5–7) constitute a clear engineering contribution that other groups can build on. Strengths include broad benchmark coverage, geometry metrics beyond pure image quality, and honest documentation of residual failure modes (thin structures, distant/occluded regions). The work is systems-oriented rather than theoretical; its significance rests on reproducibility and fair comparison rather than a new derivation.","major_comments":[{"comment":"§4.2 and Tab. 1 / Tab. 9: Combined methods (including GeoNVS) are paired with the best-performing feed-forward geometry prior per dataset. While Tab. 5 shows Adaptive Fusion is relatively robust under zero-shot prior swaps, the headline averages and SOTA claim still rest on this oracle pairing. A fixed-prior protocol (or full per-prior breakdown for every baseline) is needed so that gains are not partly attributable to prior selection rather than GS-Adapter itself.","section":null},{"comment":"§4.2 (input-level injection baseline) and Tab. 4 / Fig. 10: The input-level RGB injection baseline is implemented with a single strength s=0.2 (and a limited sweep only in the supplement). Given that this is the central foil for the claim that feature-space modulation is superior, a more complete strength sweep and, ideally, a stronger published input-level competitor trained under the same LoRA/data regime would make the comparison load-bearing rather than suggestive.","section":null},{"comment":"§5 Limitations and Fig. 26: The paper correctly notes degradation far from inputs, thin structures, and occlusions—the weakest assumption of the method. The main claims (especially long-trajectory and large-viewpoint gains) would be more credible if the main text quantified how often and how severely these failure modes occur (e.g., distance-stratified PSNR / CD, or fraction of frames with thin-structure artifacts) rather than leaving them as qualitative caveats.","section":null}],"minor_comments":[{"comment":"Notation: Ft_ref / Ft_tar and Gtar / ˆGtar / ˜Gtar are dense; a short notation table in the main text (beyond the supplement) would help.","section":null},{"comment":"Fig. 2 and Fig. 3: Some panel labels and the multi-scale fusion path are hard to parse at print size; higher-resolution or simplified diagrams would improve clarity.","section":null},{"comment":"Abstract / intro: “11.3% and 14.9% improvements” should state the metric (PSNR average) and the exact aggregation (which of the 18 settings) to avoid ambiguity.","section":null},{"comment":"Code and pretrained weights are not released with the manuscript; for a systems paper this is a presentation/reproducibility issue that should be addressed or clearly promised.","section":null},{"comment":"Eq. (1) scale alignment and the InstantSplat fitting step are important for VGGT/Pi3; a one-sentence sensitivity note would help readers who substitute other geometry models.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The central technical idea is sound and the empirical package is stronger than many concurrent generative-NVS papers. The main risk is fairness of the combined-method protocol (best prior per dataset) and incomplete characterization of the input-level baseline; both are fixable without new theory. I would not block on novelty relative to concurrent feature-lifting work if the authors clarify the distinction from Dr.Splat / LUDVIG-style distillation and from input-level injectors. Fit for a solid CV/graphics venue is good after the requested clarifications."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing is that GS-Adapter is not just another geometry-conditioned diffusion wrapper. They lift timestep-varying input-view diffusion features into 3D Gaussians with soft rendering weights, rasterize novel-view features, refine with GS positional encoding, and fuse via residual + Tanh gate so bad geometry can be down-weighted. That is cleaner than the usual “render RGB and inpaint” path, and the paper shows why: input-level injection often wrecks pose control while this one improves Terr, Rerr, and Chamfer Distance, including on non-co-visible regions.\n\nWhat is new is the finished module and the empirical regime—soft feed-forward lifting of changing diffusion features, multi-scale DPT-style aggregation, LoRA on two backbones (SEVA and CameraCtrl), zero-shot swaps across MVSplat/DepthSplat/VGGT/Pi3, plus pose-free ViPE. The ablations (naïve vs adaptive, Lfeat, GS-PE, multi-scale, soft vs hard assignment) actually move the needle. Tables cover small/large overlap, long trajectory, and geometry metrics, not only PSNR. The 11–15% average gains and the 2×/7× pose/CD claims line up with the reported numbers.\n\nSoft spots are real but proportionate. Geometry still degrades far from the references and on thin structures/occlusions; they say so in Sec. 5 and Fig. 26. Some combined baselines get the best prior per dataset, which is a fairness nit rather than a collapse of the claim. No code/checkpoints and no multi-seed bars are the usual venue gaps. Free parameters (Lfeat weight, LoRA rank, pruning rate, view-selection thresholds) exist but are not hidden.\n\nThis is for people building camera-controlled video diffusion or sparse-view NVS who care about multi-view consistency without per-scene optimization. Math is standard; data and citations look solid. I would bring it to reading group, cite it if I work in this lane, and send it to peer review. The central empirical claim holds; tighten prior-selection fairness and release artifacts and it is a clear accept-level systems contribution.","headline":"Solid systems paper: feature-space soft lifting of diffusion features into 3D-GS with adaptive residual fusion is a real, well-ablated step past noisy RGB injection, with honest limits far from the inputs.","tokens_in":32221,"tokens_out":567,"would_cite":true,"duration_ms":9560,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"GeoNVS improves novel-view synthesis by correcting video-diffusion features with 3D Gaussian geometry in feature space, not with noisy rendered images.","keywords":["novel view synthesis","video diffusion","3D Gaussian splatting","geometric consistency","camera control","feature adapter","sparse-view reconstruction"],"falsifier":"On large-baseline or heavily occluded scenes, measure pose error and Chamfer Distance of videos reconstructed from the outputs: if GS-Adapter no longer beats the pure diffusion baseline and a tuned input-level geometry-injection baseline on both metrics, the claim that feature-space geometry correction is superior fails.","tokens_in":32115,"feed_emoji":"🎥","tokens_out":982,"duration_ms":19915,"temperature":0.7,"pith_summary":"Camera-controlled video diffusion can invent new views of a scene from a few photos, but the results often warp geometry and miss the requested camera path. This paper argues that the fix is to ground the model in explicit 3D structure during denoising, not by feeding it rasterized RGB from a geometry prior. The proposed GS-Adapter lifts the diffusion model’s own input-view features into 3D Gaussians, re-renders those features from the target cameras, and adaptively blends them back so geometry can correct structure without locking in view-dependent color noise. The module is modular: it works zero-shot with several feed-forward geometry estimators and can ride on more than one diffusion backbone. Across nine scenes and eighteen settings the authors report higher image quality than strong generative baselines, roughly half the camera translation error, and up to seven times lower point-cloud reconstruction error, including gains in regions the input cameras never saw.","feed_headline":"Feature-space Gaussians cut novel-view camera error 2x","feed_subtitle":"An adapter re-renders diffusion features through 3D Gaussians, beating RGB injection on fidelity and control.","key_machinery":"The Gaussian Splat Feature Adapter (GS-Adapter): a three-stage module that soft-lifts input-view diffusion features onto 3D Gaussians, rasterizes and refines geometry-aware novel-view features, then gated-residual fuses them with the original diffusion features so structural signal can correct the denoising path without overwriting it when the prior is weak.","core_discovery":"The paper’s central claim is that explicit 3D guidance for generative novel-view synthesis is most effective when it modulates internal diffusion features rather than when it is injected as rendered images at the model input. Soft-lifting reference-view features onto 3D Gaussians, rasterizing geometry-constrained novel-view features, refining them, and adaptively fusing them yields better geometric consistency, camera controllability, and photorealism than pure video diffusion and than input-level fusion methods, while remaining plug-and-play across geometry priors.","pith_inferences":["If feature space is the right place to inject 3D structure, similar adapters could condition other controllable generators (object motion, lighting) without full diffusion retraining.","Soft lifting of timestep-varying features onto Gaussians may transfer to multi-view consistent editing and other tasks that need 3D-aligned internal features.","The stated failures on thin structures and distant regions suggest an uncertainty-weighted fusion schedule could further reduce residual geometric drift.","Coupling the adapter with pose estimation from the video itself could make generative novel-view synthesis practical for casual captures that lack ground-truth cameras."],"forward_implications":["Feature-space geometry injection can cut camera translation error by up to about 2× and Chamfer Distance by up to about 7× versus strong video-diffusion baselines.","One trained adapter works zero-shot with multiple feed-forward geometry models without retraining.","Geometry conditioned only on visible regions can still raise synthesis quality in non-co-visible regions by up to roughly 2 dB PSNR.","Input-level fusion of rasterized images can degrade camera controllability; feature-level modulation does not share that failure mode.","Dense Gaussian overhead can be cut about 2.2× by voxel pruning with little quality loss."],"fun_headline_variants":["GS-Adapter lifts features to 3D Gaussians for tighter NVS camera control","Feature-space Gaussians fix geometric drift in video novel-view synthesis","Soft-lifting diffusion features onto Gaussians cuts translation error 2x","Geometry via feature-space Gaussians beats RGB injection for novel views","Plug-and-play Gaussian feature adapter yields 7x lower Chamfer distance"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The method needs 3D Gaussians built from sparse reference views to still carry reliable structural signal after feature lifting and re-rendering, including for target cameras far from the inputs.","fun_headline_variants_meta":{"raw":{"variants":["GS-Adapter lifts features to 3D Gaussians for tighter NVS camera control","Feature-space Gaussians fix geometric drift in video novel-view synthesis","Soft-lifting diffusion features onto Gaussians cuts translation error 2x","Geometry via feature-space Gaussians beats RGB injection for novel views","Plug-and-play Gaussian feature adapter yields 7x lower Chamfer distance"]},"model":"grok-4.5","effort":"low","cost_usd":0.004986,"raw_usage":{"total_tokens":1413,"prompt_tokens":825,"num_sources_used":0,"completion_tokens":106,"cost_in_usd_ticks":49860000,"prompt_tokens_details":{"text_tokens":825,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":482,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":825,"tokens_out":106,"duration_ms":6439,"temperature":1.0,"reasoning_tokens":482,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T20:48:27.924404+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On large-baseline or heavily occluded scenes, measure pose error and Chamfer Distance of videos reconstructed from the outputs: if GS-Adapter no longer beats the pure diffusion baseline and a tuned input-level geometry-injection baseline on both metrics, the claim that feature-space geometry correction is superior fails.","supporting_citations":[],"review_version":1}