Pith. sign in

REVIEW 3 major objections 5 minor 94 references

GeoNVS improves novel-view synthesis by correcting video-diffusion features with 3D Gaussian geometry in feature space, not with noisy rendered images.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 20:48 UTC pith:LX3327OA

load-bearing objection Solid systems paper: feature-space soft lifting of diffusion features into 3D-GS with adaptive residual fusion is a real, well-ablated step past noisy RGB injection, with honest limits far from the inputs. the 3 major comments →

arxiv 2603.14965 v2 pith:LX3327OA submitted 2026-03-16 cs.CV

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

classification cs.CV
keywords novel view synthesisvideo diffusion3D Gaussian splattinggeometric consistencycamera controlfeature adaptersparse-view reconstruction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Camera-controlled video diffusion can invent new views of a scene from a few photos, but the results often warp geometry and miss the requested camera path. This paper argues that the fix is to ground the model in explicit 3D structure during denoising, not by feeding it rasterized RGB from a geometry prior. The proposed GS-Adapter lifts the diffusion model’s own input-view features into 3D Gaussians, re-renders those features from the target cameras, and adaptively blends them back so geometry can correct structure without locking in view-dependent color noise. The module is modular: it works zero-shot with several feed-forward geometry estimators and can ride on more than one diffusion backbone. Across nine scenes and eighteen settings the authors report higher image quality than strong generative baselines, roughly half the camera translation error, and up to seven times lower point-cloud reconstruction error, including gains in regions the input cameras never saw.

Core claim

The paper’s central claim is that explicit 3D guidance for generative novel-view synthesis is most effective when it modulates internal diffusion features rather than when it is injected as rendered images at the model input. Soft-lifting reference-view features onto 3D Gaussians, rasterizing geometry-constrained novel-view features, refining them, and adaptively fusing them yields better geometric consistency, camera controllability, and photorealism than pure video diffusion and than input-level fusion methods, while remaining plug-and-play across geometry priors.

What carries the argument

The Gaussian Splat Feature Adapter (GS-Adapter): a three-stage module that soft-lifts input-view diffusion features onto 3D Gaussians, rasterizes and refines geometry-aware novel-view features, then gated-residual fuses them with the original diffusion features so structural signal can correct the denoising path without overwriting it when the prior is weak.

Load-bearing premise

The method needs 3D Gaussians built from sparse reference views to still carry reliable structural signal after feature lifting and re-rendering, including for target cameras far from the inputs.

What would settle it

On large-baseline or heavily occluded scenes, measure pose error and Chamfer Distance of videos reconstructed from the outputs: if GS-Adapter no longer beats the pure diffusion baseline and a tuned input-level geometry-injection baseline on both metrics, the claim that feature-space geometry correction is superior fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Feature-space geometry injection can cut camera translation error by up to about 2× and Chamfer Distance by up to about 7× versus strong video-diffusion baselines.
  • One trained adapter works zero-shot with multiple feed-forward geometry models without retraining.
  • Geometry conditioned only on visible regions can still raise synthesis quality in non-co-visible regions by up to roughly 2 dB PSNR.
  • Input-level fusion of rasterized images can degrade camera controllability; feature-level modulation does not share that failure mode.
  • Dense Gaussian overhead can be cut about 2.2× by voxel pruning with little quality loss.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If feature space is the right place to inject 3D structure, similar adapters could condition other controllable generators (object motion, lighting) without full diffusion retraining.
  • Soft lifting of timestep-varying features onto Gaussians may transfer to multi-view consistent editing and other tasks that need 3D-aligned internal features.
  • The stated failures on thin structures and distant regions suggest an uncertainty-weighted fusion schedule could further reduce residual geometric drift.
  • Coupling the adapter with pose estimation from the video itself could make generative novel-view synthesis practical for casual captures that lack ground-truth cameras.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. GeoNVS couples feed-forward 3D Gaussian geometry priors with camera-controlled video diffusion models for sparse-view novel view synthesis. The core module, GS-Adapter, (1) soft-lifts reference-view diffusion features into 3D Gaussians via rendering-weight-weighted averaging (Eq. 4), (2) rasterizes geometry-constrained novel-view features and refines them with Gaussian positional encoding plus a lightweight ResNet (Eqs. 5–6), and (3) adaptively fuses the refined features with novel-view diffusion features via residual cross-attention and a Tanh gate (Eqs. 8–9). The design is claimed to avoid view-dependent color noise that plagues input-level RGB injection, to improve geometric consistency and camera controllability, and to support zero-shot plug-and-play with multiple geometry models (MVSplat, DepthSplat, VGGT, Pi3) and two diffusion backbones (SEVA, CameraCtrl). Training uses LoRA on the frozen diffusion model plus a cosine feature-alignment loss on reference views. Evaluation spans 9 scenes / 18 settings (small- and large-overlap set NVS, long-trajectory NVS, pose-free ViPE), reporting PSNR/SSIM/LPIPS gains and large reductions in translation error and Chamfer Distance relative to SEVA, CameraCtrl, and prior geometry-injection baselines.

Significance. If the empirical claims hold under independent reimplementation, the paper offers a practical and modular advance for generative NVS: feature-space geometry modulation that is more robust than input-level RGB injection, with demonstrated transfer across geometry priors and two diffusion backbones, plus measurable gains on non-co-visible regions (PSNRU) and camera controllability. The soft-assignment lifting, residual adaptive fusion, multi-scale aggregation, and explicit ablations (Tabs. 5–7) constitute a clear engineering contribution that other groups can build on. Strengths include broad benchmark coverage, geometry metrics beyond pure image quality, and honest documentation of residual failure modes (thin structures, distant/occluded regions). The work is systems-oriented rather than theoretical; its significance rests on reproducibility and fair comparison rather than a new derivation.

major comments (3)
  1. §4.2 and Tab. 1 / Tab. 9: Combined methods (including GeoNVS) are paired with the best-performing feed-forward geometry prior per dataset. While Tab. 5 shows Adaptive Fusion is relatively robust under zero-shot prior swaps, the headline averages and SOTA claim still rest on this oracle pairing. A fixed-prior protocol (or full per-prior breakdown for every baseline) is needed so that gains are not partly attributable to prior selection rather than GS-Adapter itself.
  2. §4.2 (input-level injection baseline) and Tab. 4 / Fig. 10: The input-level RGB injection baseline is implemented with a single strength s=0.2 (and a limited sweep only in the supplement). Given that this is the central foil for the claim that feature-space modulation is superior, a more complete strength sweep and, ideally, a stronger published input-level competitor trained under the same LoRA/data regime would make the comparison load-bearing rather than suggestive.
  3. §5 Limitations and Fig. 26: The paper correctly notes degradation far from inputs, thin structures, and occlusions—the weakest assumption of the method. The main claims (especially long-trajectory and large-viewpoint gains) would be more credible if the main text quantified how often and how severely these failure modes occur (e.g., distance-stratified PSNR / CD, or fraction of frames with thin-structure artifacts) rather than leaving them as qualitative caveats.
minor comments (5)
  1. Notation: Ft_ref / Ft_tar and Gtar / ˆGtar / ˜Gtar are dense; a short notation table in the main text (beyond the supplement) would help.
  2. Fig. 2 and Fig. 3: Some panel labels and the multi-scale fusion path are hard to parse at print size; higher-resolution or simplified diagrams would improve clarity.
  3. Abstract / intro: “11.3% and 14.9% improvements” should state the metric (PSNR average) and the exact aggregation (which of the 18 settings) to avoid ambiguity.
  4. Code and pretrained weights are not released with the manuscript; for a systems paper this is a presentation/reproducibility issue that should be addressed or clearly promised.
  5. Eq. (1) scale alignment and the InstantSplat fitting step are important for VGGT/Pi3; a one-sentence sensitivity note would help readers who substitute other geometry models.

Circularity Check

0 steps flagged

No significant circularity: empirical systems paper whose claims rest on held-out benchmarks, external baselines, and non-forced residual adaptive fusion rather than definitional or fitted identities.

full rationale

GeoNVS is a methods/systems paper, not a first-principles derivation. The load-bearing claim is that GS-Adapter (soft feature lifting into 3D-GS via Eq. 4, refinement with GS-PE + L_feat, residual adaptive fusion with Tanh gate in Eqs. 8–9) improves geometric consistency and camera controllability over pure video diffusion and over input-level RGB injection. Training uses standard latent diffusion loss plus a cosine feature-alignment loss only on reference views (Eq. 10); evaluation is on held-out public splits (9 scenes, 18 settings) against external baselines (SEVA, CameraCtrl, Difix3D, GenFusion, feed-forward geometry models). Adaptive fusion is residual and gated, so unreliable geometry is not forced into the output by construction. Zero-shot prior swaps (Tabs. 5–6) and ablations (Tab. 7) further show the gains are not tautological. Mild reuse of SEVA/CameraCtrl as backbones is ordinary engineering, not load-bearing self-citation of an unverified uniqueness theorem. No equation reduces a claimed prediction to a fitted input or self-definition; residual risk of distant/thin/occluded regions is already acknowledged as a limitation rather than hidden. Score 0 is therefore the correct, proportionate finding.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 1 invented entities

The central claim rests on standard diffusion and 3D-GS machinery plus a small set of design hyperparameters and the domain premise that sparse-view feed-forward Gaussians carry usable structure for feature-space correction. No new physical entities are postulated; GS-Adapter is an engineering module. Free parameters are training/architecture knobs, not constants fitted to invent the headline metric.

free parameters (5)
  • Lfeat loss weight
    Fixed at 0.05 in total loss L = Llatent + 0.05 Lfeat; balances feature recovery vs generation without a reported sweep in the main text.
  • LoRA rank and alpha
    Rank/α = 16 for SEVA and 4 for CameraCtrl; capacity of the trainable adapter path is hand-chosen.
  • Input-level injection strength s
    Baseline comparison uses s=0.2 (and a sweep in supplement); controls how hard the diffusion model is pinned to rasterized RGB and affects fairness of the input-level baseline.
  • Voxel Gaussian pruning rate
    Up to 96.2% pruning for 2.22× speedup; runtime/quality tradeoff parameter at inference.
  • View-selection thresholds (CLIP distance, frustum IoU, Ngroup)
    Training pair construction uses τIoU=0.4, Ngroup=21, CLIP easy/hard split; shapes the training distribution.
axioms (5)
  • domain assumption 3D Gaussian Splatting alpha-compositing (Eqs. 2–3) is a valid carrier for both color and high-dimensional diffusion features under soft weighted lifting (Eq. 4).
    Core of feature lifting; assumed throughout Sec. 3.2 and validated only empirically via soft-vs-hard ablation.
  • domain assumption Feed-forward geometry models (VGGT, Pi3, DepthSplat, MVSplat) plus optional InstantSplat fitting produce Gaussians whose structure is informative enough for novel-view feature rendering after scale alignment (Eq. 1).
    Load-bearing for all geometry-guided claims; failures on thin structures and distant views are acknowledged in limitations.
  • domain assumption Standard latent video diffusion training (SEVA/CameraCtrl backbones, LoRA on attention) remains stable when multi-scale encoder features are aggregated and residual-gated with geometry features each denoising step.
    Integration premise in Sec. 3.3; supported by training success but not theoretically derived.
  • domain assumption ViPE-reconstructed poses/point clouds are a valid proxy for geometric fidelity and camera controllability of generated videos (excluding cases with Terr or Rerr > 50).
    Used for Terr/Rerr/CD metrics in Sec. 4.4; evaluation quality depends on this external estimator.
  • standard math Linear algebra / rendering math of Gaussian projection and residual fusion (standard).
    Eqs. 1–9 use ordinary least squares scale fit, alpha blending, cross-attention, and residual addition.
invented entities (1)
  • GS-Adapter (lift–refine–adaptive-fuse pipeline with GS-PE and RefineNet) no independent evidence
    purpose: Bridge 3D Gaussians and video diffusion in feature space to correct geometrically inconsistent novel-view features without RGB-level noise.
    Primary architectural invention; independent_evidence is empirical (ablations, multi-backbone transfer), not an external physical prediction.

pith-pipeline@v1.1.0-grok45 · 35281 in / 3903 out tokens · 36103 ms · 2026-07-14T20:48:27.924404+00:00 · methodology

0 comments
read the original abstract

Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, they often suffer from geometric distortions and limited camera controllability. To overcome these challenges, we introduce GeoNVS, a geometry-grounded novel-view synthesizer that enhances both geometric fidelity and camera controllability through explicit 3D geometric guidance. Our key innovation is the Gaussian Splat Feature Adapter (GS-Adapter), which lifts input-view diffusion features into 3D Gaussian representations, renders geometry-constrained novel-view features, and adaptively fuses them with diffusion features to correct geometrically inconsistent representations. Unlike prior methods that inject geometry at the input level, GS-Adapter operates in feature space, avoiding view-dependent color noise that degrades structural consistency. Its plug-and-play design enables zero-shot compatibility with diverse feed-forward geometry models without additional training, and can be adapted to other video diffusion backbones. Experiments across 9 scenes and 18 settings demonstrate state-of-the-art performance, achieving 11.3% and 14.9% improvements over SEVA and CameraCtrl, with up to 2x reduction in translation error and 7x in Chamfer Distance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

94 extracted references · 18 linked inside Pith

  1. [1]

    In: ICCV

    Bai,J., Xia, M., Fu, X.,Wang, X.,Mu, L., Cao, J.,Liu, Z., Hu, H.,Bai, X., Wan, P., et al.: Recammaster: Camera-controlled generative rendering from a single video. In: ICCV. pp. 14834–14844 (2025)

  2. [2]

    In: CVPR

    Barron, J.T., Mildenhall, B., Verbin, D., Srinivasan, P.P., Hedman, P.: Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In: CVPR. pp. 5470–5479 (2022)

  3. [3]

    arXiv preprint arXiv:2311.15127 (2023)

    Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al.: Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127 (2023)

  4. [4]

    In: CVPR

    Cao, C., Yu, C., Liu, S., Wang, F., Xue, X., Fu, Y.: Mvgenmaster: Scaling multi- view generation from any image via 3d priors enhanced diffusion model. In: CVPR. pp. 6045–6056 (2025)

  5. [5]

    In: ICCV

    Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: ICCV. pp. 9650– 9660 (2021)

  6. [6]

    In: ICCV

    Chan,E.R.,Nagano,K.,Chan,M.A.,Bergman,A.W.,Park,J.J.,Levy,A.,Aittala, M., De Mello, S., Karras, T., Wetzstein, G.: Generative novel view synthesis with 3d-aware diffusion models. In: ICCV. pp. 4217–4229 (2023)

  7. [7]

    In: CVPR (2024)

    Charatan, D., Li, S., Tagliasacchi, A., Sitzmann, V.: pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In: CVPR (2024)

  8. [8]

    Charatan,D.,Li,S.L.,Tagliasacchi,A.,Sitzmann,V.:pixelsplat:3dgaussiansplats fromimagepairsforscalablegeneralizable3dreconstruction.In:CVPR.pp.19457– 19467 (2024)

  9. [9]

    IEEE transactions on pattern analysis and machine intelli- gence40(4), 834–848 (2017)

    Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: Deeplab: Se- mantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelli- gence40(4), 834–848 (2017)

  10. [10]

    In: ECCV

    Chen, Y., Xu, H., Zheng, C., Zhuang, B., Pollefeys, M., Geiger, A., Cham, T.J., Cai, J.: Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. In: ECCV. pp. 370–386. Springer (2024)

  11. [11]

    In: NeurIPS (2024)

    Chen, Y., Zheng, C., Xu, H., Zhuang, B., Vedaldi, A., Cham, T.J., Cai, J.: Mvs- plat360: Feed-forward 360 scene synthesis from sparse views. In: NeurIPS (2024)

  12. [12]

    In: CVPR

    Chung, J., Oh, J., Lee, K.M.: Depth-regularized optimization for 3d gaussian splat- ting in few-shot images. In: CVPR. pp. 811–820 (2024)

  13. [13]

    In: CVPR

    Deng, K., Liu, A., Zhu, J.Y., Ramanan, D.: Depth-supervised nerf: Fewer views and faster training for free. In: CVPR. pp. 12882–12891 (2022)

  14. [14]

    In: CVPR

    Duzceker, A., Galliani, S., Vogel, C., Speciale, P., Dusmanu, M., Pollefeys, M.: Deepvideomvs: Multi-view stereo on video with recurrent spatio-temporal fusion. In: CVPR. pp. 15324–15333 (2021)

  15. [15]

    In: ICCV

    Eftekhar, A., Sax, A., Malik, J., Zamir, A.: Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans. In: ICCV. pp. 10786– 10796 (2021)

  16. [16]

    In: CVPR

    Fan, H., Su, H., Guibas, L.J.: A point set generation network for 3d object recon- struction from a single image. In: CVPR. pp. 605–613 (2017)

  17. [17]

    arXiv preprint arXiv:2403.20309 (2024)

    Fan, Z., Wen, K., Cong, W., Wang, K., Zhang, J., Ding, X., Xu, D., Ivanovic, B., Pavone, M., Pavlakos, G., et al.: Instantsplat: Sparse-view gaussian splatting in seconds. arXiv preprint arXiv:2403.20309 (2024)

  18. [18]

    In: CVPR

    Fu, Y., Liu, S., Kulkarni, A., Kautz, J., Efros, A.A., Wang, X.: Colmap-free 3d gaussian splatting. In: CVPR. pp. 20796–20805 (2024) GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis 37

  19. [19]

    NeurIPS (2024)

    Gao*, R., Holynski*, A., Henzler, P., Brussee, A., Martin-Brualla, R., Srinivasan, P.P., Barron, J.T., Poole*, B.: Cat3d: Create anything in 3d with multi-view dif- fusion models. NeurIPS (2024)

  20. [20]

    In: CVPR

    Geng, D., Herrmann, C., Hur, J., Cole, F., Zhang, S., Pfaff, T., Lopez-Guevara, T., Aytar, Y., Rubinstein, M., Sun, C., et al.: Motion prompting: Controlling video generation with motion trajectories. In: CVPR. pp. 1–12 (2025)

  21. [21]

    In: CVPR

    Gu, X., Fan, Z., Zhu, S., Dai, Z., Tan, F., Tan, P.: Cascade cost volume for high- resolution multi-view stereo and stereo matching. In: CVPR. pp. 2495–2504 (2020)

  22. [22]

    In: Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers

    Gu, Z., Yan, R., Lu, J., Li, P., Dou, Z., Si, C., Dong, Z., Liu, Q., Lin, C., Liu, Z., et al.: Diffusion as shader: 3d-aware video diffusion for versatile video generation control. In: Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers. pp. 1–12 (2025)

  23. [23]

    In: CVPR

    Hanson, A., Tu, A., Singla, V., Jayawardhana, M., Zwicker, M., Goldstein, T.: Pup 3d-gs: Principled uncertainty pruning for 3d gaussian splatting. In: CVPR. pp. 5949–5958 (2025)

  24. [24]

    In: ICLR (2025)

    He, H., Xu, Y., Guo, Y., Wetzstein, G., Dai, B., Li, H., Yang, C.: Cameractrl: Enabling camera control for video diffusion models. In: ICLR (2025)

  25. [25]

    arXiv preprint arXiv:2503.10592 (2025)

    He, H., Yang, C., Lin, S., Xu, Y., Wei, M., Gui, L., Zhao, Q., Wetzstein, G., Jiang, L., Li, H.: Cameractrl ii: Dynamic scene exploration via camera-controlled video diffusion models. arXiv preprint arXiv:2503.10592 (2025)

  26. [26]

    ICLR1(2), 3 (2022)

    Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al.: Lora: Low-rank adaptation of large language models. ICLR1(2), 3 (2022)

  27. [27]

    arXiv preprint arXiv:2508.10934 (2025)

    Huang, J., Zhou, Q., Rabeti, H., Korovko, A., Ling, H., Ren, X., Shen, T., Gao, J., Slepichev,D.,Lin,C.H.,etal.:Vipe:Videoposeenginefor3dgeometricperception. arXiv preprint arXiv:2508.10934 (2025)

  28. [28]

    arXiv preprint arXiv:2601.09575 (2026)

    Huang, S.Y., Choe, J., Wang, Y.C.F., Sun, C.: Openvoxel: Training-free grouping and captioning voxels for open-vocabulary 3d scene understanding. arXiv preprint arXiv:2601.09575 (2026)

  29. [29]

    arXiv preprint arXiv:1905.00538 (2019)

    Im, S., Jeon, H.G., Lin, S., Kweon, I.S.: Dpsnet: End-to-end deep plane sweep stereo. arXiv preprint arXiv:1905.00538 (2019)

  30. [30]

    In: CVPR

    Jensen, R., Dahl, A., Vogiatzis, G., Tola, E., Aanæs, H.: Large scale multi-view stereopsis evaluation. In: CVPR. pp. 406–413. IEEE (2014)

  31. [31]

    arXiv preprint arXiv:2409.02084 (2024)

    Ji, M., Qiu, R.Z., Zou, X., Wang, X.: Graspsplats: Efficient manipulation with 3d feature splatting. arXiv preprint arXiv:2409.02084 (2024)

  32. [32]

    arXiv preprint arXiv:2505.23716 (2025)

    Jiang, L., Mao, Y., Xu, L., Lu, T., Ren, K., Jin, Y., Xu, X., Yu, M., Pang, J., Zhao, F., et al.: Anysplat: Feed-forward 3d gaussian splatting from unconstrained views. arXiv preprint arXiv:2505.23716 (2025)

  33. [33]

    splat: Directly referring 3d gaussian splatting via direct language embedding registration

    Jun-Seong, K., GeonU, K., Yu-Ji, K., Wang, Y.C.F., Choe, J., Oh, T.H.: Dr. splat: Directly referring 3d gaussian splatting via direct language embedding registration. In: CVPR (2025)

  34. [34]

    In: International Con- ference on 3D Vision (3DV)

    Keetha, N., Müller, N., Schönberger, J., Porzi, L., Zhang, Y., Fischer, T., Knapitsch, A., Zauss, D., Weber, E., Antunes, N., Luiten, J., Lopez-Antequera, M., Bulò, S.R., Richardt, C., Ramanan, D., Scherer, S., Kontschieder, P.: MapA- nything: Universal feed-forward metric 3D reconstruction. In: International Con- ference on 3D Vision (3DV). IEEE (2026)

  35. [35]

    ACM Transactions on Graphics42(4) (July 2023),���������������������������������������������������������

    Kerbl, B., Kopanas, G., Leimkühler, T., Drettakis, G.: 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics42(4) (July 2023),���������������������������������������������������������

  36. [36]

    In: ICCV

    Kerr, J., Kim, C.M., Goldberg, K., Kanazawa, A., Tancik, M.: Lerf: Language embedded radiance fields. In: ICCV. pp. 19729–19739 (2023)

  37. [37]

    In: ICCV

    Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., et al.: Segment anything. In: ICCV. pp. 4015–4026 (2023) 38 Kang et al

  38. [38]

    ACM Transactions on Graphics36(4) (2017)

    Knapitsch, A., Park, J., Zhou, Q.Y., Koltun, V.: Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics36(4) (2017)

  39. [39]

    In: CVPR

    Lee, J., Park, C., Choe, J., Wang, Y.C.F., Kautz, J., Cho, M., Choy, C.: Mosaic3d: Foundation dataset and model for open-vocabulary 3d segmentation. In: CVPR. pp. 14089–14101 (2025)

  40. [40]

    In: CVPR

    Li, L., Zhang, Z., Li, Y., Xu, J., Hu, W., Li, X., Cheng, W., Gu, J., Xue, T., Shan, Y.: Nvcomposer: Boosting generative novel view synthesis with multiple sparse and unposed images. In: CVPR. pp. 777–787 (2025)

  41. [41]

    Li, M., Yang, T., Kuang, H., Wu, J., Wang, Z., Xiao, X., Chen, C.: Control- net++: Improving conditional controls with efficient consistency feedback: Project page: liming-ai. github. io/controlnet_plus_plus. In: ECCV. pp. 129–147. Springer (2024)

  42. [42]

    Advances in Neural Information Processing Systems (2025),��������������������������������

    Li, W., Zhao, Y., Qin, M., Liu, Y., Cai, Y., Gan, C., Pfister, H.: Langsplatv2: High- dimensional 3d language gaussian splatting with 450+ fps. Advances in Neural Information Processing Systems (2025),��������������������������������

  43. [43]

    arXiv preprint arXiv:2412.10231 (2024)

    Liang, S., Wang, S., Li, K., Niemeyer, M., Gasperini, S., Lensch, H., Navab, N., Tombari, F.: Supergseg: Open-vocabulary 3d segmentation with structured super- gaussians. arXiv preprint arXiv:2412.10231 (2024)

  44. [44]

    arXiv preprint arXiv:2511.10647 (2025)

    Lin, H., Chen, S., Liew, J., Chen, D.Y., Li, Z., Shi, G., Feng, J., Kang, B.: Depth anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647 (2025)

  45. [45]

    In: CVPR

    Lin, T.Y., Dollár, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: CVPR. pp. 2117–2125 (2017)

  46. [46]

    In: CVPR

    Ling, L., Sheng, Y., Tu, Z., Zhao, W., Xin, C., Wan, K., Yu, L., Guo, Q., Yu, Z., Lu, Y., et al.: Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In: CVPR. pp. 22160–22169 (2024)

  47. [47]

    arXiv preprint arXiv:2502.06782 (2025)

    Liu, D., Li, S., Liu, Y., Li, Z., Wang, K., Li, X., Qin, Q., Liu, Y., Xin, Y., Li, Z., et al.: Lumina-video: Efficient and flexible video generation with multi-scale next-dit. arXiv preprint arXiv:2502.06782 (2025)

  48. [48]

    NeurIPS37, 133305–133327 (2024)

    Liu, X., Zhou, C., Huang, S.: 3dgs-enhancer: Enhancing unbounded 3d gaussian splatting with view-consistent 2d diffusion priors. NeurIPS37, 133305–133327 (2024)

  49. [49]

    In: ICCV

    Marrie, J., Ménégaux, R., Arbel, M., Larlus, D., Mairal, J.: Ludvig: Learning-free uplifting of 2d visual features to gaussian splatting scenes. In: ICCV. pp. 7440–7450 (2025)

  50. [50]

    arXiv preprint arXiv:2108.01073 (2021)

    Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.Y., Ermon, S.: Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073 (2021)

  51. [51]

    ACM Transactions on Graphics (ToG)38(4), 1–14 (2019)

    Mildenhall, B., Srinivasan, P.P., Ortiz-Cayon, R., Kalantari, N.K., Ramamoorthi, R., Ng, R., Kar, A.: Local light field fusion: Practical view synthesis with pre- scriptive sampling guidelines. ACM Transactions on Graphics (ToG)38(4), 1–14 (2019)

  52. [52]

    Commu- nications of the ACM65(1), 99–106 (2021)

    Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: Nerf: Representing scenes as neural radiance fields for view synthesis. Commu- nications of the ACM65(1), 99–106 (2021)

  53. [53]

    arXiv preprint arXiv:2304.07193 (2023)

    Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al.: Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023)

  54. [54]

    In: CVPR

    Qin, M., Li, W., Zhou, J., Wang, H., Pfister, H.: Langsplat: 3d language gaussian splatting. In: CVPR. pp. 20051–20060 (2024)

  55. [55]

    In: ICML

    Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: ICML. pp. 8748–8763. PmLR (2021) GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis 39

  56. [56]

    In: ICCV

    Ranftl, R., Bochkovskiy, A., Koltun, V.: Vision transformers for dense prediction. In: ICCV. pp. 12179–12188 (2021)

  57. [57]

    In: CVPR

    Reizenstein,J.,Shapovalov,R.,Henzler,P.,Sbordone,L.,Labatut,P.,Novotny,D.: Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In: CVPR. pp. 10901–10911 (2021)

  58. [58]

    In: CVPR

    Ren,X.,Shen,T.,Huang,J.,Ling,H.,Lu,Y.,Nimier-David,M.,Müller,T.,Keller, A., Fidler, S., Gao, J.: Gen3c: 3d-informed world-consistent video generation with precise camera control. In: CVPR. pp. 6121–6132 (2025)

  59. [59]

    In: CVPR

    Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: CVPR. pp. 10684–10695 (2022)

  60. [60]

    In: CVPR (2016)

    Schönberger, J.L., Frahm, J.M.: Structure-from-motion revisited. In: CVPR (2016)

  61. [61]

    In: ECCV (2016)

    Schönberger, J.L., Zheng, E., Pollefeys, M., Frahm, J.M.: Pixelwise view selection for unstructured multi-view stereo. In: ECCV (2016)

  62. [62]

    In: ICCV

    Seo, S., Chang, Y., Kwak, N.: Flipnerf: Flipped reflection rays for few-shot novel view synthesis. In: ICCV. pp. 22883–22893 (2023)

  63. [63]

    In: ICLR (2025)

    Tang, S., Ye, W., Ye, P., Lin, W., Zhou, Y., Chen, T., Ouyang, W.: Hisplat: Hierarchical 3d gaussian splatting for generalizable sparse-view reconstruction. In: ICLR (2025)

  64. [64]

    In: ICCV

    Wang, G., Chen, Z., Loy, C.C., Liu, Z.: Sparsenerf: Distilling depth ranking for few-shot novel view synthesis. In: ICCV. pp. 9065–9076 (2023)

  65. [65]

    In: CVPR

    Wang, J., Chen, M., Karaev, N., Vedaldi, A., Rupprecht, C., Novotny, D.: Vggt: Visual geometry grounded transformer. In: CVPR. pp. 5294–5306 (2025)

  66. [66]

    IEEE transactions on pattern analysis and machine intelligence43(10), 3349–3364 (2020)

    Wang, J., Sun, K., Cheng, T., Jiang, B., Deng, C., Zhao, Y., Liu, D., Mu, Y., Tan, M., Wang, X., et al.: Deep high-resolution representation learning for visual recog- nition. IEEE transactions on pattern analysis and machine intelligence43(10), 3349–3364 (2020)

  67. [67]

    arXiv preprint arXiv:2511.06457 (2025)

    Wang, S., Zhang, S., Millerdurai, C., Westermann, R., Stricker, D., Pagani, A.: Inpaint360gs: Efficient object-aware 3d inpainting via gaussian splatting for 360 {\deg}scenes. arXiv preprint arXiv:2511.06457 (2025)

  68. [68]

    In: CVPR

    Wang, S., Leroy, V., Cabon, Y., Chidlovskii, B., Revaud, J.: Dust3r: Geometric 3d vision made easy. In: CVPR. pp. 20697–20709 (2024)

  69. [69]

    In: ICCV

    Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., Shao, L.: Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In: ICCV. pp. 568–578 (2021)

  70. [70]

    Wang, Y., Zhou, J., Zhu, H., Chang, W., Zhou, Y., Li, Z., Chen, J., Pang, J., Shen, C., He, T.:π3: Scalable permutation-equivariant visual geometry learning (2025), ��������������������������������

  71. [71]

    In: SIGGRAPH

    Wang,Z.,Yuan,Z.,Wang,X.,Li,Y.,Chen,T.,Xia,M.,Luo,P.,Shan,Y.:Motionc- trl: A unified and flexible motion controller for video generation. In: SIGGRAPH. pp. 1–11 (2024)

  72. [72]

    arXiv preprint arXiv:2407.07860 (2024)

    Watson, D., Saxena, S., Li, L., Tagliasacchi, A., Fleet, D.J.: Controlling space and time with diffusion models. arXiv preprint arXiv:2407.07860 (2024)

  73. [73]

    arXiv preprint arXiv:2507.07982 (2025)

    Wu, H., Wu, D., He, T., Guo, J., Ye, Y., Duan, Y., Bian, J.: Geometry forcing: Mar- rying video diffusion and 3d representation for consistent world modeling. arXiv preprint arXiv:2507.07982 (2025)

  74. [74]

    In: CVPR

    Wu, J.Z., Zhang, Y., Turki, H., Ren, X., Gao, J., Shou, M.Z., Fidler, S., Gojcic, Z., Ling, H.: Difix3d+: Improving 3d reconstructions with single-step diffusion models. In: CVPR. pp. 26024–26035 (2025)

  75. [75]

    In: CVPR

    Wu, R., Mildenhall, B., Henzler, P., Park, K., Gao, R., Watson, D., Srinivasan, P.P., Verbin, D., Barron, J.T., Poole, B., et al.: Reconfusion: 3d reconstruction with diffusion priors. In: CVPR. pp. 21551–21561 (2024) 40 Kang et al

  76. [76]

    In: CVPR

    Wu, S., Xu, C., Huang, B., Geiger, A., Chen, A.: Genfusion: Closing the loop between reconstruction and generation via videos. In: CVPR. pp. 6078–6088 (2025)

  77. [77]

    In: CVPR

    Wu, T., Zhang, J., Fu, X., Wang, Y., Ren, J., Pan, L., Wu, W., Yang, L., Wang, J., Qian, C., et al.: Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation. In: CVPR. pp. 803–814 (2023)

  78. [78]

    In: ECCV

    Wu, W., Li, Z., Gu, Y., Zhao, R., He, Y., Zhang, D.J., Shou, M.Z., Li, Y., Gao, T., Zhang, D.: Draganything: Motion control for anything using entity representation. In: ECCV. pp. 331–348. Springer (2024)

  79. [79]

    In: CVPR

    Xia, H., Fu, Y., Liu, S., Wang, X.: Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos. In: CVPR. pp. 22378–22389 (2024)

  80. [80]

    In: CVPR

    Xu, H., Peng, S., Wang, F., Blum, H., Barath, D., Geiger, A., Pollefeys, M.: Depth- splat: Connecting gaussian splatting and depth. In: CVPR. pp. 16453–16463 (2025)

Showing first 80 references.