REVIEW 3 major objections 5 minor 94 references
GeoNVS improves novel-view synthesis by correcting video-diffusion features with 3D Gaussian geometry in feature space, not with noisy rendered images.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 20:48 UTC pith:LX3327OA
load-bearing objection Solid systems paper: feature-space soft lifting of diffusion features into 3D-GS with adaptive residual fusion is a real, well-ablated step past noisy RGB injection, with honest limits far from the inputs. the 3 major comments →
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper’s central claim is that explicit 3D guidance for generative novel-view synthesis is most effective when it modulates internal diffusion features rather than when it is injected as rendered images at the model input. Soft-lifting reference-view features onto 3D Gaussians, rasterizing geometry-constrained novel-view features, refining them, and adaptively fusing them yields better geometric consistency, camera controllability, and photorealism than pure video diffusion and than input-level fusion methods, while remaining plug-and-play across geometry priors.
What carries the argument
The Gaussian Splat Feature Adapter (GS-Adapter): a three-stage module that soft-lifts input-view diffusion features onto 3D Gaussians, rasterizes and refines geometry-aware novel-view features, then gated-residual fuses them with the original diffusion features so structural signal can correct the denoising path without overwriting it when the prior is weak.
Load-bearing premise
The method needs 3D Gaussians built from sparse reference views to still carry reliable structural signal after feature lifting and re-rendering, including for target cameras far from the inputs.
What would settle it
On large-baseline or heavily occluded scenes, measure pose error and Chamfer Distance of videos reconstructed from the outputs: if GS-Adapter no longer beats the pure diffusion baseline and a tuned input-level geometry-injection baseline on both metrics, the claim that feature-space geometry correction is superior fails.
If this is right
- Feature-space geometry injection can cut camera translation error by up to about 2× and Chamfer Distance by up to about 7× versus strong video-diffusion baselines.
- One trained adapter works zero-shot with multiple feed-forward geometry models without retraining.
- Geometry conditioned only on visible regions can still raise synthesis quality in non-co-visible regions by up to roughly 2 dB PSNR.
- Input-level fusion of rasterized images can degrade camera controllability; feature-level modulation does not share that failure mode.
- Dense Gaussian overhead can be cut about 2.2× by voxel pruning with little quality loss.
Where Pith is reading between the lines
- If feature space is the right place to inject 3D structure, similar adapters could condition other controllable generators (object motion, lighting) without full diffusion retraining.
- Soft lifting of timestep-varying features onto Gaussians may transfer to multi-view consistent editing and other tasks that need 3D-aligned internal features.
- The stated failures on thin structures and distant regions suggest an uncertainty-weighted fusion schedule could further reduce residual geometric drift.
- Coupling the adapter with pose estimation from the video itself could make generative novel-view synthesis practical for casual captures that lack ground-truth cameras.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. GeoNVS couples feed-forward 3D Gaussian geometry priors with camera-controlled video diffusion models for sparse-view novel view synthesis. The core module, GS-Adapter, (1) soft-lifts reference-view diffusion features into 3D Gaussians via rendering-weight-weighted averaging (Eq. 4), (2) rasterizes geometry-constrained novel-view features and refines them with Gaussian positional encoding plus a lightweight ResNet (Eqs. 5–6), and (3) adaptively fuses the refined features with novel-view diffusion features via residual cross-attention and a Tanh gate (Eqs. 8–9). The design is claimed to avoid view-dependent color noise that plagues input-level RGB injection, to improve geometric consistency and camera controllability, and to support zero-shot plug-and-play with multiple geometry models (MVSplat, DepthSplat, VGGT, Pi3) and two diffusion backbones (SEVA, CameraCtrl). Training uses LoRA on the frozen diffusion model plus a cosine feature-alignment loss on reference views. Evaluation spans 9 scenes / 18 settings (small- and large-overlap set NVS, long-trajectory NVS, pose-free ViPE), reporting PSNR/SSIM/LPIPS gains and large reductions in translation error and Chamfer Distance relative to SEVA, CameraCtrl, and prior geometry-injection baselines.
Significance. If the empirical claims hold under independent reimplementation, the paper offers a practical and modular advance for generative NVS: feature-space geometry modulation that is more robust than input-level RGB injection, with demonstrated transfer across geometry priors and two diffusion backbones, plus measurable gains on non-co-visible regions (PSNRU) and camera controllability. The soft-assignment lifting, residual adaptive fusion, multi-scale aggregation, and explicit ablations (Tabs. 5–7) constitute a clear engineering contribution that other groups can build on. Strengths include broad benchmark coverage, geometry metrics beyond pure image quality, and honest documentation of residual failure modes (thin structures, distant/occluded regions). The work is systems-oriented rather than theoretical; its significance rests on reproducibility and fair comparison rather than a new derivation.
major comments (3)
- §4.2 and Tab. 1 / Tab. 9: Combined methods (including GeoNVS) are paired with the best-performing feed-forward geometry prior per dataset. While Tab. 5 shows Adaptive Fusion is relatively robust under zero-shot prior swaps, the headline averages and SOTA claim still rest on this oracle pairing. A fixed-prior protocol (or full per-prior breakdown for every baseline) is needed so that gains are not partly attributable to prior selection rather than GS-Adapter itself.
- §4.2 (input-level injection baseline) and Tab. 4 / Fig. 10: The input-level RGB injection baseline is implemented with a single strength s=0.2 (and a limited sweep only in the supplement). Given that this is the central foil for the claim that feature-space modulation is superior, a more complete strength sweep and, ideally, a stronger published input-level competitor trained under the same LoRA/data regime would make the comparison load-bearing rather than suggestive.
- §5 Limitations and Fig. 26: The paper correctly notes degradation far from inputs, thin structures, and occlusions—the weakest assumption of the method. The main claims (especially long-trajectory and large-viewpoint gains) would be more credible if the main text quantified how often and how severely these failure modes occur (e.g., distance-stratified PSNR / CD, or fraction of frames with thin-structure artifacts) rather than leaving them as qualitative caveats.
minor comments (5)
- Notation: Ft_ref / Ft_tar and Gtar / ˆGtar / ˜Gtar are dense; a short notation table in the main text (beyond the supplement) would help.
- Fig. 2 and Fig. 3: Some panel labels and the multi-scale fusion path are hard to parse at print size; higher-resolution or simplified diagrams would improve clarity.
- Abstract / intro: “11.3% and 14.9% improvements” should state the metric (PSNR average) and the exact aggregation (which of the 18 settings) to avoid ambiguity.
- Code and pretrained weights are not released with the manuscript; for a systems paper this is a presentation/reproducibility issue that should be addressed or clearly promised.
- Eq. (1) scale alignment and the InstantSplat fitting step are important for VGGT/Pi3; a one-sentence sensitivity note would help readers who substitute other geometry models.
Circularity Check
No significant circularity: empirical systems paper whose claims rest on held-out benchmarks, external baselines, and non-forced residual adaptive fusion rather than definitional or fitted identities.
full rationale
GeoNVS is a methods/systems paper, not a first-principles derivation. The load-bearing claim is that GS-Adapter (soft feature lifting into 3D-GS via Eq. 4, refinement with GS-PE + L_feat, residual adaptive fusion with Tanh gate in Eqs. 8–9) improves geometric consistency and camera controllability over pure video diffusion and over input-level RGB injection. Training uses standard latent diffusion loss plus a cosine feature-alignment loss only on reference views (Eq. 10); evaluation is on held-out public splits (9 scenes, 18 settings) against external baselines (SEVA, CameraCtrl, Difix3D, GenFusion, feed-forward geometry models). Adaptive fusion is residual and gated, so unreliable geometry is not forced into the output by construction. Zero-shot prior swaps (Tabs. 5–6) and ablations (Tab. 7) further show the gains are not tautological. Mild reuse of SEVA/CameraCtrl as backbones is ordinary engineering, not load-bearing self-citation of an unverified uniqueness theorem. No equation reduces a claimed prediction to a fitted input or self-definition; residual risk of distant/thin/occluded regions is already acknowledged as a limitation rather than hidden. Score 0 is therefore the correct, proportionate finding.
Axiom & Free-Parameter Ledger
free parameters (5)
- Lfeat loss weight
- LoRA rank and alpha
- Input-level injection strength s
- Voxel Gaussian pruning rate
- View-selection thresholds (CLIP distance, frustum IoU, Ngroup)
axioms (5)
- domain assumption 3D Gaussian Splatting alpha-compositing (Eqs. 2–3) is a valid carrier for both color and high-dimensional diffusion features under soft weighted lifting (Eq. 4).
- domain assumption Feed-forward geometry models (VGGT, Pi3, DepthSplat, MVSplat) plus optional InstantSplat fitting produce Gaussians whose structure is informative enough for novel-view feature rendering after scale alignment (Eq. 1).
- domain assumption Standard latent video diffusion training (SEVA/CameraCtrl backbones, LoRA on attention) remains stable when multi-scale encoder features are aggregated and residual-gated with geometry features each denoising step.
- domain assumption ViPE-reconstructed poses/point clouds are a valid proxy for geometric fidelity and camera controllability of generated videos (excluding cases with Terr or Rerr > 50).
- standard math Linear algebra / rendering math of Gaussian projection and residual fusion (standard).
invented entities (1)
-
GS-Adapter (lift–refine–adaptive-fuse pipeline with GS-PE and RefineNet)
no independent evidence
read the original abstract
Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, they often suffer from geometric distortions and limited camera controllability. To overcome these challenges, we introduce GeoNVS, a geometry-grounded novel-view synthesizer that enhances both geometric fidelity and camera controllability through explicit 3D geometric guidance. Our key innovation is the Gaussian Splat Feature Adapter (GS-Adapter), which lifts input-view diffusion features into 3D Gaussian representations, renders geometry-constrained novel-view features, and adaptively fuses them with diffusion features to correct geometrically inconsistent representations. Unlike prior methods that inject geometry at the input level, GS-Adapter operates in feature space, avoiding view-dependent color noise that degrades structural consistency. Its plug-and-play design enables zero-shot compatibility with diverse feed-forward geometry models without additional training, and can be adapted to other video diffusion backbones. Experiments across 9 scenes and 18 settings demonstrate state-of-the-art performance, achieving 11.3% and 14.9% improvements over SEVA and CameraCtrl, with up to 2x reduction in translation error and 7x in Chamfer Distance.
Reference graph
Works this paper leans on
-
[1]
In: ICCV
Bai,J., Xia, M., Fu, X.,Wang, X.,Mu, L., Cao, J.,Liu, Z., Hu, H.,Bai, X., Wan, P., et al.: Recammaster: Camera-controlled generative rendering from a single video. In: ICCV. pp. 14834–14844 (2025)
2025
-
[2]
In: CVPR
Barron, J.T., Mildenhall, B., Verbin, D., Srinivasan, P.P., Hedman, P.: Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In: CVPR. pp. 5470–5479 (2022)
2022
-
[3]
arXiv preprint arXiv:2311.15127 (2023)
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al.: Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127 (2023)
Pith/arXiv arXiv 2023
-
[4]
In: CVPR
Cao, C., Yu, C., Liu, S., Wang, F., Xue, X., Fu, Y.: Mvgenmaster: Scaling multi- view generation from any image via 3d priors enhanced diffusion model. In: CVPR. pp. 6045–6056 (2025)
2025
-
[5]
In: ICCV
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: ICCV. pp. 9650– 9660 (2021)
2021
-
[6]
In: ICCV
Chan,E.R.,Nagano,K.,Chan,M.A.,Bergman,A.W.,Park,J.J.,Levy,A.,Aittala, M., De Mello, S., Karras, T., Wetzstein, G.: Generative novel view synthesis with 3d-aware diffusion models. In: ICCV. pp. 4217–4229 (2023)
2023
-
[7]
In: CVPR (2024)
Charatan, D., Li, S., Tagliasacchi, A., Sitzmann, V.: pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In: CVPR (2024)
2024
-
[8]
Charatan,D.,Li,S.L.,Tagliasacchi,A.,Sitzmann,V.:pixelsplat:3dgaussiansplats fromimagepairsforscalablegeneralizable3dreconstruction.In:CVPR.pp.19457– 19467 (2024)
2024
-
[9]
IEEE transactions on pattern analysis and machine intelli- gence40(4), 834–848 (2017)
Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L.: Deeplab: Se- mantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelli- gence40(4), 834–848 (2017)
2017
-
[10]
In: ECCV
Chen, Y., Xu, H., Zheng, C., Zhuang, B., Pollefeys, M., Geiger, A., Cham, T.J., Cai, J.: Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. In: ECCV. pp. 370–386. Springer (2024)
2024
-
[11]
In: NeurIPS (2024)
Chen, Y., Zheng, C., Xu, H., Zhuang, B., Vedaldi, A., Cham, T.J., Cai, J.: Mvs- plat360: Feed-forward 360 scene synthesis from sparse views. In: NeurIPS (2024)
2024
-
[12]
In: CVPR
Chung, J., Oh, J., Lee, K.M.: Depth-regularized optimization for 3d gaussian splat- ting in few-shot images. In: CVPR. pp. 811–820 (2024)
2024
-
[13]
In: CVPR
Deng, K., Liu, A., Zhu, J.Y., Ramanan, D.: Depth-supervised nerf: Fewer views and faster training for free. In: CVPR. pp. 12882–12891 (2022)
2022
-
[14]
In: CVPR
Duzceker, A., Galliani, S., Vogel, C., Speciale, P., Dusmanu, M., Pollefeys, M.: Deepvideomvs: Multi-view stereo on video with recurrent spatio-temporal fusion. In: CVPR. pp. 15324–15333 (2021)
2021
-
[15]
In: ICCV
Eftekhar, A., Sax, A., Malik, J., Zamir, A.: Omnidata: A scalable pipeline for making multi-task mid-level vision datasets from 3d scans. In: ICCV. pp. 10786– 10796 (2021)
2021
-
[16]
In: CVPR
Fan, H., Su, H., Guibas, L.J.: A point set generation network for 3d object recon- struction from a single image. In: CVPR. pp. 605–613 (2017)
2017
-
[17]
arXiv preprint arXiv:2403.20309 (2024)
Fan, Z., Wen, K., Cong, W., Wang, K., Zhang, J., Ding, X., Xu, D., Ivanovic, B., Pavone, M., Pavlakos, G., et al.: Instantsplat: Sparse-view gaussian splatting in seconds. arXiv preprint arXiv:2403.20309 (2024)
Pith/arXiv arXiv 2024
-
[18]
In: CVPR
Fu, Y., Liu, S., Kulkarni, A., Kautz, J., Efros, A.A., Wang, X.: Colmap-free 3d gaussian splatting. In: CVPR. pp. 20796–20805 (2024) GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis 37
2024
-
[19]
NeurIPS (2024)
Gao*, R., Holynski*, A., Henzler, P., Brussee, A., Martin-Brualla, R., Srinivasan, P.P., Barron, J.T., Poole*, B.: Cat3d: Create anything in 3d with multi-view dif- fusion models. NeurIPS (2024)
2024
-
[20]
In: CVPR
Geng, D., Herrmann, C., Hur, J., Cole, F., Zhang, S., Pfaff, T., Lopez-Guevara, T., Aytar, Y., Rubinstein, M., Sun, C., et al.: Motion prompting: Controlling video generation with motion trajectories. In: CVPR. pp. 1–12 (2025)
2025
-
[21]
In: CVPR
Gu, X., Fan, Z., Zhu, S., Dai, Z., Tan, F., Tan, P.: Cascade cost volume for high- resolution multi-view stereo and stereo matching. In: CVPR. pp. 2495–2504 (2020)
2020
-
[22]
In: Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers
Gu, Z., Yan, R., Lu, J., Li, P., Dou, Z., Si, C., Dong, Z., Liu, Q., Lin, C., Liu, Z., et al.: Diffusion as shader: 3d-aware video diffusion for versatile video generation control. In: Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers. pp. 1–12 (2025)
2025
-
[23]
In: CVPR
Hanson, A., Tu, A., Singla, V., Jayawardhana, M., Zwicker, M., Goldstein, T.: Pup 3d-gs: Principled uncertainty pruning for 3d gaussian splatting. In: CVPR. pp. 5949–5958 (2025)
2025
-
[24]
In: ICLR (2025)
He, H., Xu, Y., Guo, Y., Wetzstein, G., Dai, B., Li, H., Yang, C.: Cameractrl: Enabling camera control for video diffusion models. In: ICLR (2025)
2025
-
[25]
arXiv preprint arXiv:2503.10592 (2025)
He, H., Yang, C., Lin, S., Xu, Y., Wei, M., Gui, L., Zhao, Q., Wetzstein, G., Jiang, L., Li, H.: Cameractrl ii: Dynamic scene exploration via camera-controlled video diffusion models. arXiv preprint arXiv:2503.10592 (2025)
Pith/arXiv arXiv 2025
-
[26]
ICLR1(2), 3 (2022)
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al.: Lora: Low-rank adaptation of large language models. ICLR1(2), 3 (2022)
2022
-
[27]
arXiv preprint arXiv:2508.10934 (2025)
Huang, J., Zhou, Q., Rabeti, H., Korovko, A., Ling, H., Ren, X., Shen, T., Gao, J., Slepichev,D.,Lin,C.H.,etal.:Vipe:Videoposeenginefor3dgeometricperception. arXiv preprint arXiv:2508.10934 (2025)
Pith/arXiv arXiv 2025
-
[28]
arXiv preprint arXiv:2601.09575 (2026)
Huang, S.Y., Choe, J., Wang, Y.C.F., Sun, C.: Openvoxel: Training-free grouping and captioning voxels for open-vocabulary 3d scene understanding. arXiv preprint arXiv:2601.09575 (2026)
arXiv 2026
-
[29]
arXiv preprint arXiv:1905.00538 (2019)
Im, S., Jeon, H.G., Lin, S., Kweon, I.S.: Dpsnet: End-to-end deep plane sweep stereo. arXiv preprint arXiv:1905.00538 (2019)
Pith/arXiv arXiv 1905
-
[30]
In: CVPR
Jensen, R., Dahl, A., Vogiatzis, G., Tola, E., Aanæs, H.: Large scale multi-view stereopsis evaluation. In: CVPR. pp. 406–413. IEEE (2014)
2014
-
[31]
arXiv preprint arXiv:2409.02084 (2024)
Ji, M., Qiu, R.Z., Zou, X., Wang, X.: Graspsplats: Efficient manipulation with 3d feature splatting. arXiv preprint arXiv:2409.02084 (2024)
Pith/arXiv arXiv 2024
-
[32]
arXiv preprint arXiv:2505.23716 (2025)
Jiang, L., Mao, Y., Xu, L., Lu, T., Ren, K., Jin, Y., Xu, X., Yu, M., Pang, J., Zhao, F., et al.: Anysplat: Feed-forward 3d gaussian splatting from unconstrained views. arXiv preprint arXiv:2505.23716 (2025)
arXiv 2025
-
[33]
splat: Directly referring 3d gaussian splatting via direct language embedding registration
Jun-Seong, K., GeonU, K., Yu-Ji, K., Wang, Y.C.F., Choe, J., Oh, T.H.: Dr. splat: Directly referring 3d gaussian splatting via direct language embedding registration. In: CVPR (2025)
2025
-
[34]
In: International Con- ference on 3D Vision (3DV)
Keetha, N., Müller, N., Schönberger, J., Porzi, L., Zhang, Y., Fischer, T., Knapitsch, A., Zauss, D., Weber, E., Antunes, N., Luiten, J., Lopez-Antequera, M., Bulò, S.R., Richardt, C., Ramanan, D., Scherer, S., Kontschieder, P.: MapA- nything: Universal feed-forward metric 3D reconstruction. In: International Con- ference on 3D Vision (3DV). IEEE (2026)
2026
-
[35]
ACM Transactions on Graphics42(4) (July 2023),���������������������������������������������������������
Kerbl, B., Kopanas, G., Leimkühler, T., Drettakis, G.: 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics42(4) (July 2023),���������������������������������������������������������
2023
-
[36]
In: ICCV
Kerr, J., Kim, C.M., Goldberg, K., Kanazawa, A., Tancik, M.: Lerf: Language embedded radiance fields. In: ICCV. pp. 19729–19739 (2023)
2023
-
[37]
In: ICCV
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., Lo, W.Y., et al.: Segment anything. In: ICCV. pp. 4015–4026 (2023) 38 Kang et al
2023
-
[38]
ACM Transactions on Graphics36(4) (2017)
Knapitsch, A., Park, J., Zhou, Q.Y., Koltun, V.: Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics36(4) (2017)
2017
-
[39]
In: CVPR
Lee, J., Park, C., Choe, J., Wang, Y.C.F., Kautz, J., Cho, M., Choy, C.: Mosaic3d: Foundation dataset and model for open-vocabulary 3d segmentation. In: CVPR. pp. 14089–14101 (2025)
2025
-
[40]
In: CVPR
Li, L., Zhang, Z., Li, Y., Xu, J., Hu, W., Li, X., Cheng, W., Gu, J., Xue, T., Shan, Y.: Nvcomposer: Boosting generative novel view synthesis with multiple sparse and unposed images. In: CVPR. pp. 777–787 (2025)
2025
-
[41]
Li, M., Yang, T., Kuang, H., Wu, J., Wang, Z., Xiao, X., Chen, C.: Control- net++: Improving conditional controls with efficient consistency feedback: Project page: liming-ai. github. io/controlnet_plus_plus. In: ECCV. pp. 129–147. Springer (2024)
2024
-
[42]
Advances in Neural Information Processing Systems (2025),��������������������������������
Li, W., Zhao, Y., Qin, M., Liu, Y., Cai, Y., Gan, C., Pfister, H.: Langsplatv2: High- dimensional 3d language gaussian splatting with 450+ fps. Advances in Neural Information Processing Systems (2025),��������������������������������
2025
-
[43]
arXiv preprint arXiv:2412.10231 (2024)
Liang, S., Wang, S., Li, K., Niemeyer, M., Gasperini, S., Lensch, H., Navab, N., Tombari, F.: Supergseg: Open-vocabulary 3d segmentation with structured super- gaussians. arXiv preprint arXiv:2412.10231 (2024)
arXiv 2024
-
[44]
arXiv preprint arXiv:2511.10647 (2025)
Lin, H., Chen, S., Liew, J., Chen, D.Y., Li, Z., Shi, G., Feng, J., Kang, B.: Depth anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647 (2025)
Pith/arXiv arXiv 2025
-
[45]
In: CVPR
Lin, T.Y., Dollár, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: CVPR. pp. 2117–2125 (2017)
2017
-
[46]
In: CVPR
Ling, L., Sheng, Y., Tu, Z., Zhao, W., Xin, C., Wan, K., Yu, L., Guo, Q., Yu, Z., Lu, Y., et al.: Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In: CVPR. pp. 22160–22169 (2024)
2024
-
[47]
arXiv preprint arXiv:2502.06782 (2025)
Liu, D., Li, S., Liu, Y., Li, Z., Wang, K., Li, X., Qin, Q., Liu, Y., Xin, Y., Li, Z., et al.: Lumina-video: Efficient and flexible video generation with multi-scale next-dit. arXiv preprint arXiv:2502.06782 (2025)
Pith/arXiv arXiv 2025
-
[48]
NeurIPS37, 133305–133327 (2024)
Liu, X., Zhou, C., Huang, S.: 3dgs-enhancer: Enhancing unbounded 3d gaussian splatting with view-consistent 2d diffusion priors. NeurIPS37, 133305–133327 (2024)
2024
-
[49]
In: ICCV
Marrie, J., Ménégaux, R., Arbel, M., Larlus, D., Mairal, J.: Ludvig: Learning-free uplifting of 2d visual features to gaussian splatting scenes. In: ICCV. pp. 7440–7450 (2025)
2025
-
[50]
arXiv preprint arXiv:2108.01073 (2021)
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.Y., Ermon, S.: Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv preprint arXiv:2108.01073 (2021)
Pith/arXiv arXiv 2021
-
[51]
ACM Transactions on Graphics (ToG)38(4), 1–14 (2019)
Mildenhall, B., Srinivasan, P.P., Ortiz-Cayon, R., Kalantari, N.K., Ramamoorthi, R., Ng, R., Kar, A.: Local light field fusion: Practical view synthesis with pre- scriptive sampling guidelines. ACM Transactions on Graphics (ToG)38(4), 1–14 (2019)
2019
-
[52]
Commu- nications of the ACM65(1), 99–106 (2021)
Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: Nerf: Representing scenes as neural radiance fields for view synthesis. Commu- nications of the ACM65(1), 99–106 (2021)
2021
-
[53]
arXiv preprint arXiv:2304.07193 (2023)
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al.: Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023)
Pith/arXiv arXiv 2023
-
[54]
In: CVPR
Qin, M., Li, W., Zhou, J., Wang, H., Pfister, H.: Langsplat: 3d language gaussian splatting. In: CVPR. pp. 20051–20060 (2024)
2024
-
[55]
In: ICML
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: ICML. pp. 8748–8763. PmLR (2021) GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis 39
2021
-
[56]
In: ICCV
Ranftl, R., Bochkovskiy, A., Koltun, V.: Vision transformers for dense prediction. In: ICCV. pp. 12179–12188 (2021)
2021
-
[57]
In: CVPR
Reizenstein,J.,Shapovalov,R.,Henzler,P.,Sbordone,L.,Labatut,P.,Novotny,D.: Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. In: CVPR. pp. 10901–10911 (2021)
2021
-
[58]
In: CVPR
Ren,X.,Shen,T.,Huang,J.,Ling,H.,Lu,Y.,Nimier-David,M.,Müller,T.,Keller, A., Fidler, S., Gao, J.: Gen3c: 3d-informed world-consistent video generation with precise camera control. In: CVPR. pp. 6121–6132 (2025)
2025
-
[59]
In: CVPR
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: CVPR. pp. 10684–10695 (2022)
2022
-
[60]
In: CVPR (2016)
Schönberger, J.L., Frahm, J.M.: Structure-from-motion revisited. In: CVPR (2016)
2016
-
[61]
In: ECCV (2016)
Schönberger, J.L., Zheng, E., Pollefeys, M., Frahm, J.M.: Pixelwise view selection for unstructured multi-view stereo. In: ECCV (2016)
2016
-
[62]
In: ICCV
Seo, S., Chang, Y., Kwak, N.: Flipnerf: Flipped reflection rays for few-shot novel view synthesis. In: ICCV. pp. 22883–22893 (2023)
2023
-
[63]
In: ICLR (2025)
Tang, S., Ye, W., Ye, P., Lin, W., Zhou, Y., Chen, T., Ouyang, W.: Hisplat: Hierarchical 3d gaussian splatting for generalizable sparse-view reconstruction. In: ICLR (2025)
2025
-
[64]
In: ICCV
Wang, G., Chen, Z., Loy, C.C., Liu, Z.: Sparsenerf: Distilling depth ranking for few-shot novel view synthesis. In: ICCV. pp. 9065–9076 (2023)
2023
-
[65]
In: CVPR
Wang, J., Chen, M., Karaev, N., Vedaldi, A., Rupprecht, C., Novotny, D.: Vggt: Visual geometry grounded transformer. In: CVPR. pp. 5294–5306 (2025)
2025
-
[66]
IEEE transactions on pattern analysis and machine intelligence43(10), 3349–3364 (2020)
Wang, J., Sun, K., Cheng, T., Jiang, B., Deng, C., Zhao, Y., Liu, D., Mu, Y., Tan, M., Wang, X., et al.: Deep high-resolution representation learning for visual recog- nition. IEEE transactions on pattern analysis and machine intelligence43(10), 3349–3364 (2020)
2020
-
[67]
arXiv preprint arXiv:2511.06457 (2025)
Wang, S., Zhang, S., Millerdurai, C., Westermann, R., Stricker, D., Pagani, A.: Inpaint360gs: Efficient object-aware 3d inpainting via gaussian splatting for 360 {\deg}scenes. arXiv preprint arXiv:2511.06457 (2025)
arXiv 2025
-
[68]
In: CVPR
Wang, S., Leroy, V., Cabon, Y., Chidlovskii, B., Revaud, J.: Dust3r: Geometric 3d vision made easy. In: CVPR. pp. 20697–20709 (2024)
2024
-
[69]
In: ICCV
Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., Shao, L.: Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In: ICCV. pp. 568–578 (2021)
2021
-
[70]
Wang, Y., Zhou, J., Zhu, H., Chang, W., Zhou, Y., Li, Z., Chen, J., Pang, J., Shen, C., He, T.:π3: Scalable permutation-equivariant visual geometry learning (2025), ��������������������������������
2025
-
[71]
In: SIGGRAPH
Wang,Z.,Yuan,Z.,Wang,X.,Li,Y.,Chen,T.,Xia,M.,Luo,P.,Shan,Y.:Motionc- trl: A unified and flexible motion controller for video generation. In: SIGGRAPH. pp. 1–11 (2024)
2024
-
[72]
arXiv preprint arXiv:2407.07860 (2024)
Watson, D., Saxena, S., Li, L., Tagliasacchi, A., Fleet, D.J.: Controlling space and time with diffusion models. arXiv preprint arXiv:2407.07860 (2024)
Pith/arXiv arXiv 2024
-
[73]
arXiv preprint arXiv:2507.07982 (2025)
Wu, H., Wu, D., He, T., Guo, J., Ye, Y., Duan, Y., Bian, J.: Geometry forcing: Mar- rying video diffusion and 3d representation for consistent world modeling. arXiv preprint arXiv:2507.07982 (2025)
Pith/arXiv arXiv 2025
-
[74]
In: CVPR
Wu, J.Z., Zhang, Y., Turki, H., Ren, X., Gao, J., Shou, M.Z., Fidler, S., Gojcic, Z., Ling, H.: Difix3d+: Improving 3d reconstructions with single-step diffusion models. In: CVPR. pp. 26024–26035 (2025)
2025
-
[75]
In: CVPR
Wu, R., Mildenhall, B., Henzler, P., Park, K., Gao, R., Watson, D., Srinivasan, P.P., Verbin, D., Barron, J.T., Poole, B., et al.: Reconfusion: 3d reconstruction with diffusion priors. In: CVPR. pp. 21551–21561 (2024) 40 Kang et al
2024
-
[76]
In: CVPR
Wu, S., Xu, C., Huang, B., Geiger, A., Chen, A.: Genfusion: Closing the loop between reconstruction and generation via videos. In: CVPR. pp. 6078–6088 (2025)
2025
-
[77]
In: CVPR
Wu, T., Zhang, J., Fu, X., Wang, Y., Ren, J., Pan, L., Wu, W., Yang, L., Wang, J., Qian, C., et al.: Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation. In: CVPR. pp. 803–814 (2023)
2023
-
[78]
In: ECCV
Wu, W., Li, Z., Gu, Y., Zhao, R., He, Y., Zhang, D.J., Shou, M.Z., Li, Y., Gao, T., Zhang, D.: Draganything: Motion control for anything using entity representation. In: ECCV. pp. 331–348. Springer (2024)
2024
-
[79]
In: CVPR
Xia, H., Fu, Y., Liu, S., Wang, X.: Rgbd objects in the wild: Scaling real-world 3d object learning from rgb-d videos. In: CVPR. pp. 22378–22389 (2024)
2024
-
[80]
In: CVPR
Xu, H., Peng, S., Wang, F., Blum, H., Barath, D., Geiger, A., Pollefeys, M.: Depth- splat: Connecting gaussian splatting and depth. In: CVPR. pp. 16453–16463 (2025)
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.