REVIEW 3 major objections 3 minor 220 references
Sparse Input View Synthesis: 3D Representations and Reliable Priors
T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Reliable priors make sparse-input view synthesis work
desk verdict A solid, clearly written thesis that compiles four published methods; the novel ideas are mostly in the reliability of pre-training-free priors, but the evaluation has weak spots around depth validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is reliability-weighted supervision: every prior is applied only where it can be checked. For static scenes, the visibility prior is computed by warping one input view into another at several depth planes (a plane sweep volume) and thresholding the minimum per-pixel matching error, so the loss acts only on pixels judged visible in the second view. For Simple-RF, depth from reduced-capability augmented models is gated by a reprojection mask that compares each model's depth patch against the nearest training view, and the augmented models are built by lowering the positional-encoding degree, the hash-table size, or the tensor rank. For dynamic scenes, the motion field is a four-dimensional deformation stored as six factorized planes that maps every time to a canonical volume, and it is supervised by sparse keypoint correspondences that pull matched pixels to the same canonical 3D point. For temporal view synthesis, multi-plane images (a stack of images at discrete depth planes) together with masked correlation estimate object motion in 3D after nullifying camera motion.
What would settle it
A specular or textureless scene captured from only two viewpoints provides a direct test: if the visibility prior mislabels visibly matched pixels as occluded because their colors change with view, ViP-NeRF's render quality should fall below a depth-smoothness baseline on the same inputs; similarly, a dynamic multi-view clip with large inter-camera occlusion should show sparse-flow matching dragging keypoints to wrong canonical points, measurable as a rise in depth error against a dense-view reference.
Extended reading notes
Core claim
The paper's central discovery is that visibility—whether a pixel's surface appears in a second view—is a dense and reliable prior for sparse-input neural radiance fields, and that relative depth cues of this kind beat absolute-depth priors produced by pre-trained networks. It shows further that depth supervision need not come from external models: reduced-capability 'augmented' radiance fields trained alongside the main model provide better depth in smooth or Lambertian regions, and a reprojection-error check decides when each model's depth is trustworthy. For dynamic scenes, the thesis finds that dense optical flow across cameras is unreliable as a motion prior, whereas sparse keypoint matches, used to pull corresponding points to the same canonical 3D location, stabilize a factorized deformation field with only three input views. In temporal view synthesis, it shows that decoupling camera and object motion and estimating object motion in the 3D multi-plane image space improves future-frame prediction and disocclusion infilling.
Load-bearing premise
Every prior assumes photometric consistency between views: if pixel intensities change across views—specular surfaces, lighting shifts, or heavy occlusion—the matching that underpins the prior can label visible pixels as occluded or pick the wrong model's depth, pushing the radiance field in the wrong direction.
Editorial extensions
If this is right
- With two to four input views on forward-facing scenes, visibility regularization outperforms learned dense depth priors on both rendering quality and depth accuracy.
- The same 'simpler solutions' supervision recipe improves three different radiance fields—NeRF, TensoRF, and ZipNeRF—and removes floaters and duplication artifacts characteristic of sparse-input training.
- For dynamic multi-view scenes with three cameras, a factorized deformation field regularized by sparse keypoint flows surpasses a model trained with dense optical flow priors, which actively hurts performance.
- Frame-rate upsampling of rendered video improves when object motion is estimated in 3D multi-plane image space after nullifying camera motion, rather than predicted as 2D video motion.
- Reliability gating is a necessary ingredient: ablations that disable the visibility prior, the reliability masks, or the coarse-fine consistency loss all degrade performance.
Reading between the lines
- The reliability-gating principle likely transfers beyond view synthesis: any under-constrained inverse problem with a cheap geometric check (reprojection, consistency, loop closure) could audit a learned prior the same way.
- Because the visibility prior constrains relative depth ordering rather than absolute scale, combining it with a monocular absolute-depth estimate could give both robustness and metric scale without dense learned priors.
- The scene-specific augmentation recipe could extend to 3D Gaussian splatting once its sparse-input initialization is solved, potentially closing a gap this thesis leaves open for that representation.
- For dynamic scenes, the sparse keypoint ceiling could be raised by densifying correspondences with a network trained on the same scene, turning the reliable-sparse idea into a self-supervised densification loop.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The thesis addresses sparse-input novel view synthesis for static and dynamic scenes. It proposes four families of methods: ViP-NeRF, which regularizes a NeRF with a dense visibility prior computed from plane sweep volumes; Simple-RF, which trains reduced-capacity 'augmented' radiance fields (NeRF, TensoRF, ZipNeRF) in tandem with the main model and supervises via a depth-reliability mask; SF-DeRF, a fast dynamic radiance field with an explicit factorized motion field regularized by sparse SIFT-based flow priors; and DeCOMPnet, an MPI-based temporal view synthesis method that decomposes camera and object motion for frame-rate upsampling. The thesis claims state-of-the-art performance on RealEstate-10K, NeRF-LLFF, MipNeRF360, NeRF-Synthetic, N3DV, InterDigital, and MPI Sintel, and introduces the IISc VEED-Dynamic dataset.
Significance. If the empirical claims hold, the thesis makes a useful contribution by showing that reliable geometric priors can be computed in a scene-specific, pre-training-free manner: the visibility prior and the simplicity-based depth priors avoid the generalization problems of learned dense depth priors, and the sparse flow priors avoid the unreliability of dense optical flow in multi-camera dynamic scenes. The breadth of the work is a strength: it spans static and dynamic scenes, implicit and explicit radiance fields, and an application-oriented temporal view synthesis setting. The thesis is also transparent about test-set choices and metric changes, and the priors are computed independently of the supervised targets, so the central claims are not circular. The main weaknesses are that the depth-reliability mask of Simple-RF is never validated directly against ground-truth depth, all comparisons are single-run without error bars, and the dynamic-view-synthesis state-of-the-art claim rests on a narrow set of baselines.
major comments (3)
- [§4.3.1.4, Eq. (4.8); §4.4.2] The depth-reliability mask is the load-bearing component of the Simple-RF framework, but the manuscript never validates the assumption that lower reprojection MSE implies more accurate depth. The ablations in Table 4.5 show that removing the mask degrades performance, but they do not establish that the mask selects the more accurate depth; other components (capacity reduction, coarse-fine consistency) may be responsible for the gains. Critically, Sec. 4.4.2 states that NeRF-Synthetic 'ground truth depth is not provided in the dataset either,' which is incorrect: the Blender-based NeRF-Synthetic dataset includes depth maps for every view, and Simple-ZipNeRF is evaluated on this dataset in Table 4.8. The one dataset that could directly test the reliability assumption is therefore available but avoided. I request a direct evaluation of the mask against ground-truth depth on NeRF-Synthetic, or a corrected and substantiated justification of why this is infeasible.
- [§4.4.2, Tables 4.2–4.4 and 3.1–3.4] All empirical comparisons are single-run and depth quality is measured against pseudo ground truth from dense-view NeRF/ZipNeRF models rather than true depth. Given that some headline improvements are small (e.g., Table 4.3, 4-view row: Simple-NeRF vs ViP-NeRF, LPIPS 0.0847 vs 0.0892), the lack of repeated runs, confidence intervals, or significance tests makes it hard to assess whether the state-of-the-art claims are robust. I ask for at least a few seeds for the main comparisons, or an explicit discussion of training variance.
- [§5.3.3, Table 5.2] The claim that SF-DeRF 'outperforms the state-of-the-art dynamic view synthesis models with fewer input viewpoints' is supported only against K-Planes and HexPlane. Other dynamic radiance field methods discussed in Sec. 5.1 (e.g., D-NeRF, TiNeuVox) are not included in the quantitative comparisons. If the published version contains additional baselines, they should be reproduced or clearly cited in the thesis so that the state-of-the-art claim is commensurate with the evidence.
minor comments (3)
- [Table 4.7] The Simple-ZipNeRF row appears to have concatenated values ('21.030.239'), likely a formatting or rendering error; please fix the table so all entries are cleanly separated.
- [Throughout (e.g., Eq. (3.10), Eq. (4.8))] Several mathematical expressions contain glyph artifacts, such as '/x31' in place of an indicator function and '∇' used as a stop-gradient placeholder. These should be typeset properly for a camera-ready version.
- [Abstract and Sec. 1.1] The word 'meta-verse' should be 'metaverse'; also, the thesis says in Sec. 1.1.2 that source code will be released, but Sec. 1.2 only gives a publications page URL. Please include a direct pointer to the code repositories if they are available.
Circularity Check
No circular derivation found; the priors and regularizers are computed independently of the supervised targets.
full rationale
The claimed derivations do not reduce to their inputs. ViP-NeRF's visibility prior (Eq. 3.10) is computed from plane-sweep photometric error on the input images alone, and the NeRF visibility is a volume-rendered quantity supervised by that prior (Eqs. 3.7-3.8), so the supervised quantity is not defined as the prior. Simple-RF's augmented models are trained in tandem, but the reliability mask (Eq. 4.8) is computed from reprojection MSE to the nearest training view, and the mutual depth loss (Eq. 4.9) uses stop-gradient targets; this is a co-training loop rather than a by-construction equivalence. SF-DeRF obtains sparse flow priors from SIFT matches (Sec. 5.2.2) and constrains canonical-volume agreement (Eq. 5.7), independent of the optimized motion field. DeCOMPnet explicitly decouples camera and object motion and predicts the future frame from past frames; no target quantity appears in the definition of the priors. The thesis is a compilation of the author's own published papers, but the self-citations point to externally evaluated experiments, and no load-bearing uniqueness theorem or ansatz is imported from those citations. The unvalidated assumption that reprojection MSE ranks depth accuracy (Ch. 4.3.1.4) is a correctness and robustness concern, not circularity, because the final metrics are computed on held-out views. The disputed statement about ground-truth depth on NeRF-Synthetic is also a factual/evidence concern rather than a self-referential step. No equation or fitted parameter equates a prediction with its input by construction, so the circularity score is 0.
Assumptions & free parameters
free parameters (8)
- ViP-NeRF visibility threshold gamma =
10
- ViP-NeRF loss weights (lambda1-4) =
1, 0.1, 0.001, 0.1
- SimpleNeRF reduced position encoding degree l_s^p =
3
- SimpleNeRF reliability threshold e_tau and patch size k =
0.1, 5
- Simple-TensoRF augmentation capacity =
R_s=12, N_vox=160^3, b_z1^s=-0.5, N_mc=5
- Simple-ZipNeRF augmentation capacity =
T^s=2^11, s_near=0.3, e_tau=0.2
- SF-DeRF flow prior time offset and weight =
s in {t-10, t+10}, lambda_sf=1
- DeCOMPnet hyperparameters =
not enumerated in text
assumptions (5)
- domain assumption Accurate camera poses are available for input views.
- domain assumption Photometric consistency between matched pixels across views.
- domain assumption Reprojection MSE is a valid reliability criterion for depth supervision.
- domain assumption SIFT keypoint matches are reliable flow priors across cameras.
- ad hoc to paper Dense-view NeRF depth is a valid pseudo ground truth.
Cite this review
Pith. "Pith review of Sparse Input View Synthesis: 3D Representations and Reliable Priors." pith.science (2026). https://pith.science/paper/TXX77FSA
@misc{pith2026241113631,
author = {Pith},
title = {Pith review of: Sparse Input View Synthesis: 3D Representations and Reliable Priors},
year = {2026},
howpublished = {\url{https://pith.science/paper/TXX77FSA}},
note = {Machine review of arXiv:2411.13631}
}
read the original abstract
Novel view synthesis refers to the problem of synthesizing novel viewpoints of a scene given the images from a few viewpoints. This is a fundamental problem in computer vision and graphics, and enables a vast variety of applications such as meta-verse, free-view watching of events, video gaming, video stabilization and video compression. Recent 3D representations such as radiance fields and multi-plane images significantly improve the quality of images rendered from novel viewpoints. However, these models require a dense sampling of input views for high quality renders. Their performance goes down significantly when only a few input views are available. In this thesis, we focus on the sparse input novel view synthesis problem for both static and dynamic scenes. In the first part of this work, we mainly focus on sparse input novel view synthesis of static scenes using neural radiance fields (NeRF). We study the design of reliable and dense priors to better regularize the NeRF in such situations. In particular, we propose a prior on the visibility of the pixels in a pair of input views. We show that this visibility prior, which is related to the relative depth of objects, is dense and more reliable than existing priors on absolute depth. We compute the visibility prior using plane sweep volumes without the need to train a neural network on large datasets. We evaluate our approach on multiple datasets and show that our model outperforms existing approaches for sparse input novel view synthesis. In the second part, we aim to further improve the regularization by learning a scene-specific prior that does not suffer from generalization issues. We achieve this by learning the prior on the given scene alone without pre-training on large datasets. In particular, we design augmented NeRFs to obtain better depth supervision in certain regions of the scene for the main NeRF. Further, we extend this framework to also apply to newer and faster radiance field models such as TensoRF and ZipNeRF. Through extensive experiments on multiple datasets, we show the superiority of our approach in sparse input novel view synthesis. The design of sparse input fast dynamic radiance fields is severely constrained by the lack of suitable representations and reliable priors for motion. We address the first challenge by designing an explicit motion model based on factorized volumes that is compact and optimizes quickly. We also introduce reliable sparse flow priors to constrain the motion field, since we find that the popularly employed dense optical flow priors are unreliable. We show the benefits of our motion representation and reliable priors on multiple datasets. In the final part of this thesis, we study the application of view synthesis for frame rate upsampling in video gaming. Specifically, we consider the problem of temporal view synthesis, where the goal is to predict the future frames given the past frames and the camera motion. The key challenge here is in predicting the future motion of the objects by estimating their past motion and extrapolating it. We explore the use of multi-plane image representations and scene depth to reliably estimate the object motion, particularly in the occluded regions. We design a new database to effectively evaluate our approach for temporal view synthesis of dynamic scenes and show that we achieve state-of-the-art performance.
Figures
Figures from the paper (52 more)
Reference graph
Works this paper leans on
-
[3]
Adrian, R. J. (1991). Particle-imaging techniques for experimental fluid mechanics. Annual review of fluid mechanics , 23(1):261–304
1991
-
[4]
and Beeler, D
Aksoy, V. and Beeler, D. (2019). Introducing asw 2.0: Bet- ter accuracy, lower latency. https://www.oculus.com/blog/ introducing-asw-2-point-0-better-accuracy-lower-latency/ . Accessed: 24-June-2021
2019
-
[5]
Allen, B., Curless, B., and Popović, Z. (2003). The space of human body shapes: Re- construction and parameterization from range scans. ACM Transactions on Graphics (TOG), 22(3):587–594
2003
-
[6]
Anguelov, D., Srinivasan, P., Koller, D., Thrun, S., Rodgers, J., and Davis, J. (2005). SCAPE: Shape completion and animation of people. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
2005
-
[7]
Antonov, M. (2015). Asynchronous timewarp examined. https://developer. oculus.com/blog/asynchronous-timewarp-examined. Accessed: 24-June-2021. 131 132 Chapter 8. Bibliography
2015
-
[8]
Atcheson, B., Ihrke, I., Heidrich, W., Tevs, A., Bradley, D., Magnor, M., and Seidel, H.-P. (2008). Time-resolved 3D capture of non-stationary gas flows. ACM Transactions on Graphics (TOG) , 27(5)
2008
-
[9]
H., and Levine, S
Babaeizadeh, M., Finn, C., Erhan, D., Campbell, R. H., and Levine, S. (2018). Stochastic variational video prediction. In Proceedings of the International Conference on Learning Representations (ICLR)
2018
-
[10]
J., and Szeliski, R
Baker, S., Scharstein, D., Lewis, J., Roth, S., Black, M. J., and Szeliski, R. (2011). A database and evaluation methodology for optical flow. International Journal of Computer Vision (IJCV) , 92(1):1–31
2011
Show all 220 references
-
[11]
and Zollhöfer, M
Bansal, A. and Zollhöfer, M. (2023). Neural pixel composition for 3d-4d view syn- thesis from multi-views. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[12]
Bao, W., Lai, W.-S., Ma, C., Zhang, X., Gao, Z., and Yang, M.-H. (2019). Depth- aware video frame interpolation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2019
-
[13]
Bao, W., Lai, W.-S., Zhang, X., Gao, Z., and Yang, M.-H. (2021). Memc-net: Motion estimation and motion compensation driven neural network for video inter- polation and enhancement. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , 43(3):933–948
2021
-
[14]
Barnes, C., Shechtman, E., Finkelstein, A., and Goldman, D. B. (2009). Patch- match: A randomized correspondence algorithm for structural image editing. ACM Transactions on Graphics (TOG) , 28(3):24
2009
-
[15]
Barnes, R. M. (2017). A positional timewarp accelerator for mobile virtual reality devices
2017
-
[16]
T., Mildenhall, B., Tancik, M., Hedman, P., Martin-Brualla, R., and Srinivasan, P
Barron, J. T., Mildenhall, B., Tancik, M., Hedman, P., Martin-Brualla, R., and Srinivasan, P. P. (2021). Mip-NeRF: A multiscale representation for anti-aliasing 133 neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2021
-
[17]
T., Mildenhall, B., Verbin, D., Srinivasan, P
Barron, J. T., Mildenhall, B., Verbin, D., Srinivasan, P. P., and Hedman, P. (2022). Mip-NeRF 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[18]
T., Mildenhall, B., Verbin, D., Srinivasan, P
Barron, J. T., Mildenhall, B., Verbin, D., Srinivasan, P. P., and Hedman, P. (2023). Zip-NeRF: Anti-aliased grid-based neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2023
-
[19]
Beeler, D. (2016). Asynchronous spacewarp. https://developer.oculus.com/ blog/asynchronous-spacewarp. Accessed: 24-June-2021
2016
-
[20]
and Gosalia, A
Beeler, D. and Gosalia, A. (2016). Asynchronous timewarp on oculus rift. https: //developer.oculus.com/blog/asynchronous-timewarp-on-oculus-rift . Ac- cessed: 24-June-2021
2016
-
[21]
and Vetter, T
Blanz, V. and Vetter, T. (1999). A morphable model for the synthesis of 3d faces. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
1999
-
[22]
Bortolon, M., Del Bue, A., and Poiesi, F. (2022). Data augmentation for NeRF: a geometric consistent solution based on view morphing. arXiv e-prints , page arXiv:2210.04214
2022 arXiv
-
[23]
Bradley, D., Heidrich, W., Popa, T., and Sheffer, A. (2010). High resolution passive facial performance capture. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
2010
-
[24]
Bradley, D., Popa, T., Sheffer, A., Heidrich, W., and Boubekeur, T. (2008). Mark- erless garment capture. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH) . 134 Chapter 8. Bibliography
2008
-
[25]
Broxton, M., Flynn, J., Overbeck, R., Erickson, D., Hedman, P., Duvall, M., Dour- garian, J., Busch, J., Whalen, M., and Debevec, P. (2020). Immersive light field video with a layered mesh representation. ACM Transactions on Graphics (TOG) , 39(4)
2020
-
[26]
J., Wulff, J., Stanley, G
Butler, D. J., Wulff, J., Stanley, G. B., and Black, M. J. (2012). A naturalistic open source movie for optical flow evaluation. In Proceedings of the European Conference on Computer Vision (ECCV)
2012
-
[27]
Cai, S., Obukhov, A., Dai, D., and Van Gool, L. (2022). Pix2NeRF: Unsupervised conditional p-GAN for single image to neural radiance fields translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[28]
and Johnson, J
Cao, A. and Johnson, J. (2023). HexPlane: A fast representation for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[29]
A., and Seidel, H.-P
Carranza, J., Theobalt, C., Magnor, M. A., and Seidel, H.-P. (2003). Free-viewpoint video of human actors. ACM Transactions on Graphics (TOG) , 22(3):569–577
2003
-
[30]
Chai, J.-X., Tong, X., Chan, S.-C., and Shum, H.-Y. (2000). Plenoptic sampling. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
2000
-
[31]
R., Monteiro, M., Kellnhofer, P., Wu, J., and Wetzstein, G
Chan, E. R., Monteiro, M., Kellnhofer, P., Wu, J., and Wetzstein, G. (2021). Pi- GAN: Periodic implicit generative adversarial networks for 3D-aware image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR)
2021
-
[32]
Chaurasia, G., Duchene, S., Sorkine-Hornung, O., and Drettakis, G. (2013). Depth synthesis and local warps for plausible image-based navigation. ACM Transactions on Graphics (TOG) , 32(3)
2013
-
[33]
Chen, A., Xu, Z., Geiger, A., Yu, J., and Su, H. (2022a). TensoRF: Tensorial radi- ance fields. In Proceedings of the European Conference on Computer Vision (ECCV) . 135
2022
-
[34]
Chen, A., Xu, Z., Zhao, F., Zhang, X., Xiang, F., Yu, J., and Su, H. (2021). MVSNeRF: Fast generalizable radiance field reconstruction from multi-view stereo. arXiv e-prints , page arXiv:2103.15595
2021 arXiv
-
[35]
Chen, D., Liu, Y., Huang, L., Wang, B., and Pan, P. (2022b). GeoAug: Data augmentation for few-shot NeRF with geometry constraints. In Proceedings of the European Conference on Computer Vision (ECCV)
2022
-
[36]
Chen, S. E. and Williams, L. (1993). View interpolation for image synthesis. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
1993
-
[37]
Chen, Y., Xu, H., Wu, Q., Zheng, C., Cham, T.-J., and Cai, J. (2023). Explicit correspondence matching for generalizable neural radiance fields. arXiv e-prints , page arXiv:2304.12294
2023 arXiv
-
[38]
Chen, Y., Xu, H., Zheng, C., Zhuang, B., Pollefeys, M., Geiger, A., Cham, T.-J., and Cai, J. (2024). MVSplat: Efficient 3d gaussian splatting from sparse multi-view images. arXiv e-prints , page arXiv:2403.14627
2024 arXiv
-
[39]
Chibane, J., Bansal, A., Lazova, V., and Pons-Moll, G. (2021). Stereo radiance fields (SRF): Learning view synthesis for sparse views of novel scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2021
-
[40]
Cho, J.-H., Song, W., Choi, H., and Kim, T. (2017). Hole filling method for depth image based rendering based on boundary decision. IEEE Signal Processing Letters (SPL), 24(3):329–333
2017
-
[41]
Collet, A., Chuang, M., Sweeney, P., Gillett, D., Evseev, D., Calabrese, D., Hoppe, H., Kirk, A., and Sullivan, S. (2015). High-quality streamable free-viewpoint video. ACM Transactions on Graphics (TOG) , 34(4). 136 Chapter 8. Bibliography
2015
-
[42]
Collins, R. (1996). A space-sweep approach to true multi-image matching. In Pro- ceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR)
1996
-
[43]
Corder, G. W. and Foreman, D. I. (2014). Nonparametric statistics: A step-by-step approach
2014
-
[44]
Criminisi, A., Pérez, P., and Toyama, K. (2004). Region filling and object removal by exemplar-based image inpainting. IEEE Transactions on Image Processing (TIP) , 13(9):1200–1212
2004
-
[45]
de Aguiar, E., Stoll, C., Theobalt, C., Ahmed, N., Seidel, H.-P., and Thrun, S. (2008). Performance capture from sparse multi-view video. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
2008
-
[46]
Deng, K., Liu, A., Zhu, J.-Y., and Ramanan, D. (2022). Depth-supervised NeRF: Fewer views and faster training for free. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[47]
and Fergus, R
Denton, E. and Fergus, R. (2018). Stochastic video generation with a learned prior. In Proceedings of the International Conference on Machine Learning (ICML)
2018
-
[48]
Faloutsos, P., Van de Panne, M., and Terzopoulos, D. (1997). Dynamic free-form deformations for animation synthesis. IEEE Transactions on Visualization and Com- puter Graphics (TVCG) , 3(3):201–214
1997
-
[49]
Fang, J., Yi, T., Wang, X., Xie, L., Zhang, X., Liu, W., Nießner, M., and Tian, Q. (2022). Fast dynamic radiance fields with time-aware neural voxels. In Proceedings of the SIGGRAPH Asia 2022 Conference Papers
2022
-
[50]
Fehn, C. (2004). Depth-image-based rendering (DIBR), compression and transmis- sion for a new approach on 3D-TV. In Proceedings of the Stereoscopic Displays and Virtual Reality Systems XI . 137
2004
-
[51]
Finn, C., Goodfellow, I., and Levine, S. (2016). Unsupervised learning for phys- ical interaction through video prediction. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)
2016
-
[52]
Flynn, J., Neulander, I., Philbin, J., and Snavely, N. (2016). DeepStereo: Learning to predict new views from the world’s imagery. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[53]
R., Recht, B., and Kanazawa, A
Fridovich-Keil, S., Meanti, G., Warburg, F. R., Recht, B., and Kanazawa, A. (2023). K-Planes: Explicit radiance fields in space, time, and appearance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[54]
Fridovich-Keil, S., Yu, A., Tancik, M., Chen, Q., Recht, B., and Kanazawa, A. (2022). Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[55]
Gallup, D., Frahm, J.-M., Mordohai, P., Yang, Q., and Pollefeys, M. (2007). Real- time plane-sweeping stereo with multiple sweeping directions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2007
-
[56]
Gao, C., Saraf, A., Kopf, J., and Huang, J.-B. (2021). Dynamic view synthesis from dynamic monocular video. arXiv e-prints , page arXiv:2105.06468
2021 arXiv
-
[57]
Gao, C., Shih, Y., Lai, W.-S., Liang, C.-K., and Huang, J.-B. (2020). Portrait neural radiance fields from a single image. arXiv e-prints , page arXiv:2012.05903
2020 arXiv
-
[58]
Gao, H., Xu, H., Cai, Q.-Z., Wang, R., Yu, F., and Darrell, T. (2019). Disentan- gling propagation and generation for video prediction. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2019
-
[59]
Gao, Z., Dai, W., and Zhang, Y. (2024). HG3-NeRF: Hierarchical geometric, se- mantic, and photometric guided neural radiance fields for sparse view inputs. arXiv e-prints, page arXiv:2401.11711. 138 Chapter 8. Bibliography
2024 arXiv
-
[60]
J., Grzeszczuk, R., Szeliski, R., and Cohen, M
Gortler, S. J., Grzeszczuk, R., Szeliski, R., and Cohen, M. F. (1996). The lumi- graph. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
1996
-
[61]
B., and Heidrich, W
Gregson, J., Krimerman, M., Hullin, M. B., and Heidrich, W. (2012). Stochastic tomography and its applications in 3d imaging of mixing fluids. ACM Transactions on Graphics (TOG) , 31(4):1–10
2012
-
[62]
Guo, S., Wang, Q., Gao, Y., Xie, R., and Song, L. (2024). Depth-guided robust and fast point cloud fusion NeRF for sparse input views. Proceedings of the AAAI Conference on Artificial Intelligence , 38(3):1976–1984
2024
-
[63]
Guo, X., Sun, J., Dai, Y., Chen, G., Ye, X., Tan, X., Ding, E., Zhang, Y., and Wang, J. (2023). Forward flow for novel view synthesis of dynamic scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2023
-
[64]
Guskov, I., Klibanov, S., and Bryant, B. (2003). Trackable surfaces. In Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation (SCA)
2003
-
[65]
Ha, H., Im, S., Park, J., Jeon, H.-G., and Kweon, I. S. (2016). High-quality depth from uncalibrated small motion clip. In Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR)
2016
-
[66]
Hamdi, A., Ghanem, B., and Nießner, M. (2022). SPARF: Large-scale learning of 3d sparse radiance fields from few input images. arXiv e-prints , page arXiv:2212.09100
2022 arXiv
-
[67]
Han, Y., Wang, R., and Yang, J. (2022). Single-view view synthesis in the wild with learned adaptive multiplane images. In Proceedings of the ACM SIGGRAPH
2022
-
[68]
Hasler, N., Asbach, M., Rosenhahn, B., Ohm, J.-R., and Seidel, H.-P. (2006). Phys- ically based tracking of cloth. In Proceedings of the International Workshop on Vision, Modeling, and Visualization (VMV) . 139
2006
-
[69]
Hawkins, T., Einarsson, P., and Debevec, P. (2005). Acquisition of time-varying participating media. ACM Transactions on Graphics (TOG) , 24(3):812–815
2005
-
[70]
Horn, B. K. P. and Schunck, B. G. (1981). Determining optical flow. Artificial intelligence, 17(1-3):185–203
1981
-
[71]
Huang, P.-H., Matzen, K., Kopf, J., Ahuja, N., and Huang, J.-B. (2018). DeepMVS: Learning multi-view stereopsis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2018
-
[72]
and Magnor, M
Ihrke, I. and Magnor, M. (2004). Image-based tomographic reconstruction of flames. In Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Ani- mation (SCA)
2004
-
[73]
Iizuka, S., Simo-Serra, E., and Ishikawa, H. (2017). Globally and locally consistent image completion. ACM Transactions on Graphics (TOG) , 36(4):1–14
2017
-
[74]
Im, S., Jeon, H.-G., Lin, S., and Kweon, I. S. (2019). DPSNet: End-to-end deep plane sweep stereo. arXiv e-prints , page arXiv:1905.00538
2019 arXiv
-
[75]
Jain, A., Tancik, M., and Abbeel, P. (2021). Putting NeRF on a diet: Semantically consistent few-shot view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2021
-
[76]
M., Lepoittevin, Y., and Fleuret, F
Johari, M. M., Lepoittevin, Y., and Fleuret, F. (2022). Geonerf: Generalizing nerf with geometry priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[77]
Joo, H., Simon, T., and Sheikh, Y. (2018). Total capture: A 3d deformation model for tracking faces, hands, and bodies. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2018
-
[78]
K., Wang, T.-C., and Ramamoorthi, R
Kalantari, N. K., Wang, T.-C., and Ramamoorthi, R. (2016). Learning-based view synthesis for light field cameras. ACM Transactions on Graphics (TOG) , 35(6). 140 Chapter 8. Bibliography
2016
-
[79]
Kanchana, V., Somraj, N., Yadwad, S., and Soundararajan, R. (2022). Revealing disocclusions in temporal view synthesis through infilling vector prediction. In Pro- ceedings of the IEEE Winter Conference on Applications of Computer Vision (W ACV)
2022
-
[80]
Ke, Z., Wang, D., Yan, Q., Ren, J., and Lau, R. W. (2019). Dual student: Breaking the limits of the teacher in semi-supervised learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2019
-
[81]
Kerbl, B., Kopanas, G., Leimkühler, T., and Drettakis, G. (2023). 3d gaussian splat- ting for real-time radiance field rendering. ACM Transactions on Graphics (TOG) , 42(4)
2023
-
[82]
Kim, D., Woo, S., Lee, J.-Y., and Kweon, I. S. (2019). Deep video inpainting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2019
-
[83]
Kim, M., Seo, S., and Han, B. (2022). InfoNeRF: Ray entropy minimization for few-shot neural volume rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[84]
and Sheffer, A
Kraevoy, V. and Sheffer, A. (2004). Cross-parameterization and compatible remesh- ing of 3d models. ACM Transactions on Graphics (TOG) , 23(3):861–869
2004
-
[85]
and Sheffer, A
Kraevoy, V. and Sheffer, A. (2005). Template-based mesh completion. In Proceedings of the Symposium on Geometry Processing (SGP)
2005
-
[86]
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)
2012
-
[87]
Kwak, M., Song, J., and Kim, S. (2023). GeCoNeRF: Few-shot neural radiance fields via geometric consistency. arXiv e-prints , page arXiv:2301.10941. 141
2023 arXiv
-
[88]
X., Zhang, R., Ebert, F., Abbeel, P., Finn, C., and Levine, S
Lee, A. X., Zhang, R., Ebert, F., Abbeel, P., Finn, C., and Levine, S. (2018). Stochastic adversarial video prediction. arXiv e-prints , page arXiv:1804.01523
2018 arXiv
-
[89]
Lee, S., Choi, J., Kim, S., Kim, I.-J., and Cho, J. (2023a). ExtremeNeRF: Few- shot neural radiance fields under unconstrained illumination. arXiv e-prints , page arXiv:2303.11728
2023 arXiv
-
[90]
and Lee, J
Lee, S. and Lee, J. (2024). PoseDiff: Pose-conditioned multimodal diffusion model for unbounded scene synthesis from sparse inputs. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV)
2024
-
[91]
W., Won, D., and Kim, S
Lee, S., Oh, S. W., Won, D., and Kim, S. J. (2019). Copy-and-paste networks for deep video inpainting. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2019
-
[92]
Lee, Y.-C., Zhang, Z., Blackburn-Matzen, K., Niklaus, S., Zhang, J., Huang, J.- B., and Liu, F. (2023b). Fast view synthesis of casual videos. arXiv e-prints , page arXiv:2312.02135
2023 arXiv
-
[93]
Leiby, A. (2016). Interleaved reprojection now enabled for all applica- tions by default. https://steamcommunity.com/app/358720/discussions/0/ 385429254937377076/. Accessed: 12-October-2021
2016
-
[94]
and Hanrahan, P
Levoy, M. and Hanrahan, P. (1996). Light field rendering. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
1996
-
[95]
Li, D., Huang, S.-S., Shen, T., and Huang, H. (2023). Dynamic view synthesis with spatio-temporal feature warping from sparse views. In Proceedings of the ACM International Conference on Multimedia (ACM-MM)
2023
-
[96]
Li, J., Feng, Z., She, Q., Ding, H., Wang, C., and Lee, G. H. (2021a). MINE: Towards continuous depth mpi with nerf for novel view synthesis. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) . 142 Chapter 8. Bibliography
2021
-
[97]
Li, J., Zhang, J., Bai, X., Zheng, J., Ning, X., Zhou, J., and Gu, L. (2024). DNGaus- sian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth nor- malization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2024
-
[98]
Li, T., Slavcheva, M., Zollhöfer, M., Green, S., Lassner, C., Kim, C., Schmidt, T., Lovegrove, S., Goesele, M., Newcombe, R., and Lv, Z. (2022). Neural 3D video synthesis from multi-view video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[99]
Li, Z., Niklaus, S., Snavely, N., and Wang, O. (2021b). Neural scene flow fields for space-time view synthesis of dynamic scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2021
-
[100]
Lin, K.-E., Lin, Y.-C., Lai, W.-S., Lin, T.-Y., Shih, Y.-C., and Ramamoorthi, R. (2023). Vision transformer for NeRF-based view synthesis from a single input image. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV)
2023
-
[101]
Lin, K.-E., Xiao, L., Liu, F., Yang, G., and Ramamoorthi, R. (2021). Deep 3d mask volume for view synthesis of dynamic scenes. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2021
-
[102]
Liu, B., Chen, Y., Liu, S., and Kim, H.-S. (2021). Deep learning in latent space for video prediction and compression. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2021
-
[103]
A., Shih, K
Liu, G., Reda, F. A., Shih, K. J., Wang, T.-C., Tao, A., and Catanzaro, B. (2018a). Image inpainting for irregular holes using partial convolutions. In Proceedings of the European Conference on Computer Vision (ECCV)
2018
-
[104]
J., Keppo, J., Shan, Y., Qie, X., and Shou, M
Liu, J.-W., Cao, Y.-P., Mao, W., Zhang, W., Zhang, D. J., Keppo, J., Shan, Y., Qie, X., and Shou, M. Z. (2022a). Devrf: Fast deformable voxel radiance fields for 143 dynamic scenes. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)
2022
-
[105]
Liu, L., Zhang, J., He, R., Liu, Y., Wang, Y., Tai, Y., Luo, D., Wang, C., Li, J., and Huang, F. (2020). Learning by analogy: Reliable supervision from transformations for unsupervised optical flow estimation. In Proceedings of the IEEE Conference on Computer Vision and Patter...
2020
-
[106]
Liu, W., Luo, W., Lian, D., and Gao, S. (2018b). Future frame prediction for anomaly detection - a new baseline. In Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR)
2018
-
[107]
Liu, X., Kao, S.-h., Chen, J., Tai, Y.-W., and Tang, C.-K. (2023). Deceptive- NeRF: Enhancing NeRF reconstruction using pseudo-observations from diffusion mod- els. arXiv e-prints , page arXiv:2305.15171
2023 arXiv
-
[108]
Liu, Y., Peng, S., Liu, L., Wang, Q., Wang, P., Theobalt, C., Zhou, X., and Wang, W. (2022b). Neural rays for occlusion-aware image-based rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[109]
Loper, M., Mahmood, N., Romero, J., Pons-Moll, G., and Black, M. J. (2015). SMPL: A skinned multi-person linear model. ACM Transactions on Graphics (TOG) , 34(6)
2015
-
[110]
Lotter, W., Kreiman, G., and Cox, D. (2017). Deep predictive coding networks for video prediction and unsupervised learning. In Proceedings of the International Conference on Learning Representations (ICLR)
2017
-
[111]
Lowe, D. G. (2004). Distinctive image features from scale-invariant keypoints. International Journal of Computer Vision (IJCV) , 60:91–110
2004
-
[112]
Lucas, B. D. and Kanade, T. (1981). An iterative image registration technique with an application to stereo vision. 144 Chapter 8. Bibliography
1981
-
[113]
Luo, G., Zhu, Y., Li, Z., and Zhang, L. (2016). A hole filling approach based on background reconstruction for view synthesis in 3d video. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[114]
Luo, G., Zhu, Y., Weng, Z., and Li, Z. (2020). A disocclusion inpainting framework for depth-based view synthesis. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , 42(6):1289–1302
2020
-
[115]
Mark, W. R. (1999). Post-Rendering 3D Image Warping: Visibility, Reconstruc- tion, and Performance for Depth-Image Warping . The University of North Carolina at Chapel Hill
1999
-
[116]
Martin-Brualla, R., Radwan, N., Sajjadi, M. S. M., Barron, J. T., Dosovitskiy, A., and Duckworth, D. (2021). NeRF in the wild: Neural radiance fields for unconstrained photo collections. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2021
-
[117]
Mathieu, M., Couprie, C., and LeCun, Y. (2016). Deep multi-scale video prediction beyond mean square error. In Proceedings of the International Conference on Learning Representations (ICLR)
2016
-
[118]
and Bishop, G
McMillan, L. and Bishop, G. (1995). Plenoptic modeling: An image-based ren- dering system. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
1995
-
[119]
McMillan Jr, L. (1997). An Image-Based Approach to Three-Dimensional Computer Graphics. The University of North Carolina at Chapel Hill
1997
-
[120]
Meister, S., Hur, J., and Roth, S. (2018). UnFlow: Unsupervised learning of optical flow with a bidirectional census loss. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI)
2018
-
[121]
A., and Marschner, S
Miguel, E., Bradley, D., Thomaszewski, B., Bickel, B., Matusik, W., Otaduy, 145 M. A., and Marschner, S. (2012). Data-driven estimation of cloth simulation models. Computer Graphics Forum (CGF) , 31(2pt2):519–528
2012
-
[122]
P., Ortiz-Cayon, R., Kalantari, N
Mildenhall, B., Srinivasan, P. P., Ortiz-Cayon, R., Kalantari, N. K., Ramamoorthi, R., Ng, R., and Kar, A. (2019). Local light field fusion: Practical view synthesis with prescriptive sampling guidelines. ACM Transactions on Graphics (TOG) , 38(4):1–14
2019
-
[123]
P., Tancik, M., Barron, J
Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R. (2020). Nerf: Representing scenes as neural radiance fields for view synthesis. In Proceedings of the European Conference on Computer Vision (ECCV)
2020
-
[124]
G., Kelly, J., Brubaker, M
Mirzaei, A., Aumentado-Armstrong, T., Derpanis, K. G., Kelly, J., Brubaker, M. A., Gilitschenski, I., and Levinshtein, A. (2023). Spin-nerf: Multiview segmen- tation and perceptual inpainting with neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vi...
2023
-
[125]
Müller, T., Evans, A., Schied, C., and Keller, A. (2022). Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (TOG), 41(4):1–15
2022
-
[126]
Nazeri, K., Ng, E., Joseph, T., Qureshi, F., and Ebrahimi, M. (2019). EdgeCon- nect: Structure guided image inpainting using edge prediction. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) Workshop
2019
-
[127]
Ni, Z., Yang, P., Yang, W., Wang, H., Ma, L., and Kwong, S. (2024). ColNeRF: Collaboration for generalizable sparse input neural radiance field. Proceedings of the AAAI Conference on Artificial Intelligence , 38(5):4325–4333
2024
-
[128]
T., Mildenhall, B., Sajjadi, M
Niemeyer, M., Barron, J. T., Mildenhall, B., Sajjadi, M. S. M., Geiger, A., and Radwan, N. (2022). RegNeRF: Regularizing neural radiance fields for view synthesis from sparse inputs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . ...
2022
-
[129]
Nunes, H., Uzunhan, Y., Gille, T., Lamberto, C., Valeyre, D., and Brillet, P.-Y. (2012). Imaging of sarcoidosis of the airways and lung parenchyma and correlation with lung function. European Respiratory Journal, 40(3):750–765
2012
-
[130]
A., Orts- Escolano, S., Garcia-Rodriguez, J., and Argyros, A
Oprea, S., Martinez-Gonzalez, P., Garcia-Garcia, A., Castro-Vargas, J. A., Orts- Escolano, S., Garcia-Rodriguez, J., and Argyros, A. (2020). A review on deep learning techniques for video prediction. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
2020
-
[131]
T., and Martin-Brualla, R
Park, K., Henzler, P., Mildenhall, B., Barron, J. T., and Martin-Brualla, R. (2023). CamP: Camera preconditioning for neural radiance fields. ACM Transactions on Graphics (TOG) , 42(6)
2023
-
[132]
T., Bouaziz, S., Goldman, D
Park, K., Sinha, U., Hedman, P., Barron, J. T., Bouaziz, S., Goldman, D. B., Martin-Brualla, R., and Seitz, S. M. (2021). HyperNeRF: A higher-dimensional representation for topologically varying neural radiance fields. arXiv e-prints , page arXiv:2106.13228
2021 arXiv
-
[133]
Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., and Efros, A. A. (2016). Context encoders: Feature learning by inpainting. In Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[134]
and Zhang, L
Penner, E. and Zhang, L. (2017). Soft 3D reconstruction for view synthesis. ACM Transactions on Graphics (TOG) , 36(6):1–11
2017
-
[135]
and Deschaintre, V
Philip, J. and Deschaintre, V. (2023). Floaters no more: Radiance field gradi- ent scaling for improved near-camera training. In Proceedings of the Eurographics Symposium on Rendering
2023
-
[136]
Pons-Moll, G., Pujades, S., Hu, S., and Black, M. J. (2017). ClothCap: Seamless 4d clothing capture and retargeting. ACM Transactions on Graphics (TOG) , 36(4)
2017
-
[137]
Prinzler, M., Hilliges, O., and Thies, J. (2023). DINER: Depth-aware image-based 147 neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[138]
Pumarola, A., Corona, E., Pons-Moll, G., and Moreno-Noguer, F. (2021). D- NeRF: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2021
-
[139]
Ramamoorthi, R. (2023). NeRFs: The search for the best 3D representation. arXiv e-prints, page arXiv:2308.02751
2023 arXiv
-
[140]
Reiser, C., Peng, S., Liao, Y., and Geiger, A. (2021). KiloNeRF: Speeding up neural radiance fields with thousands of tiny MLPs. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2021
-
[141]
Revaud, J., De Souza, C., Humenberger, M., and Weinzaepfel, P. (2019). R2D2: Reliable and repeatable detector and descriptor. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)
2019
-
[142]
T., Mildenhall, B., Srinivasan, P
Roessle, B., Barron, J. T., Mildenhall, B., Srinivasan, P. P., and Nießner, M. (2022). Dense depth priors for neural radiance fields from sparse input views. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[143]
Sabater, N., Boisson, G., Vandame, B., Kerbiriou, P., Babon, F., Hog, M., Gendrot, R., Langlois, T., Bureller, O., Schubert, A., and Allie, V. (2017). Dataset and pipeline for multi-view light-field video. In Proceedings of the IEEE Conference on Computer Vision and Pattern Re...
2017
-
[144]
Sarkar, M., Ghose, D., and Bala, A. (2021). Decomposing camera and object motion for an improved video sequence prediction. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS) Workshop on Pre-registration in Machine Learning
2021
-
[145]
Schonberger, J. L. and Frahm, J.-M. (2016). Structure-from-motion revisited. In 148 Chapter 8. Bibliography Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[146]
Sederberg, T. W. and Parry, S. R. (1986). Free-form deformation of solid geometric models. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
1986
-
[147]
Seo, S., Chang, Y., and Kwak, N. (2023a). FlipNeRF: Flipped reflection rays for few-shot novel view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2023
-
[148]
Seo, S., Han, D., Chang, Y., and Kwak, N. (2023b). MixNeRF: Modeling a ray with mixture density for novel view synthesis from sparse inputs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[149]
Shade, J., Gortler, S., He, L.-w., and Szeliski, R. (1998). Layered depth images. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
1998
-
[150]
Shaw, R., Song, J., Moreau, A., Nazarczuk, M., Catley-Chandar, S., Dhamo, H., and Perez-Pellitero, E. (2023). Swags: Sampling windows adaptively for dynamic 3d gaussian splatting. arXiv e-prints , page arXiv:2312.13308
2023 arXiv
-
[151]
Shi, R., Wei, X., Wang, C., and Su, H. (2024). ZeroRF: Fast sparse view 360 ◦ reconstruction with zero pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2024
-
[152]
Shi, Y., Rong, D., Ni, B., Chen, C., and Zhang, W. (2022). GARF: Geometry- aware generalized neural radiance field. arXiv e-prints , page arXiv:2212.02280
2022 arXiv
-
[153]
Shih, M.-L., Su, S.-Y., Kopf, J., and Huang, J.-B. (2020). 3d photography using context-aware layered depth inpainting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 149
2020
-
[154]
Smolic, A., Mueller, K., Merkle, P., Fehn, C., Kauff, P., Eisert, P., and Wiegand, T. (2006). 3D video and free viewpoint video - technologies, applications and MPEG standards. In Proceedings of the IEEE International Conference on Multimedia and Expo (ICME)
2006
-
[155]
H., and Soundararajan, R
Somraj, N., Choudhary, K., Mupparaju, S. H., and Soundararajan, R. (2024a). Factorized motion fields for fast sparse input dynamic view synthesis. In Proceedings of the ACM Special Interest Group on Computer Graphics and Interactive Techniques (SIGGRAPH)
2024
-
[156]
Somraj, N., Karanayil, A., and Soundararajan, R. (2023). SimpleNeRF: Regular- izing sparse input neural radiance fields with simpler solutions. In Proceedings of the ACM Special Interest Group on Computer Graphics and Interactive Techniques - Asia (SIGGRAPH-Asia)
2023
-
[157]
H., Karanayil, A., and Soundararajan, R
Somraj, N., Mupparaju, S. H., Karanayil, A., and Soundararajan, R. (2024b). Simple-RF: Regularizing sparse input radiance fields with simpler solutions. arXiv e-prints
2024
-
[158]
Somraj, N., Sancheti, P., and Soundararajan, R. (2022). Temporal view synthesis of dynamic scenes through 3d object motion estimation with multi-plane images. In Proceedings of the IEEE International Symposium on Mixed and Augmented Reality (ISMAR)
2022
-
[159]
and Soundararajan, R
Somraj, N. and Soundararajan, R. (2023). ViP-NeRF: Visibility prior for sparse input neural radiance fields. In Proceedings of the ACM Special Interest Group on Computer Graphics and Interactive Techniques (SIGGRAPH)
2023
-
[160]
and Bovik, A
Soundararajan, R. and Bovik, A. C. (2013). Video quality assessment by reduced reference spatio-temporal entropic differencing. IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) , 23(4):684–694. 150 Chapter 8. Bibliography
2013
-
[161]
P., Tucker, R., Barron, J
Srinivasan, P. P., Tucker, R., Barron, J. T., Ramamoorthi, R., Ng, R., and Snavely, N. (2019). Pushing the boundaries of view extrapolation with multiplane images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2019
-
[162]
Srivastava, N., Mansimov, E., and Salakhudinov, R. (2015). Unsupervised learning of video representations using LSTMs. In Proceedings of the International Conference on Machine Learning (ICML)
2015
-
[163]
Stoll, C., Hasler, N., Gall, J., Seidel, H.-P., and Theobalt, C. (2011). Fast articu- lated motion tracking using a sums of Gaussians body model. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2011
-
[164]
Straka, Z., Svoboda, T., and Hoffmann, M. (2020). PreCNet: Next frame video prediction based on predictive coding. arXiv e-prints , page arXiv:2004.14878
2020 arXiv
-
[165]
Sun, C., Sun, M., and Chen, H.-T. (2022). Direct voxel grid optimization: Super- fast convergence for radiance fields reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
-
[166]
Sun, D., Yang, X., Liu, M.-Y., and Kautz, J. (2018). PWC-Net: CNNs for optical flow using pyramid, warping, and cost volume. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2018
-
[167]
C., Chui, S
Sun, W., Xu, L., Au, O. C., Chui, S. H., and Kwok, C. W. (2010). An overview of free view-point depth-image-based rendering (DIBR). In Proceedings of the APSIPA Annual Summit and Conference
2010
-
[168]
P., Barron, J
Tancik, M., Mildenhall, B., Wang, T., Schmidt, D., Srinivasan, P. P., Barron, J. T., and Ng, R. (2021). Learned initializations for optimizing coordinate-based neural representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 151
2021
-
[169]
and Deng, J
Teed, Z. and Deng, J. (2020). RAFT: Recurrent all-pairs field transforms for optical flow. In Proceedings of the European Conference on Computer Vision (ECCV)
2020
-
[170]
and Deng, J
Teed, Z. and Deng, J. (2021). RAFT-3D: Scene flow using rigid-motion embed- dings. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition (CVPR)
2021
-
[171]
and Yang, B
Trevithick, A. and Yang, B. (2021). GRF: Learning a general radiance field for 3D representation and rendering. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2021
-
[172]
and Snavely, N
Tucker, R. and Snavely, N. (2020). Single-view view synthesis with multiplane images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2020
-
[173]
Tulsiani, S., Tucker, R., and Snavely, N. (2018). Layer-structured 3d scene inference via view synthesis. In Proceedings of the European Conference on Computer Vision (ECCV)
2018
-
[174]
Tulyakov, S., Liu, M.-Y., Yang, X., and Kautz, J. (2018). MoCoGAN: Decompos- ing motion and content for video generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2018
-
[175]
Ulyanov, D., Vedaldi, A., and Lempitsky, V. (2018). Deep image prior. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2018
-
[176]
A., Martin-Brualla, R., Guibas, L., and Li, K
Uy, M. A., Martin-Brualla, R., Guibas, L., and Li, K. (2023). SCADE: NeRFs from space carving with ambiguity-aware depth estimates. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[177]
van Waveren, J. M. P. (2016). The asynchronous time warp for virtual reality on consumer hardware. In Proceedings of the ACM Conference on Virtual Reality Software and Technology. 152 Chapter 8. Bibliography
2016
-
[178]
Vedula, S., Baker, S., Seitz, S., and Kanade, T. (2000). Shape and motion carving in 6D. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2000
-
[179]
V., and Lee, H
Villegas, R., Pathak, A., Kannan, H., Erhan, D., Le, Q. V., and Lee, H. (2019). High fidelity video prediction with large stochastic recurrent neural networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)
2019
-
[180]
Villegas, R., Yang, J., Hong, S., Lin, X., and Lee, H. (2017). Decomposing motion and content for natural video sequence prediction. In Proceedings of the International Conference on Learning Representations (ICLR)
2017
-
[181]
Vlasic, D., Baran, I., Matusik, W., and Popović, J. (2008). Articulated mesh animation from multi-view silhouettes. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
2008
-
[182]
Vlasic, D., Peers, P., Baran, I., Debevec, P., Popović, J., Rusinkiewicz, S., and Matusik, W. (2009). Dynamic shape capture using multi-view photometric stereo. In Proceedings of the ACM Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
2009
-
[183]
Wang, C., Eckart, B., Lucey, S., and Gallo, O. (2021a). Neural trajectory fields for dynamic novel view synthesis. arXiv e-prints , page arXiv:2105.05994
2021 arXiv
-
[184]
E., Jeni, L
Wang, C., MacDonald, L. E., Jeni, L. A., and Lucey, S. (2023a). Flow supervision for deformable nerf. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[185]
C., and Liu, Z
Wang, G., Chen, Z., Loy, C. C., and Liu, Z. (2023b). SparseNeRF: Distilling depth ranking for few-shot novel view synthesis. arXiv e-prints , page arXiv:2303.16196
2023 arXiv
-
[186]
Wang, H., Liao, M., Zhang, Q., Yang, R., and Turk, G. (2009). Physically guided liquid surface modeling from videos. ACM Transactions on Graphics (TOG) , 28(3). 153
2009
-
[187]
Wang, P., Chen, X., Chen, T., Venugopalan, S., Wang, Z., et al. (2022). Is attention all that nerf needs? arXiv e-prints , page arXiv:2207.13298
2022 arXiv
-
[188]
Wang, Q., Chang, Y.-Y., Cai, R., Li, Z., Hariharan, B., Holynski, A., and Snavely, N. (2023c). Tracking everything everywhere all at once. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2023
-
[189]
P., Zhou, H., Barron, J
Wang, Q., Wang, Z., Genova, K., Srinivasan, P. P., Zhou, H., Barron, J. T., Martin- Brualla, R., Snavely, N., and Funkhouser, T. (2021b). IBRNet: Learning multi-view image-based rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2021
-
[190]
C., Sheikh, H
Wang, Z., Bovik, A. C., Sheikh, H. R., and Simoncelli, E. P. (2004). Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing (TIP), 13(4):600–612
2004
-
[191]
P., and Bovik, A
Wang, Z., Simoncelli, E. P., and Bovik, A. C. (2003). Multiscale structural sim- ilarity for image quality assessment. In Proceedings of the Asilomar Conference on Signals, Systems Computers
2003
-
[192]
Wexler, Y., Shechtman, E., and Irani, M. (2007). Space-time completion of video. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) , 29(3):463–476
2007
-
[193]
Wiles, O., Gkioxari, G., Szeliski, R., and Johnson, J. (2020). Synsin: End-to- end view synthesis from a single image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2020
-
[194]
Wimbauer, F., Yang, N., Rupprecht, C., and Cremers, D. (2023). Behind the scenes: Density fields for single view reconstruction. arXiv e-prints , page arXiv:2301.07668
2023 arXiv
-
[195]
Bibliography Xinggang, W
Wu, G., Yi, T., Fang, J., Xie, L., Zhang, X., Wei, W., Liu, W., Tian, Q., and 154 Chapter 8. Bibliography Xinggang, W. (2023). 4D gaussian splatting for real-time dynamic scene rendering. arXiv e-prints , page arXiv:2310.08528
2023 arXiv
-
[196]
P., Verbin, D., Barron, J
Wu, R., Mildenhall, B., Henzler, P., Park, K., Gao, R., Watson, D., Srinivasan, P. P., Verbin, D., Barron, J. T., Poole, B., et al. (2024). ReconFusion: 3D reconstruc- tion with diffusion priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2024
-
[197]
Wu, Y., Gao, R., Park, J., and Chen, Q. (2020). Future video synthesis with object motion prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2020
-
[198]
and Turmukhambetov, D
Wynn, J. and Turmukhambetov, D. (2023). DiffusioNeRF: Regularizing neural radiance fields with denoising diffusion models. arXiv e-prints, page arXiv:2302.12231
2023 arXiv
-
[199]
and Chen, J
Xing, W. and Chen, J. (2021). Temporal-MPI: Enabling multi-plane images for dynamic scene modelling via temporal basis learning. arXiv e-prints , page arXiv:2111.10533
2021 arXiv
-
[200]
Xiong, H., Muttukuru, S., Upadhyay, R., Chari, P., and Kadambi, A. (2023). SparseGS: Real-time 360 ◦ sparse view synthesis using gaussian splatting. arXiv e- prints, page arXiv:2312.00206
2023 arXiv
-
[201]
Xu, D., Jiang, Y., Wang, P., Fan, Z., Shi, H., and Wang, Z. (2022). SinNeRF: Training neural radiance fields on complex scenes from a single image. In Proceedings of the European Conference on Computer Vision (ECCV)
2022
-
[202]
Xu, R., Li, X., Zhou, B., and Loy, C. C. (2019). Deep flow-guided video inpainting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2019
-
[203]
Yang, C., Li, S., Fang, J., Liang, R., Xie, L., Zhang, X., Shen, W., and Tian, Q. (2024). GaussianObject: Just taking four images to get a high-quality 3D object with gaussian splatting. arXiv e-prints , page arXiv:2402.10259. 155
2024 arXiv
-
[204]
and Ramanan, D
Yang, G. and Ramanan, D. (2020). Upgrading optical flow to 3d scene flow through optical expansion. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2020
-
[205]
Yang, J., Pavone, M., and Wang, Y. (2023). FreeNeRF: Improving few-shot neural rendering with free frequency regularization
2023
-
[206]
and Pollefeys, M
Yang, R. and Pollefeys, M. (2003). Multi-resolution real-time stereo on commodity graphics hardware. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR)
2003
-
[207]
A., Zhang, Z., Shan, Q., and Huang, Q
Yang, Z., Ren, Z., Bautista, M. A., Zhang, Z., Shan, Q., and Huang, Q. (2022). FvOR: Robust joint shape and pose optimization for few-view object reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR)
2022
-
[208]
S., Kim, K., Gallo, O., Park, H
Yoon, J. S., Kim, K., Gallo, O., Park, H. S., and Kautz, J. (2020). Novel view synthesis of dynamic scenes with globally coherent depths from a monocular camera. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2020
-
[209]
Yu, A., Li, R., Tancik, M., Li, H., Ng, R., and Kanazawa, A. (2021a). Plenoctrees for real-time rendering of neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2021
-
[210]
Yu, A., Ye, V., Tancik, M., and Kanazawa, A. (2021b). pixelNeRF: Neural radiance fields from one or few images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2021
-
[211]
Á., Niinuma, K., and Jeni, L
Yu, H., Julin, J., Milacski, Z. Á., Niinuma, K., and Jeni, L. A. (2023). Cogs: Controllable gaussian splatting. arXiv e-prints , page arXiv:2312.05664
2023 arXiv
-
[212]
Yu, J., Lin, Z., Yang, J., Shen, X., Lu, X., and Huang, T. S. (2019). Free-form 156 Chapter 8. Bibliography image inpainting with gated convolution. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2019
-
[213]
Zhang, J., Yang, G., Tulsiani, S., and Ramanan, D. (2021). NeRS: Neural re- flectance surfaces for sparse-view 3D reconstruction in the wild. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS)
2021
-
[214]
Zhang, K., Riegler, G., Snavely, N., and Koltun, V. (2020). NeRF++: Analyzing and improving neural radiance fields. arXiv e-prints , page arXiv:2010.07492
2020 arXiv
-
[215]
A., Shechtman, E., and Wang, O
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. (2018). The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2018
-
[216]
Zhou, T., Tucker, R., Flynn, J., Fyffe, G., and Snavely, N. (2018). Stereo mag- nification: Learning view synthesis using multiplane images. ACM Transactions on Graphics (TOG) , 37(4)
2018
-
[217]
Zhou, T., Tulsiani, S., Sun, W., Malik, J., and Efros, A. A. (2016). View synthesis by appearance flow. In Proceedings of the European Conference on Computer Vision (ECCV)
2016
-
[218]
and Tulsiani, S
Zhou, Z. and Tulsiani, S. (2022). SparseFusion: Distilling view-conditioned diffu- sion for 3d reconstruction. arXiv e-prints , page arXiv:2212.00792
2022 arXiv
-
[219]
Zhu, B., Yang, Y., Wang, X., Zheng, Y., and Guibas, L. (2023a). VDN-NeRF: Re- solving shape-radiance ambiguity via view-dependence normalization. arXiv e-prints , page arXiv:2303.17968
2023 arXiv
-
[220]
Zhu, H., He, T., Li, X., Li, B., and Chen, Z. (2024). Is vanilla mlp in neural radiance field enough for few-shot view synthesis? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 157
2024
-
[221]
Zhu, Z., Fan, Z., Jiang, Y., and Wang, Z. (2023b). FSGS: Real-time few-shot view synthesis using gaussian splatting. arXiv e-prints , page arXiv:2312.00451
2023 arXiv
-
[222]
L., Kang, S
Zitnick, C. L., Kang, S. B., Uyttendaele, M., Winder, S., and Szeliski, R. (2004). High-quality video view interpolation using a layered representation. ACM Transac- tions on Graphics (TOG) , 23(3):600–608
2004
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.