REVIEW 5 major objections 6 minor 66 references
Dual-Camera All-in-Focus Neural Radiance Fields
T0 review · 5 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read A smartphone's second camera yields all-in-focus 3D scenes from blurry video-like captures.
desk verdict A genuine new application of dual-camera smartphone input for consistent-defocus NeRF, but the all-in-focus claim is conditional on the unmeasured sharpness of the ultra-wide camera, and the evaluation is too thin to fully support it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the defocus-aware fusion module, which turns dual-camera fusion into a learned blending problem. It renders a bokeh image from the aligned ultra-wide radiance field using a scatter operation whose circle-of-confusion radius is $r = a f |D_f - D|$, learns the blur intensity $A = af$ and focused disparity $D_f$ by matching the rendered bokeh color to the main-camera color with an SSIM loss, and computes a defocus map $D_{\text{defocus}} = A|D - D_f|$ from the NeRF disparity. That map supervises an MLP blending mask $\eta$ that weights the ultra-wide and main-camera fields during weighted volume rendering, so the sharp ultra-wide content replaces only the blurred main-camera regions. The alignment pipeline — SIFT-based homography, RAFT optical flow, a forward-backward consistency check, and histogram matching — makes the fusion meaningful across the two physically different cameras.
What would settle it
Capture a scene where the ultra-wide camera itself shows visible defocus in a region where the main camera is also blurred, such as a low-light close-up that exceeds the ultra-wide's depth of field, and check whether DC-NeRF renders that region all-in-focus; if the region stays soft, the claim that the ultra-wide camera supplies the missing sharp reference is refuted.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that view-consistent defocus blur — the case where every input view is blurred in the same places because the camera never refocuses — can be resolved by replacing the single-camera assumption of NeRF with a two-camera model. DC-NeRF takes a high-resolution main image and a lower-resolution ultra-wide image of the same view, aligns them with homography, optical flow, and histogram matching, then learns defocus parameters (blur intensity and focal disparity) by rendering bokeh from the ultra-wide field and matching it to the main-camera colors. A blending weight field predicted from the resulting defocus map fuses the sharp, deep-focus ultra-wide radiance field with the high-detail main-camera field inside volume rendering, producing all-in-focus novel views. The same learned parameters let the user refocus the rendered scene or apply split-diopter effects.
Load-bearing premise
The load-bearing premise is that the ultra-wide camera's image is sharp in every region where the main camera is blurred; if the ultra-wide shot is itself blurry or too low quality in a region, the fusion has no sharp source to draw from and the all-in-focus result degrades.
Editorial extensions
If this is right
- Users can capture an all-in-focus, multi-view-consistent NeRF with a standard smartphone without manually refocusing the main camera between shots.
- The learned defocus parameters enable post-capture refocusing, letting the focal plane be shifted to any disparity plane and the blur intensity be adjusted.
- Split-diopter-style rendering, where foreground and background stay sharp while the middle region blurs, becomes a byproduct of the same representation.
- Because DC-NeRF trains directly on the target scene, it avoids the extra defocus-deblurring dataset and pretraining that single-view dual-camera deblurring baselines require.
- Existing NeRF-based deblurring methods that assume inconsistent blur across views fail on consistent blur, so the dual-camera fusion is a necessary ingredient rather than a modest improvement.
Reading between the lines
- This suggests the same align-and-fuse design could be adapted to other paired-camera configurations, such as dual-pixel sensors or wide-plus-tele setups, whenever one camera's depth of field reliably exceeds the other's.
- Because the fusion mask is supervised only by the defocus map, the method may implicitly learn a scene-depth prior that could be reused for editing focal effects or estimating disparity from novel viewpoints.
- A natural stress test is a low-light scene where the ultra-wide image is noisy or itself blurred; the paper's own failure case already hints that such a boundary would break the assumption of a sharp ultra-wide reference.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DC-NeRF, a NeRF-based novel-view synthesis framework that inputs a smartphone dual-camera sequence: a main camera with consistent defocus blur (autofocus on the same target in all views) and an aligned ultra-wide camera. The method aligns the ultra-wide image to the main view via homography, optical flow, and histogram matching; trains separate NeRFs for the main and aligned ultra-wide images; learns defocus parameters A and D_f by fitting a scatter-based bokeh renderer to the main-camera observation; and fuses the two radiance fields with a learned blending mask to produce all-in-focus novel views. The paper also contributes a 7-scene smartphone dataset with focal-stack-derived ground truth, ablations of alignment and fusion components, and demonstrations of refocusing and split-diopter effects. The headline claim is that this is the first framework to synthesize all-in-focus NeRF from inputs without manual refocusing under view-consistent defocus blur.
Significance. If the claims hold, this is a well-motivated and useful contribution: it addresses a genuinely under-explored failure mode of NeRF (view-consistent defocus blur), exploits commodity smartphone dual-camera hardware, and the align-and-fuse pipeline is coherent and supported by component ablations. The dataset is a valuable resource, and the DoF applications show practical potential. The paper is honest about its failure cases and computational cost. However, the all-in-focus claim is stronger than the evidence: the ground-truth AiF images are synthesized from a focal stack by an external fusion method; the ultra-wide camera's sharpness is never measured and is explicitly acknowledged to fail in some regions; and the quantitative gains over strong baselines are marginal and not statistically characterized. These issues affect the central claim rather than the peripheral presentation.
major comments (5)
- [4.1, Eq. (4)] The AiF ground truth is synthesized by multi-focus image fusion of only two focal-plane main images; it is not a truly sharp sensor capture. Any artifacts of the external fusion method (e.g., focus-breathing misalignment, halos) are treated as ground truth, and the fusion objective in Eq. (17) is a similar mask-based blending of the two cameras. This creates an evaluation bias toward mask-based fusion and does not directly measure the actual sharpness of the output. Please report a sensitivity analysis with an alternative GT fusion method, or evaluate on scenes with independently known all-in-focus captures.
- [5.7, Fig. 18] The failure case states that 'the ultra-wide views may have low-quality regions on some details, so the corresponding parts of the main camera cannot be fixed.' This is exactly the condition on which the central all-in-focus claim rests: in Eq. (17), Cfuse is a weighted combination of C_w and C_m, so when C_w is blurred in a region, the learned mask has no sharp source to draw from and no loss term can recover the missing detail. The manuscript never measures the residual blur of the ultra-wide camera, and the failure case shows the assumption can fail in practice. Please add a quantitative sharpness/MTF measurement of the UWA inputs and an analysis of how often/where the assumption is violated, or explicitly re-scope the claim to 'all-in-focus when the ultra-wide camera is sharp in the region of interest.'
- [5.3, Tables 1-3] The quantitative support for the headline claim is thin. In Table 1, DC-NeRF's average PSNR advantage over EasyAIF+NeRF is 0.07 dB (24.31 vs 24.24), and on Dove DC-NeRF has lower PSNR (22.72 vs 22.84); in Table 2, DC-NeRF is not the best on average LPIPS (0.209 vs Deblur-NeRF's 0.184); and in Table 3 the per-scene PSNR is worse than Deblur-NeRF on Stadium (24.53 vs 25.38). All numbers appear to come from a single training run with no variance estimates. Please report standard deviations over multiple seeds/runs and use a statistical test or consistent per-scene wins to support 'compares favorably.'
- [4.3, Eq. (12)] The transmittance T_f(t_i) is defined as exp(-Σ_{j<i} σ_m σ_w δ_j). This is not the standard volume-rendering transmittance for two radiance fields, which should involve a sum of densities (or a weighted sum) in the exponential; a product of densities is dimensionally inconsistent and would make occlusion depend on both fields being non-zero. Please either correct the equation to the actual implementation (e.g., T_f = exp(-Σ(η σ_w + (1-η)σ_m)δ) or similar) or explain and justify the product form. This is central to the fused rendering.
- [4.3, Eqs. (10), (15), (17)] The defocus map is learned by fitting a bokeh renderer to the blurred main camera observation, and this same map defines the fusion target. Because the supervision in Eq. (15) only compares rendered bokeh to the already-blurred main ray, errors in A and D_f are not independently checked; the fusion target in Eq. (17) then inherits these errors. This is a self-supervised fitting loop rather than an independent defocus measurement. Please validate the estimated defocus map against a measured defocus/disparity (e.g., from the focal stack or an external depth sensor) or show an ablation that uses the ground-truth defocus map instead.
minor comments (6)
- [Introduction] The word 'improficiencies' appears to be a typo for 'deficiencies'.
- [4.2, Eq. (5)] Equation (5) is garbled in the manuscript; the confidence-map formula should be typeset with clear warping notation and thresholding so that the forward-backward consistency check is unambiguous.
- [4.3, Eq. (11)] The disparity formula is difficult to read because of the formatting; please rewrite it with explicit parentheses, e.g., D = 1 / (Σ_i T_i (1-exp(-σ δ_i)) t_i).
- [4.3, Eqs. (12)-(13)] The term 'volumef' is used before it is defined; please define the fused volume color explicitly before presenting the rendering equation.
- [5.2, Figs. 8-11] The main text says that the inputs of baselines are only the main-camera images for Figs. 8 and 11, while Figs. 9 and 10 show baselines trained with ultra-wide inputs; please state the input setting in each figure caption explicitly to avoid confusion.
- [5.3, Table 2] The caption acknowledges that Deblur-NeRF achieves better LPIPS and asks readers to consult qualitative results; please report the magnitude of the difference and ideally confidence intervals so the trade-off is interpretable.
Circularity Check
Defocus parameters are fit to the blurred main camera and then reused to define the fusion-supervision target; the final claim still has independent external evaluation.
-
fitted input called prediction
[Sec. 4.3-4.4, Eqs. (10), (15), (17)]
"To learn the defocus parameters A and Df , we use the ultra-wide ray and the parameters as inputs, sampling its surrounding area and implementing bokeh rendering as mentioned. We use the shallow DoF rays from the main camera as supervision for the defocus parameters estimation. ... Cfuse = Mblend·Cw + (1−Mblend)·Cm. Mblend is the defocus map Ddefocus with normalization to 0 and 1."
The defocus parameters A and Df are fit by making a bokeh rendering of the ultra-wide radiance field match the blurred main-camera rays (Eq. 15). The same fitted A and Df define the defocus map Ddefocus = A|D−Df| (Eq. 10); after normalization this map becomes Mblend, which constructs the fusion-supervision target Cfuse = Mblend·Cw + (1−Mblend)·Cm (Eq. 17). The final render is trained with Lfusion against Cfuse, so the 'all-in-focus' output is, by construction, a blend whose mask is the fitted defocus estimate. No independent measurement of defocus or sharpness is introduced; the system's ability to restore is bounded by the ultra-wide content and this fitted mask. The paper's own Sec.
full rationale
DC-NeRF is an engineering method, not a derived law, and its final novel views are measured against an independently synthesized focal-stack ground truth (Eq. 4), so the central empirical claim is not internally forced. The self-supervised loop at Eqs. (10), (15), and (17) is the main circularity-relevant step: the defocus map is not independently measured but is the output of fitting A and Df to the blurred main camera, and that same map then defines the fusion target used to supervise the final render. This makes the 'prediction' of which regions need ultra-wide replacement statistically dependent on the blurred input. However, the evaluation compares against real captured all-in-focus targets rather than the method's own Cfuse, and the paper transparently reports failure cases where ultra-wide low quality cannot be fixed. No load-bearing self-citation or imported uniqueness theorem is present; references to the authors' prior bokeh and EasyAIF work are methodological or comparative, not justifications of the central claim. Overall, the circularity is partial and confined to the training supervision design, not to the reported benchmark comparison.
Assumptions & free parameters
free parameters (3)
- A =
learned, initialized to 5
- D_f =
learned, initialized to 0.5
- beta (bokeh kernel smoothness) =
4
assumptions (3)
- domain assumption The ultra-wide camera image is sufficiently sharp to serve as an all-in-focus reference for all blurred regions of the main camera.
- domain assumption Homography plus optical flow plus histogram matching can align the ultra-wide image to the main image well enough for per-pixel fusion.
- domain assumption The NeRF-estimated disparity map (Eq 11) is accurate enough to compute meaningful defocus maps and blending masks.
Cite this review
Pith. "Pith review of Dual-Camera All-in-Focus Neural Radiance Fields." pith.science (2026). https://pith.science/paper/GLHET4JY
@misc{pith2026250416636,
author = {Pith},
title = {Pith review of: Dual-Camera All-in-Focus Neural Radiance Fields},
year = {2026},
howpublished = {\url{https://pith.science/paper/GLHET4JY}},
note = {Machine review of arXiv:2504.16636}
}
read the original abstract
We present the first framework capable of synthesizing the all-in-focus neural radiance field (NeRF) from inputs without manual refocusing. Without refocusing, the camera will automatically focus on the fixed object for all views, and current NeRF methods typically using one camera fail due to the consistent defocus blur and a lack of sharp reference. To restore the all-in-focus NeRF, we introduce the dual-camera from smartphones, where the ultra-wide camera has a wider depth-of-field (DoF) and the main camera possesses a higher resolution. The dual camera pair saves the high-fidelity details from the main camera and uses the ultra-wide camera's deep DoF as reference for all-in-focus restoration. To this end, we first implement spatial warping and color matching to align the dual camera, followed by a defocus-aware fusion module with learnable defocus parameters to predict a defocus map and fuse the aligned camera pair. We also build a multi-view dataset that includes image pairs of the main and ultra-wide cameras in a smartphone. Extensive experiments on this dataset verify that our solution, termed DC-NeRF, can produce high-quality all-in-focus novel views and compares favorably against strong baselines quantitatively and qualitatively. We further show DoF applications of DC-NeRF with adjustable blur intensity and focal plane, including refocusing and split diopter.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
Modeling and rendering architecture from photographs: A hybrid geometry-and image- based approach,
P . E. Debevec, C. J. Taylor, and J. Malik, “Modeling and rendering architecture from photographs: A hybrid geometry-and image- based approach,” in Proceedings of the 23rd annual conference on Computer graphics and interactive techniques, 1996, pp. 11–20
1996
-
[2]
The lumigraph,
S. J. Gortler, R. Grzeszczuk, R. Szeliski, and M. F. Cohen, “The lumigraph,” in Proceedings of the 23rd annual conference on Computer graphics and interactive techniques, 1996, pp. 43–54
1996
-
[3]
Light field rendering,
M. Levoy and P . Hanrahan, “Light field rendering,” in Proceedings of the 23rd annual conference on Computer graphics and interactive techniques, 1996, pp. 31–42
1996
-
[4]
Nerf: Representing scenes as neural radiance fields for view synthesis,
B. Mildenhall, P . P . Srinivasan, M. Tancik, J. T. Barron, R. Ra- mamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” in Proceedings of the European Conference on Computer Vision (ECCV) . Springer, 2020, pp. 405– 421
work page 2020
-
[5]
Deblur-nerf: Neural radiance fields from blurry images,
L. Ma, X. Li, J. Liao, Q. Zhang, X. Wang, J. Wang, and P . V . Sander, “Deblur-nerf: Neural radiance fields from blurry images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 12 861–12 870. SHELL luo et al. : DUAL-CAMERA ALL-IN-FOCUS NEURAL RADIANCE FIELDS 15
work page 2022
-
[6]
Dof-nerf: Depth-of-field meets neural radiance fields,
Z. Wu, X. Li, J. Peng, H. Lu, Z. Cao, and W. Zhong, “Dof-nerf: Depth-of-field meets neural radiance fields,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 1718– 1729
work page 2022
-
[7]
Dp-nerf: Deblurred neural radiance field with physical scene priors,
D. Lee, M. Lee, C. Shin, and S. Lee, “Dp-nerf: Deblurred neural radiance field with physical scene priors,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 12 386–12 396
2023
-
[8]
A point set generation network for 3d object reconstruction from a single image,
H. Fan, H. Su, and L. J. Guibas, “A point set generation network for 3d object reconstruction from a single image,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 605–613
work page 2017
Show all 66 references
-
[9]
Learning representations and generative models for 3d point clouds,
P . Achlioptas, O. Diamanti, I. Mitliagkas, and L. Guibas, “Learning representations and generative models for 3d point clouds,” in International conference on machine learning . PMLR, 2018, pp. 40– 49
2018
-
[10]
Unsupervised learning of 3d structure from images,
D. Jimenez Rezende, S. Eslami, S. Mohamed, P . Battaglia, M. Jader- berg, and N. Heess, “Unsupervised learning of 3d structure from images,” Advances in neural information processing systems , vol. 29, pp. 4996–5004, 2016
2016
-
[11]
Deep marching cubes: Learning explicit surface representations,
Y. Liao, S. Donne, and A. Geiger, “Deep marching cubes: Learning explicit surface representations,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 2916–2925
2018
-
[12]
Pix2vox: Context- aware 3d reconstruction from single and multi-view images,
H. Xie, H. Yao, X. Sun, S. Zhou, and S. Zhang, “Pix2vox: Context- aware 3d reconstruction from single and multi-view images,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 2690–2698
2019
-
[13]
Learning category-specific mesh reconstruction from image collections,
A. Kanazawa, S. Tulsiani, A. A. Efros, and J. Malik, “Learning category-specific mesh reconstruction from image collections,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 371–386
2018
-
[14]
Generating 3d faces using convolutional mesh autoencoders,
A. Ranjan, T. Bolkart, S. Sanyal, and M. J. Black, “Generating 3d faces using convolutional mesh autoencoders,” inProceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 704–720
2018
-
[15]
Pixel2mesh: Generating 3d mesh models from single rgb im- ages,
N. Wang, Y. Zhang, Z. Li, Y. Fu, W. Liu, and Y.-G. Jiang, “Pixel2mesh: Generating 3d mesh models from single rgb im- ages,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 52–67
2018
-
[16]
Scene representation networks: Continuous 3d-structure-aware neural scene representations,
V . Sitzmann, M. Zollhoefer, and G. Wetzstein, “Scene representation networks: Continuous 3d-structure-aware neural scene representations,” in Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch ´e- Buc, E. Fox, and R. Garnet...
-
[17]
Local light field fusion: Practical view synthesis with prescriptive sampling guidelines,
B. Mildenhall, P . P . Srinivasan, R. Ortiz-Cayon, N. K. Kalantari, R. Ramamoorthi, R. Ng, and A. Kar, “Local light field fusion: Practical view synthesis with prescriptive sampling guidelines,” ACM Transactions on Graphics (TOG), vol. 38, no. 4, pp. 1–14, 2019
2019
-
[18]
Deepvoxels: Learning persistent 3d feature em- beddings,
V . Sitzmann, J. Thies, F. Heide, M. Nießner, G. Wetzstein, and M. Zollhofer, “Deepvoxels: Learning persistent 3d feature em- beddings,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 2437–2446
2019
-
[19]
Neural volumes: Learning dynamic renderable volumes from images,
S. Lombardi, T. Simon, J. Saragih, G. Schwartz, A. Lehrmann, and Y. Sheikh, “Neural volumes: Learning dynamic renderable volumes from images,” ACM Transactions on Graphics, vol. 38, no. 4CD, pp. 65.1–65.14, 2019
2019
-
[20]
Nerf--: Neural radiance fields without known camera parameters,
Z. Wang, S. Wu, W. Xie, M. Chen, and V . A. Prisacariu, “Nerf--: Neural radiance fields without known camera parameters,” arXiv preprint arXiv:2102.07064, 2021
2021 arXiv
-
[21]
Gnerf: Gan-based neural radiance field without posed camera,
Q. Meng, A. Chen, H. Luo, M. Wu, H. Su, L. Xu, X. He, and J. Yu, “Gnerf: Gan-based neural radiance field without posed camera,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 6351–6361
2021
-
[22]
Self-calibrating neural radiance fields,
Y. Jeong, S. Ahn, C. Choy, A. Anandkumar, M. Cho, and J. Park, “Self-calibrating neural radiance fields,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , Oc- tober 2021, pp. 5846–5854
2021
-
[23]
Pharr, W
M. Pharr, W. Jakob, and G. Humphreys, Physically based rendering: From theory to implementation. Morgan Kaufmann, 2016
2016
-
[24]
Virtual dslr: High quality dynamic depth-of-field synthesis on mobile platforms,
Y. Yang, H. Lin, Z. Yu, S. Paris, and J. Yu, “Virtual dslr: High quality dynamic depth-of-field synthesis on mobile platforms,” Electronic Imaging, vol. 2016, no. 18, pp. 1–9, 2016
2016
-
[25]
Rendering natural camera bokeh effect with deep learning,
A. Ignatov, J. Patel, and R. Timofte, “Rendering natural camera bokeh effect with deep learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Work- shops, 2020, pp. 418–419
2020
-
[26]
Bggan: Bokeh-glass generative adversarial network for rendering realistic bokeh,
M. Qian, C. Qiao, J. Lin, Z. Guo, C. Li, C. Leng, and J. Cheng, “Bggan: Bokeh-glass generative adversarial network for rendering realistic bokeh,” in Computer Vision–ECCV 2020 Workshops: Glas- gow, UK, August 23–28, 2020, Proceedings, Part III 16 . Springer, 2020, pp. 229–244
2020
-
[27]
Aperture supervision for monocular depth estimation,
P . P . Srinivasan, R. Garg, N. Wadhwa, R. Ng, and J. T. Barron, “Aperture supervision for monocular depth estimation,” in Pro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 6393–6401
2018
-
[28]
Defocus to focus: Photo-realistic bokeh rendering by fusing defocus and radiance priors,
X. Luo, J. Peng, K. Xian, Z. Wu, and Z. Cao, “Defocus to focus: Photo-realistic bokeh rendering by fusing defocus and radiance priors,” Information Fusion, vol. 89, pp. 320–335, 2023
2023
-
[29]
Synthetic depth-of-field with a single-camera mobile phone,
N. Wadhwa, R. Garg, D. E. Jacobs, B. E. Feldman, N. Kanazawa, R. Carroll, Y. Movshovitz-Attias, J. T. Barron, Y. Pritch, and M. Levoy, “Synthetic depth-of-field with a single-camera mobile phone,” ACM Transactions on Graphics (TOG) , vol. 37, no. 4, pp. 1–13, 2018
2018
-
[30]
Deepfocus: Learned image synthesis for computational dis- plays,
L. Xiao, A. Kaplanyan, A. Fix, M. Chapman, and D. Lanman, “Deepfocus: Learned image synthesis for computational dis- plays,” ACM Transactions on Graphics (TOG) , vol. 37, no. 6, pp. 1–13, 2018
2018
-
[31]
Deeplens: Shallow depth of field from a single image,
L. Wang, X. Shen, J. Zhang, O. Wang, Z. Lin, C.-Y. Hsieh, S. Kong, and H. Lu, “Deeplens: Shallow depth of field from a single image,” ACM Transactions on Graphics (TOG), vol. 37, no. 6, pp. 1–11, 2018
2018
-
[32]
Interactive portrait bokeh rendering system,
J. Peng, X. Luo, K. Xian, and Z. Cao, “Interactive portrait bokeh rendering system,” in Proc. IEEE International Conference on Image Processing (ICIP). IEEE, 2021, pp. 2923–2927
2021
-
[33]
Ranking-based salient object detection and depth prediction for shallow depth- of-field,
K. Xian, J. Peng, C. Zhang, H. Lu, and Z. Cao, “Ranking-based salient object detection and depth prediction for shallow depth- of-field,” sensors, vol. 21, no. 5, p. 1815, 2021
2021
-
[34]
Depth-aware blending of smoothed images for bokeh effect generation,
S. Dutta, “Depth-aware blending of smoothed images for bokeh effect generation,” Journal of Visual Communication and Image Rep- resentation, vol. 77, p. 103089, 2021
2021
-
[35]
Bokehme: When neural rendering meets classical rendering,
J. Peng, Z. Cao, X. Luo, H. Lu, K. Xian, and J. Zhang, “Bokehme: When neural rendering meets classical rendering,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2022, pp. 16 283–16 292
2022
-
[36]
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer,
R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V . Koltun, “Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer,” IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 3, pp. 1623–1637, 2020
2020
-
[37]
Nerf in the dark: High dynamic range view synthesis from noisy raw images,
B. Mildenhall, P . Hedman, R. Martin-Brualla, P . P . Srinivasan, and J. T. Barron, “Nerf in the dark: High dynamic range view synthesis from noisy raw images,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2022, pp. 16 190–16 199
2022
-
[38]
Nerfocus: Neural radiance field for 3d synthetic defocus,
Y. Wang, S. Yang, Y. Hu, and J. Zhang, “Nerfocus: Neural radiance field for 3d synthetic defocus,” arXiv preprint arXiv:2203.05189 , 2022
2022 arXiv
-
[39]
Just noticeable defocus blur detection and estimation,
J. Shi, L. Xu, and J. Jia, “Just noticeable defocus blur detection and estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2015, pp. 657–665
2015
-
[40]
A unified approach of multi-scale deep and hand-crafted features for defocus estima- tion,
J. Park, Y.-W. Tai, D. Cho, and I. So Kweon, “A unified approach of multi-scale deep and hand-crafted features for defocus estima- tion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017, pp. 1736–1745
2017
-
[41]
Dynamic video deblurring using a locally adaptive blur model,
T. H. Kim, S. Nah, and K. M. Lee, “Dynamic video deblurring using a locally adaptive blur model,” IEEE transactions on pattern analysis and machine intelligence, vol. 40, no. 10, pp. 2374–2387, 2017
2017
-
[42]
Image and depth from a conventional camera with a coded aperture,
A. Levin, R. Fergus, F. Durand, and W. T. Freeman, “Image and depth from a conventional camera with a coded aperture,” ACM transactions on graphics (TOG), vol. 26, no. 3, pp. 70–es, 2007
2007
-
[43]
Dwdn: deep wiener deconvolu- tion network for non-blind image deblurring,
J. Dong, S. Roth, and B. Schiele, “Dwdn: deep wiener deconvolu- tion network for non-blind image deblurring,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 12, pp. 9960–9976, 2021
2021
-
[44]
Defocus deblurring using dual- pixel data,
A. Abuolaim and M. S. Brown, “Defocus deblurring using dual- pixel data,” in Proceedings of the European conference on computer vision (ECCV). Springer, 2020, pp. 111–126
2020
-
[45]
Iterative filter adaptive network for single image defocus deblurring,
J. Lee, H. Son, J. Rim, S. Cho, and S. Lee, “Iterative filter adaptive network for single image defocus deblurring,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 2034–2042
2021
-
[46]
Single image defocus deblurring using kernel-sharing parallel atrous convolutions,
H. Son, J. Lee, S. Cho, and S. Lee, “Single image defocus deblurring using kernel-sharing parallel atrous convolutions,” in Proceedings 16 MANUSCRIPT SUBMITTED TO IEEE TRANS. PATTERN ANAL YSIS & MACHINE INTELLIGENCE; SEP 2023 of the IEEE/CVF International Conference on Compute...
2023
-
[47]
Learning enriched features for fast image restoration and enhancement,
S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, M.-H. Yang, and L. Shao, “Learning enriched features for fast image restoration and enhancement,” IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 2, pp. 1934–1948, 2022
1934
-
[48]
Learning to reduce defocus blur by realistically modeling dual- pixel data,
A. Abuolaim, M. Delbracio, D. Kelly, M. S. Brown, and P . Milanfar, “Learning to reduce defocus blur by realistically modeling dual- pixel data,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 2289–2298
2021
-
[49]
Dual pixel exploration: Simultaneous depth estimation and im- age restoration,
L. Pan, S. Chowdhury, R. Hartley, M. Liu, H. Zhang, and H. Li, “Dual pixel exploration: Simultaneous depth estimation and im- age restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 4340–4349
2021
-
[50]
Restormer: Efficient transformer for high-resolution image restoration,
S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang, “Restormer: Efficient transformer for high-resolution image restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5728–5739
2022
-
[51]
Point-and- shoot all-in-focus photo synthesis from smartphone camera pair,
X. Luo, J. Peng, W. Zhao, K. Xian, H. Lu, and Z. Cao, “Point-and- shoot all-in-focus photo synthesis from smartphone camera pair,” IEEE Transactions on Circuits and Systems for Video Technology, 2022
2022
-
[52]
Dc2: Dual-camera defocus control by learning to refocus,
H. Alzayer, A. Abuolaim, L. Chun Chan, Y. Yang, Y. Chen Lou, J.-B. Huang, and A. Kar, “Dc2: Dual-camera defocus control by learning to refocus,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 21 488–21 497
2023
-
[53]
Ray tracing volume densities,
J. T. Kajiya and B. P . Von Herzen, “Ray tracing volume densities,” ACM SIGGRAPH computer graphics , vol. 18, no. 3, pp. 165–174, 1984
1984
-
[54]
Raft: Recurrent all-pairs field transforms for optical flow,
Z. Teed and J. Deng, “Raft: Recurrent all-pairs field transforms for optical flow,” in Proceedings of the European conference on computer vision (ECCV). Springer, 2020, pp. 402–419
2020
-
[55]
Guided filter-based multi- focus image fusion through focus region detection,
X. Qiu, M. Li, L. Zhang, and X. Yuan, “Guided filter-based multi- focus image fusion through focus region detection,”Signal Process- ing: Image Communication, vol. 72, pp. 35–46, 2019
2019
-
[56]
Object recognition from local scale-invariant fea- tures,
D. G. Lowe, “Object recognition from local scale-invariant fea- tures,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, vol. 2, 1999, pp. 1150–1157
1999
-
[57]
Working hard to know your neighbor’s margins: Local descriptor learning loss,
A. Mishchuk, D. Mishkin, F. Radenovic, and J. Matas, “Working hard to know your neighbor’s margins: Local descriptor learning loss,” Advances in neural information processing systems , vol. 30, 2017
2017
-
[58]
Nm-net: Mining reliable neighbors for robust feature correspondences,
C. Zhao, Z. Cao, C. Li, X. Li, and J. Yang, “Nm-net: Mining reliable neighbors for robust feature correspondences,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 215–224
2019
-
[59]
Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography,
M. A. Fischler and R. C. Bolles, “Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography,” Communications of the ACM , vol. 24, no. 6, pp. 381–395, 1981
1981
-
[60]
Full flow: Optical flow estimation by global optimization over regular grids,
Q. Chen and V . Koltun, “Full flow: Optical flow estimation by global optimization over regular grids,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2016, pp. 4706–4714
2016
-
[61]
Loss functions for image restoration with neural networks,
H. Zhao, O. Gallo, I. Frosio, and J. Kautz, “Loss functions for image restoration with neural networks,” IEEE Trans. Computational Imaging , vol. 3, no. 1, pp. 47–57, 2017. [Online]. Available: https://doi.org/10.1109/TCI.2016.2644865
2017
-
[62]
Automatic differ- entiation in pytorch,
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differ- entiation in pytorch,” in Advances in Neural Information Processing Systems Workshops (NIPSW), 2017
2017
-
[63]
Structure-from-motion revis- ited,
J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revis- ited,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4104–4113
2016
-
[64]
Adam: A method for stochastic optimiza- tion,
D. P . Kingma and J. Ba, “Adam: A method for stochastic optimiza- tion,” in Proc. International Conference on Learning Representations (ICLR), 2014
2014
-
[65]
The unreasonable effectiveness of deep features as a perceptual metric,
R. Zhang, P . Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 586–595
2018
-
[2019]
Available: https://proceedings.neurips.cc/paper/ 2019/file/b5dc4e5d9b495d0196f61d45b26ef33e-Paper.pdf
[Online]. Available: https://proceedings.neurips.cc/paper/ 2019/file/b5dc4e5d9b495d0196f61d45b26ef33e-Paper.pdf
2019
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.