Pith. sign in

REVIEW 3 minor 87 references

Light field integration followed by a conditioned vision-language model restores occluded scenes with highest benchmark accuracy.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-26 18:27 UTC pith:YP7FHKW2

load-bearing objection The paper combines light field integration with a VLM prior and multi-sample fusion for occlusion removal, and the full text supplies ablations plus baseline comparisons that make the SSIM claims credible.

arxiv 2606.19985 v1 pith:YP7FHKW2 submitted 2026-06-18 cs.CV

Vision-Reasoning-Guided Occlusion Removal from Light Fields

classification cs.CV
keywords occlusion removallight fieldsvision-language modelscomputational imagingscene recoverymulti-view integrationsemantic prior
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper presents a framework that first integrates multiple light-field views to suppress dense foreground occlusions such as vegetation and produce a visibility-enhanced image. A vision-language model then acts as a semantic prior, conditioned on those integrated measurements, to recover fine details and degraded structures. A multi-sample fusion step aggregates several model outputs to improve consistency and limit hallucinated content that would contradict the observations. The method reports the highest average SSIM on four synthetic light-field benchmark scenes and maintains performance under both structured and unstructured real-world capture conditions. This combination of physical measurement constraints with language-model reasoning is positioned as useful for tasks where visibility is severely limited.

Core claim

The vision-reasoning-guided framework achieves superior occlusion removal by first applying light field integration to suppress foreground occlusions and then conditioning a vision-language model on the integrated measurements to restore fine details, with multi-sample fusion ensuring consistency and minimizing hallucinations. This yields the highest average SSIM on four synthetic light field benchmark scenes and demonstrates strong generalization to both structured and unstructured real-world acquisitions.

What carries the argument

Multi-sample fusion of hypotheses generated by a vision-language model conditioned on light-field-integrated measurements.

Load-bearing premise

A vision-language model conditioned on the integrated measurements can restore accurate fine details without generating hallucinated structures that contradict the physical observations.

What would settle it

A quantitative test in which the final fused output shows higher error or lower SSIM than the light-field-integrated image alone on scenes where the model introduces visible structures absent from all input views.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Highest average SSIM reported across the four synthetic light field benchmark scenes.
  • Strong generalization performance on both structured and unstructured real-world acquisition settings.
  • Direct applicability to search-and-rescue and exploratory robotic navigation under severe occlusion.
  • Demonstrates that physical imaging constraints can be combined with vision-language reasoning for robust perception.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same conditioning-plus-fusion pattern could be tested on other multi-view or volumetric imaging modalities where semantic priors might complement physical data.
  • Adding explicit consistency checks between VLM outputs and raw light-field measurements might further reduce residual hallucinations.
  • Performance on dynamic or time-varying occlusions remains untested and would require extending the integration step to include temporal information.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 3 minor

Summary. The manuscript proposes a vision-reasoning-guided framework for light-field occlusion removal. Multi-view observations are integrated via light-field integration (LFI) to suppress foreground occlusions, after which a vision-language model (VLM) is conditioned on the integrated measurements to restore fine details; a multi-sample fusion step aggregates multiple VLM hypotheses to improve consistency. Experiments on the 4-Syn synthetic benchmark and real structured/unstructured captures report state-of-the-art average SSIM together with strong cross-setting generalization.

Significance. If the reported SSIM gains and generalization hold under the supplied ablations and baseline comparisons, the work shows that physical light-field constraints can be productively combined with semantic priors from VLMs to handle severe natural occlusion. The multi-sample fusion mechanism directly mitigates the hallucination risk that would otherwise undermine the VLM component, and the inclusion of implementation details plus quantitative tables strengthens the central claim for applications in robotic navigation and search-and-rescue.

minor comments (3)
  1. [§3.2] §3.2: the precise conditioning mechanism (prompt template and how LFI output is tokenized for the VLM) is described at a high level; an explicit listing of the prompt components and any learned adapters would improve reproducibility.
  2. [Table 2] Table 2: the per-scene SSIM values for the 4-Syn benchmark are summarized only as averages; adding the individual scene scores would allow readers to assess whether gains are uniform or driven by particular scenes.
  3. [Figure 4] Figure 4 caption: the visual comparison panels would benefit from an explicit statement of the quantitative metric (SSIM or PSNR) shown beneath each result.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the positive summary, significance assessment, and recommendation of minor revision. The report raises no specific major comments or criticisms, so we provide no point-by-point responses.

Circularity Check

0 steps flagged

No significant circularity in derivation chain

full rationale

The manuscript describes an empirical framework that integrates light-field integration with a vision-language model and a multi-sample fusion step, then reports quantitative results on benchmark datasets. No equations, derivations, or parameter-fitting procedures are presented that reduce to self-definitions, fitted inputs renamed as predictions, or self-citation chains. All central claims rest on external experimental measurements rather than internal consistency loops, satisfying the self-contained criterion.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Abstract-only input supplies no explicit free parameters, axioms, or invented entities.

pith-pipeline@v0.9.1-grok · 5718 in / 980 out tokens · 36226 ms · 2026-06-26T18:27:23.079926+00:00 · methodology

0 comments
read the original abstract

Occlusion-robust scene recovery remains a major challenge in computational imaging, particularly in natural environments where dense foreground vegetation severely limits visibility. We propose a vision-reasoning-guided light field occlusion removal framework that combines the visibility recovery capability of light field integration (LFI) with the semantic reasoning capacity of vision-language models (VLMs). Multi-view observations are first integrated via LFI to suppress foreground occlusions and produce an initial visibility-enhanced representation. A VLM is then incorporated as a conditional semantic prior to restore degraded structures and recover fine details, guided by the observed measurements. To improve recovery consistency and reduce hallucination artifacts, we introduce a multi-sample fusion strategy that aggregates multiple generated hypotheses into a unified estimate. Experimental results on synthetic and real-world datasets demonstrate state-of-the-art performance, achieving the highest average SSIM across four synthetic light field benchmark scenes (4-Syn) and strong generalization across structured and unstructured acquisition settings. These results highlight the effectiveness of combining physical imaging constraints with vision-language reasoning for robust perception under severe occlusion, with applicability to search-and-rescue and exploratory robotic navigation.

Figures

Figures reproduced from arXiv: 2606.19985 by Mohamed Youssef, Oliver Bimber.

Figure 1
Figure 1. Figure 1: Framework comparison between existing single-image occlusion [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the proposed hybrid occlusion-free reconstruction framework. The framework consists of four main stages to recover occlusion-free scene [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Effect of the number of averaged generated images (N) on reconstruction quality for 4-Syn benchmark scenes using Gemini 3.1 (top) and Qwen-Image [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative comparison of occlusion suppression and scene recon [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative comparison on real-world data between the proposed framework and state-of-the-art methods. The proposed approach demonstrates strong [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Example of the proposed occlusion suppression and reconstruction [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Real-world search-and-rescue scenario under severe vegetation occlusion for RGB (top) and thermal (bottom) modalities. Comparisons are shown [PITH_FULL_IMAGE:figures/full_fig_p009_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

87 extracted references · 11 canonical work pages · 3 internal anchors

  1. [1]

    Decompose to adapt: Cross-domain object detection via feature disentanglement,

    D. Liu, C. Zhang, Y . Song, H. Huang, C. Wang, M. Barnett, and W. Cai, “Decompose to adapt: Cross-domain object detection via feature disentanglement,”IEEE Transactions on Multimedia, vol. 25, pp. 1333– 1344, 2022

  2. [2]

    Salient object detection by fusing local and global contexts,

    Q. Ren, S. Lu, J. Zhang, and R. Hu, “Salient object detection by fusing local and global contexts,”IEEE Transactions on multimedia, vol. 23, pp. 1442–1453, 2020

  3. [3]

    Focal sparse convolutional networks for 3d object detection,

    Y . Chen, Y . Li, X. Zhang, J. Sun, and J. Jia, “Focal sparse convolutional networks for 3d object detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 5428– 5437

  4. [4]

    Multi- correlation filters with triangle-structure constraints for object tracking,

    W. Ruan, J. Chen, Y . Wu, J. Wang, C. Liang, R. Hu, and J. Jiang, “Multi- correlation filters with triangle-structure constraints for object tracking,” IEEE Transactions on Multimedia, vol. 21, no. 5, pp. 1122–1134, 2018

  5. [5]

    Unified transformer tracker for object tracking,

    F. Ma, M. Z. Shou, L. Zhu, H. Fan, Y . Xu, Y . Yang, and Z. Yan, “Unified transformer tracker for object tracking,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 8781– 8790

  6. [6]

    Observation- centric sort: Rethinking sort for robust multi-object tracking,

    J. Cao, J. Pang, X. Weng, R. Khirodkar, and K. Kitani, “Observation- centric sort: Rethinking sort for robust multi-object tracking,” inPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 9686–9696

  7. [7]

    SAM 3: Segment Anything with Concepts

    N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huanget al., “Sam 3: Segment anything with concepts,”arXiv preprint arXiv:2511.16719, 2025

  8. [8]

    An autonomous drone for search and rescue in forests using airborne optical sectioning,

    D. C. Schedl, I. Kurmi, and O. Bimber, “An autonomous drone for search and rescue in forests using airborne optical sectioning,”Science Robotics, vol. 6, no. 55, p. eabg1188, 2021

  9. [9]

    Search and rescue with airborne optical sectioning,

    D. C. Schedl, I. Kurmi, and O. Bimber, “Search and rescue with airborne optical sectioning,”Nature Machine Intelligence, vol. 2, no. 12, pp. 783– 790, 2020

  10. [10]

    How robot dogs see the unseeable: Improving visual interpretability via peering for exploratory robots,

    O. Bimber, K. D. von Ellenrieder, M. Haller, R. J. A. A. Nathan, G. Lunardi, M. Youssef, M. Camurri, S. M. O. Soto, and J. E. Niven, “How robot dogs see the unseeable: Improving visual interpretability via peering for exploratory robots,”arXiv preprint arXiv:2511.16262, 2025

  11. [11]

    Airborne optical sectioning for nesting observation,

    D. C. Schedl, I. Kurmi, and O. Bimber, “Airborne optical sectioning for nesting observation,”Scientific reports, vol. 10, no. 1, p. 7254, 2020

  12. [12]

    Through-foliage surface-temperature reconstruction for early wildfire detection,

    M. Youssef, L. Brunner, K. Rundhammer, G. Czech, and O. Bimber, “Through-foliage surface-temperature reconstruction for early wildfire detection,”arXiv preprint arXiv:2511.12572, 2025

  13. [13]

    Deepforest: Sensing into self- occluding volumes of vegetation with aerial imaging,

    M. Youssef, J. Peng, and O. Bimber, “Deepforest: Sensing into self- occluding volumes of vegetation with aerial imaging,”Journal of Remote Sensing, vol. 5, p. 0907, 2025

  14. [14]

    GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

    Z. Yan, J. Ye, W. Li, Z. Huang, S. Yuan, X. He, K. Lin, J. He, C. He, and L. Yuan, “Gpt-imgeval: A comprehensive benchmark for diagnosing gpt4o in image generation,”arXiv preprint arXiv:2504.02782, 2025

  15. [15]

    Is nano banana pro a low-level vision all-rounder? a comprehensive evaluation on 14 tasks and 40 datasets.arXiv preprint arXiv:2512.15110,

    J. Zuo, H. Deng, H. Zhou, J. Zhu, Y . Zhang, Y . Zhang, Y . Yan, K. Huang, W. Chen, Y . Denget al., “Is nano banana pro a low-level vision all- rounder? a comprehensive evaluation on 14 tasks and 40 datasets,”arXiv preprint arXiv:2512.15110, 2025

  16. [16]

    Patchmatch: a randomized correspondence algorithm for structural image editing,

    C. Barnes, E. Shechtman, A. Finkelstein, and D. B. Goldman, “Patchmatch: a randomized correspondence algorithm for structural image editing,”ACM Trans. Graph., vol. 28, no. 3, Jul. 2009. [Online]. Available: https://doi.org/10.1145/1531326.1531330

  17. [17]

    Context encoders: Feature learning by inpainting,

    D. Pathak, P. Kr ¨ahenb¨uhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 2536–2544

  18. [18]

    Filling-in by joint interpolation of vector fields and gray levels,

    C. Ballester, M. Bertalmio, V . Caselles, G. Sapiro, and J. Verdera, “Filling-in by joint interpolation of vector fields and gray levels,”IEEE Transactions on Image Processing, vol. 10, no. 8, pp. 1200–1211, 2001

  19. [19]

    Generative image inpainting with contextual attention,

    J. Yu, Z. Lin, J. Yang, X. Shen, X. Lu, and T. S. Huang, “Generative image inpainting with contextual attention,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 5505– 5514

  20. [20]

    Generative image inpainting with segmentation confusion adversarial training and contrastive learning,

    Z. Zuo, L. Zhao, A. Li, Z. Wang, Z. Zhang, J. Chen, W. Xing, and D. Lu, “Generative image inpainting with segmentation confusion adversarial training and contrastive learning,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 3, pp. 3888–3896, Jun. 2023. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/ article/view/25502

  21. [21]

    Repaint: Inpainting using denoising diffusion probabilistic models,

    A. Lugmayr, M. Danelljan, A. Romero, F. Yu, R. Timofte, and L. Van Gool, “Repaint: Inpainting using denoising diffusion probabilistic models,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 11 461–11 471

  22. [22]

    Imagen editor and editbench: Advancing and evaluating text-guided image inpainting,

    S. Wang, C. Saharia, C. Montgomery, J. Pont-Tuset, S. Noy, S. Pelle- grini, Y . Onoe, S. Laszlo, D. J. Fleet, R. Soricutet al., “Imagen editor and editbench: Advancing and evaluating text-guided image inpainting,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 18 359–18 369

  23. [23]

    Semantic image inpainting with deep generative models,

    R. A. Yeh, C. Chen, T. Yian Lim, A. G. Schwing, M. Hasegawa-Johnson, and M. N. Do, “Semantic image inpainting with deep generative models,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 5485–5493

  24. [24]

    Free- form image inpainting with gated convolution,

    J. Yu, Z. Lin, J. Yang, X. Shen, X. Lu, and T. S. Huang, “Free- form image inpainting with gated convolution,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 4471– 4480

  25. [25]

    Transformer based pluralistic image completion with reduced information loss,

    Q. Liu, Y . Jiang, Z. Tan, D. Chen, Y . Fu, Q. Chu, G. Hua, and N. Yu, “Transformer based pluralistic image completion with reduced information loss,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 10, pp. 6652–6668, 2024

  26. [26]

    Progressive reconstruction of visual structure for image inpainting,

    J. Li, F. He, L. Zhang, B. Du, and D. Tao, “Progressive reconstruction of visual structure for image inpainting,” in2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 5961–5970

  27. [27]

    Mat: Mask- aware transformer for large hole image inpainting,

    W. Li, Z. Lin, K. Zhou, L. Qi, Y . Wang, and J. Jia, “Mat: Mask- aware transformer for large hole image inpainting,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 758–10 768

  28. [28]

    Learning prior feature and attention enhanced image inpainting,

    C. Cao, Q. Dong, and Y . Fu, “Learning prior feature and attention enhanced image inpainting,” inEuropean conference on computer vision. Springer, 2022, pp. 306–322

  29. [29]

    Image inpainting with local and global refinement,

    W. Quan, R. Zhang, Y . Zhang, Z. Li, J. Wang, and D.-M. Yan, “Image inpainting with local and global refinement,”IEEE Transactions on Image Processing, vol. 31, pp. 2405–2420, 2022

  30. [30]

    In: Pro- ceedings of the 27th annual conference on Computer graphics and interactive tech- niques

    M. Bertalmio, G. Sapiro, V . Caselles, and C. Ballester, “Image inpainting,” inProceedings of the 27th Annual Conference on 11 Computer Graphics and Interactive Techniques, ser. SIGGRAPH ’00. USA: ACM Press/Addison-Wesley Publishing Co., 2000, p. 417–424. [Online]. Available: https://doi.org/10.1145/344779.344972

  31. [31]

    High- resolution image inpainting using multi-scale neural patch synthesis,

    C. Yang, X. Lu, Z. Lin, E. Shechtman, O. Wang, and H. Li, “High- resolution image inpainting using multi-scale neural patch synthesis,” in2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4076–4084

  32. [32]

    Image inpainting via conditional texture and structure dual generation,

    X. Guo, H. Yang, and D. Huang, “Image inpainting via conditional texture and structure dual generation,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2021, pp. 14 134–14 143

  33. [33]

    Image inpainting for irregular holes using partial convolutions,

    G. Liu, F. A. Reda, K. J. Shih, T.-C. Wang, A. Tao, and B. Catanzaro, “Image inpainting for irregular holes using partial convolutions,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 85–100

  34. [34]

    Image inpainting with learnable bidirectional attention maps,

    C. Xie, S. Liu, C. Li, M.-M. Cheng, W. Zuo, X. Liu, S. Wen, and E. Ding, “Image inpainting with learnable bidirectional attention maps,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 8858–8867

  35. [35]

    Recurrent feature reasoning for image inpainting,

    J. Li, N. Wang, L. Zhang, B. Du, and D. Tao, “Recurrent feature reasoning for image inpainting,” in2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 7757– 7765

  36. [36]

    Incremental transformer structure en- hanced image inpainting with masking positional encoding,

    Q. Dong, C. Cao, and Y . Fu, “Incremental transformer structure en- hanced image inpainting with masking positional encoding,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022

  37. [37]

    Coordfill: Efficient high-resolution image inpainting via parameterized coordinate querying,

    W. Liu, X. Cun, C.-M. Pun, M. Xia, Y . Zhang, and J. Wang, “Coordfill: Efficient high-resolution image inpainting via parameterized coordinate querying,” inAAAI, 2023

  38. [38]

    All-in-focus synthetic aperture imaging,

    T. Yang, Y . Zhang, J. Yu, J. Li, W. Ma, X. Tong, R. Yu, and L. Ran, “All-in-focus synthetic aperture imaging,” inComputer Vision – ECCV 2014, D. Fleet, T. Pajdla, B. Schiele, and T. Tuytelaars, Eds. Cham: Springer International Publishing, 2014, pp. 1–15

  39. [39]

    Synthetic aperture tracking: Tracking through occlusions,

    N. Joshi, S. Avidan, W. Matusik, and D. J. Kriegman, “Synthetic aperture tracking: Tracking through occlusions,” in2007 IEEE 11th International Conference on Computer Vision, 2007, pp. 1–8

  40. [40]

    Using plane + paral- lax for calibrating dense camera arrays,

    V . Vaish, B. Wilburn, N. Joshi, and M. Levoy, “Using plane + paral- lax for calibrating dense camera arrays,” inProceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004., vol. 1, 2004, pp. I–I

  41. [41]

    Synthetic aperture focusing using a shear-warp factorization of the viewing transform,

    V . Vaish, G. Garg, E. Talvala, E. Antunez, B. Wilburn, M. Horowitz, and M. Levoy, “Synthetic aperture focusing using a shear-warp factorization of the viewing transform,” in2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05) - Workshops, 2005, pp. 129–129

  42. [42]

    All-in-focus synthetic aperture imaging using image matting,

    Z. Pei, X. Chen, and Y .-H. Yang, “All-in-focus synthetic aperture imaging using image matting,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, no. 2, pp. 288–301, 2018

  43. [43]

    Seeing beyond foreground occlusion: A joint framework for sap-based scene depth and appearance reconstruc- tion,

    Z. Xiao, L. Si, and G. Zhou, “Seeing beyond foreground occlusion: A joint framework for sap-based scene depth and appearance reconstruc- tion,”IEEE Journal of Selected Topics in Signal Processing, vol. 11, no. 7, pp. 979–991, 2017

  44. [44]

    Synthetic aperture imaging using pixel labeling via energy minimization,

    Z. Pei, Y . Zhang, X. Chen, and Y .-H. Yang, “Synthetic aperture imaging using pixel labeling via energy minimization,”Pattern Recognition, vol. 46, no. 1, pp. 174–187, 2013. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0031320312002841

  45. [45]

    Gaussian-wiener represen- tation and hierarchical coding scheme for focal stack images,

    K. Wu, Y . Yang, Q. Liu, and X.-P. Zhang, “Gaussian-wiener represen- tation and hierarchical coding scheme for focal stack images,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 2, pp. 523–537, 2022

  46. [46]

    High performance imaging using large camera arrays,

    B. Wilburn, N. Joshi, V . Vaish, E.-V . Talvala, E. Antunez, A. Barth, A. Adams, M. Horowitz, and M. Levoy, “High performance imaging using large camera arrays,”ACM Trans. Graph., vol. 24, no. 3, p. 765–776, Jul. 2005. [Online]. Available: https://doi.org/10.1145/ 1073204.1073259

  47. [47]

    Re- constructing occluded surfaces using synthetic apertures: Stereo, focus and robust measures,

    V . Vaish, M. Levoy, R. Szeliski, C. Zitnick, and S. B. Kang, “Re- constructing occluded surfaces using synthetic apertures: Stereo, focus and robust measures,” in2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), vol. 2, 2006, pp. 2331–2338

  48. [48]

    Light field super-resolution: A critical review on challenges and opportunities,

    S. Sharma, “Light field super-resolution: A critical review on challenges and opportunities,”arXiv preprint arXiv:2510.07879, 2025

  49. [49]

    Light field depth estimation: A comprehensive survey from principles to future,

    T. Wang, H. Sheng, R. Chen, D. Yang, Z. Cui, S. Wang, R. Cong, and M. Zhao, “Light field depth estimation: A comprehensive survey from principles to future,”High-Confidence Computing, vol. 4, no. 1, p. 100187, 2024. [Online]. Available: https://www.sciencedirect.com/ science/article/pii/S2667295223000855

  50. [50]

    DeOccNet: Learning to see through foreground occlusions in light fields,

    Y . Wang, T. Wu, J. Yang, L. Wang, W. An, and Y . Guo, “DeOccNet: Learning to see through foreground occlusions in light fields,” inWinter Conference on Applications of Computer Vision (WACV), Mar 2020

  51. [51]

    Mask4d: End-to-end mask-based 4d panoptic segmentation for lidar sequences,

    R. Marcuzzi, L. Nunes, L. Wiesmann, E. Marks, J. Behley, and C. Stach- niss, “Mask4d: End-to-end mask-based 4d panoptic segmentation for lidar sequences,”IEEE Robotics and Automation Letters, vol. 8, no. 11, pp. 7487–7494, 2023

  52. [52]

    Light field occlusion removal network via foreground location and background recovery,

    S. Zhang, Y . Chen, P. An, X. Huang, and C. Yang, “Light field occlusion removal network via foreground location and background recovery,” Signal Processing: Image Communication, vol. 109, p. 116853, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S0923596522001345

  53. [53]

    I see-through you: A framework for removing foreground occlusion in both sparse and dense light field images,

    J. Hur, J. Y . Lee, J. Choi, and J. Kim, “I see-through you: A framework for removing foreground occlusion in both sparse and dense light field images,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 229–238

  54. [54]

    Mask- aware light field de-occlusion with gated feature aggregation and texture- semantic attention,

    J. Chen, P. An, X. Huang, Y . Chen, C. Yang, and L. Shen, “Mask- aware light field de-occlusion with gated feature aggregation and texture- semantic attention,”IEEE Transactions on Multimedia, vol. 27, pp. 5296–5311, 2025

  55. [55]

    Progressive multi-plane images construction for light field occlusion removal,

    S. Zhang, S. Chang, Z. Shi, and Y . Lin, “Progressive multi-plane images construction for light field occlusion removal,”IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 10, pp. 8012–8025, 2025

  56. [56]

    Effective light field de-occlusion network based on swin transformer,

    X. Wang, J. Liu, S. Chen, and G. Wei, “Effective light field de-occlusion network based on swin transformer,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 6, pp. 2590–2599, 2023

  57. [57]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021

  58. [58]

    All-in-focus synthetic aperture imaging using generative adversarial network-based semantic inpainting,

    Z. Pei, M. Jin, Y . Zhang, M. Ma, and Y .-H. Yang, “All-in-focus synthetic aperture imaging using generative adversarial network-based semantic inpainting,”Pattern Recognition, vol. 111, p. 107669, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S0031320320304726

  59. [59]

    Generative adversarial networks,

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversarial networks,” Communications of the ACM, vol. 63, no. 11, pp. 139–144, 2020

  60. [60]

    D. Wang, Q. Pan, C. Zhao, J. Hu, Z. Xu, F. Yang, and Y . Zhou, “A study on camera array and its applications**this research is funded by the state key laboratory of geo-information engineering under grant agreement no. sklgie2015-m-3-4 and is supported by national science foundation of china (grant no. 61473230), national science foundation for young scho...

  61. [61]

    Decoding, calibration and rectification for lenselet-based plenoptic cameras,

    D. G. Dansereau, O. Pizarro, and S. B. Williams, “Decoding, calibration and rectification for lenselet-based plenoptic cameras,” in2013 IEEE Conference on Computer Vision and Pattern Recognition, 2013, pp. 1027–1034

  62. [62]

    Structure-from-motion revisited,

    J. L. Sch ¨onberger and J.-M. Frahm, “Structure-from-motion revisited,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4104–4113

  63. [63]

    Simultaneous localization and mapping (slam)-based robot localization and navigation algorithm,

    J. Qiao, J. Guo, and Y . Li, “Simultaneous localization and mapping (slam)-based robot localization and navigation algorithm,”Applied Water Science, vol. 14, no. 7, p. 151, 2024

  64. [64]

    Vggt: Visual geometry grounded transformer,

    J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 5294– 5306

  65. [65]

    Depth Anything 3: Recovering the Visual Space from Any Views

    H. Lin, S. Chen, J. H. Liew, D. Y . Chen, Z. Li, G. Shi, J. Feng, and B. Kang, “Depth anything 3: Recovering the visual space from any views,”arXiv preprint arXiv:2511.10647, 2025

  66. [66]

    Significant remote sensing vegetation indices: A review of developments and applications,

    J. Xue and B. Su, “Significant remote sensing vegetation indices: A review of developments and applications,”Journal of sensors, vol. 2017, no. 1, p. 1353691, 2017

  67. [67]

    Pixel-perfect depth with semantics-prompted diffusion transformers,

    G. Xu, H. Lin, H. Luo, X. Wang, J. Yao, L. Zhu, Y . Pu, C. Chi , H. Sun, B. Wanget al., “Pixel-perfect depth with semantics-prompted diffusion transformers,”Advances in Neural Information Processing Systems, vol. 38, pp. 174 731–174 755, 2026. 12

  68. [68]

    Unsupervised monocular depth estimation from light field image,

    W. Zhou, E. Zhou, G. Liu, L. Lin, and A. Lumsdaine, “Unsupervised monocular depth estimation from light field image,”IEEE Transactions on Image Processing, vol. 29, pp. 1606–1617, 2020

  69. [69]

    Complex-valued disparity: Unified depth model of depth from stereo, depth from focus, and depth from defocus based on the light field gradient,

    J. Y . Lee and R.-H. Park, “Complex-valued disparity: Unified depth model of depth from stereo, depth from focus, and depth from defocus based on the light field gradient,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 3, pp. 830–841, 2021

  70. [70]

    Airborne optical sectioning,

    I. Kurmi, D. C. Schedl, and O. Bimber, “Airborne optical sectioning,” Journal of Imaging, vol. 4, no. 8, p. 102, 2018

  71. [71]

    Synthetic aperture imaging with drones,

    O. Bimber, I. Kurmi, and D. C. Schedl, “Synthetic aperture imaging with drones,”IEEE computer graphics and applications, vol. 39, no. 3, pp. 8–15, 2019

  72. [72]

    A statistical view on synthetic aperture imaging for occlusion removal,

    I. Kurmi, D. C. Schedl, and O. Bimber, “A statistical view on synthetic aperture imaging for occlusion removal,”IEEE Sensors Journal, vol. 19, no. 20, pp. 9374–9383, 2019

  73. [73]

    Thermal airborne optical sectioning,

    I. Kurmi, D. C. Schedl, and O. Bimber, “Thermal airborne optical sectioning,”Remote Sensing, vol. 11, no. 14, p. 1668, 2019

  74. [74]

    Fast automatic visibility opti- mization for thermal synthetic aperture visualization,

    I. Kurmi, D. C. Schedl, and O. Bimber, “Fast automatic visibility opti- mization for thermal synthetic aperture visualization,”IEEE Geoscience and Remote Sensing Letters, vol. 18, no. 5, pp. 836–840, 2020

  75. [75]

    Pose error reduction for focus enhancement in thermal synthetic aperture visualization,

    I. Kurmi, D. C. Schedl, and O. Bimber, “Pose error reduction for focus enhancement in thermal synthetic aperture visualization,”IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2021

  76. [76]

    Combined person classification with airborne optical sectioning,

    I. Kurmi, D. C. Schedl, and O. Bimber, “Combined person classification with airborne optical sectioning,”Scientific reports, vol. 12, no. 1, p. 3804, 2022

  77. [77]

    Through- foliage tracking with airborne optical sectioning,

    R. J. A. A. Nathan, I. Kurmi, D. C. Schedl, and O. Bimber, “Through- foliage tracking with airborne optical sectioning,”Journal of Remote Sensing, 2022

  78. [78]

    An autonomous drone swarm for detecting and track- ing anomalies among dense vegetation,

    R. J. Amala Arokia Nathan, S. Strand, D. Mehrwald, D. Shutin, and O. Bimber, “An autonomous drone swarm for detecting and track- ing anomalies among dense vegetation,”Communications Engineering, vol. 4, no. 1, p. 205, 2025

  79. [79]

    Reciprocal visibility for guided occlusion removal with drones,

    R. J. A. A. Nathan, S. Strand, D. Shutin, and O. Bimber, “Reciprocal visibility for guided occlusion removal with drones,”IEEE Geoscience and Remote Sensing Letters, vol. 21, pp. 1–5, 2024

  80. [80]

    Fusion of single and integral multispectral aerial images,

    M. Youssef and O. Bimber, “Fusion of single and integral multispectral aerial images,”remote sensing, vol. 16, no. 4, p. 673, 2024

Showing first 80 references.