REVIEW 3 minor 87 references
Light field integration followed by a conditioned vision-language model restores occluded scenes with highest benchmark accuracy.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-26 18:27 UTC pith:YP7FHKW2
load-bearing objection The paper combines light field integration with a VLM prior and multi-sample fusion for occlusion removal, and the full text supplies ablations plus baseline comparisons that make the SSIM claims credible.
Vision-Reasoning-Guided Occlusion Removal from Light Fields
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The vision-reasoning-guided framework achieves superior occlusion removal by first applying light field integration to suppress foreground occlusions and then conditioning a vision-language model on the integrated measurements to restore fine details, with multi-sample fusion ensuring consistency and minimizing hallucinations. This yields the highest average SSIM on four synthetic light field benchmark scenes and demonstrates strong generalization to both structured and unstructured real-world acquisitions.
What carries the argument
Multi-sample fusion of hypotheses generated by a vision-language model conditioned on light-field-integrated measurements.
Load-bearing premise
A vision-language model conditioned on the integrated measurements can restore accurate fine details without generating hallucinated structures that contradict the physical observations.
What would settle it
A quantitative test in which the final fused output shows higher error or lower SSIM than the light-field-integrated image alone on scenes where the model introduces visible structures absent from all input views.
If this is right
- Highest average SSIM reported across the four synthetic light field benchmark scenes.
- Strong generalization performance on both structured and unstructured real-world acquisition settings.
- Direct applicability to search-and-rescue and exploratory robotic navigation under severe occlusion.
- Demonstrates that physical imaging constraints can be combined with vision-language reasoning for robust perception.
Where Pith is reading between the lines
- The same conditioning-plus-fusion pattern could be tested on other multi-view or volumetric imaging modalities where semantic priors might complement physical data.
- Adding explicit consistency checks between VLM outputs and raw light-field measurements might further reduce residual hallucinations.
- Performance on dynamic or time-varying occlusions remains untested and would require extending the integration step to include temporal information.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a vision-reasoning-guided framework for light-field occlusion removal. Multi-view observations are integrated via light-field integration (LFI) to suppress foreground occlusions, after which a vision-language model (VLM) is conditioned on the integrated measurements to restore fine details; a multi-sample fusion step aggregates multiple VLM hypotheses to improve consistency. Experiments on the 4-Syn synthetic benchmark and real structured/unstructured captures report state-of-the-art average SSIM together with strong cross-setting generalization.
Significance. If the reported SSIM gains and generalization hold under the supplied ablations and baseline comparisons, the work shows that physical light-field constraints can be productively combined with semantic priors from VLMs to handle severe natural occlusion. The multi-sample fusion mechanism directly mitigates the hallucination risk that would otherwise undermine the VLM component, and the inclusion of implementation details plus quantitative tables strengthens the central claim for applications in robotic navigation and search-and-rescue.
minor comments (3)
- [§3.2] §3.2: the precise conditioning mechanism (prompt template and how LFI output is tokenized for the VLM) is described at a high level; an explicit listing of the prompt components and any learned adapters would improve reproducibility.
- [Table 2] Table 2: the per-scene SSIM values for the 4-Syn benchmark are summarized only as averages; adding the individual scene scores would allow readers to assess whether gains are uniform or driven by particular scenes.
- [Figure 4] Figure 4 caption: the visual comparison panels would benefit from an explicit statement of the quantitative metric (SSIM or PSNR) shown beneath each result.
Simulated Author's Rebuttal
We thank the referee for the positive summary, significance assessment, and recommendation of minor revision. The report raises no specific major comments or criticisms, so we provide no point-by-point responses.
Circularity Check
No significant circularity in derivation chain
full rationale
The manuscript describes an empirical framework that integrates light-field integration with a vision-language model and a multi-sample fusion step, then reports quantitative results on benchmark datasets. No equations, derivations, or parameter-fitting procedures are presented that reduce to self-definitions, fitted inputs renamed as predictions, or self-citation chains. All central claims rest on external experimental measurements rather than internal consistency loops, satisfying the self-contained criterion.
Axiom & Free-Parameter Ledger
read the original abstract
Occlusion-robust scene recovery remains a major challenge in computational imaging, particularly in natural environments where dense foreground vegetation severely limits visibility. We propose a vision-reasoning-guided light field occlusion removal framework that combines the visibility recovery capability of light field integration (LFI) with the semantic reasoning capacity of vision-language models (VLMs). Multi-view observations are first integrated via LFI to suppress foreground occlusions and produce an initial visibility-enhanced representation. A VLM is then incorporated as a conditional semantic prior to restore degraded structures and recover fine details, guided by the observed measurements. To improve recovery consistency and reduce hallucination artifacts, we introduce a multi-sample fusion strategy that aggregates multiple generated hypotheses into a unified estimate. Experimental results on synthetic and real-world datasets demonstrate state-of-the-art performance, achieving the highest average SSIM across four synthetic light field benchmark scenes (4-Syn) and strong generalization across structured and unstructured acquisition settings. These results highlight the effectiveness of combining physical imaging constraints with vision-language reasoning for robust perception under severe occlusion, with applicability to search-and-rescue and exploratory robotic navigation.
Figures
Reference graph
Works this paper leans on
-
[1]
Decompose to adapt: Cross-domain object detection via feature disentanglement,
D. Liu, C. Zhang, Y . Song, H. Huang, C. Wang, M. Barnett, and W. Cai, “Decompose to adapt: Cross-domain object detection via feature disentanglement,”IEEE Transactions on Multimedia, vol. 25, pp. 1333– 1344, 2022
2022
-
[2]
Salient object detection by fusing local and global contexts,
Q. Ren, S. Lu, J. Zhang, and R. Hu, “Salient object detection by fusing local and global contexts,”IEEE Transactions on multimedia, vol. 23, pp. 1442–1453, 2020
2020
-
[3]
Focal sparse convolutional networks for 3d object detection,
Y . Chen, Y . Li, X. Zhang, J. Sun, and J. Jia, “Focal sparse convolutional networks for 3d object detection,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 5428– 5437
2022
-
[4]
Multi- correlation filters with triangle-structure constraints for object tracking,
W. Ruan, J. Chen, Y . Wu, J. Wang, C. Liang, R. Hu, and J. Jiang, “Multi- correlation filters with triangle-structure constraints for object tracking,” IEEE Transactions on Multimedia, vol. 21, no. 5, pp. 1122–1134, 2018
2018
-
[5]
Unified transformer tracker for object tracking,
F. Ma, M. Z. Shou, L. Zhu, H. Fan, Y . Xu, Y . Yang, and Z. Yan, “Unified transformer tracker for object tracking,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 8781– 8790
2022
-
[6]
Observation- centric sort: Rethinking sort for robust multi-object tracking,
J. Cao, J. Pang, X. Weng, R. Khirodkar, and K. Kitani, “Observation- centric sort: Rethinking sort for robust multi-object tracking,” inPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 9686–9696
2023
-
[7]
SAM 3: Segment Anything with Concepts
N. Carion, L. Gustafson, Y .-T. Hu, S. Debnath, R. Hu, D. Suris, C. Ryali, K. V . Alwala, H. Khedr, A. Huanget al., “Sam 3: Segment anything with concepts,”arXiv preprint arXiv:2511.16719, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[8]
An autonomous drone for search and rescue in forests using airborne optical sectioning,
D. C. Schedl, I. Kurmi, and O. Bimber, “An autonomous drone for search and rescue in forests using airborne optical sectioning,”Science Robotics, vol. 6, no. 55, p. eabg1188, 2021
2021
-
[9]
Search and rescue with airborne optical sectioning,
D. C. Schedl, I. Kurmi, and O. Bimber, “Search and rescue with airborne optical sectioning,”Nature Machine Intelligence, vol. 2, no. 12, pp. 783– 790, 2020
2020
-
[10]
O. Bimber, K. D. von Ellenrieder, M. Haller, R. J. A. A. Nathan, G. Lunardi, M. Youssef, M. Camurri, S. M. O. Soto, and J. E. Niven, “How robot dogs see the unseeable: Improving visual interpretability via peering for exploratory robots,”arXiv preprint arXiv:2511.16262, 2025
-
[11]
Airborne optical sectioning for nesting observation,
D. C. Schedl, I. Kurmi, and O. Bimber, “Airborne optical sectioning for nesting observation,”Scientific reports, vol. 10, no. 1, p. 7254, 2020
2020
-
[12]
Through-foliage surface-temperature reconstruction for early wildfire detection,
M. Youssef, L. Brunner, K. Rundhammer, G. Czech, and O. Bimber, “Through-foliage surface-temperature reconstruction for early wildfire detection,”arXiv preprint arXiv:2511.12572, 2025
-
[13]
Deepforest: Sensing into self- occluding volumes of vegetation with aerial imaging,
M. Youssef, J. Peng, and O. Bimber, “Deepforest: Sensing into self- occluding volumes of vegetation with aerial imaging,”Journal of Remote Sensing, vol. 5, p. 0907, 2025
2025
-
[14]
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
Z. Yan, J. Ye, W. Li, Z. Huang, S. Yuan, X. He, K. Lin, J. He, C. He, and L. Yuan, “Gpt-imgeval: A comprehensive benchmark for diagnosing gpt4o in image generation,”arXiv preprint arXiv:2504.02782, 2025
work page Pith review arXiv 2025
-
[15]
J. Zuo, H. Deng, H. Zhou, J. Zhu, Y . Zhang, Y . Zhang, Y . Yan, K. Huang, W. Chen, Y . Denget al., “Is nano banana pro a low-level vision all- rounder? a comprehensive evaluation on 14 tasks and 40 datasets,”arXiv preprint arXiv:2512.15110, 2025
-
[16]
Patchmatch: a randomized correspondence algorithm for structural image editing,
C. Barnes, E. Shechtman, A. Finkelstein, and D. B. Goldman, “Patchmatch: a randomized correspondence algorithm for structural image editing,”ACM Trans. Graph., vol. 28, no. 3, Jul. 2009. [Online]. Available: https://doi.org/10.1145/1531326.1531330
-
[17]
Context encoders: Feature learning by inpainting,
D. Pathak, P. Kr ¨ahenb¨uhl, J. Donahue, T. Darrell, and A. A. Efros, “Context encoders: Feature learning by inpainting,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 2536–2544
2016
-
[18]
Filling-in by joint interpolation of vector fields and gray levels,
C. Ballester, M. Bertalmio, V . Caselles, G. Sapiro, and J. Verdera, “Filling-in by joint interpolation of vector fields and gray levels,”IEEE Transactions on Image Processing, vol. 10, no. 8, pp. 1200–1211, 2001
2001
-
[19]
Generative image inpainting with contextual attention,
J. Yu, Z. Lin, J. Yang, X. Shen, X. Lu, and T. S. Huang, “Generative image inpainting with contextual attention,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 5505– 5514
2018
-
[20]
Generative image inpainting with segmentation confusion adversarial training and contrastive learning,
Z. Zuo, L. Zhao, A. Li, Z. Wang, Z. Zhang, J. Chen, W. Xing, and D. Lu, “Generative image inpainting with segmentation confusion adversarial training and contrastive learning,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 3, pp. 3888–3896, Jun. 2023. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/ article/view/25502
2023
-
[21]
Repaint: Inpainting using denoising diffusion probabilistic models,
A. Lugmayr, M. Danelljan, A. Romero, F. Yu, R. Timofte, and L. Van Gool, “Repaint: Inpainting using denoising diffusion probabilistic models,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 11 461–11 471
2022
-
[22]
Imagen editor and editbench: Advancing and evaluating text-guided image inpainting,
S. Wang, C. Saharia, C. Montgomery, J. Pont-Tuset, S. Noy, S. Pelle- grini, Y . Onoe, S. Laszlo, D. J. Fleet, R. Soricutet al., “Imagen editor and editbench: Advancing and evaluating text-guided image inpainting,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 18 359–18 369
2023
-
[23]
Semantic image inpainting with deep generative models,
R. A. Yeh, C. Chen, T. Yian Lim, A. G. Schwing, M. Hasegawa-Johnson, and M. N. Do, “Semantic image inpainting with deep generative models,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 5485–5493
2017
-
[24]
Free- form image inpainting with gated convolution,
J. Yu, Z. Lin, J. Yang, X. Shen, X. Lu, and T. S. Huang, “Free- form image inpainting with gated convolution,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 4471– 4480
2019
-
[25]
Transformer based pluralistic image completion with reduced information loss,
Q. Liu, Y . Jiang, Z. Tan, D. Chen, Y . Fu, Q. Chu, G. Hua, and N. Yu, “Transformer based pluralistic image completion with reduced information loss,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 10, pp. 6652–6668, 2024
2024
-
[26]
Progressive reconstruction of visual structure for image inpainting,
J. Li, F. He, L. Zhang, B. Du, and D. Tao, “Progressive reconstruction of visual structure for image inpainting,” in2019 IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 5961–5970
2019
-
[27]
Mat: Mask- aware transformer for large hole image inpainting,
W. Li, Z. Lin, K. Zhou, L. Qi, Y . Wang, and J. Jia, “Mat: Mask- aware transformer for large hole image inpainting,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 758–10 768
2022
-
[28]
Learning prior feature and attention enhanced image inpainting,
C. Cao, Q. Dong, and Y . Fu, “Learning prior feature and attention enhanced image inpainting,” inEuropean conference on computer vision. Springer, 2022, pp. 306–322
2022
-
[29]
Image inpainting with local and global refinement,
W. Quan, R. Zhang, Y . Zhang, Z. Li, J. Wang, and D.-M. Yan, “Image inpainting with local and global refinement,”IEEE Transactions on Image Processing, vol. 31, pp. 2405–2420, 2022
2022
-
[30]
In: Pro- ceedings of the 27th annual conference on Computer graphics and interactive tech- niques
M. Bertalmio, G. Sapiro, V . Caselles, and C. Ballester, “Image inpainting,” inProceedings of the 27th Annual Conference on 11 Computer Graphics and Interactive Techniques, ser. SIGGRAPH ’00. USA: ACM Press/Addison-Wesley Publishing Co., 2000, p. 417–424. [Online]. Available: https://doi.org/10.1145/344779.344972
-
[31]
High- resolution image inpainting using multi-scale neural patch synthesis,
C. Yang, X. Lu, Z. Lin, E. Shechtman, O. Wang, and H. Li, “High- resolution image inpainting using multi-scale neural patch synthesis,” in2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4076–4084
2017
-
[32]
Image inpainting via conditional texture and structure dual generation,
X. Guo, H. Yang, and D. Huang, “Image inpainting via conditional texture and structure dual generation,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2021, pp. 14 134–14 143
2021
-
[33]
Image inpainting for irregular holes using partial convolutions,
G. Liu, F. A. Reda, K. J. Shih, T.-C. Wang, A. Tao, and B. Catanzaro, “Image inpainting for irregular holes using partial convolutions,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 85–100
2018
-
[34]
Image inpainting with learnable bidirectional attention maps,
C. Xie, S. Liu, C. Li, M.-M. Cheng, W. Zuo, X. Liu, S. Wen, and E. Ding, “Image inpainting with learnable bidirectional attention maps,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 8858–8867
2019
-
[35]
Recurrent feature reasoning for image inpainting,
J. Li, N. Wang, L. Zhang, B. Du, and D. Tao, “Recurrent feature reasoning for image inpainting,” in2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 7757– 7765
2020
-
[36]
Incremental transformer structure en- hanced image inpainting with masking positional encoding,
Q. Dong, C. Cao, and Y . Fu, “Incremental transformer structure en- hanced image inpainting with masking positional encoding,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022
2022
-
[37]
Coordfill: Efficient high-resolution image inpainting via parameterized coordinate querying,
W. Liu, X. Cun, C.-M. Pun, M. Xia, Y . Zhang, and J. Wang, “Coordfill: Efficient high-resolution image inpainting via parameterized coordinate querying,” inAAAI, 2023
2023
-
[38]
All-in-focus synthetic aperture imaging,
T. Yang, Y . Zhang, J. Yu, J. Li, W. Ma, X. Tong, R. Yu, and L. Ran, “All-in-focus synthetic aperture imaging,” inComputer Vision – ECCV 2014, D. Fleet, T. Pajdla, B. Schiele, and T. Tuytelaars, Eds. Cham: Springer International Publishing, 2014, pp. 1–15
2014
-
[39]
Synthetic aperture tracking: Tracking through occlusions,
N. Joshi, S. Avidan, W. Matusik, and D. J. Kriegman, “Synthetic aperture tracking: Tracking through occlusions,” in2007 IEEE 11th International Conference on Computer Vision, 2007, pp. 1–8
2007
-
[40]
Using plane + paral- lax for calibrating dense camera arrays,
V . Vaish, B. Wilburn, N. Joshi, and M. Levoy, “Using plane + paral- lax for calibrating dense camera arrays,” inProceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004., vol. 1, 2004, pp. I–I
2004
-
[41]
Synthetic aperture focusing using a shear-warp factorization of the viewing transform,
V . Vaish, G. Garg, E. Talvala, E. Antunez, B. Wilburn, M. Horowitz, and M. Levoy, “Synthetic aperture focusing using a shear-warp factorization of the viewing transform,” in2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05) - Workshops, 2005, pp. 129–129
2005
-
[42]
All-in-focus synthetic aperture imaging using image matting,
Z. Pei, X. Chen, and Y .-H. Yang, “All-in-focus synthetic aperture imaging using image matting,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 28, no. 2, pp. 288–301, 2018
2018
-
[43]
Seeing beyond foreground occlusion: A joint framework for sap-based scene depth and appearance reconstruc- tion,
Z. Xiao, L. Si, and G. Zhou, “Seeing beyond foreground occlusion: A joint framework for sap-based scene depth and appearance reconstruc- tion,”IEEE Journal of Selected Topics in Signal Processing, vol. 11, no. 7, pp. 979–991, 2017
2017
-
[44]
Synthetic aperture imaging using pixel labeling via energy minimization,
Z. Pei, Y . Zhang, X. Chen, and Y .-H. Yang, “Synthetic aperture imaging using pixel labeling via energy minimization,”Pattern Recognition, vol. 46, no. 1, pp. 174–187, 2013. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0031320312002841
2013
-
[45]
Gaussian-wiener represen- tation and hierarchical coding scheme for focal stack images,
K. Wu, Y . Yang, Q. Liu, and X.-P. Zhang, “Gaussian-wiener represen- tation and hierarchical coding scheme for focal stack images,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 2, pp. 523–537, 2022
2022
-
[46]
High performance imaging using large camera arrays,
B. Wilburn, N. Joshi, V . Vaish, E.-V . Talvala, E. Antunez, A. Barth, A. Adams, M. Horowitz, and M. Levoy, “High performance imaging using large camera arrays,”ACM Trans. Graph., vol. 24, no. 3, p. 765–776, Jul. 2005. [Online]. Available: https://doi.org/10.1145/ 1073204.1073259
-
[47]
Re- constructing occluded surfaces using synthetic apertures: Stereo, focus and robust measures,
V . Vaish, M. Levoy, R. Szeliski, C. Zitnick, and S. B. Kang, “Re- constructing occluded surfaces using synthetic apertures: Stereo, focus and robust measures,” in2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), vol. 2, 2006, pp. 2331–2338
2006
-
[48]
Light field super-resolution: A critical review on challenges and opportunities,
S. Sharma, “Light field super-resolution: A critical review on challenges and opportunities,”arXiv preprint arXiv:2510.07879, 2025
-
[49]
Light field depth estimation: A comprehensive survey from principles to future,
T. Wang, H. Sheng, R. Chen, D. Yang, Z. Cui, S. Wang, R. Cong, and M. Zhao, “Light field depth estimation: A comprehensive survey from principles to future,”High-Confidence Computing, vol. 4, no. 1, p. 100187, 2024. [Online]. Available: https://www.sciencedirect.com/ science/article/pii/S2667295223000855
2024
-
[50]
DeOccNet: Learning to see through foreground occlusions in light fields,
Y . Wang, T. Wu, J. Yang, L. Wang, W. An, and Y . Guo, “DeOccNet: Learning to see through foreground occlusions in light fields,” inWinter Conference on Applications of Computer Vision (WACV), Mar 2020
2020
-
[51]
Mask4d: End-to-end mask-based 4d panoptic segmentation for lidar sequences,
R. Marcuzzi, L. Nunes, L. Wiesmann, E. Marks, J. Behley, and C. Stach- niss, “Mask4d: End-to-end mask-based 4d panoptic segmentation for lidar sequences,”IEEE Robotics and Automation Letters, vol. 8, no. 11, pp. 7487–7494, 2023
2023
-
[52]
Light field occlusion removal network via foreground location and background recovery,
S. Zhang, Y . Chen, P. An, X. Huang, and C. Yang, “Light field occlusion removal network via foreground location and background recovery,” Signal Processing: Image Communication, vol. 109, p. 116853, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S0923596522001345
2022
-
[53]
I see-through you: A framework for removing foreground occlusion in both sparse and dense light field images,
J. Hur, J. Y . Lee, J. Choi, and J. Kim, “I see-through you: A framework for removing foreground occlusion in both sparse and dense light field images,” inProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2023, pp. 229–238
2023
-
[54]
Mask- aware light field de-occlusion with gated feature aggregation and texture- semantic attention,
J. Chen, P. An, X. Huang, Y . Chen, C. Yang, and L. Shen, “Mask- aware light field de-occlusion with gated feature aggregation and texture- semantic attention,”IEEE Transactions on Multimedia, vol. 27, pp. 5296–5311, 2025
2025
-
[55]
Progressive multi-plane images construction for light field occlusion removal,
S. Zhang, S. Chang, Z. Shi, and Y . Lin, “Progressive multi-plane images construction for light field occlusion removal,”IEEE Transactions on Visualization and Computer Graphics, vol. 31, no. 10, pp. 8012–8025, 2025
2025
-
[56]
Effective light field de-occlusion network based on swin transformer,
X. Wang, J. Liu, S. Chen, and G. Wei, “Effective light field de-occlusion network based on swin transformer,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 6, pp. 2590–2599, 2023
2023
-
[57]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y . Lin, Y . Cao, H. Hu, Y . Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021
2021
-
[58]
All-in-focus synthetic aperture imaging using generative adversarial network-based semantic inpainting,
Z. Pei, M. Jin, Y . Zhang, M. Ma, and Y .-H. Yang, “All-in-focus synthetic aperture imaging using generative adversarial network-based semantic inpainting,”Pattern Recognition, vol. 111, p. 107669, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S0031320320304726
2021
-
[59]
Generative adversarial networks,
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversarial networks,” Communications of the ACM, vol. 63, no. 11, pp. 139–144, 2020
2020
-
[60]
D. Wang, Q. Pan, C. Zhao, J. Hu, Z. Xu, F. Yang, and Y . Zhou, “A study on camera array and its applications**this research is funded by the state key laboratory of geo-information engineering under grant agreement no. sklgie2015-m-3-4 and is supported by national science foundation of china (grant no. 61473230), national science foundation for young scho...
2017
-
[61]
Decoding, calibration and rectification for lenselet-based plenoptic cameras,
D. G. Dansereau, O. Pizarro, and S. B. Williams, “Decoding, calibration and rectification for lenselet-based plenoptic cameras,” in2013 IEEE Conference on Computer Vision and Pattern Recognition, 2013, pp. 1027–1034
2013
-
[62]
Structure-from-motion revisited,
J. L. Sch ¨onberger and J.-M. Frahm, “Structure-from-motion revisited,” in2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4104–4113
2016
-
[63]
Simultaneous localization and mapping (slam)-based robot localization and navigation algorithm,
J. Qiao, J. Guo, and Y . Li, “Simultaneous localization and mapping (slam)-based robot localization and navigation algorithm,”Applied Water Science, vol. 14, no. 7, p. 151, 2024
2024
-
[64]
Vggt: Visual geometry grounded transformer,
J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 5294– 5306
2025
-
[65]
Depth Anything 3: Recovering the Visual Space from Any Views
H. Lin, S. Chen, J. H. Liew, D. Y . Chen, Z. Li, G. Shi, J. Feng, and B. Kang, “Depth anything 3: Recovering the visual space from any views,”arXiv preprint arXiv:2511.10647, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[66]
Significant remote sensing vegetation indices: A review of developments and applications,
J. Xue and B. Su, “Significant remote sensing vegetation indices: A review of developments and applications,”Journal of sensors, vol. 2017, no. 1, p. 1353691, 2017
2017
-
[67]
Pixel-perfect depth with semantics-prompted diffusion transformers,
G. Xu, H. Lin, H. Luo, X. Wang, J. Yao, L. Zhu, Y . Pu, C. Chi , H. Sun, B. Wanget al., “Pixel-perfect depth with semantics-prompted diffusion transformers,”Advances in Neural Information Processing Systems, vol. 38, pp. 174 731–174 755, 2026. 12
2026
-
[68]
Unsupervised monocular depth estimation from light field image,
W. Zhou, E. Zhou, G. Liu, L. Lin, and A. Lumsdaine, “Unsupervised monocular depth estimation from light field image,”IEEE Transactions on Image Processing, vol. 29, pp. 1606–1617, 2020
2020
-
[69]
Complex-valued disparity: Unified depth model of depth from stereo, depth from focus, and depth from defocus based on the light field gradient,
J. Y . Lee and R.-H. Park, “Complex-valued disparity: Unified depth model of depth from stereo, depth from focus, and depth from defocus based on the light field gradient,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 3, pp. 830–841, 2021
2021
-
[70]
Airborne optical sectioning,
I. Kurmi, D. C. Schedl, and O. Bimber, “Airborne optical sectioning,” Journal of Imaging, vol. 4, no. 8, p. 102, 2018
2018
-
[71]
Synthetic aperture imaging with drones,
O. Bimber, I. Kurmi, and D. C. Schedl, “Synthetic aperture imaging with drones,”IEEE computer graphics and applications, vol. 39, no. 3, pp. 8–15, 2019
2019
-
[72]
A statistical view on synthetic aperture imaging for occlusion removal,
I. Kurmi, D. C. Schedl, and O. Bimber, “A statistical view on synthetic aperture imaging for occlusion removal,”IEEE Sensors Journal, vol. 19, no. 20, pp. 9374–9383, 2019
2019
-
[73]
Thermal airborne optical sectioning,
I. Kurmi, D. C. Schedl, and O. Bimber, “Thermal airborne optical sectioning,”Remote Sensing, vol. 11, no. 14, p. 1668, 2019
2019
-
[74]
Fast automatic visibility opti- mization for thermal synthetic aperture visualization,
I. Kurmi, D. C. Schedl, and O. Bimber, “Fast automatic visibility opti- mization for thermal synthetic aperture visualization,”IEEE Geoscience and Remote Sensing Letters, vol. 18, no. 5, pp. 836–840, 2020
2020
-
[75]
Pose error reduction for focus enhancement in thermal synthetic aperture visualization,
I. Kurmi, D. C. Schedl, and O. Bimber, “Pose error reduction for focus enhancement in thermal synthetic aperture visualization,”IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1–5, 2021
2021
-
[76]
Combined person classification with airborne optical sectioning,
I. Kurmi, D. C. Schedl, and O. Bimber, “Combined person classification with airborne optical sectioning,”Scientific reports, vol. 12, no. 1, p. 3804, 2022
2022
-
[77]
Through- foliage tracking with airborne optical sectioning,
R. J. A. A. Nathan, I. Kurmi, D. C. Schedl, and O. Bimber, “Through- foliage tracking with airborne optical sectioning,”Journal of Remote Sensing, 2022
2022
-
[78]
An autonomous drone swarm for detecting and track- ing anomalies among dense vegetation,
R. J. Amala Arokia Nathan, S. Strand, D. Mehrwald, D. Shutin, and O. Bimber, “An autonomous drone swarm for detecting and track- ing anomalies among dense vegetation,”Communications Engineering, vol. 4, no. 1, p. 205, 2025
2025
-
[79]
Reciprocal visibility for guided occlusion removal with drones,
R. J. A. A. Nathan, S. Strand, D. Shutin, and O. Bimber, “Reciprocal visibility for guided occlusion removal with drones,”IEEE Geoscience and Remote Sensing Letters, vol. 21, pp. 1–5, 2024
2024
-
[80]
Fusion of single and integral multispectral aerial images,
M. Youssef and O. Bimber, “Fusion of single and integral multispectral aerial images,”remote sensing, vol. 16, no. 4, p. 673, 2024
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.