REVIEW 2 major objections 1 minor 27 references
The first real gastroscopy dataset enables evaluation of novel view synthesis methods.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-25 20:55 UTC pith:RHJRVNES
load-bearing objection The paper releases the first real gastroscopy NVS dataset and runs 3DGS baselines on it, but provides little detail on data quality or acquisition. the 2 major comments →
Gastroendoscopy View Synthesis: A New Real Dataset and Evaluation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
We present the first real gastroscopy dataset for NVS, namely the GastroNVS dataset, which contains a set of gastroscopic images, camera poses, and a point cloud for real gastroendoscopy inspection. To assess the suitability of the GastroNVS dataset, we evaluate several 3DGS methods and discuss the challenges for future development.
What carries the argument
The GastroNVS dataset, consisting of real gastroscopic images, camera poses, and a point cloud, serves as a benchmark for novel view synthesis methods in gastroendoscopy.
Load-bearing premise
The collected images, poses, and point cloud are of sufficient quality and coverage to serve as a meaningful benchmark for evaluating novel view synthesis methods in gastroendoscopy.
What would settle it
If 3DGS methods trained or tested on the dataset produce novel views that fail to match held-out real gastroscopic images in visual fidelity or geometric consistency, the dataset would not function as a useful benchmark.
If this is right
- Novel view synthesis methods can be directly tested on authentic gastroendoscopy scenes rather than synthetic or unrelated data.
- Applications such as extending the field of view of endoscopic images become possible to evaluate quantitatively.
- Digital twins for 3D archiving and endoscopist manipulation training can be developed using real captured data.
- Specific challenges in applying 3DGS to gastroendoscopy, such as narrow fields of view or lighting conditions, are now measurable.
Where Pith is reading between the lines
- The dataset could support adaptation of NVS techniques to handle specular highlights and tissue deformation common in real endoscopy.
- Combining the real GastroNVS data with existing synthetic endoscopy simulators might improve generalization of synthesis models.
- Release of the dataset could enable clinical researchers to prototype real-time 3D reconstruction tools without starting from scratch.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript claims to introduce the first real gastroscopy dataset for novel view synthesis (NVS), named GastroNVS, comprising gastroscopic images, camera poses, and a point cloud. It evaluates several 3D Gaussian splatting (3DGS) methods on the dataset and discusses domain-specific challenges for future NVS development in gastroendoscopy applications such as extended field-of-view imaging and digital twins.
Significance. If the released data prove to have adequate quality, pose accuracy, and coverage, the dataset would provide a valuable real-world benchmark in a medical imaging domain where synthetic or insufficient data have previously limited NVS research. The initial 3DGS evaluations could help identify gastroendoscopy-specific difficulties not captured in standard NVS benchmarks.
major comments (2)
- [Abstract / Dataset Description] Abstract and Dataset section: the claim that GastroNVS constitutes a suitable benchmark for NVS evaluation rests on unstated quantitative properties of the collected images, poses, and point cloud (e.g., image resolution, pose estimation reprojection error, scene overlap statistics, or coverage metrics); without these, the suitability assertion cannot be verified from the manuscript.
- [Evaluation] Evaluation section: the assessment of 3DGS methods is framed as a discussion of challenges rather than a quantitative benchmark study; if tables of metrics (PSNR, SSIM, LPIPS) are present, they should include direct comparisons to established NVS datasets to isolate gastroendoscopy-specific issues from general method limitations.
minor comments (1)
- [Abstract / Conclusion] Dataset availability statement: the note that data are available 'on request from our project page' should specify the exact access procedure, licensing terms, and any required agreements to support reproducibility.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback. We address the two major comments below regarding the need for quantitative dataset properties and the framing of the evaluation.
read point-by-point responses
-
Referee: [Abstract / Dataset Description] Abstract and Dataset section: the claim that GastroNVS constitutes a suitable benchmark for NVS evaluation rests on unstated quantitative properties of the collected images, poses, and point cloud (e.g., image resolution, pose estimation reprojection error, scene overlap statistics, or coverage metrics); without these, the suitability assertion cannot be verified from the manuscript.
Authors: We agree that explicit quantitative properties are needed to substantiate the benchmark claim. The revised manuscript will add image resolution, pose estimation reprojection error, scene overlap statistics, and coverage metrics to the Dataset section. revision: yes
-
Referee: [Evaluation] Evaluation section: the assessment of 3DGS methods is framed as a discussion of challenges rather than a quantitative benchmark study; if tables of metrics (PSNR, SSIM, LPIPS) are present, they should include direct comparisons to established NVS datasets to isolate gastroendoscopy-specific issues from general method limitations.
Authors: The manuscript already reports PSNR, SSIM, and LPIPS for 3DGS methods on GastroNVS. The evaluation intentionally emphasizes domain-specific challenges (e.g., specularities, narrow FOV, tissue deformation) rather than cross-dataset benchmarking, as direct numerical comparisons to datasets like LLFF would not isolate gastroendoscopy issues due to fundamental scene differences. We will revise the section to better clarify this intent and add a short note referencing typical metric ranges on standard datasets. revision: partial
Circularity Check
Dataset release paper with no derivations or fitted predictions
full rationale
The paper presents the GastroNVS dataset (images, poses, point cloud) and evaluates existing 3DGS methods on it to discuss challenges for NVS in gastroendoscopy. No equations, parameter fitting, predictions, or derivations are described that could reduce to inputs by construction. The central claim is the dataset release itself, stated directly without self-citation chains or ansatzes. This is a standard non-circular contribution for a benchmark dataset paper.
Axiom & Free-Parameter Ledger
read the original abstract
Novel view synthesis (NVS) is an active research topic in computer vision, owing to the success of neural radiance field (NeRF) and 3D Gaussian splatting (3DGS) methods. While NVS opens the door to potential applications in gastroendoscopy, such as extending the field of view of endoscopic images and enabling digital twins for 3D archiving and endoscopist manipulation training, the dataset is insufficient to evaluate NVS for gastroendoscopy. In this paper, we present the first real gastroscopy dataset for NVS, namely the GastroNVS dataset, which contains a set of gastroscopic images, camera poses, and a point cloud for real gastroendoscopy inspection. To assess the suitability of the GastroNVS dataset, we evaluate several 3DGS methods and discuss the challenges for future development. The dataset is available on request from our project page.
Figures
Reference graph
Works this paper leans on
-
[1]
NeRF: Representing scenes as neural radiance fields for view synthesis,
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: Representing scenes as neural radiance fields for view synthesis,” inProc. of European Conference on Computer Vision (ECCV), 2020, pp. 405–421
2020
-
[2]
3D Gaussian splatting for real-time radiance field rendering
B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3D Gaussian splatting for real-time radiance field rendering.”ACM Trans. Graph., vol. 42, no. 4, pp. 139–1–14, 2023
2023
-
[3]
ColonNeRF: High-fidelity neural reconstruction of long colonoscopy,
Y . Shi, B. Lu, J.-W. Liu, M. Li, S. Y . Yeo, and M. Z. Shou, “ColonNeRF: High-fidelity neural reconstruction of long colonoscopy,”Neurocomput- ing, vol. 657, pp. 131 445–1–15, 2025
2025
-
[4]
Realistic endoscopic illumination modeling for nerf-based data generation,
D. Psychogyios, F. Vasconcelos, and D. Stoyanov, “Realistic endoscopic illumination modeling for nerf-based data generation,” inProc. of International Conference on Medical Image Computing and Computer- Assisted Intervention (MICCAI), 2023, pp. 535–544
2023
-
[5]
Gaussian pancakes: Geometrically-regularized 3D gaussian splatting for realistic endoscopic reconstruction,
S. Bonilla, S. Zhang, D. Psychogyios, D. Stoyanov, F. Vasconcelos, and S. Bano, “Gaussian pancakes: Geometrically-regularized 3D gaussian splatting for realistic endoscopic reconstruction,” inProc. of Interna- tional Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2024, pp. 274–283
2024
-
[6]
PR-ENDO: Physically based relightable gaussian splatting for endoscopy,
J. Kaleta, W. Smolak-Dy ˙zewska, D. Malarz, D. Dall’Alba, P. Ko- rzeniowski, and P. Spurek, “PR-ENDO: Physically based relightable gaussian splatting for endoscopy,” inProc. of International Confer- ence on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2025, pp. 391–401
2025
-
[7]
H. Wang, Y . Zhou, H. Zhao, R. Wang, Q. Hu, X. Zhang, Q. Li, and Z. Wang, “Moving light adaptive colonoscopy reconstruction via illumination-attenuation-aware 3D Gaussian splatting,”arXiv preprint arXiv:2510.18739, 2025
-
[8]
Neural radiance fields for novel view synthesis in monocular gastroscopy,
Z. Jiang, Y . Monno, M. Okutomi, S. Suzuki, and K. Miki, “Neural radiance fields for novel view synthesis in monocular gastroscopy,” in Proc. of Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), 2024, pp. 1–5
2024
-
[9]
Texturing endoscopic 3D stomach via neural radiance field under uneven lighting,
Z.-Y . Lu, R.-H. Shiue, M.-L. Han, K.-C. Yen, and H. Chen, “Texturing endoscopic 3D stomach via neural radiance field under uneven lighting,” inProc. of IEEE International Conference on Image Processing (ICIP), 2025, pp. 385–390
2025
-
[10]
Stereo correspon- dence and reconstruction of endoscopic data challenge,
M. Allan, J. Mcleod, C. Wang, J. C. Rosenthal, Z. Hu, N. Gard, P. Eisert, K. X. Fu, T. Zeffiro, W. Xiaet al., “Stereo correspon- dence and reconstruction of endoscopic data challenge,”arXiv preprint arXiv:2101.01133, 2021
-
[11]
Neural rendering for stereo 3d reconstruction of deformable tissues in robotic surgery,
Y . Wang, Y . Long, S. H. Fan, and Q. Dou, “Neural rendering for stereo 3d reconstruction of deformable tissues in robotic surgery,” inProc. of International conference on medical image computing and computer- assisted intervention (MICCAI), 2022, pp. 431–441
2022
-
[12]
EndoMapper dataset of complete calibrated endoscopy proce- dures,
P. Azagra, C. Sostres, ´A. Ferr ´andez, L. Riazuelo, C. Tomasini, O. L. Barbed, J. Morlana, D. Recasens, V . M. Batlle, J. J. G ´omez-Rodr´ıguez et al., “EndoMapper dataset of complete calibrated endoscopy proce- dures,”Scientific Data, vol. 10, no. 1, pp. 671–1–16, 2023
2023
-
[13]
Structure-from-motion revisited,
J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revisited,” in Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4104–4113
2016
-
[14]
Colonoscopy 3D video dataset with paired depth from 2D- 3D registration,
T. L. Bobrow, M. Golhar, R. Vijayan, V . S. Akshintala, J. R. Garcia, and N. J. Durr, “Colonoscopy 3D video dataset with paired depth from 2D- 3D registration,”Medical Image Analysis, vol. 90, pp. 102 956–1–11, 2023
2023
-
[15]
Whole stomach 3D reconstruction and frame localization from monoc- ular endoscope video,
A. R. Widya, Y . Monno, M. Okutomi, S. Suzuki, T. Gotoda, and K. Miki, “Whole stomach 3D reconstruction and frame localization from monoc- ular endoscope video,”IEEE Journal of Translational Engineering in Health and Medicine, vol. 7, pp. 1–10, 2019
2019
-
[16]
Distinctive image features from scale-invariant keypoints,
D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, 2004
2004
-
[17]
EndoGS: Deformable endoscopic tissues reconstruction with Gaussian splatting,
L. Zhu, Z. Wang, J. Cui, Z. Jin, G. Lin, and L. Yu, “EndoGS: Deformable endoscopic tissues reconstruction with Gaussian splatting,” inProc. of International Conference on Medical Image Computing and Computer- Assisted Intervention (MICCAI), 2024, pp. 135–145
2024
-
[18]
EndoGaussian: Real-time Gaussian splatting for dynamic endoscopic scene reconstruction,
Y . Liu, C. Li, C. Yang, and Y . Yuan, “EndoGaussian: Real-time Gaussian splatting for dynamic endoscopic scene reconstruction,”arXiv preprint arXiv:2401.12561, 2024
-
[19]
Endo- 4DGS: Endoscopic monocular scene reconstruction with 4D Gaussian splatting,
Y . Huang, B. Cui, L. Bai, Z. Guo, M. Xu, M. Islam, and H. Ren, “Endo- 4DGS: Endoscopic monocular scene reconstruction with 4D Gaussian splatting,” inProc. of International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2024, pp. 197–207
2024
-
[20]
HFGS: 4D Gaussian splatting with emphasis on spatial and temporal high-frequency compo- nents for endoscopic scene reconstruction,
H. Zhao, X. Zhao, L. Zhu, W. Zheng, and Y . Xu, “HFGS: 4D Gaussian splatting with emphasis on spatial and temporal high-frequency compo- nents for endoscopic scene reconstruction,” inProc. of British Machine Vision Conference (BMVC), 2024, pp. 39–1–13
2024
-
[21]
RNNSLAM: Reconstructing the 3D colon to visualize missing regions during a colonoscopy,
R. Ma, R. Wang, Y . Zhang, S. Pizer, S. K. McGill, J. Rosenman, and J.-M. Frahm, “RNNSLAM: Reconstructing the 3D colon to visualize missing regions during a colonoscopy,”Medical Image Analysis, vol. 72, pp. 102 100–1–15, 2021
2021
-
[22]
Depth anything v2,
L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” inProc. of Advances in Neural Information Processing Systems (NeurIPS), 2024, pp. 21 875–21 911
2024
-
[23]
A hierarchical 3D gaussian representation for real-time rendering of very large datasets,
B. Kerbl, A. Meuleman, G. Kopanas, M. Wimmer, A. Lanvin, and G. Drettakis, “A hierarchical 3D gaussian representation for real-time rendering of very large datasets,”ACM Transactions on Graphics, vol. 43, no. 4, pp. 1–15, 2024
2024
-
[24]
2D Gaussian splatting for geometrically accurate radiance fields,
B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao, “2D Gaussian splatting for geometrically accurate radiance fields,” inProc. of ACM SIGGRAPH, 2024, pp. 1–11
2024
-
[25]
PGSR: Planar-based Gaussian splatting for efficient and high-fidelity surface reconstruction,
D. Chen, H. Li, W. Ye, Y . Wang, W. Xie, S. Zhai, N. Wang, H. Liu, H. Bao, and G. Zhang, “PGSR: Planar-based Gaussian splatting for efficient and high-fidelity surface reconstruction,”IEEE Transactions on Visualization and Computer Graphics, no. 9, pp. 6100–6111, 2024
2024
-
[26]
GSDF: 3DGS meets SDF for improved neural rendering and reconstruction,
M. Yu, T. Lu, L. Xu, L. Jiang, Y . Xiangli, and B. Dai, “GSDF: 3DGS meets SDF for improved neural rendering and reconstruction,” vol. 37, 2024, pp. 129 507–129 530
2024
-
[27]
NeuS: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,
P. Wang, L. Liu, Y . Liu, C. Theobalt, T. Komura, and W. Wang, “NeuS: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,” inProc. of Advances in Neural Information Processing Systems (NeurIPS), 2021, pp. 27 171–27 183
2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.