Pith. sign in

REVIEW 2 major objections 1 minor 27 references

The first real gastroscopy dataset enables evaluation of novel view synthesis methods.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-25 20:55 UTC pith:RHJRVNES

load-bearing objection The paper releases the first real gastroscopy NVS dataset and runs 3DGS baselines on it, but provides little detail on data quality or acquisition. the 2 major comments →

arxiv 2606.25427 v1 pith:RHJRVNES submitted 2026-06-24 cs.CV

Gastroendoscopy View Synthesis: A New Real Dataset and Evaluation

classification cs.CV
keywords gastroendoscopynovel view synthesis3D Gaussian splattingdatasetcamera posespoint cloudendoscopyNeRF
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper introduces the GastroNVS dataset, the first collection of real gastroscopic images paired with camera poses and a point cloud. The goal is to provide a benchmark for novel view synthesis techniques in gastroendoscopy, where such data was previously lacking. The authors evaluate several 3D Gaussian splatting methods on the dataset and identify challenges for further development. If the dataset proves suitable, it opens the way for applications like wider field of view in endoscopic images and digital twins for training.

Core claim

We present the first real gastroscopy dataset for NVS, namely the GastroNVS dataset, which contains a set of gastroscopic images, camera poses, and a point cloud for real gastroendoscopy inspection. To assess the suitability of the GastroNVS dataset, we evaluate several 3DGS methods and discuss the challenges for future development.

What carries the argument

The GastroNVS dataset, consisting of real gastroscopic images, camera poses, and a point cloud, serves as a benchmark for novel view synthesis methods in gastroendoscopy.

Load-bearing premise

The collected images, poses, and point cloud are of sufficient quality and coverage to serve as a meaningful benchmark for evaluating novel view synthesis methods in gastroendoscopy.

What would settle it

If 3DGS methods trained or tested on the dataset produce novel views that fail to match held-out real gastroscopic images in visual fidelity or geometric consistency, the dataset would not function as a useful benchmark.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Novel view synthesis methods can be directly tested on authentic gastroendoscopy scenes rather than synthetic or unrelated data.
  • Applications such as extending the field of view of endoscopic images become possible to evaluate quantitatively.
  • Digital twins for 3D archiving and endoscopist manipulation training can be developed using real captured data.
  • Specific challenges in applying 3DGS to gastroendoscopy, such as narrow fields of view or lighting conditions, are now measurable.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The dataset could support adaptation of NVS techniques to handle specular highlights and tissue deformation common in real endoscopy.
  • Combining the real GastroNVS data with existing synthetic endoscopy simulators might improve generalization of synthesis models.
  • Release of the dataset could enable clinical researchers to prototype real-time 3D reconstruction tools without starting from scratch.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The manuscript claims to introduce the first real gastroscopy dataset for novel view synthesis (NVS), named GastroNVS, comprising gastroscopic images, camera poses, and a point cloud. It evaluates several 3D Gaussian splatting (3DGS) methods on the dataset and discusses domain-specific challenges for future NVS development in gastroendoscopy applications such as extended field-of-view imaging and digital twins.

Significance. If the released data prove to have adequate quality, pose accuracy, and coverage, the dataset would provide a valuable real-world benchmark in a medical imaging domain where synthetic or insufficient data have previously limited NVS research. The initial 3DGS evaluations could help identify gastroendoscopy-specific difficulties not captured in standard NVS benchmarks.

major comments (2)
  1. [Abstract / Dataset Description] Abstract and Dataset section: the claim that GastroNVS constitutes a suitable benchmark for NVS evaluation rests on unstated quantitative properties of the collected images, poses, and point cloud (e.g., image resolution, pose estimation reprojection error, scene overlap statistics, or coverage metrics); without these, the suitability assertion cannot be verified from the manuscript.
  2. [Evaluation] Evaluation section: the assessment of 3DGS methods is framed as a discussion of challenges rather than a quantitative benchmark study; if tables of metrics (PSNR, SSIM, LPIPS) are present, they should include direct comparisons to established NVS datasets to isolate gastroendoscopy-specific issues from general method limitations.
minor comments (1)
  1. [Abstract / Conclusion] Dataset availability statement: the note that data are available 'on request from our project page' should specify the exact access procedure, licensing terms, and any required agreements to support reproducibility.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback. We address the two major comments below regarding the need for quantitative dataset properties and the framing of the evaluation.

read point-by-point responses
  1. Referee: [Abstract / Dataset Description] Abstract and Dataset section: the claim that GastroNVS constitutes a suitable benchmark for NVS evaluation rests on unstated quantitative properties of the collected images, poses, and point cloud (e.g., image resolution, pose estimation reprojection error, scene overlap statistics, or coverage metrics); without these, the suitability assertion cannot be verified from the manuscript.

    Authors: We agree that explicit quantitative properties are needed to substantiate the benchmark claim. The revised manuscript will add image resolution, pose estimation reprojection error, scene overlap statistics, and coverage metrics to the Dataset section. revision: yes

  2. Referee: [Evaluation] Evaluation section: the assessment of 3DGS methods is framed as a discussion of challenges rather than a quantitative benchmark study; if tables of metrics (PSNR, SSIM, LPIPS) are present, they should include direct comparisons to established NVS datasets to isolate gastroendoscopy-specific issues from general method limitations.

    Authors: The manuscript already reports PSNR, SSIM, and LPIPS for 3DGS methods on GastroNVS. The evaluation intentionally emphasizes domain-specific challenges (e.g., specularities, narrow FOV, tissue deformation) rather than cross-dataset benchmarking, as direct numerical comparisons to datasets like LLFF would not isolate gastroendoscopy issues due to fundamental scene differences. We will revise the section to better clarify this intent and add a short note referencing typical metric ranges on standard datasets. revision: partial

Circularity Check

0 steps flagged

Dataset release paper with no derivations or fitted predictions

full rationale

The paper presents the GastroNVS dataset (images, poses, point cloud) and evaluates existing 3DGS methods on it to discuss challenges for NVS in gastroendoscopy. No equations, parameter fitting, predictions, or derivations are described that could reduce to inputs by construction. The central claim is the dataset release itself, stated directly without self-citation chains or ansatzes. This is a standard non-circular contribution for a benchmark dataset paper.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Dataset paper with no mathematical model, fitted parameters, or new postulated entities; contribution rests on data collection and release.

pith-pipeline@v0.9.1-grok · 5689 in / 1051 out tokens · 16969 ms · 2026-06-25T20:55:05.181700+00:00 · methodology

0 comments
read the original abstract

Novel view synthesis (NVS) is an active research topic in computer vision, owing to the success of neural radiance field (NeRF) and 3D Gaussian splatting (3DGS) methods. While NVS opens the door to potential applications in gastroendoscopy, such as extending the field of view of endoscopic images and enabling digital twins for 3D archiving and endoscopist manipulation training, the dataset is insufficient to evaluate NVS for gastroendoscopy. In this paper, we present the first real gastroscopy dataset for NVS, namely the GastroNVS dataset, which contains a set of gastroscopic images, camera poses, and a point cloud for real gastroendoscopy inspection. To assess the suitability of the GastroNVS dataset, we evaluate several 3DGS methods and discuss the challenges for future development. The dataset is available on request from our project page.

Figures

Figures reproduced from arXiv: 2606.25427 by Masaki Minai, Masatoshi Okutomi, Sho Suzuki, Yusuke Monno.

Figure 1
Figure 1. Figure 1: (a) The cropping of the original image to input to [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Examples of the input images to SfM and the reconstructed 3D point cloud for each sequence. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Test views of Split-Con on Sequence B and C. on images covering the entire sequence, and therefore, it is expected to yield high-quality NVS results. However, because the test views are spatially and temporally close to the training views, Split-Reg makes it difficult to evaluate the model’s performance for really “novel” views. Thus, the second split, Split-Con, as shown in [PITH_FULL_IMAGE:figures/full_… view at source ↗
Figure 4
Figure 4. Figure 4: Visual comparisons for the test views of [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Visual comparison for the test view of Split-Con. learning surface geometry are more effective for NVS in the gastroendoscopic scene. Moreover, the results revealed an endoscopy-specific challenge caused by inter-view illu￾mination inconsistency, which leads to significant brightness discrepancies and degraded reconstruction quality in a certain sequence. We believe that GastroNVS is a valuable dataset for… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

27 extracted references · 3 canonical work pages

  1. [1]

    NeRF: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: Representing scenes as neural radiance fields for view synthesis,” inProc. of European Conference on Computer Vision (ECCV), 2020, pp. 405–421

  2. [2]

    3D Gaussian splatting for real-time radiance field rendering

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3D Gaussian splatting for real-time radiance field rendering.”ACM Trans. Graph., vol. 42, no. 4, pp. 139–1–14, 2023

  3. [3]

    ColonNeRF: High-fidelity neural reconstruction of long colonoscopy,

    Y . Shi, B. Lu, J.-W. Liu, M. Li, S. Y . Yeo, and M. Z. Shou, “ColonNeRF: High-fidelity neural reconstruction of long colonoscopy,”Neurocomput- ing, vol. 657, pp. 131 445–1–15, 2025

  4. [4]

    Realistic endoscopic illumination modeling for nerf-based data generation,

    D. Psychogyios, F. Vasconcelos, and D. Stoyanov, “Realistic endoscopic illumination modeling for nerf-based data generation,” inProc. of International Conference on Medical Image Computing and Computer- Assisted Intervention (MICCAI), 2023, pp. 535–544

  5. [5]

    Gaussian pancakes: Geometrically-regularized 3D gaussian splatting for realistic endoscopic reconstruction,

    S. Bonilla, S. Zhang, D. Psychogyios, D. Stoyanov, F. Vasconcelos, and S. Bano, “Gaussian pancakes: Geometrically-regularized 3D gaussian splatting for realistic endoscopic reconstruction,” inProc. of Interna- tional Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2024, pp. 274–283

  6. [6]

    PR-ENDO: Physically based relightable gaussian splatting for endoscopy,

    J. Kaleta, W. Smolak-Dy ˙zewska, D. Malarz, D. Dall’Alba, P. Ko- rzeniowski, and P. Spurek, “PR-ENDO: Physically based relightable gaussian splatting for endoscopy,” inProc. of International Confer- ence on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2025, pp. 391–401

  7. [7]

    Moving light adaptive colonoscopy reconstruction via illumination-attenuation-aware 3D Gaussian splatting,

    H. Wang, Y . Zhou, H. Zhao, R. Wang, Q. Hu, X. Zhang, Q. Li, and Z. Wang, “Moving light adaptive colonoscopy reconstruction via illumination-attenuation-aware 3D Gaussian splatting,”arXiv preprint arXiv:2510.18739, 2025

  8. [8]

    Neural radiance fields for novel view synthesis in monocular gastroscopy,

    Z. Jiang, Y . Monno, M. Okutomi, S. Suzuki, and K. Miki, “Neural radiance fields for novel view synthesis in monocular gastroscopy,” in Proc. of Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), 2024, pp. 1–5

  9. [9]

    Texturing endoscopic 3D stomach via neural radiance field under uneven lighting,

    Z.-Y . Lu, R.-H. Shiue, M.-L. Han, K.-C. Yen, and H. Chen, “Texturing endoscopic 3D stomach via neural radiance field under uneven lighting,” inProc. of IEEE International Conference on Image Processing (ICIP), 2025, pp. 385–390

  10. [10]

    Stereo correspon- dence and reconstruction of endoscopic data challenge,

    M. Allan, J. Mcleod, C. Wang, J. C. Rosenthal, Z. Hu, N. Gard, P. Eisert, K. X. Fu, T. Zeffiro, W. Xiaet al., “Stereo correspon- dence and reconstruction of endoscopic data challenge,”arXiv preprint arXiv:2101.01133, 2021

  11. [11]

    Neural rendering for stereo 3d reconstruction of deformable tissues in robotic surgery,

    Y . Wang, Y . Long, S. H. Fan, and Q. Dou, “Neural rendering for stereo 3d reconstruction of deformable tissues in robotic surgery,” inProc. of International conference on medical image computing and computer- assisted intervention (MICCAI), 2022, pp. 431–441

  12. [12]

    EndoMapper dataset of complete calibrated endoscopy proce- dures,

    P. Azagra, C. Sostres, ´A. Ferr ´andez, L. Riazuelo, C. Tomasini, O. L. Barbed, J. Morlana, D. Recasens, V . M. Batlle, J. J. G ´omez-Rodr´ıguez et al., “EndoMapper dataset of complete calibrated endoscopy proce- dures,”Scientific Data, vol. 10, no. 1, pp. 671–1–16, 2023

  13. [13]

    Structure-from-motion revisited,

    J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revisited,” in Proc. of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4104–4113

  14. [14]

    Colonoscopy 3D video dataset with paired depth from 2D- 3D registration,

    T. L. Bobrow, M. Golhar, R. Vijayan, V . S. Akshintala, J. R. Garcia, and N. J. Durr, “Colonoscopy 3D video dataset with paired depth from 2D- 3D registration,”Medical Image Analysis, vol. 90, pp. 102 956–1–11, 2023

  15. [15]

    Whole stomach 3D reconstruction and frame localization from monoc- ular endoscope video,

    A. R. Widya, Y . Monno, M. Okutomi, S. Suzuki, T. Gotoda, and K. Miki, “Whole stomach 3D reconstruction and frame localization from monoc- ular endoscope video,”IEEE Journal of Translational Engineering in Health and Medicine, vol. 7, pp. 1–10, 2019

  16. [16]

    Distinctive image features from scale-invariant keypoints,

    D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International Journal of Computer Vision, vol. 60, no. 2, pp. 91–110, 2004

  17. [17]

    EndoGS: Deformable endoscopic tissues reconstruction with Gaussian splatting,

    L. Zhu, Z. Wang, J. Cui, Z. Jin, G. Lin, and L. Yu, “EndoGS: Deformable endoscopic tissues reconstruction with Gaussian splatting,” inProc. of International Conference on Medical Image Computing and Computer- Assisted Intervention (MICCAI), 2024, pp. 135–145

  18. [18]

    EndoGaussian: Real-time Gaussian splatting for dynamic endoscopic scene reconstruction,

    Y . Liu, C. Li, C. Yang, and Y . Yuan, “EndoGaussian: Real-time Gaussian splatting for dynamic endoscopic scene reconstruction,”arXiv preprint arXiv:2401.12561, 2024

  19. [19]

    Endo- 4DGS: Endoscopic monocular scene reconstruction with 4D Gaussian splatting,

    Y . Huang, B. Cui, L. Bai, Z. Guo, M. Xu, M. Islam, and H. Ren, “Endo- 4DGS: Endoscopic monocular scene reconstruction with 4D Gaussian splatting,” inProc. of International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2024, pp. 197–207

  20. [20]

    HFGS: 4D Gaussian splatting with emphasis on spatial and temporal high-frequency compo- nents for endoscopic scene reconstruction,

    H. Zhao, X. Zhao, L. Zhu, W. Zheng, and Y . Xu, “HFGS: 4D Gaussian splatting with emphasis on spatial and temporal high-frequency compo- nents for endoscopic scene reconstruction,” inProc. of British Machine Vision Conference (BMVC), 2024, pp. 39–1–13

  21. [21]

    RNNSLAM: Reconstructing the 3D colon to visualize missing regions during a colonoscopy,

    R. Ma, R. Wang, Y . Zhang, S. Pizer, S. K. McGill, J. Rosenman, and J.-M. Frahm, “RNNSLAM: Reconstructing the 3D colon to visualize missing regions during a colonoscopy,”Medical Image Analysis, vol. 72, pp. 102 100–1–15, 2021

  22. [22]

    Depth anything v2,

    L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao, “Depth anything v2,” inProc. of Advances in Neural Information Processing Systems (NeurIPS), 2024, pp. 21 875–21 911

  23. [23]

    A hierarchical 3D gaussian representation for real-time rendering of very large datasets,

    B. Kerbl, A. Meuleman, G. Kopanas, M. Wimmer, A. Lanvin, and G. Drettakis, “A hierarchical 3D gaussian representation for real-time rendering of very large datasets,”ACM Transactions on Graphics, vol. 43, no. 4, pp. 1–15, 2024

  24. [24]

    2D Gaussian splatting for geometrically accurate radiance fields,

    B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao, “2D Gaussian splatting for geometrically accurate radiance fields,” inProc. of ACM SIGGRAPH, 2024, pp. 1–11

  25. [25]

    PGSR: Planar-based Gaussian splatting for efficient and high-fidelity surface reconstruction,

    D. Chen, H. Li, W. Ye, Y . Wang, W. Xie, S. Zhai, N. Wang, H. Liu, H. Bao, and G. Zhang, “PGSR: Planar-based Gaussian splatting for efficient and high-fidelity surface reconstruction,”IEEE Transactions on Visualization and Computer Graphics, no. 9, pp. 6100–6111, 2024

  26. [26]

    GSDF: 3DGS meets SDF for improved neural rendering and reconstruction,

    M. Yu, T. Lu, L. Xu, L. Jiang, Y . Xiangli, and B. Dai, “GSDF: 3DGS meets SDF for improved neural rendering and reconstruction,” vol. 37, 2024, pp. 129 507–129 530

  27. [27]

    NeuS: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,

    P. Wang, L. Liu, Y . Liu, C. Theobalt, T. Komura, and W. Wang, “NeuS: Learning neural implicit surfaces by volume rendering for multi-view reconstruction,” inProc. of Advances in Neural Information Processing Systems (NeurIPS), 2021, pp. 27 171–27 183