Pith. sign in

REVIEW 6 major objections 4 minor 83 references

PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting

T0 review · 6 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read PanSplat generates 4K panorama novel views from two 360° images in about 0.34 seconds, up to 70× faster than prior NeRF-based state of the art.

desk verdict Solid feed-forward panorama splatting system; the 512x1024 results are real, but the 4K headline is a memory-feasibility claim rather than a quality one. read the letter →

arxiv 2412.12096 v2 pith:BPLQEM3P submitted 2024-12-16 cs.CV

classification cs.CV
keywords panoramanovelviewsynthesis3DGaussiansplattingfeed-forwardrendering4KresolutionFibonaccilatticesphericalcostvolumewide-baseline360-degreeimagesvirtualreality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that feed-forward novel view synthesis for 360° panoramas can move from low-resolution, slow NeRF-style rendering to interactive 4K. PanSplat takes two posed wide-baseline panoramas, predicts a spherical 3D Gaussian pyramid, and renders a new middle view at 2048×4096 in about 0.34 seconds, roughly 70× faster than the prior NeRF-based state of the art. The design matters because VR, virtual tours, and robot navigation all want high-resolution panoramas under tight memory and latency budgets, which previous methods could not satisfy. If the claims hold, immersive panorama rendering becomes real-time enough for practical deployment.

What carries the argument

The central object is a spherical 3D Gaussian pyramid: Gaussians arranged on a Fibonacci lattice, with the number of Gaussians per level set to $n_l = \lfloor W^2/(2^l \pi) \rfloor$ so density near the equator matches image pixels while avoiding pole crowding. A hierarchical spherical cost volume estimates depth coarse-to-fine at half the input resolution, and lightweight Gaussian heads predict opacity, covariance, color, and per-Gaussian depth residuals from full-resolution image features. A cubemap renderer splits rendering into six faces and stitches them into an equirectangular image, and the two-step deferred backpropagation caches image gradients, re-renders face by face, then re-generates Gaussian parameters tile by tile, which is what makes 4K training memory feasible.

What would settle it

Render the Matterport3D 4K test set with PanSplat and compare per-view PSNR, WS-PSNR, and LPIPS at full 4K against the same model rendering at 1024×2048 and at 512×1024, each upsampled to 4K; if the full-resolution render does not clearly beat the upsampled lower-resolution renders, the claim that the 4K pipeline adds image quality collapses.

Watch

Extended reading notes

Core claim

PanSplat establishes that a feed-forward network can synthesize high-quality 4K panorama views from just two 360° images by replacing pixel-aligned Gaussians with Gaussians placed on a Fibonacci lattice and organizing them into a four-level spherical pyramid. The lattice distributes Gaussians uniformly over the sphere, cutting Gaussian count by up to 36.34% relative to pixel alignment, while the pyramid lets coarse levels carry global structure and fine levels carry texture detail. Geometry comes from a hierarchical spherical cost volume computed at 512×1024, and the Gaussian heads use full-resolution images for appearance; a cubemap renderer with two-step deferred backpropagation makes 4K training fit on a single A100 GPU. On Matterport3D, Replica, and Residential, the method reports the best or second-best quality scores among compared feed-forward baselines, and after fine-tuning on 360Loc it also outperforms the adapted perspective baseline on real-world data with the help of deferred blending.

Load-bearing premise

The strongest load-bearing assumption is that low-resolution geometry is good enough: the cost volume sees only 512×1024 inputs even when the target is 2048×4096, and the paper supports the 4K claim mostly with qualitative images rather than 4K metrics.

Editorial extensions

If this is right

  • Novel views from two wide-baseline panoramas can be rendered at 2048×4096 in about 0.34 seconds, roughly 70× faster than PanoGRF's 23.8 seconds.
  • The model generalizes from Matterport3D to Replica, Residential, and fine-tuned real-world sequences, so 4K output is not limited to synthetic data.
  • Training at 4K fits on a single A100 GPU, and inference at 4K fits on a 24GB RTX 3090, making high-resolution panorama synthesis accessible without multi-GPU setups.
  • The Fibonacci lattice cuts the required Gaussian count by up to 36.34% without hurting quality, giving a leaner representation for spherical images.
  • Separating low-resolution geometry from high-resolution appearance lets the system scale to 4K while keeping the memory cost of the cost volume fixed.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Reading beyond the paper, the same geometry-appearance separation could be applied to perspective 4K view synthesis, since the memory bottleneck is not inherently tied to equirectangular projection.
  • The paper reports only qualitative evidence at full 4K, so a natural test is whether 4K metrics beat upsampled renders from the same model at 1K; I would expect much of the visible gain to come from texture resolution rather than geometry.
  • The fixed Fibonacci lattice is a design choice that trades uniform coverage for feature-grid alignment; a promising extension would be predicting per-Gaussian density adaptively rather than fixing the count per level.
  • Deferred blending suggests a cheap route toward dynamic scenes: instead of blending whole views by input distance, per-Gaussian opacity prediction could let the network discard moving content automatically.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

6 major / 4 minor

Summary. The paper proposes PanSplat, a feed-forward network for wide-baseline panorama novel view synthesis. It introduces a spherical 3D Gaussian pyramid with Fibonacci lattice sampling, a hierarchical spherical cost volume, light-weight Gaussian heads, and a two-step deferred backpropagation scheme that enables training at up to 2048×4096 on a single A100 GPU. The method is evaluated on Matterport3D, Replica, Residential, 360Loc, and a self-captured Insta360 dataset, reporting gains over PanoGRF and MVSplat at 512×1024, a 70× inference speedup over PanoGRF, and qualitative and memory-based evidence of 4K capability.

Significance. If the 4K results are confirmed, PanSplat would be a strong systems contribution to high-resolution panorama novel view synthesis, with a clean closed-form Gaussian count formula, an informative ablation study, and public code. The paper honestly reports the design tension between low-resolution geometry and high-resolution texture, and it provides several useful engineering techniques (tiled Gaussian heads, cubemap rendering, deferred backpropagation). The main weakness is that the headline 4K claim is currently supported only by memory measurements and qualitative examples, not by quantitative image quality or latency at 2048×4096.

major comments (6)
  1. [Sec. 4.2, Table 1] The sentence "PanSplat consistently outperforms all competing methods" is contradicted by the reported numbers: at the 2.0m baseline PanoGRF has WS-PSNR 20.96 vs. PanSplat 20.56, and on Residential MVSplat has WS-PSNR 31.21 vs. PanSplat 30.97. The abstract's "state-of-the-art" claim should be restricted to the metrics and baselines where it holds, or the comparison should be extended to show superiority on a consistent metric set.
  2. [Abstract and Sec. 4.2] The 70× speedup (23.8s vs 0.34s) is reported in a paragraph whose comparisons are explicitly "all at a resolution of 512×1024". No end-to-end inference time at 2048×4096 is reported, yet the title and abstract advertise 4K synthesis. Please report latency (and, if possible, throughput) at 4K on the same GPU, including the cost volume, Gaussian heads, and renderer, or revise the headline claim to "memory-efficient 4K training/rendering".
  3. [Sec. 3.1] The Fibonacci lattice sampling formula (xj, yj) = (j/φ mod 1, j/(n−1)) is area-uniform in the equirectangular image plane, not on the sphere. With a standard equirectangular unprojection the spherical point density scales as 1/sin θ, giving higher density near the poles, which is the opposite of the claimed redundancy reduction. If the unprojection instead uses an area-preserving mapping (e.g., cylindrical equal-area or a true spherical Fibonacci lattice using z = 1−2j/n), this must be stated explicitly; otherwise the central geometric motivation of the Fibonacci arrangement is not supported.
  4. [Sec. 4.1 and Table 2] The text states 360Loc has "an average baseline of 0.47 meters", while Table 2 labels the same evaluation "360Loc (avg. 1.40m baseline)". This discrepancy affects the interpretation of the real-world results and must be resolved. In addition, Table 2 compares only with MVSplat; the claim of "state-of-the-art ... across both synthetic and real-world datasets" is therefore not established for real-world data against PanoGRF or other baselines, and should be qualified.
  5. [Sec. 4.3, Sec. G, Sec. C] The 4K capability is demonstrated solely through training/inference GPU memory (Figs. 6 and G.1) and qualitative examples (Fig. 1 and the supplementary video). Even though a 4K Matterport3D test set of 10 samples is rendered (Sec. C), no WS-PSNR/SSIM/LPIPS or inference-time numbers at 2048×4096 are provided. Without quantitative 4K image quality, the claim that geometry estimated at 512×1024 (cost volume finest level 256×512) is sufficient for 4K rendering remains an assumption. Please add 4K metrics and compare with a low-resolution baseline or a strong upsampled alternative.
  6. [Tables 1–3] No variance or repeated-seed statistics are reported. Several margins are small (e.g., Replica LPIPS: 0.069 vs 0.059 for MVSplat; WS-PSNR 30.78 vs 30.54), so it is unclear which differences are significant. Please provide standard deviations over at least three training runs or a significance analysis, particularly for claims of state-of-the-art.
minor comments (4)
  1. [Abstract] The abstract reports the 70× speedup without a resolution qualifier; specify that the runtime comparison is at 512×1024.
  2. [Sec. 3.3] The claim that "image quality relies more on texture resolution than on geometry resolution" is presented as an observation without evidence; add a supporting ablation or reference.
  3. [Fig. G.1] The legend uses "w/ Deferred BP (1 step)" and "w/ Deferred BP (16 tiles)" and notes they overlap for inference; consider separate line styles or a clearer note so readers can distinguish the two curves.
  4. [Table 1] The top1/top2/top3 formatting is not visible in the text version; ensure the formatting is unambiguous in the camera-ready version.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the claims are supported by held-out benchmark measurements and ablations, not by derivation from the paper's own assumptions.

full rationale

PanSplat is an empirical systems paper. Its central claims are benchmark numbers measured on held-out test splits (Tables 1, 2, and D.1) and memory/latency measurements (Figs. 6 and G.1) against external baselines such as PanoGRF, MVSplat, IBRNet, NeuRay, OmniSyn, and S-NeRF, so no target result is fed back into a fitting procedure. The only closed-form quantity, the Fibonacci Gaussian count n = floor(W^2/pi), is a design choice for distributing Gaussians over the sphere; the reported 36.34% reduction relative to pixel-aligned splatting is arithmetic implied by that formula rather than a fitted or empirically discovered quantity, and the claim that quality is not compromised is checked by the +Fibo ablation in Table 3. The paper's own components are isolated by ablations that remove them: Fibonacci Gaussians, 3D Gaussian pyramid, hierarchical cost volume, residual design, and monocular depth features (Tables 3 and E.1). The self-citations that exist are not load-bearing: UniMatch [71] is used only to initialize the Swin Transformer, and MVSplat [16] is adapted as a baseline or for a 2D U-Net component, neither of which is invoked as a uniqueness argument or as proof of the central result. The absence of quantitative 4K quality and latency metrics is a completeness or correctness concern, not a circularity concern, because the paper does not claim to derive 4K quality from a self-referential argument. No step in the paper reduces, by its own equations or by self-citation, to its own inputs.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central method depends on hand-set hyperparameters (Gaussian count formula, pyramid levels, depth candidates, loss weights, unspecified depth range) and several unproven but testable engineering assumptions about uniform sampling, tile equivalence, deferred-backprop gradient fidelity, and transfer from synthetic to real panoramas. No new physical or mathematical entities are introduced.

free parameters (5)
  • Gaussian count formula n = floor(W^2 / pi) = e.g., 333,806 at width 1024
    Chosen by hand to match equatorial pixel density; sets the quality/compute trade-off and is not swept in the paper.
  • Number of pyramid levels L = 4
    Set to 4; ablated only as a whole (Full vs Base), not as a level-count sweep.
  • Coarse depth candidates D = 128, then 64, 32 at finer levels
    Chosen to balance memory and depth accuracy; no sensitivity analysis.
  • Loss weights gamma, lambda, alpha = gamma=0.9, lambda=0.1, alpha=0.05
    Set without sensitivity analysis or justification.
  • Depth range [d_min, d_max] = not specified
    The paper says 'a preset range' but never gives the values, so the geometric scale of the cost volume is not reproducible from the text.
assumptions (5)
  • standard math A Fibonacci lattice uniformly distributes points on a sphere without polar clustering.
    Invoked in Sec. 3.1 to justify replacing pixel-aligned Gaussians with lattice-based Gaussians.
  • domain assumption Pre-trained monocular depth features from PanoGRF transfer to the spherical cost volume and improve depth prediction.
    Used in Sec. 3.2 and Sec. C; the paper's own ablation (w/o Mono depth) shows only marginal benefit, yet the features are frozen and included in the full model.
  • domain assumption Gaussian parameters predicted at tile boundaries with 3-pixel padding and overlap are identical to non-tiled predictions.
    Sec. B asserts pre-padding and overlap reproduce non-tiled outputs, but no formal proof or empirical verification of exact equality is provided.
  • domain assumption Two-step deferred backpropagation, which caches gradients on the image and re-renders face by face, gives gradients equivalent to full backpropagation.
    Sec. B describes the procedure but does not prove exact equivalence; it is used as a memory-saving approximation during training.
  • domain assumption The synthetic Matterport3D rendering protocol from PanoGRF is a valid proxy for real-world 4K panorama distribution.
    All training and quantitative synthetic evaluation use this data; real-world fine-tuning is only on 360Loc at 4K, so the transferability of the full pipeline is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting." pith.science (2026). https://pith.science/paper/BPLQEM3P

@misc{pith2026241212096,
  author       = {Pith},
  title        = {Pith review of: PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BPLQEM3P}},
  note         = {Machine review of arXiv:2412.12096}
}
abstract

With the advent of portable 360{\deg} cameras, panorama has gained significant attention in applications like virtual reality (VR), virtual tours, robotics, and autonomous driving. As a result, wide-baseline panorama view synthesis has emerged as a vital task, where high resolution, fast inference, and memory efficiency are essential. Nevertheless, existing methods are typically constrained to lower resolutions (512 $\times$ 1024) due to demanding memory and computational requirements. In this paper, we present PanSplat, a generalizable, feed-forward approach that efficiently supports resolution up to 4K (2048 $\times$ 4096). Our approach features a tailored spherical 3D Gaussian pyramid with a Fibonacci lattice arrangement, enhancing image quality while reducing information redundancy. To accommodate the demands of high resolution, we propose a pipeline that integrates a hierarchical spherical cost volume and Gaussian heads with local operations, enabling two-step deferred backpropagation for memory-efficient training on a single A100 GPU. Experiments demonstrate that PanSplat achieves state-of-the-art results with superior efficiency and image quality across both synthetic and real-world datasets. Code is available at https://github.com/chengzhag/PanSplat.

Figures

Figures reproduced from arXiv: 2412.12096 by the authors.

Figure 1
Figure 1. Our PanSplat can generate novel views from two 4K (2048 × 4096) panoramas. We train on rendered Matterport3D [11] data at 4K resolution (left) and can generalize to 4K real-world data (right) with a few fine-tunings on 360Loc [30] data (Zoom in for details). Please refer to the supplementary video for more results. Abstract With the advent of portable 360° cameras, panorama has gained significant attention in applic… view at source ↗
Figure 2
Figure 2. Fibonacci Gaussians. We propose a Fibonacci lattice arrangement for the Gaussians to be distributed uniformly across the sphere, avoiding information redundancy near the poles, and significantly reducing the number of required Gaussians. also making significant progress in immersive content cre￾ation. While current methods have extensively explored wide-baseline panorama view synthesis, they often strug￾gle to balan… view at source ↗
Figure 3
Figure 3. Our proposed PanSplat pipeline. Given two wide-baseline panoramas, we first construct a hierarchical spherical cost volume (Sec. 3.2) using a Transformer-based FPN to extract feature pyramid and 2D U-Nets to integrate monocular depth priors for cost volume refinement. We then build Gaussian heads (Sec. 3.3) to generate a feature pyramid, which is later sampled with Fibonacci lattice and trans￾formed to spherical 3D … view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Qualitative comparisons on synthetic datasets. We show the input panorama pairs and the ground truth novel views on the left, and compare the zoomed-in results on the right to highlight the differences. Our PanSplat generates overall sharper images with more high-frequ…
Figure 5
Figure 5. Figure 5: Qualitative comparisons of ablation study. Our Fibonacci Gaussians (+Fibo) reduces Gaussian count without compromising image quality, and our 3D Gaussian Pyramid (+3DGP) further enhances quality. Setup #Gaussian (K) WS-PSNR↑ SSIM↑ LPIPS ↓ Base 1,049 (100%) 27.07 0.895 …
Figure 6
Figure 6. Figure 6: Training GPU memory consumption at different res￾olutions, where × indicates out-of-memory errors even on a 80GB A100. Memory consumption is tested with a batch size of 1. ates sharper images with more accurate geometry. It also demonstrates that the use of Fibo does n…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

83 extracted references · 59 canonical work pages

  1. [1]

    https://www.apple.com/maps/

    Maps - apple. https://www.apple.com/maps/. 1

  2. [2]

    https : / / observablehq

    Fibonacci Lattices / Amit Sch. https : / / observablehq . com / @meetamit / fibonacci - lattices. 3

  3. [3]

    https://www.google.com/streetview/

    Explore Street View and add your own 360 images to Google Maps. https://www.google.com/streetview/. 1

  4. [4]

    https://matterport.com

    Capture, share, and collaborate the built world in immersive 3D. https://matterport.com. 1

  5. [5]

    https: //www.theasys.io/

    Theasys - 360 VR Online Virtual Tour Creator. https: //www.theasys.io/. 1

  6. [6]

    Di- verse plausible 360-degree image outpainting for efficient 3dcg background creation

    Naofumi Akimoto, Yuhi Matsuo, and Yoshimitsu Aoki. Di- verse plausible 360-degree image outpainting for efficient 3dcg background creation. In CVPR, pages 11441–11450,

  7. [7]

    Matryodshka: Real-time 6dof video view synthesis using multi-sphere images

    Benjamin Attal, Selena Ling, Aaron Gokaslan, Christian Richardt, and James Tompkin. Matryodshka: Real-time 6dof video view synthesis using multi-sphere images. In ECCV, pages 441–459. Springer, 2020. 2

  8. [8]

    360-gs: Layout-guided panoramic gaussian splatting for indoor roaming

    Jiayang Bai, Letian Huang, Jie Guo, Wen Gong, Yuanqi Li, and Yanwen Guo. 360-gs: Layout-guided panoramic gaussian splatting for indoor roaming. arXiv preprint arXiv:2402.00763, 2024. 3, 5

Show all 83 references
  1. [9]

    Immersive light field video with a layered mesh representation

    Michael Broxton, John Flynn, Ryan Overbeck, Daniel Erick- son, Peter Hedman, Matthew Duvall, Jason Dourgarian, Jay Busch, Matt Whalen, and Paul Debevec. Immersive light field video with a layered mesh representation. ACM TOG, 39(4):86–1, 2020. 1

  2. [10]

    Neusis: A compo- sitional neuro-symbolic framework for autonomous percep- tion, reasoning, and planning in complex uav search mis- sions

    Zhixi Cai, Cristian Rojas Cardenas, Kevin Leo, Chenyuan Zhang, Kal Backman, Hanbing Li, Boying Li, Mahsa Ghor- banali, Stavya Datta, Lizhen Qu, et al. Neusis: A compo- sitional neuro-symbolic framework for autonomous percep- tion, reasoning, and planning in complex uav search ...

  3. [11]

    Matterport3d: Learning from rgb-d data in indoor environments

    Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Hal- ber, Matthias Niebner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. In 2017 International Confer- ence on 3D Vision (3DV), pages 667–676. IEEE, 201...

  4. [12]

    pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction

    David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In CVPR, pages 19457–19467, 2024. 2, 3

  5. [13]

    Mvsnerf: Fast general- izable radiance field reconstruction from multi-view stereo

    Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. Mvsnerf: Fast general- izable radiance field reconstruction from multi-view stereo. In ICCV, pages 14124–14133, 2021. 2

  6. [14]

    Casual 6-dof: free-viewpoint panorama using a handheld 360 camera

    Rongsen Chen, Fang-Lue Zhang, Simon Finnie, Andrew Chalmers, and Taehyun Rhee. Casual 6-dof: free-viewpoint panorama using a handheld 360 camera. IEEE TVCG, 29(9): 3976–3988, 2022. 3

  7. [15]

    Explicit correspondence matching for generalizable neural radiance fields

    Yuedong Chen, Haofei Xu, Qianyi Wu, Chuanxia Zheng, Tat-Jen Cham, and Jianfei Cai. Explicit correspondence matching for generalizable neural radiance fields. arXiv preprint arXiv:2304.12294, 2023. 2

  8. [16]

    Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images

    Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images. In ECCV, pages 370–386. Springer,

  9. [17]

    Panogrf: generaliz- able spherical radiance fields for wide-baseline panoramas

    Zheng Chen, Yan-Pei Cao, Yuan-Chen Guo, Chen Wang, Ying Shan, and Song-Hai Zhang. Panogrf: generaliz- able spherical radiance fields for wide-baseline panoramas. NeurIPS, 36:6961–6985, 2023. 2, 3, 4, 5, 6

  10. [18]

    Splatter-360: Generalizable 360◦ gaussian splat- ting for wide-baseline panoramic images

    Zheng Chen, Chenming Wu, Zhelun Shen, Chen Zhao, Weicai Ye, Haocheng Feng, Errui Ding, and Song-Hai Zhang. Splatter-360: Generalizable 360◦ gaussian splat- ting for wide-baseline panoramic images. arXiv preprint arXiv:2412.06250, 2024. 3

  11. [19]

    Bal- anced spherical grid for egocentric view synthesis

    Changwoon Choi, Sang Min Kim, and Young Min Kim. Bal- anced spherical grid for egocentric view synthesis. InCVPR, pages 16590–16599, 2023. 3

  12. [20]

    Om- nilocalrf: Omnidirectional local radiance fields from dy- namic videos

    Dongyoung Choi, Hyeonjoong Jang, and Min H Kim. Om- nilocalrf: Omnidirectional local radiance fields from dy- namic videos. In CVPR, pages 6871–6880, 2024. 3

  13. [21]

    Guided co-modulated gan for 360° field of view extrapolation

    Mohammad Reza Karimi Dastjerdi, Yannick Hold-Geoffroy, Jonathan Eisenmann, Siavash Khodadadeh, and Jean- Franc ¸ois Lalonde. Guided co-modulated gan for 360° field of view extrapolation. In 3DV, pages 475–485. IEEE, 2022. 3

  14. [22]

    Depth-supervised nerf: Fewer views and faster train- ing for free

    Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ra- manan. Depth-supervised nerf: Fewer views and faster train- ing for free. In CVPR, pages 12882–12891, 2022. 2

  15. [23]

    Panocontext-former: Panoramic total scene understand- ing with a transformer

    Yuan Dong, Chuan Fang, Liefeng Bo, Zilong Dong, and Ping Tan. Panocontext-former: Panoramic total scene understand- ing with a transformer. In CVPR, pages 28087–28097, 2024. 3

  16. [24]

    Deterministic gaussian sampling with generalized fibonacci grids

    Daniel Frisch and Uwe D Hanebeck. Deterministic gaussian sampling with generalized fibonacci grids. In 2021 IEEE 24th International Conference on Information Fusion (FU- SION), pages 1–8. IEEE, 2021. 3

  17. [25]

    Cascade cost volume for high-resolution multi-view stereo and stereo matching

    Xiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai, Feitong Tan, and Ping Tan. Cascade cost volume for high-resolution multi-view stereo and stereo matching. In CVPR, pages 2495–2504, 2020. 4

  18. [26]

    Somsi: Spherical novel view synthesis with soft occlusion multi-sphere images

    Tewodros Habtegebrial, Christiano Gava, Marcel Rogge, Di- dier Stricker, and Varun Jampani. Somsi: Spherical novel view synthesis with soft occlusion multi-sphere images. In CVPR, pages 15725–15734, 2022. 2, 3, 5

  19. [27]

    Text2room: Extracting textured 3d meshes from 2d text-to-image models

    Lukas H ¨ollein, Ang Cao, Andrew Owens, Justin Johnson, and Matthias Nießner. Text2room: Extracting textured 3d meshes from 2d text-to-image models. InICCV, pages 7909– 7920, 2023. 3

  20. [28]

    Image quality metrics: Psnr vs

    Alain Hore and Djemel Ziou. Image quality metrics: Psnr vs. ssim. In PR, pages 2366–2369. IEEE, 2010. 6

  21. [29]

    360roam: Real-time indoor roaming us- ing geometry-aware 360◦ radiance fields

    Huajian Huang, Yingshu Chen, Tianjian Zhang, and Sai- Kit Yeung. 360roam: Real-time indoor roaming us- ing geometry-aware 360◦ radiance fields. arXiv preprint arXiv:2208.02705, 2022. 3

  22. [30]

    360loc: A dataset and benchmark for omnidirectional visual localization with 9 cross-device queries

    Huajian Huang, Changkun Liu, Yipeng Zhu, Hui Cheng, Tristan Braud, and Sai-Kit Yeung. 360loc: A dataset and benchmark for omnidirectional visual localization with 9 cross-device queries. In CVPR, pages 22314–22324, 2024. 1, 5, 2

  23. [31]

    Adversarial generation of hierarchical gaussians for 3d generative model

    Sangeek Hyun and Jae-Pil Heo. Adversarial generation of hierarchical gaussians for 3d generative model. In NeurIPS,

  24. [32]

    Ego- centric scene reconstruction from an omnidirectional video

    Hyeonjoong Jang, Andreas Meuleman, Dahyun Kang, Donggun Kim, Christian Richardt, and Min H Kim. Ego- centric scene reconstruction from an omnidirectional video. ACM TOG, 41(4):1–12, 2022. 3

  25. [33]

    Unifuse: Unidirectional fusion for 360 panorama depth estimation

    Hualie Jiang, Zhe Sheng, Siyu Zhu, Zilong Dong, and Rui Huang. Unifuse: Unidirectional fusion for 360 panorama depth estimation. IEEE Robotics and Automation Letters, 6 (2):1519–1526, 2021. 4, 3

  26. [34]

    3d gaussian splatting for real-time radiance field rendering

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM TOG, 42(4):139–1, 2023. 2, 1

  27. [35]

    Omnisdf: Scene re- construction using omnidirectional signed distance functions and adaptive binoctrees

    Hakyeong Kim, Andreas Meuleman, Hyeonjoong Jang, James Tompkin, and Min H Kim. Omnisdf: Scene re- construction using omnidirectional signed distance functions and adaptive binoctrees. In CVPR, pages 20227–20236,

  28. [36]

    xformers: A modular and hackable trans- former modelling library

    Benjamin Lefaudeux, Francisco Massa, Diana Liskovich, Wenhan Xiong, Vittorio Caggiano, Sean Naren, Min Xu, Jieru Hu, Marta Tintore, Susan Zhang, Patrick Labatut, Daniel Haziza, Luca Wehrstedt, Jeremy Reizenstein, and Grigory Sizov. xformers: A modular and hackable trans- forme...

  29. [37]

    Textslam: Visual slam with planar text fea- tures

    Boying Li, Danping Zou, Daniele Sartori, Ling Pei, and Wenxian Yu. Textslam: Visual slam with planar text fea- tures. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 2102–2108. IEEE, 2020. 1

  30. [38]

    Structdepth: Leveraging the structural regularities for self-supervised indoor depth estimation

    Boying Li, Yuan Huang, Zeyu Liu, Danping Zou, and Wenx- ian Yu. Structdepth: Leveraging the structural regularities for self-supervised indoor depth estimation. In ICCV, pages 12663–12673, 2021

  31. [39]

    Textslam: Visual slam with semantic planar text features

    Boying Li, Danping Zou, Yuan Huang, Xinghan Niu, Ling Pei, and Wenxian Yu. Textslam: Visual slam with semantic planar text features. IEEE TPAMI, 46(1):593–610, 2023

  32. [40]

    Hi-slam: Scaling-up semantics in slam with a hierarchically categorical gaussian splatting

    Boying Li, Zhixi Cai, Yuan-Fang Li, Ian Reid, and Hamid Rezatofighi. Hi-slam: Scaling-up semantics in slam with a hierarchically categorical gaussian splatting. arXiv preprint arXiv:2409.12518, 2024

  33. [41]

    Hier-slam++: Neuro-symbolic seman- tic slam with a hierarchically categorical gaussian splatting

    Boying Li, Vuong Chi Hao, Peter J Stuckey, Ian Reid, and Hamid Rezatofighi. Hier-slam++: Neuro-symbolic seman- tic slam with a hierarchically categorical gaussian splatting. arXiv preprint arXiv:2502.14931, 2025. 1

  34. [42]

    Omnisyn: Synthesizing 360 videos with wide-baseline panoramas

    David Li, Yinda Zhang, Christian H ¨ane, Danhang Tang, Amitabh Varshney, and Ruofei Du. Omnisyn: Synthesizing 360 videos with wide-baseline panoramas. In 2022 IEEE Conference on Virtual Reality and 3D User Interfaces Ab- stracts and Workshops (VRW), pages 670–671. IEEE, 2022. ...

  35. [43]

    Ggrt: Towards pose-free gen- eralizable 3d gaussian splatting in real-time

    Hao Li, Yuanyuan Gao, Chenming Wu, Dingwen Zhang, Yalun Dai, Chen Zhao, Haocheng Feng, Errui Ding, Jing- dong Wang, and Junwei Han. Ggrt: Towards pose-free gen- eralizable 3d gaussian splatting in real-time. InECCV, pages 325–341. Springer, 2024. 2, 5

  36. [44]

    Extending 6-dof vr experience via multi- sphere images interpolation

    Jisheng Li, Yuze He, Jinghui Jiao, Yubin Hu, Yuxing Han, and Jiangtao Wen. Extending 6-dof vr experience via multi- sphere images interpolation. InACM MM, pages 4632–4640,

  37. [45]

    Omnigs: Omnidirectional gaussian splatting for fast radiance field reconstruction using omnidirectional images

    Longwei Li, Huajian Huang, Sai-Kit Yeung, and Hui Cheng. Omnigs: Omnidirectional gaussian splatting for fast radiance field reconstruction using omnidirectional images. arXiv preprint arXiv:2404.03202, 2024. 3, 5

  38. [46]

    Feature pyramid networks for object detection

    Tsung-Yi Lin, Piotr Doll ´ar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In CVPR, pages 2117–2125,

  39. [47]

    Neural rays for occlusion-aware image-based render- ing

    Yuan Liu, Sida Peng, Lingjie Liu, Qianqian Wang, Peng Wang, Christian Theobalt, Xiaowei Zhou, and Wenping Wang. Neural rays for occlusion-aware image-based render- ing. In CVPR, pages 7824–7833, 2022. 2, 3, 6

  40. [48]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, pages 10012–10022, 2021. 4, 1

  41. [49]

    Nerf: Representing scenes as neural radiance fields for view syn- thesis

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In ECCV, pages 405–421. Springer, 2020. 2, 6

  42. [50]

    Reg- nerf: Regularizing neural radiance fields for view synthesis from sparse inputs

    Michael Niemeyer, Jonathan T Barron, Ben Mildenhall, Mehdi SM Sajjadi, Andreas Geiger, and Noha Radwan. Reg- nerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In CVPR, pages 5480–5490, 2022. 2

  43. [51]

    Bips: Bi-modal in- door panorama synthesis via residual depth-aided adversarial learning

    Changgyoon Oh, Wonjune Cho, Yujeong Chae, Daehee Park, Lin Wang, and Kuk-Jin Yoon. Bips: Bi-modal in- door panorama synthesis via residual depth-aided adversarial learning. In ECCV, pages 352–371. Springer, 2022. 3

  44. [52]

    A system for acquiring, pro- cessing, and rendering panoramic light field stills for virtual reality

    Ryan S Overbeck, Daniel Erickson, Daniel Evangelakos, Matt Pharr, and Paul Debevec. A system for acquiring, pro- cessing, and rendering panoramic light field stills for virtual reality. ACM TOG, 37(6):1–15, 2018. 1

  45. [53]

    How to evenly distribute points on a sphere more effectively than the canonical Fi- bonacci Lattice

    Martin Roberts. How to evenly distribute points on a sphere more effectively than the canonical Fi- bonacci Lattice. https : / / extremelearning . com.au/how-to-evenly-distribute-points- on- a- sphere- more- effectively- than- the- canonical-fibonacci-lattice/, 2020. 3

  46. [54]

    U- net: Convolutional networks for biomedical image segmen- tation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Pa...

  47. [55]

    Splatt3r: Zero-shot gaussian splatting from uncalibrated image pairs

    Brandon Smart, Chuanxia Zheng, Iro Laina, and Vic- tor Adrian Prisacariu. Splatt3r: Zero-shot gaussian splatting from uncalibrated image pairs. arXiv preprint arXiv:2408.13912, 2024. 2

  48. [56]

    The replica dataset: A digital 10 replica of indoor spaces

    Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al. The replica dataset: A digital 10 replica of indoor spaces. arXiv preprint arXiv:1906.05797,

  49. [57]

    Openvslam: A versatile visual slam framework

    Shinya Sumikura, Mikiya Shibuya, and Ken Sakurada. Openvslam: A versatile visual slam framework. In ACM MM, pages 2292–2295, 2019. 6, 2

  50. [58]

    Weighted-to-spherically- uniform quality evaluation for omnidirectional video

    Yule Sun, Ang Lu, and Lu Yu. Weighted-to-spherically- uniform quality evaluation for omnidirectional video. IEEE signal processing letters, 24(9):1408–1412, 2017. 6

  51. [59]

    Flash3d: Feed-forward gener- alisable 3d scene reconstruction from a single image

    Stanislaw Szymanowicz, Eldar Insafutdinov, Chuanxia Zheng, Dylan Campbell, Joao F Henriques, Christian Rup- precht, and Andrea Vedaldi. Flash3d: Feed-forward gener- alisable 3d scene reconstruction from a single image. arXiv preprint arXiv:2406.04343, 2024. 2

  52. [60]

    Splatter image: Ultra-fast single-view 3d recon- struction

    Stanislaw Szymanowicz, Chrisitian Rupprecht, and Andrea Vedaldi. Splatter image: Ultra-fast single-view 3d recon- struction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 10208– 10217, 2024. 2

  53. [61]

    Mvdiffusion: Enabling holistic multi- view image generation with correspondence-aware diffusion

    Shitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang, and Yasutaka Furukawa. Mvdiffusion: Enabling holistic multi- view image generation with correspondence-aware diffusion. In NeurIPS, 2023. 3

  54. [62]

    Hisplat: Hierarchical 3d gaus- sian splatting for generalizable sparse-view reconstruction

    Shengji Tang, Weicai Ye, Peng Ye, Weihao Lin, Yang Zhou, Tao Chen, and Wanli Ouyang. Hisplat: Hierarchical 3d gaus- sian splatting for generalizable sparse-view reconstruction. arXiv preprint arXiv:2410.06245, 2024. 2, 3

  55. [63]

    Sparf: Neural radiance fields from sparse and noisy poses

    Prune Truong, Marie-Julie Rakotosaona, Fabian Manhardt, and Federico Tombari. Sparf: Neural radiance fields from sparse and noisy poses. In CVPR, pages 4190–4200, 2023. 2

  56. [64]

    Stylelight: Hdr panorama generation for lighting estimation and editing

    Guangcong Wang, Yinuo Yang, Chen Change Loy, and Zi- wei Liu. Stylelight: Hdr panorama generation for lighting estimation and editing. In ECCV, pages 477–492. Springer,

  57. [65]

    360-degree panorama generation from few unregis- tered nfov images

    Jionghao Wang, Ziyu Chen, Jun Ling, Rong Xie, and Li Song. 360-degree panorama generation from few unregis- tered nfov images. In ACM MM, pages 6811–6821. ACM,

  58. [66]

    Ibr- net: Learning multi-view image-based rendering

    Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P Srinivasan, Howard Zhou, Jonathan T Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser. Ibr- net: Learning multi-view image-based rendering. In CVPR, pages 4690–4699, 2021. 2, 3, 5, 6

  59. [67]

    latentsplat: Autoencoding variational gaussians for fast generalizable 3d reconstruction

    Christopher Wewer, Kevin Raj, Eddy Ilg, Bernt Schiele, and Jan Eric Lenssen. latentsplat: Autoencoding variational gaussians for fast generalizable 3d reconstruction. arXiv preprint arXiv:2403.16292, 2024. 2, 5

  60. [68]

    Reconfusion: 3d reconstruction with diffusion priors

    Rundi Wu, Ben Mildenhall, Philipp Henzler, Keunhong Park, Ruiqi Gao, Daniel Watson, Pratul P Srinivasan, Dor Verbin, Jonathan T Barron, Ben Poole, et al. Reconfusion: 3d reconstruction with diffusion priors. In CVPR, pages 21551–21561, 2024. 2

  61. [69]

    Ipo-ldm: Depth-aided 360-degree indoor rgb panorama outpainting via latent diffusion model

    Tianhao Wu, Chuanxia Zheng, and Tat-Jen Cham. Ipo-ldm: Depth-aided 360-degree indoor rgb panorama outpainting via latent diffusion model. arXiv preprint arXiv:2307.03177,

  62. [70]

    Diversified and personalized multi-rater medical image segmentation

    Yicheng Wu, Xiangde Luo, Zhe Xu, Xiaoqing Guo, Lie Ju, Zongyuan Ge, Wenjun Liao, and Jianfei Cai. Diversified and personalized multi-rater medical image segmentation. In CVPR, pages 11470–11479, 2024. 1

  63. [71]

    Unifying flow, stereo and depth estimation

    Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger. Unifying flow, stereo and depth estimation. IEEE TPAMI, 2023. 4, 2

  64. [72]

    Murf: Multi-baseline radiance fields

    Haofei Xu, Anpei Chen, Yuedong Chen, Christos Sakaridis, Yulun Zhang, Marc Pollefeys, Andreas Geiger, and Fisher Yu. Murf: Multi-baseline radiance fields. In CVPR, pages 20041–20050, 2024. 2

  65. [73]

    Depthsplat: Connecting gaussian splatting and depth

    Haofei Xu, Songyou Peng, Fangjinhua Wang, Hermann Blum, Daniel Barath, Andreas Geiger, and Marc Pollefeys. Depthsplat: Connecting gaussian splatting and depth. arXiv preprint arXiv:2410.13862, 2024. 3

  66. [74]

    Neural render- ing in a room: amodal 3d understanding and free-viewpoint rendering for the closed scene composed of pre-captured ob- jects

    Bangbang Yang, Yinda Zhang, Yijin Li, Zhaopeng Cui, Sean Fanello, Hujun Bao, and Guofeng Zhang. Neural render- ing in a room: amodal 3d understanding and free-viewpoint rendering for the closed scene composed of pre-captured ob- jects. ACM TOG, 41(4):1–10, 2022. 3

  67. [75]

    Cost volume pyramid based depth inference for multi-view stereo

    Jiayu Yang, Wei Mao, Jose M Alvarez, and Miaomiao Liu. Cost volume pyramid based depth inference for multi-view stereo. In CVPR, pages 4877–4886, 2020. 4

  68. [76]

    No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images

    Botao Ye, Sifei Liu, Haofei Xu, Xueting Li, Marc Pollefeys, Ming-Hsuan Yang, and Songyou Peng. No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images. arXiv preprint arXiv:2410.24207, 2024. 2

  69. [77]

    Diffpano: Scalable and con- sistent text to panorama generation with spherical epipolar- aware diffusion

    Weicai Ye, Chenhao Ji, Zheng Chen, Junyao Gao, Xiaoshui Huang, Song-Hai Zhang, Wanli Ouyang, Tong He, Cairong Zhao, and Guofeng Zhang. Diffpano: Scalable and con- sistent text to panorama generation with spherical epipolar- aware diffusion. arXiv preprint arXiv:2410.24203, 2024. 3

  70. [78]

    Monosdf: Exploring monocu- lar geometric cues for neural implicit surface reconstruction

    Zehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sat- tler, and Andreas Geiger. Monosdf: Exploring monocu- lar geometric cues for neural implicit surface reconstruction. NeurIPS, 35:25018–25032, 2022. 2

  71. [79]

    Deeppanocontext: Panoramic 3d scene understanding with holistic scene con- text graph and relation-based optimization

    Cheng Zhang, Zhaopeng Cui, Cai Chen, Shuaicheng Liu, Bing Zeng, Hujun Bao, and Yinda Zhang. Deeppanocontext: Panoramic 3d scene understanding with holistic scene con- text graph and relation-based optimization. In ICCV, pages 12632–12641, 2021. 3

  72. [80]

    Taming stable diffusion for text to 360 panorama image gen- eration

    Cheng Zhang, Qianyi Wu, Camilo Cruz Gambardella, Xi- aoshui Huang, Dinh Phung, Wanli Ouyang, and Jianfei Cai. Taming stable diffusion for text to 360 panorama image gen- eration. In CVPR, pages 6347–6357, 2024. 3

  73. [81]

    Arf: Artistic radiance fields

    Kai Zhang, Nick Kolkin, Sai Bi, Fujun Luan, Zexiang Xu, Eli Shechtman, and Noah Snavely. Arf: Artistic radiance fields. In ECCV, pages 717–733. Springer, 2022. 5

  74. [82]

    Gs-lrm: Large recon- struction model for 3d gaussian splatting

    Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, and Zexiang Xu. Gs-lrm: Large recon- struction model for 3d gaussian splatting. In ECCV, pages 1–19. Springer, 2025. 2, 3

  75. [83]

    #,%,&) GTNovelViewImageLoss ImageGradientsCache CubemapRenderer Gaussian Heads Gaussian Heads DepthLoss CubemapRenderer (

    Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, pages 586– 595, 2018. 5, 6 11 PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting Supplementary Ma...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.