Pith. sign in

REVIEW 3 major objections 4 minor 61 references

A Multi-Layer System for Ultra-High-Resolution Static 360-Degree Telepresence

T0 review · 3 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Offline high-detail scans of a static scene can bring 360-degree VR telepresence close to the human visual limit, something live 8K 360 video cannot do.

desk verdict A genuine systems integration with an honest evaluation; the 24K/66.7 PPD headline is an unverified pixel-counting estimate, but the user study and architecture carry the paper. read the letter →

arxiv 2608.05570 v1 pith:XTCCDXJK submitted 2026-08-06 cs.HC cs.CV

classification cs.HCcs.CV
keywords telepresence360-degreevideovirtualrealitypan-tilt-zoomcameraultra-highresolutionimagestitchingbackgroundmattinguserstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that a static 360-degree telepresence scene can be shown with near eye-limited detail by building an offline, tile-based ultra-high-resolution background from repeated 4K pan-tilt-zoom scans and overlaying live motion and a 4K region-of-interest stream. The core claim is that this layered representation reaches an effective horizontal resolution of about 24K over the full panorama, roughly 66.7 pixels per degree, which sits just above the 64.5 pixels-per-degree threshold at which observers reach 90 percent of peak visual performance. In a 21-person VR user study, the 24K background conditions were rated significantly better than 8K conditions for seeing fine details, overall sharpness, and reading text. The paper concludes that for fixed-viewpoint, largely static setups, capture-based multi-layer refinement is a practical route to high-fidelity telepresence that learned super-resolution upscaling cannot match.

What carries the argument

The load-bearing mechanism is the virtual-PTZ alignment and Poisson fusion pipeline: for each PTZ capture, a perspective view is rendered from the equirectangular canvas at the same yaw and pitch as the physical PTZ camera, SIFT features align the 4K photo to that reference, and Poisson seamless cloning fuses it back onto the canvas. This preserves global consistency across the full panorama, unlike matching each cubemap face independently, which accumulates cross-face errors. After roughly 50 scans, the canvas is resized to the target 24K resolution and partitioned into 24 × 12 tiles of 1024 × 1024 pixels, so that region-of-interest-driven updates only need to modify nearby tiles rather than the full panorama.

What would settle it

Take the described hardware, scan a scene with dense high-frequency texture, and compute the per-tile registration error between the fused 24K canvas and independent PTZ photos not used in the fusion. If the maximum residual exceeds the 24K-equivalent pixel pitch, about 0.94 arcmin per pixel, in any scanned region, then the claimed eye-limited consistency fails; the paper's own report of visible misalignment near the ROI boundary indicates this failure mode is real.

Watch

Extended reading notes

Core claim

The central claim is that combining a consumer 8K 360-degree camera with a rotatable 4K pan-tilt-zoom camera yields a three-layer representation in which the static background reaches an effective 24K horizontal resolution in scanned regions, while live dynamics are preserved by matting-based foreground compositing and user-selected regions receive real 4K video. The key comparison is against native 8K 360 video, which provides only about 21.3 pixels per degree, well below the 64.5 pixels-per-degree threshold for 90 percent of peak visual performance. The user study found significant advantages for the 24K conditions on fine-detail, sharpness, and text-readability items, supporting the claim that the added background resolution is perceptually meaningful rather than merely numerical. The paper also acknowledges that the dynamic update layer trails the background in visual quality and that higher fidelity makes layer-boundary misalignment and artifacts more visible.

Load-bearing premise

The system assumes that a virtual perspective view rendered from the 8K panorama, when aligned by SIFT and fused by Poisson cloning, can register each 4K PTZ photo well enough that dozens of scans form a globally consistent 24K background despite different viewpoints, optics, and autofocus-induced field-of-view changes.

Editorial extensions

If this is right

  • Static telepresence scenes can offer near eye-limited detail using consumer-grade cameras, making fine text readable in VR without professional 8K-plus live capture hardware.
  • Learned super-resolution of 8K 360 video cannot reliably recover details that were never captured, so capture-based stitching is presented as the more dependable route to fine-detail fidelity.
  • The region-of-interest layer provides live 4K inspection and can locally refresh background tiles, so user-guided updates keep the ultra-high-resolution background current over time.
  • The dynamic update layer is the system's weak point: at roughly 12.76 frames per second with about 1.41 seconds of latency and coarse matting, moving subjects blend less naturally than in native video.
  • Improving fidelity also exposes artifacts: object misalignment is most noticeable near the ROI-background boundary, and participants noticed more artifacts as layers were added.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if the 64.5 pixels-per-degree anchor is correct, then even higher effective resolutions from more PTZ scans or a higher-resolution PTZ camera would yield diminishing perceptual returns, since the foveal limit is about 94 pixels per degree.
  • Beyond the paper: the virtual-PTZ alignment trick could generalize to semi-dynamic scenes by periodically capturing foreground-free reference frames during quiet moments and refreshing the UHR background automatically, reducing stale-background artifacts without user intervention.
  • Beyond the paper: a testable extension would be to measure per-tile registration error between the fused 24K canvas and independent PTZ photos; if residual errors stay below the 24K-equivalent pixel size, the eye-limited claim could be validated by a purely geometric metric rather than by subjective ratings.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a multi-layer telepresence system for largely static 360-degree scenes: an offline tile-based ultra-high-resolution (UHR) background built by stitching 4K PTZ photographs onto an 8K panorama, a real-time dynamic update layer based on background matting, and a live 4K region-of-interest (ROI) layer. The claimed outcome is an effective horizontal resolution of about 24K (66.7 PPD) in the UHR layer, just above a 64.5 PPD perceptual threshold. The evaluation consists of a system timing/latency analysis, a qualitative comparison with four super-resolution baselines, and a within-subjects user study with 21 participants comparing five conditions (8K/24K backgrounds, static/dynamic, with/without moving subject and ROI). The user study reports significant benefits for perceived detail and readability of the 24K conditions over 8K conditions, with honest discussion of confounds and artifacts.

Significance. If the claimed 24K effective resolution is supported by actual stitching accuracy, the paper would make a useful practical contribution: a consumer-camera-based route to near-eye-limited detail in static-scene VR telepresence, with a clean separation of offline background refinement from live dynamics. The user study is a genuine strength: it is within-subjects, uses 21 participants, checks ANOVA assumptions, applies Bonferroni-corrected post-hoc tests, and reports effect sizes. The manuscript is also commendably transparent about the 8DMR versus 24SMR confound, the limitations of the dynamic layer (12.76 FPS, 1.41 s latency), the absence of a 24K dynamic baseline, and the visible boundary misalignment. The main weakness is that the central '24K / 66.7 PPD' claim is currently a pixel-counting argument rather than a measured property of the final stitched panorama.

major comments (3)
  1. [Section 2.2 and Section 3.2.1] The headline claim that the UHR layer reaches approximately 24K effective horizontal resolution (about 66.7 PPD) is presented as a pixel-counting consequence of using 4K PTZ frames over a 60-degree field of view, but no measurement of the final stitched UHR panorama supports it. Figure 3(D,E) validates only the PTZ lens distortion; there is no reprojection-error or seam-consistency metric for the fused background, and Section 6.4 explicitly reports visible object misalignment at the ROI/UHR boundary. Because residual misregistration blurs high-frequency content and reduces effective resolution, the claimed position just above the 64.5 PPD threshold is unverified. Please add quantitative geometric validation (e.g., reprojection error on held-out PTZ captures, seam-consistency statistics, or a resolution-chart/contrast-transfer measurement in the final UHR layer) and report the resulting effective resolution rather than only the nominal pixel count.
  2. [Abstract/Contributions vs. Section 2.2 and Section 3.1] The paper uses conflicting statements about the coverage of the UHR layer: Section 2.2 says 'over the full panorama,' while the contributions and Section 3.1 refer to 'scanned regions' and the 'usable viewing range.' Since the UHR background is generated from a finite PTZ scan, the 24K effective-resolution claim should state explicitly whether the entire 360-degree sphere was covered by PTZ captures and, if not, what fraction of the panorama actually received UHR refinement. This is not merely a wording issue: the eye-limited-resolution argument in Section 2.2 applies only to regions actually scanned at 4K density, and unscanned regions would remain at the native 8K density of 21.3 PPD.
  3. [Section 4.2.2] The comparison with DRCT, HAT, RealBasicVSR, and Real-ESRGAN is based solely on visual inspection of Figure 6; no quantitative metric or statistical test is reported. The conclusion that the capture-based method 'reconstructs fine details more clearly' than super-resolution is therefore not supported by the presented evidence. Please add quantitative measurements (e.g., reference-based metrics on the 4K PTZ content, or a perceptual rating study with multiple raters) or explicitly reframe this section as an illustrative qualitative comparison rather than a comparative evaluation.
minor comments (4)
  1. [Section 5.4] For all post-hoc pairwise comparisons, please report test statistics, degrees of freedom, and effect sizes rather than only 'p < 0.05'; currently the reader cannot assess the magnitude or direction of the reported differences beyond the figures.
  2. [Figures 6 and 8] The reproduced images are small and may not convey the claimed text-detail differences in print or PDF; please include zoomed crops of representative text regions for the UHR and ROI comparisons.
  3. [Section 3.2.1] The statement that autofocus can change the effective FOV and that 'multiple trials within a reasonable range' may be needed should be accompanied by a concrete description of how the virtual FOV was selected in the experiments and whether the same value worked for all PTZ captures.
  4. [Section 4.2.2] The super-resolution baselines are official pre-trained models designed for perspective images; the cubemap conversion is a reasonable practical choice, but it may disadvantage those models compared with methods trained on equirectangular content, and this should be acknowledged as a limitation of the comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the 24K/66.7 PPD figure is a pixel-counting design specification, and the perceptual claims are tested empirically against external baselines and a user study.

full rationale

The paper's central quantitative claim—effective horizontal resolution of roughly 24K, or about 66.7 PPD—is computed directly from the PTZ camera's 4K resolution and its reduced 60-degree effective field of view (Section 3.1: 'a 4K PTZ image then covers about one-sixth of the full 360-degree panorama, the resulting stitched background corresponds to an effective horizontal resolution of roughly 24K'). This is an arithmetic design specification, not a parameter fitted to the outcome data, and it is not derived from the user-study results. The 64.5 PPD threshold is imported from an external source (Lloyd et al. [25]) rather than from the authors' prior work. The only self-citation, reference [6], appears in a related-work sentence and is explicitly superseded by 'a more direct criterion derived from published visual performance data'; it is not load-bearing. The evaluation is empirical and external: perceived detail is measured through user ratings (VD1–VD3, RE, ROI) and through visual comparisons against four published super-resolution methods, so the improvement claims are falsifiable independent of the resolution arithmetic. The paper also openly acknowledges validity limitations in Sections 6.4 and 6.5 (object misalignment near the ROI/UHR boundary, dynamic-layer latency, and the lack of a 24K dynamic baseline); these are correctness and generalization risks, not circular derivations. No equation or argument in the paper defines its conclusion in terms of the claim itself, and no fitted parameter is relabeled as a prediction. Therefore the derivation chain is self-contained with respect to circularity.

Assumptions & free parameters 2 free parameters · 5 assumptions · 0 invented entities

The central pipeline rests on engineering assumptions about scene stability, alignment consistency, and the adequacy of cube-face matting, rather than on fitted theoretical parameters. The only hand-tuned values are the PTZ field of view and photo count; the effective 24K figure follows from the field-of-view math, not from a fit. No new physical entities are introduced.

free parameters (2)
  • PTZ effective field of view = approximately 60 degrees
    Chosen after ChArUco distortion measurement because the native 85.5 degree FOV has corner deviations up to 3.3 pixels; this value determines the 24K effective-resolution estimate and stitching stability. The paper also notes multiple trials are needed to stabilize the virtual FOV due to autofocus.
  • Number of PTZ photos per panorama = approximately 50
    Empirical choice balancing scan time, about 25 minutes, and coverage of the usable viewing range; affects fusion completeness and the uniformity of the UHR background.
assumptions (5)
  • domain assumption The telepresence scene is largely static with a clean background reference obtainable without user effort.
    Justifies offline UHR construction and BMV2 matting; introduced in Sections 1 and 2.3.
  • domain assumption SIFT-based homography alignment between PTZ captures and virtual-PTZ renders is globally consistent despite different optics and a slight viewpoint offset.
    Core to the stitching pipeline; Section 3.2.1 acknowledges parallax and autofocus risks, and Section 6.4 reports visible cross-layer misalignment.
  • domain assumption Cube-face processing with BackgroundMattingV2 preserves sufficient geometry and fidelity for foreground compositing from equirectangular 8K video.
    Section 2.3 and Section 3.2.2 rely on this without quantitative validation of matte quality.
  • domain assumption Poisson seamless cloning and Reinhard color transfer adequately hide seams between fused PTZ views and between layers.
    Used in Sections 3.2.1 and 3.2.3; the user study still detected artifacts, smearing, and misalignment.
  • domain assumption The Lloyd et al. eye-limited resolution criterion, 64.5 pixels per degree, is the appropriate benchmark for the 24K claim.
    Section 2.2 uses this external criterion to argue that 24K reaches near eye-limited resolution, but the user study does not directly validate PPD thresholds.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Multi-Layer System for Ultra-High-Resolution Static 360-Degree Telepresence." pith.science (2026). https://pith.science/paper/XTCCDXJK

@misc{pith2026260805570,
  author       = {Pith},
  title        = {Pith review of: A Multi-Layer System for Ultra-High-Resolution Static 360-Degree Telepresence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XTCCDXJK}},
  note         = {Machine review of arXiv:2608.05570}
}
read the original abstract

360-degree video telepresence offers strong immersive potential but remains constrained by the limited resolution of current capture and display hardware. Many telepresence installations feature fixed viewpoints and largely static scenes, yet optimization strategies tailored to such setups have received limited attention. We present a multi-layer, ultra-high-resolution system for static 360-degree telepresence that combines an 8K panoramic camera with a rotatable 4K pan-tilt-zoom (PTZ) camera. Our approach builds a three-layer representation: (1) a tile-based ultra-high-resolution panoramic background, generated by offline stitching high-detail 4K PTZ scans onto the base 8K panorama to achieve effective resolution beyond native capture, and represented as a set of spatial tiles; (2) a dynamic update layer that composites foreground motions from the 8K stream via real-time high-resolution background matting; and (3) a region-of-interest 4K layer that streams a real-time PTZ view of the selected region and additionally updates the corresponding background tiles over time. We evaluate the proposed system through comparisons with representative video super-resolution approaches and a user study assessing perceived detail and immersive experience. Our results indicate that tile-based background refinement, together with user-guided updates, provides a practical way to balance panoramic fidelity and interactivity in static 360-degree telepresence.

Figures

Figures reproduced from arXiv: 2608.05570 by the authors.

Figure 1
Figure 1. System overview. A 360-degree camera and a PTZ camera jointly produce three layers that are rendered on various display [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. System pipeline diagram. pared with earlier background matting approaches such as BGM [39], it achieves substantially better detail preservation while remaining fully automatic and real time. Although background-based matting is sometimes considered lim￾ited because it requires a reference background image [21], this assump￾tion matches our target scenario well: the telepresence scene is largely static, so a clean b… view at source ↗
Figure 3
Figure 3. Hardware design of the proposed system. (A) Kandao QooCam [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (7 more)
Figure 5
Figure 5. Figure 5: Color correction comparison: (A) original overlay with visible [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 4
Figure 4. Figure 4: (A) Cubemap matching accumulates cross-face errors, causing [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 6
Figure 6. Figure 6: Visual comparison arranged in two groups: the upper group corresponds to the UHR layer, and the lower group corresponds to the ROI [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 7
Figure 7. Figure 7: Experimental setup. (A, B, C) Views of the remote room used [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 8
Figure 8. Figure 8: Visual stimuli used in the five conditions in the experiment: (a) 8DMR, (b) 8SNN, (c) 24SNN, (d) 24SNR, and (e) 24SMR. [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 9
Figure 9. Figure 9: Results for the standard questionnaires: (a) User Experience Questionnaire – Short, and (b) Igroup Presence Questionnaire. Higher is [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: Results for the single-item questionnaires. Higher is better. The vertical error bars show the standard error. The horizontal bars indicate [PITH_FULL_IMAGE:figures/full_fig_p008_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

61 extracted references · 50 canonical work pages

  1. [1]

    Agustsson and R

    E. Agustsson and R. Timofte. Ntire 2017 challenge on single image super- resolution: Dataset and study. In2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1122–1131,

  2. [2]

    Ashraf, A

    M. Ashraf, A. Chapiro, and R. K. Mantiuk. Resolution limit of the eye—how many pixels can we see?Nature Communications, 16(1):9086,

  3. [3]

    S. R. Bharadwaj and T. R. Candy. Accommodative and vergence responses to conflicting blur and disparity stimuli during development.Journal of vision, 9(11):4–4, 2009. doi: 10.1167/9.11.4 2

  4. [4]

    K. C. Chan, S. Zhou, X. Xu, and C. C. Loy. Investigating tradeoffs in real-world video super-resolution. In2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5952–5961, 2022. doi: 10.1109/CVPR52688.2022.00587 5

  5. [5]

    X. Chen, X. Wang, W. Zhang, X. Kong, Y . Qiao, J. Zhou et al. Hat: Hybrid attention transformer for image restoration.IEEE Transactions on Pattern Analysis and Machine Intelligence, 48(3):2676–2694, 2026. doi: 10.1109/TPAMI.2025.3628275 5

  6. [6]

    J. Chi, C. Neumann, C. Cruz-Neira, and D. Reiners. A modular hybrid telepresence system integrating immersive 360-degree video with on- demand high-resolution imaging. InInternational Conference on Virtual Reality and Mixed Reality, pp. 81–103. Springer, 2025. doi: 10.1007/978 -3-032-03805-0_5 2

  7. [7]

    D. B. Elliott.Clinical procedures in primary eye care E-Book. Elsevier Health Sciences, 2020. 5

  8. [8]

    Free ebooks | project gutenberg

    Gutenberg. Free ebooks | project gutenberg. https://www.gutenberg. org/, 2026. Accessed: March 2026. 6

Show all 61 references
  1. [9]

    V . K. L. Ha, R. Chai, and H. T. Nguyen. A telepresence wheelchair with 360-degree vision using webrtc.Applied Sciences, 10(1), 2020. doi: 10. 3390/app10010369 1

  2. [10]

    Hilliard, A

    J. Hilliard, A. Hilton, and J.-Y . Guillemaut. 360u-former: Hdr illumination estimation with panoramic adapted vision transformers. InEuropean Conference on Computer Vision, pp. 431–448. Springer, 2024. doi: 10. 1007/978-3-031-91838-4_26 2

  3. [11]

    Hinderks

    A. Hinderks. Design and evaluation of a short version of the user expe- rience questionnaire (ueq-s).International Journal of Interactive Multi- media and Artificial Intelligence, 2017. doi: 10.9781/ijimai.2017.09.001 6

  4. [12]

    Hsu, C.-M

    C.-C. Hsu, C.-M. Lee, and Y .-S. Chou. Drct: Saving image super- resolution away from information bottleneck. In2024 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 6133–6142, 2024. doi: 10.1109/CVPRW63382.2024.00618 5

  5. [13]

    Insta360 pro 2

    Insta360. Insta360 pro 2. https://www.insta360.com/product/ insta360-pro2, 2025. Accessed: September 2025. 1

  6. [14]

    Insta360 x4

    Insta360. Insta360 x4. https://www.insta360.com/product/ insta360-x4, 2025. Accessed: September 2025. 1

  7. [15]

    Insta360 x5

    Insta360. Insta360 x5. https://www.insta360.com/product/ insta360-x5, 2025. Accessed: September 2025. 1

  8. [16]

    Kasahara and J

    S. Kasahara and J. Rekimoto. Jackin head: immersive visual telepresence system with omnidirectional wearable camera for remote collaboration. InProceedings of the 21st ACM Symposium on Virtual Reality Software and Technology, VRST ’15, pp. 217–225. Association for Computing Ma...

  9. [17]

    Z. Ke, J. Sun, K. Li, Q. Yan, and R. W. Lau. Modnet: Real-time trimap- free portrait matting via objective decomposition. InProceedings of the AAAI conference on artificial intelligence, vol. 36, pp. 1140–1147, 2022. doi: 10.1609/aaai.v36i1.19999 3

  10. [18]

    R. S. Kennedy, N. E. Lane, K. S. Berbaum, and M. G. Lilienthal. Simulator sickness questionnaire: An enhanced method for quantifying simulator sickness.The international journal of aviation psychology, 3(3):203–220,

  11. [19]

    G. Kramida. Resolving the vergence-accommodation conflict in head- mounted displays.IEEE Transactions on Visualization and Computer Graphics, 22(7):1912–1931, 2016. doi: 10.1109/TVCG.2015.2473855 2

  12. [20]

    Lawrence, R

    J. Lawrence, R. Overbeck, T. Prives, T. Fortes, N. Roth, and B. Newman. Project starline: A high-fidelity telepresence system. InACM SIGGRAPH 2024 Emerging Technologies, SIGGRAPH ’24, art. no. 16, 2 pp. Asso- ciation for Computing Machinery, New York, NY , USA, 2024. doi: 10. ...

  13. [21]

    J. Li, J. Zhang, and D. Tao. Deep image matting: A comprehensive survey. arXiv preprint arXiv:2304.04672, 2023. doi: 10.48550/arXiv.2304.04672 3

  14. [22]

    W. Li, Q. Li, W. Tian, J. Gao, F. Wu, J. Liu et al. Mucvr: Edge computing- enabled high-quality multi-user collaboration for interactive mvr.IEEE Transactions on Parallel and Distributed Systems, 36(10):2058–2072,

  15. [24]

    Lindenberger, P.-E

    P. Lindenberger, P.-E. Sarlin, and M. Pollefeys. Lightglue: Local fea- ture matching at light speed. In2023 IEEE/CVF International Con- ference on Computer Vision (ICCV), pp. 17581–17592, 2023. doi: 10. 1109/ICCV51070.2023.01616 4

  16. [25]

    C. J. Lloyd, M. Winterbottom, J. Gaska, and L. Williams. A practical defi- nition of eye-limited display system resolution. InDisplay Technologies and Applications for Defense, Security, and Avionics IX; and Head-and Helmet-Mounted Displays XX, vol. 9470, pp. 107–115. SPIE, 20...

  17. [26]

    doi: 10.1109/TPDS.2025.3595801 2

  18. [27]

    Q. Luo, J. Zhang, Y . Xie, X. Huang, and T. Han. Comparative analysis of advanced feature matching algorithms in challenging high spatial res- olution optical satellite stereo scenarios. InIGARSS 2024 - 2024 IEEE International Geoscience and Remote Sensing Symposium, pp. 2645–2649,

  19. [28]

    Meta quest 3

    Meta. Meta quest 3. https://www.meta.com/quest/quest-3/, 2025. Accessed: September 2025. 1

  20. [29]

    M. Minsky. Telepresence. 1980. 1

  21. [30]

    D. G. Lowe. Distinctive image features from scale-invariant keypoints. International journal of computer vision, 60(2):91–110, 2004. doi: 10. 1023/B:VISI.0000029664.99615.94 4

  22. [31]

    Nassani, L

    A. Nassani, L. Zhang, H. Bai, and M. Billinghurst. Showmearound: Giving virtual tours using live 360 video. InExtended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, CHI EA ’21, art. no. 168, 4 pp. Association for Computing Machinery, New York, NY ...

  23. [32]

    Nguyen, C

    J. Nguyen, C. Smith, Z. Magoz, and J. Sears. Screen door effect reduction using mechanical shifting for virtual reality displays. In B. C. Kress and C. Peroz, eds.,Optical Architectures for Displays and Sensing in Augmented, Virtual, and Mixed Reality (AR, VR, MR), vol. 11310,...

  24. [33]

    Okuta, Y

    R. Okuta, Y . Unno, D. Nishino, S. Hido, and C. Loomis. Cupy: A numpy- compatible library for nvidia gpu calculations. InProceedings of Workshop on Machine Learning Systems (LearningSys) in The Thirty-first Annual Conference on Neural Information Processing Systems (NIPS), 2017. 4

  25. [34]

    Pejsa, J

    T. Pejsa, J. Kantor, H. Benko, E. Ofek, and A. Wilson. Room2room: En- abling life-size telepresence in a projected augmented reality environment. InProceedings of the 19th ACM Conference on Computer-Supported Cooperative Work & Social Computing, CSCW ’16, pp. 1716–1725. As- so...

  26. [35]

    Narciso, M

    D. Narciso, M. Bessa, M. Melo, A. Coelho, and J. Vasconcelos-Raposo. Immersive 360◦ video user experience: impact of different variables in the sense of presence and cybersickness.Universal Access in the Information Society, 18(1):77–87, 2019. doi: 10.1007/s10209-017-0581-5 2

  27. [36]

    Pimax crystal super

    Pimax. Pimax crystal super. https://pimax.com/pages/ pimax-crystal-super, 2025. Accessed: September 2025. 1

  28. [37]

    Reinhard, M

    E. Reinhard, M. Adhikhmin, B. Gooch, and P. Shirley. Color transfer between images.IEEE Computer Graphics and Applications, 21(5):34–41,

  29. [38]

    Schubert, F

    T. Schubert, F. Friedmann, and H. Regenbrecht. The experience of pres- ence: Factor analytic insights.Presence: Teleoper. Virtual Environ., 10(3):266–281, June 2001. doi: 10.1162/105474601300343603 6

  30. [39]

    Sengupta, V

    S. Sengupta, V . Jayaram, B. Curless, S. M. Seitz, and I. Kemelmacher- Shlizerman. Background matting: The world is your green screen. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2288–2297, 2020. doi: 10.1109/CVPR42600.2020.00236 3

  31. [40]

    Pérez, M

    P. Pérez, M. Gangnet, and A. Blake.Poisson Image Editing. Association for Computing Machinery, New York, NY , USA, 1 ed., 2023. doi: 10. 1145/3596711.3596772 4

  32. [41]

    Z. Shen, C. Lin, K. Liao, L. Nie, Z. Zheng, and Y . Zhao. Panoformer: panorama transformer for indoor 360 ◦ depth estimation. InEuropean Conference on Computer Vision, pp. 195–211. Springer, 2022. doi: 10. 1007/978-3-031-19769-7_12 2

  33. [42]

    Ultra-fast, realtime video routing for windows

    Spout. Ultra-fast, realtime video routing for windows. https://spout. zeal.co/, 2026. Accessed: July 2026. 4

  34. [43]

    J. Sun, Z. Shen, Y . Wang, H. Bao, and X. Zhou. Loftr: Detector-free local feature matching with transformers. In2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8918–8927, 2021. doi: 10.1109/CVPR46437.2021.00881 4

  35. [44]

    M. J. Tyszkiewicz, P. Fua, and E. Trulls. Disk: learning local features with policy gradient. InProceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20, art. no. 1195, 12 pp. Curran Associates Inc., Red Hook, NY , USA, 2020. 4

  36. [45]

    Varjo xr-3 - the first true mixed reality headset | varjo

    Varjo. Varjo xr-3 - the first true mixed reality headset | varjo. https: //varjo.com/products/varjo-xr-3 , 2026. Accessed: March 2026. 5

  37. [46]

    Shen, H.-H

    C.-T. Shen, H.-H. Liu, M.-H. Yang, Y .-P. Hung, and S.-C. Pei. Viewing- distance aware super-resolution for high-definition display.IEEE Transac- tions on Image Processing, 24(1):403–418, 2015. doi: 10.1109/TIP.2014. 2375639 2

  38. [47]

    Varjo xr-3: Full specification - vrcompare

    VRCompare. Varjo xr-3: Full specification - vrcompare. https:// vr-compare.com/headset/varjoxr-3 , 2026. Accessed: March 2026. 5

  39. [48]

    X. Wang, L. Xie, C. Dong, and Y . Shan. Real-esrgan: Training real- world blind super-resolution with pure synthetic data. In2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), pp. 1905–1914, 2021. doi: 10.1109/ICCVW54120.2021.00217 5

  40. [49]

    Y . Xu, Y . Dong, J. Wu, Z. Sun, Z. Shi, J. Yu et al. Gaze prediction in dynamic 360° immersive videos. In2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5333–5342, 2018. doi: 10. 1109/CVPR.2018.00559 2

  41. [50]

    S. Yan, X. Xu, R. Zhang, L. Hong, W. Chen, W. Zhang et al. Panovos: Bridging non-panoramic and panoramic views with transformer for video segmentation. In A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sat- tler, and G. Varol, eds.,Computer Vision – ECCV 2024, pp. 346–365...

  42. [51]

    Zeromq: An open-source universal messaging library

    ZeroMQ Project. Zeromq: An open-source universal messaging library. https://zeromq.org/, 2026. Accessed: July 2026. 4

  43. [52]

    J. P. Verma and A.-S. G. Abdel-Salam.Testing statistical assumptions in research. John Wiley & Sons, 2019. 7

  44. [53]

    Zhang, E

    J. Zhang, E. Langbehn, D. Krupke, N. Katzakis, and F. Steinicke. De- tection thresholds for rotation and translation gains in 360° video-based telepresence systems.IEEE Transactions on Visualization and Computer Graphics, 24(4):1671–1680, 2018. doi: 10.1109/TVCG.2018.2793679 1

  45. [54]

    Zhang, D

    W. Zhang, D. Xiao, A. Dai, Y . Liu, T. Pan, S. Wen et al. Leader360v: The large-scale, real-world 360 video dataset for multi-task learning in diverse environment, 2025. doi: 10.48550/arXiv.2506.14271 2

  46. [55]

    Zhang, F.-L

    Y . Zhang, F.-L. Zhang, Z. Zhu, L. Wang, and Y . Jin. Fast edit propagation for 360 degree panoramas using function interpolation.IEEE Access, 10:43882–43894, 2022. doi: 10.1109/ACCESS.2022.3168665 2

  47. [56]

    Zhong, X

    D. Zhong, X. Zheng, C. Liao, Y . Lyu, J. Chen, S. Wu et al. Omnisam: Omnidirectional segment anything model for uda in panoramic semantic segmentation. In2025 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 23892–23901, 2025. doi: 10.1109/ICCV51701.2025. 02215 2

  48. [58]

    Zhang, Q

    C. Zhang, Q. Cai, P. A. Chou, Z. Zhang, and R. Martin-Brualla. Viewport: A distributed, immersive teleconferencing system with infrared dot pattern. IEEE MultiMedia, 20(1):17–27, 2013. doi: 10.1109/MMUL.2013.12 2

  49. [1993]

    doi: 10.1207/s15327108ijap0303_3 6

  50. [2001]

    doi: 10.1109/38.946629 4

  51. [2017]

    doi: 10.1109/CVPRW.2017.150 1, 6

  52. [2024]

    doi: 10.1109/IGARSS53475.2024.10641727 4

  53. [2025]

    doi: 10.1038/s41467-025-64679-2 2

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.