Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Imaging for All-Day Wearable Smart Glasses

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Distributed tiny cameras approach iPhone 14 Pro image quality on smart glasses.

desk verdict Clean design-space analysis for glasses cameras, but the 'close to iPhone' claim is built on an evaluation that never degrades the detail images to the proposed 2 arcmin modules—fixable, but load-bearing. read the letter →

arxiv 2504.13060 v1 pith:SBBUCP3T submitted 2025-04-17 cs.CV

classification cs.CV
keywords smartglassesdistributedcameraarraycomputationalimagingimagesuper-resolutionopticalflowreference-baseddepthoffieldegocentric
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a camera system small enough for all-day smart glasses can still produce images of mobile-phone quality, and answers yes in a specific sense. It derives fundamental limits that tie angular resolution, lens diameter, depth of field, head motion, and available light together, concluding that a fixed-focus design at about 2 arcminutes per pixel is the sensible operating point for glasses. To overcome the size cost of high resolution, it replaces one large camera with one low-resolution wide-field guide camera plus several high-resolution narrow-field detail cameras distributed around the frame. With a reconstruction pipeline that combines optical-flow warping and reference-based super-resolution, the paper reports output that beats today's Ray-Ban Meta camera and comes close to the iPhone 14 Pro main camera on text and QR-code readability.

What carries the argument

The central object is a distributed camera array: one wide-field, low-resolution guide camera plus nine narrow-field, high-resolution detail cameras, positioned so the detail fields tile the guide field from a minimum distance onward. The load-bearing identity is the fixed-focus trade-off $H = D/(4\,\delta\theta)$, which says hyperfocal distance, lens diameter, and angular resolution cannot be chosen independently; 2 arcmin with a 1 mm entrance pupil is the recommended point because it removes autofocus and shrinks modules. The reconstruction machinery is a two-path fusion: optical-flow warping, built on RAFT with pre-warping and soft epipolar-line constraints, transfers sharp detail where correspondences are correct, and reference-based super-resolution, built on C2-Matching, fills occluded or mismatched regions reliably; a learned fusion stage then combines both outputs, with the guide image serving as the target view and as fallback for areas no detail camera sees.

What would settle it

Take a real module with about a 1 mm entrance pupil and 2-arcmin pixels, capture the same scenes used in the paper, run the pipeline with that module's measured blur and noise, and compare FLIP, PSNR, and QR recognition to the paper's simulated-degradation results; a large drop in reconstruction quality or visible stray-light or color artifacts absent from simulation would settle that the central claim holds only under the simulated model.

Watch

Extended reading notes

Core claim

The paper claims that a distributed imaging system, not a monolithic camera module, is the path to all-day wearable smart glasses with modern image quality. Under the constraints of a roughly 1 mm entrance pupil and fixed focus, the design point of about 2 arcmin angular resolution keeps a comfortable reading distance to infinity in focus while keeping modules small; details lost at that resolution are recovered from multiple 1-arcmin-class detail cameras, each imaging only a narrow field, coordinated by a guide camera that sees the whole scene. In both synthetic scenes and real captures from two prototype rigs degraded to mimic tiny-module blur and noise, the fused output outperforms the current glasses-form-factor camera and approaches the iPhone 14 Pro main camera, despite using several tiny simulated modules instead of one large module with auto-focus and stabilization.

Load-bearing premise

The load-bearing premise is that the blur and noise applied to emulate tiny modules, lens point-spread functions from optical simulation and sensor noise fitted from prototype cameras, accurately predicts what a real thumbnail-size module would produce; if real tiny modules differ in off-axis aberrations, stray light, or thermal noise, the measured phone-like quality will not transfer.

Editorial extensions

If this is right

  • A fixed-focus 2-arcmin camera array can avoid autofocus hardware, cutting size, weight, and power compared with trying to match 1-arcmin phone resolution with one module.
  • Egocentric AI and photography on glasses can use the full reconstruction as a drop-in for phone-style images, while raw detail views remain available as a truthful fallback.
  • QR-code reading and fine text, which fail on current glasses cameras, become reliable at smartphone-like pixel-per-degree levels.
  • When head motion is low or the wearer deliberately holds still, longer exposures become possible, and burst mode recovers clean images from short, noisy exposures.
  • For video, running detail cameras at a reduced frame rate while using VIO trajectories to correct epipolar geometry can reconstruct static scenes, with dynamic regions remaining a limitation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the hardest part of the claim is the simulated-degradation transfer; a natural next experiment is to build an actual roughly 1 mm aperture module and check whether its measured off-axis blur and noise match the models used here before expecting phone-like results.
  • Editorial inference: the same guide-plus-detail architecture could be made foveated, placing detail cameras near the wearer's gaze direction and low-resolution coverage elsewhere, trading reconstruction cost for power much as the human eye does.
  • Editorial inference: because the reference-based super-resolution path is generative and can hallucinate, applications that need exact scene content should trust the guide and detail raw views or a conservative fusion, even when the fused image is visually nicer.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper analyzes fundamental limits of imaging for all-day wearable smart glasses, including diffraction and depth of field (Sec. 4.1), head motion (Sec. 4.2), and photon-noise-limited signal-to-noise ratio (Sec. 4.3), from which it recommends a fixed-focus 2-arcmin design point. It then proposes a distributed camera system consisting of one low-resolution guide camera and several high-resolution narrow-field-of-view detail cameras, together with a reconstruction pipeline that combines optical-flow warping, reference-based super-resolution, and a learned fusion stage. The paper evaluates this pipeline on synthetic data and on two real prototype rigs, qualitatively comparing the output with an iPhone 14 Pro and Ray-Ban Meta Smart Glasses and claiming that the distributed system can approach phone-level image quality while keeping camera modules small enough for glasses.

Significance. The Sec. 4 analysis of the trade-off between angular resolution, depth of field, entrance-pupil diameter, head motion, and photon noise is a clear and useful contribution, and the proposed distributed architecture is a principled answer to the module-size problem. The paper is also strong in the level of implementation detail: the pipeline stages are described concretely, real egocentric motion data are used for the motion analysis, and the prototype hardware is documented in the appendix. If the central quality claim were properly supported, the work would be significant for both computational imaging and wearable-device design.

major comments (3)
  1. [Sec. 7.3.1 and Appendix B.3] The headline comparison 'close to iPhone 14 Pro' does not evaluate the proposed 2-arcmin target modules. The target detail module specified in Sec. 7.1 (f=1.925 mm, f/1.8, D=1.1 mm) has IFOV = 1.12 um / 1.925 mm ≈ 2.0 arcmin and a diffraction floor of 1.22λ/D ≈ 1.9 arcmin at 500 nm by Eq. (2), while the desk-mounted prototype used for Fig. 14 (1.25 um pixels, f=3.8 mm, f/2.8) has IFOV ≈ 1.13 arcmin and a diffraction floor ≈ 1.5 arcmin. Appendix B.3 states that no additional blur or noise is applied to the detail images from this prototype, so the 'simulated tiny cameras' in Fig. 14 are actually fed with detail images of roughly 1-arcmin angular resolution rather than the 2-arcmin target resolution. The near-iPhone result therefore reflects the prototype detail optics, not the proposed tiny modules, and the Sec. 7.3.1 remark that the experiment does not violate the optics limits does not address this mismatch. The evaluation needs either a faithful target-module PSF applied to the detail images, with an independently measured noise model, or an explicit qualification that the demonstrated quality corresponds to a 1-arcmin system rather than the recommended 2-arcmin design.
  2. [Sec. 7.2, Sec. 7.4, Table 1] The quantitative evaluation is partly circular because the fusion network is trained on synthetic data degraded with the same camera, blur, and noise models that are then used to produce the test inputs, and because the 'ground truth' is the guide image before that same degradation model is applied. Under this protocol, the PSNR/SSIM/FLIP numbers in Table 1 largely measure the pipeline's ability to invert a known degradation, not its performance on genuinely unseen tiny-camera imagery. The real-world results inherit this issue: the guide image is degraded with the target model, but the detail images are left at prototype quality (Appendix B.3), so the gap between the reconstruction and a real miniature system is not quantified. A holdout evaluation with an independently measured target degradation model, or with real miniature modules, is needed before the quantitative claims can be considered supported.
  3. [Sec. 4.4 vs. Sec. 7.3.1] The paper recommends a fixed-focus 2-arcmin design as the viable trade-off for all-day wear, but the photography comparison in Sec. 7.3.1 uses a 1-arcmin shallow-depth-of-field desk prototype and explicitly notes that this does not violate the derived limits. That means the paper has not demonstrated that the recommended 2-arcmin configuration produces images close to iPhone quality; the demonstrated quality is an upper bound from a different, higher-resolution optical design. The text should either present results for the actual 2-arcmin target configuration or clearly state that the photography comparison is a best-case demonstration rather than a validation of the recommended design point.
minor comments (5)
  1. [Sec. 4.3.1, Eq. (9)] The notation N_ph(λ) is used for both the spectral photon flux on the scene surface and the total photon count per pixel in Eq. (10); please clarify the units and the integration variables so that the dimensional relationship between Eqs. (9) and (11) is transparent.
  2. [Sec. 4.2.2] The 3 deg/s threshold for head-still behavior is selected empirically from Aria recordings; the paper should state how sensitive the conclusions of Fig. 3 are to this threshold, or at least present the threshold as a free parameter in the analysis.
  3. [Appendix B.3] The phrase 'detail images from our tiny guide camera' should presumably read 'tiny detail camera'; as written it is confusing because the guide camera and the detail cameras are distinct components in the proposed system.
  4. [Sec. 7.3.1] The sentence 'our prototype uses large cameras, and we obtain the guide and detail input images for our pipeline through simulation' is imprecise: per Appendix B.3, only the guide image is synthetically degraded, while the detail images are used as captured from the prototype without additional blur or noise. Please rephrase to state this distinction explicitly.
  5. [Sec. 7.2] Running times are reported for the research implementation, but no power or energy estimates are given; since the paper motivates the system as all-day wearable, an order-of-magnitude power estimate for the full pipeline would be a valuable addition even if hardware power is explicitly excluded from the main analysis.

Circularity Check

1 steps flagged · score 6.0 of 10

Detail images are not actually degraded to the 2-arcmin target module, so the 'tiny camera' near-iPhone result is partly forced by the prototype's higher-resolution input rather than predicted from the target design.

  1. fitted input called prediction [Appendix B.3, applied in Sec. 7.1 and Sec. 7.3.1]
    "Detail camera’s blur: The detail images from our prototypes have more blur than detail images from our tiny guide camera. The lenses has a larger f-number, so the diffraction limited spot size is larger than in our detail camera lens. As a result, we do not add additional blur to the captured raw detail images. Detail camera’s noise: As we earlier concluded the expected tiny camera noise is similar to XIMEA/Aria camera noise level, we do not add any additional noise to the captured raw detail images."

    Sec. 7.1 states that 'we degrade all ten images (nine detail images + one guide image) to simulate the expected quality and resolution of small form factor cameras,' but Appendix B.3 exempts the detail images from the target blur and noise models. The exemption is justified by pixel-domain blur, yet in angular units the prototype detail camera (f=3.8 mm, f/2.8, 1.25 μm pixels) has IFOV ≈ 1.13 arcmin and diffraction floor ≈ 1.5–1.7 arcmin, while the target detail module (f=1.925 mm, f/1.8, D=1.1 mm) has IFOV ≈ 2.0 arcmin and diffraction floor ≈ 1.9 arcmin by the paper's own Eq. (2). Thus the detail inputs already contain the ~1 arcmin information that the output is praised for, and the ground truth is defined as the same guide capture before its simulated degradation.

full rationale

The paper's physical-limit derivations (diffraction, DOF, head motion, photon budget) are standard, externally grounded results, and the motion and SNR analyses rest on external data and references rather than on a self-citation chain. No load-bearing uniqueness theorem is invoked. The substantive circularity is confined to the evaluation: Sec. 7.1 promises that all ten camera images are degraded to simulate tiny camera modules, but Appendix B.3 explicitly omits target blur and noise on the detail images. In angular units the prototype detail cameras are sharper than the target detail module, which has a roughly 2-arcmin floor by the paper's own Eq. (2) and IFOV. Consequently, the claimed 'close to iPhone' reconstruction quality is substantially forced by feeding the pipeline detail images whose angular resolution already matches the claimed output, rather than by demonstrating that the 2-arcmin target modules can produce it. This makes the central performance claim partially circular, though the underlying algorithmic fusion is real and could be validated by actually degrading the detail images to the target model or by building the target modules. Corrected that way, the evaluation would be self-contained; as presented, the headline result is not a prediction from the target design.

Assumptions & free parameters 4 free parameters · 6 assumptions · 0 invented entities

The optical limits and the SNR analysis rest on standard physics and standard photometric assumptions. The empirical parameters (head-still threshold, noise coefficients, homography depth) are calibrated by the authors from their own data, and the assumption that target tiny cameras share prototype noise behavior is introduced ad hoc for this paper's evaluation.

free parameters (4)
  • head-still motion threshold = 3 deg/sec
    Chosen by inspecting the authors' ten head-still recordings (Sec 4.2.2, Fig 22) to separate intentional fixation from saccadic motion; drives the exposure-time analysis in Fig 3.
  • noise model coefficients (lambda_shot, lambda_read) = Table 2, e.g. desk: 2.4e-4/1.5e-6
    Fitted via mean-variance plots on prototype cameras (Table 2, Fig 21); used to synthesize the noise that mimics tiny modules in the evaluation.
  • homography initialization depth = 100 m
    Used to initialize optical-flow prewarping and RSR ROI search (Sec 6.2.1, 6.2.2, Fig 11); affects the correspondence search range but not the fitted model.
  • example camera module parameters for Fig 4 = f/1.8, 1 micron pixel pitch
    Hypothetical module used for the SNR/illuminance plot; illustrative input, not fitted.
assumptions (6)
  • standard math Standard diffraction and hyperfocal distance formulas apply to compound lenses in tiny cameras.
    Used in Sec 4.1 to derive Eq 1-7.
  • domain assumption Head motion during an exposure is modelled as pure rotation with linear angular velocity within each exposure.
    Used for t_max = delta_theta / omega in Sec 4.2.1; the paper acknowledges it is an upper bound for head-still motion.
  • domain assumption The Aria Pilot Dataset IMU distribution represents general smart-glasses wearer head motion.
    Basis for the motion-blur percentages in Fig 3 and Sec 4.2.1.
  • ad hoc to paper Target tiny modules have noise similar to the XIMEA/Aria prototype cameras.
    Appendix B.2: the fitted noise model is reused for the simulated tiny cameras; specific to this validation strategy.
  • domain assumption The Alakarhu photometric model with ideal lens transmission, ideal color filters, 18% reflectance, and CIE illuminant A gives a valid upper bound on available light.
    Used in Sec 4.3 to derive N_ph and the SNR bound.
  • domain assumption Pixel pitch is small enough not to limit angular resolution in the Sec 4.1 trade-off analysis.
    Explicitly stated in Sec 4.1 to isolate the diffraction-DOF trade-off.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Imaging for All-Day Wearable Smart Glasses." pith.science (2026). https://pith.science/paper/SBBUCP3T

@misc{pith2026250413060,
  author       = {Pith},
  title        = {Pith review of: Imaging for All-Day Wearable Smart Glasses},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SBBUCP3T}},
  note         = {Machine review of arXiv:2504.13060}
}
read the original abstract

In recent years smart glasses technology has rapidly advanced, opening up entirely new areas for mobile computing. We expect future smart glasses will need to be all-day wearable, adopting a small form factor to meet the requirements of volume, weight, fashionability and social acceptability, which puts significant constraints on the space of possible solutions. Additional challenges arise due to the fact that smart glasses are worn in arbitrary environments while their wearer moves and performs everyday activities. In this paper, we systematically analyze the space of imaging from smart glasses and derive several fundamental limits that govern this imaging domain. We discuss the impact of these limits on achievable image quality and camera module size -- comparing in particular to related devices such as mobile phones. We then propose a novel distributed imaging approach that allows to minimize the size of the individual camera modules when compared to a standard monolithic camera design. Finally, we demonstrate the properties of this novel approach in a series of experiments using synthetic data as well as images captured with two different prototype implementations.

Figures

Figures reproduced from arXiv: 2504.13060 by the authors.

Figure 1
Figure 1. Full picture of the proposed distributed camera system. Our software pipeline reconstructs a superresolved color output image. Using a distributed [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Trade-off for a fixed focus camera between angular resolution, en [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Percentage of egocentric image data with minimal motion blur (move [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (19 more)
Figure 4
Figure 4. Figure 4: Illuminance levels required to achieve SNR=10 for a hypothetical [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Distributed camera layout design visualized in 2D. [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Overview of the image processing pipeline. The input is either a single set of guide and detail images or respective bursts of images. After preprocessing, [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Architecture of the fusion network. The “degrader” uses the known properties of our cameras to downsample and degrade the high-resolution images [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Comparing single image vs. burst mode. Left: capturing a single [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Effect of the soft epipolar line constraint in OFW. [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Strengths and weaknesses of the optical flow warping-based (OFW) and reference-based super resolution (RSR) pipeline. [PITH_FULL_IMAGE:figures/full_fig_p011_10.png]
Figure 11
Figure 11. Figure 11: Overview of the ROI cropping and packing during the inference [PITH_FULL_IMAGE:figures/full_fig_p011_11.png]
Figure 12
Figure 12. Figure 12: Depth estimation based on OFW. Left: Naïve averaging of individ￾ual depth maps. Middle: Weighted blend with distance transform. Right: Reconstructed image. The highlighted region illustrates the reduction of discontinuity on the flat wall compared to naïvely averaging…
Figure 14
Figure 14. Figure 14: Comparing our pipeline and other state-of-the-art commercial devices. Top: full images. Bottom: selected detail crops. The image from the guide [PITH_FULL_IMAGE:figures/full_fig_p014_14.png]
Figure 15
Figure 15. Figure 15: Percentage of QR codes successfully recognized over [PITH_FULL_IMAGE:figures/full_fig_p015_15.png]
Figure 18
Figure 18. Figure 18: Zoomed-in samples from the reconstruction of video with burst [PITH_FULL_IMAGE:figures/full_fig_p016_18.png]
Figure 16
Figure 16. Figure 16: Examples demonstrating some success and failure cases of our [PITH_FULL_IMAGE:figures/full_fig_p016_16.png]
Figure 17
Figure 17. Figure 17: Frame-by-frame processing without burst mode: temporally vary [PITH_FULL_IMAGE:figures/full_fig_p016_17.png]
Figure 20
Figure 20. Figure 20: Pictures of the prototypes. Left: desk-mounted prototype. Right: [PITH_FULL_IMAGE:figures/full_fig_p020_20.png]
Figure 21
Figure 21. Figure 21: Left: Example of a noise measurement image from XIMEA camera [PITH_FULL_IMAGE:figures/full_fig_p021_21.png]
Figure 22
Figure 22. Figure 22: Visualization of Aria rotational velocity within a set of recordings in which users wore Aria glasses and intentionally fixated on static scene objects. [PITH_FULL_IMAGE:figures/full_fig_p022_22.png]
Figure 23
Figure 23. Figure 23: Cumulative distribution function (CDF) of instantaneous rotational [PITH_FULL_IMAGE:figures/full_fig_p022_23.png]
Figure 24
Figure 24. Figure 24: Illustration of head-motion traces for Aria pilot dataset (left) and head-still recordings (right). All traces were drawn by selecting a random timestamp [PITH_FULL_IMAGE:figures/full_fig_p023_24.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

    cs.CV 2026-08 conditional novelty 7.0 of 10

    R4DSG builds an anchor-relative 4D scene graph memory from RGB egocentric video and shows improved object-centric QA accuracy over text retrieval.

Reference graph

Works this paper leans on

68 extracted references · 51 canonical work pages · cited by 1 Pith paper

  1. [1]

    Juha Alakarhu. 2007. Image Sensors and Image Quality in Mobile Phones. In 2007 International Image Sensor Workshop

  2. [2]

    Fairchild

    Pontus Andersson, Jim Nilsson, Tomas Akenine-Möller, Magnus Oskarsson, Kalle Åström, and Mark D. Fairchild. 2020. FLIP: A Difference Evaluator for Alternating Images. Proc. ACM Comput. Graph. Interact. Tech. 3, 2, Article 15 (aug 2020), 23 pages. https://doi.org/10.1145/3406183

  3. [3]

    Inessa Bekerman, Paul Gottlieb, and Michael Vaiman. 2014. Variations in eyeball diameters of the healthy adults. Journal of Ophthalmology (2014). https://doi.org/ 10.1155/2014/503645

  4. [4]

    Goutam Bhat, Martin Danelljan, Luc Van Gool, and Radu Timofte. 2021. Deep burst super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9209–9218

  5. [5]

    Taryn Bipat, Maarten Willem Bos, Rajan Vaish, and Andrés Monroy-Hernández

  6. [6]

    Tim Brooks, Ben Mildenhall, Tianfan Xue, Jiawen Chen, Dillon Sharlet, and Jonathan T. Barron. 2018. Unprocessing Images for Learned Raw Denoising.CoRR abs/1811.11127 (2018). arXiv:1811.11127 http://arxiv.org/abs/1811.11127

  7. [7]

    Matthew Brown and David G Lowe. 2007. Automatic panoramic image stitching using invariant features. International journal of computer vision 74 (2007), 59–73

  8. [8]

    Matthew Brown, David G Lowe, et al. 2003. Recognising panoramas. In ICCV, Vol. 3. 1218

Show all 68 references
  1. [9]

    Jiezhang Cao, Jingyun Liang, Kai Zhang, Yawei Li, Yulun Zhang, Wenguan Wang, and Luc Van Gool. 2022. Reference-based image super-resolution with deformable attention transformer. In ECCV. Springer, 325–342

  2. [10]

    Praneeth Chakravarthula, Jipeng Sun, Xiao Li, Chenyang Lei, Gene Chou, Mario Bijelic, Johannes Froesch, Arka Majumdar, and Felix Heide. 2023. Thin On-Sensor Nanophotonic Array Cameras. ACM Trans. Graph. 42, 6, Article 249 (Dec. 2023), 18 pages. https://doi.org/10.1145/3618398

  3. [11]

    Paul E Debevec and Jitendra Malik. 1997. Recovering high dynamic range radiance maps from photographs. In Proceedings of the 24th annual conference on Computer graphics and interactive techniques . 369–378

  4. [12]

    Jakob Engel, Thomas Schöps, and Daniel Cremers. 2014. LSD-SLAM: Large-scale direct monocular SLAM. InComputer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part II 13. Springer, 834–849

  5. [13]

    Jakob Engel, Kiran Somasundaram, Michael Goesele, Albert Sun, Alexander Gamino, Andrew Turner, Arjang Talattof, Arnie Yuan, Bilal Souti, Brighid Mered- ith, Cheng Peng, Chris Sweeney, Cole Wilson, Dan Barnes, Daniel DeTone, David Caruso, Derek Valleroy, Dinesh Ginjupalli, Dunc...

  6. [14]

    Yu Fang, Ryoichi Nakashima, Kazumichi Matsumiya, Ichiro Kuriki, and Satoshi Shioiri. 2015. Eye-Head Coordination for Visual Cognitive Processing. PLOS ONE 10, 3 (03 2015), 1–17. https://doi.org/10.1371/journal.pone.0121035

  7. [15]

    Orazio Gallo, Alejandro Troccoli, Jun Hu, Kari Pulli, and Jan Kautz. 2015. Locally non-rigid registration for mobile HDR photography. In 2015 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) . 48–55. https: //doi.org/10.1109/CVPRW.2015.7301366

  8. [16]

    Michaël Gharbi, Gaurav Chaurasia, Sylvain Paris, and Frédo Durand. 2016. Deep joint demosaicking and denoising. ACM Transactions on Graphics (ToG) 35, 6 (2016), 1–12

  9. [17]

    Gortler, Radek Grzeszczuk, Richard Szeliski, and Michael F

    Steven J. Gortler, Radek Grzeszczuk, Richard Szeliski, and Michael F. Cohen. 1996. The lumigraph. In Proceedings of the 23rd Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’96). Association for Computing Machinery, New York, NY, USA, 43–54. https:/...

  10. [18]

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, 17 Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrish- nan, Fiona Ryan, Jayant Sharma, Michael Wray, M...

  11. [19]

    Selig Hecht, Simon Shlaer, and Maurice Henri Pirenne. 1942. Energy, quanta, and vision. The Journal of general physiology 25, 6 (1942), 819–840

  12. [20]

    DC Hood and MA Finkelstein. 1986. Handbook of Perception and Human Perfor- mance. Vol. 1. Wiley Interscience

  13. [21]

    Howard and Brian J

    Ian P. Howard and Brian J. Rogers. 1995. Binocular Vision and Stereopsis . Oxford University Press

  14. [22]

    Yixuan Huang, Xiaoyun Zhang, Yu Fu, Siheng Chen, Ya Zhang, Yan-Feng Wang, and Dazhi He. 2022. Task decoupled framework for reference-based super- resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5931–5940

  15. [23]

    Photography – Digital still cameras – Determination of exposure index, ISO speed ratings, standard output sensitivity, and recommended exposure index

    ISO 12232:2019(E) 2019. Photography – Digital still cameras – Determination of exposure index, ISO speed ratings, standard output sensitivity, and recommended exposure index. Standard. International Organization for Standardization, Geneva, CH

  16. [24]

    Ophthalmic optics — Spectacle frames — Requirements and test methods

    ISO 12870:2016(E) 2016. Ophthalmic optics — Spectacle frames — Requirements and test methods. Standard. International Organization for Standardization, Geneva, CH

  17. [25]

    Yuming Jiang, Kelvin CK Chan, Xintao Wang, Chen Change Loy, and Ziwei Liu. 2021. Robust reference-based super-resolution via c2-matching. In CVPR. 2103–2112

  18. [26]

    Youngrae Kim, Jinsu Lim, Hoonhee Cho, Minji Lee, Dongman Lee, Kuk-Jin Yoon, and Ho-Jin Choi. 2023. Efficient Reference-based Video Super-Resolution (ERVSR): Single Reference Image Is All You Need. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vis...

  19. [27]

    Yong Min Kim, Sangwoo Bahn, and Myung Hwan Yun. 2021. Wearing comfort and perceived heaviness of smart glasses. Human Factors and Ergonomics in Manufacturing & Service Industries 31 (2021), 484–495. Issue 5

  20. [28]

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al

  21. [29]

    Jason Lawrence, Danb Goldman, Supreeth Achar, Gregory Major Blascovich, Joseph G Desloge, Tommy Fortes, Eric M Gomez, Sascha Häberling, Hugues Hoppe, Andy Huibers, et al. 2021. Project starline: a high-fidelity telepresence system. ACM Transactions on Graphics (TOG) 40, 6 (2021), 1–16

  22. [30]

    Junyong Lee, Myeonghee Lee, Sunghyun Cho, and Seungyong Lee. 2022. Reference-based video super-resolution using multi-camera video triplets. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion. 17824–17833

  23. [31]

    Marc Levoy and Pat Hanrahan. 1996. Light field rendering. In SIGGRAPH 96. 31–42

  24. [32]

    Barron, Dillon Sharlet, Ryan Geiss, Samuel W

    Orly Liba, Kiran Murthy, Yun-Ta Tsai, Tim Brooks, Tianfan Xue, Nikhil Karnad, Qiurui He, Jonathan T. Barron, Dillon Sharlet, Ryan Geiss, Samuel W. Hasinoff, Yael Pritch, and Marc Levoy. 2019. Handheld Mobile Photography in Very Low Light. ACM Trans. Graph. 38, 6, Article 164 (...

  25. [33]

    Bruce D Lucas and Takeo Kanade. 1981. An iterative image registration technique with an application to stereo vision. In IJCAI’81: 7th international joint conference on Artificial intelligence, Vol. 2. 674–679

  26. [34]

    Steve Mann. 2013. My Äugmediated ¨Life. IEEE Spectrum (2013)

  27. [35]

    Meta Reality Labs-R. 2022. Aria Pilot Dataset. https://www.projectaria.com/ datasets/apd/

  28. [36]

    Meta Reality Labs-R. 2023. Aria Synthetic Environments Dataset. https://www. projectaria.com/datasets/ase/

  29. [37]

    Srinivasan, and Jonathan T

    Ben Mildenhall, Peter Hedman, Ricardo Martin-Brualla, Pratul P. Srinivasan, and Jonathan T. Barron. 2022. NeRF in the Dark: High Dynamic Range View Synthesis from Noisy Raw Images. CVPR (2022)

  30. [38]

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 65, 1 (2021), 99–106

  31. [39]

    John A Mordi and Kenneth J Ciuffreda. 1998. Static aspects of accommodation: age and presbyopia. Vision Research 38, 11 (1998), 1643–1653. https://doi.org/10. 1016/S0042-6989(97)00336-2

  32. [40]

    Anastasios I Mourikis and Stergios I Roumeliotis. 2007. A multi-state constraint Kalman filter for vision-aided inertial navigation. In Proceedings 2007 IEEE Inter- national Conference on Robotics and Automation . IEEE, 3565–3572

  33. [41]

    Raul Mur-Artal and Juan D Tardós. 2017. ORB-SLAM2: An open-source slam system for monocular, stereo, and RBG-D cameras. IEEE Transactions on Robotics 33, 5 (2017), 1255–1262

  34. [42]

    Yoshikuni Nomura, Li Zhang, and Shree K Nayar. 2007. Scene collages and flexible camera arrays. In Proceedings of the 18th Eurographics conference on Rendering Techniques. 127–138

  35. [43]

    Alan Pears. 1998. Strategic study of household energy and greenhouse issues . Sus- tainable Solutions Australia

  36. [44]

    Federico Perazzi, Alexander Sorkine-Hornung, Henning Zimmer, Peter Kaufmann, Oliver Wang, Scott Watson, and Markus Gross. 2015. Panoramic video from unstructured camera arrays. In Computer Graphics Forum, Vol. 34. Wiley Online Library, 57–68

  37. [45]

    Marco Pesavento, Marco Volino, and Adrian Hilton. 2021. Attention-based multi- reference learning for image super-resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 14697–14706

  38. [46]

    Rayleigh. 1879. XXXI. Investigations in optics, with special reference to the spec- troscope. The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science 8, 49 (1879), 261–274. https://doi.org/10.1080/14786447908639684

  39. [47]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695

  40. [48]

    Ruth Rosenholtz. 2016. Capabilities and Limitations of Peripheral Vision. Annual Review of Vision Science 2, 1 (2016), 437–457. https://doi.org/ 10.1146/annurev-vision-082114-035733 arXiv:https://doi.org/10.1146/annurev- vision-082114-035733 PMID: 28532349

  41. [49]

    Sahin and Rajiv Laroia

    Furkan E. Sahin and Rajiv Laroia. 2017. Light L16 Computational Camera, In Imaging and Applied Optics 2017 (3D, AIO, COSI, IS, MATH, pcAOP). Imaging and Applied Optics 2017 (3D, AIO, COSI, IS, MATH, pcAOP) , JTu5A.20. https: //doi.org/10.1364/3D.2017.JTu5A.20

  42. [50]

    Zachary Teed and Jia Deng. 2020. Raft: Recurrent all-pairs field transforms for optical flow. In ECCV. Springer, 402–419

  43. [51]

    Tinsley, Maxim I

    Jonathan N. Tinsley, Maxim I. Molodtsov, Robert Prevedel, David Wartmann, Jofre Espigulé-Pons, Mattias Lauwers, and Alipasha Vaziri. 2016. Direct detection of a single photon by humans. Nature Communications 7, 1 (2016), 12172. https: //doi.org/10.1038/ncomms12172

  44. [52]

    Marc Comino Trinidad, Ricardo Martin Brualla, Florian Kainz, and Janne Kontka- nen. 2019. Multi-view image fusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4101–4110

  45. [53]

    Kartik Venkataraman, Dan Lelescu, Jacques Duparré, Andrew McMahon, Gabriel Molina, Priyam Chatterjee, Robert Mullis, and Shree Nayar. 2013. PiCam: an ultra-thin high performance monolithic camera array. ACM Trans. Graph. 32, 6, Article 166 (Nov. 2013), 13 pages. https://doi.or...

  46. [54]

    Cohen, and Matt Uyttendaele

    Jialiang Wang, Daniel Scharstein, Akash Bapat, Kevin Blackburn-Matzen, Matthew Yu, Jonathan Lehman, Suhib Alsisan, Yanghan Wang, Sam Tsai, Jan-Michael Frahm, Zijian He, Peter Vajda, Michael F. Cohen, and Matt Uyttendaele. 2023. A Practical Stereo Depth System for Smart Glasses...

  47. [55]

    Tengfei Wang, Jiaxin Xie, Wenxiu Sun, Qiong Yan, and Qifeng Chen. 2021. Dual- camera super-resolution with aligned attention modules. In ICCV. 2001–2010

  48. [56]

    Bovik, H.R

    Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing 13, 4 (2004), 600–612. https://doi.org/10.1109/TIP.2003.819861

  49. [57]

    Bennett Wilburn, Neel Joshi, Vaibhav Vaish, Eino-Ville Talvala, Emilio Antunez, Adam Barth, Andrew Adams, Mark Horowitz, and Marc Levoy. 2005. High performance imaging using large camera arrays. In ACM SIGGRAPH 2005 Papers . 765–776

  50. [58]

    Bartlomiej Wronski, Ignacio Garcia-Dorado, Manfred Ernst, Damien Kelly, Michael Krainin, Chia-Kai Liang, Marc Levoy, and Peyman Milanfar. 2019. Hand- held Multi-Frame Super-Resolution. ACM Trans. Graph. 38, 4, Article 28 (jul 2019), 18 pages. https://doi.org/10.1145/3306346.3323024

  51. [59]

    Xiaotong Wu, Wei-Sheng Lai, Yichang Shih, Charles Herrmann, Michael Krainin, Deqing Sun, and Chia-Kai Liang. 2023. Efficient Hybrid Zoom Using Camera Fusion on Mobile Phones. ACM Transactions on Graphics (TOG) 42, 6 (2023), 1–12

  52. [60]

    Xiaoyun Yuan, Lu Fang, Qionghai Dai, David J Brady, and Yebin Liu. 2017. Multi- scale gigapixel video: A cross resolution image matching and warping approach. In 2017 IEEE International Conference on Computational Photography (ICCP) . IEEE, 1–9. 18

  53. [61]

    Lin Zhang, Xin Li, Dongliang He, Fu Li, Errui Ding, and Zhaoxiang Zhang

  54. [62]

    Yupeng Zhang, Liyan Liu, Weitao Gong, Haihua Yu, Wei Wang, Chongying Zhao, Peng Wang, and Toshitsugu Ueda. 2018. Autofocus System and Evaluation Method- ologies: A Literature Review. Sensors and Materials 30, 5 (2018), 1165–1174

  55. [63]

    Zhifei Zhang, Zhaowen Wang, Zhe Lin, and Hairong Qi. 2019. Image super- resolution by neural texture transfer. In CVPR. 7982–7991

  56. [64]

    In Proceedings of the IEEE/CVF International Conference on Computer Vision

    LMR: A Large-Scale Multi-Reference Dataset for Reference-based Super- Resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 13118–13127

  57. [65]

    Han Zou, Liang Xu, and Takayuki Okatani. 2023. Geometry Enhanced Reference- Based Image Super-Resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6123–6132. 19 Fig. 20. Pictures of the prototypes. Left: desk-mounted prototype. Rig...

  58. [67]

    Han Zou, Masanori Suganuma, and Takayuki Okatani. 2023. RefVSR++: Exploiting Reference Inputs for Reference-based Video Super-resolution. arXiv preprint arXiv:2307.02897 (2023)

  59. [2019]

    In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19)

    Analyzing the Use of Camera Glasses in the Wild. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–8. https://doi.org/10.1145/3290605.3300651

  60. [2023]

    arXiv preprint arXiv:2304.02643 (2023)

    Segment anything. arXiv preprint arXiv:2304.02643 (2023)

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.