{"id":"56118687-2075-4040-ac32-234874f4a39a","arxiv_id":"2501.15450","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"FlatTrack shows a thin mask-based lensless camera plus a two-stage neural pipeline can estimate eye gaze with about 1.9 degrees average error and >125 fps, and introduces a 20,475-capture near-eye lensless dataset.","lead":"This paper builds an eye-tracking system from an ultra-thin lensless camera, replacing the bulky lens with a flat optical mask. It reports accuracy close to conventional trackers, about 1.9 degrees average angular error, plus a new dataset of 20,475 near-eye captures.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Parity with lensed trackers rests on a simulated lensless comparison (Sec 6.4) that omits real capture non-idealities the paper itself documents in Sec 6.5.","rationale":"The paper makes a real contribution: a 20,475-capture NIR PhlatCam gaze dataset, a two-stage reconstruction-plus-regression pipeline, and a careful within-paper comparison of reconstruction methods and regressors. The 1.92° average error and >125 fps throughput are plausible and not internally inconsistent. The load-bearing weakness is a scope-of-evidence gap rather than a mathematical error: the abstract's headline claim names conventional lens-based trackers, yet the only controlled lensed-vs-lensless comparison is the Sec 6.4 simulation built on the idealized Eq. 1 forward model. The paper's own Sec 6.5 shows that real captures suffer from illumination non-uniformity that the simulation does not encode, so the simulated parity could overstate real-world parity. I agree with the reader that this is the weakest assumption. Since the reader already marked this as the condition for the CONDITIONAL verdict, my assessment does not change the verdict; it sharpens the required condition. The concrete test—a real lensed capture arm under identical conditions—would settle whether the concern lands. If the authors cannot run that comparison, the abstract should be tempered accordingly.","tokens_in":9388,"tokens_out":8639,"duration_ms":81221,"concrete_test":"Perform a same-setup real comparison: in the FlatTrack rig, mount a conventional lensed near-eye camera with comparable resolution and field of view at the same roughly 4 cm eye distance (or capture alternating trials with the PhlatCam and a lensed camera under identical stimulus, illumination, and subject calibration). Collect data from the same 13-subject 15x15-grid protocol, train/test the identical ResNet-18 pipeline, including subject-specific fine-tuning, on both modalities, and compare average angular errors. If the real lensed-vs-lensless gap is comparable to the Table 3 ~0.1° simulation gap, the parity claim stands; if the gap is materially larger, the abstract claim should be weakened to 'competitive in absolute terms on real data, with parity demonstrated only in simulation.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—real lensless gaze tracking is on par with lens-based trackers—is supported only by the Davis-GS simulation in Sec 6.4. In that simulation, lensless measurements are generated as Y = P*X + N (Eq. 1) with a single shift-invariant PSF and added noise, then reconstructed with FlatNet. The real FlatTrack experiments (Secs 6.2–6.3) report 1.92° average angular error but include no same-setup lensed baseline. The forward model omits the spatially varying illumination, correlated sensor noise, head motion, and calibration drift that the authors themselves identify as error sources in Sec 6.5, where high-error grid points are attributed to 'harsher illumination' and 'lack of proper lighting.' Because these non-idealities are absent from the simulation, the simulated lensed-vs-lensless gap (Table 3, about 0.1°) may understate the real gap. With only three subjects and no error bars in Table 3, the simulated parity result is also statistically fragile. The abstract's 'on par' claim therefore needs a real lensed/lensless comparison on the same rig, subjects, and protocol before it can be taken as established for physical captures.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes FlatTrack, a near-eye gaze tracking system built from an NIR PhlatCam lensless camera and a two-stage pipeline: lensless reconstruction (Wiener deconvolution or FlatNet) followed by a gaze regression network (ResNet-18, EyeCOD, MobileNetV2). The authors collect a 20,475-capture dataset from 13 subjects with calibrated gaze labels, evaluate per-subject held-out rounds, report a best average angular error of 1.92 degrees for FlatNet plus ResNet-18 at over 125 fps, and present a simulated lensed-versus-lensless comparison on the Davis-GS dataset that shows roughly a 0.1 degree difference. The central claim is that lensless eye tracking performs on par with conventional lens-based trackers while enabling a flatter and more compact form factor.","tokens_in":9696,"tokens_out":7333,"duration_ms":64229,"significance":"If the parity claim holds, the contribution is significant: it demonstrates that a mask-based lensless camera placed about 4 cm from the eye can support usable gaze estimation, potentially removing the focusing-distance constraint in AR/VR headsets. The paper's concrete strengths include a new real lensless gaze dataset (the first of its kind, to the authors' knowledge), a sound per-subject held-out evaluation protocol for the real-data experiments, comparisons of two reconstruction methods and three regressors, and clearly reported inference-speed measurements. The significance is tempered by the fact that the headline parity result rests on a simulation whose fidelity to real captures is not established, and the dataset and code are promised only upon acceptance.","major_comments":[{"comment":"The abstract's statement that the proposed system 'performs on par with conventional lens-based trackers' is supported only by the simulated Davis-GS comparison in Sec. 6.4, where lensless measurements are generated via Y = P*X + N (Eq. 1) with a single PSF and added noise. The real FlatTrack experiments (Secs. 6.2-6.3) achieve 1.92 degrees average angular error but include no lensed baseline on the same setup, subjects, and protocol. The simulation omits spatially varying illumination, correlated sensor noise, head motion, and calibration drift, which Sec. 6.5 identifies as error sources in the real captures. The parity claim therefore needs a same-setup real lensed-versus-lensless comparison, or the claim should be restricted to the simulated setting.","section":"Sec. 6.4 / Abstract"},{"comment":"The lensed-versus-lensless comparison is reported for only three subjects, with differences of 0.02-0.09 degrees and no confidence intervals or per-subject variance. With n=3, the claim of 'minimal difference' is statistically fragile, especially because Subject 27 is slightly better in the lensless condition. The text also says 21 of the 27 Davis-GS subjects are used for pre-training and 'the remaining 3' for fine-tuning and evaluation, leaving 3 subjects unaccounted for. Please report variability over held-out rounds, include more test subjects, and clarify the subject allocation.","section":"Table 3 / Sec. 6.4"},{"comment":"Sec. 6.5 shows that the largest real-data gaze errors occur at grid points where reconstructed images are degraded by 'harsher illumination' and 'lack of proper lighting.' These effects are not represented in the Sec. 6.4 simulation, which adds only synthetic noise to PSF-convolved images. Consequently, the simulated lensed-versus-lensless gap of about 0.1 degrees may understate the real gap. The paper should quantify how many of the 225 grid points fall into the high-error category and, ideally, evaluate the system under controlled illumination to separate lighting artifacts from lensless-specific degradation.","section":"Sec. 6.5 / Sec. 6.4"},{"comment":"The validity of the Sec. 6.4 simulation depends on the PSF and noise model used to synthesize lensless inputs, but the manuscript does not state whether the PSF is measured from the physical PhlatCam prototype or simulated, nor how the added noise level was set to match real captures. Since the same frozen FlatNet pretrained on simulated MIRFLICKR measurements is used for reconstruction, a PSF or noise mismatch could bias the simulated comparison in favor of the proposed pipeline. Please provide the PSF calibration procedure, the noise-level justification, and a quantitative comparison of simulated and real lensless captures (e.g., residual statistics or paired examples).","section":"Sec. 6.1 / Sec. 6.4"}],"minor_comments":[{"comment":"The manuscript uses both 'Wiener' and 'Weiner' spellings; please standardize to 'Wiener' throughout.","section":"Sec. 6.2 / Table 1"},{"comment":"The 'best-case error' metric reports only the best subject's held-out error, which is not a standard evaluation statistic; please report per-subject errors or a distribution instead of, or in addition to, the best-case value.","section":"Table 2"},{"comment":"The paper should clarify that the 'lensed' condition comes from the Davis-GS dataset and describe the exact capture device for those grayscale frames; the current text reads as if a conventional lens camera and the PhlatCam were compared directly, which is not the case.","section":"Sec. 6.4"},{"comment":"The abstract describes the algorithm as 'co-designed,' but Sec. 5 freezes the reconstruction stage and trains only the gaze regressor; please either demonstrate co-design (e.g., end-to-end training or joint optimization) or reword the claim.","section":"Abstract / Sec. 5"},{"comment":"The text gives the vertical stimulus spacing as 66.3 pixels and the vertical field of view as 29.6 degrees, while the Fig. 3 caption says 65.3 pixels and 29.80 degrees; please reconcile these numbers.","section":"Sec. 4 / Fig. 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for a computational imaging or vision venue, and the dataset is potentially valuable. The main risk is overclaiming in the abstract: if a real same-setup lensed baseline cannot be added, the 'on par' statement should be softened. I would also encourage the editor to consider whether the promised dataset and algorithm release should be a condition of acceptance, since the dataset is a principal contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper delivers the first real near-eye lensless gaze dataset and a solid empirical evaluation of a two-stage pipeline on it. The headline that lensless performs on par with lensed trackers is not established by the paper's own experiments; it rests on a simulated comparison. That's the main thing to know.\n\nWhat's new: 20,475 PhlatCam captures from 13 subjects with calibrated gaze labels, plus the first real-data evaluation of lensless gaze tracking. The dataset alone is a useful contribution, and the authors are honest that prior work (EyeCOD) was simulation-only. The evaluation structure is sound: per-subject held-out rounds, fixed reconstruction, fine-tuned regressor. The numbers are concrete: FlatNet+ResNet-18 gives 1.92 degrees average angular error, EyeCOD 1.82, MobileNet 2.43, with inference times. That is a real result on real captures.\n\nThe soft spots are in the parity claim, not the core demo. The abstract's 'on par with conventional lens-based trackers' depends entirely on Sec 6.4, where lensless measurements are synthesized by convolving Davis-GS lens images with the PhlatCam PSF and adding noise. That forward model misses exactly the non-idealities the authors themselves identify in Sec 6.5—spatially varying illumination, head motion, calibration drift. The simulated gap of about 0.1 degrees is thus likely optimistic. There are also only three subjects and no error bars in Table 3, and no same-setup real lensed baseline in the real experiments, so 1.92 degrees cannot be directly benchmarked against a lens on the same rig. The dataset and code are promised but not yet public, which limits immediate reuse.\n\nNone of this is fatal. The real-data accuracy is measured honestly with held-out rounds, and the authors flag the illumination issues themselves. The overreach is in the abstract's phrasing, not in the dataset or the pipeline demonstration.\n\nWho this is for: people working on lensless imaging, computational eye tracking, and AR/VR hardware. It deserves a serious referee—the dataset is valuable, the claims are clear enough to be fixed in revision, and the field needs this kind of real-data contribution. I'd send it out.","headline":"First real lensless near-eye gaze dataset and a sound real-data evaluation; the 'parity with lensed' claim rests on simulation and should be tempered.","tokens_in":10204,"tokens_out":2030,"would_cite":true,"duration_ms":18272,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A flat, mask-based lensless camera tracks gaze on par with lensed trackers, at over 125 fps.","keywords":["lensless imaging","eye tracking","gaze estimation","PhlatCam","phase mask camera","near-eye gaze dataset","FlatNet reconstruction","augmented reality"],"falsifier":"Set up a rig with a lensed camera and the PhlatCam imaging the same eye from the same 4 cm distance on the same 15 by 15 gaze grid with the same illumination, and compare per-subject angular errors. If the real lensless error is substantially larger than the simulated 1.62 to 1.84 degrees and no longer sits within the lensed error band of roughly 1.6 to 1.8 degrees, the paper's central on-par claim fails. A cheaper check is to capture FlatTrack-style images with two NIR illuminators placed to cover the dark corners identified in Section 6.5 and see whether the average error drops below 1.92 degrees.","tokens_in":9243,"feed_emoji":"👁","tokens_out":10724,"duration_ms":85356,"temperature":0.7,"pith_summary":"This paper tries to establish that eye tracking does not need a focusing lens: a mask-based lensless camera placed a few centimetres from the eye can estimate gaze with accuracy comparable to conventional lens-based trackers, while being flat enough to sit inside an eyeglasses frame. To make this case, the authors built a near-infrared PhlatCam phase-mask camera, collected the first real lensless gaze dataset of 20,475 captures from 13 subjects with calibrated gaze directions, and paired it with a two-stage pipeline that first reconstructs an eye image and then regresses the gaze vector. On this dataset the best configuration, FlatNet reconstruction with a ResNet-18 regressor, reports an average angular error of 1.92 degrees at more than 125 frames per second. A separate simulated comparison on the Davis-GS dataset shows lensless and lensed imaging within roughly 0.1 degrees of each other, which is the evidence behind the 'on par' claim.","feed_headline":"Flat lensless eye tracker hits 1.92° gaze error in real time","feed_subtitle":"A phase-mask camera sitting just 4 cm from the eye rivals lensed trackers while staying thin enough for AR glasses.","key_machinery":"The load-bearing object is the PhlatCam phase-mask camera: a thin phase mask placed slightly under 1.5 mm from the sensor, designed for 700 nm near-infrared light, produces a point-spread function $P$ such that the measurement is $Y = P \\ast X + N$, with the scene $X$ globally multiplexed. Because the mask-to-sensor distance is tiny, the camera has a very large depth of field, so it can image the eye from roughly 4 cm away without a focusing element. The argument is carried by a two-stage pipeline: FlatNet, a learned reconstruction network trained on simulated measurements of natural images, reconstructs a visible eye image from the lensless capture; then ResNet-18 regresses a 3D gaze unit vector, projected to 2D monitor pixels where an $\\ell^1$ loss is applied. Subject-specific fine-tuning of the last two layers of the regressor handles the user calibration that real AR/VR deployments need.","core_discovery":"The central claim is that a thin phase-mask lensless camera can replace a conventional lensed camera for near-eye gaze estimation without a meaningful accuracy penalty. The paper reports that on its own FlatTrack dataset, the FlatNet plus ResNet-18 pipeline achieves 1.92 degrees average angular error (best subject 0.91 degrees) with inference at 7.81 ms, i.e., over 125 fps, and that a 5 percent accuracy gain from the EyeCOD method is not worth its 3x inference cost. To support the parity claim, the paper simulates lensless captures from the lensed Davis-GS dataset by convolving with the PhlatCam PSF and adding noise, then reconstructs with FlatNet; across three test subjects the lensed error is 1.79, 1.72, and 1.67 degrees while the lensless error is 1.84, 1.81, and 1.62 degrees. The paper interprets this as minimal difference, indicating that lensless imaging does not introduce significant performance loss while enabling a compact form factor.","pith_inferences":["The paper does not collect a real lensed-versus-lensless comparison on the same rig; its parity numbers come from simulated lensless Davis-GS images, so a same-setup real comparison could confirm or narrow the claimed parity.","Section 6.5 attributes high-error grid points to dark reconstructions with poor pupil contrast, so adding NIR illuminators aimed at those corners is a direct, testable fix that could push the 1.92 degree average lower.","The globally multiplexed raw measurements suggest a privacy property beyond the form-factor win, but the paper does not quantify how much harder identity or screen-content reconstruction becomes from FlatTrack data."],"forward_implications":["AR/VR headsets can place the eye-tracking camera inside the eyeglasses frame rather than at a lens's focusing distance, cutting thickness and weight.","At over 125 fps with 7.81 ms inference, the pipeline is fast enough for interactive gaze-based interfaces and foveated rendering.","The 20,475-capture dataset gives the lensless-gaze community a real-data training and benchmarking target instead of relying only on simulated lensless images.","User-specific fine-tuning with a small calibration set keeps accuracy near 1 to 2 degrees without per-subject retraining from scratch.","The parity result predicts that progress on lensed gaze trackers can transfer to lensless hardware through the same reconstruction-plus-regression pipeline."],"supporting_citations":[{"why":"Supplies the NIR PhlatCam phase-mask camera hardware, PSF model, and the imaging forward model used throughout.","marker":"[7]"},{"why":"Supplies FlatNet, the learned reconstruction network used as the frozen first stage of the gaze pipeline.","marker":"[13]"},{"why":"Supplies the Davis-GS near-eye dataset whose lensed images are convolved with the PSF to create the simulated lensless comparison.","marker":"[3]"},{"why":"Supplies the EyeCOD baseline, the FlatCam-based gaze estimator that the paper compares against ResNet-18.","marker":"[23]"},{"why":"Establishes the mask-based lensless imaging formulation that motivates replacing the lens with a thin coded mask.","marker":"[5]"},{"why":"OpenEDS2020 is considered as an alternative lensed gaze dataset but rejected because it lacks user-specific labels, motivating the use of Davis-GS.","marker":"[18]"},{"why":"Supplies the ResNet-18 architecture used as the gaze regressor.","marker":"[10]"},{"why":"Supplies the MIRFLICKR natural images used to pre-train FlatNet on simulated lensless measurements.","marker":"[11]"}],"fun_headline_variants":["Ultra-thin lensless camera tracks gaze at 1.92° error","FlatTrack: lensless eye tracking rivals lensed systems","Lensless camera eye tracker: 125 fps, 1.92° error","Eye tracking with lensless cameras matches lensed accuracy","Flat lensless eye tracker fits in eyeglasses, hits 125 fps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simulated lensless images used for the lensed-versus-lensless comparison, Davis-GS lens images convolved with the PhlatCam PSF plus noise, faithfully represent what a real PhlatCam would record; real captures include uneven illumination, sensor noise correlations, head motion, and calibration drift that the simulation omits.","fun_headline_variants_meta":{"raw":{"variants":["Ultra-thin lensless camera tracks gaze at 1.92° error","FlatTrack: lensless eye tracking rivals lensed systems","Lensless camera eye tracker: 125 fps, 1.92° error","Eye tracking with lensless cameras matches lensed accuracy","Flat lensless eye tracker fits in eyeglasses, hits 125 fps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1401,"prompt_tokens":938,"completion_tokens":463,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":368}},"tokens_in":554,"tokens_out":463,"duration_ms":4492,"temperature":1.0,"reasoning_tokens":368,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:16:05.518972+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Set up a rig with a lensed camera and the PhlatCam imaging the same eye from the same 4 cm distance on the same 15 by 15 gaze grid with the same illumination, and compare per-subject angular errors. If the real lensless error is substantially larger than the simulated 1.62 to 1.84 degrees and no longer sits within the lensed error band of roughly 1.6 to 1.8 degrees, the paper's central on-par claim fails. A cheaper check is to capture FlatTrack-style images with two NIR illuminators placed to cover the dark corners identified in Section 6.5 and see whether the average error drops below 1.92 degrees.","supporting_citations":[{"cited_title":"Phlatcam: Designed phase-mask based thin lensless camera","cited_arxiv_id":null,"evidence_quote":"Supplies the NIR PhlatCam phase-mask camera hardware, PSF model, and the imaging forward model used throughout."},{"cited_title":"Flatnet: Towards photorealistic scene reconstruction from lensless measure- ments","cited_arxiv_id":null,"evidence_quote":"Supplies FlatNet, the learned reconstruction network used as the frozen first stage of the gaze pipeline."},{"cited_title":"Angelopoulos, Julien N.P","cited_arxiv_id":null,"evidence_quote":"Supplies the Davis-GS near-eye dataset whose lensed images are convolved with the PSF to create the simulated lensless comparison."},{"cited_title":"Eyecod: eye tracking system acceleration via flatcam-based algorithm and accel- erator co-design","cited_arxiv_id":null,"evidence_quote":"Supplies the EyeCOD baseline, the FlatCam-based gaze estimator that the paper compares against ResNet-18."},{"cited_title":"Salman Asif, Ali Ayremlou, Aswin Sankaranarayanan, Ashok Veeraraghavan, and Richard G","cited_arxiv_id":null,"evidence_quote":"Establishes the mask-based lensless imaging formulation that motivates replacing the lens with a thin coded mask."},{"cited_title":"Deep residual learning for image recognition","cited_arxiv_id":null,"evidence_quote":"Supplies the ResNet-18 architecture used as the gaze regressor."},{"cited_title":"The mir flickr retrieval evaluation","cited_arxiv_id":null,"evidence_quote":"Supplies the MIRFLICKR natural images used to pre-train FlatNet on simulated lensless measurements."}],"review_version":1}