{"id":"9acbcccd-d333-4280-8388-5af9a5836e27","arxiv_id":"2607.05837","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":6,"one_line_summary":"A Debye CZT-based wave-optics pipeline generates lens-diverse synthetic defocus blur datasets that improve cross-device deblurring generalization over existing real and synthetic data.","lead":"This paper builds a pipeline to synthesize realistic defocus blur for diverse compound lenses using wave-optics PSF computation, depth-aware rendering, and camera ISP simulation. A smart generalist might read it to understand how synthetic data can replace hard-to-collect real photos for training robust computer vision models.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Cross-device generalization claim rests on no-reference metrics while full-reference metrics show the opposite trend; the reliability of no-reference metrics for defocus deblurring evaluation is not independently validated.","rationale":"The paper presents a well-engineered pipeline with detailed optical modeling, and the PSF computation approach is a genuine technical contribution with reasonable validation (Zemax comparison, sampling analysis, ablation). The dataset construction is thorough. However, the central empirical claim — improved cross-device generalization — is not as well-supported as it could be. The full-reference metrics on real benchmarks consistently favor the DPDD-trained model, and the paper's rebuttal (GT imperfections bias these metrics) is plausible but unvalidated. The no-reference metrics that favor CLDefocus are general-purpose image quality measures not specifically designed for deblurring evaluation, and their reliability for this task is not established. The downstream task results are too marginal (differences of 0.002–0.004) to serve as strong independent corroboration. A CONDITIONAL verdict reflects that the contribution is solid but the headline claim would be substantially strengthened by either (a) validating no-reference metric reliability via human study or task-specific metrics, or (b) showing that full-reference metric disadvantages disappear on a filtered subset with verified GT quality. The reader's ACCEPT verdict is reasonable given the engineering quality and partial evidence, but the evaluation gap warrants a more cautious assessment. I set agreement_with_reader to 'partial' because the reader correctly identified a real concern (PSF accuracy) but the more load-bearing issue for the central claim is the evaluation methodology, which the reader did not flag.","tokens_in":31420,"tokens_out":3406,"duration_ms":215295,"concrete_test":"Select 30–50 image pairs from RealDOF or RTF where ground-truth quality can be independently verified (minimal misalignment, no residual blur, consistent exposure — screened by manual inspection or a structural similarity threshold between input and GT). On this filtered subset, recompute PSNR/SSIM/LPIPS for both CLDefocus-trained and DPDD-trained models. If the CLDefocus-trained model still trails on full-reference metrics within this high-quality subset, the argument that full-reference metric disadvantages are purely GT artifacts weakens, and the reliance on no-reference metrics becomes harder to justify. Conversely, if the gap narrows or reverses, the paper's argument is strengthened.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim — that CLDefocus training improves cross-device generalization — is supported primarily by no-reference metrics (NIQE, MUSIQ, TOPIQ) on real benchmarks (RTF, RealDOF, DPDD). On full-reference metrics (PSNR, SSIM, LPIPS), the DPDD-trained model generally outperforms the CLDefocus-trained model on these same real benchmarks (Table 1, Table S4). The paper addresses this discrepancy in Sec. 5.2 by arguing that real ground-truth images contain geometric misalignment, photometric inconsistency, and residual blur, which bias pixel-wise metrics toward blur-preserving outputs. This argument is plausible and supported by qualitative examples (Fig. 4). However, the paper does not establish the converse: that no-reference metrics are reliable for evaluating defocus deblurring quality. NIQE, MUSIQ, and TOPIQ measure general image quality, not deblurring fidelity. A model that oversharpened, introduced high-frequency artifacts, or shifted color/contrast in ways preferred by these metrics could score higher without producing more faithful deblurring. The downstream task results (Tables S2, S3) provide weak corroboration — improvements over DPDD are marginal (RMSE 0.246 vs. 0.248; IoU 0.871 vs. 0.867). Additionally, on the DPDD test set itself, the DPDD-trained model wins on MUSIQ and TOPIQ (Table 2), showing that no-reference metrics do not uniformly favor CLDefocus. The claim thus hinges on an unvalidated assumption that no-reference metrics are less biased than full-reference metrics for this specific task. The reader's identified concern (PSF accuracy via Debye CZT) is valid but partially addressed by the Zemax comparison (Fig. S4) and the PSF ablation (Table 3); the evaluation methodology concern is less addressed and more directly load-bearing for the claim as stated.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This paper proposes a pipeline for synthesizing realistic defocus blur datasets for compound lenses. The pipeline integrates three components: (1) efficient wave-optics PSF computation via the Debye CZT formulation with explicit sampling criteria, (2) depth-aware defocus rendering with occlusion handling via layered compositing, and (3) blur synthesis in radiometrically linear space with camera ISP simulation. Using this pipeline, the authors generate CLDefocus, a dataset of 40,000 training pairs spanning 700 lens designs. Experiments compare models trained on CLDefocus against those trained on DPDD (real-captured) and SYNDOF (synthetic with simplified blur) across multiple real benchmarks (RTF, RealDOF, DPDD) and four deblurring architectures. The central claim is that CLDefocus training improves cross-device generalization. The paper also analyzes imperfections in real-captured ground truth that bias full-reference metrics.","tokens_in":32271,"tokens_out":1060,"duration_ms":218516,"significance":"The paper makes a solid contribution to defocus deblurring dataset synthesis. Strengths include: (1) a parameter-free sampling criterion (Eq. 2) for stable PSF computation, addressing a practical bottleneck in prior wave-optics approaches; (2) a 2500x speedup over Rayleigh-Sommerfeld (Sec. 5.3, Fig. 5) with explicit aliasing control; (3) validation across four deblurring architectures (NRKNet, Restormer, INIKNet, NAFNet) in Table S4; (4) reproducible code and dataset publicly available; (5) the smartphone evaluation (Sec. S6) and downstream task results (Sec. S7) provide additional evidence beyond standard benchmarks. The analysis of real-GT imperfections (Sec. 5.2) is a useful contribution to the evaluation methodology discussion. The lens diversity analysis (Fig. S6) provides evidence that the 700-lens collection spans a meaningful range of optical properties.","major_comments":[{"comment":"Sec. 5.1, Tables 1-2: The central claim that CLDefocus improves cross-device generalization is supported primarily by no-reference metrics (NIQE, MUSIQ, TOPIQ), while full-reference metrics (PSNR, SSIM, LPIPS) on real benchmarks (RTF, RealDOF, DPDD) generally favor the DPDD-trained model. The paper argues in Sec. 5.2 that real GT imperfections bias pixel-wise metrics toward blur-preserving outputs. This argument is plausible and supported by qualitative examples (Fig. 4), but the converse — that no-reference metrics reliably measure deblurring fidelity — is not independently established. NIQE, MUSIQ, and TOPIQ measure general image quality, not deblurring accuracy; a model that oversharpened or introduced high-frequency artifacts could score higher without producing more faithful deblurring. The downstream task results (Tables S2-S3) provide only marginal corroboration (RMSE 0.246 vs 0.0","section":null}],"minor_comments":[{"comment":"Sec. 4.2: The depth estimation relies on Depth Pro, a monocular estimator. The impact of depth estimation errors on synthesis quality is acknowledged in Sec. 6 but not quantified. A brief sensitivity analysis or discussion of failure modes would strengthen the paper.","section":null},{"comment":"Table 2: On the DPDD test set, the DPDD-trained model wins on MUSIQ and TOPIQ, showing that no-reference metrics do not uniformly favor CLDefocus. This is actually informative for the reader and could be discussed more explicitly to characterize when each training set is advantageous.","section":null},{"comment":"Sec. S4.2: The noise coefficients beta_1_ref = beta_2_ref = 1e-5 are described as much smaller than prior work. A brief justification or reference for this choice would help reproducibility.","section":null},{"comment":"Fig. 5: The runtime comparison between Debye CZT and Rayleigh-Sommerfeld is informative but the N values for R-S are only inline in the text. A small table would improve readability.","section":null},{"comment":"Sec. 5.1: The training protocol matches total iterations across datasets (350,000 for NRKNet), but datasets differ in size (40,000 vs 350 pairs). The interaction between dataset size and training duration deserves brief discussion.","section":null},{"comment":"The term 'photorealistic' is used in the abstract and throughout, but photorealism is not directly validated via human evaluation. Consider softening to 'physically grounded' or similar.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The stress-test concern about no-reference metric reliability is valid and is the main issue to address in revision. However, I assess it as addressable within minor revision: the paper already acknowledges the limitation, provides qualitative evidence, and includes corroborating downstream task results. The core pipeline contribution (efficient PSF computation with sampling criteria, ISP-aware synthesis, lens diversity) stands independently of the generalization claim and is itself a solid contribution."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for the careful and constructive review. The referee's assessment is accurate, and the recommendation of minor revision is appropriate. We address the single major comment below.","responses":[{"response":"The referee raises a valid and important concern. We agree that no-reference metrics alone do not conclusively establish deblurring fidelity, and the current manuscript does not adequately address the risk that these metrics could reward oversharpening or artifact introduction rather than faithful deblurring. We will revise the manuscript to strengthen the argument on multiple fronts. First, we will add explicit discussion acknowledging the limitation of no-reference metrics for evaluating deblurring specifically, including the oversharpening concern the referee identifies. Second, we will expand the downstream task results discussion to emphasize that these provide task-level evidence independent of image quality metrics: the depth estimation RMSE improvement (0.246 vs. 0.248, Table S2) and segmentation IoU improvement (0.871 vs. 0.867, Table S3) show that CLDefocus-trained models produce restorations that are more useful for downstream vision, which would not hold if the model were merely introducing artifacts. We acknowledge these margins are modest, and we will state this honestly. Third, we will add a brief discussion pointing to the qualitative results (Figs. 3, S12–S14) and the smartphone evaluation (Sec. S6) as additional evidence that the improvements reflect genuine deblurring rather than artifact introduction, since the smartphone PSFs differ substantially from the training distribution and artifact-only gains would not be expected to generalize. We do not claim to fully resolve the fundamental difficulty of evaluating deblurring without reliable ground truth — this is an open problem — but we will present the evidence more completely and temper the central claim accordingly.","revision_made":"partial","referee_comment":"Sec. 5.1, Tables 1-2: The central claim that CLDefocus improves cross-device generalization is supported primarily by no-reference metrics (NIQE, MUSIQ, TOPIQ), while full-reference metrics (PSNR, SSIM, LPIPS) on real benchmarks (RTF, RealDOF, DPDD) generally favor the DPDD-trained model. The paper argues in Sec. 5.2 that real GT imperfections bias pixel-wise metrics toward blur-preserving outputs. This argument is plausible and supported by qualitative examples (Fig. 4), but the converse — that no-reference metrics reliably measure deblurring fidelity — is not independently established. NIQE, MUSIQ, and TOPIQ measure general image quality, not deblurring accuracy; a model that oversharpened or introduced high-frequency artifacts could score higher without producing more faithful deblurring. The downstream task results (Tables S2-S3) provide only marginal corroboration (RMSE 0.246 vs 0.2"}],"tokens_in":31379,"tokens_out":605,"duration_ms":175939,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Two things to know up front: (1) the Debye CZT pipeline for computing compound-lens PSFs is genuinely new and well-executed, and (2) the central cross-device generalization claim is real but rests almost entirely on no-reference metrics, which is a gap the paper does not fully close.","headline":"Cross-device generalization claim rests on no-reference metrics while full-reference metrics show the opposite trend; the reliability of no-reference metrics for defocus deblurring evaluation is not independently validated.","tokens_in":32285,"tokens_out":141,"would_cite":false,"duration_ms":87197,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Synthetic lens blur from wave optics beats real-captured training data","keywords":["defocus deblurring","point spread function","wave optics","compound lens","Debye formulation","chirp Z-transform","synthetic dataset","image restoration"],"falsifier":"If the computed PSFs deviate significantly from real lens blur — particularly for off-axis fields or lenses with complex pupil shapes — the synthetic dataset would not improve cross-device generalization, and the performance gains over simpler blur models would disappear.","tokens_in":31542,"feed_emoji":"📷","tokens_out":1063,"duration_ms":86476,"temperature":0.7,"pith_summary":"This paper argues that physically accurate, lens-diverse defocus blur can be synthesized at scale by computing wave-optics point spread functions (PSFs) for real compound lens designs using the Debye formulation accelerated by the chirp Z-transform (CZT), then rendering depth-aware blur in radiometrically linear space with camera ISP simulation. The authors show that this pipeline is both physically grounded and computationally efficient — the Debye CZT computes a PSF in 0.018 seconds where a Rayleigh-Sommerfeld integral takes 45 seconds for comparable accuracy. Using 700 filtered photographic lens designs, they generate CLDefocus, a dataset of 42,000 blurred-sharp image pairs. The central experimental claim is that deblurring models trained on CLDefocus generalize better across unseen real cameras and lenses than models trained on either a real-captured dataset (DPDD, limited to one camera) or a simplified synthetic dataset (SYNDOF, using Gaussian blur kernels). The authors also demonstrate that imperfections in real-captured ground-truth images — geometric misalignment, photometric drift, residual blur — bias pixel-wise evaluation metrics toward blur-preserving predictions, which helps explain why CLDefocus-trained models score lower on PSNR against flawed ground truth while achieving higher perceptual quality.","feed_headline":"Synthetic lens blur from wave optics beats real-captured training data","feed_subtitle":"Deblurring models trained on 700 lens designs' worth of physics-based PSFs generalize better across cameras than real or simplified datasets","key_machinery":"Debye CZT","core_discovery":"The Debye CZT provides an explicit sampling criterion (N > 4NA²√(n²_t − NA²)|z|/λ) that determines the minimum grid resolution for aliasing-free PSF computation, eliminating the empirical sampling tuning required by Huygens-principle methods. This makes it feasible to compute physically accurate, lens-specific PSFs across hundreds of compound lens designs and depth configurations, producing a synthetic dataset whose optical diversity exceeds what real capture can achieve. When used to train deblurring networks, this dataset yields measurably better cross-device generalization on no-reference perceptual metrics across four benchmark datasets and smartphone images, while also improving depth-估","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Physics-based lens blur synthesis improves cross-camera deblurring","Debye CZT wave optics enables aliasing-free PSF across 700 lens designs","Compound-lens defocus dataset beats real capture for cross-device transfer","Synthetic defocus blur from wave optics generalizes better than real data","Lens-specific PSF pipeline scales defocus dataset generation beyond real capture"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The Debye formulation with scalar diffraction accurately models the blur produced by real photographic compound lenses. This is least reliable for severe off-axis fields, where the Debye approximation and imperfect ray-clipping masks can produce inaccurate PSFs.","fun_headline_variants_meta":{"raw":{"variants":["Physics-based lens blur synthesis improves cross-camera deblurring","Debye CZT wave optics enables aliasing-free PSF across 700 lens designs","Compound-lens defocus dataset beats real capture for cross-device transfer","Synthetic defocus blur from wave optics generalizes better than real data","Lens-specific PSF pipeline scales defocus dataset generation beyond real capture"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":644,"prompt_tokens":551,"completion_tokens":93,"prompt_tokens_details":null},"tokens_in":551,"tokens_out":93,"duration_ms":43332,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T22:47:17.373392+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the computed PSFs deviate significantly from real lens blur — particularly for off-axis fields or lenses with complex pupil shapes — the synthetic dataset would not improve cross-device generalization, and the performance gains over simpler blur models would disappear.","supporting_citations":[],"review_version":1}