Pith. sign in

REVIEW 4 major objections 4 minor 3 references

A single blind network can correct severe, spatially varying lens aberrations across minimalist optics, metalenses, misaligned lenses, and high-end DSLRs without per-lens PSF calibration.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 20:58 UTC pith:OHH45B2G

load-bearing objection Solid incremental advance in blind aberration correction, with a useful data library and PSF-latent guidance, but the uniformity claim is partly self-confirming and reproducibility/framing issues need attention. the 4 major comments →

arxiv 2511.17126 v5 pith:OHH45B2G submitted 2025-11-21 eess.IV cs.CVcs.LGphysics.optics

Towards Blind Lens Aberration Correction via Large LensLib Pre-training and Discrete Degradation Priors

classification eess.IV cs.CVcs.LGphysics.optics
keywords blind lens aberration correctioncomputational aberration correctionlens library pretrainingpoint spread functionvector quantizationspatially varying degradationzero-shot generalizationimage restoration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether blind lens aberration correction—restoring images from unknown lenses without per-lens PSF calibration—can generalize across the full range of real optics. The paper answers yes, provided the training lens library is balanced in both degradation severity and spatial-variation pattern, and provided the correction network is guided by a latent discrete prior learned from PSFs. On the data side it builds AODLibpro, a 3,600-lens library sampled uniformly over eighteen severity-by-pattern classes, adding aspheric surfaces and image-plane perturbation to widen coverage. On the model side it proposes LPR, a vector-quantized codebook of PSF features regularized by an optical-degradation network, so a blind encoder can pull degradation priors out of the image itself. If right, one pretrained model can replace per-lens calibration for cheap minimalist optics and high-end lenses alike.

Core claim

The paper establishes that the bottleneck in blind aberration correction is not model capacity alone but two design choices: the distribution of the lens library and the form of degradation guidance. It constructs AODLibpro by enriching the lens-source specifications and sampling uniformly over severity (a weighted PSNR/SSIM/SFR composite) and spatial-variation class (six patterns defined from per-FoV quality trends). It then learns LPR, a discrete codebook of latent PSF features via vector-quantized autoencoding, supervised by an Optical Degradation Network that must use the code to degrade a clear image into the observed one. During correction, a UNet-style model predicts and retrieves cod

What carries the argument

Two coupled objects carry the argument. AODLibpro is the data engine: a lens library generated with enriched optical design specifications and sampled with a hybrid basis crossing three severity classes (OIQ) with six spatial-variation classes (OD-Class), giving eighteen subclasses and a deliberately uniform aberration distribution. LPR is the guidance engine: a vector-quantized autoencoder turns PSF maps into a discrete codebook of latent PSF features, while an Optical Degradation Network—conditioning a clear-image encoder on quantized PSF codes to reproduce the degraded image—regularizes the codes to carry optical meaning. At inference a separate encoder predicts latent features from the d

Load-bearing premise

The load-bearing premise is that the hand-weighted PSNR/SSIM/SFR composite (OIQ) and the six-category spatial-pattern taxonomy (OD-Class), with its empirically chosen threshold, faithfully capture the optical-degradation structure that matters—if real PSF variations such as chromatic structure or non-monotonic field patterns slip through both, the library is not truly uniform and the scalability conclusions weaken.

What would settle it

Cluster AODLibpro's 3,600 training lenses by their full PSF maps rather than by OIQ summaries: if some of the eighteen OD subclasses contain near-duplicate PSF kernels while sizable PSF-shape clusters fall outside them, the hybrid sampling basis has not balanced the true degradation space and the reported scalability gain should not transfer to the missing patterns. A single counterexample lens—classified as 'spatial-uniform' but with PSFs that alternate between sagittal- and tangential-dominated lobes across the field—would directly expose the taxonomy's insufficiency.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • A single blind model can correct severe aberrations from minimalist spherical/aspheric lenses and metalenses, handle stochastic misalignment blur, and improve high-end DSLR images without per-lens PSF calibration at inference.
  • Scaling a lens library pays off only when samples are uniform in degradation severity and spatial-variation pattern; the paper finds larger improvements from 1% to 100% data with the new library than with the previous RMS-sampled one.
  • LPR guidance beats direct PSF prediction, PSF-feature prediction, and variants with only the vector-quantized codebook or only the degradation network, and its benefit grows as the library scales.
  • The frozen discrete prior makes few-shot adaptation to a new lens efficient, since the codebook transfers instead of being retrained per lens.
  • Predicted latent PSF features are discriminative across degradation patterns—the attention maps align with per-field degradation—supporting the claim that the guidance is optical rather than categorical.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The LPR codebook could be reused as a blind descriptor of optical behavior, not just a correction guide: since its entries track per-field degradation patterns, clustering unknown lenses by their retrieved codes would partition optics by effect rather than by design type.
  • If the approach transfers, optical designers could ship cheaper, aberration-heavy optics and let a universal correction model absorb the quality loss—a different cost/quality trade-off for minimalist and metalens systems than the paper spells out.
  • A direct stress test would apply the same frozen codebook to PSF structures outside the generated lens-source space (e.g., metasurfaces or diffractive elements), which the paper names as future data extensions but does not claim to cover.
  • The OIQ weights are hand-set; recomputing AODLibpro with different weights or a learned perceptual metric would probe how much of the scalability result rests on the specific proxy.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a blind lens aberration correction framework (called OmniLens++ in the main text and FoundCAC in the arXiv metadata abstract) with two main contributions. On the data side, it constructs AODLibpro by extending an EAOD-based lens-source generator with aspheric surfaces and image-distance perturbation, then samples lenses uniformly over a hybrid basis formed by an OIQ severity score (Eq. 11) and an OD-Class taxonomy of spatial-variation patterns, yielding 18 subclasses with 200 train and 3 test lenses per subclass. On the model side, it introduces LPR, a VQVAE-based discrete representation of PSF maps regularized by an ODN that simulates the optical degradation process; the learned codebook is frozen and used to guide a UNet/Swin restoration backbone via predicted latent PSF features. Experiments include a new AODLibproTest benchmark, a RealLens-Sim suite covering minimalist spherical/aspheric lenses, a metalens, misaligned smartphone lenses, and high-end lenses, plus real-snapped qualitative and non-reference quantitative evaluations. The paper reports state-of-the-art zero-shot results, with AODLibpro and LPR each contributing average PSNR gains over the OmniLens baseline on RealLens-Sim.

Significance. If the reported results hold, the paper makes a meaningful advance in a relatively underexplored area: it demonstrates that a large automatically designed lens library can be made more scalable and more uniform, and that PSF-derived priors can be injected in a fully blind manner through a learned discrete latent representation. The evaluation is broader than in prior work, spanning simulated, synthetic-benchmark, and real-captured data, and the appendix gives substantial implementation detail. The paper is also careful to ablate data-specification choices, sampling bases, representation components, and codebook size, and promises open release of code, lens design files, PSF arrays, and simulated/real images. However, the central scalability claim relies on a proxy that is partly used both to construct and to validate the library; several quantitative claims in the paper have internal inconsistencies; and the arXiv abstract claims a few-shot adaptation capability that is not evaluated in the body. These issues do not invalidate the overall framework, but they do require correction and additional validation before the paper can be accepted.

major comments (4)
  1. [Abstract / Sections 4-5] The arXiv metadata abstract states that the framework 'unlock[s] highly efficient few-shot adaptation for unseen lenses.' The body contains no few-shot experiments: Section 4 reports only zero-shot evaluations, and Section 5 explicitly lists 'a flexible fine-tuning pipeline is urged' as future work. This is a load-bearing capability claim with no supporting evidence. Either add a few-shot adaptation experiment (e.g., fine-tuning on a small number of real-lens pairs) or remove the claim from the abstract.
  2. [Section 4.4, Table 8] The percentage improvements in Table 8 are internally inconsistent. For LPR at 1% scale, PSNR decreases from 23.82 to 23.39 (about -1.8%), yet the table reports '↓10.41%'; at 10% scale PSNR increases from 25.23 to 25.72 (about +1.9%), not '↑10.67%'; at 100% scale PSNR increases from 25.92 to 26.66 (about +2.9%), not '↑15.67%'. The reported percentages appear to have been copied from the LPIPS column or computed by an unstated formula. Because this table is the primary evidence for the claim that LPR leverages AODLibpro scalability, the numbers must be recomputed and corrected.
  3. [Section 3.2, Section 4.4, Figure 5] The claim that AODLibpro 'uniformly covers' optical-degradation space is validated only through the same OIQ/OD-Class proxy used to construct the library. Figure 5(a) is a histogram over OD-Class after stratified sampling, so near-uniformity is expected by construction, and Figure 5(b) visualizes coverage with per-FoV/wavelength OIQ, i.e., the same metric. This is partly tautological. The authors should provide independent evidence of diversity in PSF/degradation space (e.g., PSF shape statistics, Zernike coefficients, or generalization across a larger set of held-out real lens designs) or explicitly qualify the coverage claim. RealLens-Sim is a useful independent check but contains only 10 lenses and cannot by itself establish the scalability conclusion.
  4. [Appendix D.2, Figure 7] The OD severity partition into Strong/Medium/Mild classes is shown only as a schematic figure with no numerical thresholds, and the OD-Class uniformity threshold is stated as 'α = 0.85 empirically.' Since the entire hybrid sampling basis depends on these values, the exact intervals and a sensitivity analysis (or at least a justification) for α are needed for reproducibility and for assessing how robust the uniformity claim is to the cutoff choices.
minor comments (4)
  1. [Title/Abstract vs. main text] The arXiv title ('...Discrete Degradation Priors') differs from the full-text title ('OmniLens++: Blind Lens Aberration Correction via Large LensLib Pre-Training and Latent PSF Representation'), and the arXiv abstract introduces the model as FoundCAC while the body consistently uses OmniLens++. This version mismatch should be resolved before publication.
  2. [Tables 2-3 and 12] All quantitative results are single-run averages without standard deviations, confidence intervals, or significance tests. Several claimed improvements over the next best method are below 0.3 dB, so reporting multiple seeds or per-lens variability would materially strengthen the SOTA claims.
  3. [Appendix G.5, Table 12] The text states that OmniLens++ 'performs better overall' on RealLens-Snap, but Table 12 is mixed: e.g., on Single-Lens-I the Universal IR model has higher CLIPIQA and MANIQA, and on several rows NIQE favors another method. An aggregate or paired comparison would be more convincing than per-lens three-metric lists.
  4. [Appendix D.2, Figure 8] The OD-Class table has rendering artifacts ('US>?', '???_???(???)') that obscure the classification criteria, and Table 8 contains a typo ('Sclae'). Please clean up the figure and text.

Circularity Check

1 steps flagged

Uniformity evidence for AODLibpro is self-definitional, but the central zero-shot correction claims rest on independent external evaluations and scaling experiments, so circularity is partial.

specific steps
  1. self definitional [§3.2 (Construction of AODLibpro) and §4.4 (Evaluation of AODLibpro) / Figure 5(a)-(b)]
    "We form 18 OD sub-classes by crossing the average OIQ category with OD-Class. From each subclass, m 1 and m 2 instances are sampled to build AODLibprotrain and test, respectively. ... the samples in AODLibpro are uniformly distributed across degradation severity and spatial variation patterns (Figure 5 (a))"

    The dataset is created by sampling a fixed number of instances from each of the 18 OIQ×OD-Class sub-classes, so uniformity over those classes is a construction constraint. Figure 5(a) then plots a histogram of the same degradation classes, making the 'uniform distribution' a restatement of the sampling key rather than independent evidence; Figure 5(b) is likewise a coverage plot in the OIQ space used to define the basis. The display therefore cannot demonstrate diversity in PSF structure that OIQ/OD-Class might miss. This does not affect the LPR supervision chain or the RealLens-Sim/Snap results, which are separate empirical evidence, but it makes the AODLibpro 'uniform coverage' validation partially circular.

full rationale

The paper's central blind-CAC claim is not a fit renamed as a prediction. LPR is learned from ground-truth PSF maps through PSF-VQVAE reconstruction and ODN forward-model consistency, and at Stage II the predicted latent PSF features are supervised by the frozen representation and by the forward degradation loss; this is internal consistency, not circularity. Self-citations (OmniLens, EAOD, OIQE) are used as baselines, data-generation components, or metric definitions rather than as unverified justifications; no uniqueness theorem is imported. The main circular element is confined to the AODLibpro coverage validation: AODLibpro is explicitly balanced over the 18 OIQ×OD-Class sub-classes, and Figure 5(a)-(b) demonstrates uniformity/coverage in exactly that same proxy space. That evidence is self-definitional and cannot independently exclude missing PSF diversity. However, the scalability claim is supported by the empirical scaling curve in Figure 5(c), and the SOTA claim is supported by independent RealLens-Sim numerical results and RealLens-Snap qualitative/quantitative results with externally sourced lens designs (Tseng et al., 2021; Eboli et al., 2022). The paper also acknowledges the same-source limitation of AODLibproTest in Appendix F.2, which is a benchmark domain-gap concern rather than a circular derivation. Overall, one partial self-definitional validation in the data section; the model and generalization claims retain independent content. Score 4.

Axiom & Free-Parameter Ledger

9 free parameters · 7 axioms · 3 invented entities

The framework's central claims rest on the realism of the simulated lens library and the proxy metrics used to balance it. Free parameters are mostly hand-set sampling/classification thresholds and loss weights; none is a deep physical constant. The main assumptions are that EAOD plus DeepLens simulations faithfully represent real optical degradation, and that OIQ/OD-Class capture the relevant variation. LPR and OD-Class are new constructs with no external falsifiable handle beyond the paper's internal benchmarks.

free parameters (9)
  • OIQ metric weights lambda_1, lambda_2, lambda_3 = 0.4, 0.3, 0.3
    Eq. 11; hand-set to normalize PSNR, SSIM, and SFR contributions in the degradation-severity score used for sampling.
  • Spatial uniformity threshold alpha = 0.85
    Figure 8; empirical cutoff separating 'spatial-uniform' OD-Class from trend-based classes; directly shapes the hybrid sampling basis.
  • Image-distance perturbation probability gamma = 25%
    Appendix D.1; empirical probability of perturbing image distance within depth of field to create additional lens variants.
  • Permissible circle of confusion delta = 24 um
    Appendix D.1; chosen for DoF-constrained image-distance perturbation.
  • OD severity class boundaries = Unspecified intervals in Figure 7
    Average OIQ range split into three classes; boundary values are hand-divided and not explicitly given.
  • VQ loss weight beta = 0.25
    Eq. 3; 'common practice' hyperparameter, not fitted to data.
  • Codebook size K = 1024
    Appendix E.1 and G.3; swept over 512/2048 with no meaningful gain, so 1024 was chosen.
  • Sampling counts m1/m2 = 200 / 3 per class
    Section 4.1; data-scale choices yielding 3,600 training and 54 test lenses; affect scalability conclusions.
  • DoF near/far limits Delta_L1, Delta_L2 = From Eq. 10
    Constrains image-distance perturbation; requires focal length, F-number, L, and delta, with delta hand-set.
axioms (7)
  • domain assumption Optical degradation is faithfully modeled by patch-wise convolution with ray-traced PSFs plus ISP and noise (DeepLens simulator).
    Used in Section 3.2 and F.1 to generate all training, benchmark, and RealLens-Sim pairs; if the simulator is unrealistic, the LensLib-PT premise fails.
  • domain assumption EAOD-generated lens designs, expanded with aspheric surfaces and image-plane perturbations, cover the real-world lens/aberration space relevant to blind CAC.
    Section 3.2 and Figure 2; the universality claim depends on the lens source distribution matching real-world cases.
  • domain assumption The OIQ/OD-Class hybrid sampling basis is a faithful proxy for degradation severity and spatial-variation patterns.
    Section 3.2, Eq. 11, Figure 8; the uniform-distribution and scalability results depend on these hand-designed metrics.
  • domain assumption A VQ-VAE codebook of K=1024 entries can represent the continuous PSF feature manifold well enough for restoration guidance.
    Section 3.3, E.1, G.3; if quantization loses discriminative PSF structure, LPR guidance degrades.
  • domain assumption The ODN degradation process (clear image plus quantized PSF features to degraded image, Eq. 4) is a sufficient forward model to regularize PSF-prior learning.
    Section 3.3 and E.2; LPR's optical-prior quality depends on this learned forward operator being faithful.
  • standard math Standard deep-network training assumptions: L1/perceptual/VQ losses are differentiable proxies and optimize restoration quality.
    Used throughout; no formal convergence guarantee, standard empirical practice in deep image restoration.
  • domain assumption Fidelity/perceptual metrics (PSNR, SSIM, LPIPS, no-reference metrics) measure correction quality.
    Used in all tables; conclusions about 'state-of-the-art' rest on these proxies.
invented entities (3)
  • Latent PSF Representation (LPR) no independent evidence
    purpose: Discrete codebook plus ODN that encodes PSF-derived optical priors to guide blind aberration correction.
    New learned construct; its usefulness is supported only by internal benchmarks and visualizations, with no external falsifiable prediction.
  • OD-Class taxonomy (six spatial-variation classes) no independent evidence
    purpose: Categorizes degradation patterns for balanced library sampling.
    Hand-defined classification in Figure 8; no independent validation that these classes are physically exhaustive or sufficient.
  • OIQ index no independent evidence
    purpose: Composite optical-image-quality severity metric for sampling.
    Hand-weighted combination of PSNR, SSIM, and SFR in Eq. 11; no external calibration against perceptual or optical truth.

pith-pipeline@v1.3.0-alltime-deepseek · 26185 in / 14352 out tokens · 127654 ms · 2026-08-03T20:58:09.798939+00:00 · methodology

0 comments
read the original abstract

Emerging deep-learning-based lens library pre-training (LensLib-PT) pipeline offers a new avenue for blind lens aberration correction by training a universal neural network, demonstrating strong capability in handling diverse unknown optical degradations. This work proposes FoundCAC, a universal foundational framework that resolves two challenges hindering the generalization of existing pipelines: the difficulty of scaling training data and the absence of prior guidance characterizing optical degradation. To improve data scalability, we expand the design specifications to increase degradation diversity and construct AODLibpro, a large-scale lens library using stratified sampling over spatial-variation patterns and degradation severity. In terms of model design, to leverage Point Spread Functions (PSFs) as guidance while maintaining the blind paradigm, we propose a multi-stage vector-quantized representation learning scheme. This paradigm is specifically designed to construct a Latent PSF Representation (LPR), explicitly encoding complex continuous PSFs into a discrete degradation prior to regularize the highly ill-posed restoration process. Through a simple yet effective codebook-freezing strategy, our framework leverages the discrete prior to elevate full-shot restoration performance and unlock highly efficient few-shot adaptation for unseen lenses. Experiments on synthetic LensLib, real-design simulations, and real-captured lenses show that our framework achieves state-of-the-art zero-shot performance under complementary evaluation protocols, while enabling highly efficient few-shot adaptation for specific lenses. The source code and datasets will be made publicly available at https://github.com/zju-jiangqi/FoundCAC.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

3 extracted references · 1 linked inside Pith

  1. [2020]

    Multimodal prompt per- ceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration

    11 Yuang Ai, Huaibo Huang, Xiaoqiang Zhou, Jiexiang Wang, and Ran He. Multimodal prompt per- ceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration. In CVPR, 2024. 3 Chaofeng Chen, Xinyu Shi, Yipeng Qin, Xiaoming Li, Xiaoguang Han, Tao Yang, and Shihui Guo. Real-world blind super-resolution via feature matching with im...

  2. [2023]

    completely blind

    11 12 Preprint. Jun Luo, Yunfeng Nie, Wenqi Ren, Xiaochun Cao, and Ming-Hsuan Yang. Correcting optical aberra- tion via depth-aware point spread functions.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. 3 Jiaqi Ma, Tianheng Cheng, Guoli Wang, Qian Zhang, Xinggang Wang, and Lefei Zhang. ProRes: Exploring degradation-aware visual promp...

  3. [2024]

    Rethinking coarse- to-fine approach in single image deblurring

    9 Sung-Jin Cho, Seo-Won Ji, Jun-Pyo Hong, Seung-Won Jung, and Sung-Jea Ko. Rethinking coarse- to-fine approach in single image deblurring. InICCV, 2021. 9 Geoffroi Cˆot´e, Jean-Franc ¸ois Lalonde, and Simon Thibault. Deep learning-enabled framework for automatic lens design starting point generation.Optics Express, 2021. 8, 22 Thomas Eboli, Jean-Michel Mo...