REVIEW 4 major objections 4 minor 3 references
A single blind network can correct severe, spatially varying lens aberrations across minimalist optics, metalenses, misaligned lenses, and high-end DSLRs without per-lens PSF calibration.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 20:58 UTC pith:OHH45B2G
load-bearing objection Solid incremental advance in blind aberration correction, with a useful data library and PSF-latent guidance, but the uniformity claim is partly self-confirming and reproducibility/framing issues need attention. the 4 major comments →
Towards Blind Lens Aberration Correction via Large LensLib Pre-training and Discrete Degradation Priors
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper establishes that the bottleneck in blind aberration correction is not model capacity alone but two design choices: the distribution of the lens library and the form of degradation guidance. It constructs AODLibpro by enriching the lens-source specifications and sampling uniformly over severity (a weighted PSNR/SSIM/SFR composite) and spatial-variation class (six patterns defined from per-FoV quality trends). It then learns LPR, a discrete codebook of latent PSF features via vector-quantized autoencoding, supervised by an Optical Degradation Network that must use the code to degrade a clear image into the observed one. During correction, a UNet-style model predicts and retrieves cod
What carries the argument
Two coupled objects carry the argument. AODLibpro is the data engine: a lens library generated with enriched optical design specifications and sampled with a hybrid basis crossing three severity classes (OIQ) with six spatial-variation classes (OD-Class), giving eighteen subclasses and a deliberately uniform aberration distribution. LPR is the guidance engine: a vector-quantized autoencoder turns PSF maps into a discrete codebook of latent PSF features, while an Optical Degradation Network—conditioning a clear-image encoder on quantized PSF codes to reproduce the degraded image—regularizes the codes to carry optical meaning. At inference a separate encoder predicts latent features from the d
Load-bearing premise
The load-bearing premise is that the hand-weighted PSNR/SSIM/SFR composite (OIQ) and the six-category spatial-pattern taxonomy (OD-Class), with its empirically chosen threshold, faithfully capture the optical-degradation structure that matters—if real PSF variations such as chromatic structure or non-monotonic field patterns slip through both, the library is not truly uniform and the scalability conclusions weaken.
What would settle it
Cluster AODLibpro's 3,600 training lenses by their full PSF maps rather than by OIQ summaries: if some of the eighteen OD subclasses contain near-duplicate PSF kernels while sizable PSF-shape clusters fall outside them, the hybrid sampling basis has not balanced the true degradation space and the reported scalability gain should not transfer to the missing patterns. A single counterexample lens—classified as 'spatial-uniform' but with PSFs that alternate between sagittal- and tangential-dominated lobes across the field—would directly expose the taxonomy's insufficiency.
If this is right
- A single blind model can correct severe aberrations from minimalist spherical/aspheric lenses and metalenses, handle stochastic misalignment blur, and improve high-end DSLR images without per-lens PSF calibration at inference.
- Scaling a lens library pays off only when samples are uniform in degradation severity and spatial-variation pattern; the paper finds larger improvements from 1% to 100% data with the new library than with the previous RMS-sampled one.
- LPR guidance beats direct PSF prediction, PSF-feature prediction, and variants with only the vector-quantized codebook or only the degradation network, and its benefit grows as the library scales.
- The frozen discrete prior makes few-shot adaptation to a new lens efficient, since the codebook transfers instead of being retrained per lens.
- Predicted latent PSF features are discriminative across degradation patterns—the attention maps align with per-field degradation—supporting the claim that the guidance is optical rather than categorical.
Where Pith is reading between the lines
- The LPR codebook could be reused as a blind descriptor of optical behavior, not just a correction guide: since its entries track per-field degradation patterns, clustering unknown lenses by their retrieved codes would partition optics by effect rather than by design type.
- If the approach transfers, optical designers could ship cheaper, aberration-heavy optics and let a universal correction model absorb the quality loss—a different cost/quality trade-off for minimalist and metalens systems than the paper spells out.
- A direct stress test would apply the same frozen codebook to PSF structures outside the generated lens-source space (e.g., metasurfaces or diffractive elements), which the paper names as future data extensions but does not claim to cover.
- The OIQ weights are hand-set; recomputing AODLibpro with different weights or a learned perceptual metric would probe how much of the scalability result rests on the specific proxy.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a blind lens aberration correction framework (called OmniLens++ in the main text and FoundCAC in the arXiv metadata abstract) with two main contributions. On the data side, it constructs AODLibpro by extending an EAOD-based lens-source generator with aspheric surfaces and image-distance perturbation, then samples lenses uniformly over a hybrid basis formed by an OIQ severity score (Eq. 11) and an OD-Class taxonomy of spatial-variation patterns, yielding 18 subclasses with 200 train and 3 test lenses per subclass. On the model side, it introduces LPR, a VQVAE-based discrete representation of PSF maps regularized by an ODN that simulates the optical degradation process; the learned codebook is frozen and used to guide a UNet/Swin restoration backbone via predicted latent PSF features. Experiments include a new AODLibproTest benchmark, a RealLens-Sim suite covering minimalist spherical/aspheric lenses, a metalens, misaligned smartphone lenses, and high-end lenses, plus real-snapped qualitative and non-reference quantitative evaluations. The paper reports state-of-the-art zero-shot results, with AODLibpro and LPR each contributing average PSNR gains over the OmniLens baseline on RealLens-Sim.
Significance. If the reported results hold, the paper makes a meaningful advance in a relatively underexplored area: it demonstrates that a large automatically designed lens library can be made more scalable and more uniform, and that PSF-derived priors can be injected in a fully blind manner through a learned discrete latent representation. The evaluation is broader than in prior work, spanning simulated, synthetic-benchmark, and real-captured data, and the appendix gives substantial implementation detail. The paper is also careful to ablate data-specification choices, sampling bases, representation components, and codebook size, and promises open release of code, lens design files, PSF arrays, and simulated/real images. However, the central scalability claim relies on a proxy that is partly used both to construct and to validate the library; several quantitative claims in the paper have internal inconsistencies; and the arXiv abstract claims a few-shot adaptation capability that is not evaluated in the body. These issues do not invalidate the overall framework, but they do require correction and additional validation before the paper can be accepted.
major comments (4)
- [Abstract / Sections 4-5] The arXiv metadata abstract states that the framework 'unlock[s] highly efficient few-shot adaptation for unseen lenses.' The body contains no few-shot experiments: Section 4 reports only zero-shot evaluations, and Section 5 explicitly lists 'a flexible fine-tuning pipeline is urged' as future work. This is a load-bearing capability claim with no supporting evidence. Either add a few-shot adaptation experiment (e.g., fine-tuning on a small number of real-lens pairs) or remove the claim from the abstract.
- [Section 4.4, Table 8] The percentage improvements in Table 8 are internally inconsistent. For LPR at 1% scale, PSNR decreases from 23.82 to 23.39 (about -1.8%), yet the table reports '↓10.41%'; at 10% scale PSNR increases from 25.23 to 25.72 (about +1.9%), not '↑10.67%'; at 100% scale PSNR increases from 25.92 to 26.66 (about +2.9%), not '↑15.67%'. The reported percentages appear to have been copied from the LPIPS column or computed by an unstated formula. Because this table is the primary evidence for the claim that LPR leverages AODLibpro scalability, the numbers must be recomputed and corrected.
- [Section 3.2, Section 4.4, Figure 5] The claim that AODLibpro 'uniformly covers' optical-degradation space is validated only through the same OIQ/OD-Class proxy used to construct the library. Figure 5(a) is a histogram over OD-Class after stratified sampling, so near-uniformity is expected by construction, and Figure 5(b) visualizes coverage with per-FoV/wavelength OIQ, i.e., the same metric. This is partly tautological. The authors should provide independent evidence of diversity in PSF/degradation space (e.g., PSF shape statistics, Zernike coefficients, or generalization across a larger set of held-out real lens designs) or explicitly qualify the coverage claim. RealLens-Sim is a useful independent check but contains only 10 lenses and cannot by itself establish the scalability conclusion.
- [Appendix D.2, Figure 7] The OD severity partition into Strong/Medium/Mild classes is shown only as a schematic figure with no numerical thresholds, and the OD-Class uniformity threshold is stated as 'α = 0.85 empirically.' Since the entire hybrid sampling basis depends on these values, the exact intervals and a sensitivity analysis (or at least a justification) for α are needed for reproducibility and for assessing how robust the uniformity claim is to the cutoff choices.
minor comments (4)
- [Title/Abstract vs. main text] The arXiv title ('...Discrete Degradation Priors') differs from the full-text title ('OmniLens++: Blind Lens Aberration Correction via Large LensLib Pre-Training and Latent PSF Representation'), and the arXiv abstract introduces the model as FoundCAC while the body consistently uses OmniLens++. This version mismatch should be resolved before publication.
- [Tables 2-3 and 12] All quantitative results are single-run averages without standard deviations, confidence intervals, or significance tests. Several claimed improvements over the next best method are below 0.3 dB, so reporting multiple seeds or per-lens variability would materially strengthen the SOTA claims.
- [Appendix G.5, Table 12] The text states that OmniLens++ 'performs better overall' on RealLens-Snap, but Table 12 is mixed: e.g., on Single-Lens-I the Universal IR model has higher CLIPIQA and MANIQA, and on several rows NIQE favors another method. An aggregate or paired comparison would be more convincing than per-lens three-metric lists.
- [Appendix D.2, Figure 8] The OD-Class table has rendering artifacts ('US>?', '???_???(???)') that obscure the classification criteria, and Table 8 contains a typo ('Sclae'). Please clean up the figure and text.
Circularity Check
Uniformity evidence for AODLibpro is self-definitional, but the central zero-shot correction claims rest on independent external evaluations and scaling experiments, so circularity is partial.
specific steps
-
self definitional
[§3.2 (Construction of AODLibpro) and §4.4 (Evaluation of AODLibpro) / Figure 5(a)-(b)]
"We form 18 OD sub-classes by crossing the average OIQ category with OD-Class. From each subclass, m 1 and m 2 instances are sampled to build AODLibprotrain and test, respectively. ... the samples in AODLibpro are uniformly distributed across degradation severity and spatial variation patterns (Figure 5 (a))"
The dataset is created by sampling a fixed number of instances from each of the 18 OIQ×OD-Class sub-classes, so uniformity over those classes is a construction constraint. Figure 5(a) then plots a histogram of the same degradation classes, making the 'uniform distribution' a restatement of the sampling key rather than independent evidence; Figure 5(b) is likewise a coverage plot in the OIQ space used to define the basis. The display therefore cannot demonstrate diversity in PSF structure that OIQ/OD-Class might miss. This does not affect the LPR supervision chain or the RealLens-Sim/Snap results, which are separate empirical evidence, but it makes the AODLibpro 'uniform coverage' validation partially circular.
full rationale
The paper's central blind-CAC claim is not a fit renamed as a prediction. LPR is learned from ground-truth PSF maps through PSF-VQVAE reconstruction and ODN forward-model consistency, and at Stage II the predicted latent PSF features are supervised by the frozen representation and by the forward degradation loss; this is internal consistency, not circularity. Self-citations (OmniLens, EAOD, OIQE) are used as baselines, data-generation components, or metric definitions rather than as unverified justifications; no uniqueness theorem is imported. The main circular element is confined to the AODLibpro coverage validation: AODLibpro is explicitly balanced over the 18 OIQ×OD-Class sub-classes, and Figure 5(a)-(b) demonstrates uniformity/coverage in exactly that same proxy space. That evidence is self-definitional and cannot independently exclude missing PSF diversity. However, the scalability claim is supported by the empirical scaling curve in Figure 5(c), and the SOTA claim is supported by independent RealLens-Sim numerical results and RealLens-Snap qualitative/quantitative results with externally sourced lens designs (Tseng et al., 2021; Eboli et al., 2022). The paper also acknowledges the same-source limitation of AODLibproTest in Appendix F.2, which is a benchmark domain-gap concern rather than a circular derivation. Overall, one partial self-definitional validation in the data section; the model and generalization claims retain independent content. Score 4.
Axiom & Free-Parameter Ledger
free parameters (9)
- OIQ metric weights lambda_1, lambda_2, lambda_3 =
0.4, 0.3, 0.3
- Spatial uniformity threshold alpha =
0.85
- Image-distance perturbation probability gamma =
25%
- Permissible circle of confusion delta =
24 um
- OD severity class boundaries =
Unspecified intervals in Figure 7
- VQ loss weight beta =
0.25
- Codebook size K =
1024
- Sampling counts m1/m2 =
200 / 3 per class
- DoF near/far limits Delta_L1, Delta_L2 =
From Eq. 10
axioms (7)
- domain assumption Optical degradation is faithfully modeled by patch-wise convolution with ray-traced PSFs plus ISP and noise (DeepLens simulator).
- domain assumption EAOD-generated lens designs, expanded with aspheric surfaces and image-plane perturbations, cover the real-world lens/aberration space relevant to blind CAC.
- domain assumption The OIQ/OD-Class hybrid sampling basis is a faithful proxy for degradation severity and spatial-variation patterns.
- domain assumption A VQ-VAE codebook of K=1024 entries can represent the continuous PSF feature manifold well enough for restoration guidance.
- domain assumption The ODN degradation process (clear image plus quantized PSF features to degraded image, Eq. 4) is a sufficient forward model to regularize PSF-prior learning.
- standard math Standard deep-network training assumptions: L1/perceptual/VQ losses are differentiable proxies and optimize restoration quality.
- domain assumption Fidelity/perceptual metrics (PSNR, SSIM, LPIPS, no-reference metrics) measure correction quality.
invented entities (3)
-
Latent PSF Representation (LPR)
no independent evidence
-
OD-Class taxonomy (six spatial-variation classes)
no independent evidence
-
OIQ index
no independent evidence
read the original abstract
Emerging deep-learning-based lens library pre-training (LensLib-PT) pipeline offers a new avenue for blind lens aberration correction by training a universal neural network, demonstrating strong capability in handling diverse unknown optical degradations. This work proposes FoundCAC, a universal foundational framework that resolves two challenges hindering the generalization of existing pipelines: the difficulty of scaling training data and the absence of prior guidance characterizing optical degradation. To improve data scalability, we expand the design specifications to increase degradation diversity and construct AODLibpro, a large-scale lens library using stratified sampling over spatial-variation patterns and degradation severity. In terms of model design, to leverage Point Spread Functions (PSFs) as guidance while maintaining the blind paradigm, we propose a multi-stage vector-quantized representation learning scheme. This paradigm is specifically designed to construct a Latent PSF Representation (LPR), explicitly encoding complex continuous PSFs into a discrete degradation prior to regularize the highly ill-posed restoration process. Through a simple yet effective codebook-freezing strategy, our framework leverages the discrete prior to elevate full-shot restoration performance and unlock highly efficient few-shot adaptation for unseen lenses. Experiments on synthetic LensLib, real-design simulations, and real-captured lenses show that our framework achieves state-of-the-art zero-shot performance under complementary evaluation protocols, while enabling highly efficient few-shot adaptation for specific lenses. The source code and datasets will be made publicly available at https://github.com/zju-jiangqi/FoundCAC.
Reference graph
Works this paper leans on
-
[2020]
Multimodal prompt per- ceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration
11 Yuang Ai, Huaibo Huang, Xiaoqiang Zhou, Jiexiang Wang, and Ran He. Multimodal prompt per- ceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration. In CVPR, 2024. 3 Chaofeng Chen, Xinyu Shi, Yipeng Qin, Xiaoming Li, Xiaoguang Han, Tao Yang, and Shihui Guo. Real-world blind super-resolution via feature matching with im...
2024
-
[2023]
11 12 Preprint. Jun Luo, Yunfeng Nie, Wenqi Ren, Xiaochun Cao, and Ming-Hsuan Yang. Correcting optical aberra- tion via depth-aware point spread functions.IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. 3 Jiaqi Ma, Tianheng Cheng, Guoli Wang, Qian Zhang, Xinggang Wang, and Lefei Zhang. ProRes: Exploring degradation-aware visual promp...
Pith/arXiv arXiv 2024
-
[2024]
Rethinking coarse- to-fine approach in single image deblurring
9 Sung-Jin Cho, Seo-Won Ji, Jun-Pyo Hong, Seung-Won Jung, and Sung-Jea Ko. Rethinking coarse- to-fine approach in single image deblurring. InICCV, 2021. 9 Geoffroi Cˆot´e, Jean-Franc ¸ois Lalonde, and Simon Thibault. Deep learning-enabled framework for automatic lens design starting point generation.Optics Express, 2021. 8, 22 Thomas Eboli, Jean-Michel Mo...
arXiv 2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.