{"id":"1c3932a3-0f86-4303-8e6a-5046c3f91890","arxiv_id":"2412.12709","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"VariLens, a physics-informed variational autoencoder, detects lensed quasars and estimates SIE lens parameters in milliseconds, yielding 42 new candidates from HSC data.","lead":"A new machine learning tool called VariLens uses a physics-informed variational autoencoder to spot strongly lensed quasars in survey images and estimate their lens properties in milliseconds. The authors apply it to about 710,000 preselected sources from the Hyper Suprime-Cam survey, producing 42 new lens candidates that await confirmation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 2σ consistency with traditional PyAutoLens modeling is not an independent validation: both methods assume SIE+γext, and the paper concedes that double-image fits are degenerate, so the millisecond parameter-estimation claim lacks a secure accuracy check.","rationale":"The paper is careful and self-critical; it reports low purity, poor shear recovery, and unconfirmed candidates. The sim-to-real gap is real, but the paper partially addresses it with 22 known lenses for classification and 20 for regression. The weakest part of the central claim, that key parameters can be estimated in milliseconds consistently with traditional modeling, is that the external benchmark is not independent: PyAutoLens is run under the same SIE+γext assumption used to generate training data, and the paper explicitly concedes degeneracy for doubles. This makes the 2σ consistency a weak test. It is the single most load-bearing issue because it affects the headline quantitative claim; if the benchmark is degenerate, the agreement could be vacuous. A concrete re-fit of the same systems with relaxed assumptions and full posterior sampling would settle it. This does not overturn the paper's contribution as a fast candidate sorter, but it should keep the verdict at CONDITIONAL: the discovery pipeline is promising while the modeling accuracy claim is not yet independently anchored.","tokens_in":33854,"tokens_out":3720,"duration_ms":38055,"concrete_test":"Take the 20 GLQD systems used in Fig. 12 and re-fit them with PyAutoLens without fixing the mass ellipticity and center to the light profile, using full MCMC or nested sampling to obtain posteriors; then compare VariLens θ_E, center, and ellipticity to those posteriors. If the traditional-model uncertainties on θ_E grow substantially (e.g., more than 2x) or become multi-modal for the doubles, the reported 2σ agreement does not establish accuracy. Where HST or AO imaging is available, use independent lens models as an additional cross-check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central parameter-estimation claim is validated mainly by comparison with PyAutoLens modeling of 20 GLQD lenses (Sect. 4.5.2, Fig. 12). This is not an independent accuracy test. VariLens is trained on SIE+γext simulations (Sect. 2.4), and the PyAutoLens benchmark uses the same SIE+γext model, with mass ellipticity and center fixed to the light profile and Sérsic index fixed to 4. More importantly, the paper states that for doubly imaged quasars the number of point-source constraints is insufficient to determine SIE+γext parameters, producing fitting degeneracy. If the traditional model's posterior is broad or multi-modal because of that degeneracy, consistency within 2σ can hold even when VariLens is biased; the two methods share the same underlying model and can be jointly wrong. Since only 20 objects are used, the headline millisecond θ_E and ellipticity claim is not tied to ground truth. The zqso comparison in Fig. 12 (R2=0.16, MAE=0.95) also shows that the external validation is not uniformly good, so parameter-specific scrutiny matters.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents VariLens, a physics-informed variational autoencoder that jointly performs image reconstruction, lens/non-lens classification, and SIE+γext parameter estimation for lensed quasars in HSC imaging. The network is trained on 137,552 simulated lensed quasar images built from real HSC galaxy cutouts and SIMQSO quasar spectra. On a held-out mock test set, the model achieves high R² values for the Einstein radius, lens centers, ellipticities, and source positions, but fails to recover external shear. The paper then compares the network's parameter estimates with PyAutoLens models for 20 known GLQD lenses and reports consistency within 2σ for θ_E<3\". Finally, applying a multiwavelength preselection (80 million to 710,966 sources) and the VariLens classifier yields 13,831 candidates, of which 42 are visually graded A or B.","tokens_in":34093,"tokens_out":4463,"duration_ms":44137,"significance":"If the accuracy claims hold, VariLens would be a valuable tool for rapid triage and preliminary modeling of lensed quasars in upcoming surveys such as LSST and Euclid, where traditional modeling is computationally expensive. The paper's strengths include the use of realistic simulations on actual HSC galaxy images, the explicit public release of candidate tables on Zenodo, and the candid discussion of known failure modes (external shear, source redshift, classifier purity). The integrated classification-plus-regression architecture is a useful engineering contribution. However, the current validation does not independently establish the headline accuracy, because the external comparison uses the same SIE+γext model shared by the training simulations and the traditional fitting, and the mock test set comes from the same simulator as the training data.","major_comments":[{"comment":"The 2σ consistency between VariLens and PyAutoLens is not an independent accuracy check, because both methods assume the same SIE+γext mass model and the paper itself states that for doubly imaged quasars the point-source constraints are insufficient, producing fitting degeneracy. Agreement within 2σ can therefore hold even if both models are jointly biased. Please validate against independently measured quantities (e.g., stellar velocity dispersions, time-delay models, or spectroscopic source redshifts) or explicitly reframe the comparison as model consistency rather than accuracy.","section":"Section 4.5.2, Fig. 12"},{"comment":"The uncertainty calibration is performed on the same test set used to report the mock R² values and the 2σ consistency in Fig. 12. Because the scaling factors are derived from the empirical errors on that set, the calibrated uncertainties are not an out-of-sample test of calibration; this circularity should be acknowledged and, ideally, the calibration should be checked on an independent validation set or through recalibration on real data.","section":"Section 3.3, Eq. (17)"},{"comment":"The transfer from simulations to real data is incomplete for several headline parameters: external shear recovery is essentially absent (R²≈0.03–0.07), source redshift is systematically overestimated with R²=0.16 and MAE=0.95, and ellipticity R² drops to 0.26–0.47 on real data. The abstract and conclusions should state these performance limits explicitly rather than describing the key parameters as reliably determined; at minimum, the 'key parameters' claim should be scoped to θ_E, lens centers, and (with caveats) ellipticity.","section":"Sections 4.4 and 4.5.2, Figs. 10 and 12"},{"comment":"The classifier's real-data performance (16 of 22 known lenses recovered, 73% completeness, ~1% purity) is substantially below the near-perfect mock AUROC of 0.998. The discovery claims should present these real-data rates in the abstract and conclusions alongside the 42 candidates, and the paper should discuss whether the 42 visually selected candidates are consistent with the expected purity and with the 68% recovery rate of the catalog-level preselection.","section":"Section 4.5.1, Fig. 11"}],"minor_comments":[{"comment":"The text reports 22 GLQD systems with HSC images for the classifier evaluation, but the regression comparison in Section 4.5.2 uses 20 systems; please clarify how the two systems are excluded from the regression analysis.","section":"Sections 4.5.1 and 4.5.2"},{"comment":"The piecewise definition of the external shear orientation uses the notation 's/2' without clear parentheses; since s is defined in Eq. (9) via an arcsine, the expression should be rewritten with explicit brackets to avoid ambiguity.","section":"Section 2.4, Eq. (8)"},{"comment":"There is a typo 'is is typically approximated', and the approximation for b_n should be described as an asymptotic or approximate expression rather than an equality over the full range of n.","section":"Appendix C, Eq. (C.6)"},{"comment":"The legend distinguishes doubles from quadruples, but the text does not discuss whether the agreement differs between these subclasses; adding one sentence would be useful because the degeneracy argument is specific to doubly imaged systems.","section":"Figure 12"},{"comment":"The deflector sample is dominated by SDSS LRGs at z≲1, a bias that is acknowledged later in the paper; I recommend stating this bias when the training set is first introduced, since it directly affects the claimed transferability of the model.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"This is a useful engineering contribution with a realistic simulation pipeline and a publicly available candidate list. The central weakness is that the validation strategy does not currently support the strongest accuracy claims: the PyAutoLens comparison is model consistency rather than ground-truth validation, and the uncertainty calibration is based on the same test set. The paper is not fatally flawed, but the claims need to be scoped and the validation strengthened or reframed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: VariLens is a real integration — VAE reconstruction, SIE+γext regression, and a transfer-learned classifier in one network — and the mock-data regression for θ_E, centers, ellipticities, and z_qso is solid (R² > 0.8). But the real-data parameter validation is much weaker than the abstract implies: external shear is essentially unconstrained, z_qso is biased, and the PyAutoLens comparison is not an independent accuracy test since both sides assume the same SIE+γext model. The detection pipeline is a reasonable incremental step, not a breakthrough.\n\nWhat is genuinely new: the single-network design that jointly reconstructs the image, classifies lens vs. non-lens, and regresses 11 physical parameters. That is not in the cited prior work, and the millisecond inference claim on a CPU is credible given the architecture and training time. The paper is also honest about its failures — shear prediction, low-z quasar redshifts, low purity — and it ships the full candidate tables and modeling outputs on Zenodo. That counts for something.\n\nThe soft spots are real but mostly acknowledged in the text. The training and test sets come from the same SIE+γext simulator, so the high R² values partly measure how well the network inverts its own training distribution. The uncertainty calibration is a post-hoc rescaling applied to the test set, not a predictive calibration. The external validation uses 20 GLQD objects, and the PyAutoLens benchmark fixes Sérsic index, ties mass ellipticity to light, and the paper itself concedes that double-image fits are degenerate. So the \"consistent within 2σ\" statement is a weak check, not a strong one. The classifier completeness on known lenses is 73%, which is decent, but purity is about 1%, so the 42 grade A/B candidates are just candidates awaiting spectroscopy.\n\nOne thing I would push on in review: the paper compares VariLens to PyAutoLens but never to their own earlier CNN/ViT detectors (Andika et al. 2023a,b). A quantitative comparison to prior classifiers, even on the same mock test set, would clarify what the VAE actually adds. Also, no code or weights are released, which limits reproducibility for a method paper.\n\nWho this is for: strong-lensing folk preparing for LSST/Euclid who want a fast triage tool. It deserves a serious referee — the architecture is novel enough and the mock-data results are strong enough to warrant a careful look — but the authors should be pushed to temper the parameter-estimation claims, add a comparison to their earlier detectors, and ideally release the trained model.","headline":"A genuinely integrated VAE-based pipeline for lensed quasar detection and SIE parameter estimation, with solid mock-data results but weaker and partly circular real-data validation; worth refereeing with revisions.","tokens_in":34716,"tokens_out":1436,"would_cite":true,"duration_ms":15981,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper presents VariLens, a physics-informed variational autoencoder that detects lensed quasars and estimates their singular-isothermal-ellipsoid mass parameters in a single forward pass, and reports 42 candidate systems from Hyper…","keywords":["gravitational lensing","lensed quasars","variational autoencoder","deep learning","Hyper Suprime-Cam","Einstein radius","SIE mass model","strong lens detection"],"falsifier":"The decisive observation is spectroscopic confirmation of the 42 grade A and B candidates plus high-resolution Einstein-radius measurements of the confirmed systems: if most of the grade A candidates turn out not to be lenses, or if confirmed systems at $\\theta_\\mathrm{E}<3$ arcsec disagree with VariLens by more than $2\\sigma$, the central claim of fast, reliable end-to-end lens modeling is not supported.","tokens_in":33629,"feed_emoji":"🔭","tokens_out":12769,"duration_ms":98615,"temperature":0.7,"pith_summary":"Strongly lensed quasars are powerful cosmological probes but are buried among tens of millions of ordinary sources. The paper presents VariLens, a physics-informed variational autoencoder in which one encoder feeds three heads: a decoder that reconstructs the five-band image, a regressor that outputs the parameters of a singular isothermal ellipsoid (SIE) mass model, and a classifier that returns the probability that an image is a lens. Trained entirely on mock lenses built by overlaying simulated quasar light on real SDSS galaxies, VariLens runs in milliseconds on a single CPU. For 22 confirmed lensed quasars with HSC images, the classifier recovers 16 (73% completeness), with a 5% false-positive rate and about 1% purity. On 20 known lensed quasars, its estimates agree with traditional lens modeling within $2\\sigma$ for systems with $\\theta_\\mathrm{E}<3$ arcsec, and screening 710,966 photometrically preselected sources yields 42 grade A and B candidates awaiting spectroscopic confirmation.","feed_headline":"VariLens finds 42 lens candidates, models each in milliseconds","feed_subtitle":"A one-network pipeline cuts 80 million HSC sources to 42 lenses, matching traditional Einstein radii within 2 sigma.","key_machinery":"The load-bearing object is VariLens, a physics-informed variational autoencoder whose 64-dimensional latent space is shared by three heads: a decoder that reconstructs the input image, a regressor that maps the latent vector to a Gaussian distribution over the SIE-plus-shear and source parameters, and a classification layer added after training. A singular isothermal ellipsoid is a standard mass profile whose Einstein radius sets the image scale of the lensed source. The physics-informed part is the Gaussian negative-log-likelihood loss on the predicted parameters, which forces the latent representation to encode the lens configuration while the reconstruction and KL terms keep it generative and regularized. This shared-latent design is what lets one forward pass return both a lens probability and a mass model in milliseconds.","core_discovery":"The paper claims that a single physics-informed variational autoencoder can replace separate detection, classification, and modeling stages in strong-lens searches. Its encoder-decoder reconstructs five-band HSC cutouts, a regressor branch predicts 11 physical parameters (lens center, complex ellipticity, Einstein radius, external shear, deflector and source redshifts, and source position) from the latent vector, and a fine-tuned classification head turns the same encoder into a lens/non-lens classifier. The physics enters through the loss: reconstruction mean squared error plus Kullback-Leibler divergence plus a Gaussian negative log-likelihood that ties the latent representation to parameters of an SIE + external-shear lens model. On simulated test data the network recovers most parameters with $R^2 \\gtrsim 0.8$, except external shear, which it pins near zero; on real data, VariLens and traditional modeling agree within $2\\sigma$ for the Einstein radius and positions of systems with $\\theta_\\mathrm{E}<3$ arcsec, and the full pipeline yields 42 grade A and B candidates from 80 million HSC sources.","pith_inferences":["The training-set bias toward bright SDSS luminous red galaxies at $z\\lesssim1$ likely limits completeness for compact, faint, or high-redshift deflectors; resampling the simulation to a uniform $\\theta_\\mathrm{E}$-$z_\\mathrm{gal}$ distribution is an implicit, testable remedy the paper leaves for future work.","A natural extension is to use VariLens's latent vector itself as a ranking feature, since the t-SNE analysis shows lenses and contaminants separate in the learned representation even before classifier fine-tuning.","If the 42 candidates are confirmed, their predicted Einstein radii and source redshifts could immediately prioritize which objects enter time-delay monitoring, connecting the discovery pipeline to $H_0$ measurements.","Running the same network on known lenses with $\\theta_\\mathrm{E}>2$ arcsec would quantify how much of the reported underestimation at large Einstein radii comes from training-set scarcity versus the SIE prior itself."],"forward_implications":["One CPU can screen survey-scale catalogs: the 80-million-source HSC parent sample is reduced to 13,831 network-ranked candidates and then to 42 visually confirmed candidate lenses, making spectroscopic follow-up feasible.","For lenses with $\\theta_\\mathrm{E}\\lesssim2$ arcsec, VariLens's Einstein radius, center, and source-position estimates can warm-start or cross-check traditional lens models, which otherwise take hours to weeks per system.","The same architecture transfers to upcoming wide surveys such as LSST and Euclid once retrained on their bandpasses, seeing, and pixel scale, with the learned Lyman-break feature giving more reliable source redshifts at $z\\gtrsim3$.","Because external shear is not recoverable from ground-based HSC images with this approach, shear-sensitive time-delay cosmography targets would still need higher-resolution or deeper data."],"supporting_citations":[{"why":"supplies the earlier CNN/ViT lens search and the lens-simulation training dataset from which VariLens's mock lenses are adapted.","marker":"Andika et al. 2023b"},{"why":"established the HSC SIE-plus-external-shear regression task with a ResNet and documented the shear-recovery difficulty that VariLens is compared against.","marker":"Schuldt et al. 2023a"},{"why":"PyAutoLens provides both the ray tracing used to create mock lensed images and the traditional lens modeling used for the 20-system comparison.","marker":"Nightingale et al. 2021"},{"why":"PyAutoGalaxy is used to fit observed deflector light with Sérsic/exponential profiles and to set the SIE parameters in the simulations.","marker":"Nightingale et al. 2018"},{"why":"underpins the variational autoencoder and reparameterization machinery that define VariLens's latent space and losses.","marker":"Kingma & Welling 2019"},{"why":"SIMQSO generates the thousand mock quasar spectra whose photometry and colors populate the simulated training lenses.","marker":"McGreer et al. 2013"},{"why":"supplies the confirmed lensed quasar sample used to evaluate classifier completeness and regression accuracy.","marker":"Lemon et al. 2023"},{"why":"the SDSS DR18 catalog provides the real deflector galaxies used as the bases of the mock lens training images.","marker":"Almeida et al. 2023"}],"fun_headline_variants":["Millisecond lens modeling: 42 quasar candidates from 80M sources","Physics-aware AI finds quasar lenses in milliseconds","Single network discovers and models 42 quasar lenses","From 80M sources to 42 lens candidates in milliseconds","AI lens finder: 42 candidates, millisecond modeling"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simulated training images—real galaxy cutouts with fake quasar point sources bent through a standard elliptical mass model—are representative enough of real HSC lensed quasars that the network's detection and parameter estimates carry over to survey data.","fun_headline_variants_meta":{"raw":{"variants":["Millisecond lens modeling: 42 quasar candidates from 80M sources","Physics-aware AI finds quasar lenses in milliseconds","Single network discovers and models 42 quasar lenses","From 80M sources to 42 lens candidates in milliseconds","AI lens finder: 42 candidates, millisecond modeling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00075,"raw_usage":{"total_tokens":3427,"prompt_tokens":1118,"completion_tokens":2309,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":734,"completion_tokens_details":{"reasoning_tokens":2235}},"tokens_in":734,"tokens_out":2309,"duration_ms":14767,"temperature":1.0,"reasoning_tokens":2235,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:48:56.106104+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The decisive observation is spectroscopic confirmation of the 42 grade A and B candidates plus high-resolution Einstein-radius measurements of the confirmed systems: if most of the grade A candidates turn out not to be lenses, or if confirmed systems at $\\theta_\\mathrm{E}<3$ arcsec disagree with VariLens by more than $2\\sigma$, the central claim of fast, reliable end-to-end lens modeling is not supported.","supporting_citations":[{"cited_title":"2021, The Journal of Open Source Software, 6, 2825","cited_arxiv_id":null,"evidence_quote":"PyAutoLens provides both the ray tracing used to create mock lensed images and the traditional lens modeling used for the 20-system comparison."},{"cited_title":"W., Dye , S., & Massey , R","cited_arxiv_id":null,"evidence_quote":"PyAutoGalaxy is used to fit observed deflector light with Sérsic/exponential profiles and to set the SIE parameters in the simulations."},{"cited_title":"D., Jiang , L., Fan , X., et al","cited_arxiv_id":null,"evidence_quote":"SIMQSO generates the thousand mock quasar spectra whose photometry and colors populate the simulated training lenses."}],"review_version":1}