Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

LTDE: The Lens Time Delay Experiment I. From pixels to light curves

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read A public benchmark challenge tests how well photometry algorithms can turn pixels of lensed quasars into accurate light curves.

desk verdict A well-designed community benchmark for lensed-quasar photometry that is honest about its terms, but whose PSF realism is unvalidated and therefore limits how far the leaderboard transfers to real LSST data. read the letter →

arxiv 2504.16249 v1 pith:H4GEGWIJ submitted 2025-04-22 astro-ph.IM astro-ph.CO

classification astro-ph.IMastro-ph.CO
keywords gravitationallensingquasarphotometrylightcurvespointspreadfunctiondeblendingsimulatedobservationsbenchmarkchallengetimedelaycosmography
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper opens a community challenge to test how well existing photometry algorithms can turn pixels of gravitationally lensed quasars into accurate light curves. It releases 108 simulated LSST-like exposures covering four image configurations, with known input light curves, known point-spread functions, and controlled noise, and invites any research team to submit extracted light curves for scoring on defined metrics. The motivation is that the upcoming LSST survey will find thousands of lensed quasars whose separations are about an arcsecond, demanding careful deblending of the quasar images from each other and from the lens and host galaxies. A sympathetic reader would care because the methods that win this benchmark could become the standard tools for time-delay cosmography and quasar microlensing studies in the LSST era.

What carries the argument

The load-bearing object is the simulation pipeline that re-creates LSST exposures from survey catalogues, injecting each lensed quasar with a super-sampled point-spread function made of a Moffat profile plus a shapelet correction, then adding Poisson photon noise and a fixed instrumental noise model. The key design choice is that the true PSF and the true input light curves are known to the challenge organizers but hidden from participants, so any mismatch between submitted and true light curves can be attributed to the photometry method rather than to the simulation. The three rounds progressively remove information: first the noiseless PSF is given, then it must be estimated from the data, and finally it varies across the field of view through an atmospheric distortion model. This isolates whether a method fails at deblending, at PSF estimation, or at handling a spatially varying PSF.

What would settle it

Re-run the challenge on a set of real LSST images of lensed quasars with independently measured time delays, using the same methods; if the method ranking on real data does not match the ranking on the 108 simulated systems, the simulation's PSF and noise models are not faithful enough to predict real performance.

Watch

Extended reading notes

Core claim

The central claim is that a public set of simulated cutouts, generated with controlled PSF, sky, and noise models, can serve as a standard benchmark for light-curve extraction from lensed quasar images. The paper establishes the benchmark by defining three rounds of increasing difficulty — a known and constant PSF, an unknown but constant PSF, and an unknown spatially varying PSF — and by specifying three metrics: a per-image $\chi^2$ test for flux leakage between quasar images, an offset-marginalised $\chi^2$ for photometric accuracy, and the absolute percentage error in reported flux uncertainties. All simulations share a damped random walk model for quasar variability and a Sérsic profile for the host and lens galaxies, so the ground-truth light curves are known exactly. If the benchmark succeeds, relative method performance on these simulations should predict relative performance on real LSST observations of lensed quasars.

Load-bearing premise

The simulated exposures capture the PSF, sky, and noise variations that actually limit photometry of lensed quasars in the LSST survey, because the benchmark's ranking of methods transfers to real data only if those models are representative.

Editorial extensions

If this is right

  • Methods that score well in all three rounds can be trusted for LSST-era monitoring of lensed quasars, assuming the simulated exposures are representative of real data.
  • The three-round design pinpoints the failure mode of each method: blending, PSF estimation, or spatial PSF variation.
  • The public dataset gives algorithm developers a fixed testbed, so future methods can be compared against the same 108 systems rather than ad hoc simulations.
  • If no method passes Round #2 comfortably, the community learns that spatially variable PSF handling needs work before LSST data arrive.
  • The metrics define an objective ranking, allowing the follow-up paper to report which families of photometry tools are best for time-delay work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the simulations omit microlensing and other astrophysical variability, the benchmark measures only measurement and deblending errors; real time-delay measurements will also face source-intrinsic microlensing, so the achievable accuracy on real data may be worse than the benchmark suggests.
  • The published data could be reused to test full time-delay estimation, not just photometry: participants' light curves could be fed to delay-estimation algorithms and the recovered delays checked against the true delays hidden in the simulations.
  • A natural extension is to add a round with realistic PSF residuals from real telescope images, which would test whether the Moffat-plus-shapelet PSF model is the limiting factor rather than the photometry itself.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript presents the Lens Time Delay Experiment (LTDE), a community challenge whose first round is a public dataset of simulated LSST-like exposures of gravitationally lensed quasars. The simulations are generated with MOLET: lensed quasar images are produced from an SIE mass profile with four configurations (double, cusp, fold, cross); the lens and host galaxies follow Sérsic profiles; quasar variability is modeled as a damped random walk; and exposures are built on LSST DP0.2 sky and atmospheric distortion data with a Moffat-plus-shapelet PSF and an LSST noise model. The authors define three rounds of increasing difficulty (known constant PSF, unknown constant PSF, unknown spatially varying PSF), specify a CSV submission format, and define three evaluation metrics: a per-image chi-squared, an offset-marginalized chi-squared on the summed light curve, and an absolute percentage error on reported uncertainties. The central deliverable is the simulated data release and the invitation to the community to submit light curves; evaluation results are deferred to a follow-up paper.

Significance. The paper's contribution is a reusable, ground-truth benchmark rather than a new photometric method. This is valuable: the simulated cutouts cover a broad range of lens configurations, brightness contrasts, and Einstein radii; the ground truth is known by construction, so the metrics are not circular; and the public data release with a fixed submission format lowers the barrier for comparing deblending and photometry algorithms. The paper is also honest about its scope: it does not claim to have solved the photometric problem, and the metrics are clearly specified. If the simulated exposures are representative of LSST data, this benchmark could help develop and validate algorithms for time-delay cosmography and microlensing. The main weakness is that the paper provides no evidence that the simulated PSF and noise are representative, and no baseline run is shown, so the transferability of the leaderboard ranking to real LSST data is currently unquantified.

major comments (3)
  1. [Section 2.3, Eqs. (2)-(5)] The realism of the PSF is the central load-bearing assumption for this benchmark, but the paper gives no numerical values for the Moffat alpha and gamma distributions, nor for the shapelet component added to the PSF. The text only says the ranges are 'consistent with observations' (citing Trujillo et al. 2001), and no comparison is made between the simulated PSF statistics (FWHM distribution, ellipticity e1,e2 and its spatial/temporal correlation, higher-order moments) and real LSST or DP0.2 PSF measurements. Because the benchmark's ranking will transfer to LSST only if the PSF model is representative, the authors should report the exact parameter ranges and add a validation subsection comparing simulated PSF statistics to observed LSST/DP0.2 values.
  2. [Section 4 and overall] No baseline photometry pipeline is run on the simulated exposures. The paper claims the experiment will 'evaluate the limits of current photometric algorithms' and defines metrics in Section 4, but it does not show that any method can recover the injected light curves, nor that the metrics are discriminatory. Even a simple baseline (e.g., forced aperture photometry or a PSF fit on a single configuration) would calibrate the expected difficulty and demonstrate that the challenge is feasible; without it, the reader cannot judge whether the benchmark is trivially easy, impossible, or affected by simulation artifacts.
  3. [Section 4, Eq. (12)] The definition of the 'true uncertainty' sigma_{n,i} is ambiguous for blended quasar images. The text states that it is the sum in quadrature of instrumental and photon noise, but the photon noise relevant to a flux measurement of one deblended image depends on the aperture or model used and includes contributions from neighboring quasar images, the lens galaxy, and sky. Without a precise statement of how sigma_{n,i} is computed from the simulation (which aperture, which components are included in the photon count), the error-recovery metric epsilon is not well defined and may not be a fair basis for ranking submissions.
minor comments (5)
  1. [Section 2.3 and Eq. (8)] The symbol gamma is used both for the Moffat scale parameter in Eq. (2) and for the normalization factor in Eq. (8); consider renaming one of them to avoid confusion.
  2. [Section 5] The phrase 'for advertising purposes' is unclear; if the intended meaning is organizational, please reword.
  3. [Section 5] The submission format should specify that flux_AB (and similar combined labels) is the sum of the fluxes of the listed images, since the current wording only says to use the suffixes.
  4. [Section 1] The text contains typographical errors such as '(de)magnificatoin' and broken ligatures in the PDF extraction (e.g., 'o ffers' and 'di fferent'); these should be cleaned up.
  5. [Section 5] The deadline dates (30 May 2025 and 30 June 2025) have passed; please update the paper to state the current status of the rounds or note that the challenge remains open.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular reasoning: the paper releases simulated images whose true light curves are inputs by design, and the scoring metrics compare submissions directly against those known inputs.

full rationale

The paper makes no claim to derive or predict the true light curves; it explicitly generates mock lensed quasars using MOLET with input damped-random-walk variability (Section 2.2), injects lensed sources into LSST-like exposures with controlled PSFs and noise (Sections 2.3 and 2.4), and then evaluates participant submissions against these known inputs using chi-squared metrics (Section 4). The central deliverable is a public simulated dataset and standardized scoring, so the ground-truth light curves are inputs by construction rather than outputs of the evaluated methods. This is a self-contained benchmark design, not a hidden reduction. The only self-citation is the use of MOLET (Vernardos 2022), a published software package by one of the authors, but it is used as a simulation tool and not as an authority invoked to force a conclusion; no uniqueness theorem or ansatz is smuggled through it. The realism of the PSF and sky models relative to real LSST data is an external-validation and transferability concern, not circularity, because the benchmark's ground truth is not built from the quantities being scored. Accordingly, no circular step is found and the score is 0.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

No new physical entities are introduced. All inputs are standard astronomical models or existing survey simulations. The free parameters are simulation design choices that define the benchmark's difficulty and realism.

free parameters (5)
  • Sersic profile parameters for lens and host galaxy = not stated
    Adopted brightness profiles in Section 2.1; simulated systems' blending depends on these choices, and no values are listed.
  • SIE Einstein radius scaling factors = 0.75 and 1.5
    Section 3: Einstein radius is varied by factors of 0.75 and 1.5; these discrete choices define the mock sample but are not justified as representative coverage.
  • Host and lens galaxy brightness contrast factors = 0.1 and 0.5
    Section 3: host and lens galaxy brightness are varied by factors 0.1 and 0.5; these factors determine the difficulty of deblending.
  • Moffat PSF parameters alpha and gamma = not specified
    Section 2.3: drawn from log-normal and uniform distributions with ranges 'consistent with observations'; exact values are not reported.
  • DRW variability parameters = not specified
    Section 2.2: quasar variations follow a damped random walk (MacLeod et al. 2010); the parameters are not given, affecting the correlation structure of the true light curves.
assumptions (5)
  • domain assumption Lens mass distribution is a singular isothermal ellipsoid (Eq. 1).
    Section 2.1: SIE is a common but non-universal model; mock image positions and fluxes depend on it.
  • domain assumption Lens and host galaxy light follow Sersic profiles.
    Section 2.1: adopted for simplicity; real galaxies have more complex profiles.
  • domain assumption Quasar intrinsic variability follows a damped random walk.
    Section 2.2: DRW is a standard empirical model, but not true for all quasars.
  • domain assumption PSF can be represented as a Moffat profile with a shapelet perturbation.
    Section 2.3: this is a model choice; if the real LSST PSF has different morphology, benchmark realism degrades.
  • domain assumption DP0.2 catalogs, sky brightness, and distortion measurements are a valid proxy for LSST observing conditions.
    Section 2.4: the benchmark inherits any biases of the DP0.2 simulation and pipeline.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LTDE: The Lens Time Delay Experiment I. From pixels to light curves." pith.science (2026). https://pith.science/paper/H4GEGWIJ

@misc{pith2026250416249,
  author       = {Pith},
  title        = {Pith review of: LTDE: The Lens Time Delay Experiment I. From pixels to light curves},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H4GEGWIJ}},
  note         = {Machine review of arXiv:2504.16249}
}
abstract

Gravitationally lensed quasars offer a unique opportunity to study cosmological and extragalactic phenomena, using reliable light curves of the lensed images. This requires accurate deblending of the quasar images, which is not trivial due to the small separation between the lensed images (typically $\sim1$ arcsec) and because there is light contamination by the lensing galaxy and the quasar host galaxy. We propose a series of experiments aimed at testing our ability to extract precise and accurate photometry of lensed quasars. In this first paper, we focus on evaluating our ability to extract light curves from simulated CCD images of lensed quasars spanning a broad range of configurations and assuming different observational/instrumental conditions. Specifically, the experiment proposes to go from pixels to light curves and to evaluate the limits of current photometric algorithms. Our experiment has several steps, from data with known point spread function (PSF), to an unknown spatially-variable PSF field that the user has to take into account. This paper is the release of our simulated images. Anyone can extract the light curves and submit their results by the deadline. These will be evaluated with the metrics described below. Our set of simulations will be public and it is meant to be a benchmark for time-domain surveys like Rubin-LSST or other follow-up time-domain observations at higher temporal cadence. It is also meant to be a test set to help develop new algorithms in the future.

Figures

Figures reproduced from arXiv: 2504.16249 by the authors.

Figure 1
Figure 1. Top: a large image separation lensed quasar J1537-3010 (Lemon et al. 2019), the left panel shows an image taken with the Hubble space telescope, in the right one taken with a ground-based telescopes. Bot￾tom: same as top but for the lensed quasar DESJ0405-3308 (Anguita et al. 2018). To obtain photometric measurements there are several meth￾ods one can choose from, each with their own strengths and lim￾itations. Aper… view at source ↗
Figure 2
Figure 2. Examples of lens configurations. The inner (solid) and outer (dashed) caustics are plotted. The position of the quasar is denoted by the x, and the multiple images generated by a circle [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Example of the brightness variations of a mock quasar following a damped random walk model. goal of this paper is to quantify how well different methods per￾form at recovering the true quasar light curves, we can safely ignore any other sources of variability (e.g. microlenisng, rever￾beration, etc.). We show an example of the variability in [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Example of a distortion field due to atmospheric effects taken from the LSST DP0.2. The direction of the distortion is indicated by the lines, which are exaggerated and aligned with the major axis of the ellipse that represents it. uniform distribution with ranges that…
Figure 6
Figure 6. Figure 6: Left: example of a simulated exposure Right: zoomed window where the lens is injected. Note the gap between the detectors which will have different orientations for each different exposure. – The host galaxy brightness by a factor of either 0.1 and 0.5. – The lensing g…
Figure 5
Figure 5. Figure 5: Diagram showing how an exposure is generated where σread-out = 8.8 electrons, σdark = 0.2 electrons, texp = 15 seconds and nexp = 2 corresponding to two back-to-back 15 sec￾onds exposures. Lastly, the photons from the actual astrophysical objects follow a Poissonian di…
Figure 7
Figure 7. Figure 7: All 108 mock cutouts. The cutouts are divided as follows. Groups of (3×3) have the same θE and lens configuration, and vary the brightness of the host (row) and lensing (column) galaxy. The top group of (3 × 9) have the same lens configuration (fold) and every 3 column…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Investigating the Dark Energy Constraint from Strongly Lensed AGN at LSST-Scale

    astro-ph.CO 2025-11 conditional novelty 6.0 of 10

    A simulated sample of 800 LSST lensed AGN, analyzed with a new hierarchical time-delay inference code, is forecast to yield ~2.5% H0 precision and a dark-energy figure of merit of 6.7 in w0waCDM.

Reference graph

Works this paper leans on

35 extracted references · 16 canonical work pages · cited by 1 Pith paper

  1. [1]

    A., O’Riordan, C

    Acevedo Barroso, J. A., O’Riordan, C. M., Clément, B., et al. 2024, arXiv e- prints, arXiv:2408.06217

  2. [2]

    Joint multiband deconvolution for Euclid and Vera C. Rubin images

    Akhaury, U., Jablonka, P., Courbin, F., & Starck, J.-L. 2025, arXiv e-prints, arXiv:2502.17177

  3. [3]

    L., Kuropatkin, N., et al

    Anguita, T., Schechter, P. L., Kuropatkin, N., et al. 2018, MNRAS, 480, 5017

  4. [4]

    Bernstein, G. M. & Jarvis, M. 2002, AJ, 123, 583 Article number, page 6 of 7 Favio Neira et al.: LTDE: The Lens Time Delay Experiment

  5. [5]

    & Arnouts, S

    Bertin, E. & Arnouts, S. 1996, A&AS, 117, 393

  6. [6]

    2019, in Astronomical Society of the Pacific Conference Series, V ol

    Bosch, J., AlSayyad, Y ., Armstrong, R., et al. 2019, in Astronomical Society of the Pacific Conference Series, V ol. 523, Astronomical Data Analysis Software and Systems XXVII, ed. P. J. Teuben, M. W. Pound, B. A. Thomas, & E. M. Warner, 521

  7. [7]

    Bretthorst, G. L. 1988, Bayesian spectrum Analysis and parameter estimation (Springer-Verlag Berlin Heidelberg)

  8. [8]

    Cooke, J. H. & Kantowski, R. 1975, ApJ, 195, L11

Show all 35 references
  1. [9]

    2005, in IAU Symposium, V ol

    Courbin, F., Eigenbrod, A., Vuissoz, C., Meylan, G., & Magain, P. 2005, in IAU Symposium, V ol. 225, Gravitational Lensing Impact on Cosmology, ed. Y . Mellier & G. Meylan, 297–303 Euclid Collaboration, Mellier, Y ., Abdurro’uf, et al. 2024, arXiv e-prints, arXiv:2405.13491 Iv...

  2. [10]

    E., et al

    Kaiser, N., Aussel, H., Burke, B. E., et al. 2002, in Society of Photo-Optical Instrumentation Engineers (SPIE) Conference Series, V ol. 4836, Survey and Other Telescope Technologies and Discoveries, ed. J. A. Tyson & S. Wol ff, 154–164

  3. [11]

    A., Auger, M

    Lemon, C. A., Auger, M. W., & McMahon, R. G. 2019, MNRAS, 483, 4242 LSST Dark Energy Science Collaboration, Abolfathi, B., Armstrong, R., et al. 2021, arXiv e-prints, arXiv:2101.04855

  4. [12]

    Lucy, L. B. 1994, in The Restoration of HST Images and Spectra - II, ed. R. J. Hanisch & R. L. White, 79

  5. [13]

    L., Ivezi´c, Ž., Kochanek, C

    MacLeod, C. L., Ivezi´c, Ž., Kochanek, C. S., et al. 2010, ApJ, 721, 1014

  6. [14]

    1998, ApJ, 494, 472

    Magain, P., Courbin, F., & Sohy, S. 1998, ApJ, 494, 472

  7. [15]

    2023, The Journal of Open Source Software, 8, 5340

    Michalewicz, K., Millon, M., Dux, F., & Courbin, F. 2023, The Journal of Open Source Software, 8, 5340

  8. [16]

    2020, A&A, 640, A105

    Millon, M., Courbin, F., Bonvin, V ., et al. 2020, A&A, 640, A105

  9. [17]

    & Marshall, P

    Oguri, M. & Marshall, P. J. 2010, MNRAS, 405, 2579

  10. [18]

    1964, MNRAS, 128, 307

    Refsdal, S. 1964, MNRAS, 128, 307

  11. [19]

    Saha, P., Sluse, D., Wagner, J., & Williams, L. L. R. 2024, Space Sci. Rev., 220, 12

  12. [20]

    Schneider, P., Ehlers, J., & Falco, E. E. 1992, Gravitational Lenses

  13. [21]

    & Schneider, P

    Seitz, C. & Schneider, P. 1997, A&A, 318, 687

  14. [22]

    Sersic, J. L. 1968, Atlas de Galaxias Australes

  15. [23]

    Stetson, P. B. 1987, PASP, 99, 191

  16. [24]

    L., Ivezi´c, Ž., & MacLeod, C

    Suberlak, K. L., Ivezi´c, Ž., & MacLeod, C. 2021, ApJ, 907, 96

  17. [25]

    Sureau, F., Lechat, A., & Starck, J. L. 2020, A&A, 641, A67

  18. [26]

    Taak, Y . C. & Treu, T. 2023, MNRAS, 524, 5446

  19. [27]

    2013, A&A, 553, A120

    Tewes, M., Courbin, F., & Meylan, G. 2013, A&A, 553, A120

  20. [28]

    Trujillo, I., Aguerri, J. A. L., Cepa, J., & Gutiérrez, C. M. 2001, MNRAS, 328, 977

  21. [29]

    2003, Acta Astron., 53, 291

    Udalski, A. 2003, Acta Astron., 53, 291

  22. [30]

    2023, arXiv e-prints, arXiv:2306.11781

    Vegetti, S., Birrer, S., Despali, G., et al. 2023, arXiv e-prints, arXiv:2306.11781

  23. [31]

    2023, arXiv e-prints, arXiv:2311.13305

    Verde, L., Schöneberg, N., & Gil-Marín, H. 2023, arXiv e-prints, arXiv:2311.13305

  24. [32]

    2022, MNRAS, 511, 4417

    Vernardos, G. 2022, MNRAS, 511, 4417

  25. [33]

    C., Suyu, S

    Wong, K. C., Suyu, S. H., Chen, G. C. F., et al. 2020, MNRAS, 498, 1420

  26. [34]

    2023, MNRAS, 520, 2328

    Zhang, T., Almoubayyed, H., Mandelbaum, R., et al. 2023, MNRAS, 520, 2328

  27. [35]

    2018, MNRAS, 481, 1149 Article number, page 7 of 7

    Zuntz, J., Sheldon, E., Samuroff, S., et al. 2018, MNRAS, 481, 1149 Article number, page 7 of 7

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.