{"id":"a24b78d1-a480-4e95-9aac-eaef0b2e63de","arxiv_id":"2411.18694","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Merging high-resolution single-band HST images with lower-resolution multi-band Legacy Survey images in a two-branch CNN improves strong-lens detection, reaching a true-positive rate of 0.55 at a false-positive rate of 10^-4.","lead":"Astronomers testing neural networks that combine sharp HST images with blurrier, multi-color Legacy Survey images found that merging the two data types detects more gravitational lenses at a low false-positive rate. The result is a stepping stone for the upcoming Euclid and Rubin/LSST surveys, which will use the same kind of multi-resolution image fusion.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline TPR gains are measured on mock lenses generated and selected by the same pipeline used for training; until the 364 real lenses in the test set are scored separately, the claim that merged HST+LS branches enhance detection over single-instrument models is not established for real data.","rationale":"The paper is a careful, internally consistent architecture comparison on mocks, and the authors are transparent about the limitations of the simulated data and the imperfect negative set. The central claim, however, is phrased as a demonstration that combining HST and LS images enhances detection. For that claim to hold on the surveys the paper targets (Euclid + LSST), the mock lenses must be representative of real lens populations. The 2-sigma LS visibility filter in Sect. 2 is the weakest link: it removes simulated lenses whose arcs are too faint for LS, which is exactly the population where adding LS might not help or could hurt. Since about 96% of test positives are such filtered mocks, the measured TPR ordering may be biased toward the merged architecture. The inclusion of 364 real lenses provides a ready, independent check, but the paper does not report aggregate scores for them. This is an addressable omission, not a fatal flaw, so the conditional verdict stands. I also note the Table 2 vs. Abstract/Sect. 4 discrepancy in the HST/LS TPR values; it does not change the merged-vs-single conclusion but reinforces the need for a corrected, real-lens-based evaluation before the headline numbers are used.","tokens_in":14739,"tokens_out":6417,"duration_ms":59340,"concrete_test":"Take the 364 real lens systems (Bolton et al. 2008; Auger et al. 2009; Pawase et al. 2014) as the positive set and the ~120,000 negative test cutouts as the negative set, and run each of the four trained networks to produce scores. Compute ROC curves, AUC, and TPR at FPR=10^-4 restricted to these real positives. If the merged-branches model does not outperform the HST-only and LS-only models on this real-positive subset (with uncertainty from N=364, e.g., bootstrap 95% CI on TPR), then the central claim that combining HST and LS significantly enhances detection is not supported for real data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison in Sect. 4 rests on a positive test set that is ~96% simulated (about 8,000 mocks vs. 364 real lenses, Table 1). Those mocks were generated with the pipeline described in Sect. 2 and then filtered so that the inserted arcs are visible at >2 sigma in all three LS bands as well as in HST. This selection builds in the condition that LS contains usable arc signal, which is precisely the regime in which a merged HST+LS model should gain. The 364 real lenses are included in the test set, but their scores are only shown individually (Fig. 13); no ROC or TPR is reported for real positives alone. Consequently, the headline improvement (TPR about 0.41 to 0.55 at FPR 10^-4) could be an artifact of the mock-selection pipeline rather than a property of real galaxy-scale lenses. This is the load-bearing assumption: if real survey lenses more often have arcs that are faint or absent in LS bands, the merged-branch advantage may shrink or vanish. A separate internal inconsistency (Table 2 lists LS=0.41, HST=0.45, while the Abstract and Sect. 4 text state HST=0.41, LS=0.45) further underlines that the quoted numbers need verification before being taken as the basis of the claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes and evaluates four ResNet-18-based CNN architectures for detecting galaxy-scale strong gravitational lenses by combining high-resolution single-band HST F814W images with lower-resolution, multi-band Legacy Survey g/r/z images. The four architectures are LS-only, HST-only, stacked (LS interpolated to HST pixel scale and processed together), and merged (separate HST and LS branches fused before the fully connected layer). Training and test data are dominated by mock lenses simulated with the GLEE pipeline using LRG deflectors and HUDF source images; only 364 real lenses are included in the test set and are not scored separately in the ROC analysis. At a false-positive rate of 10^-4, the authors report TPRs of about 0.41, 0.45, 0.51, and 0.55 for the four architectures and conclude that multi-instrument fusion significantly enhances detection. The paper also includes a control with downsampled HST images as a fourth LS filter, which supports the claim that merging is not equivalent to stacking.","tokens_in":14983,"tokens_out":7678,"duration_ms":70622,"significance":"If the reported gains hold on real survey data, the merged-branch architecture would be a useful template for combined Euclid and LSST lens searches, and the controlled comparison of fusion strategies is methodologically valuable. The downsampled-HST control and the use of a large negative set for low-FPR evaluation are strengths. However, the central quantitative claim rests on a test set that is about 96% simulated with the same pipeline used for training, and the real-lens subset is never evaluated separately. The external validity of the headline TPR improvement is therefore not yet established, although the issue is addressable within the scope of the manuscript.","major_comments":[{"comment":"The headline TPR comparison is computed on a positive test set containing roughly 8,000 mocks and only 364 real lenses. Because the mocks are generated with the same GLEE pipeline and HUDF source library used for training, and are selected (Sect. 2, Fig. 5) to have arcs visible at >2 sigma in all three LS bands, the test largely measures the networks' ability to recognize simulation-pipeline lenses rather than real galaxy-scale lenses. The paper does not report a separate ROC, TPR, or score statistics for the 364 real positives, which appear only as individual examples in Fig. 13. Please add this analysis and explicitly discuss how the 2-sigma LS-arc-visibility selection conditions the measured gain; without it, the claim that merged HST+LS branches enhance detection for real survey data is not supported.","section":"Section 4, Table 2, Fig. 10"},{"comment":"No error bars or uncertainty statements are provided for the TPR values at FPR=10^-4. At that FPR with about 120,000 negatives, the operating point is defined by only about 12 false positives, so the estimates are subject to non-negligible Poisson and threshold-sampling noise. The ten-split comparison in Fig. 10 (right) addresses only stacked versus merged and does not quantify the significance of the difference between single-instrument and fused models, which is the central 'significantly enhance' claim. Please propagate these uncertainties and, if possible, report bootstrap or split-half confidence intervals.","section":"Section 4, Table 2"},{"comment":"The reported TPR ordering for HST and LS is internally inconsistent: the abstract and Section 4 text state HST ~0.41 and LS ~0.45, while Table 2 lists LS=0.41 and HST=0.45. Because these numbers are the basis for the architecture ranking, this inconsistency must be resolved and the corrected values verified throughout the manuscript.","section":"Abstract, Section 4, Table 2"}],"minor_comments":[{"comment":"Table 1 lists positive counts summing to about 25,900, while the text states that rotation yields approximately 24,000 positive images and that the test set contains about 8,000 mocks plus 364 real lenses; please reconcile these numbers.","section":"Section 2, Table 1"},{"comment":"The description of balancing 'LS and HST contribute equally to the overall loss function' is unclear, since in the merged architecture the loss is computed after the fused feature vector; please specify the exact weighting scheme.","section":"Section 3"},{"comment":"The phrase 'ten different test sets' is not defined precisely; please state whether the models were retrained for each split, how the splits were drawn, and whether real lenses were included in each split.","section":"Section 4, Fig. 10"},{"comment":"The sentence claiming that 'the estimates from HST and LS alone are less reliable in the low-FPR regime compared to the combined data set' is unexplained, because all models are evaluated on the same negative set; please justify or remove this statement.","section":"Section 4"},{"comment":"There are minor typographical errors, including 'simulated simulated HST' in Section 2 and a duplicated article in the Fig. 13 caption ('compared to the the LS-only model').","section":"Section 2 and Fig. 13 caption"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern lands: the evaluation is dominated by mocks selected for LS arc visibility, and the real-lens subset is never scored separately. This is fixable by adding a separate real-lens analysis and uncertainty estimates, so I recommend major revision rather than rejection. The internal architecture comparison is a useful contribution if those gaps are closed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real contribution is the architectural comparison: it is, as far as I can tell, the first lens-search CNN to combine a high-resolution single-band space image with lower-resolution multiband ground images. The merged two-branch ResNet-18 is a sensible design, and the control test — downsampling HST to LS resolution and feeding it as a fourth LS filter — is a clever way to show that simple stacking does not exploit the high-resolution information. The simulation pipeline is careful: same lens parameters across resolutions, realistic PSFs, and a 2-sigma visibility cut to ensure the arcs are actually present in all three LS bands. The ten-test-set robustness check for stacked versus merged adds confidence that the ordering is real, at least for simulations.\n\nThe soft spot is in what the headline numbers actually measure. The positive test set is about 8,000 mocks plus 364 real lenses, so the TPR at FPR 10^-4 is almost entirely a mock-recognition rate. The mocks were selected to have arcs visible at >2 sigma in all three LS bands, which builds in the condition that LS carries usable signal — exactly the regime where merging should help. The 364 real lenses are only shown as individual examples, not as a separate ROC. Without a real-lens-only evaluation, the claim that merging enhances detection for actual survey data is an extrapolation. That is not fatal, but the conclusion sections states it as a demonstrated result.\n\nThere is also a concrete internal inconsistency: Table 2 lists LS=0.41 and HST=0.45, while the Abstract and the results text say HST=0.41 and LS=0.45. One digit swap, but it tells me the numbers need verification before being trusted.\n\nThe authors are honest about the limitations of the data, and the paper is a workmanlike engineering study. But the central claim would be much stronger if they scored the real lenses separately and reported ROC curves with error bars on the mock-based estimates.\n\nWho is this for? Anyone working on lens search pipelines for Euclid, LSST, or similar surveys, and anyone interested in multi-resolution data fusion. It deserves a serious referee — the method is new and the evaluation can be fixed. I would like to see a revised version with real-lens ROC and the number inconsistency resolved.","headline":"A clean architectural comparison on simulated lenses — the merged HST+LS branch wins — but the headline TPRs are measured on mocks from the same pipeline, so the claim about real survey data is not yet established.","tokens_in":15585,"tokens_out":1743,"would_cite":true,"duration_ms":17929,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Combining HST images with Legacy Survey images through a merged-branch neural network detects more gravitational lenses at a fixed false-positive rate than using either data source alone.","keywords":["strong gravitational lensing","galaxy-scale lenses","convolutional neural networks","multi-resolution image fusion","Euclid","LSST","mock image simulations","ResNet-18"],"falsifier":"Apply the four trained networks to a sample of real, spectroscopically confirmed gravitational lenses and an equal number of independent non-lens galaxies from Euclid and LSST overlap fields, and check whether the merged-branch true-positive rate at a false-positive rate of $10^{-4}$ is measurably above the HST-only and LS-only rates.","tokens_in":14513,"feed_emoji":"🔭","tokens_out":8390,"duration_ms":70612,"temperature":0.7,"pith_summary":"Strong gravitational lenses are rare, and the coming wide surveys Euclid and LSST will produce far more galaxy images than humans can inspect, so automated classifiers must recover as many lenses as possible for a fixed number of false alarms. This paper argues that the best classifier uses two complementary views at once: sharp single-band space imaging from HST, which resolves arcs and Einstein rings, and blurrier three-band ground imaging from the Legacy Survey, which adds color. The authors insert simulated lenses into real galaxy images and train four ResNet-18 classifiers, comparing them at a false-positive rate of $10^{-4}$. The merged-branch network, which processes HST and LS through separate branches and fuses their features before classification, reaches a true-positive rate near 0.55, versus about 0.41 for HST alone, 0.45 for the ground data alone, and 0.51 for a stacked four-band variant. If the ranking carries over to Euclid and LSST, surveys expected to yield roughly $10^5$ lenses would harvest a meaningfully larger share of them at no extra observing cost.","feed_headline":"Merged HST and ground images detect a third more lenses","feed_subtitle":"At one false alarm per 10,000 galaxies, the combined model catches 55% of lenses versus 41% for HST alone.","key_machinery":"The load-bearing mechanism is the merged-branches ResNet-18 architecture. One branch ingests the $200\\times200$ HST F814W cutout and the other ingests the $38\\times38$ three-band LS cutout, each through residual convolutional blocks; the flattened 512-feature vectors are concatenated into a 1024-dimensional vector, reduced by a fully connected layer, and classified. This preserves each instrument's native pixel scale instead of forcing all data onto a common grid. The training set is built by simulating singular isothermal ellipsoid lenses with GLEE, painting arcs from Hubble Ultra Deep Field sources onto real luminous red galaxy images in both HST and LS with identical geometry, keeping only systems whose arcs are detected above $2\\sigma$ in all three LS bands, and rotating the images to reach roughly 24,000 positive examples.","core_discovery":"The paper's central claim is that fusing high-resolution single-band and low-resolution multi-band images with a two-branch convolutional network detects more galaxy-scale strong lenses than either data source alone, and more than a naive stack of all bands on a common pixel grid. On a test set of roughly 8,000 simulated lenses plus 364 real lens candidates, at a false-positive rate of $10^{-4}$ the true-positive rate is about 0.41 for HST F814W alone, 0.45 for Legacy Survey $g/r/z$ alone, 0.51 when the LS bands are interpolated to HST pixel scale and processed as a four-channel stack, and 0.55 when HST and LS pass through separate ResNet-18 branches whose 512-dimensional feature vectors are concatenated before classification. The merged architecture assigns its highest-confidence scores to small-Einstein-radius systems ($\\theta_{\\rm E} \\le 0.8''$), the regime where single-instrument searches have historically struggled. A control experiment that downsampled HST to LS resolution and treated it as a fourth ground band performed like the stacked model, showing that the gain comes from preserving native resolution in its own branch.","pith_inferences":["Editorial inference: the merged-branch advantage suggests a general design rule that each instrument's native pixel scale should be processed in its own network branch and fused late, rather than interpolating all inputs onto one grid.","Editorial inference: the paper's own caveat that negative cutouts come from the same fields as the lenses means the reported false-positive rates could be optimistic, and a cleaner test would build negatives from fully independent sky regions.","Editorial inference: a testable extension is to measure the merged-branch ROC on real confirmed lenses in Euclid Early Release Observations overlapping LSST coverage, since the gain over single-instrument models could shrink when the training mocks no longer match the data.","Editorial inference: because the mock sample only keeps arcs visible above $2\\sigma$ in all three LS bands, the measured fusion gain may not apply to the very faintest arcs that only HST can see, and restricting the test set to faint arcs would bound the effect."],"forward_implications":["Applying the merged-branch design to Euclid VIS plus LSST $g/r/z$ should recover more galaxy-scale lenses at a fixed false-positive budget than either survey alone.","A true-positive rate near 0.55 at $10^{-4}$ false positives, projected onto surveys expected to contain roughly 100,000 lenses, implies tens of thousands of additional recovered systems compared with single-instrument classifiers.","Because the merged network scores low-Einstein-radius systems ($\\theta_{\\rm E} \\le 0.8''$) with high confidence, small-separation lenses become more accessible to automated discovery.","The same architecture can fuse Euclid's own VIS and NISP channels, extending multi-resolution fusion beyond the HST/LS combination tested here.","For transient follow-up programs hunting lensed supernovae, a higher true-positive rate at low false-positive rate reduces the human inspection bottleneck."],"supporting_citations":[{"why":"Supplies ResNet-18, the convolutional backbone that all four tested architectures modify.","marker":"He et al. 2016"},{"why":"Defines the mock-lens simulation procedure of painting arcs onto real images that the training set follows.","marker":"Schuldt et al. 2021"},{"why":"Provides GLEE, the lensing code used to render the simulated arcs in both HST and LS images.","marker":"Suyu & Halkola 2010"},{"why":"Supplies the K-nearest-neighbor velocity-dispersion prediction and the negative-sample construction used for training.","marker":"Rojas et al. 2022"},{"why":"Gives the luminous-red-galaxy color selection that defines the lens and negative galaxy samples.","marker":"Zhou et al. 2023"},{"why":"Provides the Hubble Ultra Deep Field source catalog, including redshifts, used as the lensed background galaxies.","marker":"Inami et al. 2017"},{"why":"Serves as the external performance baseline for a single-instrument lens search on HSC data.","marker":"Cañameras et al. 2021"},{"why":"Supplies real spectroscopically confirmed lens candidates included in the test set.","marker":"Bolton et al. 2008"},{"why":"Describes LSST, the survey whose synergy with Euclid this experiment is designed to anticipate.","marker":"Ivezić et al. 2019"}],"fun_headline_variants":["Dual-instrument AI catches 55% of simulated lenses, up from 41%","Merged HST and ground data lift lens detection to 55% from 41%","Two-branch network finds a third more lenses than single-instrument","Fusing space and ground images reveals more gravitational lenses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The mock lenses, simulated Einstein arcs pasted onto real galaxy images and kept only when visible above $2\\sigma$ in all three ground-based bands, represent the real galaxy-scale lens population that Euclid and LSST will observe, so the measured ranking of architectures transfers to survey data.","fun_headline_variants_meta":{"raw":{"variants":["Dual-instrument AI catches 55% of simulated lenses, up from 41%","Merged HST and ground data lift lens detection to 55% from 41%","Two-branch network finds a third more lenses than single-instrument","Fusing space and ground images reveals more gravitational lenses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000612,"raw_usage":{"total_tokens":2952,"prompt_tokens":1155,"completion_tokens":1797,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":771,"completion_tokens_details":{"reasoning_tokens":1726}},"tokens_in":771,"tokens_out":1797,"duration_ms":11750,"temperature":1.0,"reasoning_tokens":1726,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:58:17.578911+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the four trained networks to a sample of real, spectroscopically confirmed gravitational lenses and an equal number of independent non-lens galaxies from Euclid and LSST overlap fields, and check whether the merged-branch true-positive rate at a false-positive rate of $10^{-4}$ is measurably above the HST-only and LS-only rates.","supporting_citations":[{"cited_title":"Find the haystacks, then look for needles: The rate of strongly lensed transients in galaxy-galaxy strong gravitational lenses","cited_arxiv_id":"2407.04080","evidence_quote":"Supplies the K-nearest-neighbor velocity-dispersion prediction and the negative-sample construction used for training."},{"cited_title":"A., et al","cited_arxiv_id":null,"evidence_quote":"Gives the luminous-red-galaxy color selection that defines the lens and negative galaxy samples."}],"review_version":1}