{"id":"0c15da8f-c5cf-4819-916f-46c665e26927","arxiv_id":"2507.11635","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Galaxy-catalogue-based follow-up of gravitational wave events is valuable only for nearby events below about 300 Mpc; beyond that, simple 2D sky-map pointing performs as well or better.","lead":"This paper simulates how well galaxy catalogues can help telescopes find the light from gravitational wave events in the current LIGO/Virgo/KAGRA run, using the GROWTH-India and WINTER telescopes. It finds galaxy-targeted pointing helps only for nearby events below about 300 Mpc, and that a new mass-filling trick modestly boosts coverage.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The crossover claim is evaluated against a mass-filling model that assumes missing galaxy mass is uniform on the sky; if the real missing mass is clustered along large-scale structure, the 300–400 Mpc advantage of S_2DP over S_3DPM may disappear.","rationale":"I read the paper as a simulation-based comparison of four follow-up ranking schemes under O4-like conditions. The internal logic is consistent: the sky maps, telescope tilings, and event sample are described in enough detail that the reported ratios are plausible. The strongest support is the qualitative agreement across telescopes and the explicit dependence of the crossover on catalogue completeness. However, the evaluation metric is not an external truth; it is the mass-filled map constructed under a uniform-missing-mass assumption. This is the weakest link because the headline recommendation is a distance cut whose value is set by catalogue completeness combined with that assumption. The reader's conditional verdict already identifies this; I agree. I would not reject the paper, because the assumption is transparent and the authors flag it as future work. I would keep the verdict conditional, asking for a clustered missing-mass variant or an explicit demonstration that the crossover is insensitive to the spatial distribution of missing mass. The paper would also benefit from releasing the analysis code, but that is secondary to the ground-truth dependence. No ad hominem intended; the concern is purely about the sensitivity of the conclusions to a stated modelling choice.","tokens_in":16188,"tokens_out":5674,"duration_ms":72259,"concrete_test":"Re-run the pipeline with a direction-dependent missing-mass model: assign the missing mass in each 20 Mpc shell in proportion to the local surface density of known NED-LVS galaxies (smoothed on ~1 deg scale), keeping all telescope, tiling, and event-selection choices identical. Then recompute R = p(S_3DPM)/p(S_2DP) for GIT and WINTER as a function of distance. If the median ratio stays above unity beyond 400 Mpc, or the crossover moves significantly, the distance cut in the abstract is an artifact of the uniform-sky assumption. If the crossover is unchanged, the recommendation survives. A complementary check is to rerun with a mock catalogue that matches the NED-LVS selection function but has known true galaxy positions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparison is defined by the mass-filling model in §4.4. Missing mass per 20 Mpc shell is computed from NED-LVS completeness, then spread uniformly over healpix voxels; the resulting corrected mass \\hat m is used both as the S_3DMF ranking weight and as the ground-truth figure of merit p_3D × \\hat m against which every strategy is scored in §5. The ranking advantage of S_2DP over S_3DPM beyond ~300–400 Mpc therefore depends on the assumption that uncatalogued galaxies are distributed isotropically. In the real local universe, galaxies — including those absent from NED-LVS — trace the same large-scale structure overdensities as the catalogued galaxies, so an incomplete catalogue still contains directional information about where the missing mass is. If the missing mass follows that clustering, S_3DPM should not degrade to S_2DP as quickly, and the quoted crossover distance is not robust. The authors state the assumption explicitly in §4.4 ('we assume the missing galaxy mass to be uniform all over the sky') and defer a direction-dependent treatment to future work, so the issue is acknowledged but unresolved. A secondary, related circularity: because S_3DMF and the evaluation metric are built from the same \\hat m, the few-percent gain of S_3DMF over the other schemes is partly guaranteed by construction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper evaluates galaxy-catalogue-based optical follow-up strategies for gravitational-wave events in the O4 era, using the NED-LVS catalogue, 329 simulated BNS/NSBH events from Kiendrebeogo et al., and realistic tile grids for GIT and WINTER. Four ranking schemes are compared (2D probability tiling S_2DP, mass-only S_M, 3D probability times catalogued mass S_3DPM, and 3D probability times completeness-corrected mass S_3DMF), and performance is scored by the summed p_3D times margin over the top 100 tiles. The main empirical claims are that S_3DPM outperforms S_2DP for events within roughly 300-400 Mpc, that S_2DP is better beyond that, and that S_3DMF yields a few percent higher coverage than both. The paper recommends using galaxy catalogues for nearby events and 2D tiling for distant ones, and emphasizes the role of galaxy catalogues in candidate vetting.","tokens_in":16491,"tokens_out":9995,"duration_ms":107054,"significance":"If the central comparison were robust, this would be a practically useful result for planning electromagnetic follow-up in O4/O5, and the paper's use of a realistic public catalogue and telescope-specific tiling is a strength. The authors are transparent in stating the mass-filling assumption and the guaranteed superiority of S_3DMF. The main limitation is that the crossover distance between S_3DPM and S_2DP is evaluated against a ground-truth model that assumes missing mass is uniformly distributed on the sky; because the real missing mass is likely clustered, the headline number is not yet established. The paper also does not quantify the statistical uncertainty of the crossover. With a sensitivity analysis, the work could be a solid contribution.","major_comments":[{"comment":"The evaluation metric for all strategies is p_3D times margin, where margin is constructed by spreading the completeness-corrected missing mass uniformly over each 20 Mpc shell. Since the missing mass carries no directional information in this model, the comparison S_3DPM vs S_2DP at large distances is effectively a test of how much information the incomplete catalogue adds relative to the isotropic fill; the conclusion that S_2DP overtakes S_3DPM beyond roughly 300-400 Mpc is therefore a direct consequence of this assumption. In reality, uncatalogued galaxies trace the same large-scale structure as catalogued ones, so an incomplete catalogue still contains directional information about the missing mass; under a clustered missing-mass model, the crossover could shift to larger distances or disappear. The authors acknowledge the assumption in §4.4 and defer direction-dependent modelling to future work, but the central claim should be tested with an alternative missing-mass distribution (e.g., assigning missing mass in proportion to the angular density of catalogued galaxies, or using a lognormal mock) before the crossover is stated as a robust recommendation.","section":"§4.4 and §5"},{"comment":"Because S_3DMF's ranking weight is exactly the figure of merit p_3D times margin used to score every strategy, the statement in §5 that 'S_3DMF will always give the highest probability coverage' is a mathematical identity. The paper explicitly acknowledges this, but the abstract and discussion still present the few-percent advantage of S_3DMF as a substantive result. For example, Table 2's median coverage of 0.530 for S_3DMF versus 0.504 for S_3DPM is guaranteed by construction. I recommend reframing S_3DMF as the upper envelope of the mass-filling model rather than an independently validated strategy, and making clear that the only non-circular comparison in the paper is S_3DPM versus S_2DP (which is nevertheless scored against the same assumed margin model).","section":"§5 and Tables 2–3"},{"comment":"The crossover distance of roughly 300-400 Mpc is inferred from bin-wise medians of R = p(S_3DPM)/p(S_2DP) and from the colours of scattered points, but no uncertainty or significance level is attached to this number. The sample has 329 events with wide scatter (the violin plots show substantial spread in every distance bin), so the crossover is not well constrained. A confidence interval on the crossover, or a distance-dependent fraction of events in which S_3DPM wins, would make the headline claim in the abstract and §6 quantifiable and would also help assess whether the difference between S_3DPM and S_2DP is practically meaningful given the few-percent size of the effect.","section":"§5.2 and Figure 4"}],"minor_comments":[{"comment":"The text says 'five methods' (e.g., §5, first paragraph) and Table 1 claims 'five follow-up schemes', but only four schemes (S_2DP, S_M, S_3DPM, S_3DMF) are defined in §4 and listed in the table. Please correct the count.","section":"Table 1 and §5"},{"comment":"In the mass-filling prescription, the mass of a voxel is set to the catalogued galaxy mass if the voxel is occupied, and to the mean missing mass if empty; the mean missing mass is computed per shell. This means the total filled mass is not equal to the completeness-corrected total mass, because occupied voxels do not receive their share of the unobserved mass. The authors should clarify whether this is intentional and quantify the effect on the normalisation and ranking.","section":"§4.4"},{"comment":"The statement that the results are 'not highly sensitive' to the 3600 deg^2 99%-area cutoff is not supported by any figure or table. A supplementary test with a different cutoff (e.g., 2000 and 5000 deg^2) would be helpful.","section":"§3.3"},{"comment":"The caption uses the notation S_3DMF before it has been introduced in the text (the inset captions refer to S_3DPM and S_3DMF); consider either defining the abbreviations in the caption or reordering the caption.","section":"Figure 2 caption"},{"comment":"There is a typo in the first paragraph of §2: 'the catalogue should be as as complete as possible' should read 'as complete as possible'.","section":"§2"},{"comment":"In Table 3, the number '0,892' in the S_2DP column for d_GW<302 Mpc uses a comma as the decimal separator, inconsistent with the rest of the table (which uses periods).","section":"Table 3"},{"comment":"The sentence 'This method (hereafter S_2DP), which completely ignores galaxy catalogue information.' is missing a main verb; consider removing the comma and the period, e.g., 'This method (hereafter S_2DP) completely ignores galaxy catalogue information.'","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its assumptions, but the headline crossover result is not yet robust to the isotropic missing-mass assumption. I recommend major revision with an emphasis on sensitivity tests. The comparison of S_3DPM vs S_2DP is the core contribution; the S_3DMF numbers should be explicitly labelled as the model's upper bound. Also, the 'five methods' miscount should be fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead this one if you care about EM follow-up of LVK events. The paper is a clean, careful simulation study: it takes 329 realistic O4 BNS/NSBH events, runs them through GIT and WINTER tiling, and compares four ranking schemes. The headline result is that 3D-probability-times-mass ranking (S_3DPM) beats plain 2D tiling (S_2DP) below roughly 300-400 Mpc, and loses above that, with the crossover set by NED-LVS completeness. That is a genuinely useful, quantitative answer to a practical question that has been discussed qualitatively for a decade.\n\nThe paper earns credit for being explicit about its main assumption. Section 4.4 states that missing galaxy mass is spread uniformly per 20 Mpc shell, and Section 5 states that the mass-filled catalogue is treated as the true mass distribution. The reader's stress-test flags this as a load-bearing weakness, and I agree it is the soft spot. If the missing mass actually traces the same large-scale structure overdensities as the catalogued galaxies, then S_3DPM should degrade less quickly with distance, and the quoted crossover distance is not robust. The direction-dependent treatment is deferred to future work, so this is acknowledged but unresolved.\n\nA related but smaller issue is the circularity for S_3DMF: since the figure of merit is p_3D times the mass-filled weight, S_3DMF is guaranteed to score highest. The authors state this openly, so it is not a hidden flaw, but it does mean the few-percent gain of S_3DMF over S_3DPM is partly by construction. What is not circular is the S_3DPM vs S_2DP comparison, and that comparison drives the main recommendation.\n\nMinor points: the free parameters (99% area cutoff, 100-tile limit, shell width, distance cut) are all reasonable, and the authors note robustness checks. No code is released, which lowers reproducibility, but the method is described in enough detail to reimplement.\n\nOverall: the central claim is honest and the analysis is internally consistent, but the operational recommendation should be treated as conditional on the uniform-missing-mass assumption. A real validation would be testing against a non-uniform missing-mass distribution or a measured counterpart. I would send this to a serious referee, because the question is timely and the authors have done the work carefully; I would ask for a robustness test against clustered missing mass and release of the analysis code before accepting.\n\nRecommended: engage, with revision.","headline":"Useful O4-era simulation showing galaxy-catalogue pointing helps mainly below ~300 Mpc, with a clear but acknowledged circularity in the mass-filling metric and a uniform-sky assumption that may shift the crossover.","tokens_in":17075,"tokens_out":651,"would_cite":true,"duration_ms":9548,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper tests four telescope-pointing strategies for gravitational-wave follow-up and finds that a 3D galaxy-mass-weighted search beats plain 2D sky-map tiling for events closer than about 300–400 Mpc, while the simple 2D probability…","keywords":["gravitational waves","electromagnetic counterparts","galaxy catalogues","follow-up strategies","mass-filling","O4 observing run","multi-messenger astronomy","kilonova"],"falsifier":"Run the same simulated follow-up using a realistic mock galaxy catalogue built from a large-scale structure simulation, with galaxies placed along filaments and voids, as the ground truth, and recompute the per-strategy coverage; if the mass-filling strategy's advantage shrinks or the crossover distance moves, the uniform-filling assumption is the cause.","tokens_in":15979,"feed_emoji":"🔭","tokens_out":7043,"duration_ms":76278,"temperature":0.7,"pith_summary":"The paper asks whether galaxy-catalogue-targeted follow-up of gravitational-wave events is still worthwhile in the LIGO-Virgo-KAGRA O4 run, when localisation volumes contain thousands of galaxies and galaxy catalogues are incomplete beyond a few hundred megaparsecs. Using 329 simulated binary-neutron-star and neutron-star–black-hole events, the authors compare four tile-ranking strategies for small-field telescopes. They find that ranking by 2D sky-map probability alone covers more of the mass-weighted source probability than a 3D galaxy-mass ranking for events beyond about 300 Mpc, while the galaxy-based ranking wins for nearer events. Filling in the unobserved galaxy mass uniformly raises the measured coverage by a few percent and transitions smoothly between the two methods. The practical output is a distance-based recommendation: point at galaxies for nearby events, and fall back to plain sky-map tiling for distant ones.","feed_headline":"Galaxy catalogues win for nearby gravitational waves","feed_subtitle":"Beyond 300 Mpc, plain 2D sky-map tiling recovers more source probability; mass-filling adds a few percent.","key_machinery":"The machinery is a set of tile-ranking weights applied to a fixed healpix grid and evaluated over the top 100 tiles. $S_{2DP}$ ranks by the marginalised 2D GW probability; $S_M$ ranks by the sum of galaxy stellar masses in each tile; $S_{3DPM}$ multiplies each galaxy's mass by the 3D GW probability at its position; and $S_{3DMF}$ replaces galaxy mass with a corrected mass that adds uniformly distributed missing mass per 20 Mpc shell, so each empty voxel carries the mean missing mass of its shell. The evaluation metric is itself the total $p_{3D}\\times\\hat{m}$ covered by the top tiles, which makes the mass-filling model the ground truth that every strategy is measured against.","core_discovery":"The central discovery is a crossover in follow-up strategy driven by catalogue completeness, not by telescope field of view. For events with median distance below about 300 Mpc, weighting telescope tiles by the 3D gravitational-wave probability times galaxy mass ($S_{3DPM}$) gives higher $p_{3D}\\times\\hat{m}$ coverage than tiling by the marginalised 2D probability ($S_{2DP}$). Beyond 300–400 Mpc, where the NED-LVS galaxy catalogue is only about 70 percent complete, the plain 2D probability search outperforms the galaxy-based search by a few percent. A mass-filling scheme ($S_{3DMF}$) that distributes the missing galaxy mass uniformly within 20 Mpc shells achieves the highest coverage of all, a few percent above the best conventional method, and approaches the 2D-probability map at large distances where most mass is uncatalogued. A mass-only ranking ($S_M$) performs poorly except for nearby, well-localised events. The same qualitative trends hold for a wider-field telescope, with all ratios closer to unity.","pith_inferences":["If the missing mass is in reality clustered along filaments, a direction-dependent mass-filling scheme would likely make galaxy-based strategies perform better at larger distances than the uniform-filling assumption suggests, and the 300 Mpc crossover would shift.","The same methodology could be applied to future observing runs, where poorer localisation and larger distances will push more events into the regime where the 2D probability search is preferred.","Comparing strategies against a realistic large-scale structure mock catalogue would provide an independent ground truth that avoids evaluating each method with its own assumption.","The most massive galaxies are nearly 100 percent complete to about 450 Mpc, so a mass-weighted catalogue search may remain effective for the hosts most likely to harbour mergers even when total catalogue completeness is modest."],"forward_implications":["For O4-era follow-up with small-field telescopes, observers should use galaxy catalogues for events with median distance below about 300 Mpc and plain 2D sky-map tiling beyond that.","Mass-filling the catalogue recovers a few percent more probability coverage than either conventional method, making it a low-cost improvement worth including in scheduling.","The crossover distance is set by catalogue completeness, so as deeper catalogues become available the galaxy-based strategies will remain competitive at larger distances.","Galaxy catalogues remain essential for vetting candidate transients found in tiled surveys, even when the tiling itself does not use them."],"supporting_citations":[{"why":"Supplies the NED-LVS galaxy catalogue with masses, distances, and the distance-dependent completeness function used for galaxy selection and mass-filling.","marker":"Cook et al. 2023"},{"why":"Provides the simulated O4/O5 gravitational-wave event sample with localisation areas and distances used throughout the study.","marker":"Kiendrebeogo et al. 2023"},{"why":"Establishes the earlier galaxy-targeting result that a few tens of galaxies cover about half the source probability, the baseline this paper contrasts with for larger localisation volumes.","marker":"Gehrels et al. 2016"},{"why":"Showed that galaxy catalogues can shrink electromagnetic search regions by orders of magnitude, motivating the galaxy-targeted approach.","marker":"Nissanke et al. 2013"},{"why":"Simulated galaxy-catalogue follow-up and highlighted catalogue incompleteness as a key limitation.","marker":"Kasliwal & Nissanke 2014"},{"why":"Concluded that even incomplete catalogues outperform blind searches, providing the rationale for mass-filling incomplete catalogues.","marker":"Hanna et al. 2014"},{"why":"Simulated merger rates as a function of galaxy mass, supporting the mass-proportional prior used in the ranking weights.","marker":"Mapelli et al. 2018"},{"why":"Simulations showing merger rates correlate with galaxy mass and star formation rate, supporting the mass weighting.","marker":"Artale et al. 2019"},{"why":"Further simulation evidence for the mass dependence of compact-object merger host galaxies.","marker":"Artale et al. 2020"},{"why":"Discusses alternative approaches for accounting for missing galaxy mass in gravitational-wave follow-up, framing the mass-filling choice.","marker":"Finke et al. 2021"}],"fun_headline_variants":["Galaxy catalogues best under 300 Mpc","2D sky maps beat galaxies far away","Distance decides: galaxies or sky maps?","GW follow-up: galaxies win nearby","Mass-filling boosts GW follow-up"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Coverage is measured against the paper's own mass-filling model, which assumes the missing galaxy mass is spread uniformly across the sky within every 20 Mpc shell; if missing galaxies are actually clustered, the ranking of the strategies—and the 300–400 Mpc crossover—could change.","fun_headline_variants_meta":{"raw":{"variants":["Galaxy catalogues best under 300 Mpc","2D sky maps beat galaxies far away","Distance decides: galaxies or sky maps?","GW follow-up: galaxies win nearby","Mass-filling boosts GW follow-up"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1398,"prompt_tokens":1142,"completion_tokens":256,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":758,"completion_tokens_details":{"reasoning_tokens":190}},"tokens_in":758,"tokens_out":256,"duration_ms":3103,"temperature":1.0,"reasoning_tokens":190,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:05:22.052607+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same simulated follow-up using a realistic mock galaxy catalogue built from a large-scale structure simulation, with galaxies placed along filaments and voids, as the ground truth, and recompute the per-strategy coverage; if the mass-filling strategy's advantage shrinks or the crossover distance moves, the uniform-filling assumption is the cause.","supporting_citations":[{"cited_title":"O., Mazzarella, J","cited_arxiv_id":null,"evidence_quote":"Supplies the NED-LVS galaxy catalogue with masses, distances, and the distance-dependent completeness function used for galaxy selection and mass-filling."},{"cited_title":"W., Farah, A","cited_arxiv_id":null,"evidence_quote":"Provides the simulated O4/O5 gravitational-wave event sample with localisation areas and distances used throughout the study."},{"cited_title":"K., Kanner, J.,et al","cited_arxiv_id":null,"evidence_quote":"Establishes the earlier galaxy-targeting result that a few tens of galaxies cover about half the source probability, the baseline this paper contrasts with for larger localisation volumes."},{"cited_title":"2013, The Astrophysical Journal, 767, 124","cited_arxiv_id":null,"evidence_quote":"Showed that galaxy catalogues can shrink electromagnetic search regions by orders of magnitude, motivating the galaxy-targeted approach."},{"cited_title":"M., & Nissanke, S","cited_arxiv_id":null,"evidence_quote":"Simulated galaxy-catalogue follow-up and highlighted catalogue incompleteness as a key limitation."},{"cited_title":"2014, The As- trophysical Journal, 784, 8","cited_arxiv_id":null,"evidence_quote":"Concluded that even incomplete catalogues outperform blind searches, providing the rationale for mass-filling incomplete catalogues."},{"cited_title":"2018, Monthly Notices of the Royal Astronomical Society","cited_arxiv_id":null,"evidence_quote":"Simulated merger rates as a function of galaxy mass, supporting the mass-proportional prior used in the ranking weights."},{"cited_title":"C., Mapelli, M., Giacobbo, N.,et al","cited_arxiv_id":null,"evidence_quote":"Simulations showing merger rates correlate with galaxy mass and star formation rate, supporting the mass weighting."},{"cited_title":"C., Mapelli, M., Bouffanais, Y .,et al","cited_arxiv_id":null,"evidence_quote":"Further simulation evidence for the mass dependence of compact-object merger host galaxies."}],"review_version":1}