{"id":"dac3005e-4454-4f8c-8d9d-e19ce6ef3a10","arxiv_id":"1908.06050","paper_version":4,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A mock data challenge shows that gravitational wave standard sirens can recover an unbiased Hubble constant with galaxy catalogs as incomplete as 25%, using the gwcosmo Bayesian pipeline.","lead":"Gravitational wave detections can measure the Hubble constant independently of the cosmic distance ladder by using the signal's distance and a galaxy catalog for redshift. This paper runs mock analyses with simulated binary neutron star mergers to show that the method stays unbiased even when galaxy catalogs miss many true host galaxies.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unbiasedness claim rests on exactly-known EM selection function; the Section IV F robustness assertion is not supported by displayed evidence.","rationale":"Within the paper's explicitly simplified mock scenarios, the Bayesian derivation is coherent and every MDA recovers H0 consistent with the injected 70 km/s/Mpc, so I do not find a load-bearing mathematical error in the central methodology. The reader's weakest-assumption identification is the same one I would make: the incompleteness correction is exact only because the luminosity function and magnitude limit are taken as known. My stress-test pass sharpens this by pointing to Section IV F, where the authors assert robustness to misspecification of these quantities without presenting the supporting results. Because this robustness claim is the only bridge from the idealized mocks to realistic catalogs, it is the most load-bearing gap in the paper. However, the gap does not overturn the paper's central claim within its declared scope; it reinforces the reader's CONDITIONAL verdict. I therefore recommend no change to the verdict, with the condition that the robustness of the selection-function treatment be demonstrated or the claims be restricted to the exactly-known-selection-function scenario.","tokens_in":27314,"tokens_out":13318,"duration_ms":140532,"concrete_test":"Re-run MDA2b (50% number completeness) and MDA3 (50% luminosity completeness) with the same 249 events, but construct the mock galaxy catalog using Schechter parameters perturbed by their published measurement uncertainties, e.g. Delta(alpha) ~ 0.1 and Delta(L*) ~ 20%, while the analysis still assumes the fiducial values. Repeat with mth shifted by +/-0.5 mag and with a sky-dependent mth(Omega) of the same rms. If any resulting joint H0 posterior shifts by more than about 0.5 sigma relative to the fiducial MDA posterior, or if the 68.3% HPD intervals exclude the injected 70 km/s/Mpc at a rate above about 30% over multiple mock realizations, then the Section IV F robustness claim fails and the unconditional 'unbiased' conclusion is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of unbiased H0 recovery is established only under the assumptions stated at the start of Section III: the galaxy luminosity function and the apparent-magnitude threshold mth are known exactly, so the incompleteness correction in Eq. (9) and the out-of-catalog integral in Eq. (A.19) are computed with the true selection function. If the Schechter parameters (alpha, L*) or the effective mth differ from the assumed values, as they will for a real survey with a sky-varying depth and uncertain photometric calibration, the out-of-catalog term is misspecified and the demonstrated unbiasedness need not carry over. The paper's response is Section IV F, which asserts without displaying any table, figure, or numerical statement that the final result is insensitive to variations of alpha and L* within their measurement uncertainties and to O(1) variations in mth. This is a load-bearing unsupported assertion because it is the only evidence that the method remains unbiased outside the exactly-known-selection-function regime. The paper itself explicitly concedes that the mocks exclude redshift uncertainties, peculiar velocities, and galaxy clustering, so the published 4.4% precision and 'sufficiently unbiased' conclusion should be read as conditional on these idealized assumptions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the Bayesian framework of the gwcosmo code for estimating H0 from gravitational-wave standard sirens using both direct EM counterparts and galaxy catalogs, with explicit modeling of GW selection effects (detection threshold) and EM selection effects (apparent-magnitude-limited catalogs). The method is validated through staged mock data analyses (MDA0-3) using 249 simulated BNS detections from the First Two Years end-to-end simulation: known host galaxies, a complete catalog, three incomplete magnitude-limited catalogs with number completeness 75%, 50%, and 25%, and a luminosity-weighted catalog with approximately 50% luminosity completeness at 115 Mpc. In every configuration the combined 249-event posterior on H0 contains the injected value of 70 km s−1 Mpc−1, with fractional uncertainties ranging from 1.13% (known hosts) to 4.48% (the most realistic MDA3 weighted case), and the convergence follows the expected 1/sqrt(N) scaling. The authors conclude that the method produces sufficiently unbiased results for the tested numbers of events, while explicitly acknowledging that the mocks exclude redshift uncertainties, peculiar velocities, and galaxy clustering.","tokens_in":27513,"tokens_out":12679,"duration_ms":127042,"significance":"This is a method-validation paper rather than a new physics result, but it is a valuable one: it exercises a coded implementation that has already been used for published LIGO/Virgo standard-siren measurements (Ref. [24]). Its strengths are the end-to-end simulated GW data with full parameter estimation, the staged MDA design in which each level isolates one selection effect, the analytic derivation in Section II and the Appendix, and the fact that all five analysis configurations recover the injected H0 with the expected 1/sqrt(N) convergence. The MDA0 comparison explicitly demonstrates the bias that would arise if GW selection effects were neglected, and the out-of-catalog treatment in Eq. (9) is exercised down to 25% catalog completeness. The 4.4% precision benchmark for roughly 250 O2-like BNS events with a realistic galaxy density and a 50%-luminosity-complete catalog is a useful planning number. The paper is appropriately transparent about its idealized assumptions; the main weakness is an unsupported robustness claim in Section IV F.","major_comments":[{"comment":"The paragraph asserts that variations of the Schechter function parameters alpha and L* within their current measurement uncertainties produce variations in the final result 'small compared to the statistical uncertainties,' and that the results are 'robust against a small O(1) variation' in the threshold mth, but no table, figure, or numerical statement supporting either assertion appears in the manuscript. This matters because the unbiasedness demonstrated in Sections IV C and IV D relies on the out-of-catalog term in Eq. (9) and the integral in Eq. (A.19) being computed with the true EM selection function (luminosity function and mth known exactly, as stated in the opening of Section III). Section IV F is therefore the only displayed evidence that the method remains unbiased when the selection function is misspecified; the authors should either present the underlying robustness results or temper the claim to match what is actually shown.","section":"Section IV F (Limited Robustness Studies)"}],"minor_comments":[{"comment":"The construction of the MDA3 universe should be clarified: the statement that 'half of the original galaxies were denoted as hosts' sits oddly with the MDA1/MDA2 setup in which each of the 50,000 injected events has a corresponding galaxy, and the resulting value of beta after the factor-of-100 density increase is not stated; please specify how host and non-host luminosities were assigned and what value of beta was actually used.","section":"Section III D (MDA3 construction)"},{"comment":"The sentence 'Our dataset also allows us to to assess the convergence' contains a duplicated word, and there are minor typographical artifacts elsewhere (for example, 'Completness fraction' in the Figure 1 captions) that should be cleaned up.","section":"Section IV E (Convergence)"},{"comment":"The dashed posterior obtained when GW selection effects are neglected is shown in Figure 2 but is not quantified; quoting its MAP value and credible interval would make the size of the demonstrated selection bias concrete.","section":"Section IV A (MDA0 results)"},{"comment":"The table reports the MAP and 68.3% HPD interval for each MDA but not the injected value; adding a column indicating whether each interval contains H0 = 70 km s−1 Mpc−1 would make the consistency claim directly readable.","section":"Table II (Summary of results)"},{"comment":"The fitted 1/sqrt(N) convergence coefficient is quoted only for the known-host case (about 18%); listing the fitted coefficients for all MDAs would make the convergence comparison reproducible.","section":"Section IV E (Convergence)"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is effectively the validation paper for the gwcosmo code used in the companion LVC analysis (Ref. [24]). The central recovery results are sound as far as they go, but the robustness claim in Section IV F must be substantiated or removed; the missing evidence presumably exists in the collaboration's internal validation records, and including it would strengthen the paper considerably. I would also double-check the MDA3 host-assignment description in Section III D against the code before publication, since the text is ambiguous about the relationship between the injected-event hosts and the 'half of the original galaxies' statement. The scope of the paper is modest but appropriate for a methods journal; the idealized mocks are acknowledged by the authors themselves. I do not recommend rejection, but the unsupported robustness assertion should not appear in print."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this if you work on GW standard sirens. This is the methods paper behind gwcosmo, and the mock campaign is genuinely useful: end-to-end simulated BNS events from the First Two Years dataset, 249 detected events with real search and parameter estimation, and a ladder of mock catalogs that isolates GW selection, EM selection, and luminosity weighting. All analyses recover the injected H0 = 70; the derivation in Eqs. (6)–(10) and Appendix A is coherent; the 1/sqrt(N) convergence check is a good addition. The headline number — 4.4% precision from ~250 O2-like BNS with a 50%-complete luminosity-weighted catalog — is real, but it carries the paper's own caveats: galaxy density about three times local, no redshift uncertainties, no peculiar velocities, no clustering.\n\nThe soft spot is Section IV F. It asserts that the results are insensitive to variations in Schechter parameters and an O(1) change in mth, but shows no table, figure, or numeric statement. That matters because the unbiasedness result only holds when alpha, L*, and mth are known exactly, as the paper says at the start of Section III. If a real catalog has sky-varying depth or photometric calibration error, the out-of-catalog term in Eq. (9)/A.19 is misspecified, and these mocks do not test that regime. The stress-test concern is fair. The paper itself is honest about the idealizations, so the fix is simply to either display the robustness study or drop the claim.\n\nMinor points: MDA1/2 catalogs are roughly 35 times sparse relative to local galaxy density, so those precision numbers are not forecasts; the authors say this. MDA3 uses the same luminosity-weighting assumption in simulation and in one analysis variant, but this is a controlled validation and they run the unweighted variant as well, so I do not see a circularity problem. Citation practice is honest: they explicitly credit Chen, Fishbach, Holz and Fishbach et al. for the framework and position their contribution as implementation plus end-to-end mock validation.\n\nVerdict: send it to a serious referee. The central claim holds within the simplified scenarios it defines, and the code is already used in LVC H0 results. The referee should ask the authors to substantiate or remove the Section IV F robustness assertion. Either way this is a useful, citable methods reference.","headline":"A solid, honest validation of the gwcosmo catalog-standard-siren pipeline, with the unbiasedness claim correctly confined to exactly-known selection functions — one unsupported robustness paragraph keeps it from being stronger.","tokens_in":28144,"tokens_out":3309,"would_cite":true,"duration_ms":28092,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Bayesian analysis of mocked gravitational-wave events recovers an unbiased Hubble constant from incomplete galaxy catalogs, with 4.4% precision in the most realistic setup.","keywords":["standard sirens","Hubble constant","gravitational-wave cosmology","galaxy catalog method","selection effects","mock data analysis","binary neutron star mergers","luminosity weighting"],"falsifier":"Re-analyze the same 249 simulated events with deliberately mismatched luminosity-function parameters or a sky-varying magnitude limit, and check whether the recovered $H_0$ shifts by more than the statistical width of the posterior; the paper's own robustness checks only vary these inputs within current measurement uncertainties.","tokens_in":27094,"feed_emoji":"🌌","tokens_out":10368,"duration_ms":92163,"temperature":0.7,"pith_summary":"This paper develops and stress-tests a Bayesian method for measuring the Hubble constant from gravitational-wave standard sirens when no electromagnetic counterpart is seen. The core problem is that any realistic galaxy catalog is incomplete, so the true host galaxy of a merger may not be listed. The authors show that by modeling both the catalog's apparent-magnitude cutoff and the detectability of gravitational-wave signals, they can combine about 250 simulated binary neutron star detections and recover an unbiased $H_0=70\\,\\mathrm{km}\\,\\mathrm{s}^{-1}\\,\\mathrm{Mpc}^{-1}$ across a range of catalog completeness levels. In their most realistic mock, with a galaxy density about three times the local value and a catalog containing half the host luminosity out to a reference distance, they reach 4.4% measurement precision. The result matters because a standard-siren measurement is an independent check on the tension between local and early-universe determinations of $H_0$.","feed_headline":"Mock standard sirens recover the Hubble constant to 4.4 percent","feed_subtitle":"With 250 neutron-star mergers and a half-complete galaxy catalog, the recovered Hubble constant stays unbiased.","key_machinery":"The central object is the per-event galaxy-catalog likelihood, which splits the host into two exhaustive cases, in the catalog ($G$) and out of the catalog ($\\bar{G}$): $p(x^{GW}\\mid D^{GW},H_0)=p(x^{GW}\\mid G,\\ldots)p(G\\mid\\ldots)+p(x^{GW}\\mid \\bar{G},\\ldots)p(\\bar{G}\\mid\\ldots)$. The in-catalog term is a sum over catalog galaxies weighted by their redshifts and, optionally, their luminosities; the out-of-catalog term is an integral over galaxies dimmer than the apparent-magnitude threshold, so catalog incompleteness is corrected exactly rather than approximated. A standard power-law-plus-exponential luminosity function generates the galaxy population, and a Monte-Carlo detection efficiency $p(D^{GW}\\mid z,\\Omega,H_0)$ corrects for the fact that nearer and more favorably oriented mergers are preferentially detected. The machinery's job is to make both electromagnetic and gravitational-wave selection effects calculable within one posterior over $H_0$.","core_discovery":"For every simulated data set, the final posterior on $H_0$ contains the injected value of $70\\,\\mathrm{km}\\,\\mathrm{s}^{-1}\\,\\mathrm{Mpc}^{-1}$, even with a catalog that contains only a quarter of host galaxies. The paper's central claim is that combining a detection-efficiency term $p(D^{GW}\\mid H_0)$ with an electromagnetic selection term that separates in-catalog and out-of-catalog hosts removes the bias that magnitude-limited catalogs would otherwise introduce. It also claims that weighting candidate hosts by luminosity improves the $H_0$ constraint by a factor of 1.2 for the tested configuration. In the most realistic mock, with a galaxy density about three times the local value and a catalog that contains half the host luminosity to a reference distance, the combined result reaches 4.4% fractional precision using about 249 binary neutron star detections at second-observing-run sensitivity.","pith_inferences":["The demonstrated unbiasedness should not be expected to carry over to real catalogs automatically, because the mocks assume exact knowledge of the luminosity function and magnitude limit; a sky-varying selection function or a misspecified luminosity function could convert the out-of-catalog term into a bias.","If the luminosity-weighting improvement generalizes, catalog-based standard siren cosmology will benefit from using star-formation or stellar-mass tracers rather than a single luminosity band; this is a testable prediction for future mocks with realistic host-property correlations.","Galaxy clustering, which the mocks deliberately exclude, should help rather than hurt: even when the true host is too faint to be cataloged, a nearby cataloged galaxy can carry redshift information, so real catalogs may perform better than these mock precisions suggest.","A natural stress test is to rerun the same 249 events with photometric redshift errors and peculiar velocities included; the paper leaves that quantification to future work."],"forward_implications":["The method passes every mock data analysis: all combined posteriors are consistent with the simulated $H_0=70\\,\\mathrm{km}\\,\\mathrm{s}^{-1}\\,\\mathrm{Mpc}^{-1}$, so catalog incompleteness alone need not bias standard-siren cosmology once selection is modeled.","Precision degrades smoothly as catalogs become less complete: the fractional $H_0$ uncertainty grows from 1.13% with known hosts to 3.20% with a 25%-complete catalog, and the catalog-based analyses remain unbiased.","Weighting host probability by galaxy luminosity tightens the measurement, improving the fractional uncertainty from 5.31% to 4.48% in the most realistic mock.","The combined posterior converges roughly as $1/\\sqrt{N}$ in event count for all mock types once enough events are included, with less informative catalogs taking longer to reach that scaling."],"supporting_citations":[{"why":"introduces the idea that gravitational-wave luminosity distances combined with galaxy redshifts can measure the Hubble constant.","marker":"[1]"},{"why":"supplies the Bayesian posterior formalism for the galaxy catalog method that this paper extends.","marker":"[20]"},{"why":"provides the completeness-weighted catalog framework and the expected 1/sqrt(N) convergence used for comparison.","marker":"[6]"},{"why":"demonstrates the galaxy catalog standard siren analysis on the first binary neutron star event and motivates the out-of-catalog treatment.","marker":"[22]"},{"why":"provides the end-to-end simulated binary neutron star events from which the 249 mock detections are drawn.","marker":"[30]"},{"why":"provides the companion simulation whose parameter-estimation products are used for consistency checks on the same events.","marker":"[31]"},{"why":"provides the Bayesian parameter-estimation machinery that produces the gravitational-wave posterior samples used in the analysis.","marker":"[34]"},{"why":"supplies the luminosity function used to assign galaxy luminosities and to set host probabilities in the mock catalogs.","marker":"[50]"},{"why":"gives the local galaxy density used to set the MDA3 catalog density and to compare with real catalogs.","marker":"[49]"}],"fun_headline_variants":["Mock standard sirens pin Hubble constant to 4.4%","Unbiased H0 from 250 mock neutron-star mergers","Galaxy catalogs in mock sirens hit 4.4% H0 precision","Sirens and catalogs: robust H0 from mock data","Mock challenge proves standard siren H0 recovery"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire unbiasedness result rests on the assumption that the observer knows the galaxy luminosity function and the catalog's apparent-magnitude cutoff exactly, so the out-of-catalog correction is exactly calculable.","fun_headline_variants_meta":{"raw":{"variants":["Mock standard sirens pin Hubble constant to 4.4%","Unbiased H0 from 250 mock neutron-star mergers","Galaxy catalogs in mock sirens hit 4.4% H0 precision","Sirens and catalogs: robust H0 from mock data","Mock challenge proves standard siren H0 recovery"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000524,"raw_usage":{"total_tokens":2563,"prompt_tokens":1005,"completion_tokens":1558,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":1469}},"tokens_in":621,"tokens_out":1558,"duration_ms":11399,"temperature":1.0,"reasoning_tokens":1469,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:58:36.745190+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyze the same 249 simulated events with deliberately mismatched luminosity-function parameters or a sky-varying magnitude limit, and check whether the recovered $H_0$ shifts by more than the statistical width of the posterior; the paper's own robustness checks only vary these inputs within current measurement uncertainties.","supporting_citations":[{"cited_title":"completeness fraction","cited_arxiv_id":null,"evidence_quote":"introduces the idea that gravitational-wave luminosity distances combined with galaxy redshifts can measure the Hubble constant."},{"cited_title":"Soares-Santos, D","cited_arxiv_id":null,"evidence_quote":"supplies the Bayesian posterior formalism for the galaxy catalog method that this paper extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the Bayesian parameter-estimation machinery that produces the gravitational-wave posterior samples used in the analysis."}],"review_version":1}