{"id":"314ae27f-4735-4dd0-94a2-fb6045da8e84","arxiv_id":"2501.16788","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"In simulated LIGO/Virgo analyses, high-SNR event selection biases the inferred inclination angle and, for nonzero true values, inflates the estimated scalar dipole amplitude Ab1.","lead":"This paper uses computer injections of simulated gravitational-wave signals to show that when only loud events are analyzed, the inferred size of a scalar polarization component can be systematically overestimated. The result warns that tests of general relativity with LIGO/Virgo events need to account for selection effects before claiming evidence for non-GR polarizations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed selection bias is not actually modeled: all injections are analyzed without a detection threshold, so the reported Ab1 overestimation may be a prior-misspecification artifact rather than a demonstrated selection effect.","rationale":"I agree with the reader's weakest_assumption and identify the same load-bearing concern. The paper's construction in Sec. III equates 'event selection' with injecting loud signals at a fixed distance, but never applies a detection threshold or selection function. The reported overestimation of Ab1 is thus a demonstration of prior mismatch for a single loud event, not a demonstrated selection effect. This is a real soft spot because the paper's headline claim explicitly attributes the bias to event selection, and a realistic selection function could change the distance-inclination distribution of analyzed events and therefore the direction or magnitude of the bias. The qualitative mechanism—tensor modes favor cos ι ~ 1, scalar dipole favors cos ι ~ 0, so a biased cos ι can inflate Ab1—is plausible and supported by the pure-polarization results, which is why I do not recommend rejection or a stronger verdict. The concrete test (a population injection with an actual detection threshold) would settle whether the bias is a genuine selection effect or an artifact of the fixed-distance setup. The reader's CONDITIONAL verdict is appropriate; my read does not change it.","tokens_in":16107,"tokens_out":6524,"duration_ms":58979,"concrete_test":"Run a population-level injection study: draw sources from the volume-uniform prior of Table II with isotropic sky location and inclination, set intrinsic parameters to the Table I values, and apply a realistic detection threshold (e.g., network SNR > 12) to define the selected sample. Then perform Bayesian inference on the selected events using the same un-conditioned Table II prior, and measure the posterior median of Ab1 as a function of the injected Ab1. If the upward bias persists, the paper's conclusion is robust to realistic selection; if it vanishes, reverses, or becomes setting-dependent, the fixed-distance injection is the actual source of the claimed bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that selecting high-SNR events biases the inferred scalar dipole amplitude Ab1 upward when the true value is nonzero. But the injection scheme in Sec. III does not implement event selection. Every injection is placed at a fixed distance dL = 204 Mpc, yielding SNRs in 20–100, and then analyzed with a prior uniform in comoving volume (Table II). No detection threshold, ranking statistic, or selection function P(det|θ) is applied; all injections are automatically 'selected'. This is explicitly acknowledged in the text: 'the distributions of these 50 GW signals do not follow the prior distribution shown in Table II.' The resulting bias is therefore the standard Bayesian effect of using a prior that is inconsistent with the true source distance, not a selection effect. In a real search, the posterior for a detected event should be conditioned on detection: p(θ|d, det) ∝ p(θ) P(det|θ) p(d|θ). The paper's setup corresponds to P(det|θ)=1 for all sources. If a realistic selection function (e.g., a network SNR threshold) preferentially removes distant or face-on sources, the effective prior on dL and cos ι changes, and the reported upward bias in Ab1 could be altered or even reversed. Because the abstract and conclusions attribute the bias specifically to event selection, this missing element is load-bearing: the qualitative mechanism (tensor and scalar modes prefer opposite cos ι) is plausible, but the paper has not demonstrated that the bias arises from selection rather than from the arbitrary choice of a fixed-distance injection.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how selecting loud gravitational-wave events can bias parametrized searches for scalar-tensor polarizations. The authors construct a parametrized inspiral waveform containing tensor modes plus scalar dipole and quadrupole modes, with non-GR amplitude parameters Ab1 and Ab2. Using Bilby with nested sampling, they perform injection-recovery runs: pure tensor, pure scalar dipole, and mixed Tensor+Scalar(dipole) and Tensor+Scalar(quadrupole) injections, all at a fixed luminosity distance of 204 Mpc with SNRs in the range 20--100, while adopting a prior uniform in comoving volume. They report that the pure tensor mode favors cos iota ~ 1 while the pure scalar dipole mode favors cos iota ~ 0, and that in the Tensor+Scalar(dipole) model the scalar dipole amplitude Ab1 is overestimated when the true Ab1 is nonzero, except at iota = pi/2, with no false deviation when Ab1 = 0. The quadrupole amplitude is found to be poorly constrained. The paper concludes that event selection can bias scalar-mode amplitude estimation unless selection effects are properly modeled.","tokens_in":16389,"tokens_out":4569,"duration_ms":44607,"significance":"If the central claim is correct, the paper identifies a practically important effect for current and future tests of general relativity: using a volume-weighted distance prior on loud, selected events can bias the inferred scalar dipole amplitude, potentially creating false evidence for non-GR polarizations or exaggerating real ones. The physical mechanism is clearly explained and the pure-polarization inclination bias is supported by a simple analytic scaling argument. The paper also has positive reproducibility features: it uses standard tools (Bilby, dynamic nested sampling, LALSuite TaylorF2), states the sampler settings, and gives explicit injection parameters. The main caveats are that the selection process is not actually simulated, the mixed-polarization statistical evidence is based on only three injections per configuration without significance tests, and the quadrupole bias is asserted rather than demonstrated. These issues do not invalidate the qualitative mechanism but they do limit the strength of the quantitative claims, and they need to be addressed before the paper can be accepted.","major_comments":[{"comment":"The paper's central claim is that event selection biases the scalar dipole amplitude, but no selection function or detection threshold is actually implemented. All injections are placed at a fixed distance dL = 204 Mpc, giving SNRs of 20-100, and are then all analyzed; the paper itself states in Sec. III that \"the distributions of these 50 GW signals do not follow the prior distribution shown in Table II.\" The observed bias is therefore a prior-misspecification effect for loud events, not a demonstrated selection effect, since conditioning on detection would modify the prior by P(det|theta), which is not included. This is load-bearing because the abstract and Sec. V attribute the bias specifically to event selection. The authors should either implement an explicit selection function (e.g., a network SNR threshold applied to a population of injections drawn from the prior) or carefully reframe the results as a prior-mismatch bias for loud events and justify why this captures the selection effect of interest.","section":"Section III, Sec. III and IV"},{"comment":"The mixed-polarization overestimation claim for Ab1 is supported by only three injection realizations per parameter setting, with no significance test or quantitative measure of bias. The figures show medians and 90% credible intervals, but the reader cannot tell how many of the three estimates exceed the true value, whether the posterior mass above the injected value is significant, or how often the 90% intervals exclude the truth. To support the claim that Ab1 is systematically overestimated, the authors should either increase the number of noise realizations or report a per-injection summary statistic such as the fraction of posterior samples above the injected value, and ideally a statement about the distribution of posterior medians across realizations. This is particularly important at iota = pi/2, where the posterior is bimodal (Fig. 3) and a median may be a poor summary.","section":"Section IV B, Figs. 2, 4, 5"},{"comment":"The abstract states that the quadrupole bias \"is expected to occur also for the Tensor+Scalar(quadrupole) model,\" but Sec. IV B concludes that the scalar quadrupole amplitude results are \"uninformative\" and that the errors are constrained by the prior range [-1,1]. The paper does not demonstrate any bias in Ab2, so this expected claim is not supported by the presented simulations. Either add a simulation or analytic argument establishing the quadrupole bias, or soften the abstract and conclusions to state that the quadrupole amplitude is too poorly constrained to detect the bias.","section":"Abstract and Sec. V"},{"comment":"The analytic results p(cos iota) proportional to cos^3 iota and p(sin iota) proportional to sin^3 iota are central to the interpretation, but they are stated with a very brief derivation. The derivation should be written out more explicitly, including the assumption of an exactly degenerate likelihood in the amplitude combination and the marginalization over dL with the volume prior. This would also clarify the conditions under which the formula applies, since the full posterior with a three-detector network and finite SNR will only approximately follow this scaling.","section":"Section IV A, Eqs. (4.1) and (4.2)"}],"minor_comments":[{"comment":"There are several typographical and formatting issues, including inconsistent spacing in \"L VK\" (used for the LIGO-Virgo-KAGRA collaboration) and the notation \"rad-dec set\" in Appendix A, which should be \"ra-dec set.\" A careful proofread for these minor errors is recommended.","section":"Throughout"},{"comment":"The caption says the panels run \"from top right to bottom left,\" but the standard reading (and the layout described in the text) is from top left to bottom right. Please correct the caption.","section":"Fig. 2 caption"},{"comment":"Reference [38] is listed as \"T. Bayes, Philosophical Transactions of the Royal Society of London, 370 (1973)\"; the original paper by Bayes was published in 1763, not 1973. The reference should be corrected.","section":"Reference [38]"},{"comment":"The text says that \"p(d|theta, M) is the likelihood\" in Eq. (3.1), but Eq. (3.4) then introduces the likelihood as proportional to exp(-(1/2) sum ...). This is standard, but the notation could be harmonized so that the likelihood is not defined twice in different forms without comment.","section":"Section III"}],"recommendation":"major_revision","confidential_remarks":"The paper's reliance on the authors' previous waveform model [22,23] is not circular because the bias result is an injection-recovery property independent of whether that model is externally accepted. However, the strength of the central claim is currently disproportionate to the evidence: the absence of a modeled selection function and the tiny number of injections per configuration mean that the quantitative overestimation claim is not yet established. The paper is within the scope of a gravitational-wave data-analysis journal and the topic is timely, but the authors should be asked to either run a proper selection simulation or carefully limit their claims. I would also encourage the authors to release the injection configuration scripts to improve reproducibility, since the paper already uses public software and the community would benefit from having the exact setup."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper is a useful caution with a modest new result, but the headline claim that event selection biases the scalar dipole amplitude is not actually demonstrated as a selection effect. What is demonstrated is standard prior mismatch: inject loud events at 204 Mpc, analyze with a volume-weighted prior, and you get biased inclination recovery and, in the mixed tensor-plus-scalar-dipole case, a biased Ab1. The mechanism is clear and the direction makes sense: tensor modes prefer face-on, scalar dipole prefers edge-on, so when the prior pulls cos iota up, sin iota is underestimated and Ab1 compensates upward. The check that no false deviation appears at Ab1=0 is sensible and reassuring.\n\nWhat is actually new is the extension of the known distance-inclination bias [25-29] to the amplitude estimator for the parametrized scalar-tensor model. That is a fairly direct extension, but it is still useful for people running polarization tests with parametrized waveforms. The pure-polarization inclination histograms are informative, and the paper is honest about the quadrupole case: it explicitly says the results are uninformative and the same bias is expected but not demonstrated.\n\nNow the soft spots, in proportion. First, the setup does not actually simulate selection. All injections are loud; there is no detection threshold, ranking statistic, or P(det|theta). The stress-test note gets this right. The bias would arise for any fixed-distance injection with a mismatched prior, so calling it a selection bias is a framing choice. The quantitative size of the bias for a real detection pipeline is not established. This matters for the abstract and conclusions, though the underlying observation about prior mismatch stands. Second, the mixed-polarization claim rests on only three injections per setting, with no significance test or uncertainty on the medians. For a paper whose title says 'statistical biases,' that is thin. Third, no code or data is shipped. Fourth, the title's plural 'polarizations' overpromises, since the quadrupole claim is explicitly not demonstrated.\n\nThe waveform model is imported from earlier work, which is fine; the bias result does not depend on whether that model is externally validated, and the self-citation is not a problem here.\n\nVerdict: I would send this to a serious referee, but with the expectation that the authors either simulate a real selection function or reframe the claims as prior-mismatch effects, and that they add more injections or error bars. The paper is short, readable, and points at a real subtlety in GR tests. It deserves refereeing, not desk rejection, but it needs revision before it can be treated as a quantitative characterization of selection bias.","headline":"Useful caution about prior mismatch in scalar-tensor polarization searches, but the selection-effect framing overreaches: the bias is shown for loud injections at fixed distance, not for an actual detection selection.","tokens_in":16947,"tokens_out":2158,"would_cite":true,"duration_ms":21180,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["83C35","62F15"],"pacs":["04.30.-w","95.85.Sz"],"model":"deepseek-v4-flash","headline":"Selecting loud gravitational-wave events biases scalar-mode tests of gravity.","keywords":["gravitational waves","scalar-tensor theory","gravitational-wave polarizations","parameter estimation bias","selection effects","Bayesian inference","luminosity distance-inclination degeneracy","scalar dipole mode"],"falsifier":"Re-run the Tensor+Scalar(dipole) injections with a distance prior conditioned on the selection, for example a prior sharply peaked at the injected distance or a truncated signal-to-noise-ratio-based selection prior, and check whether the median recovered $A_{b1}$ returns to the injected value. If the overestimation disappears, the bias is a prior-mismatch effect as claimed; if it persists, the mechanism is different.","tokens_in":1533,"feed_emoji":"🌊","tokens_out":1723,"duration_ms":44018,"temperature":0.7,"pith_summary":"This paper argues that choosing only loud, high-signal-to-noise gravitational-wave events for parameter estimation can bias tests of general relativity that search for scalar polarization modes. Using injection studies with a parametrized scalar-tensor inspiral waveform, it shows that when a scalar dipole component is truly present, standard Bayesian inference with a uniform-in-comoving-volume distance prior overestimates the scalar dipole amplitude $A_{b1}$. The mechanism is that the prior pulls the inferred inclination toward face-on values ($\\cos\\iota\\sim 1$), suppressing $\\sin\\iota$, which the amplitude estimator compensates by inflating $A_{b1}$. The bias disappears when the true scalar amplitude is zero, so the analysis does not create a false scalar detection from noise alone; the same effect is expected for the scalar quadrupole mode, though generic degeneracies make that measurement too noisy to show a clean bias.","feed_headline":"Loud-event selection inflates inferred scalar-mode amplitudes","feed_subtitle":"Standard distance priors push inclination estimates, inflating the scalar dipole amplitude in GW tests.","key_machinery":"The central object is the parametrized scalar-tensor inspiral waveform with amplitude parameters $A_{b1}$ (dipole) and $A_{b2}$ (quadrupole), together with the known luminosity-distance/inclination degeneracy and the uniform-in-comoving-volume prior $p(d_L)\\propto d_L^2$. For pure tensor signals the posterior $p(\\cos\\iota)\\propto\\cos^3\\iota$ favors $\\cos\\iota\\sim1$, while for pure scalar dipole signals $p(\\sin\\iota)\\propto\\sin^3\\iota$ favors $\\cos\\iota\\sim0$. In the mixed model, tensor dominance biases the inferred inclination toward face-on, suppressing $\\sin\\iota$, and the scalar amplitude estimator compensates by overestimating $A_{b1}$.","core_discovery":"The paper claims that the amplitude of scalar dipole radiation is overestimated whenever its true value is nonzero and the event is selected as loud, except when the inclination is exactly $\\pi/2$. This overestimation arises because the tensor-mode-dominated inference, combined with a uniform-in-comoving-volume luminosity-distance prior, drives the posterior of $\\cos\\iota$ toward $1$, underestimating $\\sin\\iota$. Since the scalar dipole waveform amplitude is proportional to $A_{b1}\\sin\\iota$, the inference compensates by raising $A_{b1}$. When $A_{b1}=0$, no false deviation occurs, so the bias is conditional on a real scalar component existing. The same qualitative mechanism is argued to hold for the Tensor+Scalar(quadrupole) model, but the scalar quadrupole amplitude is too poorly constrained by the data for the bias to dominate over statistical error.","pith_inferences":["The same prior-mismatch mechanism should affect any amplitude parameter that is correlated with inclination and luminosity distance, so other parametrized post-Einsteinian searches should be checked for analogous biases.","If real catalogs select events by signal-to-noise ratio, current upper limits on scalar dipole radiation derived from individual loud events may be systematically shifted in the direction of overestimating the scalar amplitude.","A hierarchical population analysis that jointly models the detection probability and the source distribution would provide a direct test of whether this bias persists in actual observational data.","Third-generation detectors with signal-to-noise ratios roughly ten times higher may suppress the bias because the likelihood dominates the prior, but the threshold depends on source orientation and detector network."],"forward_implications":["A single loud gravitational-wave event analyzed without modeling the selection will tend to overreport the scalar dipole amplitude whenever a real scalar dipole component is present.","The bias is absent when the true scalar amplitude is zero, meaning the test does not manufacture false scalar detections from prior mismatch alone.","The inclination angle is recovered more accurately as the scalar amplitude increases, because the scalar mode's distinct angular pattern pulls the posterior back toward the true inclination.","For the scalar quadrupole model, the amplitude measurement is largely uninformative because tensor and scalar quadrupole share the same phase evolution, so statistical error exceeds the selection bias.","Multi-event statistical searches or priors that explicitly model the selection process are proposed as ways to remove the bias."],"supporting_citations":[{"why":"Supplies the parametrized scalar-tensor waveform model with the scalar amplitude parameters $A_{b1}$ and $A_{b2}$ that the injection searches use.","marker":"[23]"},{"why":"Provides the result that a uniform-in-comoving-volume distance prior induces a posterior $p(\\cos\\iota)\\propto\\cos^3\\iota$, the key bias mechanism.","marker":"[27]"},{"why":"Earlier parametrized scalar-tensor polarization search that this paper extends by studying selection effects.","marker":"[22]"},{"why":"The Bilby software used to perform the Bayesian parameter estimation and injection recovery.","marker":"[42]"},{"why":"Provides the nested-sampling settings and prior choices that define the analysis setup.","marker":"[45]"}],"fun_headline_variants":["Loud-event selection inflates scalar dipole amplitude","SNR selection cuts bias scalar amplitude estimates","Selection bias overestimates scalar dipole in loud GW events","How event selection biases scalar-mode amplitude estimates","Loud-event bias inflates scalar dipole in GW tests"],"cache_read_input_tokens":19072,"weakest_assumption_plain":"The injections place all loud events at a fixed distance of 204 Mpc with signal-to-noise ratios between 20 and 100, while the prior assumes a uniform comoving-volume distribution, and the analysis applies no actual detection threshold or selection function, so the reported bias is the consequence of prior mismatch for a loud event rather than a demonstrated selection effect.","fun_headline_variants_meta":{"raw":{"variants":["Loud-event selection inflates scalar dipole amplitude","SNR selection cuts bias scalar amplitude estimates","Selection bias overestimates scalar dipole in loud GW events","How event selection biases scalar-mode amplitude estimates","Loud-event bias inflates scalar dipole in GW tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001004,"raw_usage":{"total_tokens":4208,"prompt_tokens":867,"completion_tokens":3341,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":483,"completion_tokens_details":{"reasoning_tokens":3268}},"tokens_in":483,"tokens_out":3341,"duration_ms":22021,"temperature":1.0,"reasoning_tokens":3268,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T10:41:50.315437+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the Tensor+Scalar(dipole) injections with a distance prior conditioned on the selection, for example a prior sharply peaked at the injected distance or a truncated signal-to-noise-ratio-based selection prior, and check whether the median recovered $A_{b1}$ returns to the injected value. If the overestimation disappears, the bias is a prior-mismatch effect as claimed; if it persists, the mechanism is different.","supporting_citations":[],"review_version":1}