{"id":"a0d580aa-3668-46a6-8ce2-cc14fa206c8c","arxiv_id":"1908.01206","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Two of three NICER-detected burst oscillations in 4U 1728-34 have extremely large tail amplitudes around 46-48 percent and appear only above 6 keV, unlike all previous observations of this source.","lead":"NICER caught three X-ray bursts from the neutron star 4U 1728-34 with pulsations, and two of them oscillated about five times stronger than any previously seen burst-tail signal from this source. These large oscillations appeared only in high-energy X-rays, a combination current models struggle to explain.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 46–48% amplitudes are measured in power-maximizing intervals; selection bias of ~30% in amplitude units can reduce the true amplitudes to ~15–20%, so the headline may be an artifact.","rationale":"The reader's weakest assumption is correct and is the load-bearing point. My extreme-value calculation shows the effect is not negligible: for burst 4, the expected noise-only maximum over ~5620 trials is ~17 in Leahy power, which alone corresponds to an rms amplitude of ~33%; the selected 48% therefore contains a large upward bias. Equivalently, the observed maximum power is fully consistent with a true amplitude near 19% (burst 4) or 16% (burst 7), i.e., near the historical tail-amplitude ceiling of ~15%. This does not question the reality of the oscillations: the Monte Carlo significance estimate is a genuine strength, and the third burst's 7.7% amplitude is a useful cross-check. It does mean the paper's headline claim of unprecedented 46–48% tail amplitudes, and the theoretical discussion built on it, is not currently supported. The fix is straightforward: provide a bias-corrected amplitude estimate, e.g., from injection/recovery through the same search, or report the amplitude from a pre-specified interval/band. The reader's CONDITIONAL verdict already asks for this; my stress test confirms and sharpens the request, so no verdict change is needed.","tokens_in":10707,"tokens_out":16285,"duration_ms":161723,"concrete_test":"Run the authors' Monte Carlo search pipeline on simulated versions of bursts 4 and 7 with the same count rates and durations, injecting a sinusoid at the observed frequency with true rms amplitudes of 10%, 15%, 20%, 25%, and 30%. For each injection, identify the maximum-power 2–4 s interval and energy band exactly as in §2.1.1–2.1.3 and record the recovered amplitude. If a true amplitude of 15–20% recovers 46–48% on average, the reported large amplitudes are a selection artifact and the central claim fails; if only r_true ≥ 30% recovers them, the claim survives with a downward-revised amplitude.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1.1 states that the search attempted to maximize the power by varying the search parameters, and Section 2.1.3 measures the amplitude in exactly that power-maximizing interval and energy band. The quoted ±9% errors are fit uncertainties from the selected window and do not include the selection bias. For a Leahy power spectrum with N counts, a sinusoid of true rms amplitude r gives P ≈ 2 + N r^2, so the reported values simply restate the selected power: for burst 4, P=37.7 and N=153 give r=48%; for burst 7, P=32.2 and N≈135 give r=46%. Because P is the maximum of ~5620 trials, however, it contains a large positive noise contribution. Under the null the expected maximum Leahy power is about 2 ln(5620) ≈ 17, corresponding to an amplitude bias of √(17/153) ≈ 33%. Using the noncentral maximum distribution, a true amplitude of only ~19% (burst 4) or ~16% (burst 7) is enough to make the observed maximum power typical. These values are consistent with the previously observed ≲15% tail amplitudes, so the claim that the amplitudes are several times larger than any previous tail detection is not supported by the current analysis. The Monte Carlo study in §2.1.2 validates detection significance, not the amplitude estimator; Equations (1)–(3) are used only for upper limits in non-detected bands. No bias-corrected amplitude is provided.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports NICER observations of seven type I X-ray bursts from 4U 1728-34 and searches for burst oscillations using Leahy-normalized dynamical power spectra with overlapping 2 s windows, three broad energy bands, and an additional set of finer energy cuts and window trials. Oscillations are found in three bursts. Burst 6 is described as ordinary, with a fractional rms amplitude of 7.7 ± 1.5% in 0.3-6.2 keV, consistent with prior RXTE tail measurements. Bursts 4 and 7 are the focus of the paper: they show tail oscillations at ~362.5 Hz and ~363.7 Hz with reported fractional rms amplitudes of 48 ± 9% and 46 ± 9%, detected only in hard energy bands (6.2-9.9 keV and 6-12 keV). The authors use Monte Carlo simulations to account for correlated search trials and present a joint analysis indicating that three of seven bursts would rarely all show at least 3 sigma signals by chance. The discussion argues that existing cooling-wake and surface-mode models cannot easily explain the reported large hard-band tail amplitudes.","tokens_in":10980,"tokens_out":10370,"duration_ms":105836,"significance":"If the reported amplitudes and hard-band-only nature are correct, this would be a noteworthy discovery: previous tail oscillations in 4U 1728-34 had amplitudes below about 15%, and such large amplitudes are rare in burst tails. The paper is careful in estimating detection significances with correlated trials, makes use of public NICER data, and provides a plausible statistical case that at least some of the three detections are real. However, the two headline amplitudes are measured from the same intervals and energy bands that were selected because they maximized the search power. The quoted errors are only fit uncertainties conditional on the selected window and do not account for this selection bias. The paper's most important quantitative claim is therefore not currently supported, and the theoretical discussion is built on those uncorrected values. The result is potentially important but needs a bias-corrected amplitude measurement before it can be accepted.","major_comments":[{"comment":"The amplitude measurements in Table 2 are not independent of the detection search. Section 2.1.1 states that after finding the highest peak the authors 'attempted to maximize the power by varying the search parameters,' and Section 2.1.3 then measures the amplitude in exactly the band and interval that maximized the Leahy power. For burst 4, the reported amplitude is essentially r = sqrt((P-2)/N) = sqrt((37.7-2)/153) = 48%; for burst 7, r = sqrt((32.2-2)/135) = 46%. Because P is the maximum of about 5620 trials, it contains a positive noise contribution. Under the null hypothesis, the expected maximum Leahy power is about 2 ln(5620) ≈ 17, which alone would contribute sqrt(17/153) ≈ 33% to the rms amplitude for the 153-count burst 4 profile. The ±9% uncertainties are conditional on the chosen window and do not include this selection effect. The paper should provide a selection-bias-corrected amplitude estimate, for example by injecting sinusoids of known amplitude into the full Monte Carlo search and measuring the recovered maximum amplitude, and should not compare the raw selected-window amplitudes with literature values obtained in fixed bands and intervals.","section":"§2.1.1 and §2.1.3, Table 2"},{"comment":"The Monte Carlo simulations validate the false-positive rate of the detection pipeline, but they do not validate the amplitude estimator. The simulations reproduce the search procedure under the null hypothesis and are used to assess the rate at which a given single-trial probability is achieved by chance. To support the headline amplitudes, the same simulation infrastructure should be used with injected sinusoidal signals of known rms amplitude, running the full search and comparing the injected value with the maximum recovered amplitude. Without this, the paper has no way to quantify the upward bias in the 48% and 46% values.","section":"§2.1.2"},{"comment":"The interpretation sections treat the 48% and 46% amplitudes as established facts. For example, the discussion states that 'one would need a large temperature contrast on the surface of the star that is confined in a small region' and that canonical cooling-wake models 'cannot produce large enough temperature asymmetries to explain such large amplitudes.' These conclusions are directly built on the uncorrected, selection-maximized amplitude values. If the bias-corrected amplitudes are around 20% or lower, the claimed tension with previous tail amplitudes and with theoretical models largely disappears. The discussion should be rewritten to be conditional on the re-measured amplitudes.","section":"§3 and Abstract"},{"comment":"The claim that the oscillations are 'detected only at photon energies above 6 keV' also requires trial-aware treatment. The energy band was chosen as part of the search, so the hard/soft contrast is subject to the same selection effect. For burst 7, no upper limit in the 0.3-6 keV band is reported; for burst 4, the soft-band upper limits are quoted at 99% but are not corrected for the number of energy cuts that were tried. A quantitative comparison of hard and soft amplitudes after selection correction, and ideally a statement of the probability of obtaining the observed hard/soft contrast under the search procedure, should be included.","section":"§2.1.1, §2.1.3, Discussion"}],"minor_comments":[{"comment":"The notation f_n(Ps : Pm) = 1 - f_n(Pm : Ps) is confusing: the left-hand side is being used as a confidence function rather than a probability density, and the meaning of the colon notation should be defined explicitly.","section":"Eq. (3)"},{"comment":"The counting of trials is ambiguous: the text says '10 energy cuts (10×10 = 100 extra trials)' and then adds 370 trials to reach 5620. The logic of the multiplication and the decomposition of the 370 extra trials should be spelled out.","section":"§2.1.1"},{"comment":"The column header 'Chance Probability (Single Trial)' should be clarified, since the text distinguishes single-trial and all-trial significances; also the text gives all-trial significances only for the three detected bursts, so presenting both in the table would help the reader.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The authors are experienced and the dataset is valuable, but the central amplitude claim is currently subject to a selection bias that is large enough to change the scientific conclusion. The issue is methodological and appears fixable with the existing simulation pipeline; I would encourage a reanalysis with injected signals rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result—two burst-tail oscillations with fractional rms amplitudes of 48±9% and 46±9%, and detection only above 6 keV—is probably overstated. The amplitudes are measured in the exact time intervals and energy bands that were chosen because they maximized the search power. For burst 4, the power is P=37.7 in an interval with only 153 counts, so the amplitude is essentially sqrt((P−2)/N)≈48%. But P is the maximum of thousands of trials, and the expected maximum noise power for that many trials is about 17 in Leahy units. Subtracting that leaves a signal power of around 20, corresponding to an amplitude near 37%; accounting for the noncentral maximum distribution, the observed P is consistent with a true amplitude as low as ~19%. That would put these bursts in the same range as the previously reported <15% tail amplitudes for this source. The paper never applies a bias correction to the detected amplitude; its Equations (1)–(3) are used only for upper limits in non-detected bands.\n\nThe paper does several things right. It is the first NICER timing study of bursts from 4U 1728–34, the search procedure is standard, and the Monte Carlo simulations that handle correlated trials are a genuine strength. The authors are also honest about the modest detection significance and the source-state ambiguity. Burst 6's 7.7±1.5% rms amplitude in the soft band is a clean, ordinary detection and provides a useful consistency check.\n\nThe problem is load-bearing. The abstract and title emphasize 'large amplitude' and 'unusual.' If the true amplitudes are ~20% or less, the only remaining unusual feature is the hard-only detection, but that too is a selection from the same power-maximizing search. Combined with the small count (153 events for burst 4) and the all-trial significances—which the authors' own Monte Carlo reduces further—the evidence for a new phenomenon is thin.\n\nThis paper deserves peer review rather than desk rejection: the data are real, the methods are mostly sound, and a serious referee can request a bias-corrected amplitude estimate (for example, using the maximum distribution of the noncentral chi-square or a bootstrap) before the claim is accepted. But in its current form, the headline numbers should not be taken at face value.","headline":"The 46–48% tail amplitudes are likely a search-selection artifact; the true amplitudes could be ~15–20%, so the paper's central 'unusual' claim is not supported.","tokens_in":11570,"tokens_out":17819,"would_cite":false,"duration_ms":178652,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"NICER observations of the neutron star 4U 1728–34 reveal burst-tail oscillations with fractional rms amplitudes of 48% and 46%, far exceeding previous measurements and defying current theoretical models.","keywords":["X-ray bursts","burst oscillations","neutron stars","4U 1728-34","NICER","thermonuclear bursts","cooling wake","accreting millisecond pulsars"],"falsifier":"Re-run the search on the same data using a fixed time interval and energy band chosen independently of the detection (e.g., a pre-specified 4 s window and 6–12 keV band), and compare the resulting amplitude; if it drops to the ~10–15% level, the extreme amplitudes are selection artifacts. Alternatively, a Poisson simulation of a burst with a true 10% oscillation, searched with the same trial maximization, that produces a 46–48% amplitude in a 153-count interval would falsify the claim that such amplitudes require new physics.","tokens_in":10512,"feed_emoji":"🔭","tokens_out":4332,"duration_ms":36605,"temperature":0.7,"pith_summary":"The paper reports NICER observations of seven thermonuclear X-ray bursts from the neutron star 4U 1728–34, and finds burst oscillations in the decaying tails of three of them. Two of those oscillations have fractional rms amplitudes of $48\\pm9\\%$ and $46\\pm9\\%$ — several times larger than any previously measured tail oscillation from this source — and they appear only at photon energies above 6 keV. The third is a normal detection at $7.7\\pm1.5\\%$ in the soft band. The authors argue that these large, hard-band-only tail oscillations are difficult to reconcile with current cooling-wake and surface-mode models, and suggest that strongly anisotropic beaming or burst-induced localized accretion might be at work. If confirmed, the result would force a revision of how burst-tail oscillations are produced.","feed_headline":"Burst-tail oscillations hit 48% rms, defying models","feed_subtitle":"NICER sees 46–48% rms hard-band-only oscillations in 4U 1728's burst tails, several times past limits.","key_machinery":"The analysis uses Leahy-normalized dynamical power spectra computed on overlapping 2-second intervals in several energy bands, a search over 360–365 Hz with Monte Carlo verification of trial corrections, and folded pulse profiles fit with a sinusoid $A+B\\sin(2\\pi\\nu t-\\phi_0)$ to extract the fractional rms amplitude $|B|/(\\sqrt{2}A)$. The amplitude measurement in the maximized band is the quantity that carries the argument; it is compared against previous RXTE measurements and against predictions of cooling-wake and surface-mode models.","core_discovery":"The central claim is that two bursts (burst 4 and burst 7) observed by NICER exhibit coherent ~362.5–363.7 Hz oscillations in their decaying tails with fractional rms amplitudes of $48\\pm9\\%$ and $46\\pm9\\%$, detected only above 6 keV, while a third burst (burst 6) shows a normal $7.7\\pm1.5\\%$ oscillation below 6.2 keV. These amplitudes exceed the ~15% maximum previously seen in 4U 1728 tails and exceed the ~10% typical tail amplitudes. The authors show that standard cooling-wake models and low-amplitude surface modes cannot produce such large modulations, and they propose that strongly anisotropic beaming or a burst-triggered localized accretion event might explain the hard-band, late-tail pulsations.","pith_inferences":["If selection bias were the whole story, one would expect loudest-bin amplitudes to scatter around the true value; the fact that two independent bursts show ~47% while a third shows 7.7% hints at a real bimodality, but a dedicated Monte Carlo simulation of the search procedure over simulated bursts with a realistic 10% signal would settle the bias question.","A testable extension: search for similar high-amplitude, high-energy tail oscillations in other bursting LMXBs observed by NICER; if the phenomenon is generic, it would implicate a beaming or geometry mechanism rather than a cooling asymmetry.","The idea that bursts trigger localized accretion infall could be checked by looking for changes in the pulsed amplitude of known accreting millisecond pulsars following bursts in the same system."],"forward_implications":["If the amplitudes are real, tail oscillations can reach ~50% rms, more than three times the largest previously reported from 4U 1728 tails.","The hard-band-only detection means any model must produce a modulation that is suppressed below 6 keV, which existing oscillation models do not naturally predict.","The late-tail appearance suggests a connection to the persistent emission rather than pure burst-surface cooling, possibly linking burst oscillations to accretion-driven pulsations.","The source already shows normal (~8%) and extreme (~47%) tail oscillations in different bursts, so the mechanism must be burst- or state-dependent rather than a fixed stellar property."],"supporting_citations":[{"why":"Establishes the previous maximum tail amplitude of less than 15% for 4U 1728, the baseline that the new 48% and 46% measurements exceed.","marker":"van Straaten et al. (2001)"},{"why":"Provides the typical burst oscillation tail amplitude of about 10% and the sample context that makes the new large amplitudes unusual.","marker":"Galloway et al. (2008)"},{"why":"First detection of burst oscillations in 4U 1728 and the basis for the rotational-modulation interpretation and the 360–365 Hz search window.","marker":"Strohmayer et al. (1996)"},{"why":"Introduces the cooling-wake model for tail oscillations that the new hard-band large-amplitude detections challenge.","marker":"Cumming & Bildsten (2000)"},{"why":"Develops the thermal-wind vortex variant of the cooling-wake model, another mechanism the paper argues cannot produce the observed amplitudes.","marker":"Spitkovsky et al. (2002)"},{"why":"Presents the asymmetric-cooling phenomenological model that the authors explicitly compare against and find insufficient for the new amplitudes.","marker":"Mahmoodifar & Strohmayer (2016)"},{"why":"Provide the method for converting measured power into amplitude upper limits, used to set the soft-band limits below 7%.","marker":"Groth (1975), Vaughan et al. (1994), Watts et al. (2005)"}],"fun_headline_variants":["NICER sees 48% rms burst oscillations, defying theory","Two bursts show 46–48% rms oscillations only above 6 keV","Burst-tail oscillations at 48% rms challenge theoretical models","Hard-band-only pulsations in 4U 1728 reach 48% rms","NICER finds unusual burst oscillations with 48% rms amplitudes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quoted amplitudes are measured in the time interval and energy band that maximized the search power, and the highest-amplitude burst's measurement rests on a 2-second interval containing only 153 photons, so the 48% value could be an upward fluctuation of the fitted sinusoid rather than the true oscillation amplitude.","fun_headline_variants_meta":{"raw":{"variants":["NICER sees 48% rms burst oscillations, defying theory","Two bursts show 46–48% rms oscillations only above 6 keV","Burst-tail oscillations at 48% rms challenge theoretical models","Hard-band-only pulsations in 4U 1728 reach 48% rms","NICER finds unusual burst oscillations with 48% rms amplitudes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00061,"raw_usage":{"total_tokens":2828,"prompt_tokens":919,"completion_tokens":1909,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":1807}},"tokens_in":535,"tokens_out":1909,"duration_ms":15340,"temperature":1.0,"reasoning_tokens":1807,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:19:40.987184+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the search on the same data using a fixed time interval and energy band chosen independently of the detection (e.g., a pre-specified 4 s window and 6–12 keV band), and compare the resulting amplitude; if it drops to the ~10–15% level, the extreme amplitudes are selection artifacts. Alternatively, a Poisson simulation of a burst with a true 10% oscillation, searched with the same trial maximization, that produces a 46–48% amplitude in a 153-count interval would falsify the claim that such amplitudes require new physics.","supporting_citations":[{"cited_title":"2016, ApJ, 818, 93","cited_arxiv_id":null,"evidence_quote":"Presents the asymmetric-cooling phenomenological model that the authors explicitly compare against and find insufficient for the new amplitudes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provide the method for converting measured power into amplitude upper limits, used to set the soft-band limits below 7%."}],"review_version":1}