{"id":"43a6cb3c-3560-483e-a8c9-4431cfafb792","arxiv_id":"1908.11395","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"Stacked [C II] data from 26 z~6 quasars show a tentative, sub-3-sigma broad emission component consistent with outflows, but the detection is not statistically significant.","lead":"Using ALMA archival data, the authors stacked the [C II] spectra and cubes of 26 quasars at redshift about 6 to search for faint outflow emission. They find a tentative broad component at 1 to 2.5 sigma significance, strongest for a sub-sample of 12 sources, and conclude deeper observations are needed.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sub-sample is selected on the same wing flux later used as the detection; quoted 1.2-2.5 sigma lacks a trials correction and is not a valid significance.","rationale":"I read the paper in good faith. It is a careful observational study that honestly labels its result as tentative, and it includes useful tests: mock stacks with and without an injected broad component, a test of the velocity-rebinning procedure, an orientation Monte Carlo, and a leave-one-out check showing no single source drives the signal. The full-sample analysis is methodologically sound but yields only sub-1.5 sigma excess, so the positive case rests entirely on the sub-sample claim. The reader's weakest assumption identifies the same load-bearing issue that I find: the max sub-sample is chosen by ranking sources on the integrated flux in the line-free channels of the very stacks that are later used to measure the detection significance. No trials factor for searching 10,000 random sub-samples is applied, and the existing mock tests do not simulate the selection procedure. This is not an internal inconsistency; it is a statistical validity problem. A single concrete simulation of the full selection under the null would settle whether the 1.2-2.5 sigma significance is credible. Given the paper's already cautious framing, the appropriate verdict remains conditional rather than reject: the authors should add a selection-corrected significance or an independent validation before the sub-sample claim is accepted as evidence for outflows.","tokens_in":24490,"tokens_out":4722,"duration_ms":45542,"concrete_test":"Run the full Sect. 5.2 selection pipeline on noise-only mock data: take the best-fit single-component [C II] models of the 26 sources with no injected broad component, add Gaussian noise matching the observed rms as in Sect. 5.3, repeat the 10,000-random-sub-sample ranking and the 3-D stacking many times, and record the significance and FWHM of the resulting max sub-sample in each realization. If a non-negligible fraction of noise-only realizations (e.g., more than 5 percent) produce a max sub-sample with an excess at or above 2.5 sigma and FWHM greater than 775 km/s, the reported detection is a selection artifact. A complementary split-sample check would use half the sources to select the max sub-sample and the other half to measure the excess significance on independent data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central positive claim is the 1.2-2.5 sigma excess in the 3-D stack of the 'max sub-sample' (Sect. 5.2, Table 5). The selection of that sub-sample is the weakest link. In Sect. 5.2, sources are ranked by the integrated flux in the line-free channels (at about 2 x FWHM of the line center) of stacks of 10,000 random sub-samples; the 12 sources with the highest rank form the max sub-sample. The significance in Table 5 is then measured on the same stacked spectra that were used to select the sub-sample. This is post-hoc selection on the dependent variable: the null distribution of the maximum, over 10,000 overlapping random sub-samples, of the wing-integrated flux is not the single-stack noise distribution. The mock test in Sect. 5.3 and Fig. A.1 reports less than 17 percent false detections above 0.4 sigma, but that is for a fixed pre-defined stack, not for the maximization procedure used in Sect. 5.2. Additionally, the line-free channels at 2 x FWHM of a roughly 300-400 km/s narrow line extend to about 600-800 km/s, exactly where the claimed broad component with FWHM greater than 775 km/s has significant flux; the ranking metric is therefore not independent of the detection metric. The Fig. 5 caption also notes that the max sub-sample seems to prefer higher rms values, which is consistent with noise selection. Without a trials-corrected significance or an independent validation of the selection procedure, the 1.2-2.5 sigma excess does not support an outflow interpretation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a spectral-stacking search for broad [C II] outflow components in 26 z~6 quasars observed with ALMA. After velocity-rebinning the [C II] lines to a common width and stacking both extracted 1-D spectra and full 3-D cubes under three weighting schemes, the authors find no outflow signature in individual sources. The full-sample 3-D stack shows excess emission over a single-component fit at 1.1-1.5 sigma, while the 1-D stack from smaller regions shows only 0.4-0.7 sigma. The authors then develop a sub-sampling procedure in which 10,000 random sub-samples are stacked and ranked by the integrated flux in channels at about twice the FWHM of the main line; the 12 sources with the highest ranks form the 'max sub-sample.' The 3-D stack of this sub-sample shows a broad component with FWHM > 775 km/s and excess significance of 1.2-2.5 sigma, from which they estimate an outflow rate of 45 +/- 21 solar masses per year. Mock tests presented in Section 5.3 and Appendix A show that the stacking method can recover injected broad components and that the false-detection rate for a fixed stack is below 17 percent above 0.4 sigma. The authors conclude with evidence suggesting outflows in a sub-sample of 12 of 26 quasars, while stressing that deeper ALMA observations are needed for confirmation.","tokens_in":24827,"tokens_out":5232,"duration_ms":51250,"significance":"If the sub-sample detection were robust, this would be one of the first systematic constraints on fast outflows in z~6 quasar hosts and would support the use of spectral stacking, especially 3-D stacking over extended regions, as a tool for searching for faint broad components. The paper has clear strengths: the archive recalibration is described in detail; the authors explicitly frame the detections as tentative and below conventional thresholds; the comparison with central-pixel-only stacking (Decarli et al. 2018) usefully demonstrates that extended emission matters; and the mock tests, while not covering the selection step, show that the basic stacking and fitting pipeline does not routinely create spurious broad components. However, the central positive claim for the sub-sample is not statistically established because the sub-sample is selected on the same wing flux that is then reported as the detection, without a trials correction.","major_comments":[{"comment":"The sub-sample detection is selected on the same line-free flux that is later reported as the detection. The ranking metric is the integrated flux in channels at 2 x FWHM from the line center of stacked random sub-samples, and the same type of wing excess is then used to compute the significance in Table 5 after subtracting the one-component fit. With typical narrow-line FWHM values of 300-400 km/s, the 'line-free' channels at +/-2 x FWHM are at roughly +/-600-800 km/s, where the claimed broad component with FWHM > 775 km/s has substantial flux; the selection and detection metrics are therefore not independent. The quoted 1.2-2.5 sigma is the maximum over 10,000 random sub-samples rather than a single pre-defined stack, so the relevant null distribution is that of the maximum significance over the full search, not the single-stack noise distribution. I request a trials-corrected significance, for example by running the full sub-sampling and ranking procedure on noise-only mock data and measuring the distribution of the maximum excess, or an independent validation of the selected sub-sample.","section":"Section 5.2 and Table 5"},{"comment":"The mock tests validate the stacking of a fixed, pre-defined sample but do not exercise the sub-sample selection loop of Section 5.2. The statement that fewer than 17 percent of noise-only iterations produce an excess above 0.4 sigma refers to a fixed full-sample stack, not to the procedure that generates 10,000 candidate sub-samples, grades each by its integrated wing flux, and retains the maximum. The false-positive rate of the maximization procedure is expected to be much larger than that of a single stack, so the Appendix A thresholds cannot be used to support the significance quoted for the max sub-sample. I recommend adding a noise-only mock version of the complete Section 5.2 pipeline, including the random sub-sampling, grading, ranking, and re-fitting, and reporting the false-detection rate of the resulting 'max sub-sample' as a function of significance.","section":"Section 5.3 and Appendix A"},{"comment":"The manuscript itself notes in the Figure 5 caption that the max sub-sample 'does seem to prefer higher rms values.' This is concerning because higher-rms spectra contribute noisier stacked wings, so a selection that prefers high rms can produce an apparent broad-wing excess that is purely a noise-selection effect. The authors should quantify this: for example, compare the rms distribution of the selected 12 sources with that of random 12-source draws, and check in the mock samples whether the ranking metric correlates with individual-source rms. Without such a test, the sub-sample selection may be selecting noisy realizations rather than outflows, and the physical interpretation in Section 6.2 is not supported.","section":"Section 5.2 and Figure 5"},{"comment":"The outflow mass-rate estimate of 45 +/- 21 solar masses per year is derived from the broad-component parameters of the max sub-sample, whose statistical significance is not established and whose fitted parameters vary strongly with weighting scheme (e.g., Table 4 gives broad FWHM values from 775 +/- 116 to 1427 +/- 408 km/s depending on weight). The derived rate should be presented as an explicitly conditional estimate, or removed from the results and placed in a clearly marked 'if confirmed' discussion, until the selection and significance issues are resolved. As written, the number could be read as a measurement rather than as an upper-limit illustration.","section":"Section 5.2 and Section 6"}],"minor_comments":[{"comment":"The abstract states that the full-sample detection is only 1.1-1.5 sigma and 'tentative,' but the conclusions state 'we find evidence suggesting the presence of outflows in a sub-sample of 12 out of 26 sources.' Given that the sub-sample significance is not trials-corrected, the conclusions should be reworded to emphasize the tentative nature more strongly, for example by saying the sub-sample stack shows an excess whose significance requires confirmation with a corrected null test.","section":"Abstract and Conclusions"},{"comment":"There is a typo in the sentence about flux calibration: 'Consequently, the the fluxes extrapolated for the sources' should read 'Consequently, the fluxes extrapolated for the sources.' Also, in Table 1 the column header 'wether or not we detect' should be 'whether or not we detect.'","section":"Section 3, first paragraph"},{"comment":"The sentence 'We find excess emission due to a broad component with only a weak significance of < 3 sigma, with 75% of iterations having an excess emission of 1-2 sigma' is ambiguous because it mixes a threshold and a range. It would be clearer to state the full distribution, for example '75% of iterations have a detected excess in the 1-2 sigma range and the remaining iterations have lower or higher significance, with none above 3 sigma.'","section":"Section 5.3"},{"comment":"The phrase 'only < 17% of the time do we retrieve excess emission between > 0.4 sigma' uses contradictory inequalities. The text should state the fraction of iterations with excess significance in the specified range, such as 'only 14-17% of iterations have an excess significance > 0.4 sigma and < 2 sigma.'","section":"Appendix A, Figure A.1 caption"},{"comment":"The discussion of the test of stacking without velocity rebinning is a good control, but it would benefit from stating explicitly that the 2-5% false-positive rate above 2 sigma applies to simple (non-rebinned) stacking and that the rebinned method used in the main analysis has a different (lower) false-positive rate in that test.","section":"Section 6.1, mock tests"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and methodologically useful, but the central sub-sample claim currently rests on a selection procedure that is not properly accounted for in the quoted significance. The issue is reparable within the manuscript's scope: the authors can either run a noise-only simulation of the full Section 5.2 pipeline to obtain a trials-corrected significance, or they can downgrade the sub-sample result to an upper limit or a 'candidate requiring confirmation' and adjust the conclusions and outflow-rate estimate accordingly. The full-sample stacking methodology and the comparison with central-pixel stacking are valuable and should survive a revision. I do not see citation or novelty concerns: the parallel work by Bischetti et al. (2019) is cited and the differences are clearly stated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the take: the paper does a genuinely careful job with the full-sample stacking and the archival ALMA data, but the main positive claim—the 1.2–2.5σ broad-wing excess in a 12-source sub-sample—is not statistically valid as reported. The sub-sample is chosen by ranking sources according to the integrated flux in the line-free channels of the stacked spectra, and then the same stacked spectra are used to measure the significance of the broad component. That is post-hoc selection on the dependent variable. The quoted significance has no trials factor for the 10,000 random sub-samples searched. So the number is inflated, and the Fig. 5 caption even notes the max sub-sample leans toward higher rms, which is what noise selection looks like.\n\nWhat is genuinely new and useful: the 3-D stacking over a 2'' aperture picks up extended wing emission that central-pixel stacking (Decarli et al. 2018) misses. The re-calibration of the ALMA archive—70% of the data needed manual fixes, with flux corrections of 20–45%—is a real service. The mock tests for the 1-D stacking, with and without an injected broad component, are a reasonable sanity check, and the full-sample result is honestly framed as sub-1.5σ.\n\nThe weak spot is specific and addressable. The line-free channels sit at about 2×FWHM of the narrow line, i.e., roughly 600–800 km/s, exactly where the claimed broad component (FWHM > 775 km/s) has significant flux. So the ranking metric is not independent of the detection metric. The mock false-detection rate of <17% applies to a fixed, pre-defined stack, not to the maximization procedure used to build the max sub-sample. They would need to run the full ranking-and-maximization on null mock data and report the distribution of the maximum, or give a trials-corrected significance.\n\nThe paper reads as honest: it calls the result tentative and says deeper observations are needed. That framing is accurate for the full sample. For the sub-sample, 'tentative' is still too strong unless the selection effect is dealt with.\n\nThis is worth sending to a referee. The method development and the full-sample stacking are useful enough to justify serious review, and the sub-sample problem can be fixed. I'd tell the authors to re-run the selection on null mocks and report corrected significances.\n\nWho is this for: high-z outflow and [C II] stackers. I'd probably cite it for the 3-D stacking approach, not for the outflow claim.","headline":"A careful full-sample stacking analysis whose sub-sample outflow claim is undermined by selecting the sub-sample on the same wing flux later used as the detection.","tokens_in":25432,"tokens_out":2298,"would_cite":true,"duration_ms":21374,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.54.Aj","98.62.Nx"],"model":"deepseek-v4-flash","headline":"Stacking 26 z∼6 quasar spectra reveals a tentative broad-line outflow component, strongest for a 12-source sub-sample at 1.2–2.5σ.","keywords":["galaxy evolution","quasars at high redshift","galactic outflows","[C II] 158 μm line","spectral stacking","ALMA observations","AGN feedback","submillimeter galaxies"],"falsifier":"Run the full sub-sample selection on 10,000 noise-only mock samples built from the observed single-component lines: if the same ranking procedure produces \"max sub-samples\" with 1.2–2.5σ broad-component excesses as often as the real data do, the reported signal is a selection-noise artifact. A direct check is to observe the 12 sources with deeper ALMA integrations, where a confirmed broad component with FWHM > 775 km/s in one or more individual spectra would settle whether the outflow is real.","tokens_in":24267,"feed_emoji":"🔭","tokens_out":15402,"duration_ms":121257,"temperature":0.7,"pith_summary":"The paper asks whether fast gas outflows — a key feedback channel in galaxy evolution — are already at work in quasar host galaxies when the universe was under a billion years old (z ≈ 6). Individual [C II] 158 μm spectra of 26 ALMA-observed quasars show no outflow component, so the authors stack both extracted spectra and full spectral cubes to bring up faint broad-line wings. They find a tentative broad component in the full-sample stack (1.1–1.5σ, FWHM > 700 km/s) and a stronger one in the stack of a 12-source sub-sample (1.2–2.5σ, FWHM > 775 km/s) selected by ranking 10,000 random sub-samples. If the signal is real, roughly half of these early quasars are driving gas out of their hosts at a mean rate of about 45 solar masses per year; the authors state that the significance is below the usual detection threshold and that deeper observations are required to confirm it.","feed_headline":"Stacked ALMA spectra hint at outflow gas in 12 z~6 quasars","feed_subtitle":"The broad-line excess reaches 1.2–2.5σ with widths above 775 km/s, pointing to where deeper ALMA follow-up should look.","key_machinery":"The machinery is spectral-line stacking with velocity re-binning, applied both to extracted spectra (1-D) and to full spectral cubes pixel-by-pixel (3-D), and combined with a sub-sample ranking scheme. Because the [C II] lines have different intrinsic widths, each line is re-binned so that every FWHM spans the same number of channels, normalized to the narrowest line (270 km/s); this prevents width variations from manufacturing an artificial broad component, but ties the recovered broad-line width to the main-line width. Stacks are made with three weightings — uniform, inverse-variance (1/σ²_rms), and peak-flux normalized (1/S_peak) — and the broad-component significance is the summed residual within ±800 km/s after subtracting a single-component fit. The sub-sample selector draws 10,000 random sub-samples of 3–25 sources, grades each by the integrated flux in channels at twice the FWHM from line center over radii 0.2–3″, and ranks sources by their accumulated grades to define the 12-source \"max sub-sample\".","core_discovery":"The paper's central claim is that stacking archival ALMA [C II] observations of 26 quasars at z ≈ 6 yields a broad emission component in the stacked line that individual spectra do not show, and that this component is most prominent for a sub-sample of 12 sources. In the 3-D stacks, where emission within about 2″ of the quasar is included, the excess over a single-Gaussian fit is 1.1–1.5σ for the full sample and 1.2–2.5σ for the \"max sub-sample\", with the broad component characterized by FWHM > 775 km/s and an integrated flux above 1.2 Jy km/s; the complementary \"min sub-sample\" shows no broad component. The authors interpret the excess as high-velocity outflowing gas, estimate a mean outflow rate of 45 ± 21 solar masses per year at a radius of 11.5 kpc, and present mock tests arguing that the stacking method itself does not generate the feature; they also state clearly that the detection is tentative, that the broad component could in principle be a gas flow between interacting galaxies rather than an outflow, and that deeper ALMA observations are required to confirm the component in individual spectra.","pith_inferences":["Because the sub-sample ranking uses the same line-free-channel flux that later counts as the detection, the quoted significance should carry a trials correction for the 10,000 sub-samples searched; repeating the full ranking on noise-only mocks would calibrate the true false-alarm rate and would likely lower the reported significance.","The ranking tool could be reused as a pre-selection step: shallow ALMA data on a larger quasar sample could flag the most promising outflow hosts, so that deeper and more expensive multi-tracer follow-up (CO, [O III], X-ray) is spent on the sources most likely to show outflows.","The velocity re-binning makes the recovered broad-line width proportional to the main-line FWHM, so the absolute velocity scale of the outflow is model-dependent; re-stacking a line-width-matched subset without re-binning, or fitting a joint narrow-plus-broad model directly, would test that dependence.","If individual confirmations follow, these outflow rates would provide one of the earliest empirical anchors for AGN feedback prescriptions in simulations of galaxy formation in the first billion years of cosmic time."],"forward_implications":["A broad component with FWHM > 775 km/s and 1.2–2.5σ excess appears in the stacked [C II] line of the 12-source sub-sample for all three weighting schemes, and removing any one source leaves the fitted parameters within 20% of their original values.","The wing emission is spatially extended to about 2″ (roughly 11.5 kpc) and disappears when only the central pixel is stacked, which explains why an earlier central-pixel-only stacking search found nothing.","For the sub-sample, the implied mean outflow rate is 45 ± 21 solar masses per year at outflow velocities of roughly 337–713 km/s, within the range measured for outflows in lower-redshift AGN.","Mock tests indicate the stacking method is not fabricating the feature: with no injected broad component, fewer than 17% of stacked mock iterations show excess above 0.4σ, while with a weak injected component about 75% of iterations recover a 1–2σ excess.","Because outflow orientation is random, the true outflow fraction may exceed the detected one: the paper's Monte-Carlo test recovers a >2σ stacked signal in only 27–50% of iterations even when every source carries a broad component with FWHM above 1000 km/s."],"supporting_citations":[{"why":"Supplies most of the sample (project 2015.1.01115.S), the redshift and line catalogs used for stacking, and the central-pixel null result that this work extends by including extended emission.","marker":"Decarli et al. (2018)"},{"why":"The benchmark [C II] outflow detection at z~6 (J1148+5251) and the assumed C+ abundance and temperature used in the outflow-mass estimate.","marker":"Maiolino et al. (2012)"},{"why":"Follow-up of the J1148+5251 outflow; its broad-to-narrow flux ratio and physical assumptions anchor the comparison and the mass-rate calculation.","marker":"Cicone et al. (2015)"},{"why":"The stacked high-velocity [C II] excess seen in z~5.5 star-forming galaxies, which motivates the search and provides the alternative interacting-gas interpretation.","marker":"Gallerani et al. (2018)"},{"why":"The parallel stacking search over a wider redshift range that also reports a broad component, used as the main comparison for significance and method.","marker":"Bischetti et al. (2019)"},{"why":"Provides the equation converting the broad-component [C II] luminosity into outflowing gas mass.","marker":"Hailey-Dunsheath et al. (2010)"},{"why":"Source of part of the sample (project 2015.1.00606.S) and the photometric properties used to characterize the quasars.","marker":"Willott et al. (2017)"},{"why":"Documents the complex high-velocity structure of PJ308-21, used in the test of whether one source drives the stacked signal.","marker":"Decarli et al. (2017)"}],"fun_headline_variants":["Faint outflow signal emerges in stacked z~6 quasar spectra","Stacking ALMA data reveals tentative outflow in 12 quasars","Broad line excess hints at gas outflows in early quasars","Subtle outflow signature found by stacking 26 quasar spectra","Outflow clue in z~6 quasars: 1.2-2.5σ stacked excess"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The \"max sub-sample\" is chosen by ranking sources on the integrated flux in the line-free channels of the very stacks being measured, and the quoted 1.2–2.5σ significance does not correct for the 10,000 random sub-samples searched — if that ranking metric selects noise rather than genuine outflow emission, the sub-sample detection collapses.","fun_headline_variants_meta":{"raw":{"variants":["Faint outflow signal emerges in stacked z~6 quasar spectra","Stacking ALMA data reveals tentative outflow in 12 quasars","Broad line excess hints at gas outflows in early quasars","Subtle outflow signature found by stacking 26 quasar spectra","Outflow clue in z~6 quasars: 1.2-2.5σ stacked excess"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000737,"raw_usage":{"total_tokens":3411,"prompt_tokens":1180,"completion_tokens":2231,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":796,"completion_tokens_details":{"reasoning_tokens":2134}},"tokens_in":796,"tokens_out":2231,"duration_ms":13919,"temperature":1.0,"reasoning_tokens":2134,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:16:14.770559+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the full sub-sample selection on 10,000 noise-only mock samples built from the observed single-component lines: if the same ranking procedure produces \"max sub-samples\" with 1.2–2.5σ broad-component excesses as often as the real data do, the reported signal is a selection-noise artifact. A direct check is to observe the 12 sources with deeper ALMA integrations, where a confirmed broad component with FWHM > 775 km/s in one or more individual spectra would settle whether the outflow is real.","supporting_citations":[{"cited_title":"2015, A&A, 574, A14","cited_arxiv_id":null,"evidence_quote":"Follow-up of the J1148+5251 outflow; its broad-to-narrow flux ratio and physical assumptions anchor the comparison and the mass-rate calculation."},{"cited_title":"2018, MNRAS, 473, 1909","cited_arxiv_id":null,"evidence_quote":"The stacked high-velocity [C II] excess seen in z~5.5 star-forming galaxies, which motivates the search and provides the alternative interacting-gas interpretation."},{"cited_title":"J., et al","cited_arxiv_id":null,"evidence_quote":"Provides the equation converting the broad-component [C II] luminosity into outflowing gas mass."}],"review_version":1}