{"id":"628f1bc5-85f5-43f3-87b1-1ba0e00f8439","arxiv_id":"2607.21869","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A stacked reverberation-mapping pipeline recovers simulated C IV time lags within 1σ from as few as 2–10 DESI spectroscopic epochs per quasar, but only when the data follow the same DRW/top-hat model assumed by the fitting code.","lead":"This paper tests whether sparse DESI spectra of many quasars, paired with dense ZTF photometry, can be combined to recover average black-hole accretion time delays (reverberation lags) for high-redshift quasars. The mock-data pipeline recovers the simulated lags, suggesting an economical path to measuring quasar distances and black-hole masses at high redshift.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Closed-loop mocks (DRW+top-hat, zero scatter) leave feasibility claim untested against realistic C IV scatter and non-DRW variability.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the mocks are generated from the same DRW + top-hat model that JAVELIN assumes, and the input R–L relation has zero intrinsic scatter. My reading of the paper confirms this is the least secure condition for the central claim. The paper is otherwise careful: it uses realistic DESI/ZTF sampling distributions, performs many sensitivity tests (flux errors, photometric baseline, stacking number, spectral epochs, bin width), includes a CCF-based cross-check, and honestly lists limitations including the closed-loop issue and the unexplained systematic lag underestimation. However, none of these tests break the simulation/recovery model symmetry. The acknowledged zero-scatter assumption (§5.3) and the deferral of misspecification tests (§8.1) mean the feasibility claim is conditional on the mocks being representative. A concrete test that injects observed C IV scatter and alternative variability/transfer-function models would settle whether the pipeline remains unbiased. Since the verdict CONDITIONAL already reflects this conditionality, no change is needed.","tokens_in":42535,"tokens_out":3298,"duration_ms":32573,"concrete_test":"Run the pipeline on the same DESI/ZTF mock suite but draw each quasar's input lag from the Hoormann et al. relation plus a log-normal intrinsic scatter of 0.5 dex (as observed for C IV; Shen et al. 2024), keeping all other simulation and recovery settings identical. Check whether the stacked MAP per bin recovers the mean (or median) input lag and whether the recovered R–L relation remains within 1σ of the input. Repeat with a non-DRW continuum (e.g., a bending-power-law or CARMA(2,1) process) and a non-top-hat transfer function (e.g., Gaussian). If the recovered R–L shifts or the 1σ agreement fails, the feasibility claim is conditional on the DRW/top-hat/no-scatter assumptions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that stacked RM with DESI-like data recovers C IV lags and the R–L relation—is demonstrated only under a closed loop: the mocks are generated with a DRW continuum and a top-hat transfer function (§5.3), the same model JAVELIN assumes in recovery (§6.1), and the input R–L relation has zero intrinsic scatter. The pipeline is therefore self-consistent by construction and does not probe the regime where real quasars violate these assumptions. Real C IV lags show ~0.5 dex intrinsic scatter (Shen et al. 2024), ~10× larger than the simulated ≲0.05 dex within a bin; the authors explicitly defer testing with scatter or non-DRW variability to future work (§5.3, §8.1). The CCF validation in §9 does not break this loop: it uses the same idealized mocks and still yields an R–L relation 4σ from the input (Eq. 10), showing model misspecification can bias recovery. If the low-lag overdensity documented in §8.1/Fig. 21 is amplified by realistic scatter, the 96% within-1σ recovery rate and the unbiased R–L fit (Eq. 6) may not survive.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Using 250 mock C IV quasar light curves per luminosity–redshift bin, the authors simulate DESI sparse spectroscopy plus ZTF photometry, measure lags with JAVELIN, and stack the individual lag posteriors additively. They report 96% of stacked lags within 1σ and recover the input Hoormann et al. (2019) R–L relation as log R = (0.78±0.035)+(0.51±0.025)log(λL1350/10^44) (Eq. 6). They then map sensitivity to flux errors, photometric baseline, seasonal gaps, stack size, spectral epochs, and bin width, and cross-check with a CCF-based pipeline (Eq. 10). The abstract frames this as evidence that stacked RM with DESI-like data can extend the C IV R–L relation to high redshift.","tokens_in":42869,"tokens_out":6005,"duration_ms":62639,"significance":"Strengths: the work is transparent, builds mocks from DESI/ZTF survey characteristics, supplies public data/code, and conducts systematic parameter studies. The internal validation clearly shows the pipeline can recover lags under its assumed model. However, the feasibility claim is supported only in a closed loop: the mocks share the DRW/top-hat model assumed by JAVELIN and have zero intrinsic scatter, and the paper itself concedes these limitations. The CCF validation is valuable because it exposes a method-dependent 4σ offset (Eq. 10), indicating that realistic deviations could alter the conclusions. If the central claim can be demonstrated under misspecification, the impact is substantial.","major_comments":[{"comment":"The recovery statistics in §7 (96% within 1σ; Eq. 6 vs Eq. 5) are internal-consistency checks: the mocks are generated with DRW continuum and top-hat transfer function (§5.3), the same parametric family JAVELIN assumes (§6.1). The input R–L relation is injected with zero dispersion (§5.3), whereas observed C IV lags show ~0.5 dex intrinsic scatter (Shen et al. 2024), and real C IV may have non-DRW variability, outflows/FeII contamination, or BLR holidays. Since §8.1 explicitly defers these to future work, the paper's central feasibility conclusion is conditional. I request either additional mock runs that inject 0.3–0.5 dex scatter and/or non-DRW variability, or a substantially softened claim of feasibility outside the model assumptions.","section":"5.3, 6.1, 8.1"},{"comment":"The pipeline has a systematic tendency to underestimate lags: the averaged stacked posterior peaks at 55 days versus an expected average of 108 days (Fig. 21), and the CCF pipeline returns an R–L relation 4σ below the input (Eq. 10). The authors attribute this to a low-lag overdensity but do not correct for it or quantify its effect on the fitted slope/intercept in Eq. (6). Because the same bias is present in the main pipeline, the quoted 1σ agreement with Hoormann et al. (2019) may partly reflect compensating errors rather than unbiased recovery. Please quantify and, if possible, model or correct this bias.","section":"8.1, Fig. 21, Eq. 10"},{"comment":"The base run fixes JAVELIN's damping timescale to 700 days, restricts the top-hat width to 3–40 days, and uses a lag prior (0–500 days) that brackets the maximum simulated lag of 383 days (§5.3). Appendix A shows that the recovered lag accuracy is sensitive to the lag prior range. This tuning is acceptable for a self-consistency test, but the paper should state more explicitly how the prior choices would be made in a blind application to DESI data, where such knowledge is unavailable. I would like to see the 0–1000 day prior case included in the R–L fit and discussed.","section":"6.1, App. A"},{"comment":"The stacking procedure is described as 'admittedly... not a mathematically robust way of combining posterior distributions,' and the authors note that the prior-independence assumption is violated by the binning design. The fact that additive stacking works on the mocks is reassuring, but it cannot validate the method for data where the true lags are unknown. Since later papers will apply this to DESI, I recommend replacing or supplementing the additive stack with a hierarchical model (e.g., Brewer & Elliott 2014), or providing a formal justification/simulation study of when additive stacking yields unbiased peaks.","section":"6.2"}],"minor_comments":[{"comment":"The abstract quotes 2–10 spectral epochs; §5.3 gives a maximum of 13. These numbers should be aligned.","section":"Abstract, 5.3"},{"comment":"The example in the text says the quasar has redshift 1.63, but the sample cut is z≥1.67; either the figure or the text is inconsistent.","section":"Fig. 8"},{"comment":"The text says the median number of spectroscopic epochs is 'three days'; the unit should be epochs, not days.","section":"5.3"},{"comment":"'Elveldt et al.' should be 'Eltvedt et al.' to match the reference used elsewhere.","section":"9"},{"comment":"The redshift range is reported inconsistently as 1.48<z<5.2 and 1.67<z<5.23; make this consistent after the ZTF-band cut.","section":"5.1"}],"recommendation":"major_revision","confidential_remarks":"I believe the paper is publishable after major revision: the mock suite is valuable, the parameter study is unusually thorough, and the public data/code are assets. The main risk is overclaiming feasibility given the closed-loop design and the known low-lag bias. I would also ask the editor to check the large number of references to unpublished or in-prep works (FastSpecFit, Eltvedt et al.), since several claims rely on them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read: this is a solid feasibility forecast, not a discovery paper. The authors show that stacking JAVELIN lag posteriors from sparse DESI C IV spectra paired with ZTF photometry recovers the input R-L relation to 1σ in mocks. The result is real within their simulation world; the question is whether that world resembles actual quasars.\n\nWhat’s new: earlier stacked RM used CCFs or zDCFs. Stacking MCMC posteriors, and doing it for a DESI-like cadence with as few as two spectral epochs, is a legitimate methodological step. The sensitivity tests are thorough and sensible: flux errors, photometric baseline, number of quasars stacked, spectral epoch count, bin width. They also ran an independent CCF-based pipeline as a consistency check, which is good practice.\n\nThe soft spot is the closed loop. The mocks are generated with a DRW continuum and a top-hat transfer function, the same model JAVELIN assumes in recovery. The input R-L relation is given zero intrinsic scatter. Under those conditions the pipeline works, but the paper never tests what happens when real quasars violate them—say, with ~0.5 dex C IV scatter, non-DRW variability, outflow-contaminated line profiles, or BLR holidays. The authors acknowledge this openly and defer it to future work, so it isn’t hidden, but it does mean the feasibility claim is conditional on the simulations being representative.\n\nTwo additional things to flag. First, the systematic lag underestimation shown in Section 8.1 is real and not fully explained. The CCF cross-check doesn’t break the loop, and in fact gives an R-L relation 4σ from the input, which is a bit worrying. Second, the lag prior of 0–500 days is tuned to the simulations (max lag 383 days); the paper tests prior width only in an appendix. These are not fatal, but they add to the sense that the proof-of-concept is still fairly benign.\n\nWho should read this? People planning RM programs with large spectroscopic surveys, and anyone trying to extend the R-L relation to high redshift. It earns a serious referee—the methodology is careful and the limitations are stated. But a referee should push for misspecification tests: simulate with realistic scatter, non-DRW variability, and a different transfer function, or at least show the pipeline’s failure modes. As it stands, the paper supports a qualified yes: stacking works if the idealized assumptions hold. Real data will decide.\n\nMy recommendation: send it to peer review, with the expectation of a revision that adds these stress tests or narrows the claims accordingly.\n\nBest,","headline":"A careful mock-based feasibility study that convincingly shows the stacking method works under idealized DRW/top-hat conditions, but leaves the real-world case unproven because the simulations never break those assumptions.","tokens_in":43626,"tokens_out":2667,"would_cite":true,"duration_ms":27670,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Stacked reverberation mapping with sparse spectroscopy recovers high-redshift CIV lags and the radius–luminosity relation from DESI-style observations.","keywords":["stacked reverberation mapping","CIV emission line","radius-luminosity relation","quasar variability","damped random walk","DESI","broad line region","high-redshift quasars"],"falsifier":"Generate mock light curves with intrinsic R–L scatter of about 0.5 dex and variability that deviates from a damped random walk (or use empirical DESI+ZTF light curves with independently known lags), run the stacked pipeline, and check whether the recovered R–L slope and zero-point remain within 1σ of the input; a clear offset or smeared stacked peak would falsify the feasibility claim for realistic populations.","tokens_in":42419,"feed_emoji":"🔭","tokens_out":3783,"duration_ms":39907,"temperature":0.7,"pith_summary":"This paper tries to establish that ensemble reverberation mapping can work when individual quasars have only a handful of spectra, sometimes as few as two, provided the continuum is well sampled photometrically. Using mock light curves built to resemble quasar observations from DESI and ZTF-like photometry, the authors stack per-object lag posteriors from JAVELIN across luminosity–redshift bins. They report that 96% of the stacked lag measurements fall within 1σ of the simulated input lags, and the recovered radius–luminosity relation matches the input CIV relation to 1σ. If this holds, large cosmological surveys can double as reverberation-mapping machines, extending the R–L relation to high redshift and high luminosity without dedicated observing campaigns.","feed_headline":"Stacking sparse spectra recovers CIV quasar lags","feed_subtitle":"A DESI mock study finds the radius–luminosity relation survives with only 2–10 spectra per quasar plus dense photometry.","key_machinery":"The machinery is stacked Bayesian lag inference: each quasar's sparse CIV line light curve and dense photometric continuum light curve are fed to JAVELIN, which models the continuum as a damped random walk and the line response as a top-hat transfer function, producing a per-object lag posterior. These posteriors are additively stacked within equal-population luminosity–redshift bins, and the maximum a posteriori peak of the stacked distribution is the bin's lag. A cross-correlation peak distribution pipeline serves as an independent consistency check.","core_discovery":"The central claim is that additively stacking MCMC lag posteriors recovers average CIV lags for quasar ensembles from light curves with only 2–10 spectroscopic epochs at irregular cadence. The recovered relation log R[days] = (0.78 ± 0.035) + (0.51 ± 0.025) log(λL_1350/10^44) agrees with the input Hoormann et al. (2019) relation to 1σ. Accuracy improves with more quasars per stack up to a plateau near 400, with longer spectroscopic baselines mattering more than extra epochs, and with photometric baselines of roughly 1400 days. The method degrades at high luminosity and high redshift, and shows a systematic tendency to underestimate lags.","pith_inferences":["If real CIV scatter around the R–L relation (about 0.5 dex) and non-DRW variability are added to the mocks, the stacked peak may broaden and the recovered slope could shift; the paper leaves this misspecification untested.","The systematic low-lag overdensity visible in the averaged stacked posterior (MAP near 55 days versus an average input lag near 108 days) may contaminate real R–L fits at the low-luminosity end; separating this numerical bias from physical scatter is a natural next step.","The same stacking pipeline could be applied to MgII or Hβ where DESI's spectral coverage overlaps, giving cross-line consistency checks at high redshift.","The approach likely transfers to other wide-area programs with sparse multi-epoch spectroscopy, since the ≥2-epoch requirement is already satisfied by many existing survey designs."],"forward_implications":["DESI quasars with as few as two spectra can yield a CIV radius–luminosity relation out to z≈5 without new dedicated reverberation campaigns.","The method recovers average lags even when no single quasar's light curve is sufficient for an individual lag detection.","Stacking at least 400 quasars per bin and using photometric baselines of at least 1000 days are recommended design choices for future stacked RM programs.","Extending the spectroscopic baseline improves lag recovery more than adding extra epochs; a few spectra spread over years suffice.","The cross-correlation alternative also recovers lags but shows a stronger systematic underestimation, especially at low luminosity."],"fun_headline_variants":["Stacked reverberation mapping works with only 2–10 spectra per quasar","DESI mock lags recover CIV with sparse spectra and dense photometry","Few spectral epochs, stacked lags: feasible for DESI quasars","Stacked RM passes mock test: 2-10 spectra per quasar enough","CIV lags from sparse spectra: stacked RM feasibility shown"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The mocks assume the same damped random walk continuum and top-hat transfer function that JAVELIN uses to recover lags, and the input radius–luminosity relation has zero intrinsic scatter; real CIV lags show roughly 0.5 dex scatter and can involve outflows or BLR holidays, which would add noise and possibly bias the stacked peak.","fun_headline_variants_meta":{"raw":{"variants":["Stacked reverberation mapping works with only 2–10 spectra per quasar","DESI mock lags recover CIV with sparse spectra and dense photometry","Few spectral epochs, stacked lags: feasible for DESI quasars","Stacked RM passes mock test: 2-10 spectra per quasar enough","CIV lags from sparse spectra: stacked RM feasibility shown"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00076,"raw_usage":{"total_tokens":3271,"prompt_tokens":863,"completion_tokens":2408,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":2323}},"tokens_in":607,"tokens_out":2408,"duration_ms":16490,"temperature":1.0,"reasoning_tokens":2323,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T06:24:07.747487+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate mock light curves with intrinsic R–L scatter of about 0.5 dex and variability that deviates from a damped random walk (or use empirical DESI+ZTF light curves with independently known lags), run the stacked pipeline, and check whether the recovered R–L slope and zero-point remain within 1σ of the input; a clear offset or smeared stacked peak would falsify the feasibility claim for realistic populations.","supporting_citations":[],"review_version":1}