{"id":"2c6ae741-6335-481a-9108-4f7dcf9578bd","arxiv_id":"2412.06885","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Durbin-Watson autocorrelation and variability signal-to-noise are the best predictors of successful quasar lag measurements, and reducing first-year cadence by 40% preserves roughly 76 to 90 percent of recovered lags.","lead":"This study uses 172 high-confidence quasar lag measurements from the SDSS-RM survey to find which light-curve properties make reverberation-mapping lags detectable, and simulates how reducing the first-year observing cadence affects lag recovery. The results suggest future surveys can cut early cadence by about 40% and still recover most lags, which matters because industrial-scale black-hole mass surveys are expensive and need efficient schedules.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 90% cadence-retention estimate is conditioned on the gold sample, which was identified under the dense first-year cadence being thinned; this makes the headline an optimistic bound for general lag recovery, not the conservative bound claimed in Figure 9.","rationale":"The reader's conditional verdict identifies the same core weakness, and I agree. The central claim is useful and largely supported by the data: the Durbin-Watson statistic is a plausible, cheap screening metric, and the cadence simulation uses real light curves and a consistent analysis pipeline rather than injected mocks. The weak point is inferential: the gold sample is not exchangeable with the parent population of lags a future survey would try to recover. Because gold selection was performed on the dense first-year data, thinning the same data and measuring survival rates conditions on a high-survivability subset. The paper's own figure caption calling this conservative is inaccurate and should be corrected. A direct rerun on non-gold significant lags would settle whether the bias is quantitative. Until then, the claim should be reported as 'roughly 90% of gold lags remain significant' and '76-86% remain significant and consistent,' with a caveat that these are upper bounds for the full lag population. The reader's CONDITIONAL verdict remains appropriate; no stronger action is needed because the analysis pipeline itself is reasonable and the issue is one of scope and framing rather than internal inconsistency.","tokens_in":28727,"tokens_out":11228,"duration_ms":128579,"concrete_test":"Run the identical epoch-thinning and PyROA analysis of Section 6.2 on the non-gold but statistically significant lags from Shen et al. (2024), or on all lags passing the Section 3.2 criteria, using the same 50 realizations per object. Compare the significant-and-1-sigma-consistent recovery rate to the gold-sample rates in Figure 8. If the non-gold recovery is within about 5 percentage points, the gold-sample selection does not drive the ~90% result; if it falls substantially below, the abstract should quote the gold-sample-specific rate and the Figure 9 'conservative' claim should be removed. Report target-level fractions, such as the fraction of quasars with all or at least half of their simulations successful, alongside the simulation-level Nsig.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The cadence simulation in Section 6.2 starts from the Shen et al. (2024) gold lags, thins their first-year line epochs from 32 to about 13, and reports that roughly 90% of the resulting simulations remain statistically significant (Figure 8). This conditions on exactly the outcome the survey-design question is about. The gold sample (Section 3.3) was built by visual inspection of dense-cadence data, requiring unimodal lag PDFs, good model fits, and consistency across methods; these are the light curves most likely to survive epoch removal. The survival fraction therefore estimates P(survive | gold under the original dense cadence), not P(survive | would-be gold under a sparse design) and not the recovery rate for the general parent population of significant or non-gold detections. For a planned future survey, the relevant denominator is all objects that could yield a robust lag, and the current gold selection preferentially excludes objects whose detectability depends on the dense first year. The Figure 9 caption's statement that this makes the result 'conservative' is backwards: conditioning on the easiest detections makes the retention rate higher, not lower. Additionally, Nsig is computed across 50 simulations per target, so the 90% is a simulation-level average; the target-level recovery fraction is not reported, and the quoted uncertainty ignores clustering by quasar.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 172 high-confidence (\"gold\") reverberation-mapping lag measurements from SDSS-RM, drawn from six published studies, to identify which target and light-curve properties predict successful lag detection. The main predictors examined are continuum luminosity, equivalent width, fractional rms variability, the variability signal-to-noise ratio SNR2, and the Durbin-Watson statistic of the continuum and emission-line light curves. The paper reports that the emission-line Durbin-Watson statistic is the most significant predictor of gold-lag status, with SNR2 and the placement of H-beta on the spectrograph also playing a role. The second half of the paper simulates a 40% reduction in the first-year spectroscopic cadence (from 32 to about 13 epochs) while keeping continuum sampling fixed, and reports that roughly 90% of simulated lags remain statistically significant, with 81% (H-beta), 76% (MgII), and 86% (CIV) remaining both significant and consistent with the original lag within 1 sigma. The paper concludes that a modest reduction in first-year cadence has minimal effect on overall lag recovery if later-year sampling is uniform, providing guidance for future RM surveys.","tokens_in":28955,"tokens_out":5233,"duration_ms":56993,"significance":"The practical question addressed here is valuable: future industrial-scale RM programs need quantitative guidance on cadence allocation, and this paper uses actual multi-year SDSS-RM light curves rather than only synthetic light curves. The comparison across early-year and 7-year RM studies is a useful resource, and the paper is transparent in reporting both significance-only retention and the stricter consistency-within-1-sigma retention. If the Durbin-Watson result is robust to selection effects, it would provide a cheap and effective pre-screening statistic for RM target selection. The cadence-simulation results, if recast with proper conditioning, would inform observing-strategy decisions for programs such as 4MOST TiDES. However, the headline claims are currently overstated: the abstract's \"approximately 90%\" refers to significance-only survival, and the paper's own stricter criterion gives lower rates; moreover, both the Durbin-Watson predictor and the cadence-retention estimate are conditioned on the gold sample, whose selection plausibly inflates the reported numbers.","major_comments":[{"comment":"The claim that the emission-line Durbin-Watson statistic is the strongest predictor of gold-lag status is at risk of selection confounding. The gold sample in Section 3.3 was defined by visual inspection of light curves and lag PDFs, requiring unimodal lag PDFs, good model fits, and consistency across methods. These criteria plausibly select light curves that are smooth and positively autocorrelated, i.e., those with low dw. The logistic regression in Section 5.4 therefore may be partly restating the gold-selection criterion rather than discovering an independent predictor. To support the claim, the authors should test the predictive value of dw on a sample that did not define \"gold\", for example non-gold lags that nonetheless pass the statistical significance criteria, or on simulated light curves with known input lags. As written, the p-values reported in Section 5.4 do not address this concern.","section":"Sections 3.3 and 5.3-5.4"},{"comment":"The cadence-thinning experiment conditions on the gold sample from Shen et al. (2024), which was identified under the dense first-year cadence that the simulation thins. The reported survival fractions therefore estimate P(survive | gold under the original dense cadence), not P(recover | would-be gold under a sparse design). Because the gold-selection process preferentially kept the cleanest, most robust detections, the approximately 90% retention is an optimistic bound for the general population of recoverable lags, not the conservative bound claimed in the Figure 9 caption. The paper should repeat the thinning on the full sample of statistically significant lags (including non-gold ones) or on a simulated population with known lags, and report target-level recovery fractions. In addition, Nsig is computed over 50 simulations per target, so the quoted percentage is a simulation-level average; uncertainties should account for clustering by quasar.","section":"Section 6.2, Figures 8 and 9"},{"comment":"The abstract states that a cadence reduction to about 1.5 weeks can \"retain approximately 90% of the lag measurements\", but the paper's own stricter criterion, which requires the simulated lag to be consistent with the original lag within 1 sigma, yields 81% for H-beta, 76% for MgII, and 86% for CIV (Section 6.2, Figure 8). The abstract should either use the stricter numbers or explicitly state that the 90% figure refers only to statistical significance, since the distinction is central to the survey-design recommendation.","section":"Abstract and Section 7"}],"minor_comments":[{"comment":"The definition of SNR2 is written as \"SNR2 = p χ2 − DOF\", which is unclear; it should be rendered as SNR2 = sqrt(χ2 − DOF), with DOF explicitly defined as N-1 and χ2 computed against the mean flux.","section":"Section 5.2"},{"comment":"The Durbin-Watson statistic is defined for residuals et, but the text does not specify which model the residuals come from (e.g., residuals relative to the weighted mean, or residuals from a linear trend). The reader should know exactly how dw was computed from the PrepSpec light curves.","section":"Section 5.3, Eq. (1)"},{"comment":"The summary statement that the impact of emission-line position on detection is less strong for Mg II and C IV \"(Figure 4)\" appears to reference the wrong figure; the relevant figure is Figure 3. Figure 4 shows fractional rms variability.","section":"Section 7"},{"comment":"The column heading \"SNR2 Con.\" is undefined; it should be spelled out as \"Continuum SNR2\" and defined in the caption.","section":"Table 3"},{"comment":"The logistic regression is described only briefly; the authors should state whether predictors were standardized, how p-values were obtained, and whether any multiple-testing correction was applied across the six models.","section":"Section 5.4"},{"comment":"The text uses both \"reduce the density ... to 40%\" and \"by 40%\" to describe the same operation; these should be made consistent, since they are arithmetically opposite descriptions of the retained epoch count.","section":"Section 6.1"}],"recommendation":"major_revision","confidential_remarks":"This is a useful empirical study with a clear practical goal, and the underlying data and analysis are solid. The main risks are selection confounding in the Durbin-Watson result and conditioning on the gold sample in the cadence simulation; both are addressable with additional analysis rather than being fatal. The abstract should be aligned with the stricter consistency numbers. I believe the paper is suitable for publication after these points are resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent, useful survey-design paper that deserves peer review, but its headline retention number needs qualification. The genuinely new pieces are the Durbin-Watson statistic as a target-selection metric for RM and the cadence-thinning experiments on actual SDSS-RM gold lags rather than mocks. The logistic regression and the binned recovery maps are also useful for planning TiDES and LSST-era programs.\n\nThat said, the 90% figure in the abstract is the significance-only rate; requiring the recovered lag to be consistent with the original within 1σ drops to 81, 76, and 86 percent for Hβ, MgII, CIV. The paper does report these stricter numbers in text and figure captions, but the abstract's emphasis is misleading.\n\nThe bigger issue is conditioning. The cadence simulations start from the Shen et al. (2024) gold sample, which was selected by visual inspection of dense-cadence light curves for unimodal, well-constrained lag PDFs. Those are precisely the objects most likely to survive epoch removal. So the recovery rate estimates P(survive | gold under the original cadence), not P(recover | would-be gold under the proposed cadence). The Figure 9 caption's claim that this makes the result 'conservative' is backwards; if anything it is an optimistic bound for the general lag population. I'd ask the authors to rephrase and to present target-level recovery fractions rather than simulation-level averages, with clustering accounted for.\n\nThe Durbin-Watson result has a similar circularity smell, though lighter. The gold selection criteria favored smooth light curves, so dw < 1 may partly be a restatement of the selection rule. As a practical screening tool it is still useful, but the paper should not oversell it as an independent physical predictor.\n\nThese are not fatal. The central conclusions — that a 40% reduction in first-year cadence costs a modest fraction of lags, and that dw is a decent cheap filter — likely hold up. I would send this to a journal and ask for revisions that clarify the conditioning and the two retention criteria in the abstract.","headline":"Competent RM survey-design study with a truthful core, but the abstract's 90% retention is the significance-only number and the gold-sample conditioning makes it optimistic, not conservative.","tokens_in":29605,"tokens_out":5205,"would_cite":true,"duration_ms":48012,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Durbin-Watson statistic on emission-line light curves is the strongest predictor of reliable reverberation lags, and cutting the first-year SDSS-RM cadence by 40% retains about 90% of significant detections.","keywords":["reverberation mapping","quasar black-hole masses","Durbin-Watson statistic","survey cadence optimization","AGN broad-line region","light-curve variability","SDSS-RM"],"falsifier":"Apply the same 40% first-year cadence thinning to the full SDSS-RM parent sample rather than the gold sample and compare the fraction of still-significant lags; if the retention rate falls well below 90%, the result is specific to the cleanest detections.","tokens_in":28484,"feed_emoji":"🔭","tokens_out":9563,"duration_ms":80730,"temperature":0.7,"pith_summary":"Reverberation mapping measures the time delay between quasar continuum and emission-line variability to weigh supermassive black holes, but it is expensive: surveys monitor hundreds of quasars for years and most yield no reliable lag. This paper tries to identify, in advance, which targets and which observing plans give the best return on that investment. Using 172 high-confidence (\"gold\") lag measurements from the SDSS-RM survey, it finds that a simple autocorrelation test — the Durbin-Watson statistic computed on the emission-line light curve — is the strongest predictor of a reliable lag measurement, stronger than luminosity, line strength, or variability amplitude. It also simulates what would happen if the first year of SDSS-RM had been observed at about 40% lower cadence, and finds that roughly 90% of statistically significant lags survive as long as later-year sampling is uniform. If these results hold, future reverberation-mapping campaigns can save telescope time and analysis effort by screening targets on this statistic and by adopting a steadier, less front-loaded cadence.","feed_headline":"Cadence cut by 40% still recovers 90% of quasar lags","feed_subtitle":"An autocorrelation test on emission-line light curves is the strongest predictor of which quasars yield reliable black-hole masses.","key_machinery":"The Durbin-Watson statistic, $\\mathrm{dw} = \\sum_{t=2}^{T}(e_t - e_{t-1})^2 / \\sum_{t=1}^{T} e_t^2 \\approx 2 - 2r$ for residuals $e_t$ and first-order autocorrelation $r$, is the central object. Computed on the emission-line light curve, a value below about 1 flags the positively autocorrelated 'hook' feature that reverberation-mapping codes such as PyROA need to lock onto a time delay. The paper couples this screen with logistic regression to rank predictors and with a cadence-thinning experiment that removes first-year epochs at random 50 times per target, keeps the densely sampled photometric continuum unchanged, and re-fits the line lags with PyROA.","core_discovery":"The paper's central claim is that the Durbin-Watson statistic of the emission-line light curve, rather than the continuum, is the most consistent predictor of a gold-standard lag measurement. Across the six SDSS-RM datasets, logistic regression names the line dw statistic as the strongest predictor for the first-year and seven-year Hβ results and for the CIV results, and as the single robust predictor for the seven-year MgII results; the early four-year MgII dataset is the outlier, where only luminosity shows marginal predictive power. A dw value below about 1 flags positive first-order autocorrelation in the line light curve — a \"hook\" or inflection feature that lag-fitting algorithms need. The paper also claims that thinning the first-year SDSS-RM cadence from 32 epochs to about 13 (a 40% cut), with uniform sampling afterward, keeps 94% of Hβ, 88% of MgII, and 90% of CIV significant lag detections across 50 random thinnings per target; requiring the simulated lag to agree with the original within 1σ lowers the recovery to 81%, 76%, and 86%. Recovery is somewhat worse for faint, high-redshift quasars, so the authors recommend a modest uniform cadence of about 1.5 weeks for general-purpose surveys while noting that denser sampling is still needed at the faint end.","pith_inferences":["The dw screen could be run in real time as data accumulate, letting a survey re-allocate spectroscopic epochs mid-season to targets whose line light curves are developing the 'hook' the statistic detects; this adaptive strategy is a natural extension the paper does not test.","Because the gold sample is selected by visual inspection for clean, well-behaved light curves, the 90% retention after cadence thinning is probably an optimistic bound for the full lag population; the paper calls the simulation conservative, but the selection of easy detections works in the opposite direction.","The same dw approach could be applied to continuum reverberation mapping (accretion-disk lags) once a sufficient sample of continuum lags exists, since the logic of needing an inflection in the driving light curve carries over directly.","Combining the dw screen with a baseline-length correction for luminosity would likely sharpen target selection: the paper's luminosity dependence may be an artifact of longer lags in brighter quasars rather than an intrinsic property of those sources."],"forward_implications":["Surveys can pre-screen targets with the line-light-curve dw statistic before running expensive lag analysis, focusing effort on quasars with a high chance of a gold-quality measurement.","A front-loaded cadence (very dense first year) buys little over a uniform ~1.5-week cadence once multi-year baselines are available; future programs can redistribute epochs more evenly.","The dw<1 threshold generalizes across Hβ, MgII, and CIV, so it can serve as a universal screening metric across the redshift range of industrial-scale RM programs.","Faint and high-redshift quasars need denser sampling than the 40%-reduced cadence provides, so surveys targeting the distant quasar population should not adopt the uniform sparse cadence.","The negative luminosity correlation in the logistic regression implies that longer-lag, brighter quasars are being cut off by survey baselines, so baseline length, not target quality, is the limiting factor for those sources."],"supporting_citations":[{"why":"Supplies the first-year Hβ gold lags (26 objects) and the quality-rating scheme used to define one part of the gold sample.","marker":"Grier et al. (2017)"},{"why":"Supplies the four-year MgII gold lags (24 objects) selected by false-positive rate below 10%.","marker":"Homayouni et al. (2020)"},{"why":"Supplies the four-year CIV gold lags (16 objects).","marker":"Grier et al. (2019)"},{"why":"Supplies the 7-year gold sample (37 Hβ, 32 MgII, 37 CIV) whose light curves are the basis of the cadence-thinning simulations.","marker":"Shen et al. (2024)"},{"why":"Defines the dw statistic, the paper's central predictor of lag success.","marker":"Durbin & Watson (1950)"},{"why":"Provides the PyROA method used to re-fit lags on the cadence-reduced light curves.","marker":"Donnan et al. (2021)"},{"why":"Describes the SDSS-RM survey design and the mock-light-curve predictions against which the cadence results are compared.","marker":"Shen et al. (2015a)"},{"why":"Provides the R–L relation used to interpret the luminosity dependence of lag success.","marker":"Bentz et al. (2013)"}],"fun_headline_variants":["40% fewer cadence points still yield 90% of quasar lags","Autocorrelation test identifies quasars yielding reliable black-hole masses","Cut cadence by 40% and keep 90% of lag detections","Durbin-Watson statistic picks quasar lines for efficient RM","A smarter cadence keeps black-hole mass lags intact"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The cadence-thinning test is run on the gold sample, the easiest lags to detect, so the ~90% retention rate is likely optimistic for the full population of recoverable lags.","fun_headline_variants_meta":{"raw":{"variants":["40% fewer cadence points still yield 90% of quasar lags","Autocorrelation test identifies quasars yielding reliable black-hole masses","Cut cadence by 40% and keep 90% of lag detections","Durbin-Watson statistic picks quasar lines for efficient RM","A smarter cadence keeps black-hole mass lags intact"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000701,"raw_usage":{"total_tokens":3228,"prompt_tokens":1069,"completion_tokens":2159,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":685,"completion_tokens_details":{"reasoning_tokens":2064}},"tokens_in":685,"tokens_out":2159,"duration_ms":17563,"temperature":1.0,"reasoning_tokens":2064,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:18:17.636293+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the same 40% first-year cadence thinning to the full SDSS-RM parent sample rather than the gold sample and compare the fraction of still-significant lags; if the retention rate falls well below 90%, the result is specific to the cleanest detections.","supporting_citations":[],"review_version":1}