{"id":"ede6d61a-8693-4c44-84f2-87d61ff556a5","arxiv_id":"2608.12487","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Two of six Little Red Dots show marginal broad Halpha variability between JWST epochs, weakly supporting an AGN origin for those sources.","lead":"This paper compared new JWST spectra of six little red dots against older spectra to look for changes in their broad hydrogen emission lines. It reports marginal variability in two of six sources, a weak but meaningful hint that those objects host accreting supermassive black holes rather than pure scattering envelopes.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Both variability detections rest on the unvalidated assumption that narrow [Oiii]/[Neiii] flux is constant; a small calibrator drift would erase them.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing concern I find: the multi-epoch flux calibration via constant narrow [Oiii]/[Neiii] flux (Eq. 1, Section 3.1). The paper is unusually transparent about this fragility, explicitly noting that the constant narrow-line flux assumption 'might not be applicable in LRDs' and that for the GlimmIr the [Neiii] normalization ratio is 25x less sensitive than the [Oiii] ratio and dominates the uncertainty budget. My concern is therefore not a new objection but a sharpening of the same one: the two claimed detections are both marginal in significance, and their significances are largely set by the calibrator uncertainty rather than by intrinsically weak variability. Because the paper already labels the detections as 'marginal' and the conclusion as 'some evidence,' the CONDITIONAL verdict remains appropriate; I would not move it to REJECT or UNVERDICTED. The proposed perturbation test would settle whether the result survives a plausible systematic error in the calibrator, which is the decisive missing check. I do not see an internal inconsistency or a stronger alternative concern: the removal of the artifact-affected P63 epoch is well documented, the alternate fitting methodology is checked, and the GlimmIr's pre-selection is disclosed, though it weakens the population-level Monte Carlo comparison. Those issues support the conditional framing but do not replace the calibrator-constancy concern as the most load-bearing point.","tokens_in":22232,"tokens_out":5410,"duration_ms":53712,"concrete_test":"Recompute ΔF/F for OCEANS-100424 and the GlimmIr after perturbing the Eq. 1 normalization ratio by ±10% and by ±1σ, and identify the perturbation at which each detection falls below 2σ. If a ±10% [Oiii]/[Neiii] drift (or the 1σ normalization error itself) removes either detection, the variability claim is a systematic artifact rather than evidence for AGN. A complementary check: use the GlimmIr's P2/P3 OCEANS [Oiii] flux ratio to calibrate the RUBIES epoch instead of the P6 [Neiii] ratio; if the resulting broad-Hα difference is not significant, the 'variable' status of the GlimmIr depends entirely on the lower-sensitivity calibrator.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference that BL Hα variability in OCEANS-100424 and the GlimmIr favors a direct AGN/BLR origin is only as secure as the assumption in Eq. 1 (Section 3.1) that the narrow [Oiii] (or [Neiii]) flux is constant over ~100 rest-frame days and that slit losses are identical across epochs. The paper itself flags this as not validated in LRDs, citing Ishikawa et al. 2026 (Section 3.1), and reports that the GlimmIr's variability significance is dominated by the low-SNR [Neiii] normalization ratio (0.71±0.25, 25x less sensitive than the [Oiii] ratio; Section 4). For OCEANS-100424 the [Oiii] ratio also contributes substantially to the 2.1σ significance. Because both 'variable' sources are the two with the largest normalization uncertainties, a 10-20% intrinsic narrow-line variation or differential slit loss between epochs—well within plausible ranges for compact, high-z LRDs—could produce the reported ΔF/F values without any broad-line variability. Thus the headline conclusion is not yet robust against the scattering alternative unless the calibrator constancy is independently checked.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper searches for broad Halpha and continuum variability in six Little Red Dots (LRDs) by comparing new R~2700 OCEANS NIRSpec observations with archival R~1000 CEERS/RUBIES spectra. Using [Oiii] (or [Neiii]) narrow-line fluxes to flux-calibrate the epochs via Eq. (1) and fitting the Halpha+[Nii] complex with a narrow+broad Gaussian model, the authors report marginal broad-line variability in OCEANS-100424/RUBIES-42232 (27% at 2.1 sigma) and OCEANS-35829/RUBIES-49140, the GlimmIr (50% at 1.5 sigma), and upper limits of 4.8%--30% for the other four LRDs. They find no significant continuum variability and compare the sample to SDSS-RM quasars, concluding that the probability of reproducing two variable and four nonvariable sources is 4.71%, corresponding to a ~2 sigma departure from typical quasar variability. The paper interprets the marginal broad-line variability as evidence for a direct line of sight to an AGN broad-line region in at least some LRDs, as opposed to purely scattering-dominated models.","tokens_in":22411,"tokens_out":4298,"duration_ms":41377,"significance":"If the variability detections are robust, this is a valuable and timely result: it would be one of the first direct dynamical tests favoring an AGN/BLR origin for at least a subset of LRDs and would disfavor pure electron-scattering models for those objects. The paper has several genuine strengths: it uses an external SDSS-RM benchmark rather than tuning parameters to force detections, it performs Monte Carlo comparisons with a clearly described sample construction, it applies empirical uncertainty corrections to mitigate known NIRSpec pipeline issues, and it carefully documents and excludes a problematic epoch (RUBIES P63) using quantitative spatial-profile checks. The authors are also appropriately transparent that both detections are marginal and that the normalization uncertainty dominates. However, the central claim rests on an unvalidated assumption about narrow-line constancy, and the statistical significance is low, so the headline conclusion is not yet secure.","major_comments":[{"comment":"The flux calibration assumes that the narrow [Oiii] (or [Neiii]) flux is constant over the ~100 rest-frame day baselines and that slit losses are identical across epochs, but this is not demonstrated for LRDs. The paper itself notes that constant narrow-line fluxes in local AGN may not be applicable to LRDs (citing Ishikawa et al. 2026). This assumption is load-bearing: for OCEANS-100424 the [Oiii] ratio is 0.47+-0.06, and for the GlimmIr the [Neiii] ratio is 0.71+-0.25, with the latter dominating the uncertainty budget (Section 4). A modest 10--20% intrinsic narrow-line variation or differential slit loss between epochs could produce the reported Delta F/F values without any broad-line variability. The authors should provide an independent check of calibrator constancy, for example by comparing multiple narrow lines across epochs, or by quantifying the maximum allowable calibrator drift before the detections disappear.","section":"Section 3.1, Eq. (1)"},{"comment":"The Monte Carlo comparison includes the GlimmIr as one of the '2 variable' sources, but the GlimmIr was intentionally pre-selected for OCEANS follow-up because it was already known to vary (Lambrides et al. 2026a; stated in Section 4.2). Counting a pre-selected variable source as a random draw from the SDSS-RM variability distribution inflates the significance of the comparison. The reported 4.71% probability of reproducing '2 variable and 4 nonvariable quasars' is therefore not an unbiased test of the LRD population. The authors should either repeat the Monte Carlo excluding the GlimmIr or treat its variability as a prior, and report the resulting probability; note that the 86% consistency quoted for OCEANS-100424 suggests the independent evidence is much weaker.","section":"Section 4.2"},{"comment":"Both reported detections are below 3 sigma (2.1 sigma and 1.5 sigma), and for OCEANS-100424 the variability measurement relies on a single OCEANS epoch compared with a single RUBIES epoch after the exclusion of RUBIES P63. The Appendix convincingly justifies dropping P63 on the basis of its broader spatial profile, but the authors should demonstrate that the inferred variability is not sensitive to the choice of which RUBIES epoch is used, or at least quantify how the significance would change if P63 were included. As it stands, the 'detection' for OCEANS-100424 is a two-epoch comparison with a calibration uncertainty that is comparable to the signal.","section":"Section 4, Table 3"}],"minor_comments":[{"comment":"The 1-sigma upper limits of 4.8%--30% are quoted only as a range; please list the individual upper limits per source or refer explicitly to their values in Table 3 or Figure 6.","section":"Section 4"},{"comment":"The entries for de Graaff et al. 2025b and 2025d appear to be identical (A&A, 697, A189); please check whether one is a duplicate or should cite a different paper.","section":"References"},{"comment":"In the sentence describing Madau et al. (2026), 'viralized' should be 'virialized'.","section":"Section 1"},{"comment":"The text says the GlimmIr P6 observation has a chip gap in its [Oiii] region and later says it also has a chip gap in the 5100 Å region; please make the wording consistent about which regions are affected.","section":"Section 3.1"},{"comment":"The description of drawing '20 random spectroscopic pairs' per SDSS-RM source should clarify how pairs are drawn when a source has a limited number of epochs and whether the random draws are independent across sources.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest and well structured, but the abstract and conclusions state the central claim more strongly than the evidence supports. The 2.1 sigma and 1.5 sigma detections, the unvalidated narrow-line calibration assumption, and the pre-selection of the GlimmIr together mean that the paper currently provides only weak evidence for broad-line variability in LRDs. I would support publication after a revision that adds a calibration-robustness test and reframes the statistical claims to account for the GlimmIr preselection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a careful, honest observational analysis of six Little Red Dots using new R~2700 OCEANS spectra plus archival R~1000 data. It reports marginal broad Hα variability in two sources (27% at 2.1σ and 50% at 1.5σ) and gives upper limits for the other four. The abstract is appropriately cautious: the detections provide “some evidence” for AGN-powered broad lines, not proof.\n\nWhat the paper does well: new OCEANS spectroscopy is genuinely new, the flux-calibration and fitting steps are described in enough detail to reproduce, and the authors openly discuss the pre-selection of the GlimmIr and the large normalization uncertainties. They also compare to SDSS-RM quasars and report that the “two variable and four nonvariable” pattern is only a ~2σ departure from quasar-like variability. I appreciate that they do not oversell the result.\n\nThe soft spots are real and concentrated in one place: the assumption that narrow [Oiii] or [Neiii] flux is constant across ~100 rest-frame days and that slit losses are identical. The paper cites local-AGN constancy and then notes, following Ishikawa et al. 2026, that this might not apply to LRDs. That is the load-bearing premise. For the GlimmIr, the [Neiii] normalization ratio has 0.71±0.25, which dominates the uncertainty and makes the 1.5σ detection fragile. For OCEANS-100424, the [Oiii] ratio also contributes substantially. A 10–20% intrinsic variation in the narrow-line calibrator, or differential slit loss between epochs, would erase both detections. This is not a manufactured flaw; the stress-test concern is fair.\n\nAlso, the GlimmIr was already known to vary and was deliberately selected for OCEANS follow-up. Counting it as one of the two detections makes the “2 of 6” statistic less meaningful, though the authors are transparent about this. The nonvariable sources have weak upper limits, so the contrast between variable and nonvariable is not sharp.\n\nBottom line: this is a useful paper for the LRD community. It adds new data and an honest upper-limit framework, and it will be cited. The central claim is conditional, but the authors say so themselves. It deserves a serious referee, with particular attention to the narrow-line constancy assumption and to a version of the statistics that excludes the pre-selected GlimmIr. I would engage with it as a reviewer or citing author.","headline":"A transparent, marginal variability study of six LRDs that deserves refereeing but whose two low-significance detections rest on an unvalidated narrow-line constancy assumption, which the authors themselves flag.","tokens_in":23118,"tokens_out":1801,"would_cite":true,"duration_ms":18399,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two of six little red dots show marginal variability in their broad Hα lines, evidence that at least some of these compact JWST sources are powered by accreting black holes rather than scattered light.","keywords":["little red dots","broad-line variability","active galactic nuclei","Hα emission","JWST NIRSpec spectroscopy","supermassive black holes","electron scattering","quasar variability"],"falsifier":"Re-observe OCEANS-100424 and the GlimmIr in a third epoch with deep, high-resolution spectroscopy that covers both the [Oiii]/[Neiii] calibrators and Hα, and measure whether the narrow-line fluxes stay constant; if a calibrator line changes by more than roughly its statistical error between any two epochs, the reported 27% and 50% broad-line variability fractions are calibration artifacts, not AGN variability.","tokens_in":21975,"feed_emoji":"🔭","tokens_out":8332,"duration_ms":67095,"temperature":0.7,"pith_summary":"This paper asks whether little red dots—compact, red, JWST-discovered sources—are powered by accreting supermassive black holes or by light scattering in dense gas. The authors compare new high-resolution spectra of six little red dots with archival observations taken roughly 100–180 rest-frame days earlier. Two sources show marginal changes in their broad Hα line flux: 27% variability at 2.1σ significance for OCEANS-100424 and 50% at 1.5σ for the GlimmIr, while the other four show no broad-line change and the sample shows no continuum variability. The authors read this as evidence that at least some little red dots have a direct line of sight to a compact, virialized broad-line region near a black hole, which would favor the AGN interpretation over pure scattering models.","feed_headline":"Two little red dots show marginal flickers in Hα","feed_subtitle":"Marginal broad-line variability in 2 of 6 JWST sources points to accreting black holes, not scattered light.","key_machinery":"The load-bearing method is multi-epoch flux calibration through a narrow forbidden line, usually [Oiii] λλ4959,5007 and for the GlimmIr [Neiii] λ3869: each archival spectrum is scaled by the ratio of the narrow-line flux between epochs (the paper's Eq. 1) so that different slit orientations and aperture losses cancel. The Hα + [Nii] complex is then fit simultaneously across epochs with a shared narrow component plus a broad Gaussian using an MCMC routine, and the 5100 Å continuum is measured from the same calibrated spectra. The narrow-line ratio is doing all the work: if the narrow-line flux is not truly constant, the derived broad-line variability is an artifact, and the paper notes that the constant narrow-line assumption, well tested for local AGN, may not hold for little red dots; the GlimmIr's [Neiii] ratio has 25 times lower sensitivity than the [Oiii] ratios and dominates its uncertainty budget.","core_discovery":"The central claim, stated on the paper's own terms, is that two of six little red dots re-observed with higher-resolution JWST spectroscopy show marginal broad Hα flux variability—27% (2.1σ) in OCEANS-100424 and 50% (1.5σ) in the GlimmIr—while four others are consistent with no broad-line variability (1σ upper limits of 4.8%–30%) and none of the six shows significant continuum variability. The authors argue that broad-line variability on these ~100–180 rest-frame day baselines indicates that the broad Hα emission originates close to a central engine, with a clear line of sight undiluted by scattering or reprocessing. They additionally compare the observed pattern to low-redshift quasar variability and find that reproducing two variable and four nonvariable sources has a 4.71% probability, a ~2σ departure from typical quasar behavior, driven mainly by the extreme variability of the GlimmIr.","pith_inferences":["A decisive next step the paper does not take would be a third epoch for these two sources with deep coverage of both [Oiii] and Hα; if the narrow-line calibrator itself varies, both detections would vanish.","Measuring variability of higher-ionization lines such as Hβ or He II in the same spectra would test whether the inner broad-line region responds to continuum changes, strengthening the virial interpretation.","The authors' post-blowout or clumpy-medium explanation predicts a correlation between variability amplitude, Balmer break strength, and line-profile shape; ranking a larger sample of little red dots by these properties would test that picture.","If the 4.71% quasar-comparison result holds up in a larger sample, it would imply that little red dots differ systematically from low-redshift quasars in broad-line-region geometry, Eddington ratio, or the fraction of non-AGN contaminants among nonvariable sources."],"forward_implications":["If the variability is real, at least some little red dots contain a compact, virialized broad-line region with a direct line of sight, which would validate applying local single-epoch black-hole mass estimators to this population.","Pure electron-scattering models for the broad lines are disfavored for OCEANS-100424 and the GlimmIr, since scattering would smooth out any variability signal; both sources also lack the exponential line wings that scattering models predict.","The four nonvariable sources and the lack of continuum variability imply that little red dots are not a homogeneous population and that variability may be tied to evolutionary phase or covering fraction.","The 4.71% joint probability means the observed two-variable/four-nonvariable pattern is a ~2σ departure from typical quasar variability, with the GlimmIr as the main outlier; higher signal-to-noise re-observation of OCEANS-161695 would sharpen the nonvariable constraints.","OCEANS-100424's variability amplitude lies near the 3σ detection limit of the existing slitless survey that found no variability, which can reconcile the two results."],"supporting_citations":[{"why":"Defines the color–color LRD selection used to choose the sample.","marker":"Barro et al. 2024"},{"why":"Fixes the [Oiii] doublet flux ratio 1:2.985 used in the narrow-line fits that calibrate the epochs.","marker":"Storey & Zeippen 2000"},{"why":"Supplies the local-AGN evidence that narrow-line fluxes are constant over roughly one year, the premise behind Eq. 1.","marker":"Foltz et al. 1981"},{"why":"Cautions that constant narrow-line fluxes may not hold for little red dots, directly threatening the calibration.","marker":"Ishikawa et al. 2026"},{"why":"Previously reported variability for the GlimmIr, which motivated its selection and frames its extreme 50% change.","marker":"Lambrides et al. 2026a"},{"why":"Supplies the SDSS-RM quasar variability distribution used in the probability calculation.","marker":"Shen et al. 2024a"},{"why":"Provides the electron-scattering model that the variability detections are used to disfavor.","marker":"Rusakov et al. 2026"},{"why":"Gives the Hα profile fits and absorption properties of these sources used in the multi-epoch fitting.","marker":"Davis et al. 2026"},{"why":"The slitless survey that found no LRD variability and sets the sensitivity comparison for OCEANS-100424.","marker":"Liu et al. 2026"}],"fun_headline_variants":["Two little red dots flicker in Hα, hinting at black holes","Marginal Hα variability in 2 of 6 JWST red dots","OCEANS ripples: two LRD broad lines vary, suggesting AGN","Do little red dots hide black holes? Two show variability","JWST's little red dots: 2 of 6 have varying Hα"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on the assumption that each source's narrow [Oiii] (or, for the GlimmIr, [Neiii]) line flux is constant across the ~100–180 rest-frame days between observations, so that the ratio of narrow-line fluxes can be used to place the two epochs on the same flux scale; the paper itself notes this may not hold for little red dots, and for the GlimmIr the [Neiii] calibrator is 25 times less sensitive and dominates the uncertainty.","fun_headline_variants_meta":{"raw":{"variants":["Two little red dots flicker in Hα, hinting at black holes","Marginal Hα variability in 2 of 6 JWST red dots","OCEANS ripples: two LRD broad lines vary, suggesting AGN","Do little red dots hide black holes? Two show variability","JWST's little red dots: 2 of 6 have varying Hα"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000258,"raw_usage":{"total_tokens":1665,"prompt_tokens":1112,"completion_tokens":553,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":728,"completion_tokens_details":{"reasoning_tokens":454}},"tokens_in":728,"tokens_out":553,"duration_ms":5132,"temperature":1.0,"reasoning_tokens":454,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:07:22.757797+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-observe OCEANS-100424 and the GlimmIr in a third epoch with deep, high-resolution spectroscopy that covers both the [Oiii]/[Neiii] calibrators and Hα, and measure whether the narrow-line fluxes stay constant; if a calibrator line changes by more than roughly its statistical error between any two epochs, the reported 27% and 50% broad-line variability fractions are calibration artifacts, not AGN variability.","supporting_citations":[{"cited_title":"Spatial decomposition of Little Red Dots with JWST/NIRSpec IFU into broad-line red cores and narrow-line blue host galaxies","cited_arxiv_id":"2607.09647","evidence_quote":"Cautions that constant narrow-line fluxes may not hold for little red dots, directly threatening the calibration."}],"review_version":1}