{"id":"ec3b5295-49eb-44f6-8079-66dbb5966110","arxiv_id":"1908.03130","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors derive hypothesis-testing based contrast detection limits for kernel-phase analysis and show JWST NIRISS full-pupil observations can detect companions at contrasts near 10^3 at 200 mas with controlled false alarms.","lead":"This paper builds a statistical framework for setting detection limits in kernel-phase interferometry, then applies it to simulated JWST NIRISS images of faint brown dwarfs. It shows that contrasts of about 1,000 at 200 milliarcseconds are detectable with a 1% false-alarm probability, and that the practical test nearly matches the theoretical optimum.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Calibration covariance model is the weakest link: diagonal variance inflation from 10 OPD maps may understate correlated residuals, especially for bright targets where calibration error dominates (Sec. 3.7).","rationale":"The reader's weakest_assumption correctly identifies the noise and calibration model as the most load-bearing premise. The paper's statistical framework, including the Neyman-Pearson bound and the practical T_B test, assumes whitened kernel phases are Gaussian with a known covariance. The treatment of calibration residuals is the least secure part of that model: the authors use only 10 OPD maps, scale them to a nominal 16 nm RMS, and add a diagonal variance term, explicitly neglecting correlated residuals. For faint Y-dwarf targets the calibration contribution is small, so the approximation is defensible there, but for the bright-target case in Sec. 3.7, calibration error is 85% of the total variance, and the paper's claim of contrasts up to 10^5 depends on that approximation. If the true calibration residuals are correlated, the whitening is incorrect, the test statistics' distributions change, and the false alarm rate can no longer be guaranteed. The paper itself flags calibration errors as the most critical limitation in the conclusion, reinforcing this concern. The optimizer initialization issue is also present, but the authors report a grid-search cross-check that partially mitigates it; the calibration covariance issue is explicitly left unaddressed for exactly the regime where it matters most. A concrete test using a full covariance from many OPD realizations would settle whether the diagonal approximation is adequate. Since the authors themselves condition their conclusions on ideal performance and the reader already recommends CONDITIONAL, my assessment does not change the verdict.","tokens_in":16705,"tokens_out":5184,"duration_ms":53445,"concrete_test":"Generate a large ensemble (e.g., 1000) of wavefront drift realizations from the expected JWST drift statistics, compute the full (non-diagonal) covariance of calibration residuals on kernel phases, and re-run the detection-limit calculation for the bright-target scenario (Sec. 3.7) using this full covariance. Compare the resulting ROC curves and contrast limits at PFA=1%, PDET=68% against the diagonal-approximation results. If the nominal threshold yields a PFA that deviates from 1% by more than a factor of 2, or if the contrast limit shifts by more than a factor of 2, the central claim of an operational method with guaranteed false-alarm control is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that T_B is an operational detection test close to the Neyman-Pearson bound hinges on the noise model in Eq. (10): whitened kernel phases must be Gaussian with known covariance. The paper estimates the statistical covariance via Monte Carlo (Sec. 3.3) but handles calibration residuals by adding a diagonal variance inflation derived from only 10 OPD maps scaled to 16 nm RMS (Sec. 3.2). This diagonal approximation ignores any correlation structure in wavefront-drift-induced kernel-phase errors. The authors explicitly state they 'chose not to pursue the non-diagonal terms' (Sec. 3.3), justified only for faint targets where calibration error is small. However, for the bright-target scenario (Sec. 3.7), calibration error accounts for 85% of the total noise variance, and the paper's headline contrast limits of up to 10^5 rely directly on this approximation. If calibration residuals are correlated across kernel phases, the whitening in Eq. (8) is incorrect, so the test statistics T_NP, T_E, and T_B no longer have the assumed distributions. Consequently, the false alarm rate would not be held at the nominal 1%, and the detection limits (Fig. 8 and Fig. 11) could shift substantially. This is the most load-bearing assumption because it underpins the statistical guarantees that distinguish this work from earlier heuristic kernel-phase contrast calculations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a statistical hypothesis-testing framework for kernel-phase observables and applies it to simulated JWST NIRISS full-pupil F480M images. Three detection tests are constructed on whitened kernel phases: the Neyman-Pearson (NP) test for a known companion signature, an energy detector for a completely unknown signature, and a generalized likelihood ratio test (T_B) for a binary companion with unknown position and contrast. Analytic false-alarm and detection probabilities are derived for the NP and energy tests, while the distribution of T_B is calibrated by Monte Carlo simulation. The authors validate the theoretical ROC curves against Monte Carlo runs, show that T_B performs close to the NP upper bound, and derive contrast detection limits for representative Y-dwarf targets. The headline quantitative result is that, at PFA = 1% and PDET = 68%, companions at contrasts up to about 10^3 at 200 mas around a W2 = 14.1 target are detectable, with brighter targets allowing contrasts up to 10^5 in the idealized case.","tokens_in":16912,"tokens_out":5898,"duration_ms":60461,"significance":"If the results hold, the paper provides a principled and general method for computing and guaranteeing kernel-phase detection limits with controlled false-alarm rates, applicable to any well-sampled high-Strehl imaging instrument. The Monte Carlo validation of the theoretical distributions is a clear strength, and the use of open-source simulation tools (XARA, ami_sim) makes the numerical experiments reproducible. The paper also supplies concrete, falsifiable predictions for JWST NIRISS full-pupil observations, and it correctly identifies the NP test as an upper performance bound. The central methodological derivation—the whitened Gaussian model and the three tests—is standard and sound, and the Monte Carlo checks support the claimed distributions for T_NP and T_E.","major_comments":[{"comment":"The Monte Carlo evaluation of T_B initializes the least-squares optimizer at the true injected companion parameters ('The initialisation of the algorithm corresponds to the parameters of the injected companion'), and the paper does not state how the test is initialized under H0, where no true companion exists. Because the likelihood is multimodal (Fig. 2), the reported ROC curves and detection limits for T_B (Figs. 7-9 and 11) are conditional on an oracle start point. The assertion that a systematic grid search gives 'very similar results' is not documented with any quantitative comparison. A blind operational search can converge to local maxima, which would change both the detection power and the Monte-Carlo-calibrated threshold. Please quantify the grid-search comparison, specify the initialization under H0, and, if performance degrades, qualify the claim that T_B is an operational test close to the NP bound.","section":"3.4 and 3.5"},{"comment":"The calibration error is modeled only as a diagonal term added to the covariance, and the paper explicitly chooses not to pursue non-diagonal terms (Sec. 3.3). In the bright-target scenario this term accounts for 85% of the total noise variance (Sec. 3.7), so the whitening in Eq. (8) and the distributions of T_NP, T_E, and T_B rely on the assumption that calibration residuals are uncorrelated across kernel phases. If they are correlated, the nominal 1% false-alarm rate is not guaranteed and the bright-target limits in Fig. 11, including the 10^5 contrast value, are not robust. Please either estimate the full calibration covariance from a larger set of OPD realizations and re-run the tests, or present the bright-target limits with an explicit caveat that they hold only under the diagonal-calibration assumption.","section":"3.3 and 3.7"}],"minor_comments":[{"comment":"The caption lists 'The dashed lines represent theoretical detection limits for TE and TNP (Eq. (17) and Eq. (24))', but the equation numbers are swapped: Eq. (17) gives the TNP ROC and Eq. (24) gives the TE ROC.","section":"3.5 / Fig. 8 caption"},{"comment":"In the sentence 'will result in unaccounted residual residual errors', the word 'residual' is duplicated.","section":"3.2"},{"comment":"In the first paragraph, 'ker phases' should read 'kernel phases'.","section":"2.2"},{"comment":"The closing sentence 'The gradient descent procedure is indeed only applicable in the context of only applicable in the context of the determination of detection limits' contains a repeated phrase and should be rewritten.","section":"3.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of A&A. The two major comments above are addressable with additional simulation work. The authors should consider softening the abstract's bright-target 10^5 claim until the calibration covariance and initialization issues are treated more rigorously."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this paper brings a proper hypothesis-testing formalism to kernel-phase detection limits, and the core results hold up. The three tests—Neyman-Pearson, energy detector, and the binary GLR—are clearly derived, and the Monte Carlo validation matches the theoretical distributions. The new piece is the practical binary GLR test T_B, which lands within a factor of 2.5 of the theoretical bound in contrast. That is a solid, useful result for anyone planning JWST NIRISS observations.\n\nThe paper does well in its honesty. The authors state upfront that the non-diagonal calibration terms are ignored (Sec. 3.3), and the conclusion labels the performance as ideal. They also show the bright-target case where calibration error dominates, with a factor-of-10 degradation under a 16 nm wavefront drift. So the numbers in the abstract are not overclaimed; the limits are clearly qualified.\n\nThe softest spot is exactly what the stress-test note flags: the calibration covariance model. Adding a diagonal variance inflation from 10 OPD maps is a crude approximation. If real wavefront drifts introduce correlated errors across kernel-phases, the whitening is wrong and the constant false-alarm guarantee breaks. This is a real limitation, but it is not central to the main use case. For the faint Y-dwarf targets the paper is aimed at, calibration error is only about 14% of the total variance, so the diagonal approximation is probably fine there. For bright targets, the authors already show the impact and do not pretend otherwise. The other concern is that the T_B optimizer is initialized at the true companion parameters, but they mention checking against a grid search, so the reported detection limits are not grossly optimistic.\n\nThe paper is self-contained and the code behind it (XARA, ami_sim) is public, which helps reproducibility. The citation pattern looks reasonable, with the relevant prior work on kernel-phase and NRM detection limits cited. No signs of circularity.\n\nWho is this for? Observers planning JWST high-contrast observations, especially brown-dwarf companion searches, and anyone working on kernel-phase or closure-phase detection limits. The framework is general and not tied to NIRISS. It deserves a serious referee; the main revision request should be to tighten the calibration covariance treatment or at least state explicitly when the diagonal approximation is valid. I would bring it to reading group and cite it in the context of detection limit methodology.","headline":"A genuinely useful statistical framework for kernel-phase detection limits, with a clear-eyed caveat about calibration residuals that the authors themselves acknowledge.","tokens_in":17512,"tokens_out":2145,"would_cite":true,"duration_ms":25410,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that binary detection with kernel phases can be cast as a hypothesis-testing problem, and that a practical generalized likelihood-ratio test comes within a factor of about 2.5 in contrast of the Neyman-Pearson upper bound…","keywords":["kernel phase","hypothesis testing","Neyman-Pearson test","generalized likelihood ratio","JWST NIRISS","binary detection limits","brown dwarfs","high-contrast imaging"],"falsifier":"Take a real F480M full-pupil NIRISS observation of a binary with a $W2 \\approx 14.1$ primary and a confirmed companion at about 200 mas with contrast near $10^3$, using an 80-minute on-target sequence as simulated; if the operational test $T_B$ does not achieve $P_{\\mathrm{DET}} \\approx 68\\%$ at $P_{\\mathrm{FA}} = 1\\%$ for this companion, the predicted detection limit is not realized. A simpler numerical check is to simulate a correlated calibration term beyond the diagonal inflation and see whether $T_B$'s contrast limit falls more than a factor of 2.5 below the Neyman-Pearson bound.","tokens_in":16454,"feed_emoji":"🔭","tokens_out":6876,"duration_ms":66897,"temperature":0.7,"pith_summary":"Kernel-phase analysis infers the Fourier phase of a source while cancelling small telescope aberrations, and this paper asks how good it can be at finding faint companions. The answer is framed as a statistical hypothesis test: given whitened kernel phases, how reliably can a binary signature be distinguished from noise? The paper derives a Neyman-Pearson test that gives the theoretical best-possible detection performance, an energy detector that gives a lower bound, and a practical binary generalized likelihood-ratio test that lands within about a factor of 2.5 in contrast of the best possible bound. Applied to simulated JWST NIRISS full-pupil F480M images of Y-type brown dwarfs, it predicts that a 1000-to-1 companion at 200 mas is detectable at 68% probability with a 1% false-alarm rate, and that brighter targets could reach contrasts near $10^4$ to $10^5$. If these predictions hold, kernel-phase becomes a quantitatively predictable companion-search mode for JWST rather than an empirical calibration-dependent one.","feed_headline":"JWST NIRISS kernel phases can spot 1000-to-1 companions at 200 mas","feed_subtitle":"A generalized likelihood-ratio test lands within 2.5x of the theoretical bound, making JWST contrast limits predictable.","key_machinery":"The kernel matrix $K$ is the left nullspace of the aperture phase transfer matrix $A$, so $KA = 0$; applied to the Fourier phase $\\varphi$ it produces kernel phases $k = K\\varphi$ that are first-order insensitive to aberration phase. Whitening by $\\Sigma^{-1/2}$ decorrelates them into statistically independent observables. The operational detection test is the binary generalized likelihood ratio $T_B = 2y^{\\mathsf{T}}\\hat{x} - \\hat{x}^{\\mathsf{T}}\\hat{x}$, which equals the reduction in squared residuals between the null model and the fitted binary model, $T_{\\chi^2}(0,y) - T_{\\chi^2}(\\hat{x},y)$. The covariance $\\Sigma$ is estimated from $10^5$ Monte Carlo noise realizations, with a diagonal inflation term set by ten OPD maps scaled to 16 nm RMS to represent calibration drift. This machinery carries the argument because it converts an image into a vector with known statistics, allowing detection thresholds and contrast limits to be computed rather than guessed.","core_discovery":"The central claim is that detection of a binary by kernel phases can be fully characterized by three tests operating on whitened kernel-phase observables. The paper constructs a whitened vector $y = \\Sigma^{-1/2} k$ from kernel phases $k$, with covariance $\\Sigma$ modelled from photon, readout, and dark noise plus a diagonal calibration-inflation term. In this space the problem is $y = \\epsilon$ under the null hypothesis and $y = x + \\epsilon$ under the alternative, with $\\epsilon$ standard Gaussian. The likelihood-ratio (Neyman-Pearson) test $T_{\\mathrm{NP}} = y^{\\mathsf{T}} x$ is the most powerful possible; the energy detector $T_E = \\|y\\|^2$ uses no target information and is the least powerful; and the binary test $T_B = 2y^{\\mathsf{T}} \\hat{x} - \\hat{x}^{\\mathsf{T}} \\hat{x}$ substitutes the maximum-likelihood estimate of the binary signature (separation, position angle, contrast). Monte Carlo simulations on NIRISS F480M images show that $T_B$'s ROC curve hugs the $T_{\\mathrm{NP}}$ curve, and that the resulting contrast limits are within a factor of 2.5 of the theoretical bound at $P_{\\mathrm{FA}} = 1\\%$ and $P_{\\mathrm{DET}} = 68\\%$. For a $W2 = 14.1$ Y dwarf, companions at contrasts up to about $10^3$ at 200 mas are detectable; the brightest unsaturated targets could reach contrasts of $10^4$ to $10^5$ beyond about 500 mas, barring wavefront drift, whose uncalibrated residual accounts for 85% of the kernel noise variance in that bright regime.","pith_inferences":["The small gap between $T_B$ and the Neyman-Pearson bound implies that further algorithmic or calibration improvements can buy at most roughly a factor of 2.5 in contrast; more radical gains would need new observables such as visibility amplitudes, which the paper notes are left unused.","The constant false-alarm property suggests an operational survey design: fix thresholds from simulations once, then monitor only the empirical false-alarm rate on calibrator fields, which would also reveal whether correlated calibration errors are secretly inflating detections.","A direct test of the predictions is possible with early JWST NIRISS data on known wide binaries: measure the $T_B$ ROC empirically and compare the 68% detection contour with the paper's simulated curves, including the position-angle dependence produced by the non-centrosymmetric JWST PSF.","The position-angle dependence of the limits implies that observatories should publish per-orientation detection-limit maps rather than a single contrast curve; surveys combining multiple roll angles could then marginalize over the PSF asymmetry."],"forward_implications":["At $P_{\\mathrm{FA}} = 1\\%$ and $P_{\\mathrm{DET}} = 68\\%$, companions of contrast around $10^3$ at 200 mas are detectable around $W2 = 14.1$ Y dwarfs in F480M, which translates to a 1 Jupiter-mass companion at 1.5 AU around a 30 Jupiter-mass brown dwarf at 8 pc.","For the brightest NIRISS full-pupil targets, contrasts up to about $10^4$ (and ideally $10^5$ beyond 500 mas) are reachable if wavefront drift stays near the 16 nm RMS prediction; a drift of that size cuts bright-target performance by about a factor of 10.","Because the false-alarm probability stays constant for phase aberrations below about one radian, thresholds can be set a priori and surveys do not need to re-calibrate the false-alarm rate at every epoch.","The same three-test framework applies to any adequately sampled imaging system, including aperture-masking (NRM) data, not only NIRISS full-pupil images.","At separations below $\\lambda/D$, contrast and separation estimates are strongly correlated, so orbital fits for such binaries will need independent companion-luminosity or astrometric priors."],"supporting_citations":[{"why":"Introduces kernel phases and the linear phase model $\\varphi = \\varphi_0 + A\\phi$ whose nullspace defines the kernel matrix $K$.","marker":"(Martinache 2010)"},{"why":"Provides the likelihood-ratio test used as the theoretical upper performance bound for any detection test.","marker":"(Neyman & Pearson 1933)"},{"why":"Supplies the Gaussian likelihood-ratio and generalized-likelihood detection framework on which $T_{\\mathrm{NP}}$ and $T_B$ are built.","marker":"(Scharf & Friedlander 1994)"},{"why":"Establishes the calibration-error treatment and the first-order aberration cancellation that justify the kernel-phase noise model.","marker":"(Ireland 2013)"},{"why":"Predicts the 16 nm RMS JWST wavefront drift used to estimate the magnitude of calibration residuals.","marker":"(Perrin et al. 2018)"},{"why":"Provides the ami_sim simulator used to generate the NIRISS F480M full-pupil frames and noise realizations.","marker":"(Greenbaum et al. 2016)"},{"why":"Supplies the apodization, recentering, and sampling prescription used to extract kernel phases from simulated images.","marker":"(Laugier et al. 2019)"},{"why":"Gives prior kernel detection limits for NIRCam and NIRISS that this paper's statistical framework extends and puts on a common footing.","marker":"(Sallum & Skemer 2019)"}],"fun_headline_variants":["JWST NIRISS kernel-phase test rivals theoretical detection bound","Hypothesis testing yields precise contrast limits for JWST NIRISS","Kernel-phase detection near ideal for JWST NIRISS companions","JWST NIRISS can see 10^4:1 companions at 200 mas ultimately","Neyman-Pearson test benchmarks JWST NIRISS companion detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quoted limits assume that after whitening, kernel-phase noise is Gaussian with a known covariance and that uncalibrated wavefront drift acts only as a diagonal variance inflation estimated from ten 16 nm RMS OPD maps; if real JWST calibration errors are correlated, time-varying, or larger than assumed, the contrast limits, especially the bright-target factor-of-ten degradation, will be worse.","fun_headline_variants_meta":{"raw":{"variants":["JWST NIRISS kernel-phase test rivals theoretical detection bound","Hypothesis testing yields precise contrast limits for JWST NIRISS","Kernel-phase detection near ideal for JWST NIRISS companions","JWST NIRISS can see 10^4:1 companions at 200 mas ultimately","Neyman-Pearson test benchmarks JWST NIRISS companion detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000737,"raw_usage":{"total_tokens":3461,"prompt_tokens":1280,"completion_tokens":2181,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":896,"completion_tokens_details":{"reasoning_tokens":2083}},"tokens_in":896,"tokens_out":2181,"duration_ms":15435,"temperature":1.0,"reasoning_tokens":2083,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:22:51.478063+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a real F480M full-pupil NIRISS observation of a binary with a $W2 \\approx 14.1$ primary and a confirmed companion at about 200 mas with contrast near $10^3$, using an 80-minute on-target sequence as simulated; if the operational test $T_B$ does not achieve $P_{\\mathrm{DET}} \\approx 68\\%$ at $P_{\\mathrm{FA}} = 1\\%$ for this companion, the predicted detection limit is not realized. A simpler numerical check is to simulate a correlated calibration term beyond the diagonal inflation and see whether $T_B$'s contrast limit falls more than a factor of 2.5 below the Neyman-Pearson bound.","supporting_citations":[{"cited_title":"2010, Astrophysical Journal, 724, 464","cited_arxiv_id":null,"evidence_quote":"Introduces kernel phases and the linear phase model $\\varphi = \\varphi_0 + A\\phi$ whose nullspace defines the kernel matrix $K$."},{"cited_title":"& Pearson, E","cited_arxiv_id":null,"evidence_quote":"Provides the likelihood-ratio test used as the theoretical upper performance bound for any detection test."},{"cited_title":"& Friedlander, B","cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian likelihood-ratio and generalized-likelihood detection framework on which $T_{\\mathrm{NP}}$ and $T_B$ are built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the calibration-error treatment and the first-order aberration cancellation that justify the kernel-phase noise model."},{"cited_title":"D., Pueyo, L., Rajan, A., et al","cited_arxiv_id":null,"evidence_quote":"Predicts the 16 nm RMS JWST wavefront drift used to estimate the magnitude of calibration residuals."},{"cited_title":"2016, ami_sim","cited_arxiv_id":null,"evidence_quote":"Provides the ami_sim simulator used to generate the NIRISS F480M full-pupil frames and noise realizations."},{"cited_title":"2019, Astronomy & Astrophysics, 494 Le Bouquin, J.-B","cited_arxiv_id":null,"evidence_quote":"Supplies the apodization, recentering, and sampling prescription used to extract kernel phases from simulated images."},{"cited_title":"& Skemer, A","cited_arxiv_id":null,"evidence_quote":"Gives prior kernel detection limits for NIRCam and NIRISS that this paper's statistical framework extends and puts on a common footing."}],"review_version":1}