{"id":"73dde0cc-c369-4844-9f2a-f4b8d336bd8d","arxiv_id":"2504.19870","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The paper presents the exoALMA calibration and imaging pipeline, with public scripts and data, designed to produce high-fidelity image cubes for kinematic searches for embedded planets.","lead":"The exoALMA collaboration describes how it turns raw ALMA observations of 15 planet-forming disks into clean images of gas emission. The pipeline's careful alignment, self-calibration, and masking steps aim to prevent spurious features that could be mistaken for hidden planets.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that the final cubes are artifact-free at the level of the targeted kinematic signatures rests on an end-to-end false-positive/recovery test that is described but not reported here; a null-injection CLEAN test would settle whether the 3-sigma iterative masking creates spurious…","rationale":"The paper is a careful, well-documented pipeline description with public data and scripts, and it is honest about the role of companion papers. The reader's conditional verdict is reasonable. However, the reader's identified weakest assumption, that the target-derived self-calibration model is free of the non-Keplerian structure being sought, is not the most load-bearing point because self-calibration is applied to pseudo-continuum measurement sets with line channels flagged; line kinematics are therefore not in the model used for gain solutions. The more direct load-bearing step is the line-imaging mask and CLEAN threshold logic in Section 4.4, where the data are used to define what is signal, and cleaning to 3xRMSfin inside the mask can sculpt noise into structure that mimics kinematic deviations. The paper's own language supports this reading: it emphasizes that any imaging artifact sets a noise floor on kinematic features, and it defers the decisive validation to an RML cross-check whose reference is TBD. The proposed null-injection test is the single check that would determine whether the central claim actually holds. My assessment does not change the reader's verdict: conditional acceptance is appropriate, with the missing end-to-end false-positive/recovery validation as the condition.","tokens_in":29252,"tokens_out":7678,"duration_ms":86755,"concrete_test":"Run a null-injection test on a representative source (e.g., MWC 758 12CO): take the calibrated visibilities, replace the sky model with a Keplerian-only rotating disk plus noise at the same uv coverage and channel width, and run the full Section 4.4 imaging pipeline (7xRMSinit mask, deep CLEAN to 3xRMSfin). Then apply the same residual-velocity or channel-map analysis used for candidate planet detections and count features exceeding the detection threshold. Repeat with an injected planet-like velocity kink of the expected amplitude to measure the recovery fraction. If the null-injection false-positive rate is comparable to or larger than the candidate rate claimed in companion papers, the central claim fails; if it is much smaller, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the released line cubes are of sufficiently high quality to search for faint deviations from Keplerian rotation. The most directly load-bearing step for that claim is not the continuum self-calibration (the reader's focus), because self-calibration is performed on pseudo-continuum measurement sets with the line channels flagged, so line kinematics are not directly in the gain-solution model. The more direct threat is the line-imaging procedure in Section 4.4: a shallow CLEAN to 7xRMSinit defines a morphology-based mask; a deep iterative CLEAN then runs inside that mask down to 3xRMSfin. Because the mask is derived from the target's own emission, any noise or sidelobe peak inside the mask that exceeds 3 sigma in a given channel is added to the CLEAN model and restored with the clean beam, producing channel-to-channel structure that can mimic the very kinks the program targets. Conversely, a real faint feature that falls outside the mask is left in the residual image convolved with the dirty beam, altering its apparent morphology. The paper states that 'empirical end-to-end testing' was performed and defers the detailed cross-check to Zawadzki & Czekala (2025), whose reference is listed as TBD, but it does not report false-positive rates, recovery fractions, or null-injection statistics within this manuscript. The Table 2 excess of achieved over theoretical RMS in several continuum images (e.g., DM Tau 17.1 to 26.6 uJy/beam; RXJ1615 11.7 to 19.1 uJy/beam) is an unexplained extra noise floor, but the absence of a pipeline-level false-positive test is the more direct threat to the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents the calibration and imaging pipeline developed for the exoALMA Large Program, which observed 15 protoplanetary disks in Band 7 with ALMA in three molecular lines (12CO, 13CO, CS). The pipeline includes pseudo-continuum construction, per-execution phase self-calibration, uv-plane spatial alignment of execution blocks, flux rescaling, group-level phase and amplitude self-calibration, production of fiducial measurement sets, and continuum and line imaging using Briggs weighting and an iterative CLEAN masking procedure. The manuscript reports quantitative improvements in signal-to-noise ratio, corrections of phase decoherence, and alignment demonstrations, and it releases the calibrated measurement sets, images, and scripts. The central claim is that the resulting data are of sufficiently high quality to search for faint deviations from Keplerian rotation in protoplanetary disks.","tokens_in":29574,"tokens_out":3690,"duration_ms":39219,"significance":"If the claim holds, the released exoALMA measurement sets and image cubes become a community resource for kinematic planet searches and disk dynamics studies. The paper makes several concrete methodological contributions: an uv-plane alignment method for disks without central peaks (Section 3.3), a decoherence-correction workflow with visible improvement in amplitude-versus-baseline and waterfall diagnostics (Figures 6-8), and a non-Keplerian masking strategy for imaging (Section 4.4). Strengths include the public release of scripts and data products, machine-checkable provenance via HISTORY and exoALMA header keywords, and quantitative SNR improvements (e.g., >300% in Section 3.5.2). The main weakness is that the central fidelity claim is deferred to validation that is described but not reported in this manuscript.","major_comments":[{"comment":"The central claim that the data are 'of sufficiently high quality to look for faint deviations from Keplerian rotation' is not supported within this manuscript by quantitative validation. The iterative masking procedure (shallow CLEAN to 7×RMSinit, mask formed by convolving the shallow model with a 0.5-0.7 arcsec Gaussian, then deep CLEAN inside the mask to 3×RMSfin) can create spurious channel-to-channel structure if noise or sidelobe peaks inside the mask exceed the cleaning threshold, and it can leave real faint emission outside the mask in the residual image. The paper mentions 'empirical end-to-end testing' (Section 4.1) and defers to Zawadzki & Czekala (2025), but that reference is listed as 'TBD' and no false-positive rates, recovery fractions, or null-injection statistics are reported here. A null-injection or recovery test on representative channels should be reported in this paper, or the summary claim should be tempered to what the presented tests actually establish.","section":"Section 4.4 and Section 5"},{"comment":"The continuum images show several cases where the achieved RMS is substantially above the theoretical RMS: DM Tau (17.1 to 26.6 uJy/beam), RXJ1615.3-3255 (11.7 to 19.1 uJy/beam), RXJ1604.3-2130 (16.1 to 23.0 uJy/beam), and V4046 Sgr (14.8 to 19.7 uJy/beam). Since these continuum images are used for alignment and self-calibration and are themselves science products, the excess noise and its likely origin (residual decoherence, weighting choices, or calibration errors) should be discussed. Without an explanation, the reader cannot assess whether the excess propagates to the line cubes or affects the reliability of kinematic measurements made from those images.","section":"Table 2"},{"comment":"The amplitude and phase self-calibration at the group level uses a CLEAN model of the target itself, cleaned down to 1 sigma for the amplitude step. Because this model is derived from the same source whose structure the exoALMA program aims to measure, real continuum asymmetries or point-source features could in principle be absorbed into the gain solutions and then applied to all spectral windows, including the line data. The paper does not directly test whether the amplitude self-calibration removes real source structure. A concrete test would be to compare line cubes and continuum images made with and without the amplitude self-calibration step, or with bright continuum features masked from the model, and to report the differences in image morphology and flux.","section":"Section 3.5.3"}],"minor_comments":[{"comment":"There is a typo in 'an number of different image conditioning or post-processing methods' that should read 'a number of different image conditioning or post-processing methods'.","section":"Section 4.1"},{"comment":"The phrase 'reduced by a cosifactor' should read 'reduced by a cos i factor' for clarity.","section":"Section 3.2"},{"comment":"Several companion papers are cited with 'ApJL, TBD' (Pinte 2025; Izquierdo et al. 2025; Zawadzki & Czekala 2025; Teague et al. 2025). Since the central claim of this paper leans on Zawadzki & Czekala (2025), the reference should be complete or the validation should be included here before acceptance.","section":"References"},{"comment":"The definition of the 'exoALMA' header keyword is clear, but it would help to state explicitly that the recorded quantities include the mask smoothing kernel FWHM and the cleaning thresholds, since these are the parameters most relevant to reproducing the images.","section":"Section 4.4"},{"comment":"The 4% flux-offset threshold is motivated only as 'empirically found not to affect the resulting image'; a brief statement of what test was performed to determine this threshold would improve reproducibility.","section":"Section 3.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central claim is reasonable and the pipeline is described in unusual detail, but the paper currently defers the load-bearing validation to a companion paper with 'TBD' status. I would require either reporting null-injection/recovery statistics for the line imaging procedure or explicitly narrowing the central claim to what the present tests demonstrate. The stress-test concern about false positives from the 3-sigma iterative masking is genuine and should be addressed before the paper can serve as the calibration reference for the exoALMA science results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, this is a genuinely useful methods paper: it lays out the full exoALMA calibration and imaging pipeline, releases all scripts and data, and is unusually candid about where the choices came from. Second, the paper's central claim—that the cubes are clean enough to search for faint embedded-planet kinks in CO kinematics—is not actually evidenced inside this paper. The authors say end-to-end testing was done and point to Zawadzki & Czekala (2025), which is listed as TBD. A reader who wants to trust the released data has to take that on faith.\n\nWhat is new: the complete workflow, the LB-first alignment bootstrapping, separating spatial alignment from flux scaling, and the Band 7 decoherence correction with short solution intervals. The alignment demonstration on J1604 (Fig. 5) is convincing, and the SNR improvements from self-calibration are large (>300% in some cases) and clearly shown. The paper also does the right thing by measuring the JvM epsilon and flagging where PSFs are non-Gaussian. That is honest, reproducible work.\n\nThe soft spots are real but not fatal. The stress-test note about the line imaging is the sharper one: the morphology-based mask thresholded at 3 sigma inside the mask can, in principle, turn noise peaks into CLEAN components that resemble kinematic kinks, and real faint features outside the mask get left in the dirty-beam residual. The paper says empirical end-to-end testing was performed but does not report false-positive rates, recovery fractions, or null-injection statistics here. That is the load-bearing gap. The continuum self-calibration concern the reader raised is less threatening because self-cal was done on pseudo-continuum with line channels flagged, so line kinematics are not directly in the gain model. Also minor: Table 2 shows several continuum images with achieved RMS well above theoretical (DM Tau 17.1 vs 26.6 uJy/beam) and the discrepancy is never discussed. Some procedural choices (the 4% flux threshold, the optimal order of operations) are justified qualitatively; that is okay in a methods paper but the key quantitative validations need to appear in the companion papers. The draft date post-dating the arXiv submission is an internal inconsistency that should be cleaned up.\n\nWho is this for? Anyone in the disk community using ALMA data, especially for kinematic studies. It deserves a real referee and likely conditional acceptance; the revisions should require either reporting the null-injection tests in this paper or citing the companion paper with results instead of TBD.","headline":"A careful, openly documented calibration/imaging pipeline for a major ALMA program; the core claim about kinematic fidelity rests on companion-paper tests not reported here, but the work is real and worth refereeing.","tokens_in":30345,"tokens_out":2066,"would_cite":true,"duration_ms":19754,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The exoALMA calibration and imaging pipeline produces data clean enough to reveal faint deviations from Keplerian rotation that may mark embedded planets in protoplanetary disks.","keywords":["protoplanetary disks","ALMA interferometry","self-calibration","visibility-plane alignment","non-Keplerian kinematics","molecular line imaging","CLEAN masking","embedded planets"],"falsifier":"Take a disk observed by the program, inject a synthetic non-Keplerian perturbation of known amplitude and location into the calibrated visibilities, run the full pipeline end to end, and compare the recovered perturbation to the input; if self-calibration or CLEAN masking removes or shifts the feature, the claim that the data are artifact-free at the target sensitivity fails.","tokens_in":29031,"feed_emoji":"🔭","tokens_out":8706,"duration_ms":76706,"temperature":0.7,"pith_summary":"This paper argues that the calibration and imaging procedures built for the exoALMA Large Program deliver interferometric data of sufficient fidelity to search for faint, localized departures from Keplerian rotation in 15 protoplanetary disks. The authors combine ALMA 12-meter and Atacama Compact Array observations, correct phase decoherence with iterative self-calibration, align separate observing blocks in the visibility plane, and image with masks that do not assume Keplerian motion. If the procedures work as claimed, then any non-Keplerian kinks seen in the released channel maps are real features of the disks rather than artifacts of data reduction. The paper also documents the order-of-operations choices and quality checks that make this claim testable.","feed_headline":"Pipeline clears ALMA data of artifacts that mimic embedded planets","feed_subtitle":"Self-calibration and spatial alignment let faint non-Keplerian motions show through in 15 protoplanetary disks.","key_machinery":"The load-bearing mechanism is the ordered calibration pipeline itself, centered on self-calibration using the target as its own model. Phase-only self-calibration on individual execution blocks corrects decorrelation; a uv-plane alignment routine shifts datasets by minimizing the weighted difference of gridded visibilities where uv coverage overlaps; flux alignment then rescales blocks that differ by more than four percent; and group-level phase then amplitude-and-phase self-calibration progressively concatenates ACA, short-baseline, and long-baseline data. On the imaging side, the key device is the iterative masking CLEAN procedure in which the CLEAN model is convolved with a wide Gaussian and thresholded to build a mask from the channel's own morphology, avoiding any Keplerian assumption that would hide the very signal being searched for. The 6-sigma threshold used to build models for phase self-calibration and the 1-sigma threshold for amplitude self-calibration determine how much real source structure is allowed into the gain solutions.","core_discovery":"The central claim is that the exoALMA self-calibration and imaging pipeline, described step by step here, yields measurement sets and image cubes in which faint deviations from Keplerian motion in protoplanetary disks can be confidently assessed as real. The pipeline first phase-self-calibrates each execution block separately, aligns blocks to a common phase center by minimizing weighted differences of gridded visibilities over overlapping uv cells, rescales fluxes when offsets exceed four percent, and then iteratively self-calibrates combined ACA, short-baseline, and long-baseline data before a final amplitude-and-phase round. Line images are CLEANed with masks derived from each channel's own emission morphology rather than from an assumed Keplerian model, so real non-Keplerian features are not masked out. The authors state that the resulting datasets are of sufficiently high quality to search for faint deviations from Keplerian rotation, with a companion non-CLEAN imaging study checking that the features are not artifacts of the deconvolution method.","pith_inferences":["A natural stress test not run in this paper would inject a synthetic non-Keplerian kink into the visibilities of a cleanly Keplerian disk and check that the pipeline recovers it without distortion; that would directly test whether self-calibration absorbs real signal.","The separation of spatial alignment from flux scaling, motivated by phase decoherence, suggests that future surveys should diagnose decorrelation before deciding the order of alignment and self-calibration steps.","Because the fiducial imaging was optimized for kinematic analysis, other science goals such as accurate total flux measurements of point sources in extended emission will need different weighting or post-processing; the released measurement sets permit that.","The caution about non-Gaussian PSFs and the epsilon metric recorded in image headers gives later users a quantitative handle on where flux measurements are unreliable."],"forward_implications":["Released measurement sets and image cubes can be used to search for embedded planets and other non-Keplerian dynamical features without redoing the calibration.","Kinematic deviations reported in companion analyses of these data can be treated as real; the pipeline's imaging choices also flag where caution is needed, such as low-SNR CS emission.","The order-of-operations findings provide a template for future ALMA large programs combining ACA, short, and long baselines at Band 7.","The public release of calibration scripts and quality-assurance figures allows other teams to reproduce or modify every step of the reduction."],"supporting_citations":[{"why":"Supplies the self-calibration strategy for disk observations that exoALMA adapts and extends.","marker":"Andrews et al. 2018"},{"why":"Provides the imaging framework, the JvM correction discussion, and the epsilon metric used to quantify non-Gaussian PSFs.","marker":"Czekala et al. 2021"},{"why":"Establishes the standard practice for self-calibration, including using the target as its own model.","marker":"Brogan et al. 2018"},{"why":"Introduces the uv-plane alignment concept that the spatial alignment step independently reproduces and adapts.","marker":"Casassus & Carcamo 2022"},{"why":"Supplies the iterative masking approach used to build CLEAN masks from emission morphology without assuming Keplerian structure.","marker":"Leroy et al. 2021a"},{"why":"The companion non-CLEAN imaging study that checks whether channel maps and non-Keplerian features are recovered without CLEAN artifacts.","marker":"Zawadzki & Czekala 2025"},{"why":"Introduces the non-Gaussian PSF restoration problem that drives the pipeline's weighting choices.","marker":"Jorsater & van Moorsel 1995"},{"why":"Delivers the ALMA pipeline-calibrated measurement sets that form the input to all subsequent calibration steps.","marker":"Hunter et al. 2023b"}],"fun_headline_variants":["ALMA pipeline scrubs artifacts to reveal true planet signs","New calibration clears fake planet signals in disk data","exoALMA pipeline disentangles real from spurious disk motions","Refined ALMA imaging exposes faint planet-induced wobbles","Data pipeline removes lookalike planet artifacts from disks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The gain solutions are computed from a model built by deconvolving the target's own emission, and the argument only works if that model contains none of the faint non-Keplerian motions being searched for; otherwise self-calibration would absorb the real signal into the antenna gains.","fun_headline_variants_meta":{"raw":{"variants":["ALMA pipeline scrubs artifacts to reveal true planet signs","New calibration clears fake planet signals in disk data","exoALMA pipeline disentangles real from spurious disk motions","Refined ALMA imaging exposes faint planet-induced wobbles","Data pipeline removes lookalike planet artifacts from disks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000236,"raw_usage":{"total_tokens":1505,"prompt_tokens":950,"completion_tokens":555,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":475}},"tokens_in":566,"tokens_out":555,"duration_ms":5001,"temperature":1.0,"reasoning_tokens":475,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:41:37.406371+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a disk observed by the program, inject a synthetic non-Keplerian perturbation of known amplitude and location into the calibrated visibilities, run the full pipeline end to end, and compare the recovered perturbation to the input; if self-calibration or CLEAN masking removes or shifts the feature, the claim that the data are artifact-free at the target sensitivity fails.","supporting_citations":[{"cited_title":"2025, , TBD","cited_arxiv_id":null,"evidence_quote":"The companion non-CLEAN imaging study that checks whether channel maps and non-Keplerian features are recovered without CLEAN artifacts."}],"review_version":1}