{"id":"1716636f-80c0-470e-aa06-e28e1d906201","arxiv_id":"2607.16664","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"About 6% of edge-on disk galaxies in DESI Legacy and Pan-STARRS imaging show tidal features at r-band surface brightness of 28.6 mag/arcsec², consistent with recent TNG50 simulations.","lead":"This paper measures how often edge-on disk galaxies show faint tidal features in deep survey images, using two large catalogs with over 19,000 galaxies. It reports a ~6% detection rate and argues that newer high-resolution simulations now match that rate better than earlier ones.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The ~6% incidence claim rests on single-group visual classification with unquantified systematic bias; an independent re-classification or injection test is required.","rationale":"The reader's verdict of CONDITIONAL is appropriate. My stress-test identifies the same load-bearing concern: the incidence fraction is a measurement, and its validity hinges entirely on the reliability of the visual classification. The paper is transparent about the single-group bias but does not quantify it; the borrowed 3% disagreement rate addresses random scatter, not systematic bias. The two-catalog consistency is reassuring but does not remove the shared-bias issue because both were classified by the same team. The depth-correction to ~9.6% is a separate but related warning: it shows that the detection rate is sensitive to depth, and the extrapolation from 10.4% coverage is itself uncertain. A concrete independent re-classification of a subset and a mock-injection test would settle whether the 6% is biased. These are exactly the kinds of checks that would either validate or adjust the headline, so the conditional recommendation stands. I do not see an internal inconsistency or a reason to reject; the paper's own limitations statement (Section 6.2) is honest, and the comparison to simulations is appropriately hedged. Thus, UNCHANGED relative to the reader's CONDITIONAL verdict.","tokens_in":37594,"tokens_out":3350,"duration_ms":37638,"concrete_test":"Select a random subset of ~300 galaxies from EGIS and ~300 from EGIPS, stratified by detection status and cirrus contamination, and have an independent team classify the same processed images using the same taxonomy without access to original labels. Compute inter-group agreement (e.g., Cohen's kappa) and the difference in overall incidence. Separately, inject realistic mock tidal features (e.g., from TNG50 mock images convolved to DESI PSF and noise) into 100–200 blank DESI fields at surface brightness levels matching the reported detections (median ~25.8–27.1 mag/arcsec^2) and run the original visual pipeline to measure detection completeness and false-positive rate. If the external incidence differs by more than ~1 percentage point or the injection recovery is below ~80% for features at the reported brightness, the 6% claim should be revised; if it passes, the central result is robust","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (6.4% EGIS, 6.2% EGIPS) is produced entirely by visual classification performed by one research group. The paper explicitly concedes that 'the use of a single research group introduces the possibility of shared biases, which are difficult to quantify directly' (Section 6.2). The only uncertainty adopted is the ~3% inter-observer disagreement rate from Bridge et al. (2010), which is not a measure of systematic sensitivity at DESI depth, for edge-on galaxies, or for this specific taxonomy. The two catalogs were classified by the same team with the same training, so the agreement between EGIS and EGIPS does not control for a shared detection threshold. A systematic tendency to miss diffuse features would raise the true fraction; a tendency to over-interpret noise/cirrus would lower it. The paper's own HSC comparison shows that just 10.4% coverage yielded 19 additional detections and a depth-corrected estimate of ~9.6±1.4%, a ~60% relative increase over the headline 6% — demonstrating that the measured fraction is sensitive to imaging depth and/or classification choices. The claimed consistency with TNG50 (~7%) is thus contingent on an unvalidated detection sensitivity. Without an independent blind re-classification or a mock-injection completeness test, the headline number is only as robust as the unquantified biases of a single group.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a statistical study of low-surface-brightness (LSB) tidal structures in two large samples of edge-on disk galaxies: 5606 EGIS galaxies and 14,237 EGIPS galaxies, using DESI Legacy Imaging Surveys data, supplemented by HSC-SSP and targeted APO follow-up. After homogeneous image processing optimized for faint diffuse emission, tidal features are identified through visual inspection and classified into standard morphological categories. The authors report detection rates of 5.8% (EGIS) and 4.8% (EGIPS) in the full samples, and 6.4% and 6.2% in completeness-limited subsamples. They extrapolate that full HSC coverage would increase the fraction to 9.6±1.4% (EGIS) or 9.4±1.4% (EGIPS). The central claim is that at a typical DESI r-band surface-brightness depth of 28.6 mag arcsec^-2, the incidence of LSB tidal structures is ~6%, consistent with some recent high-resolution simulations (e.g., TNG50 at ~7%) but lower than many earlier simulation predictions of 20–40%. The paper discusses projection, classification, redshift, and stellar-mass biases, and compares with observational and theoretical literature.","tokens_in":1791,"tokens_out":1656,"duration_ms":53643,"significance":"If the ~6% incidence figure is robust, this would be an important anchor for galaxy evolution studies: a large, homogeneous census of tidal structures in edge-on galaxies, with two independent catalogs, a public classification catalog, and a direct comparison to modern cosmological simulations. The paper also highlights the critical role of imaging depth and galaxy mass in measuring merger/tidal statistics, and the APO follow-up demonstrates that shallow survey images can miss or fragment extended features. The public catalog and reproducible processing pipeline are useful community resources. The main significance rests on the accuracy of the visual classification and the representativeness of the depth corrections; these are the weakest points and need to be addressed before the headline number can be taken at face value.","major_comments":[{"comment":"The headline incidence rates (6.4% EGIS, 6.2% EGIPS) are produced entirely by visual classification by a single research group. The paper acknowledges in Section 6.2 that shared biases are difficult to quantify, but the only uncertainty quoted is a ~3% inter-observer disagreement rate from Bridge et al. (2010), which is not a measure of systematic sensitivity at DESI depth, for edge-on galaxies, or for this taxonomy. The two samples were classified by the same team, so their agreement does not control for a shared threshold. Because the central claim is a specific absolute fraction, the absence of an independent blind reclassification or a mock-injection completeness test is load-bearing. At minimum, the headline should be presented as a lower limit with an explicit statement of the unknown systematic bias.","section":"Section 5.2, Section 6.2, Table 3"},{"comment":"The HSC-SSP depth comparison uses only 10.4% of the EGIS sample (597 galaxies) to derive 19 additional detections and an extrapolated full-coverage fraction of 9.6±1.4%, a ~60% relative increase over the headline 6.4%. This extrapolation assumes the HSC-overlap subset is representative of the entire sample in galaxy mass, redshift, and morphology; no test of this assumption is provided. If the HSC footprint preferentially covers more massive or lower-redshift galaxies, the extrapolation would be biased. This uncertainty directly affects the interpretation of the result as a measurement at a fixed DESI depth versus an intrinsic incidence, and the paper should state clearly that the observed 6–7% is a depth-limited lower limit, not a corrected incidence.","section":"Section 6.2, depth extrapolation"},{"comment":"The image-processing choices for DR10 vs DR9 and PDR3 vs PDR2 are described qualitatively: for 'a subset of cases' with overly aggressive sky subtraction, the authors reverted to earlier data releases, but the number of such cases, the selection criteria, and the effect on the final images are not quantified. This makes the effective depth and the uniformity of the processed sample unclear and could affect which diffuse structures are visible. The authors should provide statistics on how many images were reverted, how the decision was made, and ideally a comparison of classifications with and without the reversion.","section":"Sections 4.1, 4.2, and 4.3"},{"comment":"The comparison with TNG50 at ~7% (Miró-Carretero et al. 2025) is used as a key conclusion, but the manuscript does not state whether the simulation mock sample matches the completeness-limited EGIS/EGIPS selection in stellar mass, redshift, inclination, or surface-brightness limit, nor whether the same classification taxonomy and visual-inspection procedure were applied. Without this information, the formal agreement between 6.4% and ~7% could be fortuitous. The authors should either specify the selection and detection methodology of the simulation comparison or soften the claim of quantitative agreement.","section":"Section 6.6, comparison to simulations"}],"minor_comments":[{"comment":"The field of view of ARCTIC is written as '7.85 arcmin2'; use 'arcmin^2' (or square arcminutes) for consistency with the rest of the paper.","section":"Section 4.3"},{"comment":"The title of Aihara et al. (2022) appears to have a typo: 'Ublications' should be 'Publications'.","section":"Reference [15]"},{"comment":"The Mann–Whitney U test p-values are reported as p=0.355 (EGIS) and p=6.4e-4 (EGIPS). The differing significance is not discussed; a brief interpretation of why EGIS does not show a significant redshift offset while EGIPS does would help the reader.","section":"Section 6.3"},{"comment":"The caption states the average difference between observed and intrinsic surface brightness is '~0.05 for EGIS' without specifying units; please clarify whether this is in mag arcsec^-2 and whether the same applies to the EGIPS sample (which is not shown in the figure).","section":"Figure 8"}],"recommendation":"major_revision","confidential_remarks":"The core measurement is potentially valuable and the paper is generally well structured, but the headline incidence fraction is under-validated because it rests on single-group visual classification without a quantified detection-sensitivity function. The depth-corrected estimate of ~9.6% underscores that the observed ~6% is likely an underestimate. I would encourage the editor to request the authors add an independent blind reclassification and/or a mock-injection test, or at minimum reframe the central claim as a lower limit with an explicit systematic caveat. The simulation comparison also needs a clearer statement of the selection function. The paper is within the journal's scope and the data release is a useful asset, so major revision is appropriate rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is what I would tell you about this paper. It measures the incidence of low-surface-brightness tidal structures in two large edge-on galaxy catalogs, EGIS (~5600 galaxies) and EGIPS (~14,000), using DESI Legacy Survey imaging. The central result is a ~6% detection rate at r-band depth ~28.6 mag arcsec-2 (6.4% for the completeness-limited EGIS subsample, 6.2% for EGIPS). That is the largest census of this kind to date, and the two catalogs were processed and classified independently enough that their close agreement is real evidence the measurement is stable.\n\nWhat it does well: careful image processing aimed at faint diffuse emission, a serious attempt to screen galactic cirrus (including IR cross-checks), a published catalog, and a small APO follow-up that demonstrates how depth changes what you see. The comparison with earlier observational studies and with simulations is level-headed; the consistency with the HSC-based CNN study of Kado-Fong et al. (5.6%) and with recent TNG50 mock observations (~7%) is a useful data point for the field. The authors also present the stellar-mass dependence, which fits the known picture.\n\nThe soft spot, as the paper itself acknowledges in Section 6.2, is that the classification is visual and done by a single research group. The only quantitative uncertainty adopted is a ~3% inter-observer disagreement rate borrowed from Bridge et al., which is not a measure of sensitivity at DESI depth, for edge-on systems, or for this taxonomy. The EGIS-EGIPS agreement does not control for a shared detection threshold since the same team with the same training classified both. And the internal HSC comparison shows the fraction would rise to ~9.6% with full deeper coverage — a 60% relative increase on the headline number. That tells you the 6% is a statement about DESI-depth detection, not about the true incidence of tidal features. The depth correction from the 10.4% HSC subset is an extrapolation, done honestly but with wide uncertainty. The DR9/DR10 sky-subtraction choices are transparent but under-specified.\n\nNone of this invalidates the paper. The claim ‘about 6% at DESI-like depth’ holds. The claim ‘this is the true cosmic incidence’ would need an independent blind reclassification or an injection/recovery test. The authors know this and say so. I would send this to referees; the community needs the large-sample measurement on record, and the limitations are addressable in a revision or follow-up work. For your own reading, it is worth a skim if you care about tidal features or LSB imaging.","headline":"Solid large-sample measurement of tidal feature incidence in edge-on galaxies; headline ~6% is well-supported at DESI depth but remains hostage to single-group visual classification, a caveat the authors themselves state.","tokens_in":38401,"tokens_out":3146,"would_cite":true,"duration_ms":32825,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"At a typical r-band depth of 28.6 mag arcsec⁻², roughly 6 percent of edge-on disk galaxies show tidal debris, and the fraction climbs with stellar mass — a result that matches modern high-resolution simulations but not older ones.","keywords":["tidal structures","low surface brightness","edge-on galaxies","galaxy interactions","cosmological simulations","galactic cirrus","visual classification","merger debris"],"falsifier":"Take a random subset of ~500 galaxies from the EGIS complete sample and have them re-classified blind by an independent team (or by the same team after a wash-out period), using the same processed images; if the disagreement in 'any tidal feature' exceeds ~3% and shifts the incidence outside ~4–8%, the 6% estimate is not robust. More decisive: inject synthetic tidal features of known surface brightness into the DESI images and measure the recovery fraction; if the recovery fraction at 28.6 mag arcsec⁻² is well below ~0.6 for the faintest detectable features, then the observed 6% is a detection","tokens_in":37491,"feed_emoji":"🌌","tokens_out":4786,"duration_ms":47672,"temperature":0.7,"pith_summary":"The paper tries to establish a reliable census of low-surface-brightness tidal structures around edge-on disk galaxies using the two largest catalogs available, processed with a homogeneous pipeline and visually classified. It finds that only about 6% of galaxies in completeness-limited subsamples of both catalogs host such features, despite earlier simulations predicting 20–40%. The paper argues this low fraction is real, not a detection failure: it rises with stellar mass (11–15% for galaxies above 10^10.5 solar masses), is consistent with previous surveys of similar depth, and matches the latest high-resolution cosmological simulations. If correct, it means survey depth, galaxy mass, and numerical resolution — not just interaction history — determine how often we see tidal debris, and that current census numbers are lower limits.","feed_headline":"Only 6% of edge-on galaxies show tidal debris at survey depth","feed_subtitle":"Two independent catalogs agree, closing the gap between deep-sky counts and simulated merger debris.","key_machinery":"The key machinery is the combination of (1) two large, independently constructed catalogs of edge-on disk galaxies (EGIS, EGIPS) selected by morphology and separately completeness-limited; (2) a uniform image-processing pipeline optimized for faint diffuse emission, with careful artifact and cirrus handling; and (3) a visually defined taxonomy of tidal features (tails, streams, shells, plumes, fans, bridges, arcs, loops, satellite debris) that is deliberately lumped into a single 'has any tidal feature' statistic. The quantitative anchor is the r-band surface-brightness depth of 28.6 mag arcsec⁻², the same depth as the Stripe 82 pilot survey and the depth at which mock observations from mode","core_discovery":"The central claim is that at an r-band surface-brightness depth of ~28.6 mag arcsec⁻², the incidence of LSB tidal features around edge-on disk galaxies is about 6% (6.4% in the EGIS completeness-limited subsample and 6.2% in EGIPS), not the 20–40% claimed by many older simulations. The paper shows that restricting to complete subsamples yields consistent values across two independent catalogs, that the fraction increases steeply with stellar mass, and that deeper HSC/APO data reveal additional features, implying the observed fraction is a lower limit. The paper concludes that modern high-resolution simulations with realistic mock observations reproduce the observed incidence, whereas older,","pith_inferences":["If survey depth is the dominant lever, then the ~6% value should not be treated as an intrinsic merger-rate measurement; instead, the paper's own extrapolation suggests the intrinsic incidence could be closer to ~10% or higher at LSST depths, and the mass-dependent trend will be sharpened.","A testable extension: run the same pipeline on mock images with known injected tidal features and measure the recovery fraction as a function of surface brightness and stellar mass; this would convert the visual-classification fractions into a completeness-corrected incidence, which the paper does not provide.","The convergence between the 6% observation and modern simulation predictions implies that older 20–40% predictions were inflated by numerical resolution and simplified physics, which in turn suggests that simulations should now be used to predict the mass- and redshift-dependence of tidal feature visibility rather than just the mean fraction.","Another consequence: if deeper data preferentially reveal coherent streams and shells in galaxies already flagged as feature hosts, then the morphological classification (tails vs shells vs loops) may be less meaningful than the paper assumes; lumping all categories into one 'any feature' statistic is a reasonable first step, but a physical decomposition awaits kinematic follow-up."],"forward_implications":["If the ~6% figure is right, then at current survey depth only about 1 in 16 edge-on disk galaxies shows detectable merger debris, making low-surface-brightness tidal features a minority phenomenon at z~0.05.","The measured increase of tidal fraction with stellar mass (to 11–15% above ~10^10.5 M_sun) means any fair comparison between surveys or simulations must control for stellar mass, not just depth.","The leap from ~6% to ~9.6% when extrapolating to HSC-like depth implies that forthcoming deeper surveys (e.g., LSST) should recover significantly more tidal features, and that published fractions are lower limits.","The agreement between the two independent catalogs and with modern high-resolution simulations supports the view that realistic galaxy-formation physics suppresses long-lived, easily detectable tidal debris, contrary to earlier theoretical predictions.","The deeper APO follow-up showing hidden extensions and new structures in individual galaxies suggests that current classification yields only a partial view of the outer stellar envelope, motivating deeper imaging of complete samples."],"fun_headline_variants":["Tidal debris found in just 6% of edge-on galaxies","Deep survey counts of tidal features: ~6%, not 20-40%","Edge-on galaxies: tidal debris rate matches modern simulations","New tally: 6% of edge-on galaxies have tidal tails","Tidal feature rate in edge-on disks: 6% at 28.6 mag/arcsec²"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire census rests on the assumption that visual inspection by one research group, with a ~3% disagreement rate borrowed from another study, reliably detects and correctly rejects tidal features at 28.6 mag arcsec⁻²; if shared biases systematically miss faint features (or flag cirrus/artifacts), the 6% fraction shifts, and the paper's own extrapolation to deeper data indicates the true fraction could be higher.","fun_headline_variants_meta":{"raw":{"variants":["Tidal debris found in just 6% of edge-on galaxies","Deep survey counts of tidal features: ~6%, not 20-40%","Edge-on galaxies: tidal debris rate matches modern simulations","New tally: 6% of edge-on galaxies have tidal tails","Tidal feature rate in edge-on disks: 6% at 28.6 mag/arcsec²"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000172,"raw_usage":{"total_tokens":1154,"prompt_tokens":831,"completion_tokens":323,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":237}},"tokens_in":575,"tokens_out":323,"duration_ms":3708,"temperature":1.0,"reasoning_tokens":237,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T20:15:25.600135+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random subset of ~500 galaxies from the EGIS complete sample and have them re-classified blind by an independent team (or by the same team after a wash-out period), using the same processed images; if the disagreement in 'any tidal feature' exceeds ~3% and shifts the incidence outside ~4–8%, the 6% estimate is not robust. More decisive: inject synthetic tidal features of known surface brightness into the DESI images and measure the recovery fraction; if the recovery fraction at 28.6 mag arcsec⁻² is well below ~0.6 for the faintest detectable features, then the observed 6% is a detection","supporting_citations":[],"review_version":1}