{"id":"0b7fab6a-054e-49e5-8714-e64fd1c2bcd7","arxiv_id":"2607.28739","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Volunteer classifications of 49,703 galaxies yield ~6,700 ram-pressure-stripping candidates, with candidate fractions rising from ~10% in groups to ~20-30% in massive clusters.","lead":"Volunteers on the Zooniverse platform sorted nearly 50,000 galaxy images and flagged about 6,700 'jellyfish' galaxies losing gas as they fall into clusters. This provides the largest homogeneous sample of ram-pressure-stripping candidates to date and confirms stripping is more common in massive clusters.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Calibration of Eq. (3) is in-sample and single-cluster; reported purity/completeness may not transfer to the full 79-cluster catalogue.","rationale":"The reader's conditional verdict is appropriately anchored to the transfer of thresholds from Abell 1644 to the full sample. I agree with that assessment, and I sharpen it further: the threshold optimisation is performed and evaluated on the same 403 galaxies, so even the reported P/C on Abell 1644 are in-sample figures. This is the most load-bearing concern because every quantitative result—the 6,739-candidate catalogue, the environmental fractions in Figs. 7-8, and the claim that citizen scientists can reliably identify RPS—depends on the validity of Eq. (3) when applied to clusters that differ from Abell 1644. The paper does have independent support: the literature cross-match in Fig. 3 shows that known RPS galaxies receive higher Fdist and Ftail, and the authors test a stricter Fdist>0.5 threshold. But neither step is a held-out purity/completeness validation. There are also smaller internal inconsistencies (e.g., the abstract states a 53% higher median clustercentric velocity while Sect. 4.3 reports ~20%; the text varies between 49,703 and 42,684 objects; the number of clusters is given as 79 and 82), which further indicate that the manuscript should be carefully revised. None of these, however, changes the conditional verdict: the catalogue is potentially very useful, but its global calibration must be demonstrated on at least one additional cluster before the catalogue is used for quantitative science.","tokens_in":24171,"tokens_out":8086,"duration_ms":87448,"concrete_test":"Apply Eq. (3) thresholds to expert labels in a second cluster not used for calibration—e.g., Abell 1367 (z=0.021, log M~14.53) using existing WINGS/OmegaWINGS expert classifications—and compute P and C. If either differs by more than about 0.1 from the Abell 1644 values (P=0.63, C=0.57), the single-cluster calibration is not transferable; the published catalogue should then be re-calibrated or the global purity/completeness should be reported as unvalidated outside Abell 1644.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The catalogue's reliability rests on Eq. (3): Fdist>=0.39 and (Fmerg<=0.23 or Ftail>=0.37), with P=0.63 and C=0.57. These thresholds were selected by Monte Carlo search and evaluated on the same 403 expert-labelled galaxies in Abell 1644 (Sect. 3.4); no sample splitting, cross-validation, or an independent held-out cluster is described. With only 403 tuning objects, the reported P/C are in-sample estimates and are likely optimistic. The thresholds are then applied unchanged to all 79 clusters spanning z~0.003-0.056 and log M~13.3-15.2, assuming that volunteer voting behaviour and the visibility of low-surface-brightness tails in DECaLS images are identical across clusters of different redshift, mass, and image quality. The cross-match with known RPS galaxies (Fig. 3) is encouraging, but it is a distributional comparison, not a quantitative purity/completeness measurement on a separate cluster; the known-RPS literature sample is itself heterogeneous. If P/C degrade elsewhere—for example, because higher redshift suppresses Ftail, or because merger/non-merger discrimination behaves differently—the derived environmental fractions (Figs. 7-8) and the claim of a 6,739-candidate homogeneous catalogue are not supported at the stated accuracy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents 'Fishing for Jellyfish Galaxies', a Zooniverse citizen-science project that visually classifies 49,703 late-type galaxies in the fields of 79 clusters/groups from DECaLS. Volunteers voted on disturbance, merger, and tail-like features; the vote fractions were debiased and calibrated against expert labels for all 403 galaxies in Abell 1644. The adopted thresholds in Eq. (3) yield a claimed purity P=0.63 and completeness C=0.57, and are applied to the full sample to produce 6,739 ram-pressure-stripping candidates (3,910 with tails), 5,430 merger candidates, and 29,729 undisturbed galaxies. The authors further report that the candidate fraction rises from ~10% in groups to ~20–30% in massive clusters, that candidates have bluer colours, elevated star-formation rates, and distinct phase-space positions, and that the catalogue is intended as a homogeneous resource for future RPS studies.","tokens_in":24556,"tokens_out":4015,"duration_ms":42663,"significance":"If the calibration transfers to the full 79-cluster sample, this would be the largest homogeneous visually selected catalogue of ram-pressure-stripping candidates to date, and a valuable training set for automated methods. The project is well matched to the journal's scope and has several clear strengths: the calibration procedure and vote-fraction definitions are described transparently; the catalogue is released; and the physical checks (phase-space, SFR, colour) are not circular, because the volunteers were not given any physical parameters. The literature cross-match in Fig. 3 and the internal consistency of the SFR/colour trends provide encouraging support for the method. However, the central quantitative claims — the size of the candidate catalogue and the environmental fractions — rest on an in-sample, single-cluster calibration that is not independently validated.","major_comments":[{"comment":"The thresholds Fdist≥0.39 & (Fmerg≤0.23 or Ftail≥0.37) are selected by a Monte Carlo search and then evaluated on the same 403 expert-labelled galaxies in Abell 1644. No sample splitting, cross-validation, or independent held-out cluster is reported, so P=0.63 and C=0.57 are in-sample estimates and are likely optimistic. These thresholds are then applied unchanged to all 79 clusters spanning z~0.003–0.056 and log M~13.3–15.2, implicitly assuming that volunteer voting behaviour and the visibility of low-surface-brightness features are constant across redshift, mass, and image quality. The cross-match with literature RPS galaxies in Fig. 3 is a distributional comparison, not a quantitative purity/completeness measurement on a separate cluster. I request an external validation (e.g., a second cluster with expert labels, or a ‘known-RPS’ sample treated as a held-out set) or, failing that, th","section":"Sect. 3.4, Eq. (3), and Fig. D.1"},{"comment":"The abstract claims a 'median clustercentric velocity 53% higher than the general cluster population', but Sect. 4.3 reports only a ~20% shift in absolute line-of-sight velocities (with K-S p-values of 0.0032 and 0.0007). No 53% value appears in the body text, and the paper does not explain how a median clustercentric velocity is derived from line-of-sight data. This is a direct numerical inconsistency in a headline result. The authors should either substantiate the 53% figure with the relevant calculation or correct the abstract.","section":"Abstract vs. Sect. 4.3"},{"comment":"The two figures use different definitions of the candidate fraction: Fig. 7 divides by all classified subjects in the cluster, while Fig. 8 divides by star-forming galaxies with M*>10^9.7 in a spectroscopic/SDSS-overlap subset. The text states both show the same trend, but the denominators and sample restrictions differ, so the quantitative fractions (10% vs 20–30%) are not directly comparable and may be affected by the different selection functions. The paper should state which definition is used in each figure and whether the trend with halo mass is robust to the denominator choice.","section":"Sect. 4.2, Figs. 7 and 8"},{"comment":"The text says SC galaxies are at larger cluster-centric radii than undisturbed galaxies by ~10%, which is in the opposite direction of the usually quoted expectation for RPS galaxies (lower cluster-centric radii, e.g., Jaffé et al. 2015). The discussion interprets the velocity difference as consistent with recent infall but does not address the radial discrepancy. If this offset is real it deserves comment; if it is a selection effect of the 4xR500 aperture or the spectroscopic membership cut, that should be stated explicitly.","section":"Sect. 4.3, Fig. 9"}],"minor_comments":[{"comment":"The abstract/conclusions state that classifications were obtained 'across 82 clusters', while Sect. 2.2 and Tables 1–2 enumerate 79 clusters. Please reconcile the number.","section":"Abstract and Conclusions"},{"comment":"The caption says the dashed lines mark thresholds '(0.39, 0.77, 0.37)', but Eq. (3) is written as Fdist≥0.39 & (Fmerg≤0.23 or Ftail≥0.37). Since 0.77 = 1−0.23, the caption is using the non-merger fraction while the text uses the merger fraction; make the relation explicit.","section":"Caption of Fig. D.1"},{"comment":"Abell 671 appears in both Table 1 and Table 2, with 186 and 11 galaxies respectively. The Table 2 note says bold entries indicate additional coverage of DR9 clusters, but Abell 671 is not bold. Clarify whether this is a duplicate entry or a deliberate additional sample.","section":"Tables 1 and 2"},{"comment":"There is a typo in 'Fig., D.1' and the sentence 'Whilst the purity level... indicates that this sample is subject to contamination' is followed by a direct comparison to a stricter threshold; consider moving the stricter-threshold test into a dedicated paragraph.","section":"Sect. 3.4"}],"recommendation":"major_revision","confidential_remarks":"This is a promising citizen-science methodology paper with a useful public catalogue. My main concern is not the method itself but the strength of the claims relative to the evidence: the purity/completeness values are in-sample and no independent validation cluster is presented. This is fixable and I would be willing to consider a revised version. The abstract/body inconsistency on the 53% velocity claim should also be resolved before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nHere's the short version: this is the first dedicated citizen-science search for ram-pressure-stripped galaxies on real imaging data, and the catalogue itself is a real resource. But the headline numbers in the abstract don't all match the body, and the calibration that everything rests on is tuned in-sample on one cluster. It deserves a serious referee, not a desk reject.\n\nWhat's new: they ran a Zooniverse project where volunteers looked at 49,703 late-type galaxies in DECaLS, with a workflow specifically designed to separate disturbed-but-not-merging galaxies from mergers and undisturbed ones. That's genuinely different from previous work, which either used simulations (Zinger et al. 2024) or repurposed Galaxy Zoo oddities (Crossett et al. 2025). The final catalogue of 6,739 stripping candidates with vote fractions for every object is a useful public resource.\n\nThey also do some things honestly. The calibration against 403 expert-labelled galaxies in Abell 1644 is described transparently, with purity 0.63 and completeness 0.57 stated plainly. They test a stricter threshold and show the trends don't change. The phase-space, colour, and SFR checks are not circular because the classifiers never saw physical parameters. The cross-match with known RPS galaxies from the literature is distributionally reassuring.\n\nNow the soft spots, in proportion.\n\nFirst, the calibration is single-cluster and in-sample. The thresholds in Eq. 3 were chosen by Monte Carlo search over the same 403 galaxies used to evaluate them. No held-out cluster, no cross-validation. Purity and completeness are likely optimistic estimates of what will happen when the same thresholds are applied to 79 clusters spanning z ~ 0.003–0.056 and a range of image quality. The assumption that volunteer voting behaviour transfers is plausible but untested. This is the kind of thing that could be fixed with an independent validation cluster, and it should have been in this paper.\n\nSecond, there's a factual inconsistency: the abstract claims a median cluster-centric velocity 53% higher than the cluster population, but Section 4.3 reports a ~20% shift in line-of-sight velocities. Those are not the same thing, and the abstract doesn't say which sample or what baseline. This needs to be resolved before the numbers are quoted.\n\nMinor things: the conclusions give 4,400 volunteers and 82 clusters while the body says 5,286 users and 79 clusters. Small but sloppy.\n\nOverall: the central idea holds up; the catalogue is valuable; the environmental trends are confirmatory and not overinterpreted. The paper is exactly what a pilot should be, except for the missing validation step and the abstract/body mismatch. A solid referee can sort those out.\n\nMy advice: send it to review. The catalogue is worth having in the literature, and the calibration issue is fixable.","headline":"A genuinely useful citizen-science RPS catalogue with transparent calibration, but the single-cluster in-sample threshold tuning and a 53% vs 20% velocity discrepancy mean the quantitative claims need a careful second look before the catalogue is used for science.","tokens_in":25072,"tokens_out":2355,"would_cite":true,"duration_ms":23251,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that thousands of untrained volunteers, each casting votes on galaxy images, can reliably identify galaxies whose gas is being stripped by the hot intracluster medium, producing a catalogue of 6,739 jellyfish candidates—th","keywords":["ram-pressure stripping","jellyfish galaxies","citizen science","galaxy clusters","morphological classification","vote fractions","galaxy evolution","Zooniverse"],"falsifier":"Classify a second cluster with full expert labels (e.g., Abell 1367) and measure the purity and completeness of the citizen criteria; if they fall far from P=0.63 and C=0.57, the transferability assumption fails. Alternatively, take a random subset of the 6,739 candidates and search for extraplanar ionized gas with integral-field spectroscopy; a low confirmation rate would weaken the physical interpretation.","tokens_in":24131,"feed_emoji":"🪼","tokens_out":4937,"duration_ms":52772,"temperature":0.7,"pith_summary":"The paper tries to show that crowd-based visual inspection can find galaxies undergoing ram-pressure stripping across a large, homogeneous sample, and that the resulting candidates behave as expected. Volunteers classified 49,703 late-type galaxies in 79 clusters; their votes were turned into fractions and calibrated against expert labels in one cluster to define stripping candidates. The final catalogue contains 6,739 candidates, and the fraction of such galaxies rises from roughly 10% in groups to 20–30% in massive clusters. The candidates are kinematically and photometrically distinct: they have higher line-of-sight velocities, bluer colours, and elevated star-formation rates. A sympathetic reader would care because this provides the largest uniform sample for studying how dense environments transform galaxies, and it demonstrates that citizen science can deliver scientifically usable morphological labels.","feed_headline":"Crowd scientists net 6,739 jellyfish galaxy candidates","feed_subtitle":"Volunteers sorted 50,000 galaxy images; stripped candidates jump from ~10% in groups to 20-30% in massive clusters.","key_machinery":"The central mechanism is the vote-fraction threshold given in Eq. (3): a galaxy is a stripping candidate when F_dist ≥ 0.39 and simultaneously F_merg ≤ 0.23 or F_tail ≥ 0.37. These thresholds were tuned with a Monte Carlo sweep over the expert-labelled Abell 1644 sample to balance purity and completeness, and they are then applied unchanged to the rest of the sample. The debiasing step—downweighting volunteers who classified fewer than ten galaxies—sharpens the vote fractions slightly but is not critical once the thresholds are tuned.","core_discovery":"By aggregating approximately ten independent volunteer classifications per galaxy into vote fractions (disturbed, merging, tail) and calibrating thresholds against expert labels of 403 galaxies in Abell 1644, the authors derive a simple selection rule—F_dist ≥ 0.39 and (F_merg ≤ 0.23 or F_tail ≥ 0.37)—that identifies ram-pressure-stripping candidates with purity 0.63 and completeness 0.57. Applied to all 79 clusters, this yields 6,739 stripping candidates (3,910 with prominent tails), 5,430 merger candidates, and 29,729 undisturbed galaxies. The candidate fraction rises from ~10% in groups to ~20–30% in massive clusters, and the candidates show higher velocities, bluer g−r colours, lower Sér","pith_inferences":["We infer that the vote-fraction thresholds could be re-calibrated per cluster with a modest number of expert labels per cluster, which would likely improve global purity and completeness beyond the single-cluster calibration; the released catalogue makes such a test straightforward.","The 5,430 merger candidates offer a way to separate gravitational from hydrodynamical transformation statistically: if the stripping fraction rises with cluster mass while the merger fraction does not, that contrast directly shows the two mechanisms respond differently to environment—an analysis the paper only sketches.","Since the project used only optical broadband images, we infer that adding H-alpha or UV data for the tail-bearing candidates would test directly whether the tails are sites of ongoing star formation, a signature the paper measures in aggregate via SDSS but not for individual tails."],"forward_implications":["The catalogue of 6,739 candidates triples or more the number of visually identified ram-pressure-stripping galaxies available for statistical study.","The rising candidate fraction with cluster mass strengthens the evidence that ram-pressure stripping is a major transformation channel in massive clusters, not just a rare phenomenon.","The candidates' higher velocities and bluer colours support the picture that they are recent infallers on radial orbits, as predicted by simulations of ram-pressure stripping.","The public release of 37,599 visually classified galaxies provides a resource for future studies of galaxy transformation, mergers, and cluster environments.","The sample can serve as a training set for automated machine-learning classifiers aimed at finding stripping candidates in even larger surveys."],"fun_headline_variants":["Volunteers reel in 6,739 jellyfish galaxies","Crowd science nets 6,739 ram-pressure stripped galaxies","Citizen scientists spot 6,739 jellyfish galaxy candidates","From 50k images: 6,739 jellyfish galaxies found by volunteers","Jellyfish galaxies: crowd science catches 6,739 candidates"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The thresholds calibrated on expert classifications of one cluster (Abell 1644) are applied unchanged to all 79 clusters, assuming that volunteer voting behaviour and the visibility of stripping signatures do not change with cluster redshift, mass, or image quality.","fun_headline_variants_meta":{"raw":{"variants":["Volunteers reel in 6,739 jellyfish galaxies","Crowd science nets 6,739 ram-pressure stripped galaxies","Citizen scientists spot 6,739 jellyfish galaxy candidates","From 50k images: 6,739 jellyfish galaxies found by volunteers","Jellyfish galaxies: crowd science catches 6,739 candidates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000634,"raw_usage":{"total_tokens":2830,"prompt_tokens":877,"completion_tokens":1953,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":1861}},"tokens_in":621,"tokens_out":1953,"duration_ms":15233,"temperature":1.0,"reasoning_tokens":1861,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T00:30:52.662247+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Classify a second cluster with full expert labels (e.g., Abell 1367) and measure the purity and completeness of the citizen criteria; if they fall far from P=0.63 and C=0.57, the transferability assumption fails. Alternatively, take a random subset of the 6,739 candidates and search for extraplanar ionized gas with integral-field spectroscopy; a low confirmation rate would weaken the physical interpretation.","supporting_citations":[],"review_version":1}