{"id":"12a0be1d-8030-4639-adf8-a6d78cd8c532","arxiv_id":"2501.09547","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new mock-observation method and a comparison of WALLABY and SIMBA H I asymmetries show hints of more highly asymmetric galaxies in the real sky, but not at a statistically significant level.","lead":"This paper compares how lopsided neutral hydrogen gas looks in real galaxies from the WALLABY survey and in simulated galaxies from SIMBA, using a new way to turn simulated gas into realistic telescope images. It finds a possible excess of very asymmetric galaxies in the real data, but the difference is not statistically significant.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Size/resolution mismatch in the §4.2 matching can bias SIMBA mock A3D downward and inflate the PQMass p-value, so the null result is not yet robust.","rationale":"I read the paper as a careful pilot comparison. The Scanline Tracing method is a genuine technical improvement for low-particle-number mocks, and Figures 1-5 support the claim that particle-based cubes contain spurious spectral shot noise that inflates asymmetries. The WALLABY asymmetry measurements and the qualitative identification of A3D > 0.5 systems with interactions, bridges, and tails are plausible and appropriately hedged. The central quantitative claim is the null PQMass result, p ~ 5.4%, and the most load-bearing condition for that claim is that the mock sample matches WALLABY in all properties that affect A3D. Section 4.2 matches M_HI, distance, inclination, and noise, but not physical size or resolution elements. Since A3D is resolution-dependent, a systematic size difference at fixed M_HI would suppress mock asymmetries and push the p-value upward. The authors themselves raise the resolution confound as a possible explanation for one discrepancy region, so this is not an external speculation; it is an admitted, untested limitation that directly affects the headline null result. The proposed test, adding size matching and re-running the PQMass comparison, would settle whether the null result survives. This concern does not invalidate the paper. It reinforces the conditional verdict: the comparison is promising and methodologically novel, but the null result should not be taken at face value until the size/resolution confound is controlled. Lack of public code for the new mock generator is a secondary reproducibility issue, not a load-bearing correctness issue for the scientific claim.","tokens_in":22205,"tokens_out":4739,"duration_ms":53679,"concrete_test":"Recompute the §4.2 mock sample with an additional matching criterion on H I size: require the SIMBA galaxy's H I radius (or, failing that, its half-mass radius or the radius from the Wang et al. 2016 size-mass relation at the matched M_HI) to fall within ±0.2 dex of the WALLABY detection's H I radius inferred from its kinematic model or size, then re-run the 20-realization PQMass comparison. If the p-value shifts from ~5% to below 1% or above 10%, or if the high-A3D, low-A1D bin excess changes sign, the reported null result is not robust to the size/resolution confound.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing assumption in §4.2 is that matching M_HI within 0.2 dex, together with distance, inclination, and noise, is sufficient to make the WALLABY and SIMBA samples comparable for A3D. A3D is resolution-dependent: at fixed beam and distance, a smaller H I disk has fewer resolution elements and a suppressed 3D asymmetry. If SIMBA disks are smaller at fixed M_HI, the mocks will be artificially smoother in the spatial dimension, depressing A3D and inflating the PQMass p-value. This biases the central null result toward agreement, not toward disagreement. The authors explicitly acknowledge that under-resolution of the SIMBA mocks may explain the moderate-A1D, low-A3D discrepancy region, but the same mechanism also suppresses the high-A3D, low-A1D tail where WALLABY shows an excess, so it can reduce the very difference the paper is trying to detect. No size or resolution-element matching is applied, and no figure or table demonstrates that the matched SIMBA and WALLABY samples have comparable angular sizes at fixed mass. Because the headline claim is a null result, the comparison needs to show that it is not biased toward null; as it stands, a false negative from the size confound cannot be excluded.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the second ASymba study, comparing H I asymmetries in WALLABY Pilot Survey detections with mock observations constructed from the SIMBA 50 Mpc cosmological simulation. The authors introduce a Scanline Tracing method for building mock H I datacubes that reduces spectral shot noise compared to particle-based MARTINI mocks. Using the 3DACS code, they measure 1D and 3D asymmetries for 116 WALLABY detections with log10(M_HI/M_sun) >= 9.2 and for 20 WALLABY-like mock realizations matched in H I mass, distance, inclination, and noise RMS. They find that high-A3D WALLABY detections are predominantly interacting systems or systems with strong tidal features. A PQMass test yields p ~ 5.4 +/- 0.2%, which the authors interpret as no statistically significant difference between the WALLABY and SIMBA asymmetry distributions, while noting hints of excess at high A3D and a possible under-resolution of SIMBA mocks in the moderate-A1D, low-A3D region.","tokens_in":22422,"tokens_out":7152,"duration_ms":71828,"significance":"If the null result is robust, the paper provides one of the first quantitative comparisons of 3D H I asymmetries between an untargeted survey and a cosmological simulation, and it offers a mock-generation method that is valuable for low-particle-number regimes. The authors give explicit sample definitions, validate the Scanline Tracing method on noiseless cubes, and apply a formal two-sample statistical test (PQMass). The main limitation is that the central claim is a null result whose robustness depends on the matched mocks being comparable in angular resolution; this is not yet demonstrated. The work is a useful methodological contribution and a step toward larger survey-simulation comparisons, but the current evidence for the null result is not fully conclusive.","major_comments":[{"comment":"The mock matching controls for H I mass (within 0.2 dex), distance, inclination, and noise RMS, but not for angular size or number of resolution elements. Because A3D is resolution-dependent, a systematic offset in H I disk size at fixed mass between SIMBA and WALLABY would suppress A3D in the mocks, biasing the PQMass p-value toward agreement. The authors acknowledge this possibility for the moderate-A1D, low-A3D region, but the same mechanism can also suppress the high-A3D, low-A1D tail where WALLABY shows an excess. In addition, the WALLABY sample is defined by the kinematic-modelling criteria (ell_maj > 2 beams or log10(S/N) > 1.25), yet these criteria are not applied to the SIMBA mocks. Please add a quantitative comparison of angular sizes (e.g., the ell_maj distribution) for the matched WALLABY and SIMBA samples, and/or perform a sensitivity test on a size-matched subsample, so that the null result is not an artifact of systematically smaller simulated disks.","section":"Section 4.2 and Figure 11"},{"comment":"The 20 mock realizations repeatedly reuse the same 789 SIMBA galaxies, so the mock points are not independent draws from the simulation population. The quoted p-value uncertainty (+/- 0.2%) reflects only the variation over different Voronoi binnings, not the correlation induced by repeated galaxies viewed from different orientations. Because the same galaxy can appear many times with correlated intrinsic morphology, the effective sample size for the mock distribution may be substantially lower than 20 times the number of WALLABY detections. Please report the number of unique SIMBA galaxies that contribute to the mock sample and assess the sensitivity of the p-value to this correlation, for example by resampling over unique galaxies rather than over individual mock cubes.","section":"Section 4.2, PQMass test"}],"minor_comments":[{"comment":"The abstract states \"p-value = 0.05\" while the text reports \"p ~ 5.4 +/- 0.2%\"; please use a consistent value and clearly state the uncertainty in the abstract.","section":"Abstract and Section 4.2"},{"comment":"The sentence \"We note that in Figure 6, the high and low mass mock cubes split along the 1:1 line...\" refers to A3D versus A1D+A2D, which is shown in Figure 7, not Figure 6; please correct the cross-reference.","section":"Section 3.2, near Figure 7"},{"comment":"The table uses \"A3D >= 0\" in some rows and \"A3D > 0\" in others; unify the notation to avoid ambiguity.","section":"Table 1"},{"comment":"Individual error bars on A1D and A3D would aid the interpretation of the distribution comparison, or the authors should at least state the typical uncertainty in the measured asymmetries.","section":"Figures 9 and 11"}],"recommendation":"major_revision","confidential_remarks":"The borderline p-value (5.4%) and the possible size confound make the null result fragile; I would not recommend acceptance before the authors demonstrate that the matched mock sample is comparable in angular resolution and account for the repeated use of SIMBA galaxies. The paper fits the journal scope well and the methodological work is promising."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Net take: this is a careful, honest pilot study whose real contribution is methodological. The Scanline Tracing mock generator fixes a genuine spectral shot-noise artifact in particle-based mocks at low particle number, and the direct WALLABY-SIMBA comparison is new. The headline astrophysical result is a null (PQMass p ≈ 5.4%), and that null is plausible but not yet robust.\n\nWhat works: the mock construction is thoughtfully matched (M_HI within 0.2 dex, same distance/inclination/noise, same SoFiA-2 mask and center), the asymmetry definitions and background correction come from 3DACS and are applied consistently to both sides, and the authors are disciplined in interpreting the p-value as a hint rather than a detection. The empirical statement that A3D > 0.5 detections in WALLABY are dominated by interactions and tidal features is useful and well supported by Figure 10. The scanline method itself is clearly derived, and the demonstration that MARTINI's particle-centric spectral sampling injects spurious asymmetry at ~1000 particles is convincing.\n\nSoft spots, in proportion. The stress-test concern lands. A3D is resolution-dependent. Matching only on M_HI does not guarantee comparable angular size at fixed beam and distance; if SIMBA disks are smaller at fixed M_HI, the mock A3D will be systematically suppressed. That suppresses the very high-A3D, low-A1D tail where WALLABY shows an excess, so the comparison may be biased toward accepting the null. The authors acknowledge under-resolution in Section 4.2 as a possible explanation for the moderate-A1D/low-A3D discrepancy, but they do not show that the matched samples have comparable angular sizes or resolution-element counts, and the same mechanism cuts both ways. For a null result, the test needs to demonstrate it could have detected a difference. This is fixable: a size-matched or resolution-element-matched comparison, or at minimum a figure showing the angular size distributions.\n\nMinor: no public code for the scanline generator (MARTINI is public, this isn't), individual asymmetry error bars aren't shown, and the p-value is borderline at 5%, which is exactly where a small confound can flip the conclusion.\n\nNo load-bearing math error that I found, and the self-citations are provenance, not circularity.\n\nWho this is for: people building HI mocks and survey-simulation morphometric comparisons. It deserves a serious referee. I'd send it to review, and ask for the size/resolution robustness check and code release before final acceptance.","headline":"A careful pilot null result with a genuinely useful new mock method, but the size/resolution matching confound means the null is not yet robust.","tokens_in":23089,"tokens_out":2844,"would_cite":true,"duration_ms":27308,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"After matching mass and noise, WALLABY and SIMBA H I asymmetries agree, and the most asymmetric WALLABY systems are interactions or tidal tails.","keywords":["galaxy asymmetries","H I morphology","mock datacubes","cosmological simulations","WALLABY survey","SIMBA simulation","3D asymmetry","spectral asymmetry"],"falsifier":"Compute the distribution of on-sky sizes ($\\mathrm{ell}_{\\rm maj}$ or $R_{\\rm HI}$) for the matched SIMBA mocks and compare it to the WALLABY detections: if SIMBA galaxies are systematically smaller at fixed H I mass, then the matched comparison is biased and the $p = 5.4\\%$ cannot be read as evidence that the simulation reproduces observed asymmetry.","tokens_in":21991,"feed_emoji":"📡","tokens_out":8909,"duration_ms":81204,"temperature":0.7,"pith_summary":"Observed and simulated galaxies should show the same kinds of lopsided neutral gas if a cosmological simulation captures the processes shaping H I. This paper makes that comparison quantitative with matched samples: 116 spatially resolved WALLABY Pilot Survey detections and 20 realizations of WALLABY-like mocks drawn from the SIMBA 50 Mpc simulation. After matching H I mass, distance, inclination and noise, the $A_{\\rm 1D}$--$A_{\\rm 3D}$ asymmetry distributions are not statistically distinguishable, with a PQMass probability of $p = 5.4 \\pm 0.2\\%$. The paper also reports that detections with $A_{\\rm 3D} > 0.5$ are almost all interacting systems or objects with strong bridges and tidal tails, making this 3D asymmetry a usable interaction finder in survey data.","feed_headline":"Matched, real and simulated gas asymmetries agree","feed_subtitle":"116 WALLABY detections and SIMBA mocks share the same asymmetry plane; extreme outliers trace mergers and tidal tails.","key_machinery":"Two objects carry the argument. The first is the 3D asymmetry $A_{\\rm 3D}$, the ratio of squared odd to squared even parts of the H I datacube about a center, with a background correction $B = 2N\\sigma^2$ that removes the Gaussian-noise contribution; the 1D asymmetry $A_{\\rm 1D}$ is the same ratio on the flux summed over the spectral axis, equivalent to the channel-by-channel asymmetry. The second is the Scanline Tracing mock generator: it interpolates the smoothed density, line-of-sight velocity and velocity-dispersion fields of the simulation on a grid, evaluates a locally Gaussian spectrum at each scanline step, and then convolves with the 30 arcsecond beam. By enforcing locally Gaussian spectra, it avoids the velocity-space shot noise of the particle-based approach and changes $A_{\\rm 3D}$ by up to roughly 0.3 for low-mass SIMBA galaxies.","core_discovery":"The paper claims that, at current sample sizes, the WALLABY Pilot Survey and the SIMBA 50 Mpc simulation produce the same population of H I asymmetries once observing parameters are controlled. To reach this claim it introduces Scanline Tracing, a mock-observation method that samples simulated gas fields along lines of sight rather than adding a per-particle Gaussian spectrum; this removes shot noise that artificially raises asymmetry in low-particle-number mocks. Each WALLABY detection is matched to a SIMBA galaxy within 0.2 dex in H I mass, placed at the same distance and inclination, given Gaussian noise at the observed RMS, and then passed through the same source finder before $A_{\\rm 1D}$ and $A_{\\rm 3D}$ are measured. The resulting PQMass test gives $p = 5.4 \\pm 0.2\\%$ that the two point clouds in the $A_{\\rm 1D}$--$A_{\\rm 3D}$ plane come from the same distribution, so the paper concludes the distributions are consistent while noting an excess of high-$A_{\\rm 3D}$ detections in WALLABY and an excess of moderate-$A_{\\rm 1D}$, low-$A_{\\rm 3D}$ mocks in SIMBA.","pith_inferences":["If the high-$A_{\\rm 3D}$ excess survives the full survey, the likely culprits are the hot SIMBA IGM suppressing cold bridges and tails, or the pilot fields' group and cluster environments preferentially selecting interactions; the paper names both candidates, and an environment-stratified comparison would separate them.","The SIMBA overpopulation at moderate $A_{\\rm 1D}$ and low $A_{\\rm 3D}$ can be tested directly by measuring the H I size--mass relation in the matched mocks; if SIMBA disks are smaller than WALLABY's at fixed $M_{\\rm HI}$, the low $A_{\\rm 3D}$ is a resolution artifact and the agreement is partly built into the matching scheme.","The same matched-mock protocol could be applied to the other SIMBA 50 Mpc feedback variants to see which subgrid physics moves the $A_{\\rm 1D}$--$A_{\\rm 3D}$ plane, turning the p-value into a physics discriminator rather than a pass/fail test."],"forward_implications":["$A_{\\rm 3D} > 0.5$ can be used as a selection cut in untargeted H I surveys to find interacting galaxies and tidal features without multi-wavelength data.","The Scanline Tracing recipe can produce WALLABY-like mocks from any SPH or MFM cosmological simulation, making asymmetry comparisons a standard morphometric test.","Full WALLABY data will decide whether the apparent excess of extreme 3D asymmetries over SIMBA is real; if it is, it may indicate missing merger or tidal physics or environmental selection in the pilot fields.","Kinematic modelling success is not strongly anti-correlated with asymmetry, so asymmetry adds independent information beyond rotation-curve fitting."],"supporting_citations":[{"why":"Provides the standard particle-based mock-cube code that the paper shows injects spectral shot noise, motivating the new Scanline Tracing method.","marker":"Oman 2019; Oman et al. 2019; Oman 2024"},{"why":"Provides the SIMBA cosmological simulation whose 50 Mpc box supplies the mock galaxies.","marker":"Davé et al. 2019"},{"why":"Releases the WALLABY Pilot Survey detections and defines the source-finding and data properties used here.","marker":"Westmeier et al. 2022"},{"why":"Adds the WALLABY Pilot Survey Phase 2 data and detections that enlarge the pilot sample.","marker":"Murugeshan et al. 2024"},{"why":"Defines the 3D asymmetry $A_{\\rm 3D}$ and the Gaussian background correction used on both data sides.","marker":"Deg et al. 2023"},{"why":"The earlier ASymba paper that sets the SIMBA galaxy selection criteria and connects spectral asymmetries to mergers and H I mass.","marker":"Glowacki et al. 2022"},{"why":"Defines the kinematic-modelling sample of WALLABY detections used as the comparison sample.","marker":"Deg et al. 2022"},{"why":"Supplies the PQMass test used to compare the WALLABY and SIMBA $A_{\\rm 1D}$--$A_{\\rm 3D}$ distributions.","marker":"Lemos et al. 2024"}],"fun_headline_variants":["HI asymmetries agree between WALLABY and SIMBA","WALLABY and SIMBA show matching gas asymmetry","Simulated and observed HI asymmetries align","Pilot survey and simulation share asymmetry patterns"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a SIMBA galaxy selected only by H I mass (within 0.2 dex) and then projected at the WALLABY distance and inclination ends up with the same spatial resolution and disk sizes as real WALLABY detections; the paper itself flags in Section 4.2 that SIMBA disks may be smaller at fixed $M_{\\rm HI}$, which would lower $A_{\\rm 3D}$ in the mocks and make the agreement look better than it is.","fun_headline_variants_meta":{"raw":{"variants":["HI asymmetries agree between WALLABY and SIMBA","WALLABY and SIMBA show matching gas asymmetry","Simulated and observed HI asymmetries align","Pilot survey and simulation share asymmetry patterns"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000176,"raw_usage":{"total_tokens":1357,"prompt_tokens":1079,"completion_tokens":278,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":695,"completion_tokens_details":{"reasoning_tokens":216}},"tokens_in":695,"tokens_out":278,"duration_ms":3407,"temperature":1.0,"reasoning_tokens":216,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:54:45.516306+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the distribution of on-sky sizes ($\\mathrm{ell}_{\\rm maj}$ or $R_{\\rm HI}$) for the matched SIMBA mocks and compare it to the WALLABY detections: if SIMBA galaxies are systematically smaller at fixed H I mass, then the matched comparison is biased and the $p = 5.4\\%$ cannot be read as evidence that the simulation reproduces observed asymmetry.","supporting_citations":[{"cited_title":"WALLABY Pilot Survey: Public data release of ~1800 HI sources and high-resolution cut-outs from Pilot Survey Phase 2","cited_arxiv_id":"2409.13130","evidence_quote":"Adds the WALLABY Pilot Survey Phase 2 data and detections that enlarge the pilot sample."}],"review_version":1}