Pith. sign in

REVIEW 4 major objections 5 minor 9 references

Improving Gamma-ray Source Search with Image Processing

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read An image-processing pipeline can seed HAWC gamma-ray sources about 60 times faster than the likelihood search, while recovering nearly all point and compact sources.

desk verdict The pipeline idea is sensible, but the paper's own numbers contradict its central claims, so it is not ready for publication as-is. read the letter →

arxiv 2507.10307 v1 pith:4QH4NYEP submitted 2025-07-14 astro-ph.HE astro-ph.IM

classification astro-ph.HEastro-ph.IM
keywords HAWCobservatoryvery-high-energygammaraysgamma-raysourcesearchimageprocessingDifferenceofGaussiansblobdetectionseedinglikelihoodfitting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is trying to establish that the slowest part of HAWC's blind gamma-ray source search can be replaced with an image-processing step. Instead of fitting a multi-source statistical model to a region directly, which takes days, the pipeline treats the significance map as an image, sharpens it with a Difference-of-Gaussians filter, and pulls out candidate source positions with a blob finder. The candidates are then handed to the usual likelihood fitter for exact localization. On 60 simulated fields the seeds recover 65 of 66 point sources and 24 of 25 compact extended sources, and the image step ran in under 26 CPU-hours total instead of the two months the likelihood search needed. If this holds, HAWC and similar instruments could scan wide regions for new or variable gamma-ray sources much faster and at lower computational cost.

What carries the argument

The machinery is a Difference-of-Gaussians (DoG) spatial band-pass filter combined with a multi-scale blob finder. A Gaussian-blurred version of the normalized significance image is subtracted from the original, so sharp high-frequency peaks survive while broad diffuse emission is removed; the blob finder then scans the residual across increasing blur scales and keeps pixels that are local extrema in several consecutive scales. Fixed parameters do the noise rejection: a 0.20° smearing radius matched to the detector point-spread function, a 3σ threshold on the residual-intensity histogram, a 5σ threshold on the original significance map, and a cut at 10 pixels from the image edge. Detected blobs with scale below 0.15° are labeled point-like and larger ones extended. This machinery reduces a many-parameter likelihood fit to a short candidate list that the fitter only has to refine.

What would settle it

Run the pipeline with those fixed parameters on a held-out set of simulations, or on a real HAWC significance map whose sources are already known, and count how many injected or catalogued point and compact sources appear in the seed list; if recovery falls well below 65 of 66 point sources and 24 of 25 compact extended sources, or if the wall-clock advantage over the likelihood search disappears in a head-to-head test, the central claim would not transfer.

Watch

Extended reading notes

Core claim

The central claim is that a standard image-processing stack recovers most gamma-ray sources from HAWC significance maps before any likelihood fitting is done, and does so far faster than the iterative multi-source search now in use. The stack clips the map at -5σ, normalizes it, applies a Difference-of-Gaussians spatial band-pass filter, and runs a multi-scale blob finder that tags local extrema as source seeds. On the 60-simulation benchmark, the pipeline recovered 65 of 66 point sources and 24 of 25 extended sources smaller than 0.5°, while recovering only 1 of 33 large diffuse sources above 0.8°, which the filter suppresses as low-frequency structure. The seeds are classified by blob scale as point-like or extended and passed to the usual HAWC likelihood fitter for final localization. The image-processing stage took under 26 CPU-hours total against more than two months for the conventional blind search, a 60× speedup (the abstract quotes up to 300×), corresponding to the reported 98.33% performance gain.

Load-bearing premise

The measured speed and recovery rates assume that the fixed 0.20° smearing radius, the 5σ detection threshold, and the other cut values will behave the same way outside the 60 simulated maps on which they are demonstrated.

Editorial extensions

If this is right

  • Source seeding for a HAWC region can drop from days to hours, leaving likelihood computation for refining a short candidate list.
  • Fainter point sources that the iterative fit misses should appear as seeds and can be added to the model manually or automatically.
  • Because the method works on significance images rather than detector-specific likelihoods, it can be ported to Fermi-LAT maps and to future LHAASO and SWGO data.
  • Large diffuse sources will need a separate search, since the Difference-of-Gaussians filter removes them along with the background.
  • The saved computation makes wide-area or repeated scans of the very-high-energy sky practical.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A decisive next test would freeze the 0.20° smearing radius and 5σ threshold on one set of simulations and measure recovery on independent simulations or real HAWC maps, since the paper reports performance on the same simulations used to demonstrate those parameters.
  • The blob scale returned by the finder could feed extended-source templates into the likelihood fit, giving the fitter a size prior that might recover some of the large diffuse sources the current filter suppresses.
  • At the reported speed, a natural deployment is a real-time or alert-driven scan that seeds candidate sources in freshly made significance maps without waiting for a full likelihood model.
  • Tuning the smearing radius to the local point-spread function of each detector would likely be needed when moving the pipeline to other instruments.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript describes an image-processing pipeline for seeding gamma-ray source candidates in HAWC significance maps. The pipeline uses a Difference-of-Gaussians (DoG) filter, a blob detector, and intensity thresholding to produce a list of source seeds that are then passed to a likelihood-based fitting tool. Performance is reported on 60 simulated maps containing 123 sources, with claims of a large speedup over the conventional blind search. The paper concludes that the algorithm achieves >98% accuracy for point sources and >95% accuracy for extended sources, with a computational speedup of about 98%. The central claim is that this pipeline can accelerate and improve HAWC source searches.

Significance. If the claims were supported, the pipeline would be a useful complement to the computationally expensive likelihood searches used by HAWC and similar observatories. The paper addresses a real operational problem and proposes a concrete, potentially transferable method. However, the reported results are undermined by internal inconsistencies between the abstract, the performance section, and the conclusion, and by the absence of any validation on independent or real data. As written, the paper does not establish its central advertised results for extended sources or the '300 times faster' speedup, so its practical significance cannot be assessed reliably.

major comments (4)
  1. [Section 3 and Section 4] The recovery numbers in Section 3 do not support the accuracy claims in Section 4. Section 3 reports recovery of 65/66 point sources, 24/25 small extended sources, and 1/33 large extended (>0.80 deg) sources. For extended sources, this is 25 of 58, or about 43%, not '>95% accuracy for extended sources' as claimed in Section 4. The failure on large extended sources is not incidental: Section 2.3 states that 'large extended sources are removed' by the DoG filter. The central claim of accurate source seeding is therefore not supported for the source class that is most challenging for the existing blind search, and the conclusion is arithmetically incompatible with the paper's own data.
  2. [Abstract and Section 3] The abstract claims the pipeline 'seeds sources accurately up to 300 times faster than the current HAWC source search pipeline,' but Section 3 reports a '60× speedup' and says nothing about 300×. The number 300 appears nowhere in the text or tables. Moreover, the timing comparison is not controlled: 'less than 26 hrs' for the image pipeline over 60 simulation sets is compared with 'over 2 months' for the conventional search without specifying the computational endpoint, hardware, region of interest, or whether the same simulations and analysis settings were used. The speedup factor and the '98.33% performance gain' should be derived from a clearly defined, matched comparison.
  3. [Section 2.3 and Section 3] The evaluation is in-sample and the key parameters are hand-picked. Section 2.3 states 'A smearing radius of 0.20 is used in the simulation plots shown in the proceeding,' and Section 2.1 sets a 5-sigma detection threshold, but no sensitivity study, cross-validation, or hold-out dataset is presented. Recovery rates and speedups are measured on the same 60 simulations used to motivate these choices. Without testing on independent simulations or on real HAWC maps, the reported accuracy and speedup cannot be interpreted as general performance measures; they are partly a product of the chosen thresholds.
  4. [Section 3 (Figure 3)] The paragraph following Figure 3 states that 'for the simulation dataset used, the image processing pipeline is fully capable of recovering the simulated sources.' This directly contradicts the quantitative recovery counts in the preceding paragraph, where only 1 of 33 large extended sources is recovered. The figure appears to show only a subset of the simulated sources, and the text should be corrected to state the overall recovery fractions, including the poor performance on large extended sources, rather than a qualitative claim of full capability.
minor comments (5)
  1. [Throughout] There are several typos: 'likeihood' in the Introduction, 'greather' in Section 2.4, 'the the pixel' in Section 2.4, and 'proceeding' should be 'proceedings' in Section 2.3. These should be corrected.
  2. [Section 2.4] The classification thresholds are given as '0.150' and '0.150' without units; presumably these are degrees, and the values should be written as '0.15°' for clarity. The blob scale definition is also not defined precisely.
  3. [Section 3] The term 'Drips' is introduced without definition or explanation. It is unclear whether this is a name for the seed list, the pipeline output, or something else; please define it.
  4. [Section 2.1] The initial peak search uses a '≥5σ intensity' threshold, but the relationship between this threshold and the later 5-sigma blob significance threshold in Section 2.4 is not clarified. Having two 5-sigma thresholds with different meanings is confusing.
  5. [Section 3] The false-positive rate of '10%' is stated without a denominator or definition. It would help to specify how false positives are counted, e.g., detected seeds not matching any simulated source within a given angular separation.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found; the DoG/blob-finder pipeline is an externally grounded image-processing method, and the paper's performance numbers are empirical benchmarks rather than reductions of the method to its inputs.

full rationale

No circularity found. The pipeline is a standard image-processing chain (contrast normalization, Difference-of-Gaussians band-pass filtering, and DoG blob detection from scikit-image, Marr & Hildreth, and Lowe) applied to HAWC significance maps. The parameters (0.20 deg smearing radius and 5 sigma threshold) are fixed operational choices, not fitted to the 60 simulations and then re-predicted; the paper reports recovery counts on those simulations as a benchmark. This is an in-sample evaluation and a generalization concern, but it is not a circular derivation because the recovery counts are measured outputs, not consequences of the parameter values by construction. The large-extended-source failure (1/33) is the expected consequence of the DoG high-pass filter, explicitly stated in Section 2.3, not a hidden circularity. The 300x abstract speedup vs 60x in Section 3 and the conclusion's '>95% accuracy for extended sources' versus the 25/58 recovery count are internal inconsistencies or correctness issues, not circularity. Citations to HAWC/3ML are background methodology, not the load-bearing justification for the new pipeline. Therefore score 0.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The central performance claim rests on a small number of hand-chosen image-processing parameters and an unvalidated assumption that simulations match real HAWC data. The free parameters are explicit in the text but none are justified beyond being used for the presented simulations.

free parameters (4)
  • smearing radius = 0.20 degrees
    Used for the initial Difference of Gaussians feature enhancement. The paper says 'A smearing radius of 0.20 is used in the simulation plots shown in the proceeding.' It is chosen by hand and not justified or cross-validated.
  • detection threshold = 5 sigma
    Peaks below 5 sigma in the clipped significance map are discarded, and blobs with significance below 5 sigma are removed. This threshold directly controls which sources are seeded and is not varied or tested.
  • boundary cut = 10 pixels
    Blobs detected within 10 pixels of the image edges are removed to reduce fake detections. The value is stated without justification or sensitivity analysis.
  • point/extended classification threshold = 0.15 degrees
    Blobs with scale less than 0.15 degrees are labeled point-like, larger blobs are labeled extended. The choice affects how seeds are passed to the likelihood fitter and is not justified.
assumptions (3)
  • standard math Gaussian convolution and Difference of Gaussians filtering provide valid feature enhancement for source detection.
    The pipeline relies on standard image processing methods cited as [6] and [7]. These are established in computer vision and are not questioned here.
  • domain assumption The simulated significance maps are representative of real HAWC data.
    There is no description of the simulation setup, background model, or PSF treatment. The transfer of performance numbers to real data is assumed rather than demonstrated.
  • ad hoc to paper The 0.20 degree smearing radius is appropriate for the PSF at the analyzed declinations.
    The text says the smearing radius equals the PSF averaged over bins, but then uses a fixed 0.20 degrees without showing how this matches different regions of the sky or different energy bins.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving Gamma-ray Source Search with Image Processing." pith.science (2026). https://pith.science/paper/4QH4NYEP

@misc{pith2026250710307,
  author       = {Pith},
  title        = {Pith review of: Improving Gamma-ray Source Search with Image Processing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4QH4NYEP}},
  note         = {Machine review of arXiv:2507.10307}
}
read the original abstract

The analysis of HAWC data is done using a likeihood-based systematic multi-source search procedure utilizing the threeML software package and the HAL Plugin. This approach was inspired by the extended source search described in the Fermi-LAT Extended Source Search Catalog. The pipeline to search for point sources and extended sources within the region of interest (ROI) is described in the recent HAWC papers. This procedure is computationally intensive and often requires multiple days to produce a final model for a region. Often this approach misses fainter sources, which need to be added manually later. This blind search could be complemented by providing a method to seed source locations, which can be assessed and evaluated by likelihood analysis, thereby significantly reducing the computational time and resources spent on finding a model.

Figures

Figures reproduced from arXiv: 2507.10307 by the authors.

Figure 1
Figure 1. HAWC Significance Sky Map showing the location and morphology of three different simulated sources, marked with their labels and extensions (in galactic coordinates). 2 [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Difference of Gaussian (DoG) Algorithm 2.4 Blob Detection and Intensity Thresholding We employ a blob detection algorithm, which is based on the DoG approach to look for blobs in the images [7]. The blob detection algorithm works by building an array of images, in which the DoG residual image is blurred with increasing smearing radii and the difference between two consecutive images are stored in a 3D array. The alg… view at source ↗
Figure 3
Figure 3. The source seeds detected from the image processing pipeline, along with the simulated source locations. The output seeds, referred to as Drips, generated by the image processing pipeline on the simulation dataset used in this study are shown in Fig.(3). We see that for the simulation dataset 5 [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

9 extracted references · 2 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION blank.sep after.quote 'output.state := FUNCTION fin.entry output.state after.quoted.block = 'skip 'add.period if write newline FUNCTION new.block output.state before.all = 'skip output.state after.quote = after.quoted.block 'output.state := after.block 'output.state := if if FUNCTION new.sentence out...

  2. [2]

    Vianello , R.J

    G. Vianello , R.J. Lauer , P. Younk , L. Tibaldo , J.M. Burgess , H. Ayala et al., The Multi-Mission Maximum Likelihood framework (3ML) , arXiv e-prints (2015) arXiv:1507.08343 https://doi.org/10.48550/arXiv.1507.08343 [ https://arxiv.org/abs/1507.08343 1507.08343 ]

  3. [3]

    HAWC collaboration, Characterizing gamma-ray sources with HAL (HAWC Accelerated likelihood) and 3ML , https://doi.org/10.22323/1.395.0828 PoS ICRC2021 (2021) 828

  4. [4]

    Ackermann , M

    M. Ackermann , M. Ajello , L. Baldini , J. Ballet , G. Barbiellini , D. Bastieri et al., Search for Extended Sources in the Galactic Plane Using Six Years of Fermi-Large Area Telescope Pass 8 Data above 10 GeV , https://doi.org/10.3847/1538-4357/aa775a Astrophysical Journal 843 (2017) 139 [ https://arxiv.org/abs/1702.00476 1702.00476 ]

  5. [5]

    TeV Analysis of a Source Rich Region with HAWC Observatory: Is HESS J1809-193 a Potential Hadronic PeVatron?

    A. Albert , R. Alfaro , C. Alvarez , J.C. Arteaga-Vel \'a zquez , D. Avila Rojas , R. Babu et al., TeV Analysis of a Source-rich Region with the HAWC Observatory: Is HESS J1809-193 a Potential Hadronic PeVatron? , https://doi.org/10.3847/1538-4357/ad59a6 AstroPhysical Journal 972 (2024) 21 [ https://arxiv.org/abs/2407.08849 2407.08849 ]

  6. [6]

    van der Walt, J.L

    S. van der Walt, J.L. S ch\"onberger, J. Nunez-Iglesias , F. B oulogne, J.D. W arner, N. Y ager et al., scikit-image: image processing in P ython , https://doi.org/10.7717/peerj.453 PeerJ 2 (2014) e453

  7. [7]

    Marr and E

    D. Marr and E. Hildreth, Theory of edge detection, https://doi.org/10.1098/rspb.1980.0020 Proceedings of the Royal Society of London. Series B. Biological Sciences 207 (1980) 187

  8. [8]

    Lowe, Distinctive image features from scale-invariant keypoints, International journal of computer vision 60 (2004) 91

    D.G. Lowe, Distinctive image features from scale-invariant keypoints, International journal of computer vision 60 (2004) 91

Show all 9 references
  1. [9]

    write newline

    " write newline "" before.all 'output.state := FUNCTION blank.sep after.quote 'output.state := FUNCTION fin.entry output.state after.quoted.block = 'skip 'add.period if write newline FUNCTION new.block output.state before.all = 'skip output.state after.quote = after.quoted.blo...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.