{"id":"3c3303e5-244b-44be-a50f-29a5ad969395","arxiv_id":"2412.07858","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A new X-ray source-detection pipeline built on the optimal Poisson matched filter detects roughly 30 percent more faint sources in Chandra images than the Chandra Source Catalog.","lead":"X-Sifter is a new software pipeline that finds faint X-ray sources using a matched filter designed for Poisson noise, the noise regime of photon-counting detectors. Tested on Chandra X-ray data, it detects about 30 percent more sources near the detection limit than the standard catalog and reports higher signal-to-noise for the same sources.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 30% sensitivity gain is not yet demonstrated because the comparison to CSC2.1 does not match false-alarm rates; the extra detections may be a threshold artifact.","rationale":"The reader's weakest assumption focuses on the per-sector constant background estimate, which is a real concern about calibration in structured backgrounds. However, I judge the more load-bearing concern to be the uncontrolled comparison against CSC2.1: the claimed 30% increase and 1.8x survey speed depend directly on comparing X-Sifter detections to a catalog produced with different detection thresholds and a different significance scale. Even with a perfectly estimated background, an unequal threshold would produce exactly such an apparent excess; conversely, if false-alarm rates were matched and the excess persisted, the background issue would be secondary. The reader's rationale does list 'match false-alarm rates' as a required check, so there is partial agreement, but the weakest_assumption field identifies background rather than threshold mismatch. My conclusion does not change the CONDITIONAL verdict: the paper presents a plausible implementation of an optimal filter and some internal validation, but the headline sensitivity gain needs a controlled comparison before it can be accepted as established. I therefore keep the verdict unchanged and ask for the matched-threshold test as the key condition.","tokens_in":12637,"tokens_out":5208,"duration_ms":54453,"concrete_test":"Run X-Sifter and CIAO wavdetect on the same set of fields with matched false-alarm probabilities: for each field, calibrate each algorithm's threshold using background-only simulations so that both expect the same number of false detections per field (e.g., 1 per 10 fields), then recompute the ratio of detected real sources and the median S/N ratio at those matched thresholds. Alternatively, perform an injection-recovery experiment with sources of known flux added to real background images, and plot completeness as a function of flux for both algorithms at equal false-positive rates. If the 30% excess and 1.3x S/N ratio reproduce at matched false-alarm rates, the survey-speed claim is supported; if not, the apparent gain is a threshold artifact rather than a sensitivity gain.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 4.2 that X-Sifter detects ~30% more real sources than CSC2.1 and reports ~1.3x higher S/N near the detection limit is based on comparing X-Sifter detections to the Chandra Source Catalog, which uses wavdetect with its own significance threshold and multi-scale selection criteria. The paper never specifies the false-alarm probability (γ) or S/N cutoff used by X-Sifter in this comparison, nor does it match the expected number of false detections per field for the two methods. If X-Sifter was run at a lower effective threshold than CSC2.1, the 55 extra sources and the higher S/N for common sources would reflect threshold mismatch rather than a genuine improvement in the detection statistic. The S/N ratio itself is also not apples-to-apples because the two pipelines convert their detection statistics to Gaussian sigma via different calibrations, and the paper admits the γ-to-S/N conversion is approximate. The false-positive check in Section 4.3 uses a different sample of five stacked fields with only 10 unmatched sources, so it does not verify the reality of the 55 extra sources in the single-observation fields. Thus the headline sensitivity gain is not quantitatively established by the current data, even if the background estimation were perfect.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents X-Sifter, a source-detection pipeline for X-ray images based on the Poisson matched filter of Ofek & Zackay (2018b). The pipeline partitions each CCD into 16x16 sectors and several energy channels, models the PSF with MARX, estimates a single background value per sector, precomputes the null-hypothesis distribution of the detection statistic to set thresholds, and filters each sector to produce a combined S/N map. The authors validate with injected-source simulations and compare detections with the Chandra Source Catalog 2.1 in ten single-observation fields. They claim ~30% more sources detected, a 1.3x higher S/N near the detection limit, and a factor 1.8 increase in survey speed. They also report a small false-positive check on five stacked fields.","tokens_in":12919,"tokens_out":6831,"duration_ms":58760,"significance":"If the sensitivity gains are real, X-Sifter would be a valuable community tool: it implements the optimal Poisson matched filter in a practical way, with publicly released code, precomputed PSF libraries, and handling of energy-dependent PSF and background variations. The injection tests demonstrate internal consistency of the implementation. However, the headline quantitative claims are not yet convincingly supported: the comparison with CSC2.1 is not performed at matched false-alarm probabilities, the S/N conversion relies on an approximate extrapolation acknowledged by the authors, and the false-positive check uses a different sample than the main comparison. The background estimator is a single median per sector, which may be biased in structured fields. These issues are addressable, but the current paper does not establish the advertised factors.","major_comments":[{"comment":"The claimed ~30% increase in detected sources and the 1.3x higher S/N are not supported by a matched-threshold comparison. The paper does not state the false-alarm probability gamma or the corresponding S/N threshold used by X-Sifter for this comparison, nor does it give the effective significance threshold used by CSC2.1/wavdetect. If X-Sifter ran at a lower threshold, the 55 extra sources near the detection limit would be a threshold artifact. The mean S/N ratio is also not apples-to-apples because X-Sifter converts its S statistic to Gaussian sigma via the gamma-S relation, which the paper itself notes in Sections 2.2.3 and 4.2 is approximate and becomes inaccurate at high sigma. The authors should present the comparison at equal expected false detections per field, report the thresholds used by both pipelines, and provide uncertainties on the ratio (55/179 carries a binomial error of roughly +/-6%).","section":"Section 4.2"},{"comment":"The false-positive check does not validate the reality of the 55 extra sources in the single-observation fields. It uses five different fields with stacked CSC detections, where only 10 X-Sifter sources lack CSC counterparts, all attributed to transients or variability. This sample is much smaller and drawn from different data. To support the claim that the 55 extra sources are real, the authors should examine a random subset of those sources directly, for example by using any other observations of the same regions, or by injecting sources into pure simulated background images and running the full pipeline with the same threshold to estimate the empirical false-positive rate.","section":"Section 4.3"},{"comment":"The background for a 16x16 sector is a single median value computed from a 4x4 grid of tiles after iterative outlier rejection. If the true background is spatially varying within a sector, or if the outlier rejection is biased by bright sources or flares, the single value of B enters both the PMF kernel (Equation 2) and the threshold (Equation 14), potentially miscalibrating the detection significance. The injection test in Section 4.1 uses the same background estimator and is therefore insensitive to this bias. The authors should test the pipeline on fields with structured background, or add a simulation with a known background gradient, to quantify the effect on completeness and false-alarm rate.","section":"Section 3.1, step 4(c)"},{"comment":"The injection test is partially circular. The injected flux Fth is derived from Equation 14 using the same background estimate and PSF that the pipeline will use, so the recovered S/N peaking at the chosen sigma is largely by construction. This demonstrates numerical correctness but not absolute sensitivity. The external comparison to CSC2.1 is meant to provide the missing calibration, but the threshold-matching issues in the first major comment prevent it from doing so. A test with sources injected at random fluxes over a range (not only Fth) would provide a more meaningful completeness curve.","section":"Section 4.1"}],"minor_comments":[{"comment":"The symbol '∼=' should be '≈' (an approximate equality sign, not a tilde-equals).","section":"Abstract"},{"comment":"The notation mixes 'Sthresh' and 'Sth'; please use one consistent symbol throughout.","section":"Section 2.2.3"},{"comment":"The phrase 'slope ranging from ∼ −1/2 and ∼ −3/2' should read 'between ∼ −1/2 and ∼ −3/2'.","section":"Section 2.2.3"},{"comment":"The background estimation description should clarify that the final representative background is the median of the tile means, not a single global fit to the sector.","section":"Section 3.1, step 4(c)"},{"comment":"'X-swifter' is a typo for 'X-Sifter'.","section":"Section 3.4"},{"comment":"'Theses two sources' should be 'These two sources'; the paper would also benefit from listing the ten OBSIDs used for the comparison, either in a table or in an appendix.","section":"Section 4.2"},{"comment":"'X=Sifter' is a typo for 'X-Sifter'.","section":"Section 4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is heavily self-referential in its theoretical foundations, citing Ofek & Zackay for both the PMF formalism and the thresholding relation, and also using the same code libraries. This is not by itself a problem, but it strengthens the need for an independent external benchmark; the current CSC2.1 comparison does not yet provide one because the thresholds are not matched. The journal fit is appropriate for PASP, and the code release is a positive feature. The main revision should focus on making the sensitivity comparison defensible at matched false-alarm rates and on quantifying the background-estimation bias."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"X-Sifter is a genuinely useful engineering contribution: it turns Ofek & Zackay's Poisson matched filter into a production pipeline that handles sector-based PSF and background variation, energy-dependent filtering, and temporal subdivision, and it comes with a code release. That alone is worth reading for anyone mining Chandra archives for faint sources or transients. The injection test is internally consistent: with a 4-sigma false-alarm setting, the recovered S/N distribution peaks at 4 sigma, which gives real confidence that the implementation matches the formalism.\n\nThe soft spots are in the external comparison. The 30% extra detections and 1.3x S/N gain over CSC2.1 are the headline, but the paper never states what false-alarm probability or S/N cutoff X-Sifter used in that comparison, nor does it match expected false detections per field between the two methods. The extra 55 sources could partly reflect a lower threshold. The paper also admits the gamma-to-S/N conversion relies on an exponential extrapolation that becomes inaccurate away from the detection limit, so the reported S/N ratio is not apples-to-apples. The false-positive check in Section 4.3 uses only five stacked fields and ten unmatched sources; it does not verify the single-observation sample where the 55 extra sources appear. The reader's worry about per-sector constant background is a real modeling premise, but it is secondary: the main issue is the uncharacterized threshold comparison.\n\nI would send this to a serious referee. The pipeline is real, the limitations are acknowledged, and the comparison can be fixed with relatively straightforward additions: state the thresholds, match false-alarm rates, add error bars, and stress-test the background estimator on crowded fields. The code release should be versioned with the exact commit used.","headline":"A worthwhile pipeline paper whose headline sensitivity gain over CSC2.1 is not yet quantitatively established because the comparison does not control for threshold differences.","tokens_in":13435,"tokens_out":2842,"would_cite":true,"duration_ms":25136,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"X-Sifter detects 30% more X-ray sources than the standard Chandra catalog.","keywords":["Poisson matched filter","X-ray source detection","transient detection","Chandra","point spread function","background estimation","survey speed"],"falsifier":"Run X-Sifter on simulated images with a known background gradient inside a single sector, such as a linear ramp across the sector with the same mean as a flat field, and compare the recovered flux and signal-to-noise of injected sources to the flat case; if the detection threshold shifts or completeness drops by more than the claimed margin, the sector-constant assumption is falsified. Alternatively, apply the code to a real field with strong structured background, such as the Galactic ridge, and compare against a deeper reference catalog: if the 30 percent gain vanishes or false positives appear, the assumption fails.","tokens_in":12461,"feed_emoji":"🔭","tokens_out":4851,"duration_ms":41729,"temperature":0.7,"pith_summary":"This paper presents X-Sifter, a software pipeline that implements the Poisson noise matched filter for detecting sources in photon-counting images. The central claim is that, when applied to single-observation Chandra fields, this filter recovers about 30 percent more real sources near the detection limit than the Chandra Source Catalog, and reports signal-to-noise ratios roughly 1.3 times higher at that limit, equivalent to a 1.8-fold increase in survey speed. The work matters because it shows that a principled likelihood-ratio statistic can outperform commonly used heuristic filters in the low-count regime, and because the pipeline includes a temporal subdivision option that can reveal short-lived transients that are diluted in stacked exposures.","feed_headline":"X-Sifter finds 30% more X-ray sources than standard catalog","feed_subtitle":"Poisson matched filter boosts S/N by 1.3x near detection limit, a 1.8x survey speed gain.","key_machinery":"The load-bearing object is the Poisson noise matched filter kernel $K_{\\mathrm{PMF}}(F)=\\ln(1+\\frac{F}{B}P)$, where $P$ is the normalized point spread function and $B$ the background, applied by cross-correlating the image with this kernel. Because the kernel depends on the unknown source flux $F$, the pipeline imposes a thresholding relation (Equation 14) that fixes the filter flux $F_{\\mathrm{th}}$ to equal the source flux at the detection limit, making the test approximately optimal for the faintest sources of interest. The implementation handles real data by partitioning each CCD into $16\\times16$ sectors with buffers, modeling or rotating PSF stamps per sector and per energy channel, estimating a constant background per sector from the median of gridded tiles, and combining independent energy channels in quadrature. Precomputed libraries of PSF stamps and of the noise distribution $P(S|H_0)$ for a grid of background values keep runtime practical.","core_discovery":"The central claim is that X-Sifter, a pipeline implementing the Poisson noise matched filter of Ofek & Zackay, recovers real Chandra sources with significantly higher sensitivity than the catalog standard. In a comparison restricted to fields observed only once with Chandra, X-Sifter detected about 30 percent more real sources than CSC2.1, all near the detection threshold. For sources detected by both methods, X-Sifter's signal-to-noise ratio was on average 1.3 times higher near the threshold, which the authors equate to a factor of 1.8 in survey speed. The paper further argues that temporal subdivision of exposures is essential for transient detection, since short bursts are buried in the background of long stacked exposures; in a check on fields observed many times, all quiescent sources found by X-Sifter had CSC2.1 counterparts, while the unmatched detections were variable on timescales shorter than the stack.","pith_inferences":["The sector-constant background assumption is likely the limiting factor in crowded fields or regions with strong background gradients; testing on such fields would show whether the claimed gain holds there.","The exponential extrapolation of the gamma-S curve used for high-significance conversion is acknowledged in the paper to lose accuracy, so X-Sifter's signal-to-noise values above about 7 sigma may be less reliable than the catalog's.","The same filtering formalism could be applied to count images in other Poisson-dominated regimes, such as gamma-ray or neutrino telescopes, with appropriate PSF and background models."],"forward_implications":["Single-observation Chandra archival searches gain roughly 30 percent more detections near the threshold at the same false-alarm rate.","The 1.3-fold signal-to-noise gain near threshold translates into a 1.8-fold increase in survey speed, allowing surveys to reach the same depth in less exposure time.","Temporal subdivision of exposures, when enabled, can expose short transients that are missed by catalogs built from stacked images, at the cost of running the pipeline multiple times.","The kernel and thresholding relation are general for any Poisson imaging instrument; the authors state the pipeline can be adapted to XMM-Newton and similar data."],"supporting_citations":[{"why":"Supplies the Poisson matched filter formalism, the kernel expression, and the flux estimator used by X-Sifter.","marker":"Ofek & Zackay 2018a"},{"why":"Provides the thresholding relation and the approach for estimating P(S|H0) via simulations.","marker":"Ofek & Zackay 2018b"},{"why":"Defines the CSC2.1 catalog used as the comparison baseline for detection sensitivity.","marker":"Evans et al. 2024"},{"why":"Describes the source detection strategy of the earlier Chandra Source Catalog, based on wavdetect.","marker":"Evans et al. 2010"},{"why":"Presents the wavdetect algorithm, the filtering method that CSC uses and that X-Sifter outperforms.","marker":"Freeman et al. 2002"},{"why":"Establishes the likelihood-ratio optimality lemma that underlies the matched filter approach.","marker":"Neyman & Pearson 1933"},{"why":"Used to compute the 95 percent confidence upper limit on the false-positive rate from the small sample of unmatched detections.","marker":"Gehrels 1986"}],"fun_headline_variants":["X-Sifter detects 30% more X-ray sources than standard catalog","X-ray survey speed up 1.8x with new Poisson filter","X-Sifter boosts S/N by 1.3x, finds 30% more sources","New X-ray filter finds 30% more sources, speeds surveys","Poisson matched filter ups X-ray detection by 30%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes the background is constant within each 16 by 16 sector and estimates it as the median of gridded tiles after outlier rejection; if the true background varies inside a sector, the filter kernel and threshold are miscalibrated and the claimed sensitivity gain may not be realized.","fun_headline_variants_meta":{"raw":{"variants":["X-Sifter detects 30% more X-ray sources than standard catalog","X-ray survey speed up 1.8x with new Poisson filter","X-Sifter boosts S/N by 1.3x, finds 30% more sources","New X-ray filter finds 30% more sources, speeds surveys","Poisson matched filter ups X-ray detection by 30%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000625,"raw_usage":{"total_tokens":2875,"prompt_tokens":910,"completion_tokens":1965,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":1867}},"tokens_in":526,"tokens_out":1965,"duration_ms":12568,"temperature":1.0,"reasoning_tokens":1867,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:28:40.780013+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run X-Sifter on simulated images with a known background gradient inside a single sector, such as a linear ramp across the sector with the same mean as a flat field, and compare the recovered flux and signal-to-noise of injected sources to the flat case; if the detection threshold shifts or completeness drops by more than the claimed margin, the sector-constant assumption is falsified. Alternatively, apply the code to a real field with strong structured background, such as the Galactic ridge, and compare against a deeper reference catalog: if the 30 percent gain vanishes or false positives appear, the assumption fails.","supporting_citations":[{"cited_title":"E., Kashyap, V., Rosner, R., & Lamb, D","cited_arxiv_id":null,"evidence_quote":"Presents the wavdetect algorithm, the filtering method that CSC uses and that X-Sifter outperforms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the likelihood-ratio optimality lemma that underlies the matched filter approach."}],"review_version":1}