{"id":"695a4d29-8ad6-4e00-9144-269de616f7f6","arxiv_id":"2501.16708","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Forward modeling of local BASS AGN through the Chandra COSMOS-Legacy survey shows that most obscured AGN are missed and low-count fits systematically overestimate column densities.","lead":"This paper simulates what the Chandra COSMOS-Legacy survey would see if it observed 2,280 nearby, well-measured AGN from the BASS catalog. It finds that over half of simulated obscured AGN would be missed, and that many sources classified as heavily obscured are actually unobscured, which suggests previous X-ray surveys may have overstated how quickly the obscured fraction grows with redshift.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline percentages (53.3%, 66.7%) are conditional on assuming the local BASS NH distribution and spectral shapes are unchanged at CCLS redshifts; the false-CT rate is prior-dominated because the simulation deliberately did not match the CCLS NH distribution.","rationale":"I read the paper in good faith as a transparent forward-modeling study with a valuable and novel population-level selection-function calculation. The 2PL degeneracy analysis, the false-CT examples, and the AXIS projections are concrete and useful. The reader's conditional verdict is appropriate: the machinery is sound, but the quantitative headline claims are conditional on a no-evolution prior for the intrinsic AGN population and its NH distribution. This is not an internal inconsistency—the authors explicitly flag it in Sec 4.1—but it is load-bearing because the false-CT majority (18/27) and the implied correction to f22(z) evolution are computed under that prior. The paper would be strengthened by quantifying the sensitivity of the headline percentages to a range of adopted intrinsic NH distributions. My concern does not move the verdict; it reinforces the reader's CONDITIONAL recommendation, so I report UNCHANGED.","tokens_in":39408,"tokens_out":5657,"duration_ms":60866,"concrete_test":"Re-run the simulation with importance weights on BASS templates drawn from an evolved intrinsic NH distribution, e.g., the Buchner et al. (2015) X-ray background synthesis model or the Aird et al. (2015)/Ueda et al. (2014) obscured-fraction evolution, keeping the same spectral shapes; recompute the headline fractions (53.3% missed obscured, 66.7% false-CT) and the f22(z) evolution in Fig 9. If the false-CT majority and the implied correction to the obscured-fraction evolution move outside the quoted ranges, the results are prior-dominated and must be reported as conditional. A complementary check is to use the CCLS* NH distribution from M16/Lanzuisi et al. (2018) as the intrinsic prior instead of the BASS NH distribution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claims—53.3% of obscured sources missed and 66.7% of best-fit CT sources being unobscured—are computed from a simulated population whose intrinsic NH distribution is the local BASS distribution, not the CCLS population. In Sec 2.2 the authors match the CCLS* luminosity, redshift, and exposure-time distributions, but explicitly do not match the CCLS NH distribution (Appendix C), instead drawing templates from local BASS. The false-CT statistic is especially fragile: it is based on only 27 best-fit CT simulations, 18 of which are unobscured, and the template pool contains a large fraction of unobscured sources (204/380 with log NH < 22). If high-redshift AGN have a higher intrinsic obscured/CT fraction, as X-ray background synthesis models and multiwavelength studies suggest, both the missed-obscured fraction and the false-CT fraction will shift. The interpretation in Sec 4.1 and Fig 9—that the apparent increase of obscured fraction with redshift is largely a measurement bias—assumes no cosmic evolution of the intrinsic NH distribution or of spectral shapes. The authors acknowledge this in Sec 4.1 but do not quantify it. Thus the headline percentages are not yet directly statements about CCLS; they are statements about a local BASS prior placed at CCLS redshifts and luminosities.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper quantifies how the Chandra COSMOS-Legacy Survey (CCLS) selects against obscured AGN by forward-modeling 380 well-measured local BASS AGN spectral templates, placing them at CCLS redshifts, luminosities, and exposure times, simulating Chandra observations with fakeit, and then fitting the simulated spectra with the same phenomenological pipeline used by Marchesi et al. (2016b). The main quantitative results are that 53.3% (563/1056) of simulated obscured AGN (log N_H >= 22) are undetected, and that among the 27 simulated sources whose best fit implies Compton-thick column densities (log N_H >= 24), 66.7% (18/27) are actually unobscured (log N_H <= 22). The paper also studies photon-index recovery, the role of double power-law fits in producing false CT classifications, refits the false CT sources with MYTorus, and repeats the detection simulation for the proposed AXIS mission. The authors conclude that previous X-ray survey results may have significantly overestimated the growth of the obscured fraction with redshift and the fraction of luminous obscured AGN.","tokens_in":39664,"tokens_out":6078,"duration_ms":61089,"significance":"If the quantitative claims are taken as statements about the CCLS population, the paper is an important, first-of-its-kind large-sample quantification of obscuration bias in a deep Chandra survey, with clear implications for interpreting existing AGN population studies and for designing future missions such as AXIS. The forward-modeling design is a genuine strength: known BASS spectra are passed through a known detector response and a well-defined external fitting pipeline, so the inference is not circular. The paper is also transparent, with simulation catalogs (Tables 4 and 5), appendices testing the fakeit procedure and background assumptions, and concrete AXIS projections. The main limitation is that the simulated intrinsic N_H distribution is deliberately taken from the local BASS parent population rather than from CCLS or from a model of the high-redshift AGN population; all headline percentages are therefore conditional on that prior. This does not invalidate the forward-modeling approach, but it does mean the headline numbers and the redshift-evolution interpretation in Sec. 4.1 are not yet direct statements about the intrinsic CCLS population.","major_comments":[{"comment":"The headline missed-obscured fraction (53.3%; 563/1056) is computed from a simulated population whose intrinsic N_H distribution is the local BASS distribution, not the CCLS distribution. The authors state in Sec. 2.2 that they matched luminosity, redshift, and exposure time but explicitly did not match the CCLS N_H distribution, and Appendix C shows that the simulated N_H distribution differs strongly from CCLS (e.g., 204 of 380 templates have log N_H < 22). Because the detection fraction in Fig. 5 varies steeply with N_H (for example, only 5.6% of sources with log N_H >= 24 are detected at log L = 43.0-43.5, versus 80.3% for log N_H < 21), the 53.3% figure depends directly on the adopted BASS N_H prior. The paper should either state the headline result as conditional on that prior and quantify the sensitivity to alternative intrinsic N_H distributions (e.g., an increasing obscured fraction with redshift), or reweight the simulation to a plausible CCLS-like N_H prior. Without this, the abstract's phrasing that Chandra would fail to detect the majority of obscured sources 'given the observed redshift and luminosity distribution of the CCLS' is incomplete, since the N_H distribution is not observed in the same sense.","section":"Sec. 2.2 and Appendix C"},{"comment":"The false-CT fraction (66.7%; 18/27) is prior-dominated and based on small numbers. The 27 best-fit CT sources are drawn from a simulation in which more than half of the templates are unobscured (204/380 with log N_H < 22), so the statement that most best-fit CT sources are actually unobscured is a property of the assumed BASS input population, not of CCLS. If the intrinsic high-redshift population has a higher obscured/CT fraction, as argued by X-ray background synthesis models and multiwavelength studies, the false-CT rate would be lower. The paper should report this fraction with an explicit binomial uncertainty (27 sources gives a 95% interval of roughly 46-83% for the 66.7% estimate) and should present the false-CT probability as a function of the assumed intrinsic obscured fraction, or at least prominently state that the number is conditional on the BASS N_H prior and is not a measured CCLS property.","section":"Sec. 3.4, Fig. 7, and abstract"},{"comment":"The detection-fraction calculation relies on an incompletely specified 'selection function' correction for sources with fewer than 30 counts. The text says that 'we included a function to randomly reduce the number of sources at the same rate that Chandra would not detect because of background,' but the exact algorithm, the background model, and the matching procedure are not described in the main text, and Appendix D only gives a qualitative discussion (e.g., ~1.4 background counts per aperture, a 3-sigma threshold of ~12 net counts). Since the 67.2% overall detection fraction and thus the 53.3% missed-obscured headline both depend on this correction, the paper should specify the exact correction function, its parameters, and the resulting count-distribution match, and should show how the headline numbers change if the correction is varied or omitted.","section":"Sec. 3.1 and Appendix D"},{"comment":"The inference that the observed increase in the obscured fraction with redshift is largely a measurement bias assumes no cosmic evolution of the intrinsic N_H distribution or of AGN spectral shapes. The manuscript acknowledges this in Sec. 4.1 ('does not account for any potential change to the intrinsic population of AGN over cosmic time'), but the discussion nevertheless concludes that 'the trend of increasing obscuration with redshift may be much less significant than previously believed' and that previous studies 'may have significantly overestimated' the redshift growth of the obscured fraction. These statements go beyond what the simulation alone can establish, because the simulated 'true' obscured fraction in Fig. 9 is itself derived from the local BASS N_H-luminosity relation. The paper should either restrict the conclusion to the conditional statement 'if the high-redshift population has the same intrinsic N_H distribution and spectral shapes as local BASS AGN, then the X-ray-measured trends overestimate the intrinsic trends,' or include a quantitative exploration of how the bias changes under simple evolutionary scenarios for the intrinsic N_H distribution.","section":"Sec. 4.1 and Fig. 9"}],"minor_comments":[{"comment":"The figure cross-references appear to be swapped: Sec. G.3 ('high secondary normalization') refers to Fig. 23, but Fig. 23 shows a very low secondary normalization (ratio = 7.03e-05), while Sec. G.4 ('low secondary normalization') refers to Fig. 22, which shows ratio = 0.503. The text and figures should be matched.","section":"Appendices G.3 and G.4"},{"comment":"The paper uses 'unobscured' in two senses: in Sec. 3.4 the 18/27 false-CT sources are described as 'unobscured' with log N_H <= 22, whereas Sec. 3.5 defines false-CT sources as those with log N_H_sim <= 20. These definitions should be reconciled or explicitly distinguished to avoid confusion about which quantity is being reported.","section":"Sec. 3.4 and Sec. 3.5"},{"comment":"The sentence 'It is notable that almost all of our fits that settled on a secondary normalization very near the upper limit turned out to be false CT source' is missing a period and should read 'false CT sources'; more generally, the caption would benefit from stating the number of false CT sources shown.","section":"Sec. 3.5 and Fig. 8 caption"},{"comment":"The statement that 'the CCLS notes that there is significant incompleteness below 20 counts' should include a specific reference (e.g., Civano et al. 2016, with section or figure) so the reader can verify the incompleteness claim.","section":"Sec. 3.1, footnote 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid forward-modeling study with clear value for the BASS and X-ray survey communities. My main concern is that the headline percentages and the redshift-evolution interpretation are presented as direct statements about CCLS, while the simulation deliberately imposes the local BASS N_H distribution and does not model cosmic evolution. This is a fixable framing/quantification issue rather than a fundamental flaw: the authors should sharpen the conditional language, add sensitivity tests to alternative N_H priors, and report small-number uncertainties. The manuscript is within the scope of the journal and the results, once properly conditioned, should be publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a solid, useful paper. The new contribution is a population-level selection function for CCLS built from 2280 simulated spectra using 380 high-fidelity BASS templates, plus the first large-sample quantification that the two-power-law fitting pipeline actually manufactures Compton-thick classifications: 11 false CT sources, with 66.7% of best-fit CT being truly unobscured. The simulation design is careful and transparent. They match the CCLS luminosity, redshift, and exposure distributions, use a conservative detection threshold, and the appendices honestly test the fakeit procedure and background effects. The direction and rough magnitude of the bias are robust: deep soft-X-ray surveys miss a large fraction of obscured AGN, and low-count spectral fitting overestimates column densities in a systematic way. This is real and worth knowing.\n\nThe soft spots are real but not fatal. The headline percentages (53.3% missed obscured, 66.7% false CT) are conditional on the intrinsic NH distribution and spectral shapes of local BASS AGN being unchanged at CCLS redshifts. The paper says this in Section 4.1 but does not quantify how much the numbers would shift under an evolving NH distribution or different spectral shapes. The false-CT fraction in particular rests on only 27 best-fit CT simulations, and the template pool is 54% unobscured, so that number is prior-dominated. A higher intrinsic obscured/CT fraction at high redshift, as many synthesis models suggest, would change the false-CT rate. Also, the authors discard fits that hit the normalization limits without reporting how many, and the headline percentages lack uncertainties. These are presentation and robustness gaps, not design flaws.\n\nMy overall take: the central argument holds. Even if the exact percentages are not directly statements about CCLS, the paper convincingly shows that previous studies likely overestimated both the redshift evolution of the obscured fraction and the luminous obscured fraction. The authors are appropriately cautious about their own assumptions. I would send this to peer review without hesitation. A serious referee can push for the robustness analysis, but the core result deserves to be in the literature.","headline":"A careful, transparent forward-model simulation that convincingly shows Chandra-like surveys severely miss obscured AGN and that low-count 2PL fits fabricate CT sources; the headline fractions are conditional on the local BASS prior but the central bias argument holds.","tokens_in":40283,"tokens_out":1441,"would_cite":true,"duration_ms":16672,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Chandra's deep COSMOS-Legacy survey misses most obscured AGN and mislabels unobscured sources as Compton-thick, so reported obscuration evolution is likely overstated.","keywords":["AGN obscuration","Compton-thick AGN","Chandra COSMOS-Legacy","BASS survey","X-ray selection bias","forward modeling","column density","obscured fraction evolution"],"falsifier":"Re-observe the 27 CCLS sources with best-fit $\\log N_H \\ge 24$ from Marchesi et al. (2016b) with NuSTAR, or stack their Chandra data at their known redshifts: if most show an unobscured soft continuum with no hard-band excess, the false-Compton-thick claim is supported; if most show flat reflected hard spectra with $\\log N_H \\ge 24$, the central inference fails.","tokens_in":39206,"feed_emoji":"🔭","tokens_out":8420,"duration_ms":74316,"temperature":0.7,"pith_summary":"The paper tests how much the deepest Chandra surveys are biased against obscured active galactic nuclei (AGN) by taking 380 well-measured, nearby AGN spectra from the BAT AGN Spectroscopic Survey, moving them to the redshifts, luminosities, and exposure times of the Chandra COSMOS-Legacy Survey, and re-running the survey's own fitting pipeline on the simulated data. It finds that Chandra would detect only 46.7% of simulated sources with obscuring column densities above $10^{22}\\ \\mathrm{cm}^{-2}$, and only 8.9% of Compton-thick ($N_H \\ge 10^{24}\\ \\mathrm{cm}^{-2}$) sources. Among the sources with enough counts to fit spectra, the fitted column density is systematically overestimated, and the majority of objects classified as Compton-thick—18 of 27—are actually unobscured. The paper concludes that reported increases of the obscured fraction with redshift, and the measured numbers of luminous obscured AGN, have been significantly inflated by these selection effects.","feed_headline":"Chandra misses most obscured black holes in deep-field survey","feed_subtitle":"Simulated local AGN at COSMOS redshifts: 53% of obscured black holes are missed and most 'Compton-thick' fits are false.","key_machinery":"The forward-modeling pipeline and the double power-law degeneracy, quantified by the primary ratio (PR). The machinery is: (1) 380 BASS X-ray spectral templates with known column density, photon index, and iron line, built from spectra with a median of 1545 counts; (2) each template placed at six CCLS-like redshifts and exposure times and passed through Chandra ACIS response files; (3) the four-step phenomenological fitting procedure from Marchesi et al. (2016b), going from an absorbed power law with fixed photon index to a free photon index, a secondary unabsorbed power law, and an iron K$\\alpha$ Gaussian; and (4) the PR diagnostic, $\\mathrm{PR} = (n_p - n_s)/(n_p + n_s)$, the fraction of detected photons belonging to the primary versus secondary component. When the energy at which the two power laws cross lies outside Chandra's 0.5–7 keV band, the PR shows that the double power law cannot distinguish an unobscured single power law from a heavily obscured source with soft leakage, which is exactly what produces the false Compton-thick classifications. The PR is what makes the bias systematic and predictable rather than a random fitting artifact.","core_discovery":"The central claim, stated as the authors would state it: at the depth of the Chandra COSMOS-Legacy Survey, the standard Chandra-based method for measuring AGN obscuration is not merely incomplete but asymmetrically biased. Given the CCLS redshift-luminosity distribution, Chandra would fail to detect 53.3% (563/1056) of obscured simulated BASS AGN ($N_H \\ge 10^{22}\\ \\mathrm{cm}^{-2}$) and 91.1% (175/192) of Compton-thick AGN ($N_H \\ge 10^{24}\\ \\mathrm{cm}^{-2}$). For detected spectra with at least 30 counts, fitting exactly as in Marchesi et al. (2016b) recovers column densities accurately on average (median $\\Delta \\log N_H = -0.10$ dex) but with a large tail of overestimates; among 27 spectra whose best fit placed them in Compton-thick territory, 18 (66.7%) were simulated from completely unobscured templates ($N_H \\le 10^{20}\\ \\mathrm{cm}^{-2}$). A double power-law model with a scattered secondary component is the culprit: it lets a low-count unobscured spectrum masquerade as a heavily absorbed one. Consequently, the measured obscured fraction increases with redshift and declines with luminosity even when the underlying simulated population is flat or declining, so earlier X-ray surveys likely overstate obscuration evolution.","pith_inferences":["If the local-template assumption holds, the same selection function can be inverted to produce a debiased obscured fraction as a function of redshift and luminosity; the paper stops at demonstrating the bias, so constructing that corrected distribution is a natural next step.","The false Compton-thick mechanism implies that multi-wavelength indicators—mid-infrared colors, X-ray-to-infrared flux ratios, or broad H$\\beta$ presence—should be used to down-select X-ray Compton-thick candidates, especially at $z > 1$, rather than trusting X-ray best fits alone.","The primary-ratio diagnostic could be turned into a survey-design tool: for any future mission bandpass and redshift window, one can precompute the range of $N_H$ values a double power-law fit can actually recover and quote completeness cells instead of point estimates.","Because the templates are local and no cosmic evolution of X-ray spectra is modeled, if high-redshift AGN have systematically weaker soft excesses or different reflection, the quantitative biases would shift; this sensitivity argues for stacking tests on real CCLS spectra to check the template assumption."],"forward_implications":["Any flux-limited Chandra survey at CCLS-like depth will miss more than half of obscured AGN and nearly all Compton-thick AGN at $z \\sim 0.5\\text{--}3$, so census numbers drawn from such surveys are lower limits.","Best-fit column densities from the standard phenomenological pipeline are unreliable for individual low-count sources; measured Compton-thick classifications need a physical torus refit, such as MYTorus, before being believed.","The reported rise of the obscured fraction with redshift and the measured abundance of luminous obscured AGN are inflated; the true evolution is weaker than X-ray fitting alone suggests.","A single power-law fit estimates $N_H$ at least as well as a double power-law fit for Chandra-quality data, so simpler models are preferable unless the data demand otherwise.","A next-generation X-ray mission with roughly ten times Chandra's effective area, such as AXIS, would detect 97.5% of the simulated sources and 51.0% of Compton-thick sources at 30 or more counts, making the bias tractable."],"supporting_citations":[{"why":"Supplies the CCLS survey definition, exposure times, detection thresholds, and source catalog that the simulations are designed to match.","marker":"Civano et al. (2016)"},{"why":"Its four-step X-ray spectral fitting procedure is replicated exactly in the simulations and defines the measured column densities.","marker":"Marchesi et al. (2016b)"},{"why":"Provides the phenomenological broadband X-ray models of BASS AGN that serve as the simulation templates.","marker":"Ricci et al. (2017c)"},{"why":"Defines the BASS sample and establishes the local AGN luminosity distribution used for template selection.","marker":"Koss et al. (2017)"},{"why":"Supplies the Compton-thick candidate analysis and the MYTorus-based refitting approach the paper uses to test false Compton-thick sources.","marker":"Lanzuisi et al. (2018)"},{"why":"Defines the MYTorus model that recovers the true column density for 10 of 11 false Compton-thick sources.","marker":"Murphy & Yaqoob (2009)"},{"why":"Provides the BNtorus model used to re-constrain the Compton-thick BASS templates.","marker":"Brightman & Nandra (2011)"},{"why":"One of the X-ray-based obscuration evolution models against which the biased measured fraction is compared in Fig. 9.","marker":"Aird et al. (2015)"},{"why":"The other evolution model in Fig. 9 whose obscured-fraction trends would be inflated by the reported fitting bias.","marker":"Peca et al. (2023)"}],"fun_headline_variants":["Chandra misses 53% of obscured AGN in COSMOS","COSMOS survey bias: false Compton-thick AGN","Simulated BASS shows Chandra selection bias","Most obscured AGN evade deep Chandra survey","CCLS obscuration bias quantified with BASS"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The intrinsic X-ray spectra of local BASS AGN—continuum slope, reflection hump, soft excess, and iron line—are exactly what high-redshift CCLS AGN look like once redshifted, with no cosmic evolution of spectral shape or of the obscuration–luminosity relation.","fun_headline_variants_meta":{"raw":{"variants":["Chandra misses 53% of obscured AGN in COSMOS","COSMOS survey bias: false Compton-thick AGN","Simulated BASS shows Chandra selection bias","Most obscured AGN evade deep Chandra survey","CCLS obscuration bias quantified with BASS"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1546,"prompt_tokens":1293,"completion_tokens":253,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":909,"completion_tokens_details":{"reasoning_tokens":177}},"tokens_in":909,"tokens_out":253,"duration_ms":3369,"temperature":1.0,"reasoning_tokens":177,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T11:15:12.547411+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-observe the 27 CCLS sources with best-fit $\\log N_H \\ge 24$ from Marchesi et al. (2016b) with NuSTAR, or stack their Chandra data at their known redshifts: if most show an unobscured soft continuum with no hard-band excess, the false-Compton-thick claim is supported; if most show flat reflected hard spectra with $\\log N_H \\ge 24$, the central inference fails.","supporting_citations":[],"review_version":1}