{"id":"7cda68d8-0840-4294-b2bd-6b3bdfcbe4f6","arxiv_id":"2507.02105","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Using identical photometric pipelines, the diffuse intracluster light fraction in TNG300 simulated clusters matches WWFI observations for most definitions, although simulated BCG+ICL systems are larger and brighter.","lead":"Scientists compared the faint glow of stray stars in galaxy clusters, called intracluster light, between a supercomputer simulation and deep telescope images, processing both identically. They find simulated clusters are about twice as spread out and brighter in surface brightness, yet the fraction of intracluster light agrees for most measurement methods, which matters for calibrating simulations against future surveys.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Correction factors from TNG300 are applied to observed fICL without independent validation; the claimed simulation-observation agreement may be partly manufactured.","rationale":"The reader's weakest_assumption identifies the same load-bearing point: correction factors are calibrated and validated only on the simulation, then applied to observations. My independent reading confirms this is the most fragile step. The abstract and Section 4.5 claim agreement for most fICL definitions, but the SB27 and de Vaucouleurs methods rely on C_flux corrections, and the 2rhalf method relies on both C_flux and C_rhalf. The paper tests the corrections only on TNG300 itself, showing the corrected SB-limited measurements recover the full simulated measurements. That is a self-consistency check, not a validation for real data. The paper also acknowledges that the de Vaucouleurs method is sensitive to the fitted region and that observed fICL scatter is dominated by uncertainties near the 30 mag limit, which underscores how sensitive the corrected values are to the extrapolation. I do not see a fatal flaw, so ACCEPT is not warranted, but the central claim is conditional on this validation. The concrete test I propose would settle whether the agreement survives without the correction or with an independent extrapolation. The reader's recommendation of CONDITIONAL is appropriate, so I keep the verdict unchanged.","tokens_in":25799,"tokens_out":1629,"duration_ms":21180,"concrete_test":"Recompute all observed fICL values without applying the correction factors (i.e., set C_flux = 1 and use SB-limited fluxes only), and compare medians to the simulation's own SB-limited fICL values. Additionally, recompute observed fICL using an independent correction derived from Sersic fits to each observed BCG+ICL profile, as in Kluge et al. (2020). If the median fICL agreement with TNG300 disappears or shifts by more than ~0.05 under either alternative, the claimed consistency is not robust. If it persists, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that median fICL agrees between TNG300 and WWFI for most definitions depends on applying simulation-derived correction factors, equations (2)-(5), to the observed images through equation (7). These factors are measured only from TNG300 raw synthetic images, and validated only by recovering the same simulation's full measurements (solid versus dashed red lines in Figures 8-11). If real clusters contain a different fraction of light below the 30 g' mag arcsec^-2 isophote, the corrected observed fICL values are biased toward the simulation. The paper itself notes that the TNG300 corrections (e.g., C_flux,circ = 1.17) are larger than those from Sersic extrapolation by Kluge et al. 2020 (1.09), implying the simulated outer profiles are shallower than observed extrapolations. Since both numerator and denominator in fICL are rescaled using these corrections, the claimed agreement near 0.3 could be an artifact of applying a simulation-specific extrapolation to real data. The paper does not present a no-correction comparison, nor an independent validation using observed data, so the central 'majority of definitions agree' claim rests on an unverified assumption. This is a correctness risk, not merely a consensus disagreement, because it is internal to the method's logic.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper forward-models 40 massive TNG300 clusters at z≈0.06 into synthetic g'-band images that mimic Wendelstein Wide Field Imager (WWFI) observations, and applies identical satellite masking, background subtraction, and photometric measurements to synthetic and real images. The authors compare BCG+ICL sizes, surface brightness profiles, and the intracluster light fraction fICL under four definitions: SB cut at 27 g' mag arcsec^-2, de Vaucouleurs excess light, light beyond 2rhalf, and fixed circular apertures of 30, 50, and 100 kpc. They report that simulated BCG+ICL systems are roughly twice as extended and ~1 g' mag arcsec^-2 brighter than observed ones, but that the median fICL is consistent between simulation and observation for the SB27, deV, and 2rhalf definitions, with the fixed-aperture method showing a ~15% excess in TNG300. The larger observed scatter in fICL is attributed primarily to observational uncertainties in the total BCG+ICL flux.","tokens_in":26090,"tokens_out":7390,"duration_ms":84381,"significance":"If the central result holds, it is valuable: it separates the ICL fraction, a relative quantity, from the absolute normalization offset of the simulated BCG+ICL light, and it demonstrates that forward-modeling with identical processing can yield meaningful apples-to-apples comparisons. The paper is transparent about its pipeline, tests the effect of image smoothing in Appendix B, and compares with prior work quantitatively. However, the main claim of fICL consistency rests on simulation-derived correction factors applied to the observed images, and on a modified de Vaucouleurs fit threshold; these choices require robustness checks before the claim can be considered fully supported.","major_comments":[{"comment":"The correction factors C_flux and C_rhalf (equations 2-5) are measured from TNG300 raw synthetic images and applied to the WWFI observations through equation (7), but they are validated only by checking that the corrected 'SB-limited' measurements recover the 'full' measurements in the simulation itself (solid versus dashed red lines in Figures 8-11). This does not test whether real clusters contain the same fraction of light below the 30 g' mag arcsec^-2 isophote. The paper itself notes that Kluge et al. (2020) obtained a flux correction of 1.09 from Sérsic extrapolation, while the TNG300-based C_flux,circ is 1.17, implying that the simulated outer profiles are shallower than observed extrapolations. Consequently, applying the TNG300 corrections to observations could bias the observed fICL toward the simulated values. The central claim of median consistency for most definitions depends on these corrections. The authors should present a version of Figures 9-12 without the corrections or with the alternative Sérsic-based corrections, and estimate the systematic uncertainty in fICL from the correction-factor calibration.","section":"Section 3.4, Eq. (7)"},{"comment":"The de Vaucouleurs inner fit threshold is changed from 23 g' mag arcsec^-2 (as in Kluge et al. 2021) to 27 g' mag arcsec^-2 because the lower threshold is too close to the TNG300 resolution limit. This change alone moves the observed fICL,deV from 0.48±0.20 (Kluge et al. 2021) to 0.26±0.19. Because the threshold is chosen partly to accommodate the simulation's resolution, the comparison is no longer based on an observationally motivated, fixed definition. To establish that the simulated-observed agreement for this method is not a coincidence of the chosen threshold, the authors should show how fICL,deV varies with the fit threshold (e.g., 23, 25, 27, 28 g' mag arcsec^-2) for both the TNG300 and WWFI samples.","section":"Section 4.2"}],"minor_comments":[{"comment":"The satellite-mask threshold T0 is changed from 0.15 (Kluge et al. 2020) to 0.3 without a stated reason; please justify this choice or show that the results are insensitive to T0.","section":"Section 2.4 (footnote 4)"},{"comment":"The sentence beginning 'The latter are somewhat larger than...' is ambiguous, because both circular and elliptical correction factors were listed immediately before; please specify which set is being compared.","section":"Section 3.4"},{"comment":"The median fICL values are reported without uncertainties on the medians; given the modest sample sizes (40 simulated and 38 observed clusters), bootstrap or jackknife confidence intervals would strengthen the comparison.","section":"Section 4.5 / Figure 12"},{"comment":"The claim that the observed scatter in fICL,SB27 is 'about an order of magnitude' larger than in the simulation is based on the 16-84 percentile range; consider also reporting a robust dispersion measure such as the median absolute deviation to ensure this is not driven by a few outliers.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-written and the forward-modeling methodology is a strength, and the authors are transparent about their assumptions. The main risk is the calibration of the correction factors: applying TNG300-derived corrections to observed data without an independent validation could bias the central agreement. The change of the de Vaucouleurs fit threshold is a related concern. I recommend major revision, asking the authors to add robustness tests for these choices, which should be straightforward to implement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The paper is a careful, transparent forward-modeling comparison of the ICL between 40 TNG300 clusters and the WWFI observations, and it does something the previous literature hasn't quite done: it applies the same satellite-masking and photometric pipeline to synthetic and real images for four different ICL definitions. The headline claim is that median fICL is consistent between simulation and observations for most definitions, settling near 0.3, while the BCG+ICL light in TNG300 is about twice as extended and about 1 mag brighter. That size discrepancy is the more robust and, to me, the more interesting result, since it is based on Petrosian radii and SB profiles that don't depend on the correction factors.\n\nWhat I like: the methodology is genuinely apples-to-apples, the sample masses are matched, the masking procedure is applied identically to both, and the smoothing-convergence test in Appendix B is a good robustness check. The paper is open about the arbitrariness of the de Vaucouleurs method and the background-subtraction choices.\n\nThe soft spot is exactly where the stress-test note points: equations (2)-(5) are correction factors measured from TNG300, applied to the WWFI images through equation (7), and validated only on the simulation itself. The paper notes that the TNG300 corrections (1.17 for flux) are bigger than Kluge's Sersic-based ones (1.09), which says the simulated outer profiles are shallower than the observed extrapolations. If real clusters have a different amount of light below the 30 g' limit, the corrected observed fICL shifts. But I don't think this is load-bearing. The observed fICL scatter is about +-0.2; a 7% change in the flux correction shifts the median fICL by maybe 0.02-0.03, which does not change the conclusion of consistency. And the factor-of-2 size difference is independent of the corrections. What the paper needs is a robustness test, a no-correction version or a version using the Kluge corrections on the observed side, and a quantitative demonstration that the scatter attribution is right. Right now the attribution to observational noise is an argument from the SB profile shapes, not a measurement.\n\nMinor but relevant: the manual masking step is not fully specified, and the synthetic images and code are not released, so full reproducibility is currently limited.\n\nThe central argument holds up. This is a solid paper for ICL researchers and for anyone benchmarking simulations against deep imaging. It deserves a serious referee; the revision should add the sensitivity test and ideally release the images. I'd send it to review.","headline":"Solid forward-modeling comparison; ICL fraction agreement holds up, but the simulation-derived correction factors deserve a sensitivity test before publication.","tokens_in":26685,"tokens_out":4221,"would_cite":true,"duration_ms":47515,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TNG300 and wide-field observations agree the intracluster light fraction of massive clusters is about 30 percent under identical photometric methods, even though the simulated central systems are twice as extended and a magnitude brighter.","keywords":["methods: numerical","techniques: image processing","galaxies: clusters: general","galaxies: haloes","galaxies: formation","galaxies: photometry","intracluster light","ICL fraction"],"falsifier":"Take deeper observations of the same clusters—longer exposures or stacking that reach 32 $g'$ mag arcsec$^{-2}$ or fainter—and measure their surface brightness profiles directly past the old 30 $g'$ mag arcsec$^{-2}$ cutoff; if the observed profiles fall below the TNG300-based extrapolations by more than the stated uncertainties, the $f_{\\rm ICL}$ agreement would shrink or vanish. A complementary check is to measure $f_{\\rm ICL}$ for the same clusters by a route that never extrapolates through the simulations, such as planetary-nebulae kinematics or globular-cluster counts.","tokens_in":25550,"feed_emoji":"🔭","tokens_out":13153,"duration_ms":119612,"temperature":0.7,"pith_summary":"This paper tries to settle whether the diffuse starlight that floats between galaxies inside massive clusters—the intracluster light, or ICL—looks the same in a large cosmological simulation as it does in deep telescope images. The authors generate synthetic $g'$-band images of 40 massive clusters from the TNG300 simulation that reproduce the pixel scale, point-spread function, noise, depth, and satellite masking of the Wendelstein Wide Field Imager survey, and then run identical photometric routines on the synthetic and real images. They find that the fraction of the central light budget held by the ICL is about 0.3 in both simulation and observations for three of the four standard ICL definitions, so the ICL fraction is not where theory and data diverge. What does diverge is the absolute light distribution: the simulated brightest-cluster-galaxy-plus-ICL systems are about twice as extended and roughly 1 magnitude per square arcsecond brighter than observed ones. The result matters because it isolates the genuine discrepancy for galaxy-formation models and indicates that measured ICL fractions near 0.3 are robust to the choice of definition.","feed_headline":"Simulated and real clusters share a 30% intracluster light fraction","feed_subtitle":"The ICL budget matches at ~30%; the real gap is size and brightness of the central systems.","key_machinery":"The carrier of the argument is a forward-modeling comparison pipeline: every TNG300 cluster is rendered as a synthetic $g'$-band image with the same pixel scale (0.2 arcsec), point-spread function (1.2 arcsec FWHM), Poisson shot noise, background noise, 30 $g'$ mag arcsec$^{-2}$ depth, and satellite-masking recipe used on the WWFI frames, so that any remaining difference between the two data sets is a physical difference rather than a methodological one. The identity doing the quantitative work is the ICL fraction $f_{\\rm ICL} = F_{\\rm ICL} / F_{\\rm BCG+ICL}$, evaluated through four competing definitions, together with the simulation-derived correction factors $C_{r_{\\rm half}} \\approx 1.3$–$1.4$ and $C_{\\rm flux} \\approx 1.15$–$1.17$ (equations 2–5 of the paper) that convert 'SB-limited' observed measurements into estimates of the total light out to $r_{\\rm 200,crit}$.","core_discovery":"On the paper's own terms, the discovery is that an 'apples-to-apples' comparison—synthetic TNG300 images run through exactly the same masking, background subtraction, and photometric analysis as the WWFI observations—yields median intracluster light fractions that are consistent between simulation and observations for most definitions: $f_{\\rm ICL} = 0.33 \\pm 0.02$ (TNG300) versus $0.34 \\pm 0.19$ (WWFI) for the 27 $g'$ mag arcsec$^{-2}$ surface-brightness cut; $0.28 \\pm 0.05$ versus $0.26 \\pm 0.19$ for the de Vaucouleurs excess; and $0.33$–$0.34$ for the $2 r_{\\rm half}$ method on both sides. The same pipeline shows that the simulated BCG+ICL is about twice as extended (median half-light radius of 74 kpc versus 34 kpc for circular apertures) and about 1 $g'$ mag arcsec$^{-2}$ brighter in surface brightness. The larger observed scatter in $f_{\\rm ICL}$ is attributed primarily to observational uncertainties in the total BCG+ICL luminosity near the 30 $g'$ mag arcsec$^{-2}$ detection limit rather than to genuine cluster-to-cluster variation in the real Universe.","pith_inferences":["My inference: the same matched-pipeline correction scheme could be applied to other cosmological simulations, or to TNG300 with a different resolution, to test whether the ~0.3 ICL fraction and the factor-of-two size offset are generic properties of current galaxy-formation models or specific to TNG300.","My inference: if the observational scatter in $f_{\\rm ICL}$ is truly dominated by the detection limit, then upcoming deeper surveys should measure scatter consistent with the simulation when pushed to comparable depth—a direct, testable consequence the paper does not itself state.","My inference: the mass-independence of $f_{\\rm ICL}$ found here, together with the transition radius near $2 r_{\\rm half}$, suggests the ICL fraction may be approximately scale-invariant down to galaxy groups; checking TNG300's lower-mass haloes would test that.","My inference: since the fractional ICL budget matches while the absolute profile does not, the ratio of diffuse to bound stellar light appears to track the dark-matter halo shape, whereas the profile normalization is set by the feedback model—a separation that could help diagnose the stellar-mass-to-halo-mass relation at the cluster scale."],"forward_implications":["If the agreement holds, the ICL fraction is a stable, method-independent observable: roughly 30% of the BCG+ICL light in massive clusters is diffuse across three independent definitions.","The factor-of-two size offset and ~1 mag arcsec$^{-2}$ brightness offset—not the ICL budget—become the target for improving galaxy-formation models at the cluster scale, pointing to the amount and distribution of accreted (ex situ) stars.","If the large observed scatter in $f_{\\rm ICL}$ is mostly measurement noise near the detection limit, deeper or stacked observations should reveal intrinsic scatter in real clusters as narrow as the simulation's.","The convergence of most definitions near $f_{\\rm ICL} \\approx 0.3$ supports $2 r_{\\rm half}$ (roughly a 100 kpc aperture in this mass range) as a practical BCG/ICL boundary that observers can adopt without profile fitting.","Fixed-aperture definitions of 30 or 50 kpc inherit the systematic size offset and should not be used for direct simulation–observation comparisons of ICL fractions."],"supporting_citations":[{"why":"Provides the 170-cluster WWFI image sample, its exposure depth and PSF, and the satellite-masking recipe applied identically to synthetic and real images.","marker":"Kluge et al. (2020)"},{"why":"Defines the SB27 and de Vaucouleurs ICL fractions used as baselines and supplies the observational Sérsic-extrapolated corrections compared with the simulation-derived ones.","marker":"Kluge et al. (2021)"},{"why":"Supplies the STATMORPH code used for 2D Sérsic fits and the adaptive-kernel synthetic image rendering method.","marker":"Rodriguez-Gomez et al. (2019)"},{"why":"Provides the GALAXEV stellar population models that convert each stellar particle into a g'-band flux in the synthetic images.","marker":"Bruzual & Charlot (2003)"},{"why":"Establishes the IllustrisTNG ICL predictions and the 2 r_half aperture convention adopted as one of the ICL definitions.","marker":"Pillepich et al. (2018b)"},{"why":"Gives the 3D stellar-mass-based f_ICL values for TNG300 clusters that the photometric 2 r_half results are shown to match.","marker":"Montenegro-Taborda et al. (2025)"},{"why":"Documents the earlier forward-modeled ICL analysis and the sensitivity of ICL fractions to adaptive smoothing, which Appendix B checks for the adopted smoothing level.","marker":"Tang et al. (2018)"}],"fun_headline_variants":["Simulation and sky agree: intracluster light is 30% of the central glow","TNG300 and real clusters both put ~30% of light in intracluster glow","Intracluster light fraction matches between simulation and actual sky","Simulated and observed galaxy clusters share a third of their light in ICL","Real and simulated clusters agree on intracluster light share of ~30%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the correction factors for light hidden below the 30 $g'$ mag arcsec$^{-2}$ limit, measured from the TNG300 synthetic images, apply to real galaxy clusters; if actual clusters place a different fraction of their light below that isophote, the corrected observed ICL fractions are biased toward agreement with the simulation.","fun_headline_variants_meta":{"raw":{"variants":["Simulation and sky agree: intracluster light is 30% of the central glow","TNG300 and real clusters both put ~30% of light in intracluster glow","Intracluster light fraction matches between simulation and actual sky","Simulated and observed galaxy clusters share a third of their light in ICL","Real and simulated clusters agree on intracluster light share of ~30%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000678,"raw_usage":{"total_tokens":3243,"prompt_tokens":1269,"completion_tokens":1974,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":885,"completion_tokens_details":{"reasoning_tokens":1874}},"tokens_in":885,"tokens_out":1974,"duration_ms":14419,"temperature":1.0,"reasoning_tokens":1874,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:38:52.744150+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take deeper observations of the same clusters—longer exposures or stacking that reach 32 $g'$ mag arcsec$^{-2}$ or fainter—and measure their surface brightness profiles directly past the old 30 $g'$ mag arcsec$^{-2}$ cutoff; if the observed profiles fall below the TNG300-based extrapolations by more than the stated uncertainties, the $f_{\\rm ICL}$ agreement would shrink or vanish. A complementary check is to measure $f_{\\rm ICL}$ for the same clusters by a route that never extrapolates through the simulations, such as planetary-nebulae kinematics or globular-cluster counts.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the SB27 and de Vaucouleurs ICL fractions used as baselines and supplies the observational Sérsic-extrapolated corrections compared with the simulation-derived ones."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the GALAXEV stellar population models that convert each stellar particle into a g'-band flux in the synthetic images."}],"review_version":1}