{"id":"15f656bd-8cdd-48ee-888a-d65451cfebf5","arxiv_id":"2506.18351","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Simulated JWST/NIRCam bulge+disk galaxies show that single and double Sérsic fits recover magnitudes within about 0.5 mag and sizes within about 0.2 dex for galaxies brighter than roughly 26 to 27 mag.","lead":"This paper simulates thousands of fake galaxies with a bulge and a disk, embeds them in real JWST images, and tests how well standard fitting methods recover their sizes and brightnesses. It finds reliable recovery for bright galaxies down to about 26 to 27 magnitudes, with concrete error budgets for future JWST studies.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-Sérsic size recovery is validated against a flux-weighted component radius (Eq. 2), not the true half-light radius of the two-component model; the reported 0.2 dex size accuracy may not be measuring what it claims.","rationale":"The reader's weakest_assumption focuses on the idealized model family and truth-initialized fitting, which are real limitations for external applicability but are largely acknowledged in §5.4 and mitigated in part by the footnote claiming consistent results with default initial parameters. The most load-bearing issue I find is internal: the single-Sérsic size recovery is compared to a flux-weighted composite radius that is not the half-light radius of the mock galaxy. Since the fitted R_eff is a half-light radius, this is a definitional mismatch, and the size-accuracy claim in the abstract and summary is therefore not properly validated. This is concrete, testable from the existing simulation setup, and does not depend on assumptions about real galaxies. I therefore recommend keeping the CONDITIONAL verdict, with an added condition that the size comparison be recomputed against the true half-light radius, and that the reported 0.2 dex thresholds be updated if the comparison shifts. The paper remains useful and the double-Sérsic component results are less affected, so a rejection is not warranted.","tokens_in":17458,"tokens_out":7542,"duration_ms":82234,"concrete_test":"For a subset of the 5,000 mocks, compute the true half-light radius of the noiseless two-component model by solving for the radius that encloses half the total flux from the analytic Sérsic cumulative profiles (or from the noiseless images), then re-derive the bottom panels of Figure 3 comparing the fitted single-Sérsic R_eff to this true R_half instead of the Eq. 2 R_weighted. If the binned mean offsets or scatters change by more than ~0.05 dex, or if the 'within 0.2 dex' threshold no longer holds at the stated magnitude limits, the headline single-Sérsic size-accuracy claim needs revision.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In §4.1.2, the single-Sérsic fits' effective radii are compared to a 'flux-weighted effective radius' defined in Eq. 2 as (R_bulge f_bulge + R_disk f_disk)/(f_bulge + f_disk). For a two-component Sérsic galaxy this is not the half-light radius of the total light distribution. The true half-light radius R_half solves the cumulative-flux condition ∫ I_bulge(r) + I_disk(r) dA = 0.5 F_total, and depends on the shapes of both Sérsic profiles—including the extended wings of a de Vaucouleurs bulge and the exponential disk's outer flux—not just on a linear flux-weighted average of the component effective radii. For typical parameters in the simulation (e.g., B/T ≈ 0.5, R_disk/R_bulge ≈ 2), the difference between R_weighted and the true R_half can be of order 0.1 dex, i.e., comparable to the quoted 0.2 dex accuracy. The single-Sérsic R_eff returned by the fit is, by definition, a half-light radius; comparing it to a non-half-light composite quantity is a unit mismatch. This directly affects the abstract and summary claims about recovering sizes within 0.2 dex down to 27 mag, and the same R_weighted is used in the galfit comparison in §5.3. The double-Sérsic component-size comparisons are not affected, but the single-Sérsic size validation is not established as stated. The paper's other limitation—that mocks and fits use the same idealized Sérsic family with fixed n_bulge = 4 and n_disk = 1—is acknowledged in §5.4 and concerns external applicability; the Eq. 2 issue concerns internal correctness of a headline result.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents controlled simulations of bulge+disk galaxies observed under CEERS/NIRCam conditions (F150W and F356W), fits them with single and double Sérsic models using the galight and galfit codes, and quantifies recovery accuracy for total and component magnitudes, effective radii, bulge-to-total ratios, model selection via the Bayesian Information Criterion, and signal-to-noise thresholds. The central numerical claims are that single Sérsic fits recover total magnitudes within about 0.5 mag and sizes within about 0.2 dex down to 27 mag, and double Sérsic fits recover component magnitudes and effective radii within 0.5 mag and 0.2 dex down to 26 mag, with SNR > 10 as a general reliability threshold.","tokens_in":17831,"tokens_out":6068,"duration_ms":53817,"significance":"If the headline results hold, the paper provides a practical reference for JWST/NIRCam morphological studies, including error budgets and BIC-based selection thresholds. The study's strengths are its large sample of 5,000 simulated galaxies, the use of realistic CEERS PSFs and background noise, explicit PSF-mismatch tests, an independent cross-check with galfit, and the SNR-based parameterization of uncertainties. However, the validation is a closed-box test: the generative model and the fitting model share the same Sérsic family with fixed bulge and disk indices, and the single-Sérsic size comparison is made against a flux-weighted radius that is not the true half-light radius of the composite system. These issues mean that the quoted accuracies are ideal-case, same-model errors unless the analysis is revised or the claims are reframed.","major_comments":[{"comment":"The single-Sérsic size recovery is validated against R_weighted = (R_bulge f_bulge + R_disk f_disk)/(f_bulge + f_disk), which is not the half-light radius of the two-component model. The true total half-light radius solves the cumulative-flux condition and depends on the full shape of both Sérsic profiles; for typical simulated parameters (e.g., B/T ≈ 0.5, R_disk/R_bulge ≈ 2) the difference can be ~0.1 dex, comparable to the quoted 0.2 dex accuracy. The abstract and §6 claims of single-Sérsic size recovery within 0.2 dex are therefore not established; the comparison should be made against the numerically computed half-light radius of the composite model, or the claim should be reframed. The same R_weighted is used in the galfit comparison in §5.3, so that consistency test does not validate the size accuracy either.","section":"§4.1.2, Eq. (2)"},{"comment":"The threshold statements 'within 0.5 mag' and 'within 0.2 dex' are based on binned means and standard deviations rather than on quantiles of the residuals. For F150W galaxies fainter than 26 mag the magnitude residual is 0.047 ± 0.323 mag; with a 1σ scatter of 0.32 mag a substantial fraction of objects must exceed 0.5 mag, so the statement in the text that 'the overall residuals remain far below 0.5 magnitudes, even for galaxies as faint as 28 mag' is not supported by the reported statistics. Please report the 68th and 90th percentiles of |Δmag| and |Δlog R_eff| and base all threshold claims on those quantiles.","section":"§4.1.2, Figure 3"},{"comment":"The abstract and §6 state that double-Sérsic effective radii are recovered within 0.2 dex for components brighter than 26 mag, but §4.2.2 reports bulge R_eff uncertainties of 0.17–0.24 dex for bright galaxies in F356W and 0.06–0.19 dex in F150W. The F356W bulge radii therefore exceed the stated 0.2 dex tolerance for part of the bright sample, and no fraction or quantile is given. The claim should be restricted to the band and component combination that actually meets the tolerance, or stated as a percentile-based fraction.","section":"§4.2.2 and Abstract"},{"comment":"The accuracy test is closed-box: the mocks are generated with the same Sérsic family used in the fits, with n_bulge fixed to 4 and n_disk fixed to 1, and the fitting is initialized at the true parameter values. The footnote in §3.2 asserts that default initialization gives consistent results, but no comparison is shown and should be presented. While §5.4 correctly lists bars, clumps, pseudo-bulges, and variable Sérsic indices as unmodeled, the abstract and summary do not carry the caveat that the quoted accuracies are ideal-case, same-model errors. I recommend adding a scope statement to the abstract and either quantifying the degradation on mocks with different indices or substructures, or explicitly stating that such a test is beyond the present scope.","section":"§3.1.1, §3.2, §5.4"}],"minor_comments":[{"comment":"The abstract says single-Sérsic total magnitudes are recovered within 0.5 mag, while §6 says within 0.2 mag for galaxies as faint as 27 mag; these numbers should be reconciled, and 'CEERs' in the abstract should be 'CEERS'.","section":"Abstract and §6"},{"comment":"The caption states that the color scale for the size panels is limited to 0.1–0.3 arcsec 'to emphasize the correlation,' but this is not the actual range of the plotted sizes and is confusing; please rephrase or use a scale that reflects the data.","section":"Figure 3 caption"},{"comment":"The text says that for single Sérsic fits 'the index value is limited to the range of 1–4,' while for double Sérsic fits the indices are fixed to 4 and 1; please clarify whether the single-Sérsic index was restricted to integer values or to the continuous range.","section":"§3.2"},{"comment":"There is a typo in 'the single Sérsic and double Sérsic models models'; also the phrase 'double Sérsic models' is used repeatedly and should be made singular where appropriate.","section":"§5.1"},{"comment":"The SNR formula uses integral-like symbols that do not render clearly; please define the integration limits and the noise model explicitly.","section":"§5.2, Eq. (3)"},{"comment":"The reference to Kelly et al. (2023) is incomplete; it should include the journal volume and article number.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the paper is a potentially useful simulation reference, but the headline accuracy claims should not be published in their current form. The main fixes are to replace the R_weighted comparison with a true composite half-light radius for single-Sérsic sizes, and to report quantile-based error statistics instead of mean±scatter claims. The closed-box nature of the test should also be stated in the abstract. I see no grounds for rejection; the issues are addressable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. The paper is a genuinely useful calibration study: 5,000 mock bulge+disk galaxies embedded in CEERS/NIRCam images, empirical PSFs and sky noise, single- and double-Sérsic fits, BIC-based model selection, SNR error budgets, and an independent galfit cross-check. That is a solid contribution for anyone doing JWST morphology. The second thing is that the headline single-Sérsic size accuracy is not established as stated. The comparison radius in Eq. 2 is a flux-weighted average of the input bulge/disk effective radii, not the half-light radius of the two-component total light distribution. For a de Vaucouleurs bulge plus exponential disk, those can differ by roughly the quoted 0.2 dex, so 'recovers size within 0.2 dex down to 27 mag' is not measuring what it claims. The stress-test note holds up.\n\nWhat the paper does well: the double-Sérsic component recovery tests compare each fitted component radius to its own input radius, and those appear internally consistent. The PSF-mismatch strategy, use of real CEERS blank sky, and agreement between galight and galfit are real strengths. The SNR>10 threshold and the BIC magnitude thresholds are practical, usable outputs. The n-B/T relation is not surprising but is quantified over a controlled grid, which is worth having.\n\nSoft spots, in proportion. The biggest is the closed-box nature: mocks are generated from the same idealized Sérsic family with n_bulge=4 and n_disk=1, and the fits use those fixed values. The authors acknowledge this in §5.4, but it means the quoted error budgets and thresholds will not simply transfer to clumpy, barred, or pseudo-bulge galaxies. The footnote saying default and truth-initialized fits agree softens the initialization concern; it does not fix the same-model circularity. There is also an internal inconsistency: the abstract and Section 4.1.2 support roughly 0.5 mag total magnitude accuracy at the faint end, while the summary claims 0.2 mag down to 27 mag; the reported scatter for mag>26 is ±0.2–0.3 mag, so 0.2 is not supported. Data and code are only 'available on reasonable request'; for a calibration paper other groups will want the mocks.\n\nBottom line: this is a reference-grade simulation paper for JWST/NIRCam bulge+disk work, not a physics discovery. With the single-Sérsic size comparison corrected or reframed and the magnitude claims reconciled, it deserves to be published and will be cited. I would send it to a serious referee; the author's own discussion shows they know the main external-validity limit, which is the part that needs more honesty in the abstract.","headline":"Useful JWST/NIRCam bulge+disk error-budget paper, but the single-Sérsic size claim rests on a flux-weighted radius comparison that does not measure half-light radius; the double-component decomposition results are the solid part.","tokens_in":18359,"tokens_out":3433,"would_cite":true,"duration_ms":33623,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Under CEERS-like JWST/NIRCam conditions, single and double Sérsic fits recover galaxy magnitudes within about 0.5 mag and sizes within about 0.2 dex, down to 27 mag for total light and 26 mag for individual bulge/disk components.","keywords":["galaxy structure","bulge+disk decomposition","Sérsic profile","JWST","NIRCam","high-redshift galaxies","Bayesian Information Criterion","signal-to-noise ratio"],"falsifier":"Take a sample of galaxies with independently known bulge and disk parameters, for example from kinematic decomposition or higher-resolution imaging, degrade them to CEERS-like JWST depth and PSF, and check whether the double Sérsic fits recover those known values within 0.5 mag and 0.2 dex; a simpler version is to rerun the same mock pipeline with pseudo-bulges or barred light and watch whether the recovery error and the BIC separation worsen substantially.","tokens_in":17281,"feed_emoji":"🔭","tokens_out":10583,"duration_ms":91827,"temperature":0.7,"pith_summary":"This paper asks whether JWST/NIRCam images can be trusted to separate a galaxy's light into a bulge and a disk. The authors generate 5,000 two-component galaxies with realistic CEERS noise and point-spread functions, then fit them with a single Sérsic model and a double Sérsic model. They find that the single model recovers total magnitudes within about 0.5 mag and effective radii within 0.2 dex down to 27 mag, and the double model recovers each component within 0.5 mag and 0.2 dex for components brighter than 26 mag. They also set a signal-to-noise threshold of 10 and give BIC-based brightness limits for when the two-component fit is preferred, providing a practical error budget for high-redshift morphology studies.","feed_headline":"Bulge-disk fits hold to 0.5 mag in JWST-like images","feed_subtitle":"A 5,000-galaxy simulation maps when single- and double-Sérsic fits are trustworthy in NIRCam data.","key_machinery":"The central object is the Sérsic profile, $I(R)=I_e\\exp[-k((R/R_{\\rm eff})^{1/n}-1)]$, with $n=4$ for the bulge and $n=1$ for the disk, convolved with real point-spread functions chosen from CEERS F150W and F356W images and embedded in real CEERS sky background with Poisson noise. The argument is carried by a controlled comparison: 5,000 simulated bulge+disk galaxies of known parameters are fit with single and double Sérsic models by a two-dimensional image-fitting pipeline, and recovery accuracy is measured against the known inputs. Two auxiliary tools do key work: the Bayesian Information Criterion decides when the extra component is justified, and the total signal-to-noise ratio is used to make the error budgets portable to other NIRCam surveys.","core_discovery":"Under CEERS-like NIRCam observing conditions, single Sérsic fits recover the total magnitude of a bulge-plus-disk galaxy within 0.5 mag and its effective radius within 0.2 dex for galaxies as faint as 27 mag, while double Sérsic fits recover the bulge and disk magnitudes within 0.5 mag and effective radii within 0.2 dex for components brighter than 26 mag. The recovered parameters are unbiased and the scatter shrinks as signal to noise increases, with SNR>10 quoted as the threshold for reliable recovery. The fitted single-Sérsic index rises monotonically with the true bulge-to-total flux ratio, so n can serve as a B/T proxy, though the relation scatters heavily near n=1 and n=4 and at magnitudes fainter than 26. For model selection, the Bayesian Information Criterion prefers the double model in 90% of systems brighter than 24.9 mag in F150W or 26.2 mag in F356W, and in 90% of comparable-flux systems with B/T between 30% and 70%.","pith_inferences":["These error budgets are likely optimistic for real high-redshift galaxies because the mocks use exactly the same Sérsic family as the fits, with a classical n=4 bulge and exponential n=1 disk and no bars, clumps, or asymmetric features; adding such substructure would probably degrade the recovery and weaken the BIC separation.","Because the nominal fits start from the true parameter values, with a footnote reporting that default initialization gave consistent results, a fully blind search over a larger survey would be a useful additional stress test of the quoted accuracies.","A direct extension would be to repeat the calibration with bulge and disk Sérsic indices left free, or with pseudo-bulges at n~2, to map how much of the 0.5 mag budget is absorbed by index flexibility.","The BIC magnitude limits translate into a redshift-dependent mass limit: at z~6 only the most luminous galaxies will be bright enough in F356W to support double-component decomposition, so B/T measurements there will be biased toward massive systems."],"forward_implications":["Surveys can treat single-Sérsic total magnitudes and sizes as reliable down to 27 mag in F150W and F356W, which extends size evolution studies to the faint high-redshift populations JWST actually detects.","Bulge and disk magnitudes can be compared at the 0.5 mag level for components brighter than 26 mag, making B/T-based tests of bulge growth versus disk accretion feasible at high redshift.","The BIC thresholds, with 90% confidence for mag<24.9 in F150W and mag<26.2 in F356W, give a simple pre-analysis rule for deciding when to attempt two-component decomposition.","The SNR>10 criterion lets other NIRCam programs compute, before fitting, which of their galaxies will yield reliable structural parameters.","Agreement between two independent fitting codes within 0.1 mag/0.1 dex for single models and 0.5 mag/0.2 dex for double models suggests these error budgets are not quirks of one implementation."],"supporting_citations":[{"why":"Defines the Sérsic profile that is both the simulated light model and the fitting function.","marker":"Sérsic 1963"},{"why":"Supplies the CEERS NIRCam images, PSFs, background noise, and depth used to embed the mock galaxies.","marker":"Finkelstein et al. 2022"},{"why":"Supplies the two-dimensional image-fitting pipeline used for all single and double Sérsic fits.","marker":"Ding et al. 2020"},{"why":"Supplies the independent fitting code used to cross-check the recovered parameters.","marker":"Peng et al. 2002"},{"why":"Underlies the image-generation engine used to create the noiseless mock galaxies.","marker":"Birrer & Amara 2018"},{"why":"Supplies the particle swarm optimization algorithm that drives the parameter fitting.","marker":"Kennedy & Eberhart 1995"},{"why":"Justifies the compact bulge sizes and high-redshift galaxy size distributions adopted in the mock parameter ranges.","marker":"van der Wel et al. 2014"},{"why":"Documents the roughly 40% clumpy fraction at z~1-3 that the paper cites as the main limitation of its smooth bulge+disk assumption.","marker":"Kalita et al. 2025"},{"why":"Shows that bars bias Sérsic decomposition, motivating the paper's acknowledged simplification to smooth two-component galaxies.","marker":"Bi et al. 2022a"}],"fun_headline_variants":["Double Sersic fits win for bulge+disk galaxies in JWST","JWST fits: 0.5 mag accuracy for bulges and disks down to 26 mag","Sersic index traces bulge fraction in simulated NIRCam data","Model selection favors double components in 90% of bright JWST galaxies","SNR>10 ensures reliable galaxy structure fits in JWST surveys"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the mock galaxies are generated from the same Sérsic model family used in the fits, with the bulge index fixed at 4 and the disk index at 1, and that the fits are initialized with the true parameter values; if real high-redshift galaxies contain pseudo-bulges, bars, clumps, or Sérsic indices outside those fixed values, the quoted recovery accuracies and BIC thresholds will not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Double Sersic fits win for bulge+disk galaxies in JWST","JWST fits: 0.5 mag accuracy for bulges and disks down to 26 mag","Sersic index traces bulge fraction in simulated NIRCam data","Model selection favors double components in 90% of bright JWST galaxies","SNR>10 ensures reliable galaxy structure fits in JWST surveys"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001368,"raw_usage":{"total_tokens":5629,"prompt_tokens":1109,"completion_tokens":4520,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":725,"completion_tokens_details":{"reasoning_tokens":4420}},"tokens_in":725,"tokens_out":4520,"duration_ms":29175,"temperature":1.0,"reasoning_tokens":4420,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:50:27.383266+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of galaxies with independently known bulge and disk parameters, for example from kinematic decomposition or higher-resolution imaging, degrade them to CEERS-like JWST depth and PSF, and check whether the double Sérsic fits recover those known values within 0.5 mag and 0.2 dex; a simpler version is to rerun the same mock pipeline with pseudo-bulges or barred light and watch whether the recovery error and the BIC separation worsen substantially.","supporting_citations":[{"cited_title":"L., Bagley M., Ferguson H","cited_arxiv_id":null,"evidence_quote":"Supplies the CEERS NIRCam images, PSFs, background noise, and depth used to embed the mock galaxies."},{"cited_title":"pp 1942--1945","cited_arxiv_id":null,"evidence_quote":"Supplies the particle swarm optimization algorithm that drives the parameter fitting."}],"review_version":1}