{"id":"0c61fdac-8063-4eca-872f-0055974d0123","arxiv_id":"2501.10919","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"New WISE W1 color-to-mass-to-light models reproduce Spitzer, S4G and Dustpedia stellar masses, but conflict with the Jarrett et al. GAMA-G23-based calibration by about 0.2 dex.","lead":"Astronomers built new color-based recipes to estimate galaxy stellar masses from WISE infrared brightness, and checked them against Spitzer and other surveys. The recipes mostly agree with earlier methods, but disagree with one rival WISE calibration by about 0.2 dex.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The W1-vs-Spitzer mass agreement cannot validate the empirical AGB prescription, because both filters use the same SML models; the AGB calibration is anchored at K and only propagated to W1, leaving the 0.2 dex Jarrett offset as an unresolved test of the central claim.","rationale":"The paper is careful and mostly honest about the AGB problem, and the SML program has a track record. The new element is the W1 application, not the population models. My stress-test focuses on the weakest link in the chain from cluster colors to W1 masses. The internal W1-versus-Spitzer comparison (Fig. 4) is a photometric and color-relation check, not a stellar-population check, because both channels use the same Upsilon*(color) relations and the same AGB correction. The Jarrett et al. (2023) offset is the only large-sample independent tension; the paper's rebuttal uses gas-fraction plausibility, which is an indirect argument and depends on the same stellar mass scale. The cluster-color test I propose would directly test whether the AGB prescription is valid at W1, and would separate a Upsilon* calibration error from a W1 photometry error. This does not overturn the paper's conclusion that W1 can substitute for Spitzer for many purposes, but it does mean the conditional acceptance should remain, with the AGB-at-W1 calibration identified as the condition to be checked.","tokens_in":15095,"tokens_out":5156,"duration_ms":58121,"concrete_test":"Use the same 116 star clusters with known ages and metallicities: measure their W1 and [3.6] fluxes from AllWISE/unWISE and Spitzer, and compare the model-predicted g-W1 and W1-[3.6] colors (from the empirical AGB tracks) to the observed cluster colors as a function of age and metallicity, especially in the 0.1-1 Gyr range where AGB dominates. If the W1 colors are reproduced within ~0.05 mag, the AGB calibration is anchored at W1; if they deviate systematically, the W1 Upsilon* models inherit an uncalibrated AGB error and the 0.2 dex Jarrett offset is not resolved. An additional quantitative check: recompute the Fig. 4 comparison using only galaxies with independent SED-based masses from Dustpedia and compare residuals to the Jarrett offset to see whether the offset is stellar-population or photometric.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central product is a color-Upsilon* relation for WISE W1. The underlying stellar-population models are calibrated by an empirical AGB prescription derived from 116 Magellanic Cloud and Milky Way star clusters (Section 2). Two features of the argument make this calibration load-bearing. First, the prescription is built at K band: the clusters provide V-K colors, and the K-[3.6] and then W1 relations are applied through color-color relations; there is no direct cluster test of W1 or g-W1 colors. Second, the headline validation in Fig. 4 compares W1 masses to Spitzer 3.6um masses produced with the same SML models and the same AGB correction, so any systematic error in Upsilon* largely cancels. The genuinely independent checks (S4G with Eskew et al. 2012; Dustpedia SED fits; Cluver et al. 2014) are small samples or share near-IR calibration assumptions. The largest independent comparison, Jarrett et al. (2023), shows a ~0.2 dex offset that the paper rejects on gas-fraction plausibility (Fig. 3) rather than with a direct stellar-population test. If the W1 AGB treatment is off by even half that amount, the claimed 0.1 dex accuracy for low-mass blue galaxies fails, and the TF conclusions in Section 6 shift.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents new stellar mass-to-light ratio (Upsilon*) models for converting WISE W1 fluxes into stellar masses, extending the authors' earlier SML models calibrated on Magellanic Cloud and Milky Way star clusters. The models are built from assumed star formation histories and chemical evolution scenarios, with an empirical AGB correction. The authors compare masses derived from W1 photometry with those from Spitzer 3.6um, S4G, Dustpedia, Cluver et al. (2014), and Jarrett et al. (2023), and also compare W1 and 3.6um surface brightness and mass density profiles. They find good agreement with most comparisons and present WISE luminosity and stellar mass Tully-Fisher relations for the SPARC sample.","tokens_in":15459,"tokens_out":3643,"duration_ms":40658,"significance":"If the models are reliable, they provide a practical path to stellar masses from all-sky WISE data, which is valuable now that Spitzer is no longer operational. The paper makes useful comparisons to S4G and Dustpedia that are partially independent of the authors' population synthesis, and it provides tabulated masses in machine-readable form. However, the central validation is weakened by the fact that the W1-to-Spitzer mass comparison shares the same stellar population models, and by an unresolved ~0.2 dex offset from Jarrett et al. (2023). The claimed accuracy of 0.1 dex for blue, low-mass galaxies therefore rests on assumptions that are not fully tested.","major_comments":[{"comment":"The W1-to-Spitzer mass comparison is not an independent validation of the Upsilon* scale because both masses are derived from the same SML stellar population models and the same empirical AGB prescription. Any systematic error in Upsilon* largely cancels in this comparison. The agreement in Fig. 4 confirms photometric consistency and the color-color relations, but it cannot validate the absolute stellar mass scale. The statement that 'the scatter is completely explained by photometric errors' is too strong. Please either rephrase the claim or provide a test that uses a fully independent Upsilon* determination, such as dynamical masses or SED fits that do not share the same AGB treatment.","section":"Section 4, Fig. 4"},{"comment":"The empirical AGB prescription is anchored using V-K colors of 116 LMC, SMC, and Milky Way star clusters, and then propagated to W1 through K-[3.6] and [3.6]-W1 color-color relations. There is no direct test of the models against W1 or g-W1 colors of star clusters or of resolved stellar populations. This is a critical extrapolation because the paper's central product is the g-W1 color-Upsilon* relation. Please quantify the systematic uncertainty introduced by this propagation chain, and if possible test the final W1 Upsilon* models on clusters with W1 photometry. If no such test is possible, state clearly that the W1 AGB behavior is an assumption.","section":"Section 2"},{"comment":"The disagreement with Jarrett et al. (2023) of roughly 0.2 dex is dismissed primarily through an argument about gas-fraction plausibility (Fig. 3). This is an indirect test: the low-Upsilon* values may indeed imply high gas fractions, but this is not a direct measurement of stellar population properties. Please provide a direct comparison with independent stellar mass estimates for the same galaxies, such as high-quality optical SED fits or dynamical mass constraints where the gas fraction is known. The unresolved 0.2 dex offset is significant relative to the claimed 0.1 dex accuracy and needs to be addressed quantitatively before the models can be adopted with confidence.","section":"Section 3, Fig. 3"},{"comment":"The comparison with Cluver et al. (2014) yields an RMS of 0.29 dex in log stellar mass for 58 galaxies. This scatter is much larger than the ~0.1 dex uncertainty quoted in the abstract for the Upsilon* models. The paper does not discuss whether this scatter is consistent with the quoted photometric and model errors, or whether it indicates an additional systematic floor in the W1-W2 method. Please quantify the expected scatter from the stated errors and compare it with the observed RMS.","section":"Section 3, Fig. 2"},{"comment":"The claim that the stellar mass TF 'improves the linearity of the entire correlation' is not supported by any quantitative fit statistics in the text. Please report the fitted slope, intercept, scatter, and the number of galaxies for both the luminosity TF and the stellar mass TF, so the reader can judge whether the improvement is significant. As written, the comparison with Ristea et al. (2023) is only qualitative.","section":"Section 6, Fig. 7"}],"minor_comments":[{"comment":"The phrase 'color-$\\Upsilon_*$ (mass-to-light) models' is redundant; 'mass-to-light ratio' is clearer. Please use consistent terminology throughout.","section":"Abstract"},{"comment":"The sentence 'Our stellar population models is based on...' should be 'Our stellar population models are based on...'.","section":"Section 2, paragraph 3"},{"comment":"The sentence 'This may signal a systematic difference in our W1 fluxes (unlikely) or a systematic underestimate of near-IR $\\Upsilon_*$ values from optically determined SED fits' contains an unsupported parenthetical 'unlikely'. Please present both possibilities neutrally.","section":"Section 3, paragraph 4"},{"comment":"The phrase 'A optimal course' should be 'An optimal course'.","section":"Section 5, paragraph 2"},{"comment":"The sample size for the W1-vs-Spitzer comparison is not clearly stated. Section 5 says 111 galaxies have good W1 and IRAC images, and 81 have SDSS imaging, but the number used in Fig. 4 is not given. Please specify the sample size in the text.","section":"Section 4"},{"comment":"The reference 'Duey et al. 2024' is mentioned as Paper I but does not appear in the reference list. Please add the full citation.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper is part of a well-known series from this group, and the models are likely to be used by the community. The main concern is that the validation of the W1 mass scale is partly circular, and the unresolved Jarrett et al. offset is a load-bearing issue for the claimed accuracy. The gas-fraction argument is suggestive but not a substitute for a direct stellar population test. If the authors can provide an independent check or at least a thorough quantitative discussion of the systematic uncertainty, the paper would be suitable for publication. I do not see evidence of misconduct or citation problems, but the reference list should be checked for completeness."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a practical calibration paper, not a new population-synthesis result. The authors are explicit that the W1 models are a simple interpolation of their SML models, so the novelty is in the package: W1 color-Υ* relations in three scenarios, total stellar masses for the SPARC sample, mass density profiles, and a WISE Tully-Fisher relation. That fills a real niche now that Spitzer is gone, and the comparisons to S4G, Dustpedia and Cluver are the right checks.\n\nThe 0.02 ± 0.18 dex offset against their own Spitzer masses is fine as a photometric cross-check, but it does not validate the stellar-population models, since both filters go through the same SML machinery and the same AGB correction. The stress-test note has this right. The unresolved 0.2 dex disagreement with Jarrett et al. is where the underlying AGB calibration is actually tested, and the paper's response—gas-fraction plausibility using Illustris—is suggestive but not a direct stellar-population test. I also think the claim that the scatter is completely explained by photometric errors overreaches, given there is no quantitative error budget in the paper.\n\nOn the positive side, the authors are transparent about the model lineage, the data products are real, and the agreement with independent estimates (Eskew et al., Dustpedia SED fits) is encouraging. The TF section is minor and mostly consistent with earlier work.\n\nFor the right reader—someone needing WISE-based masses or checking TF systematics—this is a usable, honest calibration reference. It deserves a serious referee; the referee's main job is to make the error budget explicit and to force a direct engagement with Jarrett rather than plausibility arguments.","headline":"A practical WISE W1 stellar-mass calibration package, honest about being an interpolation of the authors' own SML models, but with an unresolved 0.2 dex tension against Jarrett et al. that deserves referee attention.","tokens_in":15954,"tokens_out":2486,"would_cite":true,"duration_ms":29192,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"New color-based mass-to-light models allow WISE W1 fluxes to yield galaxy stellar masses consistent with Spitzer-based estimates.","keywords":["mass-to-light ratio","WISE W1 photometry","stellar mass","Tully-Fisher relation","stellar population models","AGB stars","galaxy scaling relations"],"falsifier":"A resolved-star census of blue dwarf galaxies that yields stellar masses systematically lower than the $g-W1$ color-$\\Upsilon_*$ prediction by more than about 0.2 dex would falsify the empirical AGB calibration, since the paper identifies blue, star-forming galaxies as the regime with the largest model uncertainty.","tokens_in":14902,"feed_emoji":"🌌","tokens_out":10847,"duration_ms":97660,"temperature":0.7,"pith_summary":"This paper tries to establish that WISE W1 infrared fluxes, converted through new color-based mass-to-light models, give galaxy stellar masses as reliable as those from Spitzer 3.6 micron photometry. The models are calibrated on star cluster colors and applied to the SPARC galaxy sample, and the resulting masses agree with independent literature estimates to a mean offset of $0.02 \\pm 0.18$ dex. This matters because WISE covers the whole sky, so if the conversion is sound, stellar masses for large galaxy samples no longer depend on pointed Spitzer observations. The paper also shows that the same masses improve the high-mass end of the stellar-mass Tully-Fisher relation.","feed_headline":"WISE starlight now yields galaxy masses matching Spitzer's to 0.02 dex","feed_subtitle":"A color-based recipe lets the all-sky WISE archive stand in for Spitzer stellar-mass work.","key_machinery":"The central object is the color-$\\Upsilon_*$ relation: a mapping from $g-W1$ color to the stellar mass-to-light ratio in the WISE W1 band. It is generated from composite stellar population models that combine specified star formation histories and chemical enrichment scenarios, with the AGB contribution calibrated empirically on star cluster colors. This relation is what converts observed flux into stellar mass, so every comparison in the paper, whether against Spitzer, independent samples, or the Tully-Fisher relation, depends on it.","core_discovery":"The central claim is that a simple color, $g-W1$, plus a choice of bulge/disk decomposition, determines the WISE W1 mass-to-light ratio well enough to recover stellar masses consistent with Spitzer 3.6 micron masses and with independent literature stellar-mass estimates. The new $\\Upsilon_*$ models replace uncertain theoretical AGB treatment with an empirical calibration on 116 LMC, SMC, and Milky Way star clusters, and they are built for three star formation history scenarios: pure disk, pure bulge, and a bulge+disk hybrid. Using these models on SPARC galaxies gives a mean offset of $0.02 \\pm 0.18$ dex against Spitzer-based masses and an RMS scatter of 0.29 dex against an independent W1-W2 color relation. The authors convert WISE luminosity profiles into stellar mass density profiles that track Spitzer profiles, and they construct WISE luminosity and stellar-mass Tully-Fisher relations in which the stellar-mass version is more linear at high masses.","pith_inferences":["If the AGB calibration holds up, near-IR stellar masses for dwarf galaxies may need to be revised upward relative to color-blind WISE luminosity relations, which would remove the tension between low $\\Upsilon_*$ values and the stability of spiral density waves.","The same color-$\\Upsilon_*$ machinery could be tested at high redshift by applying it to rest-frame near-IR JWST photometry, although the paper only gestures at this link.","A stronger test than the paper offers would compare W1-based stellar masses against dynamical masses in galaxies where the dark matter fraction is measured independently, since that would bypass stellar population assumptions entirely.","The 0.2 dex offset with the alternative total-flux relation could indicate that optically calibrated SED-fitting masses inherit too-low near-IR $\\Upsilon_*$ values; the paper supports this indirectly through gas fraction plausibility rather than by a direct stellar population test."],"forward_implications":["The all-sky WISE archive can replace pointed Spitzer 3.6 micron observations for stellar and baryonic mass studies, including total masses and radial mass profiles.","Stellar mass Tully-Fisher relations built from WISE masses are more linear at the high-mass end because bulge-dominated galaxies receive higher mass-to-light ratios.","Color-blind W1 luminosity-to-mass relations that use a constant $\\Upsilon_*$ will systematically underpredict stellar masses in gas-rich, low-luminosity galaxies, inflating their gas fractions.","Optical-to-near-IR colors such as $g-W1$ are sufficient to determine $\\Upsilon_*$; full SED fitting did not improve agreement in the tested comparison samples."],"supporting_citations":[{"why":"Supplies the base stellar population models, star formation history prescriptions, and the empirical AGB calibration that the W1 models extend.","marker":"SML"},{"why":"Establishes the near-IR color-to-mass-to-light technique and the motivation for empirically treating AGB stars.","marker":"Schombert & McGaugh (2014)"},{"why":"Provides the main-sequence relation used to set the disk star formation history and its final strength.","marker":"Speagle et al. (2014)"},{"why":"Offers the alternative W1-W2 color-based stellar mass relation against which the new models are compared.","marker":"Cluver et al. (2014)"},{"why":"Supplies the total-flux W1 mass relation whose 0.2 dex offset with the new models is the main discrepancy discussed.","marker":"Jarrett et al. (2023)"},{"why":"Supplies the S4G stellar masses used as an independent comparison sample.","marker":"Muñoz-Mateos et al. (2015)"},{"why":"Provides the 3.6 micron mass-to-light conversion that the S4G comparison sample applied.","marker":"Eskew et al. (2012)"},{"why":"Supplies the full SED-fitting stellar masses used as another independent comparison.","marker":"Clark et al. (2018)"}],"fun_headline_variants":["Empirical WISE W1 mass recipe matches Spitzer's to 0.02 dex","Star-cluster calibration puts Spitzer-quality masses on all-sky WISE","Bulge/disk separation is the new limiting factor for WISE masses","WISE W1 profiles now deliver Spitzer-like stellar masses for Tully-Fisher","All-sky WISE now rivals Spitzer for stellar-mass extraction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method stands or falls on the assumption that the AGB prescription calibrated on 116 Milky Way, LMC, and SMC star clusters applies to every galaxy type in the SPARC sample.","fun_headline_variants_meta":{"raw":{"variants":["Empirical WISE W1 mass recipe matches Spitzer's to 0.02 dex","Star-cluster calibration puts Spitzer-quality masses on all-sky WISE","Bulge/disk separation is the new limiting factor for WISE masses","WISE W1 profiles now deliver Spitzer-like stellar masses for Tully-Fisher","All-sky WISE now rivals Spitzer for stellar-mass extraction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001283,"raw_usage":{"total_tokens":5236,"prompt_tokens":930,"completion_tokens":4306,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":4203}},"tokens_in":546,"tokens_out":4306,"duration_ms":33311,"temperature":1.0,"reasoning_tokens":4203,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:49:56.343412+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A resolved-star census of blue dwarf galaxies that yields stellar masses systematically lower than the $g-W1$ color-$\\Upsilon_*$ prediction by more than about 0.2 dex would falsify the empirical AGB calibration, since the paper identifies blue, star-forming galaxies as the regime with the largest model uncertainty.","supporting_citations":[{"cited_title":"S., Steinhardt, C","cited_arxiv_id":null,"evidence_quote":"Provides the main-sequence relation used to set the disk star formation history and its final strength."},{"cited_title":"A New WISE Calibration of Stellar Mass","cited_arxiv_id":"2301.05952","evidence_quote":"Supplies the total-flux W1 mass relation whose 0.2 dex offset with the new models is the main discrepancy discussed."},{"cited_title":"C., Sheth, K., Regan, M., et al","cited_arxiv_id":null,"evidence_quote":"Supplies the S4G stellar masses used as an independent comparison sample."},{"cited_title":"2012, AJ, 143,","cited_arxiv_id":null,"evidence_quote":"Provides the 3.6 micron mass-to-light conversion that the S4G comparison sample applied."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the full SED-fitting stellar masses used as another independent comparison."}],"review_version":1}