{"id":"dab32917-0561-4800-a8fc-b42619aae064","arxiv_id":"2506.07845","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"DR19 provides ASPCAP stellar parameters and abundances for 964,989 stars, with validated precision of 50-100 K in Teff, 0.07-0.09 dex in log g, and 0.02-0.04 dex for several elements.","lead":"The SDSS-V Milky Way Mapper team reports stellar temperatures, surface gravities, metallicities, and 21 element abundances for 964,989 stars in Data Release 19, and compares them against independent reference catalogs. A generalist reader should care because this is one of the largest uniform spectroscopic catalogs of the Milky Way, and the paper calibrates how reliable those measurements are for Galactic archaeology.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abundance 'accuracy' metric is circular: the SNSM zero-point sample both defines the zero point and is used as the accuracy/quality reference, so the Table 5 'excellent' accuracy claims are guaranteed rather than tested.","rationale":"The reader's weakest_assumption is precisely the SNSM zero-point assumption, and my concern is the same one, sharpened and traced through to Table 5. The paper's strongest claims—precision estimates and catalog size—are separately supported by open clusters, wide binaries, and external comparisons, so the central precision claims are not the weakest point. The accuracy claim is the weak spot because the SNSM serves double duty as both calibrator and accuracy reference. I agree with the reader that the paper is honest and the catalog is useful; the issue is not fraud or sloppiness but a methodological circularity that makes the 'excellent' accuracy category for 10 elements a construct rather than a measurement. The concrete test—an external, unused reference sample—would settle whether the SNSM choice is benign. If the external zero points agree with Table 2, the concern evaporates and the paper could move to ACCEPT; if they disagree, the accuracy claims need qualification. Either way, the CONDITIONAL verdict is appropriate pending this check.","tokens_in":61205,"tokens_out":2170,"duration_ms":24202,"concrete_test":"Recompute the Table 2 zero-point offsets using an independent solar-metallicity reference that was not used in the calibration—for example, GALAH DR4 stars with |[Fe/H]| < 0.05 and Teff, log g in the Table 1 range, cross-matched to MWM. If the inferred zero points for Al, Mn, Cu, V, or Ce shift by more than ~0.1 dex relative to Table 2, then the SNSM-based accuracy categories in Table 5 overstate the accuracy of those elements and the calibration is not robust to reference-sample choice.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Sections 4.1 and 4.4 define abundance accuracy through zero-point offsets computed by forcing the Solar Neighborhood Sample (SNSM; Table 1, Section 4.1) to mean [X/H] = 0, then assign quality categories (Table 5) using those same offsets with thresholds (<0.05 dex for 'excellent'). For C and N the offsets are not applied, and their quality is judged by scatter; but for all other elements, an element with a large raw offset is rescued by construction because the calibration removes exactly that offset. The Table 5 statement that 10 elements have 'excellent' accuracy with precision 0.02–0.1 dex is therefore not an independent validation: it is a restatement of the calibration assumption that the solar-metallicity, distance-limited SNSM stars have solar [X/H] for every element. If the local thin disk is genuinely non-solar in Al, Cu, Mn, or other elements (as several abundance studies of the solar neighborhood suggest), every calibrated abundance inherits that offset, and the 'accuracy' column of Table 5 is shifted for all stars. The paper's external comparisons (GALAH, Gaia-ESO, GBS) provide partial independent checks, but Table 5 is the headline accuracy claim and it rests on the SNSM assumption. This is a real soft spot, not a fatal flaw; the catalog remains useful, but the accuracy categories are not as secure as presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper describes the SDSS-V DR19 ASPCAP data products: atmospheric parameters and abundances for 964,989 stars observed with APOGEE, with 894,256 stars below 8000 K. It validates raw Teff against IRFM photometric scales and benchmark stars, log g against APOKASC/TESS asteroseismology and other surveys, and [M/H] against GALAH, Gaia-ESO, GBS, and cluster samples. For individual abundances, zero-point offsets are derived from the solar-neighborhood sample (Table 2), temperature-dependent corrections are derived from open clusters (Eq. 3), and precision estimates are obtained from the solar-neighborhood sample, open clusters, and wide binaries (Table 4). The paper concludes that 10 elements are of excellent quality, with Teff precision of 50-70 K for giants and 70-100 K for dwarfs, log g precision of 0.07-0.09 dex for giants, and abundance precision better than 0.1 dex for at least 10 elements.","tokens_in":61457,"tokens_out":7612,"duration_ms":86976,"significance":"If the accuracy and precision claims hold, this is an important public catalog for Galactic archaeology, and the paper's broad comparison campaign against independent references is a genuine strength. The paper is also unusually candid in flagging unreliable regions (M dwarfs below 4500 K, P, V, Cu, 12C/13C, and a problematic cool carbon-rich group) and in warning that reported formal uncertainties are underestimated. However, the headline accuracy assessment in Table 5 is weakened by a calibration-validation circularity: the same solar-neighborhood sample is used both to define the zero-point offsets and to grade the accuracy of those offsets. The independent external comparisons are the real accuracy tests and should be incorporated into the quality summary. The paper also contains internal inconsistencies between the quality categories in Section 4.4 and the conclusions, and the abstract precision claim is not fully supported by Table 4. These issues are fixable within the scope of the manuscript, so revision rather than rejection is appropriate.","major_comments":[{"comment":"The 'Accuracy' column of Table 5 is built from the zero-point offsets Δ in Table 2, but §4.1 defines those Δ by forcing the solar-neighborhood sample (SNSM) to have mean [X/H]=0. Consequently, the 'excellent' (<0.05 dex) accuracy flags for 10 elements are a restatement of the calibration assumption rather than an independent test: any element would appear accurate on the SNSM by construction, and if the local thin disk is non-solar in, e.g., Al or Cu, the entire calibrated scale inherits that offset. The paper does contain genuinely independent accuracy checks (GALAH DR4, Gaia-ESO DR5, GBS, open and globular clusters, and asteroseismic and IRFM comparisons for the atmospheric parameters), but these are not used to set the accuracy categories in Table 5. Please relabel the first criterion as a 'zero-point consistency with the assumed solar-neighborhood scale,' and either compute the Table 5 accuracy categories from the independent comparisons or present the independent offsets (e.g., MWM − GALAH and MWM − Gaia-ESO medians) alongside Table 5 so users can judge accuracy without relying on the circular metric.","section":"§4.1 and Table 5"},{"comment":"The quality summary is internally inconsistent. Section 4.4 and Table 5 classify Na, Ti, Co, Ce, and Nd as 'fair,' but conclusion item 4 states that these same elements are 'considered to have poor quality'; this contradicting sentence in the conclusions should be corrected. In addition, Table 5 lists C and N as having 'excellent' accuracy even though §4.1 explicitly excludes C and N from the zero-point analysis because they are not calibrated, and Table 6 marks their accuracy entries with '· · ·'; a non-applicable quantity should not be placed in the <0.05 dex accuracy bin. These issues bear directly on how users will select elements for their science, so they should be fixed before publication.","section":"§4.4, Table 5, and §6"},{"comment":"The abstract states that 'the precision of at least 10 elements is better than 0.1 dex,' and Table 5 gives a Precision rating of Excellent (<0.1 dex) for 10 elements. Table 4, however, shows that several of those elements have at least one independent scatter estimate above 0.1 dex: Nglobal 0.113, Nwindows 0.138, S 0.067–0.151, K 0.081–0.106, Ti 0.073–0.149, Cr 0.077–0.178, and Mn 0.035–0.077. Only α, Mg, Al, Si, Ca, and Ni have all three independent estimates at or below 0.1 dex, while Fe appears in Table 5 but has no row in Table 4. Please specify the precise statistic (e.g., the minimum, the mean, or a giant-only estimate) behind the '10 elements' claim, add the supporting data for Fe and [M/H], and adjust the abstract if the claim is not supported.","section":"Abstract and Tables 4–5"}],"minor_comments":[{"comment":"The text 'NWM DR19-APOGEE DR17 common sample' appears to be a typo for 'MWM DR19-APOGEE DR17.'","section":"§5.11"},{"comment":"The word 'aseisimic' in the abstract should be 'asteroseismic.'","section":"Abstract"},{"comment":"The caption contains the typo 'neighboorhod'; it should read 'neighborhood.'","section":"Figure 16 caption"},{"comment":"The optimal-region entries for Nglobal and Nwindows list 'nowhere' for dwarfs; if this is intentional, please state explicitly that no reliable dwarf region exists for nitrogen, and if it is not intentional, please correct the entries.","section":"Table 6"},{"comment":"The M dwarf caveat (Teff < 4500 K, log g > 4) is clearly documented in Sections 3.1.1–3.3.1 and Section 4.4, but the abstract does not mention it; a brief caveat in the abstract would help users who rely only on the summary.","section":"§3.1.1–§3.3.1"}],"recommendation":"major_revision","confidential_remarks":"This is an important survey data-release paper, and the central issues are fixable. The circularity in the Table 5 accuracy metric is real but not fatal, because the paper does include independent external comparisons; the problem is that those comparisons are not reflected in the headline quality table. The internal inconsistencies between Section 4.4, Table 5, and the conclusions, and the unsupported precision claim in the abstract, should be corrected before acceptance. The paper also depends heavily on Casey et al. (in preparation) for calibration details, which is typical for SDSS papers but limits standalone verification; the authors should make clear which quantities are defined in this paper and which are deferred to the companion paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid data release paper. The genuinely new thing is the catalog itself: 964,989 stars with ASPCAP parameters, 336,511 of them new APO observations, all prior APOGEE data reanalyzed under the Astra framework. That alone makes it a resource the community will use. The paper also does what a data release paper should: extensive comparisons against IRFM temperatures, asteroseismic gravities, GALAH, Gaia-ESO, GBS, open and globular clusters, and it says clearly which regions are unreliable (M dwarfs below 4500 K, P, V, Cu). The precision estimates, 50–70 K for giants, 0.07–0.09 dex in log g, and 0.02–0.04 dex for the best abundances, are broadly plausible and supported by the comparisons.\n\nThe main soft spot is the abundance accuracy metric. The paper calibrates zero-point offsets by forcing a solar-neighborhood sample (SNSM) to solar [X/H], then uses the magnitude of those offsets to assign the 'accuracy' column in Table 5. So that column isn't a test of the calibrated abundances; it's a measure of how far the raw values were from the SNSM assumption. That assumption—that local thin-disk stars at solar metallicity have exactly solar [X/H] for every element—is not established, and if it's wrong (which is plausible for Al, Cu, Mn), all calibrated abundances inherit a constant shift. The external comparisons partially hedge this, but they're limited for exactly the elements with large offsets. I'd like the paper to say more plainly that the accuracy categories are diagnostics of the calibration, not independent validation.\n\nAlso, there's a straight internal contradiction: the conclusions call Na, Ti, Co, Ce, and Nd 'poor', while Section 4.4 and Table 5 put them at 'fair'. Someone needs to fix that. And key pieces of the calibration (log g, Teff, E(B-V), the uncertainty derivation) are deferred to Casey et al., in prep. That's common for survey papers, but it makes DR19 less self-contained than I'd like; the accuracy claims can't be fully audited until that companion appears.\n\nNone of this is fatal. The catalog is real, the caveats are honest, and the precision numbers look defensible. This deserves a serious referee. The fixes are clarifications, not re-analysis: resolve the fair/poor contradiction, reframe the accuracy claims as calibration diagnostics, and note the dependence on the companion paper. Send it to review.","headline":"A solid, honestly caveated data release paper for the new SDSS-V APOGEE sample; the catalog is a real resource, but the abundance accuracy metric is partly circular and the quality categories have an internal contradiction.","tokens_in":62172,"tokens_out":4157,"would_cite":true,"duration_ms":47935,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The SDSS-V DR19 catalog delivers measured uncertainties for nearly a million stars, with ten elements accurate below 0.1 dex.","keywords":["ASPCAP","SDSS-V DR19","stellar atmospheric parameters","chemical abundances","zero-point calibration","precision assessment","Milky Way stellar populations","FGKM stars"],"falsifier":"Compare the DR19 calibrated abundances for the 42,376 solar-neighborhood calibration stars against an independent, non-LTE optical analysis of the same stars; if the mean residual for any element exceeds the claimed 0.02–0.04 dex precision (for example, a mean [Al/H] offset near the 0.17 dex raw correction), the zero-point assumption is falsified.","tokens_in":60959,"feed_emoji":"🌟","tokens_out":7016,"duration_ms":76576,"temperature":0.7,"pith_summary":"This paper is the science verification of a nearly one-million-star catalog of stellar parameters and chemical abundances released by the fifth phase of the Sloan Digital Sky Survey. It claims that the survey's pipeline measures effective temperatures to about 50–70 K for giants and 70–100 K for dwarfs, surface gravities to 0.07–0.09 dex for giants, and abundances to 0.02–0.04 dex for the best elements, with at least ten elements better than 0.1 dex. These claims matter because the catalog puts a uniform, all-sky chemical map of the Milky Way within reach, letting astronomers trace where and when elements were produced. The paper also flags real cracks: gravities sit 0.09–0.18 dex above asteroseismic values, cool M-dwarf parameters can be off by up to a dex, and the abundance zero points assume the local solar-neighborhood sample is exactly solar in every element.","feed_headline":"965,000 stars get a precision check","feed_subtitle":"Ten elements hit abundance errors under 0.1 dex, though cool M dwarfs and the solar-neighborhood zero point need care.","key_machinery":"The load-bearing machinery is ASPCAP, the survey's spectral-fitting pipeline, which pseudo-continuum-normalizes H-band spectra and uses the FERRE interpolator to chi-square fit a grid of MARCS-model synthetic spectra, first for eight global parameters (Teff, log g, [M/H], microturbulence, macroturbulence or v sin i, [α/M], [C/M], [N/M]) and then for individual element abundances in element-specific wavelength windows. Accuracy is anchored by zero-point offsets computed from a solar-neighborhood, solar-metallicity sample; systematics are mapped with open-cluster stars; and precision is cross-checked with the same solar-neighborhood sample, open clusters, and wide binaries.","core_discovery":"On its own terms, the paper establishes that the DR19 ASPCAP products are science-ready for FGKM stars: effective temperatures agree with the infrared-flux-method scale for dwarfs and sit within −63 to −80 K of it for giants; surface gravities are precise to 0.07–0.09 dex for red giants yet systematically offset from asteroseismic values by 0.09–0.18 dex; and the calibrated abundances reach 0.02–0.04 dex precision for [M/H], [α/M], [Mg/H], and [Si/H], with ten elements rated excellent quality overall. It provides zero-point offsets for 18 elements, temperature-dependent correction coefficients for giant stars derived from open clusters, and element-by-element quality tables that tell users where each abundance can be trusted.","pith_inferences":["If the solar-neighborhood sample is not exactly solar in elements like Al or Cu, the published 'accuracy' is really precision-plus-zero-point; a testable extension is to compare the DR19 zero-point offsets against NLTE-corrected optical abundances of the same stars.","The open-cluster temperature-correction coefficients could be incorporated directly into the next data release, removing the need for users to apply them externally.","The cool carbon-rich group that ASPCAP misfits suggests a specific grid deficiency, likely molecular line opacities or missing carbon-enhanced model atmospheres, worth targeting in future synthetic grids.","The M-dwarf gravity failure implies that H-band spectra alone cannot anchor log g below 4500 K; combining ASPCAP with Gaia parallaxes and radii, or with isochrone priors, is a natural next step."],"forward_implications":["The paper's quality ratings give users a direct recipe: [M/H], [α/M], C, N, O, Mg, Si, Ca, Fe, and Ni can be trusted at the 0.02–0.1 dex level across most of the surveyed parameter space.","Because the internally reported uncertainties (median 0.001–0.008 dex) are several times smaller than external scatter estimates, science using DR19 abundances should adopt the paper's external precision values rather than the pipeline errors.","Giant-star abundances can be improved by applying the provided Teff-dependent corrections, which are not baked into the DR19 files.","For dwarfs cooler than 4500 K, the reported Teff, log g, and [M/H] can be badly off, so abundance work on M dwarfs should wait for the isochrone-calibrated gravity values or independent analyses.","The ten excellent-quality elements make the catalog suitable for Galactic chemical evolution and stellar population studies, while P, V, and Cu should be avoided."],"supporting_citations":[{"why":"Supplies the ASPCAP/FERRE chi-squared fitting method that derives stellar parameters from the observed spectra.","marker":"García Pérez et al. (2016)"},{"why":"Provides the synthetic spectral grids and element-window strategy that DR19 reuses for abundance measurements.","marker":"Jönsson et al. (2020)"},{"why":"The APOGEE DR17 catalog, the main internal comparison sample and the prior data release whose grids carry into DR19.","marker":"Abdurro'uf et al. (2022)"},{"why":"The APOKASC3 asteroseismic catalog is the reference for quantifying surface-gravity offsets and precision.","marker":"Pinsonneault et al. (2024)"},{"why":"Supplies the J−Ks and V−Ks infrared-flux-method temperature scales used to validate Teff.","marker":"González Hernández & Bonifacio (2009)"},{"why":"Supplies the BP−RP infrared-flux-method temperature scale used for an independent Teff comparison.","marker":"Casagrande et al. (2021)"},{"why":"The Gaia FGK benchmark stars provide fundamental temperatures from angular diameters, anchoring the accuracy assessment.","marker":"Soubiran et al. (2024)"},{"why":"GALAH DR4 is the cross-survey comparison sample for metallicity and individual element abundances.","marker":"Buder et al. (2024)"},{"why":"Dartmouth isochrones are used to expose the M-dwarf log g discrepancy below 4500 K.","marker":"Dotter et al. (2008)"},{"why":"Documents the MARCS model atmosphere grids and earlier calibration procedures that ASPCAP builds on.","marker":"Holtzman et al. (2018)"}],"fun_headline_variants":["965k stars get precision-tested abundances","Ten elements hit 0.1 dex precision in 965k stars","SDSS-V DR19 validates parameters for 965k stars","Milky Way Mapper: precise stellar parameters for 965k stars","965k stars: 21 elements, 10 at <0.1 dex precision"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The zero-point calibration assumes that a selected set of nearby solar-metallicity stars has exactly the Sun's composition for every calibrated element; if that assumption is wrong for any element, all published abundances of that element are shifted by the same amount.","fun_headline_variants_meta":{"raw":{"variants":["965k stars get precision-tested abundances","Ten elements hit 0.1 dex precision in 965k stars","SDSS-V DR19 validates parameters for 965k stars","Milky Way Mapper: precise stellar parameters for 965k stars","965k stars: 21 elements, 10 at <0.1 dex precision"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000613,"raw_usage":{"total_tokens":2901,"prompt_tokens":1044,"completion_tokens":1857,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":1766}},"tokens_in":660,"tokens_out":1857,"duration_ms":15231,"temperature":1.0,"reasoning_tokens":1766,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:23:54.942543+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the DR19 calibrated abundances for the 42,376 solar-neighborhood calibration stars against an independent, non-LTE optical analysis of the same stars; if the mean residual for any element exceeds the claimed 0.02–0.04 dex precision (for example, a mean [Al/H] offset near the 0.17 dex raw correction), the zero-point assumption is falsified.","supporting_citations":[{"cited_title":"H., Zinn, J","cited_arxiv_id":null,"evidence_quote":"The APOKASC3 asteroseismic catalog is the reference for quantifying surface-gravity offsets and precision."},{"cited_title":"D., et al","cited_arxiv_id":null,"evidence_quote":"Supplies the BP−RP infrared-flux-method temperature scale used for an independent Teff comparison."}],"review_version":1}