{"id":"dd477eef-0766-4e6b-b907-a6140bca1a33","arxiv_id":"1908.02717","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 35-star empirical stellar spectral library observed with MUSE provides slit-loss-free optical spectra with verified continuum shapes and Lick indices.","lead":"Astronomers present a new library of 35 high-quality stellar spectra taken with the MUSE instrument on the VLT, spanning 4800 to 9300 Angstroms. These spectra avoid slit losses and multi-order stitching, so their continuum shapes should be more reliable for galaxy and stellar modeling.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Synthetic-SDSS 'tightness' test is blind to the dominant error mode for continuum shapes: a smooth, reddening-like response or extinction error shifts all MUSE stars coherently and leaves the color-color sequence tight.","rationale":"Good-faith reading: this is an honest, useful data-release paper, and the central instrument choice is sound — an IFU genuinely removes slit losses and order-stitching, and the mean-of-six-exposures with r.m.s. demonstrates repeatability. The single most load-bearing weakness is under-verification of the exact quantity the paper promises: continuum-shape fidelity at the <1% level. Section 4's Fig. 5 compares the tightness of synthetic SDSS colors of MUSE versus XSL for the same stars. That is a relative smoothness test. It cannot detect a common-mode smooth slope error, which is the most plausible failure of the §3 assumption that non-photometric extinction and the instrument response are gray across 4800–9300 Å: thin cirrus extinction and standard-star SED uncertainties are approximately monotonic in wavelength, so the resulting tilt shifts every star coherently in color-color space and can even tighten the sequence, making the 'slightly tighter than XSL' statement true while shapes are still wrong. The XSL comparison cannot adjudicate because the Fig. 4 ratios (10–15% smooth trends) mix MUSE and XSL errors; relative tightness implies smoothness, not accuracy. The authors' own Section 5 limitation — external photometry for only about a quarter of the sample, with no synthetic-versus-observed comparison presented — is exactly the missing test. The proposed check (catalog V and I_C versus synthetic, both bands fully inside the MUSE wavelength range) is inexpensive, uses existing bright-star photometry, and would decisively reveal a reddening-like tilt if present. If it passes, the concern is retired; if it fails, the <1% claim and the Table B.1 corrections need revision. I therefore recommend CONDITIONAL rather than ACCEPT: the library is accepted as a valuable product and the method is promising, but the headline slope-fidelity claim must be verified directly or explicitly softened. This agrees with the reader's weakest-assumption analysis and sharpens why the existing test cannot detect the error.","tokens_in":23060,"tokens_out":14049,"duration_ms":166380,"concrete_test":"Run the direct photometric check that Fig. 5 omits. For all 35 stars, collect catalog Johnson-Cousins V and I_C photometry; both bandpasses lie fully inside the MUSE range, and bright-star photometry exists (e.g., Ducati 2002 catalogue, Hipparcos/Tycho compilations). Compute synthetic V and I_C from each MUSE spectrum with the same pyphot tool used in §4, form residuals Δ(V−I_C) = synthetic − observed, and fit Δ(V−I_C) against observed (V−I_C) or spectral type. A monotonic trend exceeding ≈0.01–0.02 mag (a ≈1% flux tilt across 4800–9300 Å) falsifies the <1% slope-fidelity claim and demonstrates that the tightness test was blind to a reddening-like response/extinction error. Residuals consistent with photometric errors and showing no temperature trend would confirm the §3 gray-extinction/response assumption and settle the concern in the paper's favor.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim — that the 35 MUSE spectra have continuum shapes reliable at the <1% level and 'more reliable' than XSL (§4) — rests on two supports: the §3 physical argument that under non-photometric conditions the extinction and the instrument response are gray across 4800–9300 Å, and the §4/Fig. 5 external check, which computes synthetic SDSS colors of the MUSE and XSL spectra. The load-bearing weakness is that Fig. 5 compares the tightness of two synthetic color sequences; it never compares synthetic to observed photometry. A smooth, monotonic (reddening-like) wavelength-dependent error, whether from non-gray cloud extinction, a standard-star SED error, or residual response curvature, shifts every star coherently in color-color space: it does not broaden the sequence, and if the error correlates with temperature it can even tighten it. The comparison against XSL cannot break this degeneracy, because the smooth 10–15% XSL/MUSE ratios in Fig. 4 could in part be MUSE's own error; relative tightness establishes smoothness, not accuracy. The r.m.s. from the six (or twelve) individual exposures (Fig. A.1) tests repeatability only, since all exposures share the same response and the same sky. The authors themselves flag the gap in §5: external photometry exists for only about a quarter of the sample, and no direct synthetic-versus-observed comparison is presented anywhere. The <1% slope-fidelity claim is therefore untested against exactly the error mode — a smooth slope tilt — that most degrades continuum shapes, and the Table B.1 corrections inherit the same blind spot.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a library of 35 high signal-to-noise MUSE stellar spectra covering 4800-9300 Å, selected from the X-shooter Spectral Library sample to span the Hertzsprung-Russell diagram. The central claim is that these spectra have reliable continuum shapes because integral-field observations avoid slit losses and the order-stitching required by cross-dispersed echelle spectrographs. The authors support this claim with the internal repeatability of six (or twelve) individual exposures, aperture-extraction experiments that bound aperture-induced slope changes below 1%, a comparison with XSL spectra that reveals smooth 10-15% continuum discrepancies (attributed to XSL) with polynomial corrections listed in Table B.1, synthetic SDSS color-color diagrams, and Lick indices. The paper acknowledges in §5 that external photometry is available for only about a quarter of the sample.","tokens_in":23401,"tokens_out":10710,"duration_ms":111327,"significance":"If the continuum-shape fidelity is confirmed, this small library would be a valuable reference for the absolute shape of stellar continua and a useful cross-calibrator for XSL and similar multi-order libraries, and it would demonstrate the IFU approach as a promising route for building empirical spectral libraries. The paper has genuine strengths: the same-spaxel placement of science targets and spectrophotometric standards, the careful aperture-loss tests, the public machine-readable spectra (Table A.2), the XSL/MUSE correction coefficients (Table B.1), and the Lick-index table (Table C.1). The internal consistency is good and the data themselves are useful. However, the external verification is weaker than the central claim requires: the synthetic color comparison in Fig. 5 is internal to the two spectral libraries, not a check against observed photometry, and a coherent smooth slope error would evade it.","major_comments":[{"comment":"The claim that the MUSE spectra 'have more reliable shapes' than the XSL spectra rests on the synthetic SDSS color-color comparison in Fig. 5, but both sequences are computed from the spectra themselves, so the test is not external to the two libraries. A smooth, monotonic wavelength-dependent error in the MUSE flux calibration (non-gray cloud extinction, a standard-star SED error, or residual response curvature) shifts all stars coherently in color-color space and leaves the sequence tight; the comparison against XSL cannot break this degeneracy because the 10-15% sloped XSL/MUSE ratios in Fig. 4 could in part reflect MUSE's own systematic error. The r.m.s. of the individual exposures (Fig. A.1) tests repeatability only, since all exposures share the same response and the same sky. The paper's own §5 admits that external photometry exists for only about a quarter of the sample, and no synthetic-versus-observed comparison is shown anywhere. I request a direct, quantitative test for the stars that do have Gaia/SDSS photometry (for example, residuals of synthetic minus observed colors with stated uncertainties), or alternatively a rephrasing that restricts the claim to internal consistency and order-scale agreement with XSL.","section":"§4, Fig. 5"},{"comment":"The argument in §3 that non-photometric observations preserve the 'true' intrinsic shape implicitly assumes that the atmospheric extinction and the response transfer are effectively gray across 4800-9300 Å and that the standard-star calibration introduces no smooth slope error. This assumption is not quantitatively tested. Table A.1 shows differential airmasses between targets and their spectrophotometric standards of up to about 1.3 (for example, HD 100733 at sec z 2.39-2.52 with GD 108 at 1.06), so a small non-gray component of the atmospheric extinction or an error in the standard-star SED would produce exactly the smooth slope error that the r.m.s. of the individual exposures and the aperture-radius experiments cannot reveal. A feasible check would be to extract the observed standard-star spectra and compare them with their tabulated SEDs, or to redo the flux calibration with an alternate standard SED and report the resulting slope change across the MUSE band.","section":"§3, Table A.1"}],"minor_comments":[{"comment":"The text contains several typos: 'build with the MUSE' should be 'built with the MUSE', 'homogenious' should be 'homogeneous', and 'observaitons' should be 'observations'.","section":"Abstract, §1"},{"comment":"'A example of the data products is plotted in Fig. 2' should read 'An example', and the sentence 'the error is the r.m.s. of that averaging' would be clearer as 'the uncertainty is the r.m.s. scatter of the individual exposures about the mean'.","section":"§4"},{"comment":"The caption contains garbled text ('Larg er open circles', 'datab ase', 'and although many statrs are variable') and does not state which SDSS filters and color axes are plotted; both the axes and the filters should be identified explicitly.","section":"Fig. 5 caption"},{"comment":"The phrase 'and the rest and the rest are equivalent widths' is garbled and should be corrected.","section":"Table C.1 caption"},{"comment":"The phrase 'the high blue wavelength limit of MUSE' is ambiguous; it should state that MUSE does not cover wavelengths shortward of about 4800 Å.","section":"§5"},{"comment":"The statement that the MUSE sequences are 'slightly tighter' is not quantified; reporting the RMS dispersion of each sequence about the Lenz et al. (1998) reference lines, and the number of stars used in each case, would make the comparison reproducible.","section":"§4, Fig. 5"}],"recommendation":"major_revision","confidential_remarks":"This is a useful data paper and the IFU approach is a promising way to build spectral libraries, but the central claim about continuum-shape fidelity is currently validated only in a relative sense. The main risk is that the community will take the 'high-fidelity shapes' claim at face value even though the external check compares synthetic to synthetic colors. I recommend that the editor convey that a direct comparison of synthetic MUSE colors with observed photometry for the subset that has it, or an explicit downgrading of the claim, is needed before final acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a useful data paper, not a physics discovery. Using MUSE to build a slit-loss-free empirical library of 35 well-chosen stars is genuinely new, and the spectra plus the polynomial corrections in Table B.1 are a real community resource. The authors are also honest about limitations: they flag the small sample, the missing bluest wavelengths, and the lack of external photometry for three-quarters of the sample.\n\nWhat the paper does well: the data curation is careful (aperture experiments, telluric corrections, binary decontamination), the internal repeatability across six exposures is shown per star, and the Lick index comparisons are a useful sanity check. Releasing the spectra in machine-readable form is exactly right. The comparison with XSL is instructive and the measured 10-15% smooth residuals are a legitimate cautionary tale about order stitching.\n\nWhere it gets soft: the central claim is that these spectra have reliable continuum shapes, at the <1% level, and are \"more reliable\" than XSL. That claim depends on the external verification in Fig. 5, which compares synthetic SDSS colors of MUSE and XSL spectra against each other. This does not test absolute shape accuracy. A smooth, reddening-like response error would shift every star coherently in color-color space, leaving the sequence just as tight. The comparison with XSL cannot assign blame for the smooth 10-15% differences in Fig. 4; the error could be in MUSE's own response or standard-star SED. The r.m.s. of individual exposures tests repeatability only. So the <1% slope-fidelity claim is untested against exactly the smooth error mode that matters most, and the authors even acknowledge the right fix (broadband photometry for more stars) but do not have it here.\n\nThat said, this is not a fatal flaw. The library remains valuable as an independent set of high-S/N spectra with good internal consistency, and the XSL corrections are usable if one treats them as relative rather than absolute. I would not block publication over this, but I would push the authors to either soften the fidelity claim or add a direct synthetic-versus-observed check for the few stars that do have photometry.\n\nWho is this for? People doing stellar population synthesis, template fitting, or testing echelle reduction pipelines. It deserves a serious referee; the data product itself is worth having in the literature. I would recommend acceptance with revisions: tone down the <1% language, make clear the external verification is relative, and ideally add the observed-photometry comparisons that are currently missing.","headline":"A genuinely useful IFU-built stellar library, but the continuum-shape fidelity claim goes beyond what the external checks can support.","tokens_in":23926,"tokens_out":2038,"would_cite":false,"duration_ms":26141,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper supplies 35 high-signal-to-noise MUSE stellar spectra with continuum shapes it argues are more reliable than those of the X-shooter Spectral Library, because the integral field unit avoids slit losses and order stitching.","keywords":["stellar spectral library","integral field unit","MUSE","continuum shape","slit losses","X-shooter Spectral Library","Lick indices","synthetic photometry"],"falsifier":"Compare the continuum of any one library star with a space-based spectrum that has no atmospheric extinction and an independent flux calibration; a smooth blue-to-red slope difference of 1% or more would refute the claimed fidelity.","tokens_in":22918,"feed_emoji":"⭐","tokens_out":7391,"duration_ms":72988,"temperature":0.7,"pith_summary":"This paper aims to establish that an integral-field spectrograph can produce a stellar spectral library whose continuum shapes are more trustworthy than those assembled from cross-dispersed echelle spectra. The authors present 35 high-signal-to-noise MUSE spectra covering roughly 4800–9300 Å for stars spanning the Hertzsprung-Russell diagram, with the central argument that an IFU avoids two known corruptions: slit losses and the stitching together of many spectral orders. If the claim holds, these spectra serve as a shape reference for other libraries and as templates for galaxy and stellar population studies, and the paper's polynomial correction coefficients can be used to repair the continuum of the X-shooter library.","feed_headline":"35 MUSE spectra offer slit-loss-free continuum shapes","feed_subtitle":"Synthetic SDSS colors are tighter than XSL's, and Table B.1 corrects the older library's slope.","key_machinery":"The central object is the integral-field unit itself: MUSE records a two-dimensional field at every wavelength, so all of the star's light that falls in the field is captured and each wavelength is measured contiguously rather than reassembled from separate orders. The paper uses this property to argue that aperture and slit losses change the spectrum's overall slope by less than 1% from blue to red, and it tests the resulting shapes with synthetic SDSS colors, aperture-radius experiments, telluric correction, and Lick indices.","core_discovery":"The central claim is that the MUSE spectra have more reliable continuum shapes than the X-shooter Spectral Library, because the integral-field unit eliminates slit losses and produces continuous wavelength coverage without order stitching. The paper supports this with direct XSL-to-MUSE ratios that show gradual 10–15% slope differences within the MUSE range, and with synthetic SDSS colors that are \"slightly tighter\" for MUSE than for XSL, confirming the shape claim. It also reports Lick indices measured from the new spectra and lists second-order polynomial coefficients in Table B.1 that quantify and correct the XSL/MUSE ratio for individual stars. The intended result is a high-fidelity empirical library that can anchor continuum shapes across roughly 4800–9300 Å.","pith_inferences":["Inference: the paper's internal consistency checks, such as repeat exposures, aperture experiments, and XSL ratios, would not detect a smooth wavelength-dependent response error, so an independent space-based comparison is the cleanest test of the less-than-1% slope claim.","Inference: if the same IFU approach were applied blueward of 4800 Å, the gray-response assumption would need to be re-proved; the current library cannot certify continuum shapes at shorter wavelengths.","Inference: after applying Table B.1 corrections, XSL and MUSE should agree in narrow features and disagree mainly in broad-band slope, so parameters derived from index-based versus color-based methods would flag which library to trust."],"forward_implications":["Users of the X-shooter Spectral Library can apply the second-order polynomial coefficients in Table B.1 to correct the continuum slope of each overlapping star.","Galaxy stellar-population models can use the 35 MUSE spectra as shape-accurate templates across 4800–9300 Å.","Synthetic colors derived from MUSE spectra should reproduce observed broad-band colors more closely than XSL spectra do for the same stars.","The measured Lick indices place these 35 stars on the standard index system, allowing direct index-based stellar and galaxy analysis."],"supporting_citations":[{"why":"Supplies the 35 sample stars and the X-shooter Spectral Library spectra that the MUSE shapes are compared against and corrected for.","marker":"Chen et al. 2014"},{"why":"Introduces the MUSE instrument whose integral-field design is the paper's claimed remedy for slit losses and order stitching.","marker":"Bacon et al. 2010"},{"why":"Defines the Lick index system used to measure the new spectra and to compare with the standard index locus.","marker":"Worthey et al. 1994"},{"why":"Provides the molecfit telluric-line removal used on each individual science exposure before averaging.","marker":"Smette et al. 2015"},{"why":"Describes the Reflex reduction environment that ran the MUSE pipeline for the data cubes.","marker":"Freudling et al. 2013"},{"why":"Supplies the solar-abundance dwarf and giant color-color sequences used to judge the synthetic SDSS colors.","marker":"Lenz et al. 1998"}],"fun_headline_variants":["35 MUSE spectra deliver slit-loss-free continuum shapes","35 MUSE spectra correct XSL's continuum slopes","IFU library: verified continuum shapes, no slit losses","MUSE spectra: reliable shapes from IFU, no stitching","New MUSE library corrects XSL's 10-15% slope errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that, on the non-photometric nights, the Earth's atmosphere and the instrument's response changed the spectrum by the same gray factor at every wavelength, so after dividing by the same-spaxel standard star the measured shape is the true stellar shape.","fun_headline_variants_meta":{"raw":{"variants":["35 MUSE spectra deliver slit-loss-free continuum shapes","35 MUSE spectra correct XSL's continuum slopes","IFU library: verified continuum shapes, no slit losses","MUSE spectra: reliable shapes from IFU, no stitching","New MUSE library corrects XSL's 10-15% slope errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000766,"raw_usage":{"total_tokens":3392,"prompt_tokens":938,"completion_tokens":2454,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":2369}},"tokens_in":554,"tokens_out":2454,"duration_ms":19430,"temperature":1.0,"reasoning_tokens":2369,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:36:22.381661+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the continuum of any one library star with a space-based spectrum that has no atmospheric extinction and an independent flux calibration; a smooth blue-to-red slope difference of 1% or more would refute the claimed fidelity.","supporting_citations":[{"cited_title":"2010, in , Vol","cited_arxiv_id":null,"evidence_quote":"Introduces the MUSE instrument whose integral-field design is the paper's claimed remedy for slit losses and order stitching."},{"cited_title":"D., Newberg , J., Rosner , R., Richards , G","cited_arxiv_id":null,"evidence_quote":"Supplies the solar-abundance dwarf and giant color-color sequences used to judge the synthetic SDSS colors."}],"review_version":1}