{"id":"073b88c9-d4d2-49a4-8a25-c145f05eef7d","arxiv_id":"2608.10250","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An automated pipeline and public database for massive star spectra recovers literature stellar parameters for OB stars and red supergiants, with the red-supergiant temperature check relying on the same code used to produce the reference values.","lead":"This paper presents Astro+, a public web database for massive star spectra, together with HiLineThere, a fully automated pipeline that classifies stars and measures stellar parameters without human intervention. The authors report that for OB stars the automated parameters agree with expert measurements to about 800 K in temperature, 0.1 dex in gravity, and 10 km/s in rotation, and that radial velocities for red supergiants agree to about 7 km/s.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"OB validation is partly in-sample: the pipeline was iteratively tuned on the Holgado et al. (2018) benchmark, so the claimed Teff/logg/vsini accuracy needs an independent test.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing concern: the OB validation is in-sample because the code was iteratively adjusted to reproduce the Holgado et al. (2018) results. This matters because the tool's advertised value is to replace human inspection for large new surveys; if the quoted errors are tuned to a single benchmark, they may not generalize. The paper's own admission in Sect. 3.1 makes this concrete, not hypothetical. The RSG Teff circularity is a related but secondary issue; the Vr validation against Dorda et al. (2018) is independent and supports the red-path radial velocity claim. The TLUSTY experiment is a useful sanity check but does not test generalization to real, heterogeneous observed spectra. The proposed out-of-sample test on an independent OB sample directly adjudicates whether the pipeline's accuracy holds beyond the tuning set. Since the reader already recommends CONDITIONAL and our concern aligns with that recommendation, the verdict should remain CONDITIONAL (i.e., no change from the reader's verdict).","tokens_in":37629,"tokens_out":2610,"duration_ms":24741,"concrete_test":"Select a sample of ~50 OB stars with published stellar parameters from a survey not used in the pipeline's development (e.g., VFTS or IACOB OB stars analyzed by independent authors). Run the unmodified HiLineThere pipeline on their reduced spectra and compute the median absolute differences in Teff, logg, and vsini relative to the literature values. If these differences exceed the claimed ~800 K, 0.1 dex, or 10 km/s respectively, the in-sample tuning has inflated the reported accuracy. Separately, repeat the RSG Teff comparison against an independent analysis not based on SteParSyn to break the circularity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that HiLineThere autonomously recovers stellar parameters within Teff~800 K, logg~0.1 dex, and vsini~10 km/s for OB stars. The validation in Sect. 3.1 uses the Holgado et al. (2018) sample, and the paper explicitly states that an 'iterative comparison process led to progressively more complex and rigorous versions of the code, ensuring that the automated reduced chi-squared minimization reproduces expert-level parameters.' This is tuning on the benchmark: the algorithm's line weights, thresholds, and selection logic were adjusted until agreement with this specific sample was achieved. The reported differences are therefore partly in-sample and may understate the errors on genuinely new spectra. The TLUSTY experiment in Sect. 3.3 is an internal model-to-model consistency check, not an independent observational benchmark. The red-supergiant Teff comparison in Sect. 3.2 is also circular because it uses the same SteParSyn code as the reference Negueruela et al. (2018) analysis; only the Vr comparison against Dorda et al. (2018) is independent. Without a held-out sample, the headline accuracy figures are not yet established for the 'fully autonomous' use case advertised for large surveys.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Astro+, a web-based database for massive-star spectra, together with HiLineThere, an automated analysis pipeline. HiLineThere routes spectra through three paths: a blue path using a large FASTWIND grid for OB stars, a yellow path using KURUCZ/ATLAS9 models for intermediate temperatures, and a red path using MARCS models and the SteParSyn code for cool stars and red supergiants. The pipeline automatically measures radial velocity, vsini, and macroturbulence from diagnostic lines, then derives Teff and logg by reduced chi-square fitting. Validation is carried out against the O-star sample of Holgado et al. (2018) and a small set of early-B stars from Nieva & Przybilla (2014), reporting differences of Teff ~ 800 K, logg ~ 0.1 dex, and vsini ~ 10 km/s for OB stars; against Dorda et al. (2018) for RSG radial velocities (within 7 km/s); and against Negueruela et al. (2018) for RSG Teff (within ~300 K). A TLUSTY model-injection experiment is added to estimate systematic errors of the blue path. The paper claims that the tool can homogeneously process large future surveys in a completely autonomous way.","tokens_in":37998,"tokens_out":6861,"duration_ms":72159,"significance":"If the accuracy figures are confirmed on independent data, Astro+ and HiLineThere would be a timely and valuable contribution: they address a real bottleneck for upcoming multi-object surveys of massive stars, provide a public database, and include a transparent hierarchical decision tree with quantitative comparisons against large, carefully studied samples. The authors are also commendably explicit about several limitations, such as the deferred validation of the yellow path, the use of the same code in the RSG Teff comparison, and the non-like-for-like nature of the TLUSTY test. However, the headline accuracy claims are not yet established for genuinely new spectra because the OB validation sample was used iteratively to tune the algorithm, and the RSG Teff validation is circular. The significance of the paper therefore hinges on the additional independent validation requested below.","major_comments":[{"comment":"The paragraph beginning 'This iterative comparison process led to progressively more complex and rigorous versions of the code' explicitly states that the Holgado et al. (2018) sample was used as a tuning set during development. Both Test 1 and Test 2 are applied to this same sample, so the reported agreement in Teff, logg, and vsini is partly in-sample and may not represent performance on new spectra. This directly affects the abstract's central claim of 'completely autonomous' analysis with quoted differences of ~800 K, ~0.1 dex, and ~10 km/s. Please add a genuinely independent validation: for example, hold out a subset of the Holgado et al. (2018) stars during all tuning and report residuals on that subset, or use another large independent sample such as IACOB or VFTS. Alternatively, if such a test is not feasible, the quoted differences should be explicitly reframed as internal consistency with the development benchmark, not as expected errors for survey data.","section":"Sect. 3.1, O-type stars"},{"comment":"The Teff comparison against Negueruela et al. (2018) is circular because the same SteParSyn implementation and MARCS grid were used in both the reference analysis and the present pipeline. The text acknowledges this in the sentence 'we must keep in mind that we are using the same code as in the original paper,' but the abstract and conclusions still present ~300 K as a validation result. The only independent red-path check in the paper is the radial-velocity comparison against Dorda et al. (2018). Please remove the Teff comparison from the validation claims, or supplement it with an independent set of RSG parameters derived with a different code or method. As currently presented, the evidence for the accuracy of the red-path Teff, logg, and [Fe/H] is insufficient.","section":"Sect. 3.2, Late-type stars"},{"comment":"The TLUSTY experiment is a model-to-model consistency check rather than an observational validation. The injected spectra are noiseless synthetic models, so the test does not exercise the pipeline's response to normalization errors, cosmic-ray residuals, weak blends, wind variability, or other real-data artifacts that contribute to the actual error budget. The final cautionary sentence of the section is appropriate, but the earlier statement that this experiment gives 'an upper limit on the systematic differences... should be below 2,000 K' can easily be read as an accuracy claim. Please present this test strictly as a code-consistency sanity check and avoid using it to bound the systematic error of the pipeline on real spectra.","section":"Sect. 3.3, Systematic errors on blue path"}],"minor_comments":[{"comment":"The displayed expression for the macroturbulence kernel appears to contain a typo: the proposed delta-function term '(-v pi^(1/2)/zeta_RT) delta(-v^2/zeta_RT^2 - 1)' is dimensionally inconsistent and is not a standard radial-tangential expression. Please check the formula; as written it cannot be implemented literally.","section":"Sect. 2.1.3, Eq. (8)"},{"comment":"The comparison with Holgado et al. (2018) would benefit from explicit summary statistics for the whole sample (mean and rms or median absolute differences for Test 1 and Test 2), rather than only qualitative statements and quoted subgroup values. This would make it easier for readers to assess the headline numbers.","section":"Sect. 3.1"},{"comment":"The Teff axis of Fig. 15 is labeled in kK but the plotted range is 3000-7000 K, so the units are inconsistent. In addition, the text does not provide the rms scatter for the 11-star sample; please include it.","section":"Sect. 3.2 and Fig. 15"},{"comment":"The phrase 'the full spectrum with all the features considered by the code' is misleading: Test 2 uses all detected H, Hei, and Heii diagnostic lines from Table 1, not the entire spectral range. Consider rewording to 'all available diagnostic lines.'","section":"Sect. 3.1, Test 2 description"},{"comment":"The headers 'eTeff' and 'e(logg)' are not self-explanatory; please use standard notation such as 'sigma(Teff)' or a dedicated error column. Also, the table lists only 11 stars, while the text just says 'small sample'; please specify the sample selection from Negueruela et al. (2018).","section":"Table C.3"},{"comment":"The phrasing 'No point closer to line center than Gaussian sigma deviates by >3 sigma fit' is awkward; a clearer formulation is 'no point within one Gaussian sigma of the line center deviates from the fit by more than 3 sigma.' Also, Section 5 states that the database is public but the HiLineThere source code is not; given the paper's aim of community use, please provide a code repository or state clearly that the code is available on request.","section":"Sect. 2.1.2, Condition IV and Sect. 5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of A&A and addresses a timely problem. The main issue is that the validation strategy is partly in-sample and partly circular, and this is load-bearing for the central accuracy claims. The authors themselves acknowledge the circularity for the RSG Teff comparison and partially acknowledge the iterative tuning for the OB path. I do not see this as grounds for rejection, because the pipeline is described in enough detail that a held-out or independent validation appears feasible, or the claims could be reframed honestly. I recommend major revision and ask the editor to ensure that the authors either provide an independent benchmark or substantially temper the statements in the abstract and conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real infrastructure contribution with an honest limitations section, and the main accuracy claims for the blue path are credible enough to deserve referee time, but the headline numbers partly come from tuning on the benchmark, and the red-path Teff validation is circular. The yellow path is explicitly unvalidated.\n\nWhat's actually new: Astro+ is the first public multiwavelength database dedicated to massive-star spectra, and HiLineThere's hierarchical temperature routing (FASTWIND blue, KURUCZ/ATLAS9 yellow, MARCS/SteParSyn red) is new as an integrated autonomous tool. The paper ships real code, a public database, and a detailed validation campaign. The OB comparison with roughly 80 Holgado stars plus the Nieva-Przybilla early-B stars is substantial, and the residual scatter is consistent with the stated ~800 K, ~0.1 dex, ~10 km/s errors. The treatment of binaries is sensible. The TLUSTY experiment is a reasonable cross-code consistency check, and the authors correctly warn that it is not like-for-like.\n\nSoft spots: the stress-test note is right. The OB validation is partly in-sample: Sect. 3.1 says the code was iteratively refined until it reproduced the Holgado sample, so the quoted accuracy is not fully out-of-sample. That doesn't kill the claim, but it caps how much trust the numbers deserve. The RSG Teff comparison against Negueruela et al. (2018) is explicitly circular — same SteParSyn code — as the authors acknowledge in Sect. 3.2. Only the Vr comparison against Dorda is independent. The yellow path is central to the full-temperature-range claim but has no systematic validation; the paper defers it to Paper II, which is honest but leaves a big gap. Minor: microturbulence is fixed at 10 km/s and the grid is solar-metallicity only, so the current blue path is intentionally narrow.\n\nVerdict: the paper overstates what is established — the 'fully autonomous across the HR diagram' framing is stronger than what the validation supports — but it is honest about the circularity and the tuning. This looks like a solid foundation paper, not a finished accuracy claim. I'd send it to peer review with a request for an out-of-sample test (e.g., synthetic spectra or a held-out subset) and a sharper distinction between validated and unvalidated paths. For the red path, an independent Teff benchmark is needed. The paper is worth engaging seriously; I'd bring it to reading group and would cite it once the database is stable.","headline":"A genuine infrastructure paper with honest caveats: the blue-path accuracy claims are plausible but partly in-sample, the red-path Teff check is circular, and the yellow path is unvalidated — still worth a serious referee.","tokens_in":38460,"tokens_out":1528,"would_cite":true,"duration_ms":14525,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fully automatic pipeline, HiLineThere, can derive stellar parameters for massive OB stars and red supergiants, matching expert literature values within about 800 K in effective temperature, 0.1 dex in surface gravity, and 10 km/s in…","keywords":["Stars: massive","Astronomical databases","Techniques: spectroscopic","Methods: data analysis","FASTWIND models","OB-type stars","red supergiants","automated stellar parameters"],"falsifier":"Run HiLineThere on a sample of massive-star spectra that were never used in its development and whose parameters were determined independently, or on synthetic spectra from an independent atmosphere code such as CMFGEN or PoWR; if the differences systematically exceed about $800$ K in $T_{\\rm eff}$, $0.1$ dex in $\\log g$, or $10$ km s$^{-1}$ in $v\\sin i$, the claimed autonomy overstates its accuracy.","tokens_in":37417,"feed_emoji":"🔭","tokens_out":6540,"duration_ms":56872,"temperature":0.7,"pith_summary":"The paper is trying to establish that human expertise is no longer required to extract basic stellar parameters from massive-star spectra. It presents Astro+, an open database of over 4,000 spectra, and HiLineThere, a fully automatic Python pipeline that classifies each spectrum and fits effective temperature, surface gravity, rotation, and radial velocity. The validation claim is that the automated results agree with expert literature values within about $800$ K in $T_{\\rm eff}$, $0.1$ dex in $\\log g$, and $10$ km s$^{-1}$ in $v\\sin i$ for OB stars, and within about $7$ km s$^{-1}$ in radial velocity and $300$ K in $T_{\\rm eff}$ for red supergiants. If this holds, the pipeline can digest the tens of thousands of spectra expected from upcoming multi-object surveys without the manual line selection that currently bottlenecks analysis.","feed_headline":"Fully automatic spectrum analysis matches expert stellar parameters","feed_subtitle":"HiLineThere derives temperature, gravity, and rotation for OB stars within ~800 K, 0.1 dex, and 10 km/s of literature values.","key_machinery":"The load-bearing mechanism is a hierarchical decision tree combined with reduced chi-squared minimization over large model grids. A line-detection routine identifies Balmer, He I, He II, and metal lines; their absence or presence routes the spectrum to MARCS-based SteParSyn (red path), KURUCZ/ATLAS9 (yellow path), or FASTWIND (blue path). For blue-path stars, the projected rotation $v\\sin i$ is set from the first zero of the Fourier transform of a line profile, macroturbulence $\\zeta$ is then fitted, and a reduced chi-squared comparison to $46\\,000$ solar-metallicity FASTWIND models yields $T_{\\rm eff}$ and $\\log g$, with a subroutine that discards H$\\alpha$ when wind emission invalidates it.","core_discovery":"The paper's central claim is that a fully automated program, HiLineThere, can take an optical spectrum of a massive star and, without any human input, classify it and determine effective temperature, surface gravity, projected rotational velocity, and radial velocity. For OB-type stars analyzed against a grid of FASTWIND models, comparison with the expert-analysis sample of Holgado et al. (2018) and early-B stars from Nieva & Przybilla (2014) yields differences within about $800$ K in $T_{\\rm eff}$, $0.1$ dex in $\\log g$, and $10$ km s$^{-1}$ in $v\\sin i$; for red supergiants, radial velocities agree within $7$ km s$^{-1}$ and temperatures within about $300$ K. The pipeline's hierarchical logic routes spectra into blue, yellow, or red analysis paths depending on whether Balmer lines are present and whether the star is hotter or cooler than about $15\\,000$ K. The authors claim this replaces human inspection for large spectroscopic surveys while providing strictly homogeneous parameters.","pith_inferences":["If the pipeline's accuracy holds on independent data, the bottleneck in massive-star spectroscopy shifts from parameter fitting to quality control: the inspection graph becomes a spot-check rather than the analysis itself.","A sharper validation would feed the pipeline synthetic spectra from an independent grid that includes winds, line blanketing, and non-solar abundances; the TLUSTY test already points in this direction but mixes resolution effects with atmosphere-code differences.","The yellow path (7,500-15,000 K) is implemented but not systematically validated, so claims of full OBAFGKM coverage currently rest on the blue and red path tests plus the early-B transition sample.","Because the blue-path grid is solar-metallicity with microturbulence fixed at 10 km/s, the quoted accuracies should be treated as conditional on those assumptions; extending the grid is a testable way to see how much of the residual scatter they explain."],"forward_implications":["Surveys such as WEAVE can have their OB and red supergiant spectra processed automatically and homogeneously, with no expert line selection.","Single-lined binaries and binaries with faint secondaries are analyzed without degrading parameter accuracy, since their snapshot spectra behave like single stars.","Spectra with chemical peculiarities, such as the ON supergiant HD 105056, are correctly recovered when the pipeline is allowed to select lines freely rather than using a fixed line list.","Red supergiant radial velocities from the CaT region are recovered to within about 7 km/s, making the database useful for kinematic studies of cool supergiants.","The model-grid experiment with TLUSTY spectra places an upper bound of roughly 2,000 K on temperature systematics for O-type stars, though part of that difference reflects known code-to-code atmosphere physics."],"supporting_citations":[{"why":"Supplies the expert O-star sample, line-by-line analysis, and benchmark parameters the pipeline is tuned against and compared with.","marker":"Holgado et al. (2018)"},{"why":"Supplies the early-B star sample used to test the pipeline's behavior near the lower-temperature edge of the blue path.","marker":"Nieva & Przybilla (2014)"},{"why":"Supplies the red supergiant sample and radial velocities used to validate HiBandThere.","marker":"Dorda et al. (2018)"},{"why":"Supplies the cool supergiant sample used to validate red-path effective temperatures.","marker":"Negueruela et al. (2018)"},{"why":"Provides SteParSyn, the Bayesian code used for red-path stellar parameter determination.","marker":"Tabernero et al. (2022)"},{"why":"Introduces FASTWIND, the atmosphere code behind the blue-path grid.","marker":"Santolaya-Rey et al. (1997)"},{"why":"Provides the TLUSTY model grid used in the systematic-error experiment.","marker":"Lanz & Hubeny (2003)"}],"fun_headline_variants":["Automated spectrum analysis matches expert stellar parameters","Fully automatic stellar parameter extraction wins accuracy test","Robotic star spectroscopy rivals human expert analysis","Massive star spectra decoded automatically within error bars","No manual fitting: automated pipeline hits stellar parameters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The validation benchmarks are independent of how the algorithm was built: the code was iteratively adjusted until it reproduced the Holgado et al. (2018) sample, so the quoted agreement on that sample may overstate performance on truly new spectra.","fun_headline_variants_meta":{"raw":{"variants":["Automated spectrum analysis matches expert stellar parameters","Fully automatic stellar parameter extraction wins accuracy test","Robotic star spectroscopy rivals human expert analysis","Massive star spectra decoded automatically within error bars","No manual fitting: automated pipeline hits stellar parameters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000678,"raw_usage":{"total_tokens":3149,"prompt_tokens":1077,"completion_tokens":2072,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":693,"completion_tokens_details":{"reasoning_tokens":2003}},"tokens_in":693,"tokens_out":2072,"duration_ms":15883,"temperature":1.0,"reasoning_tokens":2003,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:23.275645+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run HiLineThere on a sample of massive-star spectra that were never used in its development and whose parameters were determined independently, or on synthetic spectra from an independent atmosphere code such as CMFGEN or PoWR; if the differences systematically exceed about $800$ K in $T_{\\rm eff}$, $0.1$ dex in $\\log g$, or $10$ km s$^{-1}$ in $v\\sin i$, the claimed autonomy overstates its accuracy.","supporting_citations":[],"review_version":1}