REVIEW 7 cited by
SpectraFM: Tuning into Stellar Foundation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Machine learning models in astrophysics are often limited in scope and cannot adapt to data from new instruments or tasks. We introduce SpectraFM, a Transformer-based foundation model architecture that can be pre-trained on stellar spectra from any wavelength range and instrument. SpectraFM excels in generalization by combining flexibility with knowledge transfer from pre-training, allowing it to outperform traditional machine learning methods, especially in scenarios with limited training data. Our model is pre-trained on approximately 90k examples of synthetic spectra to predict the chemical abundances (Fe, Mg, O), temperature, and specific gravity of stars. We then fine-tune the model on real spectra to adapt it to observational data before fine-tuning it further on a restricted 100-star training set in a different wavelength range to predict iron abundance. Despite a small iron-rich training set of real spectra, transfer learning from the synthetic spectra pre-training enables the model to perform well on iron-poor stars. In contrast, a neural network trained from scratch fails at this task. We investigate the Transformer attention mechanism and find that the wavelengths receiving attention carry physical information about chemical composition. By leveraging the knowledge from pre-training and its ability to handle non-spectra inputs, SpectraFM reduces the need for large training datasets and enables cross-instrument and cross-domain research. Its adaptability makes it well-suited for tackling emerging challenges in astrophysics, like extracting insights from multi-modal datasets.
Forward citations
Cited by 7 Pith papers
-
Microlensing Detection and Inference via Learned Bayes Factors
A unified transformer-based pipeline detects 99.9% of recoverable simulated microlensing events and outperforms literature hard cuts in the short-duration finite-source regime with amortized neural posterior inference.
-
Emulation of non-linear 1D spectral models: relativistic X-ray reflection
A modular operator-learning emulator (RTFAST2) reproduces the relativistically convolved reflection spectrum of reltrans to O(0.1)% precision with 4–10× speed-up and unbiased posterior recovery on simulated spectra.
-
How Low Can We Go? Minimum Spectroscopic Requirements For Supernova Subtype Classification
ABC-SN classifies ten supernova subtypes with no performance loss down to R_λ=50 and SNR=5, and only minimal loss at R_λ=25.
-
Generalization from Low- to Moderate-Resolution Spectra with Neural Networks for Stellar Parameter Estimation: A Case Study with DESI
Pre-trained MLPs on LAMOST low-resolution spectra generalize to DESI medium-resolution spectra for [Fe/H] and [α/Fe], outperforming the DESI SP pipeline in zero-shot and improving with modest fine-tuning.
-
ABC-SN: Attention Based Classifier for Supernova Spectra
ABC-SN, a transformer-based classifier, reaches 82.5% macro F1 on ten supernova subtypes, outperforming a retrained DASH at 58.9% on the same test set.
-
Foundation Models for Astrophysics
Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...
-
From stellar light to astrophysical insight: automating variable star research with machine learning
An invited review of machine learning for automated variable star research, covering data cleaning, variability classification, stellar parameter inference, and foundation models.
Discussion (0). Continue with ORCID to comment.