Pith. sign in

REVIEW 3 major objections 4 minor 150 references

A foundation model is defined by whether its representation transfers, not by its architecture or scale.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 04:27 UTC pith:PEYNXB6S

load-bearing objection A well-argued review that defines foundation models by transfer and gives astro-ML a useful evaluation agenda; the main weak spot is the non-auditable claim that transfer demonstrations are comparatively rare. the 3 major comments →

arxiv 2608.02573 v1 pith:PEYNXB6S submitted 2026-08-03 astro-ph.IM

Foundation Models for Astrophysics

classification astro-ph.IM
keywords foundation modelsrepresentation learningtransfer learningself-supervised learningastronomyastronomical surveysfew-shot learningdomain gap
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that the term 'foundation model' should be reserved for networks whose learned representation transfers to new tasks, instruments, and populations, working with few or no task-specific labels. It develops the idea of a transferable representation from first principles, then reads the astronomical literature through that lens. The paper's central observation is that many astronomy models borrow the transformer-and-self-supervision recipe without demonstrating such transfer; clear demonstrations of transfer across instruments, populations, and tasks remain comparatively rare. The stakes are practical: astronomy has abundant unlabeled data, scarce labels, and a persistent gap between simulations and observations, so a representation that genuinely transfers would be far more valuable than one that merely fits a single survey.

Core claim

The paper's core claim is definitional and evidentiary. A foundation model is not identified by its components — a transformer, a self-supervised objective, large-scale pretraining — but by the property those components are meant to produce: a representation, the internal vector an encoder builds, that remains useful when the instrument, the population, or the task changes. The paper calls this the 'transferable representation' and gives it an operational meaning: a target should be readable from the representation with little labeled data, ideally zero-shot or few-shot. Surveying foundation-model attempts in astronomy by modality — light curves, spectra, images, multimodal data, and physica

What carries the argument

The central object is the learned representation — the vector an encoder produces from raw observations, which is meant to separate physical factors such as temperature, gravity, composition, and redshift from nuisances such as noise, calibration, and instrument response. The load-bearing property is transferability, measured operationally as the ease with which a linear probe reads a target from a frozen embedding using few labels. Around this object the paper organizes the machinery: architectures and inductive biases, self-supervised objectives (autoregressive prediction, masked autoencoding, joint embedding), scaling laws, fine-tuning strategies from linear probes to low-rank adaptation,

Load-bearing premise

The load-bearing premise is that the small set of representative models surveyed in Section 3 is a fair sample of the field; if that selection is biased, the conclusion that real transfer is comparatively rare could be an artifact of the survey rather than a property of astronomy.

What would settle it

Take the models the paper surveys and evaluate each frozen encoder on data from an instrument, population, or signal-to-noise range excluded from its pretraining, using only a linear probe and a handful of labels. If most retain useful accuracy, the paper's 'comparatively rare' claim would be disproven; if one model without demonstrated transfer is nevertheless widely accepted as a foundation model, the definitional claim would need revision.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the definition holds, many current astronomical models should be described as task-specific pretrained networks until transfer across instruments or populations is demonstrated.
  • Evaluation practice would change: random train/test splits are insufficient; transfer claims require holding out a whole instrument, population, or signal-to-noise range, with sparse tails handled separately.
  • Scale alone is unlikely to close the gap in astronomy because the field is data-limited; inductive biases and pretext tasks matched to physical structure should matter more than model size.
  • A pretext task tied to physics — for example, tokenizing spectra by line profiles rather than fixed bins — is a concrete opening the paper identifies for better transfer.
  • If real transfer is achieved, astronomy would become a testing ground for physical AI, because a representation can be checked against known physics and across independent instruments.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A cheap empirical test of the paper's claim would be to freeze an existing pretrained encoder and apply a linear probe on data from an instrument or population excluded at pretraining; the resulting accuracy gap would quantify how rare true transfer is.
  • The paper's survey is informal, so its scarcity conclusion would be on firmer footing with a systematic, inclusion-explicit census of astronomical pretrained models.
  • If the transfer-based definition is adopted, model releases would need to document evaluation splits and data provenance, not merely release weights, before the label can be audited.
  • The same criterion would apply beyond astronomy: many scaled-up language and vision models might lose 'foundation model' status if their fluent generation is not matched by transferable measurement, which is exactly why the authors propose physics as a tougher proving ground.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This chapter reviews foundation models for astrophysics from the standpoint of transferable representations. It defines a foundation model by its learned representation transferring to new tasks, instruments, or populations with little or no task-specific data, and it repeatedly distinguishes this goal from the toolkit of transformers, self-supervised objectives, and large-scale pretraining. The first half develops the conceptual machinery: representation quality, inductive biases, self-supervised pretraining, multimodal alignment, scaling laws, and fine-tuning. The second half surveys attempts in time domain, spectroscopy, imaging, multimodal, other messengers, and physical simulations, then offers a cautious assessment: many systems adopt foundation-model architecture, but clear demonstrations of transfer across instruments, populations, and tasks in astronomy remain comparatively rare. The chapter concludes with open questions and directions, emphasizing that transfer evidence must be built into evaluation rather than read from leaderboards.

Significance. If accepted, the paper's conceptual contribution is valuable: it gives the astronomical community a clear, transfer-centered definition of foundation models and a principled account of why architecture and pretraining alone are insufficient. The honest, self-critical treatment of limitations—including the authors' own SpecCLIP and scaling-law results—is a notable strength, as is the explicit call for held-out-instrument/population evaluation rather than random splits. The chapter is also well suited as a pedagogical introduction. However, the central empirical claim that transfer demonstrations are 'comparatively rare' is the load-bearing premise for the chapter's cautious verdict, and that claim is not backed by a systematic, auditable survey. This is the main weakness and, in my view, the reason the manuscript needs a revision before publication.

major comments (3)
  1. [Section 3, Table 2, and Section 3.3] The central claim that clear demonstrations of transfer are 'comparatively rare' rests on an informal, representative selection of models. Table 2 has no stated inclusion criteria, no exhaustive search protocol, no denominator, and no quantitative counts of how many candidate models were considered or how many exhibited the transfer the authors define. Consequently, the scarcity conclusion is not checkable: a different selection—for example, one weighted toward transfer-focused evaluations—could change the verdict. This is load-bearing because the chapter's recommendation to demand stronger evidence depends on the scarcity being a real property of the field rather than an artifact of selection. Please either conduct a systematic, reproducible audit (search strategy, inclusion/exclusion criteria, explicit classification of each model by transfer evidence) or explicitly reframe the claim a
  2. [Section 3.3, §3.3] The limitations list 'Few clear wins over supervised baselines' and 'The synthetic gap is rarely closed' with only a few citations, mainly [73] and [128], and no comparative data. These are themselves empirical assertions about the literature. As written, a reader cannot tell how many studies used matched baselines, how many held out whole instruments or populations, or how many succeeded or failed. At minimum, please tabulate the surveyed studies with columns for domain, transfer test, baseline matching, and outcome. If such a table is not feasible, these assertions should be softened to reflect that they are qualitative impressions rather than established findings.
  3. [Section 4 and Section 5] The chapter states that robust transfer is 'still uncommon even in vision and the wider physical sciences' and that 'astronomy will not reach it by scale alone.' These are broad empirical and prognostic claims made without systematic citation or survey. The first may be defensible, but as stated it is not documented; the second is presented as a conclusion from the preceding review, yet it depends on the same non-auditable sample. Please either support these statements with a focused literature comparison or mark them clearly as the authors' interpretation rather than established consensus.
minor comments (4)
  1. [References] References [120] and [121] appear to be duplicates: both are 'Universal Spectral Tokenization via Self-Supervised Panchromatic Representation Learning' by Shen, Lanusse, Parker, et al. with the same arXiv identifier 2510.17959. Please merge or correct.
  2. [Table 2, physical simulations row] The 'Physical simulations' row lists MPP and Walrus, which are trained on PDE and continuum-dynamics problems rather than astronomical survey data. The text does clarify this is a separate line, but the table could mislead a reader into counting these as astronomical foundation models. Consider relabeling the row or adding a footnote that these are cross-system physical surrogates, not astronomical data models.
  3. [Figure 3 caption] The vertical axis is described as 'accuracy on a downstream task, a stand-in for how well that task is done, classification accuracy or an equivalent score for a regression.' For regression tasks, accuracy is not a natural notion; suggest using 'performance' or 'inverse loss' for clarity.
  4. [Section 2.5, 'about d/2r times fewer parameters'] This sentence is mathematically correct but could be clearer: the factor compares the number of trainable parameters in LoRA (2dr) with the number in the frozen matrix (d^2). Consider adding a brief derivation or a parenthetical for readers unfamiliar with rank decomposition.

Circularity Check

0 steps flagged

No circular derivation: transfer-based definition is operational and the scarcity claim is empirical; self-citations are supporting, not load-bearing.

full rationale

The chapter's central claim is not a derived prediction. Its definition of a foundation model — 'The presence of a transformer, a self-supervised objective, and large-scale pretraining does not by itself make a model a foundation model, since the defining property is that the learned representation transfers, as tested by its ability to work on new tasks with little or no task-specific training data' — is a definitional criterion, not an equation fitted to data. The scarcity claim that 'clear demonstrations of transfer across instruments, populations, and tasks remain comparatively rare' is an empirical assessment supported by the narrative survey in Section 3 and Table 2. That selection is informal and not auditable, which is a methodological weakness and a correctness risk, but it is not circular: no specific result in the paper is equivalent by construction to its input. The paper's own earlier works ([91, 97, 111, 128, 147, 148, 149, 151]) are cited for background or illustrative claims — Cycle-StarNet for domain adaptation, scaling laws for light curves and spectra, FALCO and SpecCLIP as representative models, and [148] even documents SpecCLIP's degradation in sparse tails, which is self-critical. The limitation about reconstruction as a weak pretext task cites [128], the author's own review, but that point is one supporting critique among several and is not the load-bearing justification for the transfer-scarcity conclusion. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via citation. Under the required standard of exhibiting a specific reduction (e.g., Eq. X = Eq. Y by construction, or a fitted parameter renamed as prediction), no circular step can be identified. The self-citations are minor and supporting, so the score is 2 rather than 0, but the central argument has independent content.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 0 invented entities

The chapter makes no empirical fits, so free_parameters is empty. The axioms listed are the definitional and methodological premises the central thesis leans on. The most consequential is the definition of foundation model by transfer, followed by the assumption that Table 2 fairly represents the literature. The invented_entities list is empty because the chapter introduces no new physical or formal objects; 'transferable representation' is a pre-existing concept from the cited literature.

axioms (6)
  • domain assumption A model earns the label 'foundation model' only if its representation transfers to new tasks, instruments, or populations, regardless of architecture or pretraining objective.
    Section 1 and the Abstract state this definition; the entire literature assessment depends on it. If one instead defines foundation models by architecture, the paper's central scarcity claim dissolves.
  • domain assumption Table 2's 'representative models' and the cited evaluations are a fair sample of the astronomical foundation-model literature.
    The claim that transfer demonstrations are 'comparatively rare' is not backed by an exhaustive or quantitative survey; the conclusion depends on the representativeness of the author-selected examples.
  • domain assumption Transfer is best tested by holding out whole instruments, populations, or signal-to-noise ranges rather than random splits.
    Section 3.3 asserts that a random split preserves selection effects; this evaluation standard is adopted to judge existing models and to define what would count as evidence.
  • standard math Unsupervised representation learning is fundamentally non-identifiable without additional structural assumptions.
    The paper invokes Locatello et al. [81] in Section 4 to argue that physical axes cannot be recovered from data alone; this mathematical identifiability result underpins the recommendation that physical meaning must be earned with labels or interventions.
  • domain assumption Neural scaling laws are empirical regularities, not consequences of a first-principles theory.
    Section 2.4 states this explicitly; the conclusion that 'astronomy will not reach it by scale alone' relies on treating scaling laws as observations rather than guarantees.
  • domain assumption Grokking is evidence that training can move a network from memorization to a general rule that transfers.
    Section 1.1 uses grokking to motivate the transferable-representation view; if grokking were only a small-scale curiosity, the language-success analogy would be weaker.

pith-pipeline@v1.3.0-daily-deepseek · 31680 in / 13492 out tokens · 125846 ms · 2026-08-04T04:27:01.422734+00:00 · methodology

0 comments
read the original abstract

Foundation models are high-capacity networks pretrained once on broad data and then reused across many tasks. This chapter introduces them through the idea of a transferable representation, the internal description a network forms during training, which, rather than the fitted task, is what carries over to new problems. We develop the idea from first principles for an astronomical reader, starting from why a representation matters and what makes one useful, and then surveying the architectures, self-supervised objectives, scaling, adaptation, and cross-modal learning that produce one. A theme throughout is the distinction between these methods and the goal they serve. The presence of a transformer, a self-supervised objective, and large-scale pretraining does not by itself make a model a foundation model, since the defining property is that the learned representation transfers, as tested by its ability to work on new tasks with little or no task-specific training data (few-shot and zero-shot learning). We then consider astronomy, where data are abundant but labels are scarce and simulations often stand in for ground truth. Here we offer a cautious reading of the current literature, in which many models adopt the architecture of foundation models while clear demonstrations of transfer across instruments, populations, and tasks remain comparatively rare. This is to be expected, since robust transfer beyond language is still uncommon even in vision and the wider physical sciences, and whether further scaling or a different account of representation will close the gap remains an open question. We close by placing the goal within the broader aim of machine intelligence and outlining the evidence that would mark real progress.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

150 extracted references · 79 linked inside Pith

  1. [1]

    Flamingo: a Visual Language Model for Few-Shot Learning.arXiv e-prints, page arXiv:2204.14198, April 2022

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Ruther- ford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Ja- cob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkow...

  2. [2]

    Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture.arXiv e-prints, page arXiv:2301.08243, January 2023

    Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture.arXiv e-prints, page arXiv:2301.08243, January 2023

  3. [3]

    V-JEPA 2: Self-Supervised Video Models Enable Understand- ing, Prediction and Planning.arXiv e-prints, page arXiv:2506.09985, June 2025

    Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba, Komeili, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bo- janowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krish- nakumar, Yong Li,...

  4. [4]

    Foundational Models Defining a New Era in Vision: A Survey and Outlook.arXiv e-prints, page arXiv:2307.13721, July 2023

    Muhammad Awais, Muzammal Naseer, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, and Fahad Shahbaz Khan. Foundational Models Defining a New Era in Vision: A Survey and Outlook.arXiv e-prints, page arXiv:2307.13721, July 2023

  5. [5]

    Neural Machine Translation by Jointly Learning to Align and Translate.arXiv e-prints, page arXiv:1409.0473, September 2014

    Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural Machine Translation by Jointly Learning to Align and Translate.arXiv e-prints, page arXiv:1409.0473, September 2014

  6. [6]

    Neural networks and principal component analysis: Learning from examples without local minima.Neural Networks, 2(1):53–58, 1989

    Pierre Baldi and Kurt Hornik. Neural networks and principal component analysis: Learning from examples without local minima.Neural Networks, 2(1):53–58, 1989

  7. [7]

    LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics.arXiv e-prints, page arXiv:2511.08544, November 2025

    Randall Balestriero and Yann LeCun. LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics.arXiv e-prints, page arXiv:2511.08544, November 2025

  8. [8]

    Multimodal Machine Learning: A Survey and Taxonomy.arXiv e-prints, page arXiv:1705.09406, May 2017

    Tadas Baltru ˇsaitis, Chaitanya Ahuja, and Louis-Philippe Morency. Multimodal Machine Learning: A Survey and Taxonomy.arXiv e-prints, page arXiv:1705.09406, May 2017

  9. [9]

    John Wiley & Sons, 2011

    David J Bartholomew, Martin Knott, and Irini Moustaki.Latent variable models and factor analysis: A unified approach. John Wiley & Sons, 2011

  10. [10]

    Bartlett, Harry Desmond, and Pedro G

    Deaglan J. Bartlett, Harry Desmond, and Pedro G. Ferreira. Exhaustive Symbolic Regres- sion.arXiv e-prints, page arXiv:2211.11461, November 2022

  11. [11]

    Battaglia, Jessica B

    Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andrew Ballard, Justin Gilmer, George Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Langston, Chris Dyer, Nicolas Heess, Daan Wierstra...

  12. [12]

    Representation learning: A review and new perspectives.IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8):1798–1828, August 2013

    Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives.IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8):1798–1828, August 2013

  13. [13]

    Bishop.Pattern Recognition and Machine Learning

    Christopher M. Bishop.Pattern Recognition and Machine Learning. Information Science and Statistics. Springer, 2006. Foundation Models for Astrophysics 31

  14. [14]

    Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S

    Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Bryn- jolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jared Quincy Davis, Dora Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, St...

  15. [15]

    RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.arXiv e-prints, page arXiv:2307.15818, July 2023

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov...

  16. [16]

    Bronstein, Joan Bruna, Taco Cohen, and Petar Veli ˇckovi´c

    Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veli ˇckovi´c. Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges.arXiv e-prints, page arXiv:2104.13478, April 2021

  17. [17]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwi...

  18. [18]

    Deep Multimodal Representation Learning for Stellar Spectra.arXiv e-prints, page arXiv:2410.16081, October 2024

    Tobias Buck and Christian Schwarz. Deep Multimodal Representation Learning for Stellar Spectra.arXiv e-prints, page arXiv:2410.16081, October 2024

  19. [19]

    Emerging Properties in Self-Supervised Vision Transformers.arXiv e-prints, page arXiv:2104.14294, April 2021

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging Properties in Self-Supervised Vision Transformers.arXiv e-prints, page arXiv:2104.14294, April 2021

  20. [20]

    Multitask learning.Machine learning, 28(1):41–75, 1997

    Rich Caruana. Multitask learning.Machine learning, 28(1):41–75, 1997

  21. [21]

    Chameleon: Mixed-Modal Early-Fusion Foundation Models.arXiv e- prints, page arXiv:2405.09818, May 2024

    Chameleon Team. Chameleon: Mixed-Modal Early-Fusion Foundation Models.arXiv e- prints, page arXiv:2405.09818, May 2024

  22. [22]

    A Simple Framework for Contrastive Learning of Visual Representations.arXiv e-prints, page arXiv:2002.05709, February 2020

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A Simple Framework for Contrastive Learning of Visual Representations.arXiv e-prints, page arXiv:2002.05709, February 2020. 32 X. Zhao & Y .-S. Ting

  23. [23]

    Group equivariant convolutional networks

    Taco Cohen and Max Welling. Group equivariant convolutional networks. In Maria Florina Balcan and Kilian Q. Weinberger, editors,Proceedings of The 33rd International Confer- ence on Machine Learning, volume 48 ofProceedings of Machine Learning Research, pages 2990–2999, New York, New York, USA, 20–22 Jun 2016. PMLR

  24. [24]

    The frontier of simulation-based infer- ence.Proceedings of the National Academy of Science, 117(48):30055–30062, December 2020

    Kyle Cranmer, Johann Brehmer, and Gilles Louppe. The frontier of simulation-based infer- ence.Proceedings of the National Academy of Science, 117(48):30055–30062, December 2020

  25. [25]

    Interpretable Machine Learning for Science with PySR and SymbolicRe- gression.jl.arXiv e-prints, page arXiv:2305.01582, May 2023

    Miles Cranmer. Interpretable Machine Learning for Science with PySR and SymbolicRe- gression.jl.arXiv e-prints, page arXiv:2305.01582, May 2023

  26. [26]

    Discovering Symbolic Models from Deep Learning with Inductive Biases.arXiv e-prints, page arXiv:2006.11287, June 2020

    Miles Cranmer, Alvaro Sanchez-Gonzalez, Peter Battaglia, Rui Xu, Kyle Cranmer, David Spergel, and Shirley Ho. Discovering Symbolic Models from Deep Learning with Inductive Biases.arXiv e-prints, page arXiv:2006.11287, June 2020

  27. [27]

    Sparse Autoencoders Find Highly Interpretable Features in Language Models.arXiv e-prints, page arXiv:2309.08600, September 2023

    Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse Autoencoders Find Highly Interpretable Features in Language Models.arXiv e-prints, page arXiv:2309.08600, September 2023

  28. [28]

    G. Cybenko. Approximation by superpositions of a sigmoidal function.Mathematics of Control, Signals, and Systems, 2(4):303–314, December 1989

  29. [29]

    Imant Daunhawer, Alice Bizeul, Emanuele Palumbo, Alexander Marx, and Julia E. V ogt. Identifiability Results for Multimodal Contrastive Learning.arXiv e-prints, page arXiv:2303.09166, March 2023

  30. [30]

    Zico Kolter

    Filipe de Avila Belbute-Peres, Kevin Smith, Kelsey Allen, Josh Tenenbaum, and J. Zico Kolter. End-to-end differentiable physics for learning and control. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, editors,Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018

  31. [31]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.arXiv e-prints, page arXiv:1810.04805, October 2018

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.arXiv e-prints, page arXiv:1810.04805, October 2018

  32. [32]

    Donoso-Oliva, I

    C. Donoso-Oliva, I. Becker, P. Protopapas, G. Cabrera-Vives, M. Vishnu, and H. Vardhan. ASTROMER. A transformer-based embedding for the representation of light curves.Astron- omy and Astrophysics, 670:A54, February 2023

  33. [33]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.arXiv e-prints, page arXiv:2010.11929, October 2020

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.arXiv e-prints, page arXiv:2010.11929, October 2020

  34. [34]

    Understanding Emergent Abilities of Language Models from the Loss Perspective.arXiv e-prints, page arXiv:2403.15796, March 2024

    Zhengxiao Du, Aohan Zeng, Yuxiao Dong, and Jie Tang. Understanding Emergent Abilities of Language Models from the Loss Perspective.arXiv e-prints, page arXiv:2403.15796, March 2024

  35. [35]

    Scalable Pre-training of Large Autoregressive Image Models.arXiv e-prints, page arXiv:2401.08541, January 2024

    Alaaeldin El-Nouby, Michal Klein, Shuangfei Zhai, Miguel Angel Bautista, Alexander To- shev, Vaishaal Shankar, Joshua M Susskind, and Armand Joulin. Scalable Pre-training of Large Autoregressive Image Models.arXiv e-prints, page arXiv:2401.08541, January 2024

  36. [36]

    Why does unsupervised pre-training help deep learning?Journal of Machine Learning Research, 11(19):625–660, 2010

    Dumitru Erhan, Yoshua Bengio, Aaron Courville, Pierre-Antoine Manzagol, Pascal Vincent, and Samy Bengio. Why does unsupervised pre-training help deep learning?Journal of Machine Learning Research, 11(19):625–660, 2010

  37. [37]

    Mellier, Abdurro’uf, J

    Euclid Collaboration, Y . Mellier, Abdurro’uf, J. A. Acevedo Barroso, A. Ach ´ucarro, J. Adamek, R. Adam, G. E. Addison, N. Aghanim, M. Aguena, V . Ajani, Y . Akrami, A. Al- Bahlawan, A. Alavi, I. S. Albuquerque, G. Alestas, G. Alguero, A. Allaoui, S. W. Allen, V . Allevato, A. V . Alonso-Tetilla, B. Altieri, A. Alvarez-Candal, S. Alvi, A. Amara, L. Amen-...

  38. [38]

    Neocognitron: A self-organizing neural network model for a mech- anism of pattern recognition unaffected by shift in position.Biological cybernetics, 36(4):193–202, 1980

    Kunihiko Fukushima. Neocognitron: A self-organizing neural network model for a mech- anism of pattern recognition unaffected by shift in position.Biological cybernetics, 36(4):193–202, 1980

  39. [39]

    Garc ´ıa P´erez, Carlos Allende Prieto, Jon A

    Ana E. Garc ´ıa P´erez, Carlos Allende Prieto, Jon A. Holtzman, Matthew Shetrone, Szabolcs M´esz´aros, Dmitry Bizyaev, Ricardo Carrera, Katia Cunha, D. A. Garc ´ıa-Hern´andez, Jen- nifer A. Johnson, Steven R. Majewski, David L. Nidever, Ricardo P. Schiavon, Neville Shane, Verne V . Smith, Jennifer Sobeck, Nicholas Troup, Olga Zamora, David H. Wein- berg, ...

  40. [40]

    Draper, and Andrew Sheinis

    Sankalp Gilda, Yuan-Sen Ting, Kanoa Withington, Matthew Wilson, Simon Prunet, William Mahoney, Sebastien Fabbro, Stark C. Draper, and Andrew Sheinis. Astronomical Image Quality Prediction based on Environmental and Telescope Operating Conditions.arXiv e- prints, page arXiv:2011.03132, November 2020

  41. [41]

    Hamiltonian Neural Networks.arXiv e-prints, page arXiv:1906.01563, June 2019

    Sam Greydanus, Misko Dzamba, and Jason Yosinski. Hamiltonian Neural Networks.arXiv e-prints, page arXiv:1906.01563, June 2019

  42. [42]

    Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, R ´emi Munos, and Michal Valko

    Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, R ´emi Munos, and Michal Valko. Boot- strap your own latent: A new approach to self-supervised Learning.arXiv e-prints, page arXiv:200...

  43. [43]

    Mamba: Linear-Time Sequence Modeling with Selective State Spaces.arXiv e-prints, page arXiv:2312.00752, December 2023

    Albert Gu and Tri Dao. Mamba: Linear-Time Sequence Modeling with Selective State Spaces.arXiv e-prints, page arXiv:2312.00752, December 2023

  44. [44]

    Efficiently Modeling Long Sequences with Structured State Spaces.arXiv e-prints, page arXiv:2111.00396, October 2021

    Albert Gu, Karan Goel, and Christopher R ´e. Efficiently Modeling Long Sequences with Structured State Spaces.arXiv e-prints, page arXiv:2111.00396, October 2021

  45. [45]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Deli Chen, Dongjie Ji, ...

  46. [46]

    Language Models Represent Space and Time.arXiv e-prints, page arXiv:2310.02207, October 2023

    Wes Gurnee and Max Tegmark. Language Models Represent Space and Time.arXiv e-prints, page arXiv:2310.02207, October 2023

  47. [47]

    Self- supervised Representation Learning for Astronomical Images.ApJL, 911(2):L33, April 2021

    Md Abul Hayat, George Stein, Peter Harrington, Zarija Luki ´c, and Mustafa Mustafa. Self- supervised Representation Learning for Astronomical Images.ApJL, 911(2):L33, April 2021

  48. [48]

    Masked Autoencoders Are Scalable Vision Learners.arXiv e-prints, page arXiv:2111.06377, Novem- ber 2021

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked Autoencoders Are Scalable Vision Learners.arXiv e-prints, page arXiv:2111.06377, Novem- ber 2021

  49. [49]

    Momentum Contrast for Unsupervised Visual Representation Learning.arXiv e-prints, page arXiv:1911.05722, November 2019

    Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum Contrast for Unsupervised Visual Representation Learning.arXiv e-prints, page arXiv:1911.05722, November 2019

  50. [50]

    Bartlett, Nicolas Chartier, Carolina Cuesta-Lazaro, Simon Ding, Axel Lapel, Pablo Lemos, Christopher C

    Matthew Ho, Deaglan J. Bartlett, Nicolas Chartier, Carolina Cuesta-Lazaro, Simon Ding, Axel Lapel, Pablo Lemos, Christopher C. Lovell, T. Lucas Makinen, Chirag Modi, Viraj Pandya, Shivam Pandey, Lucia A. Perez, Benjamin Wandelt, and Greg L. Bryan. LtU-ILI: An All-in-One Framework for Implicit Inference in Astrophysics and Cosmology.The Open Journal of Ast...

  51. [51]

    Ming-Feng Ho, Simeon Bird, and Christian R. Shelton. Multifidelity emulation for the matter power spectrum using Gaussian processes.MNRAS, 509(2):2551–2565, January 2022

  52. [52]

    Long short-term memory.Neural Comput., 9(8):1735–1780, November 1997

    Sepp Hochreiter and J ¨urgen Schmidhuber. Long short-term memory.Neural Comput., 9(8):1735–1780, November 1997

  53. [53]

    Rae, Oriol Vinyals, and Laurent Sifre

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre...

  54. [54]

    Approximation capabilities of multilayer feedforward networks.Neural Net- works, 4(2):251–257, 1991

    Kurt Hornik. Approximation capabilities of multilayer feedforward networks.Neural Net- works, 4(2):251–257, 1991

  55. [55]

    Analysis of a complex of statistical variables into principal components

    Harold Hotelling. Analysis of a complex of statistical variables into principal components. Journal of educational psychology, 24(6):417, 1933

  56. [56]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models. arXiv e-prints, page arXiv:2106.09685, June 2021. Foundation Models for Astrophysics 35

  57. [57]

    Modality Com- petition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Prov- ably).arXiv e-prints, page arXiv:2203.12221, March 2022

    Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang, and Longbo Huang. Modality Com- petition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Prov- ably).arXiv e-prints, page arXiv:2203.12221, March 2022

  58. [58]

    Position: the platonic representation hypothesis

    Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola. Position: the platonic representation hypothesis. InProceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024

  59. [59]

    Nonlinear independent component analysis: Existence and uniqueness results.Neural networks, 12(3):429–439, 1999

    Aapo Hyv ¨arinen and Petteri Pajunen. Nonlinear independent component analysis: Existence and uniqueness results.Neural networks, 12(3):429–439, 1999

  60. [60]

    ˇZeljko Ivezi ´c, Steven M. Kahn, J. Anthony Tyson, Bob Abel, Emily Acosta, Robyn Alls- man, David Alonso, Yusra AlSayyad, Scott F. Anderson, John Andrew, James Roger P. Angel, George Z. Angeli, Reza Ansari, Pierre Antilogus, Constanza Araujo, Robert Arm- strong, Kirk T. Arndt, Pierre Astier, ´Eric Aubourg, Nicole Auza, Tim S. Axelrod, Deborah J. Bard, Je...

  61. [61]

    MIT Press, Cambridge, MA, 1997

    Frederick Jelinek.Statistical Methods for Speech Recognition. MIT Press, Cambridge, MA, 1997

  62. [62]

    Bronstein, and Hagai B

    Ilay Kamai, Alex M. Bronstein, and Hagai B. Perets. Machine Learning Inference of Stellar Properties Using Integrated Photometric and Spectroscopic Data.ApJ, 994(1):110, Novem- ber 2025. 36 X. Zhao & Y .-S. Ting

  63. [63]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling Laws for Neural Language Models.arXiv e-prints, page arXiv:2001.08361, January 2020

  64. [64]

    Physics-informed machine learning.Nature Reviews Physics, 3(6):422–440, 2021

    George Em Karniadakis, Ioannis G Kevrekidis, Lu Lu, Paris Perdikaris, Sifan Wang, and Liu Yang. Physics-informed machine learning.Nature Reviews Physics, 3(6):422–440, 2021

  65. [65]

    Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics.arXiv e-prints, page arXiv:1705.07115, May 2017

    Alex Kendall, Yarin Gal, and Roberto Cipolla. Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics.arXiv e-prints, page arXiv:1705.07115, May 2017

  66. [66]

    OmniSpectra: A Unified Foundation Model for Native Resolution Astronomical Spectra.arXiv e-prints, page arXiv:2601.15351, January 2026

    Md Khairul Islam and Judy Fox. OmniSpectra: A Unified Foundation Model for Native Resolution Astronomical Spectra.arXiv e-prints, page arXiv:2601.15351, January 2026

  67. [67]

    OpenVLA: An Open-Source Vision-Language-Action Model.arXiv e-prints, page arXiv:2406.09246, June 2024

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. OpenVLA: An Open-Source Vision-Language-Action Model.arXiv e-prints, page arXiv:2406.09...

  68. [68]

    Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment Anything.arXiv e-prints, page arXiv:2304.02643, April 2023

  69. [69]

    SpectraFM: Tuning into Stellar Foundation Models.arXiv e-prints, page arXiv:2411.04750, November 2024

    Nolan Koblischke and Jo Bovy. SpectraFM: Tuning into Stellar Foundation Models.arXiv e-prints, page arXiv:2411.04750, November 2024

  70. [70]

    Simon Kornblith, Jonathon Shlens, and Quoc V . Le. Do Better ImageNet Models Transfer Better?arXiv e-prints, page arXiv:1805.08974, May 2018

  71. [71]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks.Commun. ACM, 60(6):84–90, May 2017

  72. [72]

    Synthe spectrum synthesis programs and line data.Kurucz CD-Rom, 1993

    Robert L Kurucz. Synthe spectrum synthesis programs and line data.Kurucz CD-Rom, 1993

  73. [73]

    Lastufka, O

    E. Lastufka, O. Bait, M. Drozdova, V . Kinakh, D. Piras, M. Audard, M. Dessauges-Zavadsky, T. Holotyak, D. Schaerer, and S. V oloshynovskiy. Examining vision foundation models for classification and detection in optical and radio astronomy.A&A, 703:A217, November 2025

  74. [74]

    Towards Understanding Inductive Bias in Trans- formers: A View From Infinity.arXiv e-prints, page arXiv:2402.05173, February 2024

    Itay Lavie, Guy Gur-Ari, and Zohar Ringel. Towards Understanding Inductive Bias in Trans- formers: A View From Infinity.arXiv e-prints, page arXiv:2402.05173, February 2024

  75. [75]

    Lecun, L

    Y . Lecun, L. Bottou, Y . Bengio, and P. Haffner. Gradient-based learning applied to document recognition.Proceedings of the IEEE, 86(11):2278–2324, 1998

  76. [76]

    A path towards autonomous machine intelligence

    Yann LeCun. A path towards autonomous machine intelligence. OpenReview, 2022

  77. [77]

    Rediscover- ing orbital mechanics with machine learning.Machine Learning: Science and Technology, 4(4):045002, December 2023

    Pablo Lemos, Niall Jeffrey, Miles Cranmer, Shirley Ho, and Peter Battaglia. Rediscover- ing orbital mechanics with machine learning.Machine Learning: Science and Technology, 4(4):045002, December 2023

  78. [78]

    Leung and Jo Bovy

    Henry W. Leung and Jo Bovy. Towards an astronomical foundation model for stars with a transformer-based model.MNRAS, 527(1):1494–1520, January 2024

  79. [79]

    Jiadong Li, Mingjie Jian, Yuan-Sen Ting, and Gregory M. Green. Differentiable Stellar At- mospheres with Physics-Informed Neural Networks.arXiv e-prints, page arXiv:2507.06357, July 2025

  80. [80]

    Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning.arXiv e-prints, page arXiv:2203.02053, March 2022

    Weixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung, and James Zou. Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning.arXiv e-prints, page arXiv:2203.02053, March 2022

Showing first 80 references.