Pith. sign in

REVIEW 5 cited by

Scaling Laws for Galaxy Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.02973 v1 pith:2CM7Y7QZ submitted 2024-04-03 cs.CV astro-ph.GA

Scaling Laws for Galaxy Images

classification cs.CV astro-ph.GA
keywords imagesgalaxydownstreamscalingperformancetasksachieveacross
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We present the first systematic investigation of supervised scaling laws outside of an ImageNet-like context - on images of galaxies. We use 840k galaxy images and over 100M annotations by Galaxy Zoo volunteers, comparable in scale to Imagenet-1K. We find that adding annotated galaxy images provides a power law improvement in performance across all architectures and all tasks, while adding trainable parameters is effective only for some (typically more subjectively challenging) tasks. We then compare the downstream performance of finetuned models pretrained on either ImageNet-12k alone vs. additionally pretrained on our galaxy images. We achieve an average relative error rate reduction of 31% across 5 downstream tasks of scientific interest. Our finetuned models are more label-efficient and, unlike their ImageNet-12k-pretrained equivalents, often achieve linear transfer performance equal to that of end-to-end finetuning. We find relatively modest additional downstream benefits from scaling model size, implying that scaling alone is not sufficient to address our domain gap, and suggest that practitioners with qualitatively different images might benefit more from in-domain adaption followed by targeted downstream labelling.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Star-forming clump detection in nearby galaxies using Faster R-CNN and $ugrizy$ imaging data from CLAUDS and HSC-SSP

    astro-ph.IM 2026-07 conditional novelty 6.5

    A multi-band Faster R-CNN with Zoobot backbone detects star-forming clumps in low-z galaxies at ≥0.9 completeness and ≥0.8 purity on simulated injections, yielding ~1.5M candidates.

  2. Star-forming clump detection in nearby galaxies using Faster R-CNN and $ugrizy$ imaging data from CLAUDS and HSC-SSP

    astro-ph.IM 2026-07 conditional novelty 6.0

    A six-band Faster R-CNN with the Zoobot backbone detects star-forming clump candidates in ~700,000 local galaxies, claiming ~90% completeness and ~80% purity for clumps brighter than the surveys' detection limits.

  3. The Edge-on Galaxies in the DESI survey (EGIDE): sample building and photometry

    astro-ph.GA 2026-06 unverdicted novelty 6.0

    The EGIDE project releases a tenfold larger catalogue of edge-on galaxies with griz photometry, stellar masses, redshifts and star formation rates, finding that red-sequence galaxies are thicker than blue-cloud ones a...

  4. The Edge-on Galaxies in the DESI survey (EGIDE): sample building and photometry

    astro-ph.GA 2026-06 conditional novelty 6.0

    EGIDE provides 149,215 edge-on galaxy candidates from DESI DR10 with homogeneous photometry, masses, and redshifts, ten times larger than EGIPS.

  5. Leveraging Multimodality for Real-Time Classification of Transients and Variables found by the Zwicky Transient Facility

    astro-ph.IM 2026-06 unverdicted novelty 5.0

    ORACLE-2 multimodal classifiers raise macro F1 from 0.52-0.66 (light-curve only) to 0.73 on ZTF Bright Transient Survey data and reach 0.88 on simulated ELAsTiCC data.