{"id":"097615f2-981a-4c65-a92d-a13f87c7003c","arxiv_id":"2607.20797","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Contrastive alignment of SPHEREx spectra with DESI Legacy images improves star-galaxy separation, lifting image-only stellar purity from 76% to 93% and forecasting sub-percent SPHEREx contamination at p>0.7.","lead":"Using paired SPHEREx spectra and DESI Legacy images, this paper shows that contrastive alignment makes star-galaxy separation markedly better, especially for images. It then forecasts that a stricter classification threshold keeps SPHEREx stellar contamination below one percent over most of the sky.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Synthetic-only validation: the sub-percent contamination forecast is an idealized bound; real-data test needed to confirm transfer.","rationale":"The reader's weakest_assumption identifies the synthetic nature of the evaluation as load-bearing. My analysis agrees and further pinpoints two specific mechanisms: (i) the blank-sky injection avoids blending, which is acknowledged in §5.4 but not quantified; (ii) the photo-z selection in §5.1 uses templates matching those used to generate the galaxy SEDs, potentially over-suppressing the stellar tail and making the σ_z/(1+z)<0.2 cut appear more effective than in real data. These factors directly affect the headline sub-percent contamination forecast. However, the paper's internal experiments are internally consistent, the ablation studies support the relative claim that alignment improves linear separability, and the authors clearly identify the limitation in §5.4 and the need for real-data validation in §6. Thus the appropriate verdict remains CONDITIONAL: the central claims are plausible but require end-to-end validation on real SPHEREx spectra and DESI-LS sources before the sub-percent forecast can be accepted. The concrete test I propose is the decisive experiment: run the frozen pipeline on real labeled data and measure contamination. This would settle whether the synthetic mocks are more separable than reality. I do not find grounds to reject the paper; the concern is about generalization, not internal inconsistency.","tokens_in":22260,"tokens_out":5909,"duration_ms":66639,"concrete_test":"Take a sample of DESI-LS sources in the COSMOS field that have real SPHEREx spectra (from early survey data), with labels independently established by Gaia DR3 astrometry for stars and spectroscopic/photometric redshifts from COSMOS2020 or DESI for galaxies. Without retraining any encoder or classifier, apply the aligned spectrum/image embeddings and XGBoost models, then select p_gal>0.7 and σ_z/(1+z)<0.2 as in §5.2, and measure the contamination fraction. Compare to the 0.4% median forecast from Fig. 12. If the measured contamination is >1%, the sub-percent claim is not supported; if it remains <1%, the forecast is plausibly robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the synthetic evaluation pipeline is representative of real SPHEREx/DESI-LS data. Every quantitative headline—the image-based classification gains (Table 1) and the sub-percent contamination forecast (§5.3)—is measured on mocks: galaxies are template SEDs from COSMOS photometry (Feder et al. 2023), and stars are BaSeL spectra injected at blank-sky positions at least 1.5'' from any source (§2.2). Three effects could make these mocks more separable than reality. (1) Blending: real faint sources overlap with neighbors; injected stars deliberately avoid this. §5.4 acknowledges this but does not quantify the impact. (2) Photo-z circularity: the σ_z/(1+z)<0.2 selection in §5.1 is applied using a template-fitting code on spectra generated from the same template grid used to create the galaxies. This can artificially suppress the stellar photo-z tail, which is the primary lever that reduces contamination in the forecast. (3) The stellar-density correction is hand-calibrated (∼50% correction near COSMOS, §4.2.1) and then extrapolated across the sky via Gaia, which may not trace the star types that leak into galaxy samples. Since the central claim is a forecast of a few-tenths-percent contamination, a factor-of-2 error in the stellar probability tail materially changes the conclusion. The paper is explicit about these caveats (§5.4) and does not claim real-data validation, but the abstract statement ('can be controlled at the sub-percent level') is only supported by an idealized upper bound, not by observed data.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a CLIP-style contrastive alignment between DESI Legacy Survey images and synthetic SPHEREx spectrophotometry, with the goal of improving star–galaxy separation for the SPHEREx galaxy survey. Using galaxy spectra generated from COSMOS template fits and synthetic stars injected into DESI-LS cutouts, the authors find that aligned embeddings outperform unaligned ones, especially for image-based classification: Table 1 reports galaxy purity 0.9908 and galaxy completeness 0.9614 for aligned image embeddings versus 0.9810 and 0.8349 unaligned. The paper also shows that aligned embeddings are more linearly separable (Section 4.2.3), and it uses photo-z precision cuts to forecast stellar contamination over the full SPHEREx footprint, concluding that sub-percent contamination is achievable with p_gal>0.7 (Section 5.3). The evaluation is entirely synthetic, with caveats acknowledged in Section 5.4.","tokens_in":22636,"tokens_out":3866,"duration_ms":41263,"significance":"If the results transfer to real data, the paper would provide a practical method for improving star–galaxy separation in SPHEREx and Rubin LSST, directly relevant to f_NL science goals. The study is carefully designed: it uses a clean paired image–spectrum construction, compares aligned versus unaligned embeddings under controlled conditions, and includes classifier ablation tests that strengthen the interpretation that alignment restructures the embedding space. The use of publicly available synthetic spectral data (Feder et al. 2023 on Zenodo) and standard open-source tools supports reproducibility. However, the central quantitative claims—the gain in image-based classification and the sub-percent contamination forecast—rest on the realism of the mocks and on the photo-z selection; both need stronger support before the abstract-level claims can be accepted.","major_comments":[{"comment":"The sub-percent contamination forecast is derived entirely from synthetic data: galaxies are template SEDs from Feder et al. (2023) and stars are injected at blank-sky positions with a minimum separation of 1.5″ (Section 2.2.3). The caveats in Section 5.4 are qualitative; no quantitative estimate is given for how source blending, spatial noise inhomogeneity, or PSF errors would change η_star. Since the abstract states that contamination 'can be controlled at the sub-percent level across most of the extragalactic sky,' the authors should either validate with real SPHEREx/DESI-LS data (e.g., using Gaia/Euclid cross-checks) or provide a mock-degradation study that quantifies the impact of these effects. Without this, the abstract overstates the strength of the forecast.","section":"§5.3, Fig. 12"},{"comment":"The photo-z precision cut σ_z/(1+z)<0.2 is computed with a template-fitting code whose model grid matches the same template grid used to generate the synthetic galaxy SEDs. This creates a circularity: stars whose spectra resemble template galaxies at moderate redshift may be assigned large photo-z errors and removed preferentially, which is the main lever that suppresses contamination in the forecast. Please test sensitivity to an independent photo-z method (e.g., a different template set, a neural photo-z estimator, or real stellar SEDs) and report how the contamination fractions in §5.3 change. This is a load-bearing point because a factor-of-two change in the high-redshift stellar tail would materially alter the conclusion.","section":"§5.1, Fig. 10"},{"comment":"The classification metrics are reported without error bars, even though the 5-fold cross-validation procedure yields a distribution. The image-based improvement claims rely on differences of ~17 percentage points in stellar purity and ~13 percentage points in galaxy completeness; reporting fold-to-fold scatter or bootstrap uncertainties is necessary to confirm that the improvement is not driven by a particular split. The same applies to the AUC values in Figure 8 and the completeness/purity curves in Figures 6 and 11.","section":"Table 1, §4.2.2"},{"comment":"The stellar-density correction is hand-calibrated to a ~50% overprediction near COSMOS and then extrapolated to the full sky using Gaia star counts. This correction is applied as a fixed factor, and its uncertainty is not propagated into the contamination forecast. Moreover, as the authors note in Section 5.4, Gaia may not trace the same stellar populations that leak into galaxy samples. Please provide a sensitivity analysis—e.g., varying the correction factor by ±50%—and show how the median and the 90/99th-percentile contamination fractions in Figure 12 change.","section":"§4.2.1 and §5.3"}],"minor_comments":[{"comment":"The subsample R² values are computed on the set of sources that are misclassified before alignment and correctly classified after. The sample size and the restricted range of this subset should be stated; R² on small, selected subsets can be noisy and should be interpreted with caution.","section":"§4.2.4, Table 2"},{"comment":"The sentence 'η_star has a median of 1.3%, with 90% of the footprint having η_star<2.2% for 90% of HEALPix tiles and <3.6% for 99% of tiles' is confusing. Please rephrase, e.g., 'the median contamination fraction is 1.3%; 90% of the footprint has η_star<2.2% and 99% has η_star<3.6%.'","section":"§5.3"},{"comment":"The text describes the 102-channel SPHEREx data as 'photometry' in the sentence 'mapping the 102-channel photometry to a 6-dimensional latent space.' These are low-resolution spectral channels, not broadband photometry; please adjust the terminology for clarity.","section":"§3.2"},{"comment":"The caption contains a typo: 'In the top bottom' should read 'In the bottom row.'","section":"Figure 3"},{"comment":"The definition of the blank-sky injection separation θ_min=1.5″ is clear, but the paper does not state how many injected stars were rejected for lack of valid blank-sky positions. A sentence on the success rate would help assess potential selection biases.","section":"§2.2.3"}],"recommendation":"major_revision","confidential_remarks":"This is a solid methods paper with a clean internal comparison, but the abstract's sub-percent contamination claim is a forecast from synthetic data. The authors are transparent about the caveats, but the photo-z circularity and the unquantified mock-fidelity issues are load-bearing. I recommend major revision rather than rejection because the concerns are addressable with a degraded-mock or real-data test on one hand, or a more cautious framing on the other. No issue with self-citation or novelty; the core methodological question is well posed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the Nguyen et al. paper. My take: the core comparison is clean and the central empirical claim holds on its own terms — contrastive alignment of SPHEREx spectrophotometry with DESI-LS imaging substantially improves image-based star/galaxy separation, and the aligned embeddings are much less sensitive to classifier choice. The authors do a nice job connecting the image gains to a mechanism: aligned image embeddings are better organized along redshift and IR spectral features, which matters for the sources that were previously misclassified. That is genuinely new relative to AstroCLIP/SpecCLIP, which established contrastive alignment but not this specific application or the classifier-robustness result.\n\nThe soft spots are real and they are exactly where the reader's report says. Everything is synthetic. The galaxies are COSMOS templates convolved with SPHEREx bandpasses, the stars are BaSeL SEDs injected at blank-sky positions that avoid blending, and the headline sub-percent contamination forecast in Section 5.3 is an extrapolation from that mock pipeline. The paper is explicit about these caveats in Section 5.4, which I give it credit for. But the abstract's 'can be controlled at the sub-percent level' overstates what is shown; it should read as an idealized upper bound.\n\nOne subtle worry that I think is worth flagging to the referee: the photo-z selection in Section 5.1 uses template-fitting code with a grid matching the Feder et al. synthetic spectra, and the galaxies are generated from that same grid. That creates potential circularity in the stellar high-redshift tail, which is exactly the population that drives the contamination forecast. A factor-of-two error in that tail would change the headline. This is not easy to fix with the current data, but it should be acknowledged more explicitly than the current text does.\n\nAlso, Table 1 has no error bars, and the stellar density correction is hand-calibrated and extrapolated via Gaia, which may not trace the faint stars that leak into the galaxy samples. These are all minor if the paper stays in the 'methodology on mocks' lane; they become important if the abstract is read as a prediction for real SPHEREx operations.\n\nWho gets value: anyone working on SPHEREx f_NL systematics, and people doing multimodal foundation-model applications in astronomy. It deserves a serious referee — the experiments are reproducible in principle, the writing is clear, and the caveats are mostly in the right place. My recommendation: send it to peer review, ask for the abstract to match the evidence, and ideally for a small validation on real Gaia/DESI-LS sources, even if real SPHEREx spectra are not yet available.","headline":"A solid synthetic proof-of-concept that alignment helps star/galaxy separation for SPHEREx-like data, but the sub-percent contamination forecast is an idealized bound from mocks, not a measured result.","tokens_in":23146,"tokens_out":2756,"would_cite":true,"duration_ms":26123,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Contrastively aligning SPHEREx spectra with DESI-LS images substantially improves star–galaxy separation and predicts sub-percent stellar contamination over most of the extragalactic sky.","keywords":["star–galaxy separation","contrastive learning","SPHEREx","DESI Legacy Surveys","multimodal embeddings","stellar contamination","large-scale structure","photometric classification"],"falsifier":"Run the same aligned-image classifier on real SPHEREx spectra and DESI-LS cutouts with labels from Gaia and Euclid across 18<z_AB<22.5 and measure stellar purity at p=0.5; if it falls from the predicted ~0.93 toward the unaligned ~0.76, or if the median footprint contamination at p>0.7 exceeds 1%, the paper's forecast is refuted.","tokens_in":22144,"feed_emoji":"🌌","tokens_out":5819,"duration_ms":48435,"temperature":0.7,"pith_summary":"Galaxy clustering analyses aiming at sigma(f_NL) ~ 1 need galaxy samples contaminated by fewer than one star per hundred; stars trace Galactic structure and can fake cosmological power. This paper asks whether forcing SPHEREx near-infrared spectrophotometry and DESI Legacy Survey images into a shared embedding, via contrastive learning, makes stars and galaxies easier to tell apart. It finds yes: aligned image embeddings raise stellar purity from 0.756 to 0.931 and galaxy completeness from 0.835 to 0.961 at a probability threshold of 0.5, and aligned spaces stay nearly as accurate under simple classifiers. Extrapolating to the full SPHEREx footprint with a photo-z precision cut, it forecasts a median stellar contamination of 0.4% when requiring galaxy probability above 0.7. The gain seems to come from alignment reorganizing image embeddings so redshift and near-infrared spectral shape become linearly accessible.","feed_headline":"Star contamination forecast below 1% after spectrum–image alignment","feed_subtitle":"Pairing SPHEREx spectra with optical imaging makes galaxy samples clean enough for the survey's primordial non-Gaussianity science.","key_machinery":"A CLIP-style contrastive alignment: an InfoNCE loss with cosine similarity pulls image and spectrum embeddings of the same source together and pushes different sources apart in a shared latent space. Images are encoded by a pre-trained transformer and spectra by a 1D convolutional autoencoder, with multi-head cross-attention projection heads generating 1024-dimensional embeddings; encoders are frozen, only the projection heads train. Downstream, XGBoost classifiers operate on these embeddings, and linear probes measure how accessible redshift and spectral fluxes are. The alignment is what reorganizes the embedding geometry; the paper's key evidence is that this reorganization, not new inform","core_discovery":"The paper establishes that multimodal contrastive alignment reorganizes the embedding space along dimensions better suited to source classification. Using a CLIP-style InfoNCE objective to align a pre-trained image encoder with a spectrum encoder, the authors obtain shared embeddings on which a single XGBoost classifier separates stars from galaxies. On the magnitude-limited COSMOS-like sample, the aligned image embeddings are the strongest representation: stellar purity 0.9308 vs 0.7559 unaligned and galaxy completeness 0.9614 vs 0.8349 at p_gal=0.5, with AUC above 0.992 for every classifier tested. The aligned space is also markedly more linearly separable, with logistic regression within","pith_inferences":["The same contrastive recipe could be applied to other imaging surveys paired with any spectrophotometric catalog; the gain should be largest wherever morphology and broadband colors alone are degenerate, since alignment imports spectral discriminants into the image embedding.","Because alignment mostly exposes information already in the imaging, it is a cheaper alternative to adding new bands: no new observations are needed, only a paired spectral catalog at train time. A testable prediction is that gains shrink as survey depth or PSF quality degrades.","The mechanism suggests a broader design principle: contrastive alignment can act as a prior that linearizes classification boundaries, so for any high-dimensional astronomical classification task with paired modalities, aligned embeddings should outperform raw embeddings under simple models.","A direct test would be to check whether the 1.6-micron H- opacity minimum, fixed in observed wavelength for stars and redshifted for galaxies, is what the aligned image embeddings encode; if so, galaxy sub-samples selected by aligned-image probability should have redshift distributions matching those implied by that feature."],"forward_implications":["A DESI-LS image-only classifier, after alignment, reaches stellar purity 0.931 and galaxy completeness 0.961 on the z<22.5 mock sample, making imaging a much stronger star–galaxy separator than raw spectra alone.","Adopting p_gal>0.7 with a sigma_z/(1+z)<0.2 cut yields a predicted median stellar contamination of 0.4% over the SPHEREx extragalactic footprint, with 90% of tiles below 0.9%, meeting sub-percent requirements for sigma(f_NL)~O(1) analyses.","Aligned embeddings degrade little when the classifier is simplified: a logistic regression stays within about 0.003 AUC of the full XGBoost model, implying contamination control will be less sensitive to classifier complexity and likely to systematic perturbations.","Alignment disproportionately fixes the hardest subpopulations: galaxies at z~0.7–0.9 with 1.4<r-z<2.2, which overlap the stellar locus, are exactly the ones whose image embeddings become linearly separable.","Completeness losses from aggressive purity cuts concentrate at low redshift, where SPHEREx samples remain sample-variance limited, so the purity/completeness tradeoff is affordable for clustering analyses."],"fun_headline_variants":["AI alignment slashes star contamination to sub-1%","Contrastive learning pairs spectra and images to purify galaxy samples","Multimodal alignment beats single-modality for star–galaxy separation","SPHEREx spectra plus imaging cut star contamination to sub-1%","Multimodal alignment reduces stellar contamination to sub-percent"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the mock SPHEREx spectra and injected stars are as hard to separate as real data; if real spectra, source blending, PSF errors, or noise inhomogeneity make stars and galaxies less separable, the measured gains and the sub-percent contamination forecast will not hold.","fun_headline_variants_meta":{"raw":{"variants":["AI alignment slashes star contamination to sub-1%","Contrastive learning pairs spectra and images to purify galaxy samples","Multimodal alignment beats single-modality for star–galaxy separation","SPHEREx spectra plus imaging cut star contamination to sub-1%","Multimodal alignment reduces stellar contamination to sub-percent"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000723,"raw_usage":{"total_tokens":3114,"prompt_tokens":810,"completion_tokens":2304,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":2216}},"tokens_in":554,"tokens_out":2304,"duration_ms":13880,"temperature":1.0,"reasoning_tokens":2216,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:20:19.167741+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same aligned-image classifier on real SPHEREx spectra and DESI-LS cutouts with labels from Gaia and Euclid across 18<z_AB<22.5 and measure stellar purity at p=0.5; if it falls from the predicted ~0.93 toward the unaligned ~0.76, or if the median footprint contamination at p>0.7 exceeds 1%, the paper's forecast is refuted.","supporting_citations":[],"review_version":1}