{"id":"ae04596f-c932-46eb-9595-f65b560a79f0","arxiv_id":"2606.07771","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Conformal methods achieve near-nominal 90% coverage on galaxy property regression with AION-1 embeddings while LVD additionally delivers local validity, outperforming Deep Ensembles and MC Dropout.","lead":"This paper benchmarks seven uncertainty quantification methods for predicting galaxy properties like redshift and stellar mass from AION-1 foundation model embeddings. A smart generalist might read it to learn which methods give reliable error bars for scientific use in astronomy surveys.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Coverage and local validity measured only w.r.t. PROVABGS-derived labels treated as exact ground truth","rationale":"This is precisely the reader's weakest_assumption and is load-bearing because the entire evaluation pipeline (including the distinction between marginal and local validity) is conditioned on treating the derived labels as noiseless targets. The data-split exchangeability concern is secondary once label noise is acknowledged. No other internal inconsistency is visible from the abstract and claim description.","tokens_in":1780,"tokens_out":348,"duration_ms":16863,"concrete_test":"Recompute all coverage and local-validity metrics after adding zero-mean Gaussian noise to each PROVABGS label with variance matching the reported PROVABGS posterior widths; if LVD's local-validity advantage or the ~1 pp marginal coverage disappears for any property, the headline claim is label-noise dependent.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central empirical claim is that conformal methods (esp. LVD on AION-1 embeddings) achieve marginal coverage within ~1 pp of 90% and finite-sample local validity, while baselines do not. All reported coverages and validity statements are computed against PROVABGS labels, which are themselves posterior summaries from a separate SED-fitting pipeline subject to modeling assumptions, parameter degeneracies, and noise. Conformal guarantees are with respect to the observed label distribution; if label errors are non-negligible or vary with galaxy type/redshift, the intervals cover the noisy labels but need not cover the underlying physical quantities at the stated rates. The paper provides no propagation of PROVABGS uncertainties into the UQ benchmark or sensitivity analysis on label noise.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript benchmarks seven uncertainty quantification methods for regressing galaxy properties (redshift, stellar mass, age, metallicity, sSFR) from Legacy Survey photometry/DESI spectra using frozen AION-1 foundation-model embeddings and PROVABGS-derived labels. It claims that distribution-free conformal methods achieve marginal coverage within ~1 pp of the nominal 90% level across properties, that non-conformal baselines (Deep Ensembles, MC Dropout) fail to calibrate, that CQR performs best in poorly predicted bins, and that only the Locally Valid and Discriminative (LVD) framework—especially on AION-1 embeddings—delivers finite-sample local validity in addition to marginal coverage.","tokens_in":1933,"tokens_out":658,"duration_ms":26117,"significance":"If the central empirical claims hold after addressing label noise, the work would be significant for astro-ph.IM: it supplies a concrete, multi-property comparison of UQ methods on foundation-model embeddings and identifies LVD as providing both marginal and local validity guarantees. The reproducible experimental setup on public survey data and the explicit contrast between marginal and local validity constitute strengths that could guide adoption of conformal methods for uncertainty-aware downstream inference.","major_comments":[{"comment":"The evaluation computes all coverage and local-validity statistics against PROVABGS-derived labels treated as exact ground truth (Abstract and Results). These labels are themselves posterior summaries from an SED-fitting pipeline subject to modeling assumptions, parameter degeneracies, and noise; no propagation of PROVABGS uncertainties or sensitivity analysis on label noise is reported. Because conformal guarantees are with respect to the observed label distribution, this directly affects whether the reported intervals can be interpreted as reliable for the underlying physical quantities, which is load-bearing for the claim that LVD provides “uncertainty-aware inference” suitable for scientific use.","section":"Abstract and Results"},{"comment":"The abstract states that LVD “provides finite-sample local validity” when operating on AION-1 embeddings, yet the manuscript supplies no explicit definition of the local-validity metric, the binning or conditioning procedure used to verify it, or the precise implementation details that distinguish it from standard conformal methods. Without these, it is impossible to confirm that the reported local-validity advantage is not an artifact of the chosen evaluation protocol or data splits.","section":"Abstract and Methods"}],"minor_comments":[{"comment":"The abstract refers to “seven UQ methods” and “the bin with the poorest model predictions” without naming the methods or defining the binning criterion; these should be stated explicitly in the opening paragraph for clarity.","section":"Abstract"},{"comment":"Notation for the seven methods (Deep Ensembles, MC Dropout, CQR, LVD, etc.) should be introduced consistently in a table or methods subsection rather than only in the abstract.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a methods paper in astro-ph.IM and appears to fit the journal scope, but the citation list should be checked for completeness on prior conformal-prediction applications in astronomy."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments, which help clarify the interpretation of our results and the presentation of the LVD method. We address each point below and will revise the manuscript accordingly.","responses":[{"response":"We agree that PROVABGS labels are subject to SED-fitting uncertainties and that conformal coverage is formally with respect to the observed label distribution. In the revised manuscript we will add a sensitivity analysis that perturbs the labels by draws from their reported posterior uncertainties, recomputes coverage and local-validity statistics, and discusses the distinction between coverage w.r.t. the observed labels versus the underlying physical quantities. This will be presented in a new subsection of Results.","revision_made":"yes","referee_comment":"[Abstract and Results] The evaluation computes all coverage and local-validity statistics against PROVABGS-derived labels treated as exact ground truth (Abstract and Results). These labels are themselves posterior summaries from an SED-fitting pipeline subject to modeling assumptions, parameter degeneracies, and noise; no propagation of PROVABGS uncertainties or sensitivity analysis on label noise is reported. Because conformal guarantees are with respect to the observed label distribution, this directly affects whether the reported intervals can be interpreted as reliable for the underlying physical quantities, which is load-bearing for the claim that LVD provides “uncertainty-aware inference” suitable for scientific use."},{"response":"We will expand the Methods section with an explicit definition of the local-validity metric (empirical coverage conditioned on local difficulty), the binning/conditioning procedure (quantiles of absolute residual or embedding-nearest-neighbor distance), and the precise algorithmic differences between LVD and standard conformal methods (including pseudocode). These additions will make the local-validity claims fully reproducible and allow direct verification against the evaluation protocol.","revision_made":"yes","referee_comment":"[Abstract and Methods] The abstract states that LVD “provides finite-sample local validity” when operating on AION-1 embeddings, yet the manuscript supplies no explicit definition of the local-validity metric, the binning or conditioning procedure used to verify it, or the precise implementation details that distinguish it from standard conformal methods. Without these, it is impossible to confirm that the reported local-validity advantage is not an artifact of the chosen evaluation protocol or data splits."}],"tokens_in":1525,"tokens_out":496,"duration_ms":12952,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper benchmarks seven UQ approaches on frozen AION-1 embeddings for five galaxy properties and reports that distribution-free conformal predictors reach marginal coverage near the nominal 90 percent while Deep Ensembles and MC Dropout do not. Among the conformal options, LVD stands out for also delivering finite-sample local validity that adapts to each galaxy.\n\nWhat is new is the specific application to AION-1 embeddings rather than any new method. The comparison is empirical and uses real survey data plus PROVABGS labels, which is the right setting for this kind of work.\n\nThe results look solid on their own terms: conformal methods calibrate where the baselines do not, and LVD improves on standard conformal quantile regression for the hardest bins. That part of the story is useful for anyone already working with foundation-model embeddings in astrophysics.\n\nThe main limitation is the treatment of the labels. PROVABGS values come from a separate SED-fitting pipeline that carries its own modeling assumptions and degeneracies. The coverage and local-validity numbers are computed against those derived labels, not against independent truth. If label noise varies with redshift or galaxy type, the reported guarantees apply only to the noisy targets. The paper does not appear to propagate those uncertainties or run sensitivity checks, so the practical strength of the local-validity claim is harder to judge.\n\nThis is the kind of targeted benchmark that belongs in the astro foundation-model literature. Readers who care about uncertainty reporting on large surveys will get something concrete from it. It deserves a serious referee because the experimental setup is reproducible and the central comparison is falsifiable, even though the label-noise issue will need attention in revision.","headline":"Conformal methods especially LVD give better calibrated intervals than ensembles or dropout on AION-1 embeddings, but the whole comparison treats PROVABGS labels as exact ground truth.","tokens_in":2431,"tokens_out":416,"would_cite":false,"duration_ms":10608,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Conformal prediction methods achieve reliable 90% coverage and local validity for galaxy property estimates from foundation model embeddings, while standard baselines do not.","keywords":["conformal prediction","uncertainty quantification","galaxy properties","foundation models","local validity","regression calibration","astrophysics"],"falsifier":"A new collection of galaxies drawn from the same distribution where the conformal prediction intervals cover the true property values at a rate materially below the nominal 90% would falsify the coverage claim.","tokens_in":2692,"feed_emoji":"📊","tokens_out":680,"duration_ms":12922,"temperature":0.7,"pith_summary":"The paper benchmarks seven uncertainty quantification approaches on regression tasks that predict galaxy redshift, stellar mass, age, metallicity, and star-formation rate from Legacy Survey photometry and DESI spectra, using frozen embeddings from an astronomical foundation model. Distribution-free conformal techniques reach marginal coverage within about one percentage point of the nominal 90% target across all five properties. Non-conformal methods such as deep ensembles and Monte Carlo dropout do not calibrate reliably. Only the Locally Valid and Discriminative framework, especially when run on the foundation-model embeddings, also supplies finite-sample local validity that adapts interval width to each galaxy's individual prediction difficulty rather than relying solely on marginal guarantees.","feed_headline":"Conformal methods hit 90% coverage on galaxy properties from embeddings","feed_subtitle":"Non-conformal baselines fail to calibrate while LVD supplies local validity that adapts to each galaxy.","key_machinery":"The Locally Valid and Discriminative (LVD) framework, which supplies finite-sample local validity guarantees when applied to foundation-model embeddings for regression tasks.","core_discovery":"Distribution-free conformal methods achieve marginal coverage within ∼1 pp of the nominal 90% across all properties, while non-conformal baselines fail to calibrate reliably. Among conformal approaches, Conformalized Quantile Regression delivers the best coverage in the bin with the poorest model predictions. Only the Locally Valid and Discriminative framework—particularly when operating on the foundation-model embeddings—also provides finite-sample local validity, producing intervals that adapt to each galaxy's local prediction difficulty.","pith_inferences":["The same conformal workflow could be tested on embeddings from other astronomical foundation models to check whether local validity transfers.","Local validity might reduce systematic errors in downstream analyses that combine many galaxy property estimates, such as population studies.","If the method is applied to new photometric surveys, the coverage guarantees would still hold provided the exchangeability assumption between calibration and test sets remains reasonable."],"forward_implications":["Conformalized Quantile Regression yields the tightest reliable intervals in the regions where point predictions are weakest.","Locally Valid and Discriminative intervals adapt their width to each galaxy's individual prediction difficulty rather than using a single marginal width.","Conformal prediction becomes the preferred uncertainty framework for downstream inference that uses foundation-model embeddings in astrophysics.","Local validity guarantees remain available even when the underlying point predictor is a frozen foundation model."],"fun_headline_variants":["Conformal methods achieve 90% coverage on AION-1 galaxy properties","LVD provides local validity on AION-1 embeddings","Non-conformal baselines fail to calibrate UQ reliably","Conformalized quantile regression covers in poor prediction bins","Local validity only from LVD on foundation model embeddings"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The evaluation treats the derived labels as accurate ground truth and assumes the chosen data splits allow the conformal coverage and local validity guarantees to hold.","fun_headline_variants_meta":{"raw":{"variants":["Conformal methods achieve 90% coverage on AION-1 galaxy properties","LVD provides local validity on AION-1 embeddings","Non-conformal baselines fail to calibrate UQ reliably","Conformalized quantile regression covers in poor prediction bins","Local validity only from LVD on foundation model embeddings"]},"model":"grok-4.3","cost_usd":0.005972,"raw_usage":{"total_tokens":2849,"prompt_tokens":705,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":59724500,"prompt_tokens_details":{"text_tokens":705,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2064,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":705,"tokens_out":80,"duration_ms":12994,"temperature":1.0,"reasoning_tokens":2064,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T20:39:30.178742+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new collection of galaxies drawn from the same distribution where the conformal prediction intervals cover the true property values at a rate materially below the nominal 90% would falsify the coverage claim.","supporting_citations":[],"review_version":1}