{"id":"3f66cba3-ff54-49cd-a33c-01487911b9be","arxiv_id":"2411.18054","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Transfer learning from a broad photometric-redshift sample to a spectroscopic sample cuts bias and RMS error for galaxy redshift prediction on the spectroscopic sample, but degrades performance on the broad sample.","lead":"Astronomers trained a neural network on photometric redshifts from a broad survey, then refined it on precise spectroscopic redshifts, reducing bias and error on the spectroscopic sample. The same approach slightly hurt performance on the broad survey, revealing a trade-off in combining different ground truths for galaxy distance estimation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed generalization gain is not established because target metrics are measured on the same survey used for fine-tuning, with no same-framework GalaxiesML-only baseline; the reported improvement may just reflect fitting the target labels.","rationale":"The reader correctly identifies a missing same-framework GalaxiesML-only baseline and notes that target gains are in-domain. My stress test agrees with that structural critique but places the weight on the attribution problem rather than on COSMOS2020 photo-z reliability per se. Even if lp_zPDF were perfect, the current experiment would not demonstrate that the source data improves generalization, because the control group is a model deliberately trained on a different/noisy label set and never exposed to target labels. The decisive experiment is the GalaxiesML-only baseline plus an external held-out evaluation. Since the reader's verdict is already CONDITIONAL and the concern is addressable with additional experiments, I recommend no change to the verdict. The paper has independent value in its dataset release and careful data cuts, but the headline generalization claim should be reframed as 'fine-tuning on target spec-z after photo-z pretraining improves target metrics' until the control is run.","tokens_in":11771,"tokens_out":4724,"duration_ms":44842,"concrete_test":"Train a GalaxiesML-only model (NN-GML) with the exact NN-Base architecture, loss, optimizer, and hyperparameter-selection procedure on the GalaxiesML training split. Evaluate NN-GML, NN-Base, NN-TL, and NN-Combo on the same GalaxiesML test set and on the TransferZ holdout. If NN-GML matches or beats NN-TL/NN-Combo on GalaxiesML, the claimed source-data benefit is not established. Additionally, evaluate all four models on an external spectroscopic sample not used in any training (e.g., the COSMOS2020 spec-z validation sample or a separate field) and compare bias, RMS, and catastrophic outlier rate; the generalization claim requires NN-TL/NN-Combo to improve over NN-GML there as well. Report whether the 'meet cosmological requirements' statement holds for NN-GML under the same thresholds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that combining TransferZ (COSMOS2020 photo-z labels) with GalaxiesML (spectroscopic redshifts) via transfer learning or joint training improves generalization and can meet cosmological requirements. The load-bearing step is the comparison in Table 2 and Figure 3: every 'improvement' on the target is computed on the GalaxiesML test set, which is the same survey whose spectroscopic redshifts were used to fine-tune NN-TL and to train NN-Combo. NN-Base, by contrast, was trained only on TransferZ and never saw spec-z. So a large drop in GalaxiesML bias/RMS after adding GalaxiesML to training is expected even if TransferZ contributes nothing useful; it is the trivial effect of fitting the target labels. Without a same-architecture, same-loss model trained only on GalaxiesML, the reader cannot tell whether the source data helped generalization or whether the gains are entirely due to target supervision. The paper's own source-side results (NN-TL bias 10.7x higher and RMS 1.26x higher on TransferZ, Section 4) show the learned features do not transfer back, further undercutting a generalization story. The comparison to J24 is not a substitute because it uses a different architecture and loss. The stated LSST/cosmology conclusion therefore rests on an unsupported attribution of the gain to the photo-z source data, independent of whether the COSMOS2020 labels themselves are unbiased.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces TransferZ, a dataset of 116,335 galaxies with five-band HSC photometry and COSMOS2020 LePhare photometric redshifts, and uses it together with the spectroscopic-redshift GalaxiesML dataset to train photometric redshift networks. Three models are compared: NN-Base trained only on TransferZ, NN-TL initialized on TransferZ and fine-tuned on GalaxiesML, and NN-Combo trained on the combined dataset. On the GalaxiesML test set, both NN-TL and NN-Combo are reported to reduce bias by about 5x, RMS by about 1.5x, and catastrophic outlier rate by about 1.3x relative to NN-Base. The paper also reports that source-side metrics on TransferZ degrade for NN-TL (bias and RMS) and are roughly unchanged for NN-Combo, and concludes that the proposed approaches can meet cosmological requirements for LSST.","tokens_in":12107,"tokens_out":4815,"duration_ms":45923,"significance":"The question addressed is timely and practically relevant: whether broad but less precise photometric-redshift labels can supplement narrow but precise spectroscopic samples for training photo-z models for LSST. The paper's concrete contributions are a publicly released TransferZ dataset (Zenodo DOI) and an evaluation protocol with 100 random initializations, which is a reproducible and honest way to report metric uncertainties. If the generalization claim were supported by a controlled comparison, the result would be valuable for survey preparation. However, the central attribution of the reported gains to the TransferZ source data is not established by the current experimental design, because the target improvements are measured on the same survey used for fine-tuning and no same-framework GalaxiesML-only baseline is provided.","major_comments":[{"comment":"The headline comparison on GalaxiesML compares NN-TL and NN-Combo, both of which are trained or fine-tuned on GalaxiesML spectroscopic redshifts, against NN-Base, which never sees GalaxiesML. The reported 5x bias reduction and 1.5x RMS reduction are therefore expected even if TransferZ contributes nothing, simply from fitting the target labels. A same-architecture, same-loss, same-training-schedule model trained only on GalaxiesML must be added as a baseline before the improvement can be attributed to combining ground truths. Without this control, the 'generalization' claim is unsupported.","section":"Section 4, Table 2"},{"comment":"The source-side results undercut the generalization narrative: NN-TL on TransferZ shows bias increasing from -0.69e-3 to 7.45e-3 and RMS increasing from 22.6e-3 to 28.5e-3 relative to NN-Base. This indicates catastrophic forgetting of the source features rather than transfer of broadly useful representations. The paper should either provide a mechanism for this degradation, demonstrate generalization on an external survey not used in fine-tuning, or substantially temper the claim that TransferZ improves generalization to the broader galaxy population.","section":"Section 4, Table 2, source rows"},{"comment":"The text contains a direct contradiction: it first states 'Our transfer learning model performs better than the one from [23]', then states 'The model from J24 achieved a bias of one order magnitude lower than our approach in NN-TL and NN-Combo evaluated on the GalaxiesML.' Please clarify which metric(s) are meant and present a quantitative side-by-side comparison to J24, which is the closest existing target-only baseline and therefore important for interpreting the results.","section":"Section 4, J24 comparison"},{"comment":"The paper notes that 500 galaxies are common to TransferZ and GalaxiesML and assumes the impact is negligible. It is not stated explicitly that these overlapping objects are removed from the test sets before splitting. If any of them appear in the GalaxiesML test set and also in the TransferZ training set (or the Combo training set), the target metrics will be optimistically biased. Please state clearly how the overlap was handled in the train/validation/test split, or quantify the effect.","section":"Section 2, overlap handling"}],"minor_comments":[{"comment":"References [22] and [23] refer to the same Jones et al. paper and should be consolidated to avoid confusion.","section":"References"},{"comment":"The label 'Catastrophc Outlier Rate' contains a typo, and the 'LSST Requirements' lines in the figure are not defined in the text; please specify the numerical thresholds used.","section":"Figure 3"},{"comment":"The header 'Redshift Median Redshift i-band mag No. Sources 90th percentile Uncertainty 90th percentile Filters' is difficult to parse; please reformat so each column is clearly labeled.","section":"Table 1"},{"comment":"The statement that photometry is 'normalized separately for each training stage' should specify whether validation and test sets are normalized using training-set statistics; otherwise metric comparisons can be affected by a preventable inconsistency.","section":"Section 3"},{"comment":"Figure 2 shows predictions for the GalaxiesML test set only; given the paper's focus on generalization, showing the analogous TransferZ test panels or explicitly limiting the figure to the target set would improve clarity.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a workshop-style contribution with a useful dataset release and a reproducible evaluation protocol. The main missing control is a same-framework GalaxiesML-only baseline; without it, the abstract's generalization claim is overstated. I believe this is fixable within a revision and is not a reason for outright rejection, but the authors should be asked to provide the additional experiment or clearly reframe the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a tidy empirical paper with a genuinely useful new public dataset, but its headline claim—that adding COSMOS2020 photo-z labels improves generalization—is not supported by the experiments as presented. The stress-test concern is right. Every target-side improvement in Table 2 is measured on GalaxiesML, the same survey used to fine-tune NN-TL and to train NN-Combo. The gains over NN-Base (which never saw spec-z) are exactly what you would expect from fitting the target labels. There is no same-architecture, same-loss model trained only on GalaxiesML. The J24 comparison does not fill that gap because it uses a different architecture and loss.\n\nWhat is actually new: TransferZ, a 116k-galaxy dataset matching COSMOS2020 photo-z labels (35-band LePhare) to HSC five-band photometry, with explicit quality cuts and a Zenodo DOI. That is a real resource for survey teams, especially for LSST-like five-band work. The paper also documents a clean dataset-construction workflow, and the internal comparisons are consistent. Credit where due: the source-side degradation is reported honestly rather than buried. The NN-TL bias on TransferZ is 7.45e-3 versus -0.69e-3 for NN-Base, which undercuts any story that learned features transfer back, but at least it is transparent.\n\nSoft spots, in order. (1) The missing GalaxiesML-only baseline is the load-bearing gap. Without it, the attribution of the gain to TransferZ is untestable; the improvement could be entirely from target supervision. This is fixable and should be added. (2) The \"can meet cosmological requirements\" line is a stretch. The comparison to LSST requirements leans on one scatter estimate while their own bias on GalaxiesML is about an order of magnitude worse than J24; the abstract overstates what the evidence supports. (3) The transfer-learning configuration (freeze pattern, learning rate 5e-10) is under-justified; they should at least report sensitivity to those choices. (4) Using photo-z as ground truth is acknowledged and handled with reasonable quality cuts; I do not see that as circular, just standard supervised evaluation on labels with larger uncertainty.\n\nWho is this for? Practitioners who want to try leveraging broad photo-z catalogs for pretraining and need a ready-made dataset. It deserves a serious referee. I would accept it but require the missing baseline and a toned-down abstract. If the gain survives that baseline, this is a solid recipe paper; if it does not, the dataset still has value on its own.","headline":"A useful new dataset and clean experiments, but the central generalization claim is not established until the authors add a same-framework GalaxiesML-only baseline.","tokens_in":12625,"tokens_out":1761,"would_cite":true,"duration_ms":16934,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Mixing photometric and spectroscopic redshifts cuts photo-z bias 5x.","keywords":["photometric redshifts","transfer learning","ground truth combination","COSMOS2020","GalaxiesML","neural networks","cosmological surveys","LSST"],"falsifier":"An independent check would be to evaluate NN-TL and NN-Combo on a spectroscopic sample that is distinctly fainter or redder than GalaxiesML (e.g., a deeper spectroscopic survey in the COSMOS field). If the hybrid models show the same improvements over NN-Base on that held-out population, the gains reflect genuine generalization; if the gains shrink or vanish, they are specific to GalaxiesML's color-magnitude range. A second check is to retrain after removing the ~500 galaxies that appear in both TransferZ and GalaxiesML to see whether the reported improvements depend on this overlap.","tokens_in":11601,"feed_emoji":"🌌","tokens_out":7862,"duration_ms":61822,"temperature":0.7,"pith_summary":"The paper claims that photometric redshift neural networks can be made to generalize better by training on two complementary sources of real ground truth at once: a broad but imprecise photometric redshift sample (TransferZ, built from the 35-band COSMOS2020 catalog) and a narrower but precise spectroscopic sample (GalaxiesML). Two recipes—transfer learning, where a network pretrained on TransferZ is fine-tuned on GalaxiesML, and joint training on the combined dataset—both improve bias, RMS error, and catastrophic outlier rate on GalaxiesML by roughly 5x, 1.5x, and 1.3x relative to a baseline trained only on TransferZ. The paper argues these gains are enough to meet the redshift accuracy requirements of upcoming surveys like LSST. The cost is a modest worsening of bias and RMS on the TransferZ sample itself, while catastrophic outlier rates on that sample improve.","feed_headline":"Mixing photometric and spectroscopic redshifts cuts photo-z bias 5x","feed_subtitle":"Pretraining on broad photo-z labels and refining with spectroscopy meets LSST-era requirements.","key_machinery":"The central object is TransferZ, a dataset that pairs five-band HSC grizy photometry from HSC PDR2 with photometric redshift labels taken from the 35-band COSMOS2020 catalog (specifically the LePhare lp_zPDF median of the likelihood). These labels are ~100 times less precise than spectroscopy but cover a much wider and fainter galaxy population. The mechanism that carries the argument is the combination of this broad, imprecise label source with the narrow, precise spectroscopic labels of GalaxiesML, realized in two ways: transfer learning, in which the base network trained on TransferZ is fine-tuned on GalaxiesML with most layers frozen and a learning rate of $5\\times10^{-10}$, and joint training, in which a network is trained on the concatenated Combo dataset of 402,408 galaxies. Both recipes force the model to keep the wide coverage learned from photometric redshifts while sharpening its predictions using spectroscopy.","core_discovery":"The central discovery is that combining ground truths from different sources—rather than using only the most precise labels available—improves photometric redshift estimation. The authors construct TransferZ, a dataset of 116,335 galaxies with five-band HSC photometry paired with photometric redshifts from the 35-band COSMOS2020 survey (median uncertainty ~0.03), as a source that covers a wider range of galaxy types, magnitudes, and colors than spectroscopic samples. They pair it with GalaxiesML, 286,401 galaxies with spectroscopic redshifts (median uncertainty ~0.0002). Three networks are trained: a base model on TransferZ alone, a transfer-learned model that fine-tunes the base on GalaxiesML, and a combined model trained on both datasets at once. On the GalaxiesML test set both hybrid approaches reduce bias by ~5x, RMS by ~1.5x, and catastrophic outlier rate by ~1.3x compared to the base model, and the paper reports that these results meet cosmological requirements. The combined model slightly outperforms transfer learning on bias and RMS, while transfer learning gives better catastrophic outlier control.","pith_inferences":["The reported degradation on the TransferZ test set suggests a trade-off surface between accuracy on the precise narrow sample and accuracy on the broad photometric sample; weighting the two losses or using domain-adaptive training might recover both.","The same pretrain-on-photometric, fine-tune-on-spectroscopic recipe could be applied to other surveys where a multi-band photometric catalog with template redshifts overlaps a small spectroscopic calibration sample; the improvement should be tested there.","A strong test of the generalization claim is to use the trained models to predict redshifts for galaxies within known clusters: cluster members share a redshift but span many galaxy types, so tight, unbiased predictions would independently validate that the hybrid training generalizes beyond the training color-magnitude range.","The assumption that the ~500 overlapping galaxies have negligible impact can be checked directly; if their removal changes the reported factors, part of the apparent improvement is leakage rather than generalization."],"forward_implications":["Photometric redshift models for LSST and similar surveys can be built without waiting for a complete spectroscopic sample: the broad, imprecise photometric labels supply coverage, and a relatively small precise spectroscopic sample anchors accuracy.","The two recipes give comparable results, so the choice between them can be driven by the science goal: NN-Combo gives lower bias and RMS on the target sample, while NN-TL gives a lower catastrophic outlier rate.","Using photometric redshifts as training labels, despite being ~100x less precise than spectroscopy, improves the model's performance on the spectroscopic test set relative to training on the broad sample alone, indicating that training-set representativeness matters as much as label precision.","The released TransferZ dataset lets the community reproduce the hybrid training results and test variations directly."],"supporting_citations":[{"why":"Supplies the photometric redshift ground-truth labels for TransferZ.","marker":"[51]"},{"why":"Provides the spectroscopic ground-truth dataset used for target training and evaluation.","marker":"[15]"},{"why":"Gives the five-band grizy photometry used as model features.","marker":"[1]"},{"why":"The LePhare template-fitting method that produced the COSMOS2020 photometric redshifts used as labels.","marker":"[19]"},{"why":"The spectroscopic-only neural network baseline that the hybrid models are compared against.","marker":"[23]"},{"why":"The neural network architecture that the models are based on.","marker":"[24]"},{"why":"Source of the custom loss function used for training.","marker":"[47]"},{"why":"Supplied the photometric-redshift quality-cut criteria used to build TransferZ.","marker":"[45]"}],"fun_headline_variants":["Combining photo-z and spec-z labels cuts photo-z bias 5x","Two-label training reduces photometric redshift bias fivefold","Transfer learning with mixed ground truths improves photo-z fivefold","Pretrain on photometric redshifts, fine-tune on spectroscopy: bias cut 5x"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The premise that the photometric redshifts from COSMOS2020 are good enough to act as training truths for the broad galaxy population, despite being about a hundred times less precise than spectroscopy and possibly carrying systematic biases, is load-bearing: if those labels are systematically wrong for some galaxy types, the claimed improvement in generalization could be an illusion.","fun_headline_variants_meta":{"raw":{"variants":["Combining photo-z and spec-z labels cuts photo-z bias 5x","Two-label training reduces photometric redshift bias fivefold","Transfer learning with mixed ground truths improves photo-z fivefold","Pretrain on photometric redshifts, fine-tune on spectroscopy: bias cut 5x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000242,"raw_usage":{"total_tokens":1557,"prompt_tokens":1011,"completion_tokens":546,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":469}},"tokens_in":627,"tokens_out":546,"duration_ms":4624,"temperature":1.0,"reasoning_tokens":469,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:33:12.106141+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent check would be to evaluate NN-TL and NN-Combo on a spectroscopic sample that is distinctly fainter or redder than GalaxiesML (e.g., a deeper spectroscopic survey in the COSMOS field). If the hybrid models show the same improvements over NN-Base on that held-out population, the gains reflect genuine generalization; if the gains shrink or vanish, they are specific to GalaxiesML's color-magnitude range. A second check is to retrain after removing the ~500 galaxies that appear in both TransferZ and GalaxiesML to see whether the reported improvements depend on this overlap.","supporting_citations":[{"cited_title":"Photometric Redshifts for Cosmology: Improving Accuracy and Uncertainty Estimates Using Bayesian Neural Networks","cited_arxiv_id":"2202.07121","evidence_quote":"The neural network architecture that the models are based on."},{"cited_title":"Machine Learning Classification to Identify Catastrophic Outlier Photometric Redshift Estimates","cited_arxiv_id":null,"evidence_quote":"Supplied the photometric-redshift quality-cut criteria used to build TransferZ."}],"review_version":1}