{"id":"68dac822-bd7b-4be4-b113-442e33c7d920","arxiv_id":"2607.26822","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"turboPDZ extracts optimized photo-z point estimates and reliability scores from PDZs that outperform HSC catalog z_best and risk/confidence flags across six pipeline-layer combinations.","lead":"A machine-learning pipeline turns photometric-redshift probability curves into better single redshifts and a data-driven reliability score. It beats catalog defaults on HSC data and is especially useful when template-fitting quality flags fail.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection beyond the reader's already-flagged generalization risk; the in-sample comparative claims hold.","rationale":"The reader's strongest claim is narrowly empirical and comparative on labeled HSC sets; the tables, Mizuki qualitative failure of analytic flags, and complementary PCA/location vs PCA/peak importances make fabrication or circularity unlikely. Correctness risk on what was measured is low. The generalization premise (spec-z cross-match → full catalog / science samples) is the genuine soft underbelly, but the reader already made it the weakest_assumption and correctly issued CONDITIONAL rather than ACCEPT. No tighter internal inconsistency (loss design, calibration leakage, or metric definition) rises to the same load-bearing level: log-space MAE plus post-hoc multiplicative calibration is explicitly a ranking device, not a probabilistic uncertainty, and the paper does not over-claim otherwise. I therefore leave the verdict unchanged and mark full agreement with the reader on the critical assumption.","tokens_in":19611,"tokens_out":562,"duration_ms":13815,"concrete_test":"Apply the released Wide-layer models to a magnitude- or color-reweighted test split that matches the full photometric catalog's i-band and color distribution (or to an external deep-field spec-z set with disjoint training overlap); recompute Table 3 metrics and Table 6 AUCs. If σ_NMAD/η gains or AUC advantage shrink by ≳30% relative to the paper's splits, the VAC-default claim weakens; stability would clear the residual concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparative claim (z_ml beats z_best on σ_NMAD/η_0.15; r_ml beats catalog risk/conf on AUC filtering, including the Mizuki failure mode) is supported by six independent pipeline–layer tables on held-out labeled splits, with feature-importance patterns that are physically complementary rather than circular. The only load-bearing soft spot is the one the reader already named: both models are supervised on the spectroscopic cross-match, so claims that the released VAC quantities are optimal for the full photometric catalog (or for cosmology samples with different mag/z selection) rest on representativeness rather than on a blind photometric-only or out-of-distribution check. The paper itself notes possible HSC-pipeline training overlap with the test set and magnitude-distribution mismatch (§5.1), so the fairness of the z_best baseline and the external validity of r_ml are not fully secured. That does not undermine the tabulated in-sample gains.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper presents turboPDZ, a survey-independent ML pipeline that extracts an optimized photometric-redshift point estimate z_ml and a reliability score r_ml directly from photo-z PDZs. PDZs from the three HSC-SSP PDR3 pipelines (DEmP, DNNz, Mizuki) in Wide and DUD layers are PCA-compressed, augmented with hand-crafted descriptors, and fed to Optuna-tuned MLPs. z_ml is trained under a composite σ_NMAD/η_0.15 objective; σ_ml is trained in log-space with post-hoc calibration and converted to r_ml by percentile ranking. On held-out spectroscopic test splits, z_ml improves σ_NMAD and η_0.15 over catalog z_best in all six pipeline–layer combinations (Table 3), and r_ml yields lower AUC under the σ_NMAD and η_0.15 versus retained-fraction curves than photoz_risk_best and photoz_conf_best (Table 6, Fig. 9), with a particularly large gain for Mizuki where catalog indicators fail. Feature-importance analyses show complementary roles for PCA/location versus PCA/peak statistics. Code and a value-added catalog are released.","tokens_in":19963,"tokens_out":1432,"duration_ms":38498,"significance":"If the reported gains hold under realistic use, the work is a useful methodological and practical contribution to photo-z post-processing for Stage-III/IV surveys. Strengths include: (i) consistent tests across six independent PDZ products rather than a single pipeline; (ii) held-out stratified evaluation with separate calibration for σ_ml; (iii) a clear demonstration that analytic risk/confidence can fail for template-fitting PDZs while a data-driven ranker still works; (iv) physically interpretable grouped permutation importances; and (v) public code plus a planned VAC of z_ml, σ_ml, and r_ml. These make the paper more than a one-off catalog tweak and support reuse on other surveys after retuning.","major_comments":[{"comment":"§5.1 explicitly notes that HSC pipeline training may overlap the held-out test set, so z_best metrics in Table 3 may be optimistic and the reported z_ml gains are a lower bound. That caveat is appropriate but under-developed relative to the central comparative claim. Please quantify or bound the overlap if feasible (e.g., via public training-set documentation or a rough membership estimate), and state clearly in the abstract/results that the baseline comparison is not fully blind. Without that, readers may over-read Table 3 as a fully independent bake-off.","section":"§5.1, Table 3"},{"comment":"The strongest application claim is release of z_ml/σ_ml/r_ml for all PDR3 objects (§6, abstract). Both models are supervised only on the spectroscopic cross-match after magnitude/redshift cuts (§2.4, Table 1), and §5.1 already notes possible magnitude-distribution mismatch with the full catalog. The manuscript does not show that ranking quality or point-estimate gains persist under reweighting to the photometric selection, in faint bins with sparse spec-z, or on a held-out spectroscopic survey not used in labeling. This does not invalidate the in-sample tables, but it is load-bearing for the VAC. Please add a dedicated limitations subsection with concrete guidance (e.g., recommended r_ml thresholds only where the labeled prior is representative) and, if possible, one simple stress test (magnitude reweighting or bright/faint split) for r_ml AUC and z_ml metrics.","section":"§2.4, §5.1, §6"},{"comment":"r_ml is defined from a log-space MAE error scale on |Δz_ml| plus percentile ranking (Eqs. 10–14), not from a calibrated probabilistic posterior; the text correctly disclaims NLL-style uncertainty. Table 5 and Fig. 8 support improved ranking versus s_best, but science users may still treat σ_ml as a Gaussian width for error propagation. Please state more prominently (abstract or §4.2 opening) that σ_ml/r_ml are ranking/filter products, not drop-in replacements for PDZ moments in forward modeling, and avoid language that equates σ_ml with a standardized uncertainty without that caveat.","section":"§4.2, Eqs. (10)–(14), Table 5"}],"minor_comments":[{"comment":"Eq. (1): the composite objective normalizes by training-split z_best metrics; briefly note whether fold-wise renormalization was tried and whether equal 1/2 weights were ablated, even if only in the appendix.","section":"§3.3, Eq. (1)"},{"comment":"Figure 1 is illustrative but has no quantitative summary of PDZ diversity (e.g., typical multimodality rates). A short sentence or small table would help motivate PCA component counts in Table 2.","section":"Fig. 1, Table 2"},{"comment":"Table 3 highlights best results in red; ensure the monochrome/print version remains readable (bold or boxes).","section":"Table 3"},{"comment":"§4.1 odds discussion (Eqs. 7–9) is clear but slightly long relative to its use; consider tightening or moving part to an appendix.","section":"§4.1"},{"comment":"Data availability: confirm whether the VAC DOI/release path will be frozen at acceptance and whether model cards list the exact train/test object IDs for reproducibility.","section":"Data availability"},{"comment":"Minor typography: “photo-zor” / “spec-zorz” spacing in the introduction; “Optunaunder” and similar missing spaces in the abstract PDF text; unify “Deep/UltraDeep” vs “DUD”.","section":"Abstract, §1"}],"recommendation":"minor_revision","confidential_remarks":"Sound applied-methods paper for PASJ; central in-sample claims look solid. The main editorial risk is over-selling the full-catalog VAC without stronger external-validity language. I would not require new surveys or a full probabilistic recalibration for acceptance. No concerns about scope or conduct."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: on held-out spectroscopic cross-matches from all three HSC-SSP PDR3 pipelines (DEmP, DNNz, Mizuki) in both Wide and DUD, their z_ml improves σ_NMAD and η_0.15 over catalog z_best, and r_ml filters more efficiently than photoz_risk_best / photoz_conf_best by AUC on the retained-fraction curves. For Mizuki the catalog flags actually get worse as you cut, while r_ml still works—that is the cleanest qualitative check in the paper.\n\nWhat is actually new is not another photo-z engine. It is a systematic extraction layer: PCA compression of the PDZ plus hand descriptors, Optuna-tuned MLP under a composite scatter/outlier objective for the point estimate, then a second net trained in log-error space with post-hoc multiplicative calibration and percentile ranking into r_ml. Prior work already poked at point estimates and odds/risk; this packages the optimization, ships the code, and releases the VAC quantities for the full PDR3.\n\nThey do the empirical work carefully. Six independent settings, stratified splits, separate calibration split, feature-importance patterns that invert sensibly between the two tasks (location stats matter for z_ml; peak/morphology stats matter for reliability). They flag the possible overlap between HSC pipeline training and their test set, and they note magnitude-distribution mismatch with the full catalog. That honesty helps.\n\nSoft spots are real but proportionate. Everything is supervised on the spec-z cross-match, so claims that the released VAC is optimal for the unlabeled photometric sample (or for cosmology cuts with different mag/z selection) rest on representativeness, not on a blind photometric-only check. The z_best baseline may be slightly optimistic for the same reason. No error bars on the metrics. None of that invents the tabulated gains; it just bounds how hard you should lean on the VAC for Stage-IV systematics without further validation.\n\nMath and citation pattern look fine—standard robust photo-z metrics, sensible references to Tanaka, Nishizawa, Schmidt, Desprez, etc. This is for people who actually cut gold samples or build photo-z VACs. I would send it to referees; it is solid methods work, not desk-reject material. Engage if you touch HSC or photo-z quality flags; skim the tables and the Mizuki section if you only need the takeaway.","headline":"Practical, well-executed ML stack that consistently beats HSC catalog point estimates and quality flags on six PDZ sets, with public code and a VAC; main caveat is labeled-spec generalization.","tokens_in":20553,"tokens_out":618,"would_cite":true,"duration_ms":20856,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A machine-learning pipeline that reads photometric-redshift probability distributions yields better point estimates and reliability scores than the catalog defaults across six HSC pipeline–layer combinations.","keywords":["photometric redshifts","redshift PDFs","point estimates","reliability scores","HSC-SSP","machine learning","large-scale structure","catalogs"],"falsifier":"On a fully held-out spectroscopic sample with no overlap to any training or calibration set used by either turboPDZ or the original HSC pipelines, check whether z_ml still beats z_best in σ_NMAD and η_0.15 and whether r_ml still yields smaller AUC under the quality-versus-retained-fraction curves than the catalog risk and confidence flags.","tokens_in":20489,"feed_emoji":"🔭","tokens_out":985,"duration_ms":21989,"temperature":0.7,"pith_summary":"Photometric redshifts are delivered as full probability curves, yet large-scale structure work usually collapses each curve to a single number and a quality flag. This paper argues that those two numbers can be optimized directly from the curve with a survey-independent neural pipeline called turboPDZ. On Hyper Suprime-Cam data, the optimized redshift beats the catalog “best” estimate in scatter and outlier rate for every one of the six pipeline-and-depth combinations tested. The companion reliability score filters galaxies more efficiently than the catalog risk and confidence indicators, and it still works when the catalog flags fail—most dramatically for the template-fitting pipeline. The practical payoff is cleaner gold samples for weak lensing and clustering without throwing away as many galaxies.","feed_headline":"ML pulls better redshifts from photo-z probability curves","feed_subtitle":"On six HSC setups, optimized estimates and reliability scores beat catalog defaults and rescue failed flags.","key_machinery":"turboPDZ: each PDZ is PCA-compressed and joined to magnitude and summary descriptors; one multilayer perceptron, tuned under a composite σ_NMAD–η_0.15 objective, outputs z_ml; a second network trained in log-error space and post-hoc calibrated outputs σ_ml, which is turned into the percentile-rank reliability score r_ml.","core_discovery":"Across all six HSC-SSP PDR3 pipeline–layer combinations, an optimized point estimate z_ml extracted from each galaxy’s redshift probability distribution improves σ_NMAD and η_0.15 relative to the catalog z_best, and a data-driven reliability score r_ml, derived from a calibrated uncertainty σ_ml, filters galaxies more efficiently than photoz_risk_best and photoz_conf_best as measured by the area under the quality-versus-retained-fraction curves. For the Mizuki template-fitting pipeline the catalog indicators fail dramatically (AUC up to about ten times larger), while r_ml still identifies unreliable objects across redshift regimes.","pith_inferences":["Because point estimation leans on location statistics while reliability leans on peak morphology, joint multi-task training of z and σ might trade a little redshift accuracy for still-stronger ranking of bad objects.","The method’s dependence on a representative spectroscopic anchor implies that deep, sparse fields (or high-z tails) will need explicit domain-adaptation checks before r_ml cuts are trusted at the same thresholds as in the training domain.","Surveys that already ship full PDZs but weak quality flags (including Stage-IV optical and near-IR catalogs) are the natural next stress tests of whether the Mizuki-style rescue generalizes."],"forward_implications":["Gold photometric samples for weak lensing and clustering can keep more galaxies at fixed scatter and outlier rate by cutting on r_ml instead of catalog risk or confidence.","Template-fitting PDZs that are sharply peaked at the wrong redshift can still be flagged as unreliable by r_ml even when analytic risk and confidence fail.","The same PCA-plus-descriptor architecture can be retuned for other surveys’ PDZs without redesigning the pipelines that produced those PDZs.","Value-added z_ml, σ_ml, and r_ml columns for the full HSC-SSP PDR3 catalog become available for immediate downstream use."],"fun_headline_variants":["turboPDZ extracts sharper redshifts from photo-z PDZs","Optimized z_ml beats catalog z_best on all six HSC setups","ML reliability score r_ml filters galaxies better than catalog flags","PCA-compressed PDZs yield improved photo-z point estimates","Data-driven r_ml rescues Mizuki where catalog indicators fail"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Models trained on the spectroscopic cross-match must still rank and predict well on the full photometric catalog and on science samples whose magnitude and redshift mix differ from the labeled set.","fun_headline_variants_meta":{"raw":{"variants":["turboPDZ extracts sharper redshifts from photo-z PDZs","Optimized z_ml beats catalog z_best on all six HSC setups","ML reliability score r_ml filters galaxies better than catalog flags","PCA-compressed PDZs yield improved photo-z point estimates","Data-driven r_ml rescues Mizuki where catalog indicators fail"]},"model":"grok-4.5","effort":"low","cost_usd":0.003527,"raw_usage":{"total_tokens":1301,"prompt_tokens":961,"num_sources_used":0,"completion_tokens":75,"cost_in_usd_ticks":35268000,"prompt_tokens_details":{"text_tokens":961,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":265,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":961,"tokens_out":75,"duration_ms":5412,"temperature":1.0,"reasoning_tokens":265,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T20:04:19.626336+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a fully held-out spectroscopic sample with no overlap to any training or calibration set used by either turboPDZ or the original HSC pipelines, check whether z_ml still beats z_best in σ_NMAD and η_0.15 and whether r_ml still yields smaller AUC under the quality-versus-retained-fraction curves than the catalog risk and confidence flags.","supporting_citations":[],"review_version":1}