{"id":"05d0b937-e738-4121-a96e-a1eb76e5f1b8","arxiv_id":"2509.08189","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A neural network with two-stage MCMC solves multi-band contact binary light curves and recovers stellar and spot parameters, demonstrated on 3,541 OGLE systems.","lead":"This paper trains a neural network to quickly estimate the physical properties of contact binary stars from multi-band light curves. The model also fits spot parameters and was applied to more than 3,500 binaries from the OGLE survey.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reliability claim depends on uncalibrated posteriors: Table 2's ±0.000 uncertainties on q and i for real light curves imply the MCMC likelihood ignores photometric noise; the OGLE catalog error bars are therefore unsupported.","rationale":"The reader's verdict is CONDITIONAL, and I agree with that status. The most load-bearing issue is not only the generic PHOEBE-to-real fidelity, but the internal inconsistency of the reported uncertainties: Table 2's near-zero posterior widths for observed light curves are physically implausible and indicate that the MCMC likelihood ignores observational noise. This matters because the paper's product is a catalog of 3,541 systems with error bars (Table 3), and the central claim explicitly includes reliable parameter determination. The synthetic validation checks accuracy on noise-free, same-simulator data but does not calibrate uncertainty. The comparison with Wang+2024 shows good agreement in q, i, T2, and radii for several targets, which supports the forward-model inversion concept, but the spot-parameter mismatches and zero uncertainties show that the error estimates are not trustworthy. A coverage test with injected noise is the direct, decisive check: it would show whether the reported credible intervals are calibrated. I would keep the conditional verdict pending that test rather than rejecting the work, because the architecture and the broad agreement on geometric parameters indicate the method has real promise.","tokens_in":10491,"tokens_out":5590,"duration_ms":76746,"concrete_test":"Run a posterior calibration test: generate 100 synthetic V/I light curves from PHOEBE with parameters drawn from the training prior, add Gaussian noise at the per-point level matching the observed OGLE scatter, and run the full two-stage MCMC pipeline described in Section 3. Measure the empirical coverage of the 68% credible intervals for q, i, T2, and the four spot parameters. If coverage is not ≈68% — or if the credible intervals are orders of magnitude narrower than the injected-noise floor — the published uncertainties are likelihood artifacts and Table 3's error bars are not scientifically usable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — that the model reliably determines physical parameters, including their uncertainties — requires the two-stage MCMC posteriors in Table 2 to be meaningful. The paper never specifies the likelihood or noise model used in the MCMC (Section 3). For eight observed systems, it quotes q uncertainties of ±0.000, i uncertainties of ±0.0 deg, and T2 uncertainties of ±1 K, which are 10–1000× smaller than the W-D uncertainties from Wang+2024 listed in the same table (Δq ≈ 0.001–0.002, Δi ≈ 0.1–0.5°, ΔT2 ≈ 20–40 K). Real photometry contains noise and the contact-binary parameter space has known degeneracies; such delta-function posteriors indicate the likelihood effectively assigns near-zero noise to the data. This is not rescued by the 1000-set synthetic validation, because those LCs are generated by the same PHOEBE forward model and, as described, without injected observational noise. Since Table 3 reports the same kind of nearly singular uncertainties for all 3,541 OGLE systems, the scientific reliability of the catalog's error bars is unestablished. The spot-parameter comparisons further undercut the 'remarkable consistency' claim: e.g., V0394 Cam gives r_s = 11° vs 25°, T_s = 0.89 vs 0.94; V0737 Cep gives T_s = 0.90 vs 0.96. These discrepancies are much larger than the quoted zero uncertainties.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a neural-network surrogate model for contact-binary light-curve analysis. A fully connected network is trained on ~437,000 PHOEBE-generated light curves over a 10-dimensional parameter space that includes temperature ratio, mass ratio, inclination, potential, fillout, third light, and four starspot parameters, and is then fine-tuned by transfer learning to 19 photometric bands. A two-stage MCMC procedure is used to invert observed light curves. The model is validated on 1000 synthetic PHOEBE light curves and on eight systems from Wang et al. (2024), and is applied to OGLE V/I data, yielding physical parameters for 3,541 systems. The software is archived as CBLA.exe at China-VO.","tokens_in":10952,"tokens_out":4062,"duration_ms":46019,"significance":"If the inferred parameters and their uncertainties are reliable, this would be a practically valuable tool for the large photometric surveys now producing hundreds of thousands of contact-binary light curves. The paper's concrete strengths are its large training set, multi-band transfer-learning approach, inclusion of all four spot parameters, publicly archived executable, and the reported speed gain (82 s vs. 4.8 days for a three-band light curve). However, the central reliability claim is currently not supported: the MCMC uncertainties reported for real systems are essentially zero, the likelihood/noise model is never stated, and the synthetic validation is performed on noise-free PHOEBE light curves. The advertised spot-parameter capability is also contradicted by the comparison in Table 2.","major_comments":[{"comment":"The MCMC likelihood and noise model are never specified. For the eight real systems, Table 2 reports uncertainties of +0.000/-0.000 for q, +0.0/-0.0 deg for i, and +1/-2 K for T2 for most targets, whereas the comparison Wang et al. (2024) values have finite uncertainties (q ~0.001-0.002, i ~0.1-0.5 deg, T2 ~20-40 K). Such delta-function posteriors can arise only if the likelihood effectively assigns near-zero noise to the photometric data. The 1000-set synthetic validation (Figure 2) does not rescue this because those light curves are generated by the same PHOEBE model with no stated injected observational noise. Since Table 3 reports the same kind of nearly singular uncertainties for all 3,541 OGLE systems, the catalog error bars are unsupported. Please specify the likelihood/noise model and validate the two-stage MCMC on noisy synthetic light curves with known injected noise.","section":"Section 3, Table 2"},{"comment":"The text states that the model and Wang et al. (2024) show 'remarkable consistency,' but the spot parameters in Table 2 contradict this. For example, V0394 Cam gives r_s = 11 deg vs. 25 deg and lambda = 317 deg vs. 351 deg; J055741 gives lambda = 63 deg vs. 10 deg and r_s = 5 deg vs. 8 deg; several T_s values differ by 0.04-0.05 (e.g., V0394 Cam 0.89 vs. 0.94; V0737 Cep 0.90 vs. 0.96). Given that the quoted uncertainties are zero to the displayed precision, these are many-sigma discrepancies. Since simultaneous determination of all four spot parameters is a central advertised capability, this comparison does not support the claimed accuracy.","section":"Table 2, spot parameters"},{"comment":"The OGLE application selects among four configurations (phase shift 0 or 0.5, with or without third light) by choosing the highest R^2, and spots are included only if the two maxima differ by more than 0.01 mag in both bands. This model-selection procedure has no penalty for extra degrees of freedom and is not cross-validated on held-out observed data. With real photometric noise, the 'best R^2' configuration can absorb noise, especially when spot parameters are free. Please clarify whether R^2 is computed on the binned, normalized fluxes and whether the reported goodness of fit is compared to the unbinned original data. A cross-validation or information-criterion comparison would strengthen the catalog-level claims.","section":"Section 4, OGLE model selection"}],"minor_comments":[{"comment":"Typo: 'between our model-derived physical parameters and and the true values' should read 'and the true values.'","section":"Section 3"},{"comment":"The statement that transfer learning requires '5,000 to 50,000' training samples per band is vague. Please report the actual training set size and fine-tuning details for each band (e.g., in a table or appendix).","section":"Section 2"},{"comment":"The description of second-stage priors is unclear: 'their uncertainties set to the range of the corresponding Gaussian distributions' appears circular. Define how the prior width is computed from the first-stage chain.","section":"Section 3"},{"comment":"Several entries in Table 2 are empty (e.g., L1g/LTg for NSVS 503993 and NSVS 2561806). Please state explicitly which quantities were not fitted or not available for those targets.","section":"Table 2"},{"comment":"Column 23 heading: 'Equival volume radius' should be 'Equivalent volume radius.'","section":"Table 3"},{"comment":"The outlier removal ('data points below the 1st percentile and above the 99th percentile') is ambiguous: is this applied per star, per band, or globally? Please clarify.","section":"Section 4"},{"comment":"The caption does not define the discrepancy variable or the units. Please specify what is plotted (e.g., inferred minus true value) and whether outliers are truncated.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The eight-system benchmark (Wang et al. 2024) shares a co-author, which reduces the independence of the real-data validation; an external benchmark or a comparison with a non-overlapping sample would substantially strengthen the paper. The uncertainty-calibration issue in Table 2 is the main blocker: if the reported error bars are not meaningful, the OGLE catalog cannot be used for statistical population studies. This is fixable within the manuscript's scope by specifying the noise model and re-validating with noise-injected synthetic light curves."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the capability: a neural network that fits contact binary light curves simultaneously in multiple bands and optimizes all four spot parameters. That is a real extension over Ding et al. 2022 and Li & Wang 2025, which were single-band or incomplete on spots. Training on 437,000 PHOEBE models and using transfer learning to cover 19 filters is sensible. The reported speed gain—82 seconds versus 4.8 days per three-band target—makes this a practical tool for large surveys. Validation on 1000 synthetic PHOEBE light curves shows good parameter recovery, and the eight real systems from Wang et al. 2024 mostly land near the comparison values. The executable is archived, which helps if someone wants to try it.\n\nThe soft spot is the error bars, and it is load-bearing. Table 2 quotes q uncertainties of ±0.000 and i uncertainties of ±0.0 degrees for real observed systems, while the independent W-D analyses in the same table quote ±0.001–0.002 and ±0.1–0.5 degrees. Real photometry has noise, and contact binaries have known degeneracies. Those delta-function posteriors imply the MCMC likelihood is effectively assigning near-zero noise to the data. The paper never specifies the likelihood or noise model in Section 3. The synthetic validation does not fix this, because those light curves were generated without injected observational noise—so they only confirm that the inversion works in the noiseless PHOEBE world. The spot-parameter discrepancies reinforce the problem: V0394 Cam gives r_s = 11° vs 25° and T_s = 0.89 vs 0.94, far outside the quoted zero uncertainties. The claim of “remarkable consistency” is therefore overstated, and the error bars in the 3,541-system OGLE catalog are unsupported.\n\nI still think the basic capability is real. As a fast first-pass parameter estimator, the model is useful. But as a source of reliable physical parameters with trustworthy uncertainties, the paper needs a proper noise model, calibrated posteriors on noisy synthetic data, and independent validation—ideally with code and not just an executable.\n\nThis paper deserves peer review because the method is important and the core approach is sound, but the error analysis needs major revision before the catalog can be used scientifically.","headline":"Useful multi-band NN tool for contact binary light curves, but the quoted near-zero uncertainties on real systems are unsupported and the 3,541-system catalog error bars need serious revision.","tokens_in":11376,"tokens_out":1749,"would_cite":false,"duration_ms":21584,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural network trained on synthetic light curves recovers contact-binary parameters from multi-band survey data in about a minute.","keywords":["Contact binary stars","Eclipsing binary stars","Light curve analysis","Neural networks","Starspots","Fundamental parameters of stars","Astronomy data analysis","Astronomy software"],"falsifier":"Compare the model's mass ratios and temperature ratios for a sample of contact binaries against independent double-lined spectroscopic radial-velocity orbits; a systematic offset larger than the quoted uncertainties in q or T2/T1 would show the synthetic-training premise fails. A practical mid-step is to fit a subset of the 3,541 OGLE systems with a conventional W-D/PHOEBE analysis and check whether parameter differences are consistent with the claimed precision.","tokens_in":10409,"feed_emoji":"⭐","tokens_out":8531,"duration_ms":84002,"temperature":0.7,"pith_summary":"The paper sets out to replace slow, iterative light-curve fitting of contact binaries with a neural network that maps multi-band light curves directly to physical parameters. It claims the network recovers temperature ratio, mass ratio, inclination, potential, fillout, radii, luminosities, third light, and all four starspot parameters from phased, normalized flux curves. This matters because all-sky surveys have identified over a million contact binaries, far too many for traditional physics-based fitting codes. The model is validated on synthetic light curves and on eight previously studied systems, then applied to OGLE V and I data to derive parameters for 3,541 contact binaries.","feed_headline":"Neural network decodes 3,541 contact binaries from survey light curves","feed_subtitle":"Model trained on synthetic light curves recovers mass ratio, temperature, radii, and starspots in about a minute.","key_machinery":"The central machinery is a neural-network surrogate for the PHOEBE eclipsing-binary light-curve model, trained band by band using transfer learning from a pre-trained V-band model. The network's split architecture—first five physical parameters, then the concatenated spot and third-light parameters—is what allows it to cover the full ten-parameter space including starspots. A two-stage MCMC inverts the network to recover parameters from observed light curves, and a separate radius/potential subnetwork takes advantage of the geometric dependence on only mass ratio and fillout.","core_discovery":"The central claim is that a fully connected neural network can serve as a fast surrogate for both light-curve synthesis and parameter inversion in contact binaries. The architecture splits the ten input parameters into two groups: the first five (primary temperature, temperature ratio, mass ratio, inclination, fillout) pass through early layers, then the four spot parameters plus third light are concatenated before the final layers. Separate small networks output luminosities, radii, and potential, exploiting the fact that radii and potential depend only on mass ratio and fillout. Training data are uniformly sampled parameter sets used to generate synthetic light curves with PHOEBE, with a s","pith_inferences":["Editorial inference: the 3,541 OGLE parameters are mostly unvalidated against independent geometric or spectroscopic solutions; a robust test would compare the network's mass ratios and fillouts with radial-velocity mass ratios for a statistically meaningful subset.","Editorial inference: if real starspot patterns are more complex than the single-circular-spot PHOEBE model, the inferred spot latitude, longitude, radius, and temperature may absorb unmodeled asymmetries and should be treated cautiously until checked against spot-mapping observations.","Editorial inference: the transfer-learning recipe from the V band could plausibly extend to semi-detached or other eclipsing-binary geometries, though the training ranges, parameter split, and output heads would need re-tuning.","Editorial inference: the speed opens the possibility of on-the-fly parameter estimation during observing runs, allowing immediate follow-up decisions for newly discovered candidates."],"forward_implications":["Survey-scale parameter catalogues become feasible: the same pipeline applied to ASAS-SN, ZTF, TESS, or Gaia data could produce physical parameters for hundreds of thousands of contact binaries.","Analysis time drops from about 4.8 days to 82 seconds for a three-band, 729-point light curve, making million-sample studies practical on ordinary hardware.","Because the model includes all four starspot parameters, fast analyses no longer have to ignore the O'Connell effect or spot-induced light-curve asymmetries.","The packaged executable supports 19 standard filters, so observers can fit multi-band light curves from both large surveys and individual telescopes without running slow fitting codes.","Processing bands simultaneously lets the model exploit colour information across filters, which can help break degeneracies between parameters such as temperature ratio and inclination."],"supporting_citations":[{"why":"Introduces PHOEBE, the eclipsing-binary light-curve code used to generate the synthetic training light curves.","marker":"A. Prša & T. Zwitter 2005"},{"why":"PHOEBE 2, the version of the code that produced the synthetic light curves for this work.","marker":"A. Prša et al. 2016"},{"why":"Earlier machine-learning plus MCMC model for rapid contact-binary parameter estimation; this paper's baseline.","marker":"X. Ding et al. 2022"},{"why":"Latest prior model that added spot parameters; the new model is positioned as surpassing it by handling multiple bands and all four spot parameters.","marker":"K. Li & L.-H. Wang 2025"},{"why":"Source of the eight well-studied systems used to validate the model's inferred parameters and light-curve fits.","marker":"L.-H. Wang et al. 2024"},{"why":"OGLE contact-binary catalogue providing the periods, epochs, and target list used in the Section 4 application.","marker":"I. Soszyński et al. 2016"},{"why":"Stellar atmosphere models adopted when generating the synthetic training light curves.","marker":"F. Castelli & R. L. Kurucz 2004"},{"why":"Transfer-learning method used to train the 19 band-specific models from the pre-trained V-band model.","marker":"S. J. Pan & Q. Yang 2010"}],"fun_headline_variants":["Neural net fits contact binary light curves in seconds","AI model handles multi-band light curves of 3,541 binaries","Fast neural network solves contact binary parameters","One-minute neural network for contact binary light curves","Neural network tackles massive contact binary surveys"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The network learns only from PHOEBE-generated synthetic light curves drawn from chosen parameter ranges, so the whole pipeline inherits any systematic difference between that synthetic model and real contact-binary photometry.","fun_headline_variants_meta":{"raw":{"variants":["Neural net fits contact binary light curves in seconds","AI model handles multi-band light curves of 3,541 binaries","Fast neural network solves contact binary parameters","One-minute neural network for contact binary light curves","Neural network tackles massive contact binary surveys"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000144,"raw_usage":{"total_tokens":1045,"prompt_tokens":809,"completion_tokens":236,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":163}},"tokens_in":553,"tokens_out":236,"duration_ms":2999,"temperature":1.0,"reasoning_tokens":163,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T21:05:19.366868+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the model's mass ratios and temperature ratios for a sample of contact binaries against independent double-lined spectroscopic radial-velocity orbits; a systematic offset larger than the quoted uncertainties in q or T2/T1 would show the synthetic-training premise fails. A practical mid-step is to fit a subset of the 3,541 OGLE systems with a conventional W-D/PHOEBE analysis and check whether parameter differences are consistent with the claimed precision.","supporting_citations":[],"review_version":1}