{"id":"7609591d-fd45-4969-93cf-7888f3928555","arxiv_id":"2607.20260","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A unified transformer-based pipeline detects 99.9% of recoverable simulated microlensing events and outperforms literature hard cuts in the short-duration finite-source regime with amortized neural posterior inference.","lead":"This paper trains a shared transformer network to both detect gravitational microlensing events and infer their parameters from simulated Roman Space Telescope light curves. On simulated data it reports catching 99.9% of recoverable events, including short finite-source events that standard threshold-based searches miss.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Detection rates are measured on events that already pass §2.2 recoverability filters (peak SNR>5, sampling criteria); the paper does not quantify the filter pass rate, so 99.9% efficiency and ~95% in the ρ≳5 bin are conditional on a selected high-quality sub-population, not survey-level.","rationale":"The reader's weakest_assumption already identifies the recoverability filters and the Gaussian-only, no-confounder simulation as the main condition on the headline numbers. I agree and find this to be the single most load-bearing concern: the paper's central claim — that the Evidence Network provides calibrated detection for Roman's microlensing survey and recovers finite-source events that hard cuts miss — depends on the simulation faithfully representing the survey. The paper is honest about several limitations (no astrophysical false positives, box priors, fixed 20-day window), and it provides code and a transparent comparison with oracle hard cuts, which strengthens the relative claim. However, the absolute detection numbers are conditioned on an in-distribution test set that has already passed the §2.2 filters. The paper does not report the filter pass rate, so a reader cannot convert the reported TPR into a survey-level sensitivity. This is not an internal inconsistency: the network genuinely outperforms hard cuts on the filtered test set, and the hard-cut comparison is conservative (oracle parameters). But it is a real external-validity gap. The posterior calibration issue flagged by the reader (per-parameter 68% coverages ranging 0.55–0.83 in Table 4) is secondary to the detection claim and would not change the verdict. A single Monte Carlo calculation of the filter pass rate would settle whether the headline numbers are representative of the prior population. If the pass rate is high, the concern lands weakly; if low, the abstract's 99.9% should be qualified as a conditional sub-population figure. Since the reader already returned CONDITIONAL for essentially this reason, I recommend no change to the verdict.","tokens_in":17168,"tokens_out":12949,"duration_ms":107102,"concrete_test":"Run a Monte Carlo over the §2.1 priors (t0, u0, log tE, log ρ, fs) using the §2.2 augmentation pipeline (seasonal gaps, dropout, noise) to compute the fraction of simulated events that satisfy the recoverability filters: (i) ≥5 points within 1.5tE of t0, (ii) ≥5 points beyond 3tE, (iii) peak SNR >5. Report this pass rate overall and specifically for the ρ≳5 bin. Then compute the effective survey detection rate as (pass rate) × (reported TPR in that bin). If the pass rate in the ρ≳5 bin is materially below 1 (e.g., <50%), the headline 99.9% and ~95% are upper bounds for a clean, selected sub-population rather than survey-level sensitivities.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline detection numbers are computed on test light curves that pass the §2.2 recoverability filters: ≥5 points within 1.5tE of the peak, ≥5 points beyond 3tE, and peak SNR >5. These filters are applied during training and, per §3.1, to the test set as well. The paper never reports what fraction of events drawn from the §2.1 priors actually pass these filters. This matters quantitatively: in the extreme finite-source regime (ρ≳5), the peak magnification is small (A_peak ≈ 1 + 2/ρ², so ~1.08 for ρ=5). With per-point noise σ drawn uniformly from [0.001, 0.02], an event with A_peak = 1.08 needs σ < 0.016 to satisfy SNR >5; a substantial fraction of the prior volume will fail this criterion. Thus the reported ~95% detection rate in the ρ≳5 bin, and the overall 99.9%, apply only to events already selected for high SNR and adequate sampling. The network is also trained and evaluated on the same simulation pipeline, making the test set in-distribution. Roman survey data will include lower-SNR events, events with unfavorable gap placement, and astrophysical contaminants; the paper's own §4.2 acknowledges sensitivity to distribution shift but does not quantify it. If the recoverability filter pass rate is low, the effective survey-level detection efficiency is the reported rate multiplied by the pass rate — a number that could be substantially below 99.9%. This is the main gap between the paper's conditional claims and the broader claim of providing detection for Roman's microlensing survey.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a unified framework for gravitational microlensing detection and parameter inference. Detection is cast as Bayesian model comparison via Evidence Networks trained on binary signal/noise labels to estimate Bayes factors, and inference is performed with Neural Posterior Estimation; both share a transformer encoder that handles irregularly sampled light curves without imputation. On simulated Roman Space Telescope-like FSPL light curves with data augmentation and Gaussian noise, the Evidence Network achieves 99.9% detection efficiency at a detection threshold that yields zero false positives among 5,000 noise events (95% upper limit 6e-4), outperforming literature and re-tuned hard-cut detectors. The gains are largest in the extreme finite-source regime (rho ≳ 5), where reported detection efficiency is ~95% versus ~65% for hard cuts. The paper also validates the Evidence Network on a toy model with analytic Bayes factors, performs a shared-encoder ablation, and benchmarks NPE wall-clock speed against MCMC.","tokens_in":17619,"tokens_out":6148,"duration_ms":49493,"significance":"If the reported numbers hold, this is a useful and timely contribution: it is the first application of Evidence Networks to irregularly sampled astrophysical time series, and the combination of amortized detection and inference with a shared transformer encoder is a natural fit for Roman's microlensing survey. The toy-model validation against analytic Bayes factors (Appendix A) is a strong check that binary labels alone can recover calibrated Bayes-factor estimates. The hard-cut comparison is conservative, since the hard cuts are evaluated with oracle (true) parameters, which strengthens the claimed advantage in the finite-source regime. The shared-encoder ablation (Appendix C) and the wall-clock benchmark (Appendix D) are also valuable practical additions. However, the headline detection efficiency and the posterior calibration claims are conditional on in-distribution, recoverability-filtered, Gaussian-noise simulations; these conditionality issues need to be quantified and reconciled before the results can support the broader survey-level claims made in the abstract.","major_comments":[{"comment":"The 99.9% detection efficiency (and the ~95% rate in the rho≳5 bin) is measured on test events that already pass the §2.2 recoverability filters: at least 5 points within 1.5tE of the peak, at least 5 points beyond 3tE, and peak SNR > 5. The paper never reports the fraction of events drawn from the §2.1 priors that actually pass these filters after the §2.2 augmentation. This matters quantitatively: in the extreme finite-source regime the peak magnification is only ~1+2/rho^2, so a substantial fraction of the prior volume will fail the peak-SNR>5 criterion unless the noise realization is unusually favorable. As written, the 99.9% and ~95% figures are conditional on a pre-selected, high-quality sub-population, not survey-level detection rates. Please report the filter pass rate as a function of the model parameters (especially rho), and either report an unconditional efficiency that inclu","section":"§2.2, §3.1"},{"comment":"The claim of 'good calibration' from the TARP coverage test is in tension with the per-parameter coverage reported in Table 4. The 68% credible-interval coverage is 0.826 for t0, 0.636 for u0, 0.661 for log10 tE, 0.553 for log10 rho, and 0.643 for fs — all below the nominal 0.68, and for log10 rho substantially so. TARP coverage is a joint coverage statistic and can pass even when individual parameters are miscalibrated in different directions. The manuscript should reconcile the TARP statement with these per-parameter numbers, report coverage separately in degenerate regions, and avoid the unqualified 'well-calibrated posteriors' conclusion unless the per-parameter miscalibration is corrected or explicitly discussed as a known limitation.","section":"§3.2, Table 4, Appendix E"},{"comment":"The sentence 'we adopt a detection threshold of log10 K > 0.8 (equivalently, K > 6.3, or approximately 1 in a million chance of misclassification under ideal conditions)' is incorrect as written. With equal prior odds, K=6.3 corresponds to a posterior signal probability of roughly 0.86, not 10^-6. The threshold was actually selected by requiring zero false positives on the validation set; the '1 in a million' phrasing appears to conflate a binomial upper limit on the validation false-positive count with a per-event misclassification probability. Please correct this sentence and state explicitly how the operating point was selected and what it means probabilistically.","section":"§3.1"}],"minor_comments":[{"comment":"The discussion correctly acknowledges that the reported false-positive rate is not survey-level because astrophysical contaminants and non-Gaussian noise are absent from the training and test sets. It would be helpful to also acknowledge that the test set is generated by the same simulation pipeline as the training set, so the reported detection efficiency is an in-distribution upper bound. A brief statement on expected sensitivity to distribution shift would make the limitations section more complete.","section":"§4.2"},{"comment":"The caption states that re-tuned hard cuts are shown as dashed lines, but the legend in the figure body does not clearly distinguish as-published and re-tuned versions. Please ensure the line styles and legend entries match the caption.","section":"Figure 2 caption"},{"comment":"The text says '95% credible-interval coverage remaining between 0.90 and 0.97 throughout,' which matches Table 4. However, the 68% coverage values are much lower and deserve a direct comment in the main text, not only in the appendix. This is related to the second major comment.","section":"Appendix E, Table 4"},{"comment":"The speedup factor of ~16,000x is reported for a specific MCMC configuration (32 walkers, 1,000-step burn-in) with a CPU-side likelihood evaluator. The text acknowledges that optimizing MCMC could reduce the gap. It would be useful to state explicitly that the wall-clock comparison is not a tuned MCMC benchmark, to avoid over-generalization.","section":"§2.5"}],"recommendation":"major_revision","confidential_remarks":"The recoverability-filter issue and the per-parameter calibration discrepancy are both concrete and addressable within the manuscript's scope. The core methodology is sound, and the toy-model validation and oracle hard-cut comparison are strong points. I would be comfortable with acceptance after these two issues are quantified and reconciled."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the main relative claim in this paper holds up. On simulated Roman-like light curves, a transformer-based Evidence Network detects finite-source microlensing events at 99.9%, versus 95–97% for literature hard cuts, and the comparison is designed the right way: the hard cuts are handed the true lensing parameters and their thresholds are re-tuned on the same validation generator, and they still lose. The advantage concentrates in the ρ≳ 5 regime, which is exactly where free-floating-planet science sits. That is a new result, and it is not an artifact of the evaluation.\n\nWhat is well done: the toy-model validation with analytic Bayes factors comes within 0.001 AUC of the oracle; the Appendix C ablation shows the shared encoder is doing real work, not decoration; the posterior discussion is honest about the u0–ρ degeneracy; and §4.2 says plainly that the false-positive rate is not a survey-level number. The code is public. This is a carefully built pipeline paper.\n\nSoft spots, in order. First, the 99.9% headline (and the ~95% in the ρ≳ 5 bin) is computed on events that passed the §2.2 recoverability filters — at least five points within 1.5tE, five beyond 3tE, peak SNR>5 — and the paper never reports the filter pass rate. The stress-test note is right: at ρ=5 the peak magnification is about 1.08, so SNR>5 requires σ≲ 0.016; with σ drawn uniformly up to 0.02, roughly a fifth of the prior fails on noise alone, and the sampling criteria will remove many short-tE events on top of that. The survey-level efficiency is the reported rate times an unknown pass rate, and that product could be far below 99.9%. Second, “calibrated posteriors” overstates Table 4: the 68% credible-interval coverage for u0, log10 ρ, and fs is between 0.55 and 0.64, which is real overconfidence at the 1σ level; the TARP curve in Figure 3 hides it. Third, the test set comes from the same generator as training, so this is an in-distribution result; the paper acknowledges distribution-shift sensitivity but does not quantify transfer to real Roman noise and contaminants. None of this breaks the relative claim, but it should stop anyone from quoting 99.9% as a survey-level detection rate.\n\nThis paper is for people building Roman-era detection pipelines and anyone applying Evidence Networks to irregular time series. It deserves a serious referee. I would send it out with two requests: report the recoverability-filter pass rate, and re-state the abstract's efficiency and calibration claims so they match the conditional numbers.","headline":"Solid methods paper: the relative claim (Evidence Network beats oracle hard cuts on finite-source microlensing) holds, but the 99.9% headline is conditional on a recoverability filter whose pass rate is never reported.","tokens_in":18065,"tokens_out":5581,"would_cite":true,"duration_ms":43529,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a single neural pipeline, trained only on binary labels, can detect gravitational microlensing events via calibrated Bayes factors and infer their parameters in milliseconds, recovering low-magnification finite-source","keywords":["gravitational microlensing","Bayes factor estimation","simulation-based inference","transformer encoder","amortized posterior estimation","finite-source events","free-floating planets","Roman Space Telescope"],"falsifier":"Run the trained Evidence Network on a simulated test set that adds stellar flares, eclipsing binaries, cataclysmic variables, and correlated noise to the noise-only distribution, and measure the false-positive rate at the same log10 K > 0.8 threshold; if it rises above about 6e-4, the survey-level claim fails. Alternatively, repeat detection on a test set that excludes the recoverability filters (fewer than 5 points near peak or peak SNR below 5) to measure how much of the 99.9% efficiency depends on those cuts.","tokens_in":17045,"feed_emoji":"🔭","tokens_out":4308,"duration_ms":34741,"temperature":0.7,"pith_summary":"The paper aims to replace the hard, threshold-based criteria used by current microlensing surveys with a unified neural framework. It trains an Evidence Network to estimate Bayes factors for signal versus noise directly from binary-labeled simulated light curves, and shares its transformer encoder with a Neural Posterior Estimator so detected events get amortized parameter posteriors. On simulated Roman Space Telescope data, the detector reaches 99.9% detection efficiency at a false-positive rate below 6e-4 on clean Gaussian noise, with the largest gains in the finite-source regime (ρ≳5) that dominates short free-floating planet events. If these numbers hold on real data, the framework would both recover events that threshold cuts systematically miss and deliver posteriors fast enough for real-time survey analysis.","feed_headline":"Neural detector catches 99.9% of simulated microlensing events","feed_subtitle":"Shared transformer also returns calibrated parameter posteriors in milliseconds, recovering finite-source events hard cuts miss.","key_machinery":"The Evidence Network with the leaky parity-odd power (l-POP) exponential loss, which turns binary labels into calibrated Bayes factors; and the shared transformer encoder with a classification token that aggregates irregularly-sampled light-curve points via attention masking, feeding both the detection head and the masked autoregressive flow (Neural Posterior Estimation) for amortized posterior inference.","core_discovery":"The central claim is that microlensing detection can be framed as Bayesian model comparison and solved by learning the Bayes factor from binary labels, and that this detection can share a single transformer embedding with parameter inference. The Evidence Network, trained with the l-POP exponential loss on equal signal/noise simulations, outputs calibrated Bayes factors that outperform literature hard cuts (including oracle-parameter versions) across the finite-source point-lens parameter space, especially for large source radius ρ≳5 where it holds about 95% detection versus about 65% for hard cuts. The shared encoder plus a masked autoregressive flow yields posteriors that pass TARP coverag","pith_inferences":["Editorial: If real Roman fields contain stellar flares, eclipsing binaries, or correlated noise, the network's false-positive rate will almost certainly rise above 6e-4; the paper's published upper bound applies only to a clean Gaussian-noise test set, so survey-level performance depends on retraining with contaminant classes.","Editorial: The recoverability filters used when generating training data (minimum points near peak, baseline coverage, peak SNR >5) mean the reported 99.9% efficiency characterizes a pre-selected recoverable sub-population; survey-level completeness over all events is an open question.","Editorial: The same Evidence Network + NPE architecture could be applied to other transient searches (supernovae, exoplanet transits, variable-star classification) whenever a forward simulator can generate binary-labeled training data.","Editorial: Because the model uses box-uniform priors rather than a Galactic population model, per-event Bayes factors and posteriors may be biased when applied to the real event population; hierarchical population-level inference would be needed to convert them into survey-yield predictions."],"forward_implications":["If the network's performance holds on real data, Roman's microlensing survey will recover a larger fraction of short-duration, finite-source events (free-floating planets) than threshold-based pipelines.","Detection becomes a calibrated Bayesian comparison: thresholds can be set to target false-positive rates or incorporate prior odds, replacing ad-hoc Δχ² cuts.","Parameter inference is amortized: once trained, posteriors for new events are produced in milliseconds rather than minutes, enabling real-time analysis of billions of light curves.","The shared transformer encoder means features transfer from inference to detection: a detection head trained on the frozen inference encoder reaches the same detection rate at a fraction of training cost.","Because the transformer handles gaps and irregular sampling without imputation, the same pipeline can be applied to other irregularly-sampled time-domain surveys."],"fun_headline_variants":["Learned Bayes factors catch 99.9% of simulated microlensing","Neural net recovers finite-source microlensing hard cuts miss","Bayesian detector for microlensing improves finite-source events","Transformer-based Bayes factors for microlensing detection","Microlensing detection via learned Bayes factors hits 99.9%"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline detection and false-positive numbers assume the test light curves contain only microlensing signals plus Gaussian noise, and that the same simulation pipeline, including the recoverability filters, produces both training and test data; real Roman fields will include flares, eclipsing binaries, variables, and instrumental systematics that the network has never seen.","fun_headline_variants_meta":{"raw":{"variants":["Learned Bayes factors catch 99.9% of simulated microlensing","Neural net recovers finite-source microlensing hard cuts miss","Bayesian detector for microlensing improves finite-source events","Transformer-based Bayes factors for microlensing detection","Microlensing detection via learned Bayes factors hits 99.9%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000374,"raw_usage":{"total_tokens":1816,"prompt_tokens":708,"completion_tokens":1108,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":452,"completion_tokens_details":{"reasoning_tokens":1020}},"tokens_in":452,"tokens_out":1108,"duration_ms":8329,"temperature":1.0,"reasoning_tokens":1020,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T10:18:40.463848+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the trained Evidence Network on a simulated test set that adds stellar flares, eclipsing binaries, cataclysmic variables, and correlated noise to the noise-only distribution, and measure the false-positive rate at the same log10 K > 0.8 threshold; if it rises above about 6e-4, the survey-level claim fails. Alternatively, repeat detection on a test set that excludes the recoverability filters (fewer than 5 points near peak or peak SNR below 5) to measure how much of the 99.9% efficiency depends on those cuts.","supporting_citations":[],"review_version":1}