{"id":"b52d9b17-7068-44bf-8369-28e1759ab009","arxiv_id":"2504.21085","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new automated pipeline detects 107 binary microlensing events with well-separated bumps in OGLE-IV data, split into 59 binary-lens and 48 binary-source candidates.","lead":"This paper builds an automated pipeline to find binary microlensing events in ten years of OGLE bulge data, selecting light curves with two well-separated brightness bumps. It reports 107 such events, half modeled as binary lenses and half as binary sources, as a step toward measuring the Milky Way bulge's initial mass function.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported tE peak is prior- and cut-dominated: a Gaussian tE prior from Mróz et al. (2020) is imposed in the fits and then echoed in the Fig. 7 histogram, while the q and flux-ratio flatness are raw, efficiency-uncorrected sample statistics.","rationale":"The paper's strongest claim combines detection success with quantitative distributions. The benchmark concern identified by the reader is legitimate, but the paper has partial independent support: most events were alerted by OGLE, MOA, or KMTNet, and the 12 visually identified false positives were removed. The place where the argument is least secure is the conversion of fitted posteriors into the claimed tE and ratio distributions. Section 4 explicitly imposes a Gaussian tE prior; Section 5 states that the histogram follows it. Therefore the 35–40 d peak is not an empirical finding about the sample unless the prior is shown to be non-influential. The post-fit tE cuts and the model-preference tie-breaker make the prior influence larger, and the absence of efficiency corrections means the flat q and flux-ratio histograms are at best raw sample summaries. I would keep the reader's CONDITIONAL verdict: the sample may be genuine, but the abstract and Fig. 7 need to be reframed as prior-informed raw statistics, with the prior dependence explicitly tested. No verdict change is needed because the reader already conditioned acceptance on selection-function caveats and on correcting overstatements in the abstract.","tokens_in":23680,"tokens_out":7376,"duration_ms":81271,"concrete_test":"Re-run the full fitting pipeline on the identical 107-event sample with the tE prior replaced by a flat prior in log tE over [1, 1000] d, and with model selection made without the 'derived tE closer to the peak of the prior' tie-breaker; then recompute the Fig. 7 histograms with and without the tE < 200/120 d cuts. If the 35–40 d peak shifts by more than about 5 d, or if the q and source-flux-ratio histograms change shape beyond the 1σ Poisson band, the headline distributions are prior- and selection-dominated. If the histograms are statistically unchanged, the concern is retired. Optionally, inject simulated 1L2S and 2L1S events over a grid in q, flux ratio, tE, and separation to compute efficiency weights and apply them to the raw histograms.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In the fitting section (Section 4), the authors state: 'A Gaussian prior in tE follows the distribution derived in Mróz et al. (2020), which peaks around 30 d.' Section 5 then reports that the resulting tE histograms in Fig. 7 'closely follow the adopted prior' and takes the 'peak around 35–40 d' as a result. Because the prior enters the likelihood for every event and also enters the tie-breaking rule between 1L2S and 2L1S models ('derived tE closer to the peak of the prior'), the headline tE distribution is not measured from the data; it is a posterior-weighted reflection of the input prior. The post-fit cuts tE < 200 d (Table 2) and tE < 120 d (Table 3) additionally truncate the tail, and the q and source-flux-ratio histograms are raw observed histograms with no detection-efficiency correction and no propagation of model-classification uncertainty. The paper explicitly defers efficiency corrections to future work, so the abstract's phrasing 'with a distribution of Einstein timescales around 35–40 d and flat distributions for mass ratio and source flux ratio' presents sample statistics as though they were population results. This is internal circularity, not a disagreement with an external prior. The benchmark-tuning issue identified by the reader is real, but the independent alert cross-match and visual inspection partially mitigate it; the prior issue directly affects the numerical claims in the abstract and Fig. 7.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a fully automated pipeline for detecting binary microlensing events with two well-separated bumps in ten years of OGLE-IV bulge data. The detection algorithm modifies the Mróz et al. (2017, 2019) event finder with two moving windows, baseline slope correction, and time binning, with thresholds tuned on a 29-event benchmark sample. The authors fit 1L2S and 2L1S models to each candidate with MCMC and nested sampling, report 107 events (59 with preferred 2L1S and 48 with preferred 1L2S), and cross-match against OGLE, MOA, and KMTNet alerts. They report tE distributions peaking around 35–40 d and flat mass-ratio and source-flux-ratio distributions, and state that after efficiency corrections these will constrain the bulge IMF.","tokens_in":23923,"tokens_out":5949,"duration_ms":60597,"significance":"The paper's main value, if the sample is accepted, is as a large, homogeneously selected catalog of well-separated binary microlensing events, with a documented pipeline and public fit parameters. The cross-match with external alert systems and the explicit description of detection modifications are strengths. However, the headline population claims are not supported as stated: the tE peak is prior-dominated, and the flat ratio histograms are raw sample statistics without efficiency corrections. Because the catalog is useful and the issues can be fixed by reframing the statistical claims, the paper merits revision rather than rejection.","major_comments":[{"comment":"The tE histograms are prior-dominated rather than measured. The text states that a Gaussian prior in tE follows the Mróz et al. (2020) distribution peaking around 30 d, and the 1L2S/2L1S tie-breaker explicitly favors the solution with tE closer to the prior peak. The paper itself notes in Section 5 that the histograms \"closely follow the adopted prior,\" and the post-fit cuts tE < 200 d and tE < 120 d (Tables 2 and 3) further truncate the tail. The abstract's \"distribution of Einstein timescales around 35–40 d\" therefore reports a prior-weighted posterior summary, not a measured distribution. Please report the data-constrained information (e.g., likelihood profiles or fits with a weakly informative prior) or explicitly reframe the statement as a consistency check with the adopted prior, and adjust the abstract accordingly.","section":"Section 4; Section 5, Fig. 7; Abstract"},{"comment":"The flat mass-ratio and source-flux-ratio distributions are presented as results without detection-efficiency corrections or model-classification uncertainty. The paper explicitly defers efficiency corrections to future work, and the classification of several events with |Δχ2/dof| < 0.001 required additional evidence, blending-flux, and tE criteria. The abstract's \"flat distributions\" thus overstates what has been measured. Please present these as observed-sample histograms, add uncertainty estimates (e.g., Poisson uncertainties or classification probabilities), and state clearly that population-level flatness will be tested once efficiency corrections are applied.","section":"Section 5, Fig. 7; Tables 7–8; Abstract"},{"comment":"The detection thresholds were optimized to recover all 29 benchmark events of Table 1, and those same benchmark events are included in the reported 107-event sample. This makes the claim that the tools were effective partly circular: recovery of the benchmark subset is guaranteed by construction. Please quantify the detection statistics for benchmark versus non-benchmark events, or use a held-out validation set, and wherever possible compare the final sample against external alert catalogs independently of the tuned thresholds.","section":"Section 3.2, Tables 2 and 3"}],"minor_comments":[{"comment":"The axis labels contain truncated text: \"ag/yr\" should be \"mag/yr,\" and the magnitude axis label appears to be missing an \"n\".","section":"Fig. 2"},{"comment":"The word \"analy ed\" should be \"analyzed.\"","section":"Fig. 1 caption"},{"comment":"The reference for Skowron and Gould (2012) is listed as arXiv:2310.12069, which does not match the 2012 work (ApJ, 744, 129; arXiv:1203.1034); please correct it.","section":"References"},{"comment":"The text contains the typo \"Y eeet al.\" and should read \"Yee et al.\"","section":"Section 5, paragraph on K2/Spitzer data"},{"comment":"Only the first 15 rows of each table are printed; please state explicitly that the full tables are available in the journal archive and consider also providing a machine-readable version in the arXiv source.","section":"Tables 5 and 6"},{"comment":"The caption explains the HJD format, but the text should state consistently that times are HJD−2450000, since this convention is used in later tables and fits.","section":"Table 1 caption"}],"recommendation":"major_revision","confidential_remarks":"The central issue for me is that the abstract sells prior-echoing histograms as measurements. I would ask the authors to add data-driven checks (e.g., uninformative-prior fits or likelihood-based intervals) or to revise the abstract and conclusions so that the population claims are clearly deferred to the efficiency-corrected follow-up paper. The benchmark-tuning circularity is secondary but should be quantified rather than asserted. The paper is otherwise a solid methods/catalog contribution and fits the journal; with these revisions I would be happy to see it published."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read of arXiv:2504.21085. Bottom line: this is a useful paper and should go to review, but the abstract's tE claim is not as clean as it reads.\n\nWhat's new and good: it adapts the Mróz et al. single-bump search to hunt for two well-separated bumps, adds slope correction and time binning, and runs the modified pipeline over 10 years of OGLE-IV fields. The result is a catalog of 107 binary events (59 2L1S, 48 1L2S), most without prior fits in the literature. That is a solid, homogeneous sample and probably the largest of this morphology for the bulge. The detection steps are described with enough detail to reimplement: threshold tables, bump criteria, post-fit cuts. The modeling uses standard public tools (MulensModel, emcee, UltraNest) and the event tables include parameter posteriors. The cross-match against OGLE/MOA/KMTNet alerts gives real confidence that these are genuine microlensing events. Credit where due: the catalog alone is a contribution.\n\nNow the soft spots, in rough order of importance.\n\nThe tE histogram in Fig. 7 is not a measurement. Section 4 states a Gaussian prior on tE following Mróz et al. (2020), peaking around 30 d, and Section 5 says the resulting histograms 'closely follow the adopted prior.' Then the abstract reports 'a distribution of Einstein timescales around 35–40 d' as if it came from the data. That is internal circularity. The prior is also used in tie-breaking between 1L2S and 2L1S models. So the abstract overstates. The fix is straightforward: report it as a posterior or prior-dependent result, and soften the abstract.\n\nRelated: the q and source-flux-ratio flatness are raw observed histograms, no efficiency correction. The paper explicitly defers that, which is fine, but the abstract should not present sample statistics as population statements. Same caveat.\n\nThe benchmark circularity is real but mild. Thresholds were tuned to recover 29 visually selected benchmark events; those events then appear in the final 107. That means detection completeness is not blind. But the independent alert cross-match and the large number of newly found events make the sample bona fide. The 'fully-automated' wording is a bit generous given visual inspection and the manual removal of 12 candidates.\n\nMinor: the Skowron and Gould (2012) reference has a wrong arXiv number. Fix it.\n\nReproducibility is limited by not releasing detection code; not a killer but worth asking.\n\nWho this is for: anyone working on bulge binaries, microlensing population statistics, or preparing for Roman. The catalog is the product; the population distributions need the promised efficiency corrections before they should be used.\n\nRecommendation: yes, send to a serious referee. Requires minor-to-moderate revision, mostly language in the abstract and a clear statement that the tE distribution is prior-influenced. The core detection and modeling work holds up.","headline":"A genuinely useful catalog and pipeline paper; the headline tE peak is prior-driven, so treat the abstract's numbers as sample stats, not measurements.","tokens_in":24613,"tokens_out":2872,"would_cite":true,"duration_ms":29138,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An automated search of ten years of OGLE-IV data finds 107 binary microlensing events with flat mass-ratio and flux-ratio distributions.","keywords":["binary microlensing","OGLE-IV","Galactic bulge","initial mass function","mass ratio distribution","automated event detection","Einstein timescale","binary source"],"falsifier":"Re-run the detection on the same 121 fields with the secondary-bump significance threshold lowered (e.g., from χ>60 to χ>40 in high-cadence fields) and inspect all newly added candidates by eye; if the added events are numerous and shift the mass-ratio or timescale distributions, the reported flat distributions are an artifact of the tuned cuts, and if they are absent the sample is robust to small threshold changes.","tokens_in":1818,"feed_emoji":"🔭","tokens_out":6176,"duration_ms":118887,"temperature":0.7,"pith_summary":"The paper claims that binary microlensing events with two well-separated bumps can be found automatically, without visual scanning, in ten years of OGLE-IV bulge observations, and that the resulting sample is large enough to begin measuring bulge binary statistics. It reports 107 such events from 121 fields, with 59 better fit by a binary-lens model and 48 by a binary-source model. The Einstein timescales of the sample cluster near 35–40 days, and the mass-ratio and source-flux-ratio distributions are consistent with being flat between 0 and 1. This matters because the bulge initial mass function can be recovered from luminosity functions only if the unresolved binary fraction and mass-ratio distribution are known; the authors position this catalog as the ingredient that supplies those statistics after detection-efficiency corrections.","feed_headline":"Automated search bags 107 binary microlensing events from OGLE-IV","feed_subtitle":"Double-bump light curves yield flat mass-ratio and flux-ratio distributions for bulge IMF studies.","key_machinery":"The load-bearing mechanism is a two-bump detection pipeline built on the previous single-bump OGLE search, with three additions: two moving 360-day windows chosen to minimize baseline scatter, a linear slope correction to the baseline flux, and time-binned bump durations with iterative bump removal (up to three bumps). Detection thresholds on significance, amplitude, and duration are tuned separately for high- and low-cadence fields to recover all 29 benchmark events. The fitting half of the pipeline uses an algebraic mapping from two independent PSPL fits to the starting parameters of a 2L1S model — mass ratio q = (tE,2)^2 / (tE,1)^2, combined tE, separation s, and trajectory angle α — and then refines both 1L2S and 2L1S solutions with MCMC and nested sampling, choosing between them by χ² difference aided by Bayesian evidence. This combination converts a search for 'two bumps in a noisy light curve' into a catalog of physically parameterized binary events.","core_discovery":"The central discovery is a method and its demonstration: a modified version of the OGLE-IV event-finding algorithm, retargeted to find two bumps instead of one and tuned against 29 visually selected benchmark events, recovers a bona-fide sample of 107 binary microlensing events with well-separated bumps and no caustic crossings. For every candidate the pipeline fits both a single-lens/binary-source (1L2S) and a binary-lens/single-source (2L1S) model using MCMC and nested sampling, choosing the preferred model by goodness of fit supplemented by Bayesian evidence, blending flux, and timescale. The paper's headline results are that the two model classes split nearly evenly (59 vs 48), that the Einstein timescale distribution peaks around 35–40 days in both classes, and that the mass ratio (for lenses) and source flux ratio (for sources) are approximately flat over the range 0–1. These distributions are explicitly preliminary: the authors state that detection-efficiency corrections will be applied in later papers before the binary fraction and mass-ratio statistics are used to constrain the bulge IMF.","pith_inferences":["If the flat mass-ratio distribution survives efficiency corrections, it would imply that the bulge's binary lens population has essentially no preference for equal masses over the range 0–1, a property that could be compared with local binary surveys to test whether bulge binaries resemble disk binaries.","The same pipeline, applied to other wide-field surveys with different cadence or to caustic-crossing light curves, could produce a homogeneous binary fraction across the bulge instead of the mixed literature sample used here.","The slope-correction step may systematically add long-baseline events that previous searches missed; if so, reapplying it to single-lens searches would revise the high-tE tail of the ordinary event sample.","The nearly equal 1L2S/2L1S split suggests that binary-source events contribute as much as binary-lens events to the separated-bump morphology; a future measurement of the 1L2S fraction could directly probe the binary fraction of bulge source stars, a separate population from the lenses."],"forward_implications":["The 107-event catalog itself is the largest homogeneous OGLE-IV sample of well-separated-bump binary events; its fitted parameters are the raw material for binary fraction and mass-ratio measurements.","A flat mass-ratio distribution over q in (0,1) means that, among detectable wide-separation binary lenses, low-mass companions are as common as equal-mass companions, contrary to a peaked preference.","The 59/48 split between 2L1S and 1L2S shows that for well-separated bumps the two physical interpretations are nearly equally frequent, so population studies that ignore binaries must account for both.","Because the tE distribution peaks at 35–40 days, matching previous PSPL results, binary events should not skew the overall event-timescale distribution once counted.","After the planned detection-efficiency corrections, the binary fraction and mass-ratio distribution can be combined with the observed bulge luminosity function to constrain the low-mass end of the bulge IMF."],"supporting_citations":[{"why":"Supplies the OGLE-IV photometric data and instrument description underlying the search.","marker":"Udalski et al. 2015"},{"why":"Provides the event-finding algorithm and the 121-field bulge sample that this work modifies.","marker":"Mróz et al. 2019"},{"why":"Documents the original single-bump selection algorithm and high-cadence field definitions inherited by the pipeline.","marker":"Mróz et al. 2017"},{"why":"Provides MulensModel, the package used to compute binary lens and binary source magnifications.","marker":"Poleski and Yee 2019"},{"why":"Supplies the emcee MCMC sampler used for posterior fitting of the models.","marker":"Foreman-Mackey et al. 2013"},{"why":"Supplies UltraNest nested sampling used to evaluate Bayesian evidence for model selection.","marker":"Buchner 2021"},{"why":"Provides the polynomial root solver used in binary-lens magnification calculations.","marker":"Skowron and Gould 2012"},{"why":"Supplies the Gaussian prior on the Einstein timescale and the peak timescale used for comparison.","marker":"Mróz et al. 2020"},{"why":"Provides the bulge luminosity function measurement that the binary statistics are meant to constrain.","marker":"Calamida et al. 2015"}],"fun_headline_variants":["Automated pipeline finds 107 binary microlensing events in OGLE-IV","107 binary microlensing events unearthed by automated OGLE-IV search","Automated search yields 107 double-bump microlensing events","OGLE-IV automated search nets 107 binary microlensing events"],"cache_read_input_tokens":26496,"weakest_assumption_plain":"The 29 benchmark events, chosen by eye from an internal alert compilation, are a correct and representative set of binary microlensing events; the detection thresholds were tuned to recover exactly these events, so any error or incompleteness in that list propagates into the 107-event sample and the reported distributions.","fun_headline_variants_meta":{"raw":{"variants":["Automated pipeline finds 107 binary microlensing events in OGLE-IV","107 binary microlensing events unearthed by automated OGLE-IV search","Automated search yields 107 double-bump microlensing events","OGLE-IV automated search nets 107 binary microlensing events"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0007,"raw_usage":{"total_tokens":3237,"prompt_tokens":1098,"completion_tokens":2139,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":714,"completion_tokens_details":{"reasoning_tokens":2061}},"tokens_in":714,"tokens_out":2139,"duration_ms":13520,"temperature":1.0,"reasoning_tokens":2061,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:14:24.597178+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the detection on the same 121 fields with the secondary-bump significance threshold lowered (e.g., from χ>60 to χ>40 in high-cadence fields) and inspect all newly added candidates by eye; if the added events are numerous and shift the mass-ratio or timescale distributions, the reported flat distributions are an artifact of the tuned cuts, and if they are absent the sample is robust to small threshold changes.","supporting_citations":[{"cited_title":"2019, Astronomy and Computing, 26,","cited_arxiv_id":null,"evidence_quote":"Provides MulensModel, the package used to compute binary lens and binary source magnifications."},{"cited_title":"2015, ApJ, 810,","cited_arxiv_id":null,"evidence_quote":"Provides the bulge luminosity function measurement that the binary statistics are meant to constrain."}],"review_version":1}