{"id":"e4b709b2-fc68-42fe-9936-4e93981010be","arxiv_id":"2607.14450","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Low-acceleration wide binary stars show a MOND-type velocity boost even after applying the stricter data cuts and triple-star models proposed by critics.","lead":"This paper re-examines Gaia wide-binary star data to see whether an apparent MOND-type gravity anomaly survives stricter data-quality cuts and better modeling of hidden companion stars. It finds the anomaly persists and matches numerical MOND predictions, and it argues that prior 'no anomaly' results came from biased cuts or small samples.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The hidden-companion fraction f_trip is calibrated only in the Newtonian bin and assumed constant; if f_trip grows with r_p/r_M in the MOND-regime subsample, the rising \\tilde v profile could be contamination rather than gravity.","rationale":"The reader's weakest-assumption diagnosis matches my own reading. The central claim that data-quality cuts and multiple-star modeling cannot remove the low-acceleration anomaly depends on f_trip being constant across the Newtonian-to-MOND transition. The paper provides valuable internal consistency checks — especially the d<150 pc sample with f_trip ≈ 0.03 and the demonstration that a global f_trip = 0.41 cannot simultaneously describe Newtonian and MOND regimes — but these do not close the loophole of a per-bin f_trip(r_p) increase induced by sample selection. The proposed slope test is the most direct way to decide whether hidden companions can account for the signal. I do not think this concern forces rejection or a stronger verdict than CONDITIONAL; the existing CONDITIONAL verdict is appropriate because the anomaly is plausible and partly supported by multiple tests, but independent confirmation of f_trip constancy is needed before full acceptance.","tokens_in":46098,"tokens_out":9017,"duration_ms":89726,"concrete_test":"Take the PSS d<150 pc, ruwe<1.2, ipd_frac_multi_peak=0, \\tilde v<1.5 sample used in Figs. 43–53 and re-run the acceleration-plane/\\tilde v-histogram fits with f_trip parameterized as f_trip = f0 + s(Δx0), where Δx0 is the bin offset from the Newtonian calibration bin (x0≈−8.0). Fit f0 and s jointly to the full data, then compare Δχ² relative to the constant-f_trip model and check whether the best-fit MOND excess vanishes (Newton becomes acceptable) when s is allowed to be nonzero. If s is >3σ and Newton fits, the anomaly is contaminated; if s is consistent with 0 and QUMOND is still preferred, the concern is settled.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.2 calibrates f_trip using the PSS effective triple model (Eq. 18) in one Newtonian bin (x0 ≈ −8.0) and then fixes it across all bins (Figs. 25, 29). The paper asserts f_trip is independent of outer separation because it characterizes a member star, but the actual sample cuts (ruwe<1.2, ipd_frac_multi_peak=0, distance limits) need not be f_trip-neutral as a function of r_p. Figure 9 explicitly shows f_flyby grows with r_p, so the sample composition changes across the probed range; the same could happen for hidden companions through, e.g., distance/angular-resolution covariances. The highest-quality d<150 pc sample has f_trip ≈ 0.03, and Figure 53 shows a single global f_trip = 0.41 overfits the Newtonian regime while fitting the MOND median, but this does not test a per-bin f_trip(r_p) increase. A sufficiently strong r_p-dependent f_trip could mimic δ_obs−newt > 0 at g_N ≲ 1e−9 without modified gravity. This is the weakest link in the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reexamines the wide-binary gravity tests from Gaia DR3 that have yielded conflicting conclusions about a low-acceleration anomaly. It focuses on two challenges raised by null-result studies: data quality cuts (especially the Banik cut on the uncertainty of the normalized velocity) and the modeling of hierarchical systems with hidden companions. The authors implement the Pittordis et al. (2025) 'PSS' triple model via an effective algorithm (Eq. 18), calibrate the hidden-companion fraction f_trip in the Newtonian regime, and then run acceleration-plane, v-tilde-distribution, and median-v-tilde-profile tests on three samples (Chae 2023, Banik et al. 2024, PSS 2025). Their central claim is that the low-acceleration anomaly (δ_obs-newt > 0 at g_N ≲ 1e-9 m/s^2) survives all reasonable choices of data quality cuts and multiple-star modeling, with a combined significance >5σ, and that the observed trend agrees with the realistic QUMOND orbit solutions of Pflamm-Altenburg (2025). They further argue that the Banik cut introduces an r_p-dependent bias, that Cookson et al. (2026) had too few MOND-regime binaries (N=61) to discriminate models, and that the PSS preference for Newton is traceable to an overestimated f_trip and an inaccurate MOND transition-regime model.","tokens_in":46415,"tokens_out":7950,"duration_ms":85800,"significance":"If the main claim holds, the paper is a significant contribution to the wide-binary gravity debate: it addresses the two most prominent criticisms (data quality and hidden companions) with a unified modeling framework and shows that a MOND-type anomaly remains. Its strengths include the calibration of f_trip in the Newtonian regime rather than fitting the low-acceleration signal, the use of external numerical QUMOND solutions rather than a model tuned to the data, and the explicit power calculation in Figure 39 showing that Cookson et al.'s sample is too small to distinguish Newton from boosted gravity. The Pflamm-Altenburg comparison is a genuinely new element that goes beyond earlier approximate MOND treatments. The main caveat, discussed below, is that the r_p-independence of f_trip is asserted rather than directly tested; this is the weakest link in an otherwise well-constructed analysis.","major_comments":[{"comment":"The central assumption that f_trip is independent of the outer projected separation is asserted at the end of §5.2.2, but it is not tested. f_trip is calibrated in one Newtonian bin (x0 ≈ -8.0) and then held fixed across all bins. However, the sample-defining cuts (ruwe<1.2, ipd_frac_multi_peak=0, CMD cut, distance limit) are applied to member stars, and the probability of passing those cuts could correlate with distance and stellar mass distributions that vary with r_p. Figure 9 demonstrates that a different contaminant fraction, f_flyby, changes strongly with r_p, so an r_p-dependent f_trip is not implausible. A per-bin increase of f_trip from ~0.1 to ~0.4 could in principle produce the observed rise in median v-tilde without modified gravity. Figure 53 excludes a single global f_trip=0.41 because it destroys the Newtonian-regime agreement, but it does not exclude a per-bin f_trip(r_p)","section":"§5.2, Eq. (18), Figs. 25, 29, 53"},{"comment":"The headline 'combined statistical significance of >5σ' is not defined. Multiple bins, three different tests, and several overlapping samples are used, so the effective number of independent trials is unclear. As written, the >5σ claim is not falsifiable because the reader cannot tell which comparisons are being combined and how. Please specify the exact combination rule — for example, a single pre-specified bin/test, a meta-analysis with a stated correlation model, or a false-discovery-rate control — and provide the resulting p-value. This is particularly important because the paper itself reports different significances in different tests (e.g., 2.3σ in one configuration of Figure 36, 4.6σ in Figure 31, etc.).","section":"§7 (summary bullets; Figs. 25, 29, 43)"},{"comment":"The paper states that I. Banik et al. (2024) used an erroneous Newtonian benchmark, with their ⟨v⟩=0.64 for α=1 being 0.55 and their 'Newtonian' consequently boosted by ≈1.37 in the effective gravitational constant. This is a serious claim about a published null result and is used in §5.4 to explain why Banik et al. preferred Newton. It is not needed for the internally calibrated measurement of the anomaly, but it is central to the paper's dismissal of one of the main contrary studies. Please provide a step-by-step reproduction of the relevant calculation, or a numerical table with the same inputs, so that the 0.55 value and the boost factor can be independently checked. If this is simply taken from previous papers, the derivation should be restated here.","section":"§3, Fig. 13 and §5.4"}],"minor_comments":[{"comment":"There are typographical issues: 'T ests' in the typeset title and 'f pb' in §5.1 where 'η_phot' is introduced. Please proofread the final PDF version.","section":"Title and §5.1"},{"comment":"The figure mixes theoretical prediction bands, sample-specific colored bands, and individual points with several curves. Please separate the Newtonian benchmark comparison into a dedicated panel or expand the legend, as the current figure is difficult to read.","section":"Fig. 13"},{"comment":"The acronyms 'PSS' and 'CMD' are used without expansion at first appearance. The Gaia parameter 'ipd_frac_multi_peak' is typeset inconsistently; define it once and use a uniform notation.","section":"§5.1 and abstract"},{"comment":"The paper honestly states that the distinction between realistic and approximate MOND solutions is 'not conclusive from the present studies.' This caveat should also appear in the abstract, where the agreement with Pflamm-Altenburg (2025) is currently presented more strongly than the body supports.","section":"§7"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a lengthy point-by-point rebuttal of Banik et al. (2024), Pittordis et al. (2025), and Cookson et al. (2026). The technical core is strong, but the r_p-independence of f_trip is the key assumption that needs to be defended or relaxed. The claim about Banik et al.'s Newtonian benchmark error is serious and should be independently verifiable. The paper is heavily self-cited, and several cited works are 'in press' or 'in preparation'; the editor may wish to confirm that the dependence on these works does not create circularity. The planned public release of the effective triple-model code is welcome; it should be available before publication to make the methodological claims reproducible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Ky, here's my read. The paper does two genuinely useful things. First, it implements the Pittordis triple model in Chae's acceleration-plane test and recalibrates f_trip using Newtonian-regime data, getting lower values (0.07–0.12) than Pittordis's 0.185. That is a real methodological advance, and the effective model reproduces the Pittordis v-tilde distributions well. Second, it applies the Cookson quality framework to a sample eight times larger and shows the flat median-v-tilde profile in Cookson is consistent with small-sample fluctuation; the R.A. quadrant split is a nice demonstration. The mock simulations showing the Banik cut biases v-tilde scaling are also solid and worth remembering.\n\nThe core result—an anomaly at g_N < 1e-9, matching Pflamm-Altenburg's QUMOND rather than Newton—is internally consistent. The Newtonian-regime calibration is a genuine check: if f_trip were wildly wrong, the high-acceleration bins would not line up. So the central claim is plausible.\n\nWhere the paper is softer: the biggest assumption is that f_trip is constant from the Newtonian bin to the MOND regime. The authors assert this because f_trip characterizes a member star, but the sample selection (ruwe, distance limits, angular resolution) is not obviously f_trip-neutral. The paper itself shows f_trip drops from ~0.12 in the full PSS sample to ~0.03 for d<150 pc, so distance-dependent selection is real. Whether it also depends on r_p is the open question. The authors could have tested this by splitting the Newtonian-regime calibration bin by r_p, but they don't. That is a real soft spot, but not fatal: an r_p-dependent f_trip large enough to explain the whole rising profile would likely break the Newtonian-regime agreement, so there is indirect support for constancy.\n\nSecond, the '>5σ combined' headline is misleading because it combines overlapping samples; individual significances are 3–5σ. Third, no code or data are shipped yet, only a promise of Zenodo. Given the debate, independent replication with public data is essential before accepting the quantitative MOND match.\n\nOverall: for someone working on wide binaries or MOND tests, this is worth a careful read. It deserves a serious referee—likely with requests for the f_trip robustness check and for the code. I'd engage with it, but I wouldn't treat the combined significance as the headline.","headline":"Careful, mostly persuasive re-analysis showing the low-acceleration anomaly survives the Pittordis triple model and Cookson quality cuts; main weaknesses are an untested f_trip constancy assumption and no shipped code/data.","tokens_in":46925,"tokens_out":2888,"would_cite":true,"duration_ms":29503,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The low-acceleration gravitational anomaly in wide binaries survives both stringent data-quality control and realistic modeling of hidden companion stars, confirming a MOND-type velocity boost at more than 5σ significance.","keywords":["wide binaries","MOND","modified gravity","Gaia DR3","hierarchical systems","low acceleration","quality control bias","galactic dynamics"],"falsifier":"Measure the hidden-companion fraction directly in the 3–30 kau separation range—for example, by searching each primary for a resolved or astrometric companion via high-resolution imaging and radial-velocity monitoring rather than inferring it statistically—and check whether the fraction equals the Newtonian-calibrated value. If the hidden-companion fraction rises steeply with projected separation, or if the astrometric quality cut admits more short-period inner binaries at large separations, the rising v-tilde profile could be partly spurious; if it stays flat, the anomaly is gravitational.","tokens_in":45971,"feed_emoji":"🔭","tokens_out":6192,"duration_ms":59526,"temperature":0.7,"pith_summary":"This paper re-examines the two main objections raised against earlier wide-binary evidence for modified gravity: that poor data quality—especially a cut on the uncertainty of the normalized velocity v-tilde—and unseen companion stars in hierarchical systems could fake a low-acceleration anomaly. The authors implement the most realistic triple-star model currently proposed and the full quality-control framework advocated by skeptical studies, then calibrate the hidden-companion fraction in the high-acceleration Newtonian regime before testing low accelerations. They find that neither objection removes the anomaly: the median v-tilde still rises with separation, and the deviation from the Newtonian prediction is positive at internal accelerations below about 10^-9 m/s^2 with combined significance greater than 5σ. The shape of the deviation matches recent realistic QUMOND two-body orbit solutions rather than Newtonian gravity, and studies reporting Newtonian consistency are traced to an uncalibrated triple fraction, a biased velocity-error cut, or too few low-acceleration binaries.","feed_headline":"Wide binaries confirm MOND-type gravity anomaly at >5σ","feed_subtitle":"Reanalysis of Gaia binaries shows low-acceleration velocity boosts cannot be blamed on hidden companions or data cuts.","key_machinery":"The central object is the normalized sky-plane velocity v-tilde = v_p / v_c(r_p), the observed 2D relative velocity divided by the Newtonian circular speed at the projected separation; in Newtonian gravity its median should be nearly flat in separation, so a rising v-tilde profile is a gravity signal. Two tools carry the argument: the acceleration-plane test, which projects each binary's logarithmic Newtonian and empirical acceleration onto a diagonal coordinate and measures the orthogonal deviation from Newton, and a forward model of apparent binaries with hidden companion stars, encoded by an effective lower limit on the inner orbit semi-major axis that reproduces the proposed 'realistic'","core_discovery":"The central claim is that wide binary stars with projected separations larger than a few thousand astronomical units show a relative velocity that rises above the Newtonian prediction as the internal acceleration drops below roughly 10^-9 m/s^2, and that this rise survives the two main challenges raised against it. Implementing the most sophisticated triple-star model currently proposed, the authors calibrate the hidden-companion fraction using high-acceleration Newtonian binaries; with that calibration, the low-acceleration bins still show a positive deviation from Newton, with combined significance exceeding 5σ. Applying the full set of quality cuts advocated by skeptical studies—distance","pith_inferences":["If the constancy of the hidden-companion fraction across separations is confirmed, the v-tilde profile versus r_p/r_M could be used to map the MOND interpolation function in the external-field-dominated regime, where current constraints are weakest.","Targeted follow-up of wide binaries in the 3–30 kau range—high-resolution imaging and radial-velocity monitoring to find hidden companions directly rather than statistically—would settle whether the calibrated triple fraction truly holds at large separations.","The bias analysis suggests a practical prescription for the next generation of astrometric surveys: use absolute proper-motion velocity errors and direct uncertainty propagation, and verify Newtonian predictions in the high-acceleration regime before interpreting low-acceleration bins.","A direct 3D orbit analysis of the small d<150 pc sample with radial velocities could independently confirm the statistical result and potentially distinguish AQUAL from QUMOND in the transition regime."],"forward_implications":["The low-acceleration anomaly is not an artifact of hidden triples: a nearly pure nearby subsample with a fitted triple fraction of about 0.03 still shows the velocity boost.","A quality cut based on the uncertainty of v-tilde is biased: it selectively removes lower-mass and higher-velocity systems at large separations, so future studies should use absolute velocity-error cuts or direct uncertainty propagation.","The hidden-companion fraction must be calibrated in the Newtonian regime free of chance alignments; using a value fitted to the low-acceleration region biases the test toward Newtonian gravity.","The observed velocity boost, corresponding to γ ≈ 1.3–1.6, agrees with realistic two-body QUMOND predictions under the Galactic external field, while approximate test-particle MOND models overpredict the effect in the transition regime.","Reports of no anomaly based on samples with fewer than about 100 MOND-regime binaries are statistically unable to distinguish Newton from a γ = 1.4 boost, so their null result is expected even if modified gravity is correct."],"fun_headline_variants":["Wide binaries still defy Newton at low acceleration: >5σ","Hidden companion stars cannot erase MOND signal in binaries","Reanalysis confirms MOND-type gravity boost in wide binaries","Low-accel gravity anomaly in wide binaries confirmed at >5σ"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The fraction of apparent binaries hiding an unseen companion star, measured in the high-acceleration Newtonian regime, is assumed to stay constant as binaries move to larger separations; if it actually grows with separation, some of the rising velocity signal could be contamination rather than gravity.","fun_headline_variants_meta":{"raw":{"variants":["Wide binaries still defy Newton at low acceleration: >5σ","Hidden companion stars cannot erase MOND signal in binaries","Reanalysis confirms MOND-type gravity boost in wide binaries","Low-accel gravity anomaly in wide binaries confirmed at >5σ"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1275,"prompt_tokens":835,"completion_tokens":440,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":369}},"tokens_in":579,"tokens_out":440,"duration_ms":4846,"temperature":1.0,"reasoning_tokens":369,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:02:37.669564+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the hidden-companion fraction directly in the 3–30 kau separation range—for example, by searching each primary for a resolved or astrometric companion via high-resolution imaging and radial-velocity monitoring rather than inferring it statistically—and check whether the fraction equals the Newtonian-calibrated value. If the hidden-companion fraction rises steeply with projected separation, or if the astrometric quality cut admits more short-period inner binaries at large separations, the rising v-tilde profile could be partly spurious; if it stays flat, the anomaly is gravitational.","supporting_citations":[],"review_version":1}