{"id":"6599963d-0c73-453e-8e45-c4b72a1eb9de","arxiv_id":"2412.14542","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Most reported discrepancies between orbital fits for wide companions trace to data selection, conventions, and posterior sampling, not to the F19/F23 method itself, which the paper shows is equivalent to orvara.","lead":"A team reanalyzed radial velocity and astrometry data for nine stars with wide-orbit companions and concluded that most reported disagreements between orbital solutions come from differences in data, conventions, and sampling, not from a fundamental flaw in the fitting method. The work resolves a public dispute about the F19/F23 astrometry pipeline and offers updated orbits including a new solution for eps Ind A b.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §2 equivalence between F19/F23 and orvara is asserted from a mean-proper-motion identity, but the two pipelines differ in data representation and nuisance parameters; without a blind injected-signal comparison, the 'data-related not methodology-related' conclusion is not fully secured.","rationale":"The reader's CONDITIONAL verdict is reasonable. The strongest_claim has two parts: (i) F19/F23 and orvara are physically equivalent, and (ii) discrepancies are primarily data-related. Part (i) is the load-bearing foundation: if different methods give different answers on identical data, then methodology cannot be exonerated. The paper's analytic derivation in §2 shows that the mean proper motion is parameterized consistently, but it does not prove likelihood equivalence given the different data products and nuisance models. The real-data comparisons provide positive evidence but no guarantee. The HD 28185 concern is real, but the paper's own Model 2 demonstrates that the outer-orbit discrepancy is explained by RV baseline even without the inner-companion astrometric signal; hence that concern is not the one on which the central claim rests. A controlled injection test would settle part (i) decisively, and it is cheap to run. Until then, CONDITIONAL is the right verdict; the paper should add this test or clearly acknowledge the residual risk. No ad hominem is intended: the issue is structural, not about the authors' conduct. The out-of-sample eps Ind A b prediction is genuinely strong evidence for the pipeline's practical reliability, but it does not replace a controlled equivalence test.","tokens_in":24587,"tokens_out":11697,"duration_ms":103903,"concrete_test":"Blind injection: simulate a star with a known wide companion (P ≈ 10 yr, e ≈ 0.3, I ≈ 70°) and a 386-day inner companion (K ≈ 164 m/s, e ≈ 0.06), generate synthetic Hipparcos IAD and Gaia DR2/DR3 data with realistic cadence, sky-plane coverage, and noise, and add the same synthetic RVs. Run F19/F23 and orvara with identical priors and likelihood settings, then compare recovered posteriors against truth and against each other. If either pipeline misses the injected values or the two pipelines differ by more than sampling noise, the §2 equivalence claim fails; if they agree, the 'data-related' conclusion is strengthened. This directly tests whether jitter/offset freedoms or HGCA calibration choices change the inferred orbit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2's equivalence argument reduces to Eq. (4): the modeled mean proper motion μ̂_HG equals μ_b + μ_HG, matching Brandt et al.'s definition. The text then concludes that 'the difference between F19 and orvara lies solely in the parameterization and not in the physical content.' This is not sufficient. F19/F23 fit catalog positions, proper motions, and parallax with free astrometric jitter and barycentric offset parameters (Table 1), while orvara with HGCA uses calibrated proper motions at H/G epochs and the mean, with no equivalent freedom. F23 additionally incorporates GOST epoch data and multiple Gaia releases. These are different observables and different nuisance models: the extra offsets and jitters can absorb real astrometric signal, and epoch data contain information beyond the compressed HGCA summary. The paper's real-data agreements (HD 38529, 14 Her, HD 211847, GJ 680, HD 111031) are suggestive but are not controlled tests. The HD 28185 parallax-offset assumption flagged by the reader is less load-bearing: Model 2, without the inner companion's astrometric signal, already reproduces V24's outer orbit when the same RV baseline is used, so the 'data-related' explanation for the F22/V24 discrepancy does not depend on the Δϖ attribution. The load-bearing assumption is the claimed equivalence itself.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reanalyzes a set of wide-orbit systems (HD 28185, GJ 229, HD 62364, HD 38529, 14 Her, eps Ind A, HD 211847, HD 111031, and GJ 680) to address discrepancies between the F19/F22/F23 astrometry-plus-RV pipeline and the orvara/HGCA-based analyses reported by Venner et al. (2024). It argues that the F19 and orvara methods are equivalent (Section 2), that the HD 28185 discrepancy is driven by RV dataset differences and by the astrometric signal of the inner year-long companion (Section 3), and that other discrepancies arise from conventions, posterior sampling, inner companions, and RV baseline coverage (Sections 4 and 5). It also presents an out-of-sample prediction for eps Ind A b's position that is matched by JWST/MIRI imaging (Section 5).","tokens_in":24910,"tokens_out":7010,"duration_ms":57435,"significance":"The paper provides a useful collection of case studies and a clear list of pitfalls for wide-orbit fitting, and it demonstrates that several apparent F22 discrepancies disappear when the same data and conventions are used. The eps Ind A b prediction without imaging data is a genuine, falsifiable out-of-sample test, and the many side-by-side F23/orvara comparisons on real targets are a practical strength. However, the central claim that discrepancies are 'primarily data-related rather than methodology-related' is not fully established: the Section 2 equivalence argument is incomplete, and the paper's own list of causes includes methodological items such as conventions and posterior sampling. A controlled injection test or a more carefully qualified claim would be needed to justify the headline.","major_comments":[{"comment":"The equality hat-mu_HG = mu_b + mu_HG shows that the mean proper motion modeled by F19 matches the HGCA definition, but it does not establish that the two fitting pipelines are inferentially equivalent. F19/F23 fit catalog positions and proper motions with free astrometric jitter, barycentric offsets, and Gaia error-inflation factors (Table 1), while orvara uses calibrated HGCA proper motions without equivalent freedom; F23 additionally fits GOST epoch data and multiple Gaia releases. Free nuisance parameters can absorb real astrometric signal, so the statement that 'the difference between F19 and orvara lies solely in the parameterization and not in the physical content' is not supported by the derivation alone. The real-data agreements in Sections 4-5 are suggestive but are not a controlled test; the paper should either soften this claim or add a blind injected-signal comparison of the two pipelines on identical data.","section":"Section 2, Eq. (4)"},{"comment":"The abstract's claim that 'the discrepancies are primarily data-related rather than methodology-related' is in tension with the manuscript's own list of causes, which includes 'clear definitions of conventions' and 'efficient posterior sampling' (abstract, and the first two reasons listed in Section 6). Conventions and posterior sampling are methodological. The categorization should be clarified, or the claim should be qualified so that the headline matches the evidence presented in the paper.","section":"Abstract / Section 6"},{"comment":"The claim that the inner companion of HD 28185 induces a 'significant parallax offset' and improves the fit by Delta log-likelihood 8.4 (Delta BIC 12) depends on modeling the year-long companion's astrometric signal as a constant parallax-like offset Delta-vari in the combined Hipparcos/Gaia data. Since the 385.9-day orbital period is not exactly commensurate with the annual parallax cycle, this is an approximation, and the paper does not verify its adequacy (for example, by injecting a synthetic signal at the fitted orbit and checking recovery). Given that Delta-vari is already a fitted parameter and that jitter and error-inflation freedom are present, the BIC improvement alone is not a fully robust demonstration; this point should be checked, or the wording of the 'importance of parallax modeling' lesson should be moderated.","section":"Section 3, Table 1 and Figs. 3-4"}],"minor_comments":[{"comment":"The caption states that the vectors AD and AF 'denote the observed proper motions at the Gaia and Hipparcos reference epochs,' but AF is labeled r_H,o, which is a reference position; it should presumably be mu_H,o.","section":"Figure 1 caption"},{"comment":"The text says 'we added 51 HARPS RVs from the ESO archive' without giving the program IDs or a precise data source; please provide this information for reproducibility.","section":"Section 4.4"},{"comment":"The sentence 'The dataset of HARPSpost2 were released by Barbieri (2023)' has a subject-verb agreement error and should read 'The dataset ... was released.'","section":"Section 5, paragraph 2"},{"comment":"The phrase 'The shade regions of Panel (a)' should be 'The shaded regions of Panel (a)'.","section":"Figure 8 caption"},{"comment":"The sentence 'First, the use different conventions' should read 'First, the use of different conventions.'","section":"Section 6, final paragraph"},{"comment":"The Venner et al. (2024) bibliography entry lacks a volume and page range; please complete the reference.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a direct rebuttal to the criticisms of Venner et al. (2024) and contains useful lessons for wide-orbit fitting. The main risk is the overstatement of the equivalence of the two pipelines and of the 'primarily data-related' narrative; if the authors add a controlled test or qualify the claim, the paper would be suitable for publication. There is also a minor internal inconsistency between the abstract's claim and the paper's own list of causes that should be reconciled."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does real work: nine systems reanalyzed with both F23 and orvara, and eps Ind A b with a 29.2-year RV baseline predicts the MIRI position without using imaging. That out-of-sample check is the strongest part. It is genuinely new, and the agreement is not algorithm-dependent. The HD 28185 Model 1, which includes the inner companion's astrometric signal as a parallax-like offset, is also new and makes a plausible case that the inner planet matters for the Hipparcos/Gaia parallax discrepancy.\n\nThe soft spot is Section 2. The mean-proper-motion identity in Eq. (4) shows a structural similarity, but it does not establish that the pipelines are equivalent. F19/F23 fit catalog positions and proper motions with free barycentric offsets and astrometric jitter; orvara uses calibrated HGCA proper motions and the mean. Those nuisance parameters can absorb real signal, and the GOST epoch data F23 uses contain information beyond the compressed HGCA summary. The paper's own list of discrepancy causes includes conventions, posterior sampling, and companion modeling, which are methodological items. So the abstract's 'primarily data-related' is overreach; 'not solely due to a fundamental methodological flaw' would be fair. The real-data agreements are suggestive but not controlled tests; an injected-signal or blinded comparison between the pipelines would settle it. On the HD 28185 parallax-offset point, the reader's worry is minor: Model 2 already matches V24 when the same RV baseline is used, so that conclusion does not depend on the Delta-vari attribution.\n\nNo code or data artifacts are shipped, which is a limitation for a methods-comparison paper, though the data are public.\n\nWho is this for: anyone working on long-period companion orbits with Hipparcos/Gaia astrometry, or on the V24/F22 dispute. It deserves serious peer review, but the referee should push on the equivalence claim and ask for the abstract to match the evidence.","headline":"A solid reanalysis with a genuine out-of-sample check, but the 'methodology vs data' conclusion is broader than the evidence; referee it.","tokens_in":25466,"tokens_out":2603,"would_cite":true,"duration_ms":23720,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the disputed orbital solutions for wide companions are explained by differences in data, conventions, and sampling rather than by a flawed astrometric method, and that the F19/F23 and orvara pipelines are physically…","keywords":["radial velocity","astrometry","wide-orbit companions","orbital fitting","Hipparcos-Gaia astrometry","parallax modeling","posterior sampling","eps Ind A b"],"falsifier":"Fit the HD 28185 system with the F23 model but replace the constant parallax offset $\\Delta\\varpi$ by an explicit Keplerian orbit for the inner planet evaluated at the individual Hipparcos epoch abscissae and Gaia scan times, keeping the same RV data. If the log-likelihood gain of 8.4 relative to the model without the inner planet vanishes, or if the fitted $\\Delta\\varpi$ goes to zero while the orbit parameters stay fixed, the claim that the inner planet's annual signal acts as a measurable parallax-like bias would be falsified. A second test: run F23 and orvara on the same synthetic wide-companion dataset with known injected parameters and identical conventions; disagreement beyond the quoted uncertainties would contradict the 'solely parameterization' equivalence.","tokens_in":24357,"feed_emoji":"🪐","tokens_out":9733,"duration_ms":74699,"temperature":0.7,"pith_summary":"Disagreements between published orbital solutions for wide-orbit giant planets have been taken as evidence that one fitting pipeline is unreliable. This paper argues the opposite: the disputed solutions differ mainly because the studies used different radial-velocity datasets, different orbital conventions, different amounts of posterior sampling, and different assumptions about additional companions, not because the underlying astrometric methods are physically different. It demonstrates that the F19/F23 approach of directly modeling Hipparcos and Gaia catalog data and the orvara approach based on calibrated proper motions are equivalent in physical content, differing only in parameterization. The reanalysis of nine systems, with HD 28185 as the worked example, is meant to show that once data and conventions are aligned the methods converge. If the paper is right, previously reported conflicts do not undermine the astrometry-plus-RV route to detecting and characterizing long-period planets.","feed_headline":"Wide-companion orbit clashes traced to data, not methods","feed_subtitle":"Reanalysis of nine systems shows rival orbit-fitting codes agree once data, conventions, and sampling are aligned.","key_machinery":"The load-bearing object is the F19/F23 astrometric model: it fits barycentric motion (reference position and proper motion) plus reflex-motion terms, astrometric jitter, and catalog offsets directly to Hipparcos and Gaia catalog data, and in the F23 version fits Hipparcos plus Gaia DR2 and DR3 simultaneously with parallax contributions to GOST-generated abscissae. Its equivalence to orvara rests on the identity $\\hat{\\boldsymbol{\\mu}}_{HG} \\equiv (\\boldsymbol{r}_G - \\boldsymbol{r}_H)/\\Delta T = \\boldsymbol{\\mu}_b + \\boldsymbol{\\mu}_{HG}$, which shows that a position-based fit and a proper-motion-based fit contain the same physical information. The second key mechanism is the fitted parallax offset $\\Delta\\varpi$ in the HD 28185 model, which absorbs the annual astrometric signal of the inner planet and converts it into a concrete, testable contribution to the Hipparcos versus Gaia parallax difference. Together these mechanisms let the paper separate genuine physical signals from apparent disagreements caused by conventions, data selection, sampling, and companion modeling.","core_discovery":"The paper's central claim is that the discrepancies between its earlier wide-companion solutions and those of a competing analysis are primarily data-related rather than methodology-related. Concretely, the authors assert that the difference between F19 and orvara lies solely in the parameterization and not in the physical content: both solve for the same reflex motion of the host star, one by fitting Hipparcos/Gaia catalog positions and proper motions together with astrometric jitter and offsets, the other by fitting the calibrated Hipparcos-Gaia proper-motion accelerations, with the identity $\\hat{\\boldsymbol{\\mu}}_{HG} = \\boldsymbol{\\mu}_b + \\boldsymbol{\\mu}_{HG}$ connecting the two. For HD 28185 they show that including the astrometric signal of the year-long inner planet, modeled through a fitted parallax-like offset, improves the fit by a log-likelihood of 8.4 (ΔBIC = 12) and resolves the roughly 2–3σ Hipparcos/Gaia parallax discrepancy, and that without this signal their solution matches the competing one. Across the other targets the disagreements are attributed to four named causes: a 180-degree convention difference in the longitude of ascending node, incomplete sampling of the two inclination modes, unmodeled inner companions, and radial-velocity baselines too short to cover the orbital turn-over of decades-long companions. The eps Ind A b case is used to show that with a roughly 29-year RV baseline, F23 and orvara produce consistent orbits and a position matching the JWST/MIRI image.","pith_inferences":["If the HD 28185 mechanism is general, stars with inner planets on roughly one-year orbits should show systematic Hipparcos-versus-Gaia parallax offsets of order 0.1–0.4 mas; Gaia DR4 epoch astrometry could test this by fitting the annual signal explicitly instead of as a constant offset.","The equivalence between position-based and proper-motion-based astrometric fits suggests a practical diagnostic: whenever two pipelines disagree on the same epochs and weights, the disagreement is a flag for a data or sampling problem, not a reason to discard one method.","The eps Ind A b result implies that some previously published shorter-period solutions for other long-period companions may be artifacts of short RV baselines, and that reanalysis with a longer baseline or imaging constraints could shift their periods and masses upward.","Some RV-plus-astrometry-only classifications of wide companions as brown dwarfs may need revisiting with relative astrometry, since without it the mass-period degeneracy allows low-mass stellar companions to be underestimated."],"forward_implications":["The F19/F23 pipeline can be used with confidence on the same data as orvara; reported differences for HD 28185, HD 38529, 14 Her, GJ 229, HD 62364, HD 211847, GJ 680, and HD 111031 are not signs of a methodological flaw.","For long-period companions, a radial-velocity baseline that does not cover the orbital turn-over leaves a mass-period degeneracy; adding relative astrometry from imaging or extending the baseline is required to break it.","Multi-modal inclination posteriors are expected for astrometric orbits, so convergence diagnostics like $\\hat{R}<1.1$ alone do not guarantee that all modes were found; multiple samplers or chains with different starting points are needed.","Year-long inner companions can bias parallax measurements at the level of the Hipparcos/Gaia discrepancies, so multi-companion fits should model their astrometric signal rather than averaging it away.","The longitude of ascending node reported by different studies can differ by 180 degrees purely from convention, so apparent node discrepancies should be checked against the adopted convention before being interpreted physically."],"supporting_citations":[{"why":"The competing study whose disputed orbital solutions and its claim that the F19 method is unreliable define the problem this paper addresses.","marker":"Venner et al. (2024)"},{"why":"Introduces the F19 direct modeling method for Hipparcos/Gaia catalog data whose equivalence to orvara is the paper's central methodological claim.","marker":"Feng et al. (2019b)"},{"why":"Describes orvara and the HGCA-based proper-motion model that F19 is compared against; supplies the mean-proper-motion formulation and priors.","marker":"Brandt et al. (2021a)"},{"why":"The F23 update that fits Hipparcos, Gaia DR2, and DR3 with GOST-generated abscissae; used for all the reanalyses in this paper.","marker":"Feng et al. (2023)"},{"why":"The earlier wide-companion survey whose solutions were called into question and are reexamined here.","marker":"Feng et al. (2022)"},{"why":"Supplies RV plus relative astrometry solutions for HD 211847, GJ 680, HD 111031, and eps Ind A b, used to show that missing imaging data explains F22's discrepancies.","marker":"Philipot et al. (2023a,b)"},{"why":"JWST/MIRI direct imaging observation of eps Ind A b; its measured position provides the external check on the predicted orbits.","marker":"Matthews et al. (2024)"},{"why":"Earlier GJ 229 B solution whose mass is reproduced once the same data are used, demonstrating the role of inner companions.","marker":"Brandt et al. (2021c)"},{"why":"Supplies the 14 Her inclinations that missed one of the two mirrored modes, illustrating insufficient posterior sampling.","marker":"Bardalez Gagliuffi et al. (2021)"},{"why":"The HD 38529 solution used to trace the 180-degree convention difference in the longitude of ascending node and the DR2-versus-EDR3 data effect.","marker":"Xuan et al. (2020)"}],"fun_headline_variants":["Rival orbit fits agree once data is handled properly","Orbit clashes traced to data, not methods","Data issues, not methodology, cause orbit discrepancies","Wide-companion orbits: fix data, get consistent fits","Nine systems show orbit solutions differ due to data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The worked example for HD 28185 assumes that the astrometric wobble from the inner, roughly one-year planet can be adequately captured by a single constant offset in parallax across the combined Hipparcos and Gaia data; if that shortcut is wrong, the claimed improvement in fit and the explanation of the parallax discrepancy would not hold.","fun_headline_variants_meta":{"raw":{"variants":["Rival orbit fits agree once data is handled properly","Orbit clashes traced to data, not methods","Data issues, not methodology, cause orbit discrepancies","Wide-companion orbits: fix data, get consistent fits","Nine systems show orbit solutions differ due to data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1572,"prompt_tokens":1067,"completion_tokens":505,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":429}},"tokens_in":683,"tokens_out":505,"duration_ms":4551,"temperature":1.0,"reasoning_tokens":429,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:09:20.093863+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit the HD 28185 system with the F23 model but replace the constant parallax offset $\\Delta\\varpi$ by an explicit Keplerian orbit for the inner planet evaluated at the individual Hipparcos epoch abscissae and Gaia scan times, keeping the same RV data. If the log-likelihood gain of 8.4 relative to the model without the inner planet vanishes, or if the fitted $\\Delta\\varpi$ goes to zero while the orbit parameters stay fixed, the claim that the inner planet's annual signal acts as a measurable parallax-like bias would be falsified. A second test: run F23 and orvara on the same synthetic wide-companion dataset with known injected parameters and identical conventions; disagreement beyond the quoted uncertainties would contradict the 'solely parameterization' equivalence.","supporting_citations":[],"review_version":1}