{"id":"8dcaa486-5be2-4bda-817f-e7eb952bb304","arxiv_id":"1908.07547","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":17,"one_line_summary":"Basilisk fits raw satellite positions and velocities in a hierarchical Bayesian model and recovers unbiased central galaxy halo mass distributions and orbital anisotropy in mock surveys.","lead":"This paper introduces Basilisk, a Bayesian method that uses the motions of small satellite galaxies to measure how galaxies connect to dark matter halos, without binning galaxies into stacks. It tests the method on simulated surveys and argues it can reach halo mass precision comparable to gravitational lensing while also measuring the shape of satellite orbits.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unbiasedness claim rests on Tier-3 subhalo mocks, but those mocks do not fully violate the Jeans/Gaussian assumptions: satellites are drawn from the most massive (likely most relaxed) subhaloes at z=0, and the analysis still imposes spherical NFW and Gaussian LOSVD.","rationale":"Reading the paper in good faith, the method is clearly specified and the three-tier validation is a real strength; the Tier-3 mock, in particular, provides independent evidence beyond self-consistency. However, the central claim is broader than the validation. The paper's own Section 7 lists Assumptions I-IV and then cites work showing real satellite subhaloes are not in steady state or spherical equilibrium; the response is the Tier-3 mock, but that mock's subhalo selection (highest M_peak) and the imposed spherical NFW/Gaussian LOSVD in the analysis leave the key premise only partially tested. This does not make the paper wrong, but it makes the abstract's unconditional wording ('unbiased', 'accurately recovers the full PDF') stronger than the evidence. The reader's weakest_assumption points to Assumption II/III; I agree and would sharpen it by noting that Tier-3 is the only non-self-consistent validation and that its construction may select the most relaxed subhaloes. The proposed hydro or random-subhalo test would settle whether the Jeans/Gaussian approximation is the bottleneck. The satellite-CLF biases seen in Table 2 reinforce the conditional nature of the claim, though they do not by themselves invalidate the central P(M|Lc) recovery. Overall the CONDITIONAL verdict is appropriate; I would not move it.","tokens_in":48201,"tokens_out":12796,"duration_ms":572420,"concrete_test":"Construct a Tier-4 mock from a cosmological hydrodynamical simulation (e.g., IllustrisTNG or EAGLE) at z approximately 0.1: assign central luminosities with the fiducial CLF, but place satellites at the positions and velocities of simulated satellite galaxies, deliberately including recently accreted and backsplash populations. Run the Basilisk pipeline exactly as in Section 5.3, including the fibre-collision treatment, and compare the posterior predictive P(<M|Lc) at Lc = 10^9.8, 10^10.3, and 10^10.8 h^-2 L_sun and the central-CLF parameters (M1, L0, gamma1, gamma2, sigma12, sigma14) with the input values. If any parameter or any of the three cumulative distributions falls outside the 95% credible interval, the unbiasedness claim is conditional on Jeans equilibrium.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (abstract and Section 7) is that Basilisk yields unbiased constraints on the galaxy-halo connection and the full P(M|Lc) using the likelihood of Eqs. (5)-(11) under Assumptions I-IV. The load-bearing premise is Assumption II/III: satellites are a virialized, steady-state tracer of a spherical NFW potential with Gaussian LOSVD (Eq. 45). The only validation that does not generate satellite phase space with the same Jeans/Gaussian model is the Tier-3 mock (Section 5.3), so the entire unbiasedness claim reduces to how well Tier-3 approximates reality. Two features make that approximation favorable rather than neutral. First, satellites are assigned to the N_sat subhaloes with the highest M_peak; these are likely the oldest, most tidally relaxed subhaloes, not a representative sample including recent infall and backsplash. When N_sat exceeds the resolved subhalo count, phase-space coordinates are taken from subhaloes of other haloes of similar mass, breaking the physical host-satellite correlation. Second, the analysis model still enforces spherical NFW with zero-scatter concentration-mass relation and Gaussian P(DeltaV|Rp,M,z), which are exactly the assumptions Section 7 admits are violated in reality (citing Wang et al. 2017; Adhikari et al. 2019). The Tier-3 success is therefore evidence of robustness for one subhalo population, not for the full range of disequilibrium and asphericity expected in real satellite galaxies. The paper's own Section 5.2 also notes that some satellite-CLF parameters (e.g. alpha12 in Table 2) are recovered with input values outside the 95% posterior in Tier-3, so 'unbiased constraints on the galaxy-halo connection' is broader than the validation supports.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces Basilisk, a Bayesian hierarchical method that constrains the central and satellite conditional luminosity functions using the raw projected phase-space coordinates of primary-secondary galaxy pairs, without binning or velocity-dispersion summary statistics. Satellites are modelled as a relaxed tracer population in spherical NFW haloes with a Gaussian line-of-sight velocity distribution (Eqs. 44-50), and the likelihood marginalizes over halo mass as a latent variable (Eqs. 5-11). The method is validated on three mock tiers: idealized model mocks (Tier-1), N-body-based mocks with analytic satellite phase-space, realistic interlopers and fibre collisions (Tier-2), and N-body mocks in which satellites are placed on resolved subhaloes (Tier-3). The paper reports unbiased recovery of the central CLF parameters and the full P(M|Lc), plus an anisotropy constraint, and claims precision competitive with galaxy-galaxy lensing.","tokens_in":48749,"tokens_out":10538,"duration_ms":106955,"significance":"The method is conceptually a step forward: it avoids luminosity binning, exploits the Delta-V-Rp correlation, uses flux-limited samples, and includes the number of secondaries per primary as a constraint, which Appendix B shows is crucial. The three-tier validation, especially the Tier-3 subhalo mocks and the tests against central velocity bias and radial-profile errors, is a genuine strength. However, the load-bearing validation of the Jeans/Gaussian assumptions is weaker than the abstract's general unbiasedness claim, and the presented 'predicted halo masses' are not full posterior estimates. If the missing tests are supplied or the claims are appropriately qualified, this would be a valuable methods paper.","major_comments":[{"comment":"The Tier-3 mock is the only validation in which satellite phase space is not drawn from the same Jeans/Gaussian model used in the likelihood, and it is therefore load-bearing for the paper's unbiasedness claim. As described in §5.3, satellites are assigned to the N_sat subhaloes with the highest M_peak, which selects the oldest, most tidally relaxed subhaloes; when N_sat exceeds the resolved subhalo count, the missing phase-space coordinates are taken from subhaloes of other, similar-mass haloes, which breaks the physical host-satellite correlation. Moreover, the analysis still imposes a spherical NFW profile, a zero-scatter concentration-mass relation, and a Gaussian LOSVD (Assumptions I-III in §7). The paper's own §7 acknowledges that real satellite populations violate these assumptions via recent accretion, asphericity, and backsplash (citing Wang et al. 2017; Adhikari et al. 2019). A successful recovery on the most relaxed subhalo subset does not establish unbiasedness for the full range of disequilibrium expected in real data; the authors should either relax the subhalo selection (e.g., random subhaloes including recent infall), add a mock that violates the Jeans/Gaussian assumptions more strongly, or soften the general claim in the abstract.","section":"§5.3 and §7"},{"comment":"The 'predicted halo mass' M_pred used in Figs. 4(f), 7(f), and 9(f) is computed from P(M|Lc,zc,Ns) alone, without conditioning on the satellite phase-space data (Delta-V, Rp). It is therefore not the full posterior mean of the latent halo mass under the Basilisk model, and the small offsets quoted (⟨log(Mpred/Mtrue)⟩ = 0.10-0.15 with scatter 0.31-0.37) cannot be used to validate the kinematic likelihood or to support the statement in §1 that Basilisk yields individual halo-mass estimates as a by-product. If individual halo masses are part of the claimed output, the paper should sample or approximate the latent masses from the full posterior and report the bias of those estimates; if not, Eq. (54) and the related text should be re-labelled as the population-level conditional expectation.","section":"Eq. (54) and §5.1-§5.3"},{"comment":"The reported posterior percentiles are conditional on the radial profile parameters (R,gamma) being fixed at their best-fit values; the MCMC is run separately for the best-fit (gamma,R) pair rather than marginalizing over these parameters. Table 2 therefore presents conditional intervals that understate the parameter uncertainty. The sensitivity analysis in §6.2 shows that the best-fit CLF parameters shift by less than the conditional 95% intervals when R and gamma are varied, which is reassuring, but it does not provide the fully marginalized posteriors. The authors should either marginalize over (R,gamma) in the MCMC or explicitly state that all quoted intervals are conditional and provide a marginalization check for at least one tier.","section":"§4.2.3, §5, and Table 2"},{"comment":"The abstract claims 'unbiased constraints on the galaxy-halo connection' without restricting the claim to the central component, but Table 2 shows that the satellite CLF slope alpha_12 is not recovered in the Tier-2 or Tier-3 mocks: the input value -1.20 is outside the 95% intervals (-1.21,-0.32) and (-1.01,-0.37) respectively. The text in §5.2 openly acknowledges that parameters characterizing Phi_s(L|M) are often inconsistent with the input. Because the satellite component is part of the galaxy-halo connection, the paper should either qualify the unbiasedness claim to the central part Phi_c(L|M) or explain why the alpha_12 bias is not a violation of the central claim.","section":"Table 2 and §5.2"}],"minor_comments":[{"comment":"The phrase 'the only available method that simultaneously solves for halo mass and orbital anisotropy' is too strong; Wojtak & Mamon (2013) already constrained anisotropy from satellite kinematics, as the paper itself notes later. Suggest rewording.","section":"§1"},{"comment":"For the Tier-3 mock the best-fit gamma is 0, at the edge of the prior range; the authors should note that the recovery of a cored profile is partly prior-limited.","section":"Figure 3"},{"comment":"No convergence diagnostics (e.g., Gelman-Rubin statistics or autocorrelation lengths) are reported for the MCMC chains; a brief statement would help the reader assess the quoted confidence intervals.","section":"§3.5"},{"comment":"The rows for log[ra/rs] report percentiles from a separate OM-model MCMC, but the table caption does not state whether these are conditional on the fixed (R,gamma) values; clarify.","section":"Table 2"},{"comment":"The text says the Tier-2 mock uses the measured concentration of each halo while the analysis assumes a zero-scatter concentration-mass relation; this is a strength of the test and could be stated more prominently in the summary of the validation strategy.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"I see no reason to doubt the authors' good faith. The main editorial questions are whether the abstract's generality ('unbiased constraints on the galaxy-halo connection') is supported given the alpha_12 bias, and whether the validation can be made less favorable to the method by testing a less relaxed subhalo sample. I recommend major revision to address these coverage issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Basilisk is a real step forward for satellite kinematics. The likelihood in Eqs. (5)-(11) treats each central-satellite pair as raw data, marginalizes over halo mass, and includes interloper and fibre-collision terms. No stacking, no velocity-dispersion summary statistic, and explicitly flux-limited. That combination is new relative to the More/Lange pipelines, and it matters because it opens a clean route to constrain P(M|Lc) and orbital anisotropy simultaneously. The three-tier validation is the right structure. Tier 1 is acknowledged as a self-consistency check. Tier 2 adds realistic halo masses, interlopers, and incompleteness. Tier 3 puts satellites on actual subhalo phase-space coordinates. In all three tiers the central CLF parameters, the quantities the method is mainly for, are recovered at the few-percent level, and the full P(M|Lc) is reproduced. That is a solid, honest demonstration.\n\nNow the soft spots. The \"unbiased constraints on the galaxy-halo connection\" claim is broader than what the validation actually shows. The satellite CLF parameters, especially alpha12, are biased in the realistic mocks, sometimes outside the 95% posterior. The paper notes this and attributes it to interlopers and impurity, but it still means the headline should be \"unbiased central galaxy-halo connection\" for now. Second, the Tier-3 mock is less adversarial than it first appears. Satellites occupy the most massive, likely oldest and most relaxed, subhaloes, and the analysis model still enforces spherical NFW with zero-scatter concentration and a Gaussian LOSVD. The paper cites the very studies showing those assumptions are violated in reality. So Tier-3 is evidence of robustness for one subhalo population, not for the full range of disequilibrium, asphericity, and concentration scatter expected in real data. I do not think this kills the method. The central CLF appears insensitive to these details by construction, but the framing should be more careful. The abstract's \"precision rivals galaxy-galaxy lensing\" is an extrapolation; no lensing comparison is actually made here. And no code or data are released. For a method paper of this complexity, independent verification is difficult without them.\n\nWho is this for? Anyone building likelihood-based probes of the galaxy-halo connection, and observers planning satellite kinematics with BOSS or DESI. It deserves a serious referee. I would accept it for peer review, with revisions that tighten the scope of the unbiasedness claim and ask for code and mock release. It is careful, technically serious work.","headline":"Basilisk is a genuinely new likelihood for satellite kinematics with a solid three-tier validation, but the unbiasedness claim runs ahead of what the tests actually show and the Tier-3 mock is less adversarial than it looks; still, this deserves a serious referee.","tokens_in":49366,"tokens_out":2188,"would_cite":true,"duration_ms":147969,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Bayesian hierarchical likelihood that keeps raw satellite kinematics can recover unbiased halo-mass distributions and orbital anisotropy from galaxy survey data.","keywords":["satellite kinematics","galaxy-halo connection","conditional luminosity function","Bayesian hierarchical inference","orbital anisotropy","halo mass","Jeans equation","mock validation"],"falsifier":"Run Basilisk on a hydrodynamical cosmological simulation where satellite galaxies may be out of equilibrium, treating the simulation as a mock redshift survey, and compare the recovered $P(M|L_c)$ with the true distribution of its central galaxies; a significant offset would falsify the steady-state Jeans assumption.","tokens_in":47968,"feed_emoji":"🔭","tokens_out":5365,"duration_ms":51703,"temperature":0.7,"pith_summary":"Basilisk is a Bayesian hierarchical method for using the positions and line-of-sight velocities of satellite galaxies to infer how galaxies populate dark matter haloes. It claims to extract this information from raw central-satellite pairs, without stacking galaxies in luminosity bins and without reducing the data to a velocity dispersion, and it can use flux-limited rather than volume-limited samples. The paper argues this makes satellite kinematics competitive with galaxy-galaxy lensing while also recovering the full probability distribution of halo mass at fixed central luminosity and the orbital anisotropy of the satellites. Validation on three tiers of mock data is presented as evidence that the method returns unbiased constraints on the central galaxy-halo connection even when the analysis model is simplified.","feed_headline":"Satellite motions can pin down halo masses without stacking","feed_subtitle":"A Bayesian likelihood on raw pair data recovers the full distribution of halo masses and the orbital anisotropy from flux-limited surveys.","key_machinery":"The engine is the hierarchical satellite-kinematics likelihood $L_{\\rm SK}$ of Eqs. (5)-(11): for each primary, the number of secondaries and their $(\\Delta V, R_p)$ coordinates constrain a latent halo mass through Bayes' theorem, with the halo-occupation statistics supplied by a conditional luminosity function and the kinematics supplied by a spherical Jeans-equation model of satellites with a Gaussian line-of-sight velocity distribution. The number of secondaries per primary enters as data, which is what lets the method constrain the scatter in the galaxy-halo relation and avoid the satellite-weighting bias. Interlopers are modeled with a parametric effective bias, and fibre collisions are corrected by down-weighting the expected secondary count. Together these pieces replace stacking and velocity-dispersion summary statistics with a full data likelihood.","core_discovery":"The central claim is that the full likelihood for the projected phase-space data, with each central's halo mass treated as a latent variable and marginalized over, yields unbiased constraints on the galaxy-halo connection. Specifically, Basilisk accurately recovers the conditional luminosity function parameters and the full PDF $P(M|L_c)$ for central luminosity versus halo mass, simultaneously constrains the satellite orbital anisotropy, and does so without binning or summary statistics. The validation shows this holds for idealized mocks, for mocks built from N-body haloes with realistic interlopers and fibre collisions, and for mocks in which satellites follow the phase-space distribution of subhaloes. The paper concludes that stacking-based analyses that ignore mass-mixing or assume isotropic orbits can be superseded by this approach.","pith_inferences":["If Basilisk's unbiased recovery survives on real surveys, satellite kinematics could serve as an independent check of galaxy-galaxy lensing without needing shape measurements, potentially probing gravity by comparing dynamical and lensing masses.","The framework extends naturally to stellar-mass-based occupation statistics, and the hierarchical setup could absorb uncertainty in stellar-mass estimates.","The paper's demonstrated insensitivity to the satellite radial profile suggests that adding satellite luminosities to the likelihood may be a low-risk way to tighten the satellite component of the conditional luminosity function.","One open extension is to let halo concentration enter the occupation model, since concentration correlates with satellite number and could feed back on the inferred halo masses."],"forward_implications":["Satellite kinematics can be applied to flux-limited surveys, enlarging the usable sample and dynamic range relative to volume-limited stacking analyses.","The method simultaneously constrains halo mass and orbital anisotropy, removing the need to assume isotropic satellite orbits.","Because scatter in the galaxy-halo connection is modeled, the inferred $P(M|L_c)$ is not biased by mass-mixing.","Individual halo masses for primaries are produced as by-products, allowing tests of how halo mass depends on secondary properties.","Mild central velocity bias and modest errors in the inferred satellite radial profile do not bias the galaxy-halo connection inference."],"supporting_citations":[{"why":"Predecessor satellite-kinematics analysis that established the selection criteria and fibre-collision corrections Basilisk builds on.","marker":"Lange et al. 2019a"},{"why":"Forward-modelling study whose volume-limited estimators and covariance matrix calibrate the comparison and inform the corrections.","marker":"Lange et al. 2019b"},{"why":"Introduced the adaptive cylindrical selection criteria and Jeans-based satellite modeling that the method uses.","marker":"van den Bosch et al. 2004"},{"why":"Diagnosed the satellite-weighting/mass-mixing bias that the hierarchical treatment of $N_s$ is designed to avoid.","marker":"More et al. 2009b"},{"why":"Earlier stacking analysis whose inferred galaxy-halo connection supplies the fiducial conditional luminosity function parameters for the mocks.","marker":"More et al. 2011"},{"why":"Supplies the zero-scatter concentration-mass relation assumed for the NFW host haloes.","marker":"Macciò et al. 2008"},{"why":"Introduced the conditional luminosity function that parameterizes halo occupation in the likelihood.","marker":"Yang et al. 2003"},{"why":"Provides the halo mass function used to set the prior on latent halo masses.","marker":"Tinker et al. 2008"}],"fun_headline_variants":["Satellite motions unlock full halo-mass PDF without stacking","Bayesian method reveals halo mass and orbit from raw data","No stacking, no binning: unbiased halo-mass PDF from satellites","Simultaneous halo mass and anisotropy from satellite kinematics","Hierarchical inference of galaxy-halo connection from satellite data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that satellite galaxies behave as a virialized, steady-state tracer population of a spherical dark-matter halo, so their line-of-sight velocities follow the spherical Jeans equation with a Gaussian velocity distribution; if real satellites are recently accreted, aspherical, or out of equilibrium, the inferred halo masses could be biased.","fun_headline_variants_meta":{"raw":{"variants":["Satellite motions unlock full halo-mass PDF without stacking","Bayesian method reveals halo mass and orbit from raw data","No stacking, no binning: unbiased halo-mass PDF from satellites","Simultaneous halo mass and anisotropy from satellite kinematics","Hierarchical inference of galaxy-halo connection from satellite data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000445,"raw_usage":{"total_tokens":2264,"prompt_tokens":975,"completion_tokens":1289,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":1208}},"tokens_in":591,"tokens_out":1289,"duration_ms":13378,"temperature":1.0,"reasoning_tokens":1208,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:04:56.466178+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Basilisk on a hydrodynamical cosmological simulation where satellite galaxies may be out of equilibrium, treating the simulation as a mock redshift survey, and compare the recovered $P(M|L_c)$ with the true distribution of its central galaxies; a significant offset would falsify the steady-state Jeans assumption.","supporting_citations":[],"review_version":1}