{"id":"e6294ad4-3e0e-4045-bc85-a69bfd8f0634","arxiv_id":"2607.10905","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"HCβ replaces conventional (1-ht) leverage powers in HC estimators with a data-driven Beta-CDF adjustment, improving finite-sample size and coverage while remaining asymptotically consistent.","lead":"The paper introduces HCβ, a heteroskedasticity-consistent covariance estimator for OLS that replaces fixed leverage powers with a Beta-CDF adjustment whose parameters are fit to the sample leverages. Applied researchers get more stable standard errors and confidence intervals when high-leverage points are present, without the explosive corrections of HC4-type methods.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the free constants already flagged by the reader.","rationale":"The reader correctly isolates the only soft spot: the four simulation-tuned constants (c1=7, c2=0.75, truncation bounds 0.01/0.99, shrinkage 50). That concern is real but already priced into the CONDITIONAL verdict; it does not rise to a load-bearing flaw that would require a harsher adjustment. The asymptotic justification is transparent, the Monte Carlo evidence is standard and favorable under the designs studied, the empirical illustrations are consistent with the claimed avoidance of overshooting, and the open-source package makes independent verification immediate. No deeper technical inconsistency (e.g., failure of the Beta-CDF construction under high leverage, or breakdown of the claimed asymptotic equivalence) is present. Therefore the reader’s CONDITIONAL verdict with high confidence stands unchanged.","tokens_in":23163,"tokens_out":561,"duration_ms":5128,"concrete_test":"Using the publicly available hcinfer package, recompute the null rejection rates of Table 2 (second scenario, λ≈50, n=50 and n=100) under a 3×3 grid of (c1,c2) around the recommended defaults (e.g., c1∈{5,7,9}, c2∈{0.5,0.75,1.0}); if size remains within 1–2 percentage points of the reported figures for all nine combinations, the defaults are robust and the claim is unaffected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (accurate finite-sample size/coverage under leverage while retaining asymptotic HC0-like behavior) rests on the construction in §3: gt = [n/(n-p)] * [1/F_Beta(wt; ã, b̃)]^{c1/n^{c2}} with moment-estimated Beta parameters, truncation of wt to [0.01,0.99], and shrinkage ζ = n/(n+50). The asymptotic argument (gt → 1 as n→∞ with p fixed) is elementary and holds regardless of the specific finite values of c1,c2. The Monte Carlo design (Tables 1–4) and four empirical illustrations document the claimed finite-sample gains relative to HC0/HC3/HC4/HC4m under the designs examined. The only free numerical choices are precisely the four constants already identified by the reader; they are not hidden assumptions that invalidate the construction, only provisional defaults. No internal inconsistency, missing regularity condition, or unacknowledged bias source appears that would overturn the claim under the paper’s own stated conditions.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes HCβ, a new heteroskedasticity-consistent covariance-matrix estimator for OLS. It replaces the usual powers of (1-h_t) with an adjustment factor built from a Beta CDF whose shape parameters are estimated by method of moments from the observed leverages (after truncation of w_t=1-h_t to [0.01,0.99] and shrinkage toward the uniform). The resulting g_t multiplies the squared residual by n/(n-p) times [1/F_Beta(w_t; ã, b̃)] raised to the decaying power c1/n^{c2} (recommended defaults c1=7, c2=0.75). Asymptotically g_t\to1 so the estimator recovers the HC0 sandwich; Monte Carlo experiments (two designs, three heteroskedasticity strengths, n=50/100/200, 10 000 replications) and four empirical illustrations claim improved size and coverage relative to HC0/HC3/HC4/HC4m, especially under strong leverage, while an accompanying R package implements the method.","tokens_in":23482,"tokens_out":1163,"duration_ms":9806,"significance":"If the finite-sample gains hold under a broader range of designs, HCβ would be a useful practical addition to the HC toolkit: it supplies a single, data-adaptive correction that moderates the well-documented “overshooting” of HC4/HC4m without requiring the user to choose among several fixed-exponent rules. The asymptotic argument is elementary and clean, the Monte Carlo design is standard and transparent, the empirical examples give concrete g_t-versus-h_t plots that make the overshooting phenomenon visible, and the open-source package lowers the barrier to adoption. These are genuine strengths. The contribution remains incremental rather than foundational, because the functional form and the four free constants are chosen by Monte Carlo search rather than derived from a formal optimality criterion.","major_comments":[{"comment":"Section 3 (paragraphs following the definition of g_t) and the Monte Carlo section: the constants c1=7, c2=0.75, the truncation bounds 0.01/0.99, and the shrinkage constant 50 are selected solely by “extensive Monte Carlo simulations” on designs that closely resemble those later used for performance evaluation. This introduces a mild but load-bearing circularity. The paper should either (i) report a systematic sensitivity analysis (tables or figures showing size/coverage for a grid of nearby constants) or (ii) re-estimate the constants on a hold-out design family and then re-evaluate Tables 1–4, so that the claimed superiority is not partly an artifact of in-sample tuning.","section":null},{"comment":"Section 4, Tables 1–2: under every design the HCβ test is mildly to moderately conservative (null rejection rates 3.6–4.8 % at n=100–200). While the authors note this as “protection against false positives,” the power comparison in Table 3 is performed with size-adjusted critical values. Without size-adjusted power (or an explicit discussion of the size–power trade-off under the unadjusted asymptotic critical values that practitioners actually use), it is difficult to judge whether the improved size control is purchased at an unacceptable power cost. A short size-adjusted power panel or a brief remark on the practical implications of the conservativeness would strengthen the central claim.","section":null}],"minor_comments":[{"comment":"Section 3: the claim that “the asymptotic behavior … is controlled entirely by the decay rate of the exponent c1/n^{c2}, not by the estimated parameters” is correct, yet the truncation prevents the moment estimators from converging to their theoretical limits. A one-sentence clarification that the truncation bias vanishes in the product that defines g_t would remove any residual ambiguity.","section":null},{"comment":"Figures 1, 3–5: the panels are informative, but the vertical scales differ dramatically across estimators; a common log-scale or an inset for the extreme points would make the visual comparison of “overshooting” more immediate.","section":null},{"comment":"Section 5.4 (orthorexia data): the fitted Beta parameters (ã≈101.7, b̃≈1.67) are extreme relative to the earlier applications. A brief remark on whether the moment estimators remain numerically stable for such large shape values would be useful for practitioners.","section":null},{"comment":"References: the recent comprehensive review by Farrar et al. (2025) is cited; a short sentence locating HCβ relative to the bias-adjusted or residual-based estimators surveyed there would help readers place the contribution.","section":null},{"comment":"Typographical: “COV ARIANCE” in the running title; “quasi-ttest” in the keywords; occasional missing spaces after periods in the abstract and introduction.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid, well-executed incremental contribution that fits a methods journal. The free constants are the only real soft spot; once the authors supply a sensitivity check or a hold-out calibration, the manuscript should be ready for acceptance. No concerns about novelty disclosure or citation patterns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new piece is the adjustment factor itself: they treat 1-ht as the argument of a Beta CDF whose parameters are moment-estimated from the observed leverages, shrink toward Uniform(1,1), truncate to [0.01,0.99], and raise the reciprocal to a decaying power c1/n^c2. That is not in the HC0–HC5m literature. It is a genuine technical tweak, not a paradigm shift, but it is cleanly engineered and it works on the designs they study.\n\nWhat they do well is concrete. The asymptotic argument is elementary and correct: with p fixed the exponent goes to zero so gt\to1 and you recover HC0-like behavior. The Monte Carlo (10k reps, two designs, three λ levels, n=50/100/200) shows clear size and coverage gains over HC0/HC3/HC4/HC4m under strong heteroskedasticity and high leverage; the tests are mildly conservative rather than liberal, which is the safer direction. The four empirical examples, especially the gt-versus-ht plots, make the “overshooting” problem of HC4/HC4m visible and show that HCβ keeps the max gt modest (around 4–7 instead of tens or thousands). Shipping the hcinfer package with the estimator as default is real reproducibility, not window dressing.\n\nThe soft spots are exactly the free constants the reader flagged: c1=7, c2=0.75, the 0.01/0.99 truncation, and the shrinkage constant 50. They were chosen by Monte Carlo search on designs similar to the ones later used for performance claims. That is mild circularity, not fatal. There is no higher-order analytic expansion, and the paper does not explore sensitivity of the defaults across a wider class of designs. Those are real limitations, but they do not overturn the central claim under the conditions the paper actually states.\n\nThis is for people who already use sandwich estimators and care about finite-sample size under leverage. A serious editor should send it to referees; the construction is sound enough and the evidence is sharp enough to deserve that time. I would cite it when I next need a stable HC option under high leverage, and I would bring the package and the gt plots to reading group.","headline":"Clean, usable new HC estimator that tames overshooting under leverage; free constants are the only real soft spot, and the package makes it immediately checkable.","tokens_in":24039,"tokens_out":557,"would_cite":true,"duration_ms":6238,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62J05","62F12","62F03"],"pacs":[],"model":"grok-4.5","headline":"A Beta-fitted leverage adjustment yields more stable heteroskedasticity-robust standard errors for OLS without the explosive overshoot of older HC methods.","keywords":["Beta distribution","heteroskedasticity","leverage","linear regression","HC estimator","quasi-t test","sandwich covariance"],"falsifier":"Re-run the same Monte Carlo designs (or new designs with different leverage and heteroskedasticity patterns) using a systematically different triple of constants (c1, c2, shrinkage) and check whether the size and coverage advantage of HCβ over HC3/HC4/HC4m disappears or reverses.","tokens_in":24061,"feed_emoji":"📐","tokens_out":635,"duration_ms":6392,"temperature":0.7,"pith_summary":"Ordinary least squares remains consistent under unknown heteroskedasticity, but the usual sandwich covariance estimators can either understate variance or explode when a few observations have high leverage. This paper replaces the fixed or piecewise powers of (1-h) used by HC2–HC4m with a single data-driven factor taken from the cumulative distribution function of a Beta distribution whose two shape parameters are estimated (and lightly shrunk) from the sample leverages themselves. The resulting HCβ estimator automatically adapts to asymmetric or heterogeneous leverage patterns, keeps the correction factors from becoming huge, and still converges to the classical White form as the sample grows. Simulations and four real data sets show that the associated quasi-t tests and confidence intervals stay closer to their nominal levels than the leading alternatives, especially when leverage is strong and n is only moderate. An accompanying R package makes the estimator immediately usable.","feed_headline":"Beta-fitted leverages tame explosive HC standard errors","feed_subtitle":"Data-driven adjustment keeps size accurate under strong leverage without overshooting","key_machinery":"The HCβ adjustment factor gt = [n/(n-p)] × [1 / F_Beta(wt; ã, b̃)]^(c1/n^c2), where the Beta parameters are method-of-moments estimates of the truncated leverages, lightly shrunk toward (1,1), and the exponent decays with sample size.","core_discovery":"Replacing the uniform-based term (1-ht) that appears in classical HC estimators by the CDF of a Beta distribution fitted to the observed leverages produces a heteroskedasticity-consistent covariance matrix whose finite-sample size and coverage are more accurate and whose adjustment factors remain bounded, while the estimator remains asymptotically equivalent to White’s HC0.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Beta CDF of leverages bounds HC adjustments for accurate size","Fitted Beta replaces (1-h) terms in heteroskedasticity-consistent estimators","Data-driven Beta correction yields stable HC covariance under leverage","Beta-fitted leverages keep HC standard errors accurate without overshoot","Leverage Beta fit produces bounded HC estimator asymptotically like HC0"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The numerical constants that control truncation, shrinkage strength, and the decay rate of the exponent are chosen solely by Monte Carlo trial-and-error and are presented as universal defaults.","fun_headline_variants_meta":{"raw":{"variants":["Beta CDF of leverages bounds HC adjustments for accurate size","Fitted Beta replaces (1-h) terms in heteroskedasticity-consistent estimators","Data-driven Beta correction yields stable HC covariance under leverage","Beta-fitted leverages keep HC standard errors accurate without overshoot","Leverage Beta fit produces bounded HC estimator asymptotically like HC0"]},"model":"grok-4.5","effort":"low","cost_usd":0.004304,"raw_usage":{"total_tokens":1206,"prompt_tokens":685,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":43040000,"prompt_tokens_details":{"text_tokens":685,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":449,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":685,"tokens_out":72,"duration_ms":4640,"temperature":1.0,"reasoning_tokens":449,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T08:23:15.806011+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same Monte Carlo designs (or new designs with different leverage and heteroskedasticity patterns) using a systematically different triple of constants (c1, c2, shrinkage) and check whether the size and coverage advantage of HCβ over HC3/HC4/HC4m disappears or reverses.","supporting_citations":[],"review_version":1}