{"id":"9bdf46cf-1ee4-4676-9ac0-db89ef4cfffd","arxiv_id":"1908.03390","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new tail-index estimator blends censored-data Hill estimates with expert guesses through an entropy-perturbed likelihood, and reduces to a simple weighted harmonic mean when the expert is correct.","lead":"The paper introduces a penalized likelihood estimator that combines censored insurance claim data with expert opinions about the tail index, producing a simple weighted combination of the classical Hill estimator and the expert guess. Actuaries and statisticians working with heavy-tailed, heavily censored data get a principled way to incorporate expert knowledge without extra assumptions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Corollary 4.4's variance reduction for λ=1 is contingent on β=α exactly; any fixed expert misspecification makes the combined estimator inconsistent, and no diagnostic is offered.","rationale":"The reader's verdict already conditions on expert quality; my analysis sharpens that condition into a precise asymptotic statement. The paper's own Theorem 4.1 supplies the centering (λα2/β+1)/(λα2+α), so for λ=1 the reciprocal estimator converges to p/α+(1-p)/β, not 1/α unless β=α. Hence the variance reduction of Corollary 4.4 is not a free lunch: it buys variance reduction by trading away consistency when the expert is wrong. Since Remark 5.1 rules out cross-validation and no data-driven λ is proposed, a practitioner has no diagnostic to tell whether the method is in the favourable regime. This matches the reader's weakest assumption. The theoretical derivation is internally consistent, and the paper states its clairvoyant condition, so the concern is about the practical default recommendation rather than a mathematical contradiction. The proposed simulation with ±10% expert error would settle whether the finite-sample MSE advantage survives realistic misspecification. No change to the CONDITIONAL verdict is needed.","tokens_in":16281,"tokens_out":12520,"duration_ms":124384,"concrete_test":"Simulate the exact Pareto model of Section 5 with n=200, true α=2, censoring fraction 1-p≈0.5, and fixed expert values β∈{1.8,2.0,2.2} (i.e., ±10% error); for k=20,50,100 and Nsim=1000, compare empirical MSE of 1/α̂^P_k (λ=1) with 1/α̂^MLE_k from (21). If the combined estimator's MSE exceeds the censored Hill MSE for any k once β differs from α by 10%, the variance-reduction claim of Corollary 4.4 is not robust to realistic expert misspecification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The pivotal result is Corollary 4.4: when λ=1, β=α and δ=0, 1/α̂^P_k = p̂_k(1/α̂^MLE_k)+(1-p̂_k)/β, with asymptotic variance 1/[kp(α+α2)^2], smaller than 1/(kpα²) from (21). This is local to β=α. Theorem 4.1 centers the reciprocal estimator at (λα2/β+1)/(λα2+α), so for λ=1 the probability limit is p/α+(1-p)/β (since p̂_k→p and H_k→1/(α+α2)=p/α). Thus, even with δ=0, any fixed expert misspecification β≠α leaves a permanent bias (1-p)(1/β−1/α); the combined estimator is inconsistent while censored Hill remains consistent. Its AMSE then tends to a positive constant instead of 0, so for all sufficiently large k the claimed enhancement reverses. No safeguard is offered: Remark 5.1 rejects cross-validation for extreme-value applications, and the simulation only uses a centered expert guess with σ=0.2 around the true ξ. The case study derives β from the ultimates of the same dataset, so it cannot validate independence of the expert information. The practical recommendation λ=1 therefore rests on the expert being effectively clairvoyant.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a penalized-likelihood estimator for the tail index of Pareto-type data under random right censoring, where each censored observation is accompanied by an expert-supplied tail index. The penalty is the relative entropy between Pareto densities, which yields an explicit estimator of Hill type that interpolates between the censored maximum-likelihood (Hill) estimator and the expert information as the penalization strength λ varies. Section 3 gives a Bayesian/gamma-prior interpretation, Section 4 derives asymptotic normality under the Hall class, and Corollary 4.4 shows that for λ = 1, β = α, and δ = 0, the reciprocal estimator is a weighted average of the censored Hill estimator and the expert tail index, with asymptotic variance 1/[kp(α + α2)^2]. Section 5 reports simulations against censored Hill and two Bayesian estimators, and applies the method to motor third-party liability insurance claims.","tokens_in":16623,"tokens_out":4191,"duration_ms":44975,"significance":"If the main claims hold, the paper provides a simple, explicit, and interpretable rule for combining expert tail information with censored data, together with a Bayesian analogy and a quantile extension. The derivations in Section 2 and Appendix A are largely explicit, and the asymptotic variance formula in Theorem 4.1 is a useful technical contribution. However, the advertised variance reduction and the λ = 1 recommendation are proven only under a clairvoyant expert condition (β = α) and a vanishing second-order bias condition (δ = 0), which the paper itself acknowledges only in passing. The simulation evidence conditions on expert guesses that are centered at the true value, and the case study derives β from the same dataset, so the empirical support is weaker than the wording suggests. The contribution is therefore valuable as a conditional combination tool, but the paper needs to address misspecification and provide safeguards before the practical recommendation can be accepted.","major_comments":[{"comment":"The centering in Theorem 4.1 shows that for λ = 1 and δ = 0 the probability limit of 1/α̂^P_k is (1 + α2/β)/(α + α2) = p/α + (1 − p)/β, which equals 1/α only when β = α. For any fixed expert misspecification β ≠ α, the estimator 1/α̂^P_k has permanent asymptotic bias (1 − p)(1/β − 1/α), so its AMSE tends to a positive constant while the censored Hill estimator remains consistent. The variance reduction claimed in Corollary 4.4 therefore holds only in the clairvoyant case, and the paper does not provide an analysis of the bias–variance trade-off for misspecified experts. I would ask the authors to add an explicit misspecification analysis, at least for the Hall class with δ = 0, and to temper or qualify the recommendation of λ = 1 accordingly.","section":"§4, Theorem 4.1 and Corollary 4.4"},{"comment":"The simulations generate the expert input by drawing a single random number from a Gaussian distribution centered at the true ξ with standard deviation 0.2, and the Bayesian competitor is given the true variance by moment matching. This makes the comparison favorable by construction. The text in §5.1 concedes that changed conditions can make both the perturbed and Bayesian estimators perform much worse, but no such scenario is reported. Since the central empirical claim is that the estimator 'often outperforms' the benchmarks, the paper should include misspecified-expert scenarios, such as β drawn from a distribution with a nonzero mean shift or larger variance, and report bias and MSE relative to the censored Hill estimator.","section":"§5.1, simulation design"},{"comment":"The paper rejects cross-validation for selecting λ in Remark 5.1 but offers no alternative data-driven safeguard, and the conclusion reiterates that no tuning parameter selection is needed. Given that λ = 1 is only formally justified when β = α and δ = 0, the practical recommendation is incomplete: a practitioner with an uncertain expert guess has no guidance on when the combined estimator is preferable to censored Hill. I would ask for at least a diagnostic based on the discrepancy between the expert value and the censored-Hill estimate, or a clear statement identifying the conditions under which λ = 1 is safe to use.","section":"§5, Remark 5.1 and §6"}],"minor_comments":[{"comment":"The text refers to the 'Weismann estimator'; the correct spelling is Weissman, as in the references.","section":"§5.1"},{"comment":"The references to 'Theorem 4.4' should refer to Corollary 4.4, which is the statement being used.","section":"§5.2 and §6"},{"comment":"The final expression in Eq. (29) uses \\\\hat\\\\xi^P_k, but that notation is only introduced later in Eq. (31); please define \\\\hat\\\\xi^P_k before the quantile formula.","section":"§4, Eq. (29)"},{"comment":"The statement that the estimator is 'unbiased' should explicitly say 'asymptotically unbiased under the conditions β = α and δ = 0', since the unbiasedness is not a finite-sample property and depends on the clairvoyance assumption.","section":"§4, Corollary 4.4"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real new thing here is the entropy-perturbed likelihood estimator for Pareto-type tails with random censoring and expert guesses on the tail index. It is not a repackaged Hill estimator: the explicit formula, the asymptotic normality under Hall-class assumptions, and especially Corollary 4.4's combination identity are original and cleanly derived. The Bayesian-prior analogy is a nice interpretive bonus, and the paper is honest that it is not a true Bayesian method.\n\nThe derivations hold up. I checked Theorem 4.1's algebra in outline; the variance and bias expressions are plausible and the cancellation that kills the r2 term is real. The proof sketch in the appendix is thin in places but the ingredients are standard, and the omitted algebra is routine. The simulation study is fair to the competitors: the Bayesian gamma estimator is given extra information (true variance) while the proposed estimator uses only λ=1, and the new estimator still wins often. That is real evidence.\n\nThe soft spot is exactly what the stress-test note says: the λ=1 rule is justified only when β=α and second-order bias δ=0. With any fixed misspecification β≠α, the combined estimator has a permanent asymptotic bias (1−p)(1/β−1/α) and is inconsistent, while censored Hill stays consistent. So the variance reduction advertised in Corollary 4.4 is local to a clairvoyant expert. The paper acknowledges the dependence on expert quality in Section 5 but offers no diagnostic or data-driven safeguard; Remark 5.1 dismisses cross-validation for good reasons, but that leaves practitioners with no way to check whether the expert guess is in the right ballpark. Also, the case study uses β derived from the same dataset's ultimates, which cannot validate the independence of the expert information. These are addressable issues: a sensitivity analysis with biased experts, or a data-driven check on β, would materially strengthen the paper.\n\nOverall, this is a solid methodological contribution for actuarial heavy-tail estimation. The central estimator is useful, the math is sound under stated assumptions, and the limitations are honestly disclosed. A serious referee should engage with it. I would not desk-reject, and I would cite it if I worked on censored extremes.","headline":"A genuinely useful censored-tail estimator with an explicit expert combination formula, though the λ=1 rule leans on a clairvoyant expert and no safeguard is offered.","tokens_in":17092,"tokens_out":695,"would_cite":true,"duration_ms":10172,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G32","62N01","62P05"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an entropy-penalized likelihood can blend the censored Hill estimator with expert tail-index guesses, and that with a correct expert guess the blend is asymptotically normal and has lower variance than censored Hill…","keywords":["tail index estimation","random censoring","expert information","entropy-perturbed likelihood","Hill estimator","regular variation","extreme value statistics","liability insurance"],"falsifier":"Simulate from a known Hall-class model with tail index $\\alpha$ and censoring tail index $\\alpha_2$, feed the estimator the true value $\\beta=\\alpha$ and $\\lambda=1$, and check the empirical variance of the reciprocal estimator against $1/[kp(\\alpha+\\alpha_2)^2]$ for large $k$; if the realized variance does not fall below the censored Hill variance $1/(kp\\alpha^2)$, Corollary 4.4 is not confirmed.","tokens_in":16083,"feed_emoji":"📊","tokens_out":15872,"duration_ms":124557,"temperature":0.7,"pith_summary":"In liability insurance, claim sizes are often right-censored because claims stay open for years, but experts maintain projections for those open claims. This paper proposes using that expert information in heavy-tail estimation by adding an entropy-based penalty to the Pareto-type likelihood. Its main result is that for penalty strength $\\lambda=1$ the reciprocal of the new estimator is exactly a weighted average of the reciprocal of the censored Hill estimator and $1/\\beta$, with weights the observed proportions of uncensored and censored observations among the largest $k$ claims. Under the Hall second-order condition and a correct expert guess ($\\beta=\\alpha$) the reciprocal of the estimator is asymptotically normal with variance $1/[kp(\\alpha+\\alpha_2)^2]$, smaller than the censored Hill estimator's variance. The paper thereby offers a parameter-free way to combine closed-claim data with expert judgment, and a simulation study plus a motor third-party liability insurance dataset show it often beats censored Hill and two Bayesian benchmarks.","feed_headline":"Blend of censored Hill and expert guesses lowers tail-index variance","feed_subtitle":"With a trustworthy expert guess, the reciprocal estimate is a censoring-weighted average with smaller asymptotic variance.","key_machinery":"The engine is the entropy-perturbed likelihood: multiply the usual censored Pareto likelihood by $e^{-\\lambda(1-e_i)D(\\alpha,\\beta_i)}$, where $D$ is the Kullback-Leibler divergence between a Pareto density with tail index $\\alpha$ and the expert's Pareto density with tail index $\\beta_i$. That specific divergence makes the maximizer explicit, turns the penalty into a gamma-type prior whose strength is governed by the censoring indicators, and at $\\lambda=1$ collapses to the weighted-average identity in Corollary 4.4. The asymptotic results then flow from the Hall-class expansion of the censored tail quantile function and from the known joint limit behaviour of the Hill statistic and the censoring proportion.","core_discovery":"Working with exact Pareto distributions first, the authors perturb the censored-data likelihood by a factor $e^{-\\lambda(1-e_i)D(\\alpha,\\beta_i)}$ with $D$ the Kullback-Leibler divergence between Pareto densities, so the log-likelihood becomes $\\sum_i e_i\\log f_\\alpha(Z_i)+\\sum_i(1-e_i)\\log\\bar F_\\alpha(Z_i)-\\lambda\\sum_i(1-e_i)D(\\alpha,\\beta_i)$. The maximizer is explicit: $\\hat\\alpha^P(\\lambda)=\\sum_i(e_i+\\lambda(1-e_i))/\\sum_i(\\log(Z_i/x_0)+\\lambda(1-e_i)/\\beta_i)$, which reduces to the censored Hill estimator as $\\lambda\\to0$ and to the experts' harmonic mean as $\\lambda\\to\\infty$. For Pareto-type tails in the Hall class, Theorem 4.1 establishes asymptotic normality of the reciprocal estimator, and Corollary 4.4 shows that when $\\lambda=1$, the second-order bias vanishes ($\\delta=0$), and the expert guesses equal the true index ($\\beta=\\alpha$), one has $1/\\hat\\alpha^P_k=\\hat p_k/\\hat\\alpha^{MLE}_k+(1-\\hat p_k)/\\beta$; that is, the inverse estimate is the censoring-weighted average of the inverse censored-Hill estimate and the inverse expert tail index. Its asymptotic variance is $1/[kp(\\alpha+\\alpha_2)^2]$, improving on the censored Hill benchmark. The same combination is used to define a quantile estimator, and simulations plus an MTPL insurance case study support the method.","pith_inferences":["Because the variance ratio is $\\alpha^2/(\\alpha+\\alpha_2)^2$, the largest gains from expert information occur exactly in the high-censoring regime that motivates the paper; a natural test would be to measure the estimator's advantage as a function of the censoring fraction.","The same entropy-penalty mechanism could replace the censored Hill baseline with bias-reduced or trimmed Hill estimators inside the combination formula, which the paper flags as future work on trimming.","The Bayesian reading suggests an explicit hierarchical extension: treat each expert $\\beta_i$ as a draw from a prior with its own uncertainty and let $\\lambda$ control the prior's weight, which would yield a diagnostic for when expert information is too noisy to use.","Since the paper gives no data-driven way to choose $\\lambda$, a practical extension is to select $\\lambda$ by the stability of $\\hat\\alpha^P_k$ across thresholds $k$, or by comparing it with the censored Hill estimator to detect expert bias."],"forward_implications":["With a reliable expert guess, the estimator replaces tuning by a fixed rule: the reciprocal estimate is $\\hat p_k$ times the reciprocal censored-Hill estimate plus $(1-\\hat p_k)/\\beta$, so the data weight is exactly the observed fraction of uncensored observations among the largest $k$ claims.","In Hall-class models with a clairvoyant expert, the asymptotic variance of the reciprocal estimator is $1/[kp(\\alpha+\\alpha_2)^2]$, strictly smaller than the censored Hill variance $1/(kp\\alpha^2)$, and the reduction grows as the censoring tail becomes heavier.","The same weighted-average rule carries over to quantile estimation: the combined high quantile is a weighted geometric combination of the Kaplan-Meier based extrapolation and the expert-based extrapolation, yielding a stable compromise between the two.","As $\\lambda\\to0$ the estimator collapses to the censored Hill estimator and as $\\lambda\\to\\infty$ it collapses to the expert's harmonic mean, so the method forms a continuum with the proposed $\\lambda=1$ as a natural no-tuning choice.","In the paper's simulations across Pareto, Burr, and Fréchet tails, the combined estimator frequently has lower bias and MSE than censored Hill and the two Bayesian benchmarks, especially for heavy tails and for quantiles under non-Pareto tails."],"supporting_citations":[{"why":"Introduces the censored Hill estimator that the perturbed estimator generalizes and reduces to as the penalty vanishes.","marker":"[2]"},{"why":"Establishes the asymptotic distribution of extreme value estimators under random censoring, including the joint behaviour used in Theorem 4.1.","marker":"[3]"},{"why":"Defines the classical Hill estimator for the tail index, the basis of the censored version.","marker":"[8]"},{"why":"Provides the Hall-class second-order condition that the asymptotic bias and variance expansions rely on.","marker":"[12]"},{"why":"Supplies the tail quantile extrapolation used to derive the combined quantile estimator.","marker":"[13]"},{"why":"One of the Bayesian tail-index estimators under random censoring used as a benchmark in the simulations.","marker":"[5]"},{"why":"Describes the liability insurance setting and supplies the motor third-party liability dataset used in the case study.","marker":"[1]"},{"why":"Gives the Kaplan-Meier estimator used to construct the base quantile in the combined quantile estimator.","marker":"[14]"},{"why":"Standard reference for the asymptotic theory of Hill-type estimators under second-order conditions.","marker":"[10]"},{"why":"Provides the trimming-based threshold selection that produces the expert tail index used in the insurance application.","marker":"[15]"}],"fun_headline_variants":["Hill plus expert: lower tail variance under censoring","Censored Hill + expert guesses: smaller tail-index variance","Weighted blend of Hill and expert guess lowers tail variance","Expert-augmented censored Hill estimator: lower variance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the expert's tail-index guess $\\beta$ is actually right, because the clean unbiasedness and variance reduction only hold for $\\beta=\\alpha$; with a biased or uninformative guess, the $\\lambda=1$ estimator can be more biased than the censored Hill estimator that ignores the expert entirely.","fun_headline_variants_meta":{"raw":{"variants":["Hill plus expert: lower tail variance under censoring","Censored Hill + expert guesses: smaller tail-index variance","Weighted blend of Hill and expert guess lowers tail variance","Expert-augmented censored Hill estimator: lower variance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000954,"raw_usage":{"total_tokens":4141,"prompt_tokens":1089,"completion_tokens":3052,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":705,"completion_tokens_details":{"reasoning_tokens":2985}},"tokens_in":705,"tokens_out":3052,"duration_ms":25384,"temperature":1.0,"reasoning_tokens":2985,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:15:25.677907+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate from a known Hall-class model with tail index $\\alpha$ and censoring tail index $\\alpha_2$, feed the estimator the true value $\\beta=\\alpha$ and $\\lambda=1$, and check the empirical variance of the reciprocal estimator against $1/[kp(\\alpha+\\alpha_2)^2]$ for large $k$; if the realized variance does not fall below the censored Hill variance $1/(kp\\alpha^2)$, Corollary 4.4 is not confirmed.","supporting_citations":[{"cited_title":"Bias reduced tail estimation for censored pareto type distributions","cited_arxiv_id":null,"evidence_quote":"Introduces the censored Hill estimator that the perturbed estimator generalizes and reduces to as the penalty vanishes."},{"cited_title":"Statistics of extremes under random censoring","cited_arxiv_id":null,"evidence_quote":"Establishes the asymptotic distribution of extreme value estimators under random censoring, including the joint behaviour used in Theorem 4.1."},{"cited_title":"A simple general approach to inference about the tail of a distribution","cited_arxiv_id":null,"evidence_quote":"Defines the classical Hill estimator for the tail index, the basis of the censored version."},{"cited_title":"On some simple estimates of an exponent of regular variation","cited_arxiv_id":null,"evidence_quote":"Provides the Hall-class second-order condition that the asymptotic bias and variance expansions rely on."},{"cited_title":"Estimation of parameters and large quantiles based on the k largest observations","cited_arxiv_id":null,"evidence_quote":"Supplies the tail quantile extrapolation used to derive the combined quantile estimator."},{"cited_title":"Bayesian estimation of the tail index of a heavy tailed distribution under random censoring","cited_arxiv_id":null,"evidence_quote":"One of the Bayesian tail-index estimators under random censoring used as a benchmark in the simulations."},{"cited_title":"John Wiley & Sons, 2017","cited_arxiv_id":null,"evidence_quote":"Describes the liability insurance setting and supplies the motor third-party liability dataset used in the case study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the Kaplan-Meier estimator used to construct the base quantile in the combined quantile estimator."},{"cited_title":"Statistics of Extremes: Theory and Applications","cited_arxiv_id":null,"evidence_quote":"Standard reference for the asymptotic theory of Hill-type estimators under second-order conditions."},{"cited_title":"Threshold selection and trimming in extremes","cited_arxiv_id":"1903.07942","evidence_quote":"Provides the trimming-based threshold selection that produces the expert tail index used in the insurance application."}],"review_version":1}