Pith. sign in

REVIEW 4 major objections 6 minor 8 references

Nonparametric modeling cash flows of insurance company

T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This paper claims that a nonparametric qED estimator applied to censored insurance registers reconstructs the joint distribution of claim size and reporting delay, from which net premium, claims frequency, IBNR and outstanding claims…

desk verdict A conditional but useful extension of qED reserving to deductibles; referees should press on the censoring-independence assumption. read the letter →

arxiv 1908.05200 v1 pith:XNIJT5QW submitted 2019-08-14 q-fin.RM q-fin.MF

classification q-fin.RMq-fin.MF MSC 62G0562N0162P05
keywords claimsreservesIBNRnetpremiumcensoreddataqEDestimatormultivariatedistributionfunctionnon-lifeinsuranceadditiveestimates
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that the joint distribution of claim size and reporting delay can be reconstructed nonparametrically from the censored registers insurers already keep, and that this single distribution is enough to derive net premium, claims frequency, IBNR reserves, outstanding claims reserves, and their interval estimates. On Russian non-life data covering 271,674 contracts, the resulting claims reserve estimate is about 135.6 million roubles, inside the 116.7–147.6 million range produced by chain-ladder, frequency-severity and Bornhuetter-Ferguson methods. Because each estimate is a sum over individual claim days, the method can redistribute reserves across portfolio segments and time horizons. A simulation with log-normal claim sizes and gamma reporting delays reports nearly unbiased IBNR estimates, with error variance falling sharply as portfolios grow. A sympathetic reader would take the paper's contribution to be a single distributional engine that replaces several separate reserving and pricing procedures.

What carries the argument

The load-bearing object is the qED estimator, a generalized maximum-likelihood estimate of a multivariate distribution from censored data, computed by an EM-type iterative algorithm on censoring sets C_k = C^s_k × C^tau_k. Each censoring set contains the true (claim size, reporting delay) pair even when the claim is unreported, below a deductible, or not yet settled — so the sample is censored but not truncated. The resulting estimate of F(s,tau) is then turned into expected claim amounts, variances, and claim probabilities by formulas that sum over individual exposure days, and the additivity of those sums is what allows reserve redistribution by segment or time interval.

What would settle it

Compare average settled claim sizes across reporting-delay buckets in the same portfolio: if short-delay claims are systematically larger or smaller than long-delay claims (beyond sampling noise), the assumed independence of censoring time from (s,tau) breaks down and the reserve and premium estimates from the qED procedure are biased. A simpler calendar check is whether inflation-adjusted average claim size rises with occurrence date within a fixed delay bucket.

Watch

Extended reading notes

Core claim

The central claim is that a censored sample of two-dimensional vectors (claim size, reporting delay) still contains complete information about their joint distribution, provided censoring sets record all possible values consistent with what is observed. Using the qED estimator — an EM-type maximum-likelihood analogue of the empirical distribution for censored data — the authors estimate F(s, tau) from policy and claim registers, including deductibles, unreported claims, and zero claims. From F(s, tau) they compute additive estimates of the net premium, claims frequency, IBNR and outstanding claims reserves, and they report that the reserve estimate of 135,556,927 roubles falls within the 116.7–147.6 million roubles range of traditional reserving methods. The authors would state that the main discovery is not a new estimator but a demonstration that one nonparametric joint distribution, estimated once from raw registers, yields the whole suite of cash-flow indicators an insurer needs.

Load-bearing premise

The censoring time (time from claim occurrence to the reporting date) is assumed independent of claim size and reporting delay, so the joint distribution factorizes; if big claims are reported faster or inflation raises claim sizes over time, the estimated distribution and all derived reserves and premiums are biased.

Editorial extensions

If this is right

  • Net premium and claims frequency for any insurance period follow directly from the marginal claim distribution; for the studied data the one-year figures are roughly 4,067 roubles and 5.68 percent.
  • The claims reserve, IBNR and outstanding claims reserve are obtained from the same joint distribution, with the point estimate (about 135.6 million roubles) landing inside the range of chain-ladder, frequency-severity and Bornhuetter-Ferguson estimates.
  • Because the estimates are additive over sample elements, reserves can be redistributed across portfolio segments or future reporting intervals without re-estimating the model.
  • Interval estimates can be built from a normal approximation to the reserve distribution; simulation results show the approximation is close for large portfolios (98% tolerance intervals within about 1–6% of simulated intervals).
  • The accuracy study indicates IBNR estimates are nearly unbiased, with a slight upward bias for portfolios above 500 policies, which the authors interpret as a built-in actuarial margin.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own homogeneity caveat points to a testable extension: on real data, check whether average settled claim size varies with reporting delay; if it does, the independence assumption fails and the reserve point estimate would need inflation or dependence corrections.
  • Replacing the reporting delay by time to settlement in the joint distribution is a natural extension, though the paper notes it makes censoring sets less informative; a numerical comparison on the same portfolio would show the accuracy tradeoff.
  • Because all estimates are additive, the same distribution could be used for segment-level capital allocation, reinsurance pricing, or regulatory reporting without rebuilding the model.
  • The method could serve as a benchmark against parametric or survival-analysis alternatives on censored insurance data, comparing reserve estimates and their variability.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a nonparametric methodology for modeling insurance cash flows by estimating the joint distribution F(s, τ) of claim size and reporting delay from censored insurance data using the quasi-empirical distribution (qED) estimator. The estimated distribution is used to derive additive point and interval estimates of net premium, claims frequency, IBNR, and incurred-but-outstanding reserves. The method is illustrated on real data from Russian non-life insurers, where the resulting reserve estimate falls within a range of estimates from traditional reserving methods, and on simulations that assess the accuracy of IBNR estimates.

Significance. If the methodological assumptions hold, the paper offers a flexible, fully nonparametric alternative to traditional reserving that yields additive estimates, which is genuinely useful for segment-level reserving and cash-flow projection. The authors explicitly connect the statistical estimates to actuarial quantities and show how censoring sets arise from the regulatory data structure. Strengths include the real-data demonstration, the simulation study of IBNR accuracy, and the explicit acknowledgment of the homogeneity assumption's limitations in Section 6.

major comments (4)
  1. [Section 2, Eq. (6)–(8)] The paper's central identification step is the assertion that the censoring time t_k = t − τ^1_k is independent of (s, τ), giving F(s, τ, t) = F(s, τ)F(t). This assumption is load-bearing because the qED estimator in Eq. (9) requires non-informative censoring; if large claims are reported faster or if inflation makes claim size vary with occurrence date, the censored sample is length-biased and F(s, τ) is not identified. The authors acknowledge this in Section 6, but the simulation in Section 5 generates data under the same independence assumption and the real-data validation does not test it. Please add a sensitivity analysis (for example, comparing estimated marginal distributions across occurrence-year cohorts) or a formal test of the independence assumption.
  2. [Section 4, Eqs. (10) and surrounding text] The real-data validation is a single point-in-range comparison: the proposed reserve estimate (135.6 million rubles) falls within the 116.7–147.6 million range from traditional methods. This is weak evidence because the traditional range is wide and a wide range of distributions could produce a reserve in that interval. To support the central claim, please report additional comparisons for other estimated quantities (e.g., IBNR, claims frequency, net premium) and ideally provide uncertainty intervals for the proposed estimate, not just a point value.
  3. [Section 5, Eq. (20) and simulation description] The simulation study is under-specified. Eq. (20) uses the same symbol K for both the number of simulation samples and the error ratio, which is confusing. The text does not fully specify how the censored sample is constructed for zero claims, limitation periods, or the grouping into time intervals, nor how the 98% tolerance intervals are computed. The simulation also assumes independence between claim size and reporting delay, so it cannot detect the main identifiability concern. Please provide complete algorithmic details and report results for claims frequency and OCR as well as IBNR.
  4. [Section 3, Eq. (9)] The consistency, asymptotic normality, and maximum-likelihood interpretation of the qED estimator are cited from the authors' own prior work (refs [1], [4]) without stating the sufficient conditions under which Eq. (9) is a valid estimator for the specific censoring structure used here (with truncation sets and zero-claims mass). Since all downstream estimates inherit the properties of this estimator, the paper should either state the relevant theorem and verify its conditions for the present setup or provide an accessible statement of the required conditions.
minor comments (6)
  1. [Section 2, Eq. (8)] In the third case of Eq. (8), the condition 'tk≤tk' appears to be a typo; it should presumably be 'tk≤t' or the intended inequality involving the reporting date.
  2. [Section 5, Eq. (20)] The symbol K is overloaded: it denotes the number of simulation samples in the text and the relative error ratio in Eq. (20). Please use a different symbol for the ratio, such as ε.
  3. [Section 6] The sentence 'For one-column wide figures use' appears to be a LaTeX instruction accidentally left in the text and should be removed.
  4. [Abstract and Section 1] The word 'inshurance' appears in the abstract and in Section 1; it should be 'insurance'.
  5. [Section 4] The claim amount 11.141814 is reported without units; the text later uses 'P' for rubles, but consistent notation would improve readability.
  6. [Table 3] The table is dense and does not clearly indicate how the grouped counts correspond to the censoring sets defined in Eq. (8); a short explanatory paragraph or footnote would help.

Circularity Check

1 steps flagged · score 2.0 of 10

No structural circularity; minor self-citation burden from relying on the authors' own qED estimator for the method's statistical guarantees.

  1. self citation load bearing [Section 1 (Introduction), p. 3; Section 3, Eq. (9)]
    "This paper applies the ideas presented in one of the authors’ works [2], [3] and generalizes them for the case of insurance policies with deductibles, applying qED estimate of the multivariate distribution function for censored data [4]."

    The central statistical engine is the qED estimator, whose consistency and asymptotic normality are asserted by reference to previous work by the same author ([1], [4]) rather than proved or independently validated in this manuscript. All downstream reserve, IBNR and OCR estimates are functionals of the resulting F(s,tau), so the theoretical pedigree of the outputs rests on this self-citation chain. However, the numerical predictions are not themselves defined by the citation: they are new functionals (Eqs. 10-18) applied to new data, and the Section 5 simulation independently validates the whole pipeline against a separate generative model. Thus the self-citation is load-bearing for the method's pedigree but does not make the empirical claims tautological.

full rationale

The paper does not fit its parameters to the IBNR or reserve targets. The joint distribution F(s,tau) is estimated from the censored sample by the qED algorithm, and net premium, claims frequency, IBNR and OCR are then computed as explicit functionals of that distribution (Eqs. 10-18). The reserve calculation in Eq. (10) is an accounting identity (expected ultimate claims minus paid to date), and it is benchmarked against independent traditional reserving methods (chain-ladder, frequency-severity, Bornhuetter-Ferguson) rather than derived from them. The Section 5 simulation uses a separate generative model (log-normal severities, gamma reporting delays) and evaluates the estimator's relative error against the simulated true IBNR, so it is a genuine check of the estimation pipeline rather than a refit of the same target. The only circularity-adjacent feature is the reliance on the authors' own prior work for the qED estimator's theoretical properties; this is a self-citation burden but not a reduction of the predictions to the inputs. The paper also explicitly concedes in Section 6 that the homogeneity/independence assumption may fail under inflation or claim-size/settlement-time correlation; that is an honest correctness limitation, not a circularity, because the derivation is conditional on that assumption rather than hiding it inside the inference.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper's central claim rests on the stated independence and homogeneity assumptions, on the cited properties of the qED estimator, and on a modeling convention that converts policies into daily exposures. No free parameters are fitted to the target data, and no new entities are introduced.

assumptions (5)
  • domain assumption The censoring time t_k is independent of the claim amount s and reporting delay tau, so F(s, tau, t) = F(s, tau)F(t).
    Stated in Section 2 to justify reducing the 3D problem to 2D. If violated (e.g., inflation or faster reporting of large claims), the estimated distribution and reserves are biased.
  • domain assumption The insurance portfolio is random and homogeneous over time, so the sample of daily claim indicators is identically distributed.
    Stated in Section 6 as the main assumption and disadvantage: claim sizes must not depend on calendar time (no inflation) and no settlement-time dependence.
  • ad hoc to paper Each policy-day interval contains at most one claim, and unreported claims within the limitation period are modeled as zero claims with tau = infinity.
    Introduced in Section 2 to turn policies into daily exposure units. This is a modeling convention specific to this paper.
  • standard math The qED estimator P*_n is consistent and asymptotically normal under the given censoring scheme.
    Properties are cited from Baskakov's prior work (refs [1], [4]) and not proven here. The paper relies on these to justify the estimates.
  • domain assumption The distribution of the total IBNR reserve is approximately Gaussian for large portfolios.
    Used in formula (19) for tolerance intervals; supported by simulation results in Section 5 but not proven analytically.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Nonparametric modeling cash flows of insurance company." pith.science (2026). https://pith.science/paper/XNIJT5QW

@misc{pith2026190805200,
  author       = {Pith},
  title        = {Pith review of: Nonparametric modeling cash flows of insurance company},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XNIJT5QW}},
  note         = {Machine review of arXiv:1908.05200}
}
read the original abstract

The paper proposes an original methodology for constructing quantitative statistical models based on multidimensional distribution functions constructed on the basis of the insurance companies' data on inshurance policies (including policies with deductible) and claims incurred. Real data of some Russian insurance companies on non-life insurance contracts illustrate some opportunities of the proposed approach. The point and interval estimates of net premium, claims frequency, claims reserves including IBNR and OCR, are thus obtained. The resulting estimate of claims reserves falls in the range of reasonable estimates calculated on the basis of traditional reserving methods (the chain-ladder method, the frequency-severity method and the Bornhuetter-Ferguson method). The proposed methodology is based on additive estimates of a company's financial indicators, in the sense that they are calculated as a sum of estimates built separately for each element of the sample (claim). This allows using the proposed methodology to model insurance companies' financial flows and, in particular, to solve the problems of reserve redistribution between particular segments of insurance portfolio and/or time intervals; to adjust risk as part of financial reporting under IAS 17 Insurance Contracts; and to deal with many other tasks. The accuracy of insurance companies' financial parameters estimate based on the proposed methods was tested by statistical modeling. IBNR was used as the test parameter. The modeling results showed a satisfactory accuracy of the proposed reserve estimates.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

8 extracted references · 8 canonical work pages

  1. [1]

    Baskakov V, Bartunova A (2019) Nonparametric estimation of multivariate distribution function for truncated and censored lifetime data. Eur. Actuar. J. 9:209–239 24 Valery Baskakov, Nikolay Sheparnev & Evgeny Yanenko 0 5 10 15 20 τ* 25 50 75 100 s k=5 k=1 k=12 k=11 Fig. 9 Censoring sets of the two-dimensional vector (s,τ ∗)

  2. [4]

    J Math Sci 81(4):2779–2785

    Baskakov V (1996) On an analog of empirical distribution for multivariate censored data. J Math Sci 81(4):2779–2785

  3. [2]

    Actuary (Russian) 5:21–25

    Baskakova A, Baskakov V (2014) IBRN reserves estimate on the bases of multivariate censored data of an insurance company. Actuary (Russian) 5:21–25

  4. [3]

    Actuary (Russian) 4:37–41

    Baskakov V, Baskakov I (2010) On ratemaking and other tasks in non-life insurance. Actuary (Russian) 4:37–41

  5. [5]

    Benjamin B (1977) General Insurance, Heinemann, London

  6. [6]

    J R Stat Soc, Ser B 39:1–38

    Dempster A, Laird N, Rubin D (1977) Maximum likelihood estimation from incomplete data. J R Stat Soc, Ser B 39:1–38

  7. [7]

    Springer, New York, p 536

    Klein JP, Moeschberger ML (2003) Survival analysis: techniques for censored and trun- cated data. Springer, New York, p 536

  8. [8]

    Variance Advancing the Science of Risk 2(1):85–110

    Schmidt KD, Zocher M (2008) The Bornhuetter-Ferguson principle. Variance Advancing the Science of Risk 2(1):85–110

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.