Pith. sign in

REVIEW 3 major objections 6 minor 27 references

Bayesian approach to Lorenz curve using time series grouped data

T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Bayesian state-space model tightens estimates of income inequality

desk verdict Useful state-space extension of the Dirichlet pseudo-likelihood for Lorenz curves, but the headline efficiency gain rests on an unvalidated precision calibration that could make the intervals overconfident. read the letter →

arxiv 1908.06772 v1 pith:2OTOLYVK submitted 2019-08-19 stat.ME

classification stat.ME MSC 62F1562M1062P20
keywords LorenzcurveGinicoefficientDirichletdistributionstatespacemodelgroupeddataincomeinequalityBayesianinferencetimeseries
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Governments publish income data as grouped shares rather than individual records, and estimating inequality from a single survey period leaves wide uncertainty. This paper argues that modelling a whole sequence of grouped Lorenz-curve observations as a state space model, with the Lorenz curve parameters drifting through time as an autoregressive or random-walk process, lets each period borrow strength from its neighbours. The authors show in simulation and in Japanese monthly survey data that this shrinks credible intervals for the Gini coefficient and for the Lorenz curve by a large margin, with bias comparable to period-by-period Dirichlet estimation. Their central proposition is that the Dirichlet pseudo-likelihood for income shares, with precision proportional to the survey sample size, can serve as the observation equation of a time series model.

What carries the argument

The load-bearing object is the Dirichlet pseudo-likelihood, f(q_t|theta_t, lambda_t) = Gamma(lambda_t) prod_k q_{tk}^{lambda_t (L(p_{tk}|theta_t)-L(p_{t,k-1}|theta_t))-1} / Gamma(lambda_t (L(p_{tk}|theta_t)-L(p_{t,k-1}|theta_t))), which treats the income shares as a Dirichlet draw whose mean vector is the vector of Lorenz-curve increments. Around it, the paper builds a state space model: a link-transformed parameter vector u_t drives the Lorenz curve, evolves by AR(1) or random walk, and the Dirichlet precision lambda_t = n_t exp(psi) ties sampling noise to the survey size. The mechanism that produces the efficiency gain is the borrowing of information across time through the latent process, together with the sample-size-adapted precision, which stabilizes what was previously a nuisance parameter.

What would settle it

Simulate grouped income shares from a data-generating process that respects the Lorenz-curve means but assigns the shares a variance-covariance structure different from the Dirichlet's (for example, a Dirichlet-multinomial with an extra dispersion parameter, or a survey sampling scheme with clustering), then estimate the proposed state-space model and record the empirical coverage of the 95% credible intervals for the Gini coefficient over many replications. If coverage falls substantially below 95% while the single-period Dirichlet estimator maintains coverage, the claimed efficiency improvement is an artifact of the Dirichlet variance assumption.

Watch

Extended reading notes

Core claim

The paper's central claim is that the Dirichlet pseudo-likelihood of Chotikapanich and Griffiths, where each income share's expectation is the difference in Lorenz-curve heights between consecutive population shares, can be embedded as the observation equation of a state space model. The transformed parameters of the chosen Lorenz curve (using log or logit links) evolve under either an AR(1) process with |rho|<1 or a random walk, and the Dirichlet precision at time t is set to lambda_t = n_t exp(psi), so sampling variability scales with the survey's sample size. The authors maintain that this joint model yields posterior distributions for the Gini coefficient, the Lorenz curve, and the income-distribution parameters that are far more concentrated than those from fitting the Dirichlet model period by period, while retaining comparable relative bias. On Japanese Family Income and Expenditure Survey data, the Kakwani Lorenz curve with a random walk latent process gives the lowest posterior predictive loss, and the estimated Gini coefficient declines after 2008.

Load-bearing premise

The Dirichlet pseudo-likelihood with precision lambda_t = n_t exp(psi) correctly describes the sampling variability of the grouped income shares; if the actual share variance does not scale this way, the narrow credible intervals are overconfident and the efficiency gain is an artifact.

Editorial extensions

If this is right

  • The same framework can be applied to any parametric Lorenz curve or income distribution whose parameters can be link-transformed, not just the six families considered in the paper.
  • Inequality measures such as the Gini coefficient can be monitored monthly or quarterly with much narrower uncertainty, making trend breaks and turning points more detectable.
  • The sample-size scaling lambda_t = n_t exp(psi) offers a template for combining surveys of different sizes, so national and regional surveys could be pooled in one analysis.
  • Model comparison via posterior predictive loss selects among Lorenz curves and latent processes; here the Kakwani curve with random walk wins for Japanese data.
  • Because precision is estimated from all periods jointly, the paper's approach resolves, at least within its model, the lack of guidance on choosing the Dirichlet precision that plagued single-period estimation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the Dirichlet variance calibration is wrong, the headline reduction in credible-interval length may be overconfidence: a natural check is to generate grouped shares from a sampling mechanism with overdispersion and measure coverage of the proposed intervals.
  • The model's improvement will likely be largest when the true latent process is smooth; for rapidly changing inequality, an AR(1) with strong persistence may oversmooth genuine breaks, and regime-switching or shrinkage extensions would be worth testing.
  • The precision parameter psi could be interpreted as an effective survey design effect; letting psi vary by survey (rather than one global value) might capture changes in survey methodology over long panels.
  • Since the Lorenz curve is location-free, the approach estimates relative inequality only; combining it with a location model, as the authors note, would recover the full income distribution.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a Bayesian state-space model for estimating Lorenz curves and inequality measures from time series grouped income data. The observation equation uses a Dirichlet pseudo-likelihood in which the expected income shares are the differences of a parametric Lorenz curve at consecutive class boundaries, and the Dirichlet precision is parameterized as lambda_t = n_t exp(psi) so that survey sample sizes inform sampling variability. The transformed Lorenz-curve parameters evolve according to independent AR(1) or random-walk latent processes, and the posterior is explored with a Gibbs/Metropolis-Hastings sampler. The method is illustrated with a simulation based on the Singh-Maddala distribution and with monthly Japanese Family Income and Expenditure Survey data from 2000-2018. The central claim is that the proposed time-series model yields substantially narrower credible intervals for the inequality measures than period-wise Dirichlet estimation, with comparable bias.

Significance. If the efficiency claim is validated with calibrated uncertainty, the paper offers a practical and flexible toolkit for inequality measurement from grouped time series data, which is a common data-release format in many countries. The state-space formulation, the sample-size-dependent precision, and the posterior predictive loss comparison across six Lorenz families are useful contributions. The paper is transparent about the MCMC algorithm and reports inefficiency factors, although no code or data are provided. The main value would be for applied researchers who want to estimate Gini coefficients and Lorenz ordinates with uncertainty from grouped survey data over time.

major comments (3)
  1. [Section 3.1 and Equations (2), (4), (7)] The central claim of improved efficiency rests entirely on the fact that the proposed credible intervals are much shorter than those from the separate Dirichlet approach, but the paper never checks whether those intervals have valid coverage. The Dirichlet pseudo-likelihood is not the true sampling model for the simulated data, since the data are generated by drawing individual incomes and then aggregating into shares (Section 3.1, steps 3-4). The posterior mean of psi reported in Table 1 is 4.428, which with lambda_t = n_t exp(psi) implies an observation variance roughly 84 times smaller than the variance of a binomial share from n_t independent draws; this suggests the model may be treating sampling noise as signal. A coverage analysis over the T = 500 simulated periods, or a posterior predictive check of the income shares, would tell whether the narrower intervals are calibrated or merely overconfident. Without such a check, the headline efficiency gain is not yet established.
  2. [Section 3.1 and Table 1] There is an internal inconsistency in the reported simulation design. The text states "We set eta1 = (1.25, 0.8, 0.015)' and eta2 = (0.4, 0.8, 0.02)'", implying rho2 = 0.8, but Table 1 reports the true value of rho2 as 0.5, with a posterior mean of 0.537 that matches 0.5 rather than 0.8. This ambiguity makes it impossible to verify the simulation evidence as reported. The authors should correct either the text or the table, and ideally report the actual true values used in the data-generating process.
  3. [Section 3.1, Figures 1-2] The efficiency comparison is presented only through boxplots of relative bias and credible interval lengths, with no numerical summaries of interval lengths or, more importantly, of coverage rates. The statement that the credible intervals under the proposed approach are "immeasurably narrower" does not substitute for a quantitative check of whether the intervals attain their nominal level. Since the Dirichlet pseudo-likelihood is an approximation, the authors should report empirical coverage of the credible intervals for alpha_t, gamma_t, the Lorenz ordinates, and the Gini coefficient in the simulation, and should also compare the estimated lambda_t with the actual sampling variance of the observed shares.
minor comments (6)
  1. [Equation (6)] The random-walk state equation appears to contain a typo: "u_{tj} = u_{t,j-1} + e_{tj}" should presumably read "u_{tj} = u_{t-1,j} + e_{tj}".
  2. [Section 2.1] In the sentence defining q_k, the text says "q_k = y_k - y_{k-1} is the income share for the jth income class"; the index should be k, not j.
  3. [Section 2.1] The phrase "the cumulative distribution function and probability density function of the hypothetical income distribution in the ith area" contains leftover notation from another application; the "ith area" should be removed or clarified.
  4. [Section 3.2 and Table 2] The reference "Kl08" in the paragraph discussing the Dagum distribution should be spelled out as a proper citation, e.g., Kleiber (2008). Also, in the reference list "Ecnometrica" should be "Econometrica".
  5. [Figure 5] The caption says "posterior distributions of log lambda_t obtained from the proposed approach with RW and LNDIR"; since LNDIR is a separate, non-state-space model, the caption should clarify that the right panel is from the separate Dirichlet approach, not from the proposed RW model.
  6. [Abstract and Section 2.2] The phrase "the parameters of the Dirichlet likelihood are set to the differences between the Lorenz curve ... for the consecutive income classes" is unclear; the intended meaning is that the expected income shares are set to those differences. Please rephrase for clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the model parameters are estimated from data, and the efficiency claim is an empirical simulation result, not a fitted input renamed as a prediction.

full rationale

The central claim is that the proposed state-space Dirichlet model yields more efficient estimates of inequality measures than per-period Dirichlet fits. In the simulation, the parameters (latent processes, ψ) are estimated from the generated grouped income-share data via the likelihood and priors; the reported credible intervals for α_t, γ_t, and the Gini coefficient are posterior outputs, not quantities that were fitted or imposed. The Dirichlet pseudo-likelihood (Eq. 2) is adopted from Chotikapanich and Griffiths (2002, 2005), an external source, and the state-space dynamics (Eqs. 5-6) plus the precision specification λ_t = n_t exp(ψ) (Eq. 7) are stated modeling assumptions, not consequences of the efficiency claim. No inequality measure or credible interval length appears as an input to its own estimation. The self-citations to Kobayashi and Kakamu (2019) are used only to note the lack of guidance on λ and to illustrate prior sensitivity in the single-period Dirichlet approach; they are not invoked as a uniqueness theorem or as the justification for the model choice, so they are not load-bearing. The concern that the Dirichlet pseudo-likelihood may not be calibrated to the true sampling variance, and that no coverage check is performed, is a validity or robustness risk about interval coverage, not a circularity: the model does not define its own success criterion in terms of the fitted parameters. Under the rules requiring a specific reduction (Eq. X = Eq. Y by construction, or a fitted parameter renamed as a prediction), no circular step can be exhibited.

Assumptions & free parameters 2 free parameters · 7 assumptions · 0 invented entities

The method's contribution rests on several assumptions that are not derived from first principles: the Dirichlet pseudo-likelihood, the proportionality of precision to sample size, and the state space dynamics. No new physical or conceptual entities are introduced.

free parameters (2)
  • psi (log precision scaling) = 4.43 (posterior mean in simulation, Table 1)
    Controls Dirichlet precision through lambda_t = n_t exp(psi). It is estimated rather than derived from survey design, so it absorbs any mismatch between the pseudo-likelihood and true sampling variance.
  • c (random walk initial variance scale) = 10^5
    Set by hand in Section 3.2 for RW specifications; makes the initial state prior diffuse.
assumptions (7)
  • domain assumption The Dirichlet pseudo-likelihood (Eq. 2/4) adequately approximates the sampling distribution of grouped income shares.
    Assumed following Chotikapanich and Griffiths (2002). The paper does not validate coverage or calibration of the resulting credible intervals.
  • domain assumption Income shares q_t are conditionally independent across time given the latent parameters and precision.
    Stated in Section 2.2: 'The model (4) assumes the conditional independence of qt given theta_t and lambda_t.'
  • domain assumption The transformed Lorenz curve parameters follow AR(1) or random walk processes (Eqs. 5-6).
    The dynamics are assumed, not tested against alternatives such as regime-switching or more general ARMA structures.
  • domain assumption The Dirichlet precision is proportional to the survey sample size: lambda_t = n_t exp(psi) (Eq. 7).
    The functional form is asserted with reference to small area estimation practice; it is not derived from the survey sampling mechanism.
  • domain assumption The number of income classes K and the parametric Lorenz curve family are known and fixed over time.
    Stated in Section 2.2. Model selection is done separately via PPL, but the chosen family is treated as fixed within estimation.
  • domain assumption The prior distributions are weakly informative and do not drive the conclusions.
    Priors N(0,1), IG(3,0.1), N(0,100) are chosen from empirical ranges; the paper checks robustness for mu_j and psi, but not for all parameters.
  • standard math The Lorenz curve and Gini formulas for the parametric families are correct (e.g., L(p)=Phi(Phi^{-1}(p)-sigma), G=2Phi(sigma/sqrt(2))-1).
    Used in the likelihood and for inequality measures; standard results from the income distribution literature.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bayesian approach to Lorenz curve using time series grouped data." pith.science (2026). https://pith.science/paper/2OTOLYVK

@misc{pith2026190806772,
  author       = {Pith},
  title        = {Pith review of: Bayesian approach to Lorenz curve using time series grouped data},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2OTOLYVK}},
  note         = {Machine review of arXiv:1908.06772}
}
read the original abstract

This study is concerned with estimating the inequality measures associated with the underlying hypothetical income distribution from the times series grouped data on the Lorenz curve. We adopt the Dirichlet pseudo likelihood approach where the parameters of the Dirichlet likelihood are set to the differences between the Lorenz curve of the hypothetical income distribution for the consecutive income classes and propose a state space model which combines the transformed parameters of the Lorenz curve through a time series structure. Furthermore, the information on the sample size in each survey is introduced into the originally nuisance Dirichlet precision parameter to take into account the variability from the sampling. From the simulated data and real data on the Japanese monthly income survey, it is confirmed that the proposed model produces more efficient estimates on the inequality measures than the existing models without time series structures.

Figures

Figures reproduced from arXiv: 1908.06772 by the authors.

Figure 1
Figure 1. The boxplots of the relative bias (left) and lengths of the credible intervals (right) for [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. The boxplots of the relative bias (left) and lengths of the credible intervals (right) for the Lorenz curve [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Time series plot of the income shares of the income classes [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: The posterior means (solid lines) and 95% credible intervals (shaded areas) of the parameters and [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 5
Figure 5. Figure 5: The posterior distributions of log λt under the RW models (left) and LNDIR (right) for t = 50, 100, 150, 200 with the priors ψ, log λt ∼ N(0, 100) (solid) and N(0, 10) (dotted) 17 [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 6
Figure 6. Figure 6: The posterior means of the income shares for [PITH_FULL_IMAGE:figures/full_fig_p018_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 26 canonical work pages

  1. [1]

    Chib, S. (2001). Markov chain Monte Carlo methods: computation and inference. In Heckman, J. J. and

  2. [2]

    (ed.) (2008)

    Chotikapanich, D. (ed.) (2008). Modeling Income Distributions and Lorenz Curves , Springer: New York

  3. [3]

    and Griffiths, W.E

    Chotikapanich, D. and Griffiths, W.E. (2002). Estimating Lorenz curves using a Dirichlet distribution. Journal of Business & Economic Statistics , 20, 290–295

  4. [4]

    and Griffiths, W.E

    Chotikapanich, D. and Griffiths, W.E. (2005). Averaging Lorenz curves.Journal of Income Inequality, 3, 1–19

  5. [5]

    Dagum, C. (1977). A new model of personal income distribution: Specification and estimation. Economie Appliqu´ee, 30, 413–437

  6. [6]

    and Nagaraja, H.N

    David, H.A. and Nagaraja, H.N. (2003). Order Statistics, 3rd ed., Wiley: New York

  7. [7]

    Doornik, J. (2007). Ox: object oriented matrix programming , Timberlake Consultants Press, London

  8. [8]

    Gelfand, A. E. and Ghosh, S. K. (1998). Model choice: a minimum posterior predictive loss approach. Biometrika 85, 1–11. Griffiths, W. E. and Hajargasht, G. (2015). On GMM estimation of distributions from grouped data.Economics Letters, 126, 122–126

Show all 27 references
  1. [9]

    E., Brice, J., Rao, D.S

    Hajargasht, G., Griffiths, W. E., Brice, J., Rao, D.S. P. and Chotikapanich, D. (2012). Inference for Income Distributions Using Grouped Data. Journal of Business & Economic Statistics , 30, 563–575

  2. [10]

    and Kozumi, H

    Hasegawa, H. and Kozumi, H. (2002). Estimation of Lorenz curves: A Bayesian nonparametric approach. Journal of Econometrics, 115, 277–291

  3. [11]

    Kakamu, K. (2016). Simulation studies comparing Dagum and Singh–Maddala income distributions. Comput Econ, 48, 593–605

  4. [12]

    and Nishino, H

    Kakamu, K. and Nishino, H. (2018). Bayesian estimation of beta-type distribution parameters based on grouped data. Computational Economics, DOI:10.1007/s10614-018-9843-4

  5. [13]

    Kakwani, N. C. (1980). On a Class of Poverty Measures. Econometrica, 48, 437–446

  6. [14]

    and Podder, N

    Kakwani, N.C. and Podder, N. (1976). Efficient estimation of the Lorenz curve and associated inequality mea- sures from grouped observations. Econometrica, 44, 137–148. 22

  7. [15]

    Kleiber, C. (2008). A guide to the Dagum distributions. Ib Modeling Income Distributions and Lorenz Curves , Springer: New York

  8. [16]

    and Kotz, S

    Kleiber, C. and Kotz, S. (2003).Statistical Size Distributions in Economics and Actuarial Science . Wiley: New York

  9. [17]

    and Kakamu, K

    Kobayashi, G. and Kakamu, K. (2019). Approximate Bayesian computation for Lorenz curves from grouped data. Computational Statistics, 34, 253–279

  10. [18]

    McDonald, J.B. (1984). Some generalized functions for the size distribution of income. Econometrica, 52, 647–663

  11. [19]

    and Xu, Y .J

    McDonald, J.B. and Xu, Y .J. (1995). A generalization of the beta distribution with applications. Journal of Econometrics, 66, 133–152

  12. [20]

    and Kakamu, K

    Nishino, H. and Kakamu, K. (2011). Grouped data estimation and testing of Gini coefficients using lognormal distributions. Sankhya Series B, 73, 193–210

  13. [21]

    and Oga, T

    Nishino, H., Kakamu, K. and Oga, T. (2012). Bayesian estimation of persistent income inequality using the lognormal stochastic volatility model. Journal of Income Distribution, 21, 88–101

  14. [22]

    and Kakamu, K

    Nishino, H. and Kakamu, K. (2015). A random walk stochastic volatility model for income inequality. Japan and the World Economy, 36, 21–28

  15. [23]

    A., Lodoux, M., and Garcia, A

    Ortega, P., Fernandez, M. A., Lodoux, M., and Garcia, A. (1991). A new functional form for estimating Lorenz curves. Review of Income and Wealth, 37, 447–452

  16. [24]

    (2008) Parametric Lorenz curves: Models and applications

    Sarabia, J.M. (2008) Parametric Lorenz curves: Models and applications. in Chotikapanich, D. (ed.) Modeling Income Distributions and Lorenz Curves . Springer: New York, 167–190

  17. [25]

    H., Gaffney, J., Koo, A., and Obst, N

    Rasche, R. H., Gaffney, J., Koo, A., and Obst, N. (1980). Functional Forms for Estimating the Lorenz Curve. Econometrica, 48, 1061–1062

  18. [26]

    and Slottje, D.J

    Ryu, H.K. and Slottje, D.J. (1996) Two flexible functional form approaches for approximating the Lorenz curve. Journal of Econometrics, 72, 251–274

  19. [27]

    and Maddala, G.S

    Singh, S.K. and Maddala, G.S. (1976). A function for size distribution ofIncomes. Ecnometrica, 44, 963–970. 23

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.