Pith. sign in

REVIEW 3 major objections 5 minor 49 references

A nonparametric Bayesian approach to the rare type match problem

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read For a rare type match, the likelihood ratio reduces to a two-parameter quotient, giving about 40,000 support for the prosecution on European Y-STR data.

desk verdict A clean and honest derivation of a new closed-form LR for the rare type match, but the headline number rests on the controversial assumption that profile labels are uninformative, and the validation is in-sample. read the letter →

arxiv 1908.02954 v4 pith:HRKJMBZA submitted 2019-08-08 stat.AP

classification stat.AP MSC 62F1562P10
keywords forensicstatisticslikelihoodratioraretypematchBayesiannonparametrictwo-parameterPoissonDirichletY-STRPitmansamplingformulaEmpiricalBayes
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The rare type match problem arises when a suspect's DNA profile matches a crime stain but that profile has never been seen in the reference database, so its population frequency is unknown. The paper argues that if the ranked population proportions are given a two-parameter Poisson Dirichlet prior and the names of the DNA profiles are discarded, the likelihood ratio for such a match is approximately $(n+1+\theta_{\mathrm{MLE}})/(1-\alpha_{\mathrm{MLE}})$, where $n$ is the database size and $\alpha_{\mathrm{MLE}},\theta_{\mathrm{MLE}}$ are fitted from the database. On the European Y-STR reference data used here ($n=18{,}925$) this gives $\log_{10}\mathrm{LR}=4.59$, meaning the reduced data are about 40,000 times more likely under the prosecution hypothesis than under the defence. The point of the paper is that a problem previously handled by simulations and conservative bounds can be reduced to a simple, immediately usable formula.

What carries the argument

The load-bearing object is the two-parameter Poisson Dirichlet distribution $\mathrm{PD}(\alpha,\theta)$ over ranked population proportions, a distribution on ordered infinite lists of frequencies with power-law tails. Its Chinese restaurant representation makes prediction simple: a new customer (DNA type) occupies a new table with probability $(\theta+k\alpha)/(n+\theta)$ and joins an existing table of size $n_i$ with probability $(n_i-\alpha)/(n+\theta)$. This yields the one-step probabilities (10) and (11), and the sufficiency property from Zabell (2005) ensures that new-type and old-type probabilities depend only on counts, which is what allows the profiles to be reduced to partitions. The Pitman sampling formula (8) is the likelihood that lets the parameters $\alpha$ and $\theta$ be estimated from the database by maximum likelihood.

What would settle it

Simulate a Y-STR population under a known mutation and inheritance model that produces close relatives, then for many rare matches compare the likelihood ratio computed from the full labelled profiles with the value from equation (13); if the two diverge systematically, the reduction to partitions loses information and Assumption 2 fails.

Watch

Extended reading notes

Core claim

The central claim is equation (13): in a rare type match, the likelihood ratio is approximately $\mathrm{LR}=(n+1+\theta_{\mathrm{MLE}})/(1-\alpha_{\mathrm{MLE}})$. The derivation starts from the two-parameter Poisson Dirichlet model and the reduction of data to partitions: prosecution and defence agree on the distribution of the database enlarged by the suspect's new profile, and disagree only on whether the crime stain joins that profile's class. Under the defence the joining probability is $(1-\alpha)/(n+1+\theta)$; under the prosecution it is 1. Averaging over the posterior distribution of $(\alpha,\theta)$ and then replacing that expectation by the plug-in maximum likelihood values is justified empirically by the near-Gaussian, symmetric log-likelihood around the MLE. Applied to the 7-locus European Y-STR subset with $\alpha_{\mathrm{MLE}}=0.51$ and $\theta_{\mathrm{MLE}}=216$, the paper obtains $\log_{10}\mathrm{LR}=4.59$, about 40,000 in favour of the prosecution.

Load-bearing premise

The load-bearing premise is Assumption 2, that the names of DNA profiles contain no relevant information; if Y-STR profile similarity signals shared ancestry, then the simple quotient (13) is the likelihood ratio for the reduced partition data, not for the full genetic evidence, and the 40,000 figure could misstate the weight of the match.

Editorial extensions

If this is right

  • A forensic analyst facing a rare Y-STR match can compute the likelihood ratio directly from $n$ and the fitted $(\alpha,\theta)$, without modelling the allele structure; on the European Y-STR data this gives $\log_{10}\mathrm{LR}=4.59$.
  • In simulation using a Dutch subpopulation of size 2037, the plug-in approximation tracks the true partition-based likelihood ratio: the difference $\log_{10}\mathrm{LR}-\log_{10}\mathrm{LR}|p$ has standard deviation about 0.126 and stays between -0.146 and 0.381 in the cases studied.
  • The method transfers to any categorical forensic characteristic whose type frequencies show power-law behaviour, such as shoe marks or glass fragments, for which the rare type match problem also arises.
  • The paper does not claim the plug-in approximation is safe for small reference databases: it explicitly notes that the Gaussian shape underlying (13) is not empirically supported for samples of size 100, so such cases require exact posterior computation rather than the simple formula.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct consequence of Assumption 2 is that equation (13) is the likelihood ratio for the partition evidence only; if closeness between Y-STR profiles reflects shared ancestry, a richer summary such as counts of one-step neighbours would be expected to change the LR, and that change is a measurable test of the assumption.
  • The formula can be read as a two-parameter Good-Turing correction: the unseen matching type is assigned a positive probability derived from $(\alpha,\theta)$ rather than zero, so comparing (13) with classical Good-Turing estimates on the same database would be a natural external check that the paper does not perform.
  • A hold-out calibration exercise (fit $(\alpha,\theta)$ on one part of the European database, predict rare-match LR on another) would test whether the fitted power-law prior is predictively accurate, not just descriptive of the full database.
  • A genealogy-aware simulation with father-son mutation could quantify how much of the evidence is lost by discarding profile labels; the paper's own Diff1 only measures the gap between the plug-in LR and the partition-based ideal LR|p, not the gap to the full-data likelihood ratio.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a Bayesian nonparametric method for the rare type match problem in forensic DNA casework. The model assigns a two-parameter Poisson-Dirichlet prior to the ranked population proportions of DNA types and, under Assumption 2, discards the names of the types so that the data reduce to a random partition. Using the Chinese restaurant representation and a Bayesian network lemma, the authors derive a simple likelihood-ratio formula, Eq. (13), of the form LR ≈ (n + 1 + θ_MLE)/(1 − α_MLE), where α_MLE and θ_MLE are maximum-likelihood estimates of the Poisson-Dirichlet parameters. Applied to a European YHRD database of 18,925 Y-STR profiles, the formula gives log10 LR = 4.59. The paper also reports a simulation study on Dutch population data to compare the Bayesian LR with the 'true' LR obtained when the population proportions are known.

Significance. If its assumptions held, the paper would make a significant contribution: it connects nonparametric species-sampling theory to forensic likelihood ratios and yields an unusually simple, interpretable formula that could be applied by practitioners. The mathematical derivation of the LR identity from the Pitman sampling formula and the Bayesian network lemma is clean and correct within the stated model, and the paper is honest about several limitations, including the acknowledged controversy over discarding the genetic structure of Y-STR profiles. However, the practical significance for real Y-STR casework is not established because the key reduction to partitions and the empirical-Bayes plug-in are not validated against data or models that use the full genetic information.

major comments (3)
  1. [Section 3.2, Assumption 2; Section 2.2] Assumption 2 is load-bearing, and the paper provides no sensitivity analysis for it. The reduction of Y-STR profiles to equality classes removes the allele-distance information on which relatedness inference relies; the paper itself cites Andersen and Balding (2017) as finding that 95% of matching Y-STR profiles are separated by only 50–100 meioses and that relatedness is a very influential factor. Equation (13) is therefore the likelihood ratio for the reduced partition π[n+2], not for the full genetic evidence (E,B), and the headline log10 LR = 4.59 would not be the likelihood ratio for the observed Y-STR evidence if Assumption 2 fails. The statement 'we believe in the accuracy of our method' does not quantify the gap. The manuscript should either provide a formal or empirical comparison with methods that use the full genetic structure (e.g., Discrete Laplace or models based on Andersen and Balding's approach) or explicitly restrict the claim to the reduced data and warn practitioners against interpreting Eq. (13) as the likelihood ratio for full Y-STR evidence.
  2. [Section 5.4, Tables 1 and 2] The simulation study is in-sample and cannot validate the method for real applications. The Dutch population is used both to estimate α_MLE and θ_MLE and to define the 'true' population proportions p for generating rare type match cases; the resulting Diff1 and Diff2 therefore measure error only within the assumed Poisson-Dirichlet model class on the same data that produced the parameter estimates. This does not assess the discrepancy between the partition-based likelihood ratio and the likelihood ratio for the full genetic evidence. Moreover, the text explicitly states that the Gaussian shape justifying approximation (13) is 'not empirically supported for small databases of size n = 100', even though the simulation study uses samples of size n = 100. The empirical support for Eq. (13) as a general approximation is thus weaker than the paper's main text suggests.
  3. [Section 5.2, Eq. (13) and Figure 6] The empirical-Bayes approximation E[(1−A)/(n+1+Θ) | π[n+1]] ≈ (1−α_MLE)/(n+1+θ_MLE) is justified only by the visual Gaussian symmetry of one observed log-likelihood (Figure 6). The text itself says 'one could safely make this approximation if one believed that this symmetry would also be true in the real data situation at hand', which is a conditional belief statement rather than a demonstrated property. No sensitivity analysis is given for the choice of hyperprior, and no measure is provided for how far the posterior mean can be from the mode in realistic settings. Since Eq. (13) is the central practical result, the paper needs either a more principled justification of the plug-in approximation, a sensitivity analysis over plausible hyperpriors, or an explicit statement that the approximation is heuristic and may be unreliable outside the specific database analyzed.
minor comments (5)
  1. [Section 4] The definition of Φ in the sentence preceding Eq. (13) is garbled: the printed expression 'Φ = n 1−A n + 1 + Θ' lacks parentheses and appears to include an unexplained factor n. It should be written unambiguously, e.g., Φ = (1−A)/(n+1+Θ), with the relation to LR = 1/E(Φ) made explicit.
  2. [Section 5.3 and 5.4] The name 'Metropolis Hashting' appears in two places and should be 'Metropolis-Hastings'.
  3. [Table 2] The caption of Table 2 refers to 'Diff1, Diff2, and Diff3', but only Diff1 and Diff2 are defined in Section 5.4; this is a typographical error.
  4. [Section 5.3] The notation πn+1 is used both for a partition of the integer n+1 and for the partition of the enlarged database, which is confusing given the earlier distinction between partitions of [n] and partitions of n. Please use distinct notation.
  5. [Section 3.3] The phrase 'sufficientness property' should be 'sufficiency property'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the likelihood-ratio formula follows from the stated Pitman/Chinese-restaurant model, and the MLE plug-in is presented as an explicit approximation rather than as a prediction forced by the fit.

full rationale

The central derivation is self-contained. Equations (10) and (11) are the Chinese-restaurant transition probabilities for the two-parameter Poisson-Dirichlet prior, and the likelihood ratio formula (13) is obtained by applying Corollary 3.1 and then replacing a posterior expectation by its MLE plug-in. This is an explicit empirical-Bayes approximation, not a quantity that equals the fitted parameters by construction: the quoted result is a ratio involving n, alpha_MLE, and theta_MLE, and it is not merely a restatement of the data or of the fitted values. The paper is also transparent that Assumption 2 reduces the evidence to a partition, so the computed LR is the LR for D = pi[n+2], and any gap between that and the full Y-STR evidence is a stated modeling limitation rather than a circular derivation. Self-citations to Cereda (2017a,b) and Anevski et al. (2017) are contextual or algorithmic and are not load-bearing; the Pitman sampling formula and the two-parameter Poisson-Dirichlet distribution are external mathematical results. The in-sample fitting of alpha and theta to the YHRD database and subsequent plug-in is empirical Bayes, which may weaken extrapolation, but it does not make the derived likelihood ratio an input of the model. No circular step can be exhibited by the paper's own equations.

Assumptions & free parameters 3 free parameters · 8 assumptions · 0 invented entities

The central result is not derived from first principles: it depends on the infinite-type idealization, the no-name reduction, the Poisson Dirichlet prior, the random-sample database assumption, and the MLE plug-in approximation. Only alpha and theta are explicitly fitted scalar parameters; the truncation length m is an unspecified computational choice. The validation benchmark is in-sample because the Dutch data are used both to fit the model and to define the true population frequencies.

free parameters (3)
  • alpha (discount parameter of PD prior) = 0.51 (European YHRD), 0.62 (Dutch)
    Estimated by maximum likelihood from Y-STR database via the Pitman sampling formula; the LR formula depends on 1-alpha in the denominator.
  • theta (concentration parameter of PD prior) = 216 (European YHRD), 22 (Dutch)
    Estimated by maximum likelihood; enters the LR formula through n+1+theta in the numerator.
  • finite truncation length m for population vector p
    The Metropolis-Hastings computation of LR|p in Section 5.3 assumes only the first m entries of p are positive; m is not specified and the results depend on it through constraints (15) and (16).
assumptions (8)
  • domain assumption There are infinitely many different DNA types in Nature (Assumption 1, Section 3.2).
    A convenient idealization that allows Bayesian nonparametric species sampling; the authors state the haplotype space is so large it can be treated as infinite.
  • ad hoc to paper The names of DNA types carry no relevant information, so data reduce to a partition of equality classes (Assumption 2, Section 3.2).
    This is the key modeling choice and is explicitly controversial because Y-STR profiles are inherited and similar profiles share recent ancestry.
  • domain assumption The reference database is a random sample from the population of potential perpetrators (Section 3.4, eq. (3)).
    Forensic databases are convenience collections; the authors acknowledge this limitation in Section 2.2.
  • domain assumption The ranked population proportions follow a two-parameter Poisson Dirichlet distribution (Section 3.3).
    Chosen for power-law behavior and sufficiency; the fit is checked only in-sample on the data used to estimate alpha and theta.
  • ad hoc to paper The posterior expectation of Phi = n(1-A)/(n+1+Theta) can be replaced by its MLE plug-in (Section 5.2, eq. (13)).
    Justified by visual Gaussian symmetry of one log-likelihood; the authors state it is not empirically supported for n=100.
  • standard math The Pitman sampling formula (8) and the sufficiency property of the PD model (Section 3.5).
    Established results cited to Pitman (1992, 1995), used to derive the distribution of the reduced data.
  • standard math Lemma 3.1 and Corollary 3.1 on likelihood factorization in Bayesian networks (Section 3.7).
    Elementary probability manipulation, proved in the paper, used to express the LR as a ratio of posterior expectations.
  • domain assumption The Metropolis-Hastings chain for the latent map chi converges to the target distribution (Section 5.3).
    No convergence diagnostics are reported; burn-in and thinning are fixed without verification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A nonparametric Bayesian approach to the rare type match problem." pith.science (2026). https://pith.science/paper/HRKJMBZA

@misc{pith2026190802954,
  author       = {Pith},
  title        = {Pith review of: A nonparametric Bayesian approach to the rare type match problem},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HRKJMBZA}},
  note         = {Machine review of arXiv:1908.02954}
}
read the original abstract

The "rare type match problem" is the situation in which the suspect's DNA profile, matching the DNA profile of the crime stain, is not in the database of reference. The evaluation of this match in the light of the two competing hypotheses (the crime stain has been left by the suspect or by another person) is based on the calculation of the likelihood ratio and depends on the population proportions of the DNA profiles, that are unknown. We propose a Bayesian nonparametric method that uses a two-parameter Poisson Dirichlet distribution as a prior over the ranked population proportions, and discards the information about the names of the different DNA profiles. This fits very well the data coming from European Y-STR DNA profiles, and the calculation of the likelihood ratio becomes quite simple thanks to a justified Empirical Bayes approach.

Figures

Figures reproduced from arXiv: 1908.02954 by the authors.

Figure 1
Figure 1. Bayesian network showing the conditional dependencies of the relevant random variables in our model. H is a dichotomous random variable that represents the hypotheses of interest and can take values h ∈ {hp, hd}, according to the prosecution or the defense, respectively. A uniform prior on the hypotheses is chosen for mathematical convenience since it will not affect the likelihood ratio (the variable H being in the… view at source ↗
Figure 2
Figure 2. Simplified version of the Bayesian network in [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Conditional dependencies of the random variables of Lemma 3.1 Lemma 3.1. Given four random variables Z, H, X and Y , whose conditional dependen￾cies are represented by the Bayesian network of [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Log scale ranked frequencies from the database (thick line) are compared to the relative frequencies of samples of size n = 180925 obtained from several realizations of PD(αMLE, θMLE) (thin lines). Asymptotic power-law behavior is also displayed (dotted line). We can s…
Figure 5
Figure 5. Figure 5: Log scale ranked frequencies from the two-parameter Poisson Dirichlet distribution with α = 0.51, θ = 216 approximated through a Chinese restaurant seating plan, each with its ownnumber of costumers, corresponding to the different thickness of the lines. θ φ −100 −100 …
Figure 6
Figure 6. Figure 6: Relative log-likelihood for φ = n 1−α n+1+θ and θ compared to a Gaussian distribution displayed with 95% and 99% confidence intervals 16 [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Bayesian network for the case in which p is known. The likelihood ratio in this case can be obtained using again Corollary 3.1, where now X1, ..., Xn+1 play the role of Z. Indeed, now that p is known, the unobservable part of the model are the ranks of the types in the…
Figure 8
Figure 8. Figure 8: Log scale ranked frequencies from the Dutch database (thick line) compared to the rela￾tive frequencies of samples of size n = 2037 obtained from several realizations of PD(αMLE, θMLE) (thin lines). Asymptotic power-law behavior is also displayed (dotted line). The dis…
Figure 9
Figure 9. Figure 9: (a) comparison between the distribution of log10 LR, log10 LR|p, and log10 LRf . (b) the error log10 LR − log10 LR|p and log10 LR − log10 LRf . In [PITH_FULL_IMAGE:figures/full_fig_p021_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

49 extracted references · 49 canonical work pages

  1. [1]

    Interpreting Evidence: Evaluating Forensic Science in the Courtroom ; John Wiley & Sons : Chichester, 1995

    Robertson, B.; Vignaux, G.A. Interpreting Evidence: Evaluating Forensic Science in the Courtroom ; John Wiley & Sons : Chichester, 1995

  2. [2]

    Interpreting DNA evidence: Statistical Genetics for Forensic Scientists ; Sinauer Associates: Sunderland, 1998

    Evett, I.; Weir, B. Interpreting DNA evidence: Statistical Genetics for Forensic Scientists ; Sinauer Associates: Sunderland, 1998

  3. [3]

    Statistics and the Evaluation of Evidence for Forensics Scientists ; John Wiley & Sons: Chichester, 2004

    Aitken, C.; Taroni, F. Statistics and the Evaluation of Evidence for Forensics Scientists ; John Wiley & Sons: Chichester, 2004

  4. [4]

    Weight-of-evidence for Forensic DNA Profiles ; John Wiley & Sons : Chichester, 2005

    Balding, D. Weight-of-evidence for Forensic DNA Profiles ; John Wiley & Sons : Chichester, 2005

  5. [5]

    Bayesian Networks and Probabilistic Inference in Forensic Science ; John Wiley & Sons: Chichester, 2006

    Taroni, F.; Aitken, C.; Garbolino, P.; Biedermann, A. Bayesian Networks and Probabilistic Inference in Forensic Science ; John Wiley & Sons: Chichester, 2006

  6. [6]

    Bayesian approach to LR in case of rare type match

    Cereda, G. Bayesian approach to LR in case of rare type match. Statistica Neerlandica 2017 , 71 , 141--164

  7. [7]

    Impact of model choice on LR assessment in case of rare haplotype match (frequentist approach)

    Cereda, G. Impact of model choice on LR assessment in case of rare haplotype match (frequentist approach). Scandinavian Journal of Statistics 2017 , 44 , 230--248

  8. [8]

    Fundamental problem of forensic mathematics -- The evidential value of a rare haplotype

    Brenner, C.H. Fundamental problem of forensic mathematics -- The evidential value of a rare haplotype . Forensic Science International: Genetics 2010 , 4 , 281--291

Show all 49 references
  1. [9]

    An investigation of the potential of DIP-STR markers for DNA mixture analyses

    Cereda, G.; Biedermann, A.; Hall, D.; Taroni, F. An investigation of the potential of DIP-STR markers for DNA mixture analyses. Forensic Science International: Genetics 2014 , 11 , 229--240

  2. [10]

    Essai Philosophique sur les Probabilites ; Mme

    Laplace, P. Essai Philosophique sur les Probabilites ; Mme. Ve Courcier: Paris, 1814

  3. [11]

    The performance of universal coding

    Krichevsky, R.; Trofimov, V. The performance of universal coding. IEEE Transactions on Information Theory 1981 , 27 , 199--207

  4. [12]

    What's wrong with adding one? Corpus-Based Research into Language

    Gale, W.A.; Church, K.W. What's wrong with adding one? Corpus-Based Research into Language. Rodolpi , 1994

  5. [13]

    The population frequencies of species and the estimation of population parameters

    Good, I. The population frequencies of species and the estimation of population parameters . Biometrika 1953 , 40 , 237--264

  6. [14]

    On Modeling Profiles Instead of Values

    Orlitsky, A.; Santhanam, N.P.; Viswanathan, K.; Zhang, J. On Modeling Profiles Instead of Values . Uncertainty in Artificial Intelligence , 2004 , pp. 426--435

  7. [15]

    E stimating a probability mass function with unknown labels

    Anevski, D.; Gill, R.D.; Zohren, S. E stimating a probability mass function with unknown labels. Annals of Statistics 2017 , in Press

  8. [16]

    Nonparametric B ayes estimation of the probability of discovering a new species

    Tiwari, R.C.; Tripathi, R.C. Nonparametric B ayes estimation of the probability of discovering a new species. Communications in statistics: Theory and methods 1989 , A18 , 877--895

  9. [17]

    Bayesian nonparametric estimation of the probability of discovering new species

    Lijoi, A.; Mena, R.H.; Pruenster, I. Bayesian nonparametric estimation of the probability of discovering new species . Biometrika 2007 , 94 , 769--786

  10. [18]

    Are G ibbs-Type Priors the Most Natural Generalization of the D irichlet P rocess

    De Blasi, P.; Favaro, S.; Lijoi, A.; Mena, R.H.; Pruenster, I.; Ruggiero, M. Are G ibbs-Type Priors the Most Natural Generalization of the D irichlet P rocess. IEEE Transactions on Patterns Analysis ans Machine Intelligence 2015 , 37 , 212--229

  11. [19]

    Bayesian nonparametric inference for species variety with a two parameter Poisson-Dirichlet process prior

    Favaro, S.; Lijoi, A.; Mena, R.H.; Pruenster, I. Bayesian nonparametric inference for species variety with a two parameter Poisson-Dirichlet process prior. Journal of the Royal Statistical Society: Series B (Methodological) 2009 , 71 , 993--1008

  12. [20]

    B ayesian nonparametric inference for discovery probabilities: credible intervals and large sample asymptotics

    Arbel, J.; Favaro, S.; Nipoti, B.; Teh, Y.W. B ayesian nonparametric inference for discovery probabilities: credible intervals and large sample asymptotics. Statistica Sinica 2017 , 27 , 839--859

  13. [21]

    R ediscovery of G ood- T uring estimators via B ayesian nonparametrics

    Favaro, S.; Nipoti, B.; Teh, Y.W. R ediscovery of G ood- T uring estimators via B ayesian nonparametrics. Biometrics 2016 , 72 , 136--145

  14. [22]

    Hierarchical D irichlet processes

    Teh, Y.W.; Jordan, M.I.; Beal, M.J.; Blei, D.M. Hierarchical D irichlet processes. Journal of the American Statistical Association 2006 , 101 , 1566--1581

  15. [23]

    Power laws, Pareto distributions and Zipf's law

    Newman, M. Power laws, Pareto distributions and Zipf's law . Contemporary P hysics 2005 , 46 , 323--351

  16. [24]

    No shortcut solutions to the problem of Y-STR match probability calculation

    Caliebe, A.; Jochens, A.; Willuweit, S.; Roewer, L.; Krawczak, M. No shortcut solutions to the problem of Y-STR match probability calculation. Forensic Science International: Genetics 2015 , 15 , 69--75

  17. [25]

    Modelling the dependence structure of Y-STR haplotypes using graphical models

    Andersen, M.M.; Curran, J.M.; de Zoete, J.; D., T.; Buckleton, J. Modelling the dependence structure of Y-STR haplotypes using graphical models. Forensic Science International: Genetics 2018 , 37 , 29--36

  18. [26]

    Estimation of Y haplotype frequencies with lower order dependencies

    Andersen, M.M.; Caliebe, A.; K, K.; M., K.; N., V.; Curran, J.M. Estimation of Y haplotype frequencies with lower order dependencies. Forensic Science International: Genetics 2019

  19. [27]

    DNA profile match probability calculation: how to allow for population stratification, relatedness, database selection and single bands

    Balding, D.J.; Nichols, R.A. DNA profile match probability calculation: how to allow for population stratification, relatedness, database selection and single bands . Forensic Science International 1994 , 64 , 125--140

  20. [28]

    How convincing is a matching Y-chromosome profile? PLOS Genetics 2017 , 13 , 1--16

    Andersen, M.M.; Balding, D.J. How convincing is a matching Y-chromosome profile? PLOS Genetics 2017 , 13 , 1--16

  21. [29]

    Y-profile evidence: Close paternal relatives and mixtures

    Andersen, M.M.; Balding, D.J. Y-profile evidence: Close paternal relatives and mixtures. Forensic Science International: Genetics 2019 , 38 , 48 -- 53

  22. [30]

    Estimating Haplotype Frequency and Coverage of Databases

    Egeland, T.; Salas, A. Estimating Haplotype Frequency and Coverage of Databases. PLoS ONE 2008 , 3 , e3988--e3988

  23. [31]

    Y chromosome STR typing in crime casework

    Roewer, L. Y chromosome STR typing in crime casework. Forensic Science, Medicine, and Pathology 2009 , 5 , 77--84

  24. [32]

    The interpretation of lineage markers in forensic DNA testing

    Buckleton, J.; Krawczak, M.; Weir, B. The interpretation of lineage markers in forensic DNA testing. Forensic Science International: Genetics 2011 , 5 , 78--83

  25. [33]

    Y-STR Frequency Surveying Method: A critical reappraisal

    Willuweit, S.; Caliebe, A.; Andersen, M.M.; Roewer, L. Y-STR Frequency Surveying Method: A critical reappraisal. Forensic Science International: Genetics 2011 , 5 , 84--90

  26. [34]

    Inferences from DNA data: population histories, evolutionary processes and forensic match probabilities

    Wilson, I.J.; Weale, M.E.; Balding, D.J. Inferences from DNA data: population histories, evolutionary processes and forensic match probabilities. Journal of the Royal Statistical Society: Series A (Statistics in Society) 2003 , 166 , 155--188

  27. [35]

    The discrete Laplace exponential family and estimation of Y-STR haplotype frequencies

    Andersen, M.M.; Eriksen, P.S.; Morling, N. The discrete Laplace exponential family and estimation of Y-STR haplotype frequencies . Journal of Theoretical Biology 2013 , 329 , 39--51

  28. [36]

    Y chromosome haplotype reference database (YHRD): Update

    Willuweit, S.; Roewer, L. Y chromosome haplotype reference database (YHRD): Update . Forensic Science International: Genetics 2007 , 1 , 83--87

  29. [37]

    a ler, G.; Wiest, T.; Berger, B.; Niederst \

    Purps, J.; Siegert, S.; Willuweit, S.; Nagy, M.; Alves, C.; Salazar, R.; Angustia, S.M.T.; Santos, L.H.; Anslinger, K.; Bayer, B.; Ayub, Q.; Wei, W.; Xue, Y.; Tyler-Smith, C.; Bafalluy, M.B.; Mart \' nez-Jarreta, B.; Egyed, B.; Balitzki, B.; Tschumi, S.; Ballard, D.; Court, D....

  30. [38]

    THE NUMBER OF ALLELES THAT CAN BE MAINTAINED IN A FINITE POPULATION

    Kimura, M. THE NUMBER OF ALLELES THAT CAN BE MAINTAINED IN A FINITE POPULATION. G enetics 1964 , 49 , 725--738

  31. [39]

    Bayesian Nonparametrics ; Cambridge University Press: Cambridge, 2010

    Hjort, N.; Holmes, C.; M \"u ller, P.; Walker, S. Bayesian Nonparametrics ; Cambridge University Press: Cambridge, 2010

  32. [40]

    Fundamentals of Nonparametric Bayesian Inference ; Cambridge University Press, 2017

    Ghosal, S.; Van der Vaart, A. Fundamentals of Nonparametric Bayesian Inference ; Cambridge University Press, 2017

  33. [41]

    The two-parameter Poisson-Dirichlet distribution derived from a stable subordinator

    Pitman, J.; Yor, M. The two-parameter Poisson-Dirichlet distribution derived from a stable subordinator . Annals of Probability 1997 , 25 , 855--900

  34. [42]

    The Poisson-Dirichlet Distribution and Related Topics: Models and Asymptotic Behaviors ; Springer: Berlin, 2010

    Feng, S. The Poisson-Dirichlet Distribution and Related Topics: Models and Asymptotic Behaviors ; Springer: Berlin, 2010

  35. [43]

    A B ayesian view of the P oisson- D irichlet process

    Buntine, W.; Hutter, M. A B ayesian view of the P oisson- D irichlet process. arXiv:1007.0296, 2012

  36. [44]

    Combinatorial Stochastic Processes ; \'E cole D' \'E t \'e de Probabilit \'e s de Saint-Flour XXXII - 2002, Springer: Berlin, 2006

    Pitman, J. Combinatorial Stochastic Processes ; \'E cole D' \'E t \'e de Probabilit \'e s de Saint-Flour XXXII - 2002, Springer: Berlin, 2006

  37. [45]

    243--274

    Zabell, S.L., The Continuum of Inductive Methods Revisited; Cambridge Studies in Probability, Induction and Decision Theory, Cambridge University Press, 2005; pp. 243--274

  38. [46]

    The two-parameter generalization of E wens' random partition structure

    Pitman, J. The two-parameter generalization of E wens' random partition structure. Technical report 345, Department of Statistics U.C. Berkeley CA, 1992

  39. [47]

    Exchangeable and partially exchangeable random partitions

    Pitman, J. Exchangeable and partially exchangeable random partitions. Probability Theory and Related Fields 1995 , 102 , 145--158

  40. [48]

    Exchangeability and Related Topics ; Vol

    Aldous, D.J. Exchangeability and Related Topics ; Vol. 1117, \'E cole D' \'E t \'e de Probabilit \'e s de Saint-Flour , Springer-Verlag: New York, 1985

  41. [49]

    Information-Theoretical Assessment of the Performance of Likelihood Ratio Computation Methods

    Ramos, D.; Gonzales-Rodriguez, J.; Zadora, G.; Aitken, C. Information-Theoretical Assessment of the Performance of Likelihood Ratio Computation Methods. Journal of Forensic Sciences 2013 , 58 , 1503--1517

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.