REVIEW 3 major objections 6 minor 15 references
A standard two-latent-class model for diagnostic tests can identify the wrong latent class — the majority measurand — when most tests do not measure the target condition.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Standard two-latent-class diagnostic models can mislabel latent classes when tests measure different measurands; modeling measurands explicitly via DAGs yields unbiased estimates and reinterprets two published analyses.
T0 review reviewed 2026-08-02 challenge →
load-bearing objection A genuinely useful framing for latent class analysis—DAG-driven measurands—with a solid simulation, but the applied examples and identifiability reporting need cleanup before the estimates can be trusted. the 3 major comments →
Improving interpretation of latent class models for diagnostic tests by recognizing their measurands via directed acyclic graphs (DAGs)
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
On the paper's own terms: a conventional two-latent-class model does not always identify the target condition. The latent classes it recovers are determined by the combination of observed tests; if the majority of tests measure a measurand distinct from the target condition, the classes are effectively M+ and M- (measurand positive/negative), not D+ and D- (target condition positive/negative). This happens because the tests' shared measurand induces conditional dependence that the 2LC model can only absorb by relabeling the latent variable. When the model is expanded so that latent classes are all combinations of the target condition and the measurands (as dictated by a DAG), the likelihood
What carries the argument
The central device is the distinction between a test's measurand (the quantity it is designed to measure, such as IgM antibodies or M. tuberculosis bacteria) and the target condition of interest (the disease one wants to diagnose). The paper represents both, plus observed tests and covariates, in a directed acyclic graph (DAG). The DAG reveals shared and nonshared measurands, which implies conditional dependence among tests with the same measurand and determines how many latent classes exist (combinations of target condition and measurands). The expanded likelihood then factors each observed test's result through its measurand: P(T_j | M_p) times P(M_p | D), so test accuracy and measurand ac
Load-bearing premise
The whole approach rests on the DAG being right: the chosen measurands, the dropped nodes, and the 'impossible' latent class combinations must genuinely reflect how the tests work; if any of these structural assumptions is wrong, the expanded model's class labels and accuracy estimates are just as unreliable as the 2LC model's.
What would settle it
If, in a dataset where the majority of tests are known (by design or external gold standard) to measure something other than the target condition, the 2LC model's estimated latent class prevalence still equals the target-condition prevalence (and test sensitivities match the target-condition values) rather than the measurand's prevalence, the paper's central claim would be false. Concretely: simulate with five tests where four detect a proxy M and one detects the disease D, with D and M only weakly correlated; the 2LC model should return M prevalence as its class prevalence. Any simulation tha
If this is right
- In any 2LC analysis, the latent class label is not guaranteed to be the target condition; it is an emergent property of which measurand dominates the test set.
- When a majority of tests share a non-target measurand, prevalence and test-accuracy estimates from the 2LC model describe the measurand, not the disease.
- Expanding the model to include measurand-based latent classes restores interpretable labels and separates test accuracy with respect to the measurand from accuracy with respect to the target condition.
- The expanded models may fail local identifiability, but informative expert priors on one or two specificities can make them estimable.
- Re-analysis of the leptospirosis data shows the previously 'low sensitivity' MAT estimate was an artifact: the 2LC classes were IgM+, not Leptospirosis+.
Where Pith is reading between the lines
- The same mislabeling hazard applies to any finite mixture model applied to measurements that tap different constructs; the paper's DAG-plus-measurand recipe is a general diagnostic for mixture label validity, not just for medical tests.
- A testable extension: when the majority of tests share a proxy measurand, the 2LC model's estimated prevalence should track the measurand's prevalence even if the target-condition prevalence is known from external sources; this can be checked in datasets with verified disease status.
- The structural constraints (which latent class combinations are 'impossible') are assumptions; when those are wrong, the expanded model can be as mislabeled as the 2LC model. The paper does not sensitivity-analyze those constraints, but a robustness check that relaxes one constraint at a time would quantify the risk.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that conventional two-latent-class (2LC) latent class models for diagnostic test accuracy can misidentify the latent variable when tests do not all measure the same target condition. It introduces directed acyclic graphs (DAGs) to explicitly separate each test's measurand from the target condition, and proposes expanded latent class models with classes defined by combinations of target condition and distinct measurands. The authors derive the mixture likelihood in Eq. (1), discuss identifiability in terms of degrees of freedom and the Jacobian rank, and use simulations to show that when the majority of tests measure a non-target measurand, the 2LC model tends to label that measurand rather than the target condition. They re-analyze two published datasets (pediatric pulmonary tuberculosis and leptospirosis) and report that the 2LC model in the TB example overestimates active TB prevalence (0.27 vs 0.22 under the proposed 4LC model) and that in the leptospirosis example the 2LC latent class appears to correspond to IgM rather than Leptospirosis infection.
Significance. The methodological message is important and clearly demonstrated by the simulation: ignoring the measurand structure of diagnostic tests can lead to incorrectly labeled latent classes and biased accuracy/prevalence estimates. The DAG-based framework is a useful tool for making these assumptions transparent and for structuring the latent class model. The applied examples are engaging and the face validity checks (e.g., treatment patterns in the TB data) lend some support. However, the applied conclusions depend critically on structural-zero assumptions that are not clinically justified and are not subjected to sensitivity analysis, and the models are only locally identifiable with the help of informative priors. These limitations materially weaken the paper's central applied claims. If the structural assumptions were relaxed, the numerical findings could change substantially; the paper does not provide evidence of robustness.
major comments (3)
- [§2.1, Table 1, Figure 1b] The reduced DAG drops 'Other respiratory disease' and relabels the x-ray measurand as 'Intrathoracic abnormalities due to TB'. Table 1 then declares the latent class (Active TB=0, TB Infection=0, Intrathoracic abnormalities=1) as not possible (NP). In a pediatric cohort presenting with respiratory symptoms, this is a strong and clinically questionable restriction: chest x-ray abnormalities frequently occur without active TB (e.g., pneumonia). Because no observed test measures 'other respiratory disease', children with such abnormalities are forced into class 4 (0,0,0), directly inflating the estimated class-4 prevalence (0.52, 95% CrI 0.34–0.64) and biasing x-ray sensitivity with respect to its measurand (reported 0.95, 0.82–1). The manuscript provides no sensitivity analysis to this structural zero. The headline prevalence difference (0.22 vs 0.27) is therefore not robust to a clinicall
- [§2.2, Table 3] The leptospirosis model suffers an analogous problem: the DAG in Figure 2b drops 'Other infection' because it is 'not estimable', and Table 3 restricts the model so that IgM/IgG arise only from Leptospirosis or from the target condition. The informative priors that the specificity of IgM exceeds 90% and that the specificity of IgG is 100% effectively impose the absence of other causes of these antibodies. In a hospital setting in Tanzania, this assumption is not clinically defensible. The resulting 6LC estimates—and the conclusion that the 2LC latent class corresponds to IgM rather than Leptospirosis—are conditional on this untestable exclusion. A sensitivity analysis that allows nonzero prevalence of other infections, even with weak prior information, is necessary before these applied conclusions can be considered reliable.
- [§5.1, Identifiability] The authors report that the TB model's Jacobian rank is 12 with 14 unknown parameters, requiring at least 2 informative priors (Section 5.1). The structural-zero restrictions in Table 1 are additional, non-estimated constraints. This combination means the posterior estimates—including the active TB prevalence of 0.22 and the measurand-specific accuracies—depend on both the choice of which parameters receive informative priors and the validity of the structural restrictions. The manuscript does not report any sensitivity analysis varying either the prior specifications or the structural-zero set. Given that the local identifiability condition is not met, the applied estimates may be largely driven by prior and structural assumptions rather than by the data, and the paper should demonstrate stability of its conclusions under plausible perturbations.
minor comments (6)
- [§4.1] The simulation uses the expected dataset rather than simulated random datasets, which is fine for illustrating mean behavior, but the sample size N is never stated. Please specify the sample size used to scale the expected frequencies, and consider reporting Monte Carlo standard errors or additional randomly simulated replicates to assess sampling variability.
- [Table 2] Table 2 is difficult to read: the column headers do not clearly indicate which columns are sensitivities versus specificities and with respect to which target/measurand. The reader has to reconstruct the structure from the prose. Please reorganize the table so that each test's accuracy with respect to each latent variable is presented unambiguously.
- [§2.1] There is a typo: 'thus illustrating they are are conditionally dependent' should read 'they are conditionally dependent'.
- [§5.2] In the leptospirosis results, the credible intervals for the prevalence of Leptospirosis under the 6LC model are extremely wide (0.28, 95% CrI 0.02–0.93; Table 4). Even the point estimates of sensitivity for several tests with respect to Leptospirosis have intervals spanning a large range. The conclusion that the 2LC model identifies IgM rather than Leptospirosis is partly based on point estimates with substantial uncertainty; the overlap of intervals should be acknowledged or handled more carefully in the interpretation.
- [References] The Wasserman reference is incomplete ('2008-2010', 'Chapter 18') and should be updated to a proper citation. Also, 'Mycobaterium' is a typo in Section 2.1.
- [§5.1] The text says '13 unknown parameters' and then later 'total number of unknown parameters was 19' after accounting for prevalence covariates and a continuous measurand. The derivation of the 19 total parameters (beyond the 13) is not shown and should be clarified, since the reader is referred to the supplementary material for details that are not available in the main text.
Circularity Check
No significant circularity; the measurand-DAG framework is an explicit modeling assumption, and the 2LC-bias demonstrations are simulation-based rather than definitional.
full rationale
The paper's central derivation is a probability decomposition from stated latent-variable assumptions (Section 3.1, Equation 1); no target quantity is defined in terms of the conclusion. The simulation study generates data from an explicit DAG-based model with fixed prevalence and accuracy values, then fits a deliberately misspecified 2LC model and compares estimates to the true values; this is a standard misspecification illustration, not a fitted parameter relabeled as a prediction. The applied examples impose structural constraints (Tables 1 and 3) and explicit informative priors, including priors taken from earlier publications by the same group (Schumacher et al. 2016). These are transparent modeling assumptions and prior inputs, not conclusions derived from the data; the paper even discloses local non-identifiability (rank 12 vs 14 parameters) and states that informative priors are needed. The self-citations are used for prior values and modeling conventions, but the measurand claim itself is supported by the likelihood derivation and simulations, not by authority. The acknowledged limitations (dropping 'Other respiratory disease', 'Other infection', dichotomous tests, missing data) concern model validity and robustness, not circularity. No step was found where a prediction is equivalent by construction to its inputs.
Axiom & Free-Parameter Ledger
axioms (6)
- domain assumption Tests are conditionally independent given their measurand and the target condition (Eq. 1).
- domain assumption DAG structure for TB: Culture, Xpert, Smear share the measurand 'Detectable M.tb bacteria'; X-ray measures intrathoracic abnormalities due to TB; TST measures immune response to M.tb.
- domain assumption Impossible latent class restrictions in TB: active TB without latent TB infection is impossible; intrathoracic abnormalities due to TB without active TB are impossible.
- ad hoc to paper Reduced DAG drops 'DNA of M.tb' and 'Other respiratory disease'/'Other infection' nodes because they are not statistically identifiable with available data.
- domain assumption DAG structure for leptospirosis: three rapid tests measure IgM; MAT measures both IgG and IgM; IgM may be caused by other infections.
- domain assumption Expert priors: specificity of IgM > 90%; specificity of IgG = 100% (and culture/Xpert bounds in TB).
Cite this review
Pith. "Pith review of Improving interpretation of latent class models for diagnostic tests by recognizing their measurands via directed acyclic graphs (DAGs)." pith.science (2026). https://pith.science/paper/EZG4WLGU
@misc{pith2026260714473,
author = {Pith},
title = {Pith review of: Improving interpretation of latent class models for diagnostic tests by recognizing their measurands via directed acyclic graphs (DAGs)},
year = {2026},
howpublished = {\url{https://pith.science/paper/EZG4WLGU}},
note = {Machine review of arXiv:2607.14473}
}
read the original abstract
Summary: In the absence of a perfect diagnostic test for a target condition, multiple imperfect tests may be used to arrive at a clinical diagnosis. Latent class analysis can be used to model such data with the objective of estimating test accuracy and target condition prevalence. Such models typically assume two latent classes - target condition positive and target condition negative. However, as we will illustrate in this manuscript, this would be an oversimplification if the different tests do not share the target condition as their measurand. We show how a Directed Acyclic Graph (DAG) can be used to illustrate the relationships between the relevant variables - the observed imperfect test results, their latent measurands, the latent target condition of interest and observed covariates - revealing any conditional dependence relations. The DAG helps determine the number of latent classes, underlying the observed data, and their labels. We show how the likelihood function changes due to incorporating the measurand of each test. We study the impact on identifiability of the model. Using simulation studies we show how ignoring the measurand of an imperfect test, when it is distinct from the target condition, can lead to biased estimates of test accuracy and prevalence. We illustrate the value of the proposed approach by re-analyzing two datasets used in previously published latent class analyses of tests for pediatric tuberculosis and leptospirosis.
Reference graph
Works this paper leans on
-
[1]
Albert, P. S. and Dodd, L. E. (2004). A Cautionary Note on the Robustness of Latent Class Models for Estimating Diagnostic Error without a Gold Standard - Albert - 2004 - Biometrics - Wiley Online Library.Biometrics60,427–435
2004
-
[2]
Bossuyt, P. M. (2020). Testing COVID-19 tests faces methodological challenges.Journal of Clinical Epidemiology126,172–176
2020
-
[3]
and Huynh, M
Collins, J. and Huynh, M. (2014). Estimation of diagnostic test accuracy without full verification: A review of latent class methods.Statistics in Medicine33,4141–4169
2014
-
[4]
and Joseph, L
Dendukuri, N. and Joseph, L. (2001). Bayesian approaches to modeling the conditional dependence between multiple diagnostic tests.Biometricspages 158–167
2001
-
[5]
Goodman, L. a. (1974). Exploratory latent structure analysis using both identifiable and unidentifiable models.Biometrika61,215
1974
-
[6]
Greenland, S., Pearl, J., and Robins, J. (1999). Causal diagrams for epidemiologic research. Epidemiology1,37–48
1999
-
[7]
Gustafson, P. (2005). On model expansion, model contraction, identifiability and prior information: Two illustrative scenarios involving mismeasured variables.Statistical Science20,111–140
2005
-
[8]
MacLean, E., Kohli, M., K¨ oppel, L., Schiller, I., Sharma, S., Pai, M., et al. (2022). Bayesian latent class analysis produced diagnostic accuracy estimates that were more interpretable than composite reference standards for extrapulmonary tuberculosis tests.Diagnostic and Prognostic Research6:11,
2022
-
[9]
(2022).rjags: Bayesian Graphical Models using MCMC
Plummer, M. (2022).rjags: Bayesian Graphical Models using MCMC. R package version 4-13
2022
-
[10]
Qu, Y., Tan, M., and Kutner, M. (1996). Random effects models in latent class analysis for evaluating accuracy of diagnostic tests.Biometrics3,797–810. 24Biometrics, January2023 R Core Team (2022).R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria
1996
-
[11]
R., Maze, M
Schofield, M. R., Maze, M. J., Crump, J. A., Rubach, M. P., Galloway, R., and Sharples, K. J. (2021). On the robustness of latent class models for diagnostic testing with no gold standard.Statistics in Medicine40,4751–4763
2021
-
[12]
G., van Smeden, M., Dendukuri, N., Joseph, L., Nicol, M
Schumacher, S. G., van Smeden, M., Dendukuri, N., Joseph, L., Nicol, M. P., Pai, M. P., and Zar, H. J. (2016). Diagnostic Test Accuracy in Childhood Pulmonary Tuberculosis: A Bayesian Latent Class Analysis.American journal of epidemiology. van Smeden, M., Naaktgeboren, C. A., Reitsma, J. B., Moons, K. G. M., and de Groot, J. A. H. (2014). Latent Class Mod...
2016
-
[13]
Walter, S. D. and Irwig, L. M. (1988). Estimation of test error rates, disease prevalence and relative risk from misclassified data: a review.Journal of Clinical Epidemiology41, 923–937
1988
-
[14]
Y., Dendukuri, N., and Joseph, L
Wang, Z. Y., Dendukuri, N., and Joseph, L. (2016). Understanding the effects of conditional dependence in research studies involving imperfect diagnostic tests.Statistics in medicine
2016
-
[15]
(2008-2010)
Wasserman, L. (2008-2010). Directed graphical models.Carnegie Mellon University Chapter 18,. Received Nov2023 DAGs for interpretation of latent class analysis25 (a) Complete version (b) Reduced version Figure 1: Directed Acyclic Graph (DAG) for pediatric pulmonary tuberculosis tests Active TB disease refers to pediatic pulmonary tuberculosis 26Biometrics,...
2008
This paper was first reviewed by deepseek-v4-flash on August 2, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.