Pith. sign in

REVIEW 3 major objections 6 minor 15 references

A standard two-latent-class model for diagnostic tests can identify the wrong latent class — the majority measurand — when most tests do not measure the target condition.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Standard two-latent-class diagnostic models can mislabel latent classes when tests measure different measurands; modeling measurands explicitly via DAGs yields unbiased estimates and reinterprets two published analyses.

T0 review reviewed 2026-08-02 challenge →

load-bearing objection A genuinely useful framing for latent class analysis—DAG-driven measurands—with a solid simulation, but the applied examples and identifiability reporting need cleanup before the estimates can be trusted. the 3 major comments →

arxiv 2607.14473 v1 pith:EZG4WLGU submitted 2026-07-16 stat.ME

Improving interpretation of latent class models for diagnostic tests by recognizing their measurands via directed acyclic graphs (DAGs)

classification stat.ME MSC 62P1062F1562H30
keywords latent class analysisdiagnostic test accuracymeasuranddirected acyclic graphconditional dependenceidentifiabilityBayesian inferenceprevalence estimation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that in diagnostic accuracy studies without a gold standard, the conventional two-latent-class (2LC) model silently assumes every test measures the same target condition. In reality each test has its own measurand — the biological quantity it actually detects — which may be only related to the disease. The authors show that when most observed tests share a non-target measurand, the 2LC model's latent classes are labeled by that measurand rather than by the disease, biasing prevalence and sensitivity/specificity estimates. They propose drawing a DAG to expose each test's measurand, expanding the model to the resulting latent classes, and fitting that expanded model with Bayesian inference. Re-analysis of pediatric tuberculosis and leptospirosis data illustrates that the expanded model gives clinically more plausible class labels and different accuracy estimates than the original 2LC analyses.

Core claim

On the paper's own terms: a conventional two-latent-class model does not always identify the target condition. The latent classes it recovers are determined by the combination of observed tests; if the majority of tests measure a measurand distinct from the target condition, the classes are effectively M+ and M- (measurand positive/negative), not D+ and D- (target condition positive/negative). This happens because the tests' shared measurand induces conditional dependence that the 2LC model can only absorb by relabeling the latent variable. When the model is expanded so that latent classes are all combinations of the target condition and the measurands (as dictated by a DAG), the likelihood

What carries the argument

The central device is the distinction between a test's measurand (the quantity it is designed to measure, such as IgM antibodies or M. tuberculosis bacteria) and the target condition of interest (the disease one wants to diagnose). The paper represents both, plus observed tests and covariates, in a directed acyclic graph (DAG). The DAG reveals shared and nonshared measurands, which implies conditional dependence among tests with the same measurand and determines how many latent classes exist (combinations of target condition and measurands). The expanded likelihood then factors each observed test's result through its measurand: P(T_j | M_p) times P(M_p | D), so test accuracy and measurand ac

Load-bearing premise

The whole approach rests on the DAG being right: the chosen measurands, the dropped nodes, and the 'impossible' latent class combinations must genuinely reflect how the tests work; if any of these structural assumptions is wrong, the expanded model's class labels and accuracy estimates are just as unreliable as the 2LC model's.

What would settle it

If, in a dataset where the majority of tests are known (by design or external gold standard) to measure something other than the target condition, the 2LC model's estimated latent class prevalence still equals the target-condition prevalence (and test sensitivities match the target-condition values) rather than the measurand's prevalence, the paper's central claim would be false. Concretely: simulate with five tests where four detect a proxy M and one detects the disease D, with D and M only weakly correlated; the 2LC model should return M prevalence as its class prevalence. Any simulation tha

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • In any 2LC analysis, the latent class label is not guaranteed to be the target condition; it is an emergent property of which measurand dominates the test set.
  • When a majority of tests share a non-target measurand, prevalence and test-accuracy estimates from the 2LC model describe the measurand, not the disease.
  • Expanding the model to include measurand-based latent classes restores interpretable labels and separates test accuracy with respect to the measurand from accuracy with respect to the target condition.
  • The expanded models may fail local identifiability, but informative expert priors on one or two specificities can make them estimable.
  • Re-analysis of the leptospirosis data shows the previously 'low sensitivity' MAT estimate was an artifact: the 2LC classes were IgM+, not Leptospirosis+.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same mislabeling hazard applies to any finite mixture model applied to measurements that tap different constructs; the paper's DAG-plus-measurand recipe is a general diagnostic for mixture label validity, not just for medical tests.
  • A testable extension: when the majority of tests share a proxy measurand, the 2LC model's estimated prevalence should track the measurand's prevalence even if the target-condition prevalence is known from external sources; this can be checked in datasets with verified disease status.
  • The structural constraints (which latent class combinations are 'impossible') are assumptions; when those are wrong, the expanded model can be as mislabeled as the 2LC model. The paper does not sensitivity-analyze those constraints, but a robustness check that relaxes one constraint at a time would quantify the risk.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper argues that conventional two-latent-class (2LC) latent class models for diagnostic test accuracy can misidentify the latent variable when tests do not all measure the same target condition. It introduces directed acyclic graphs (DAGs) to explicitly separate each test's measurand from the target condition, and proposes expanded latent class models with classes defined by combinations of target condition and distinct measurands. The authors derive the mixture likelihood in Eq. (1), discuss identifiability in terms of degrees of freedom and the Jacobian rank, and use simulations to show that when the majority of tests measure a non-target measurand, the 2LC model tends to label that measurand rather than the target condition. They re-analyze two published datasets (pediatric pulmonary tuberculosis and leptospirosis) and report that the 2LC model in the TB example overestimates active TB prevalence (0.27 vs 0.22 under the proposed 4LC model) and that in the leptospirosis example the 2LC latent class appears to correspond to IgM rather than Leptospirosis infection.

Significance. The methodological message is important and clearly demonstrated by the simulation: ignoring the measurand structure of diagnostic tests can lead to incorrectly labeled latent classes and biased accuracy/prevalence estimates. The DAG-based framework is a useful tool for making these assumptions transparent and for structuring the latent class model. The applied examples are engaging and the face validity checks (e.g., treatment patterns in the TB data) lend some support. However, the applied conclusions depend critically on structural-zero assumptions that are not clinically justified and are not subjected to sensitivity analysis, and the models are only locally identifiable with the help of informative priors. These limitations materially weaken the paper's central applied claims. If the structural assumptions were relaxed, the numerical findings could change substantially; the paper does not provide evidence of robustness.

major comments (3)
  1. [§2.1, Table 1, Figure 1b] The reduced DAG drops 'Other respiratory disease' and relabels the x-ray measurand as 'Intrathoracic abnormalities due to TB'. Table 1 then declares the latent class (Active TB=0, TB Infection=0, Intrathoracic abnormalities=1) as not possible (NP). In a pediatric cohort presenting with respiratory symptoms, this is a strong and clinically questionable restriction: chest x-ray abnormalities frequently occur without active TB (e.g., pneumonia). Because no observed test measures 'other respiratory disease', children with such abnormalities are forced into class 4 (0,0,0), directly inflating the estimated class-4 prevalence (0.52, 95% CrI 0.34–0.64) and biasing x-ray sensitivity with respect to its measurand (reported 0.95, 0.82–1). The manuscript provides no sensitivity analysis to this structural zero. The headline prevalence difference (0.22 vs 0.27) is therefore not robust to a clinicall
  2. [§2.2, Table 3] The leptospirosis model suffers an analogous problem: the DAG in Figure 2b drops 'Other infection' because it is 'not estimable', and Table 3 restricts the model so that IgM/IgG arise only from Leptospirosis or from the target condition. The informative priors that the specificity of IgM exceeds 90% and that the specificity of IgG is 100% effectively impose the absence of other causes of these antibodies. In a hospital setting in Tanzania, this assumption is not clinically defensible. The resulting 6LC estimates—and the conclusion that the 2LC latent class corresponds to IgM rather than Leptospirosis—are conditional on this untestable exclusion. A sensitivity analysis that allows nonzero prevalence of other infections, even with weak prior information, is necessary before these applied conclusions can be considered reliable.
  3. [§5.1, Identifiability] The authors report that the TB model's Jacobian rank is 12 with 14 unknown parameters, requiring at least 2 informative priors (Section 5.1). The structural-zero restrictions in Table 1 are additional, non-estimated constraints. This combination means the posterior estimates—including the active TB prevalence of 0.22 and the measurand-specific accuracies—depend on both the choice of which parameters receive informative priors and the validity of the structural restrictions. The manuscript does not report any sensitivity analysis varying either the prior specifications or the structural-zero set. Given that the local identifiability condition is not met, the applied estimates may be largely driven by prior and structural assumptions rather than by the data, and the paper should demonstrate stability of its conclusions under plausible perturbations.
minor comments (6)
  1. [§4.1] The simulation uses the expected dataset rather than simulated random datasets, which is fine for illustrating mean behavior, but the sample size N is never stated. Please specify the sample size used to scale the expected frequencies, and consider reporting Monte Carlo standard errors or additional randomly simulated replicates to assess sampling variability.
  2. [Table 2] Table 2 is difficult to read: the column headers do not clearly indicate which columns are sensitivities versus specificities and with respect to which target/measurand. The reader has to reconstruct the structure from the prose. Please reorganize the table so that each test's accuracy with respect to each latent variable is presented unambiguously.
  3. [§2.1] There is a typo: 'thus illustrating they are are conditionally dependent' should read 'they are conditionally dependent'.
  4. [§5.2] In the leptospirosis results, the credible intervals for the prevalence of Leptospirosis under the 6LC model are extremely wide (0.28, 95% CrI 0.02–0.93; Table 4). Even the point estimates of sensitivity for several tests with respect to Leptospirosis have intervals spanning a large range. The conclusion that the 2LC model identifies IgM rather than Leptospirosis is partly based on point estimates with substantial uncertainty; the overlap of intervals should be acknowledged or handled more carefully in the interpretation.
  5. [References] The Wasserman reference is incomplete ('2008-2010', 'Chapter 18') and should be updated to a proper citation. Also, 'Mycobaterium' is a typo in Section 2.1.
  6. [§5.1] The text says '13 unknown parameters' and then later 'total number of unknown parameters was 19' after accounting for prevalence covariates and a continuous measurand. The derivation of the 19 total parameters (beyond the 13) is not shown and should be clarified, since the reader is referred to the supplementary material for details that are not available in the main text.

Circularity Check

0 steps flagged

No significant circularity; the measurand-DAG framework is an explicit modeling assumption, and the 2LC-bias demonstrations are simulation-based rather than definitional.

full rationale

The paper's central derivation is a probability decomposition from stated latent-variable assumptions (Section 3.1, Equation 1); no target quantity is defined in terms of the conclusion. The simulation study generates data from an explicit DAG-based model with fixed prevalence and accuracy values, then fits a deliberately misspecified 2LC model and compares estimates to the true values; this is a standard misspecification illustration, not a fitted parameter relabeled as a prediction. The applied examples impose structural constraints (Tables 1 and 3) and explicit informative priors, including priors taken from earlier publications by the same group (Schumacher et al. 2016). These are transparent modeling assumptions and prior inputs, not conclusions derived from the data; the paper even discloses local non-identifiability (rank 12 vs 14 parameters) and states that informative priors are needed. The self-citations are used for prior values and modeling conventions, but the measurand claim itself is supported by the likelihood derivation and simulations, not by authority. The acknowledged limitations (dropping 'Other respiratory disease', 'Other infection', dichotomous tests, missing data) concern model validity and robustness, not circularity. No step was found where a prediction is equivalent by construction to its inputs.

Axiom & Free-Parameter Ledger

0 free parameters · 6 axioms · 0 invented entities

The paper's applied results rest on a sequence of unverified structural choices: the DAG topology, the restriction of certain latent classes as impossible, the dropping of measurands/nodes that cannot be estimated, and informative expert priors (IgG specificity 1, IgM specificity >0.9, culture/Xpert specificity bounds). No constants are fitted to force the central conclusion; the simulation inputs are design choices rather than fitted values. The likelihood itself is a standard conditional-independence mixture.

axioms (6)
  • domain assumption Tests are conditionally independent given their measurand and the target condition (Eq. 1).
    Basis of the likelihood; if unmodeled conditional dependence remains (e.g., shared laboratory batches), likelihood is misspecified.
  • domain assumption DAG structure for TB: Culture, Xpert, Smear share the measurand 'Detectable M.tb bacteria'; X-ray measures intrathoracic abnormalities due to TB; TST measures immune response to M.tb.
    Imported from clinical knowledge (Schumacher 2016), not empirically identified in this paper; determines latent classes in Table 1.
  • domain assumption Impossible latent class restrictions in TB: active TB without latent TB infection is impossible; intrathoracic abnormalities due to TB without active TB are impossible.
    Structural zeros in Table 1; if wrong, latent class labels and estimates change.
  • ad hoc to paper Reduced DAG drops 'DNA of M.tb' and 'Other respiratory disease'/'Other infection' nodes because they are not statistically identifiable with available data.
    Post hoc simplification in Section 2.1; the paper offers no sensitivity analysis to this reduction.
  • domain assumption DAG structure for leptospirosis: three rapid tests measure IgM; MAT measures both IgG and IgM; IgM may be caused by other infections.
    Clinical assumption (Schofield 2021); used to define the 6LC model.
  • domain assumption Expert priors: specificity of IgM > 90%; specificity of IgG = 100% (and culture/Xpert bounds in TB).
    Informative priors applied to achieve identifiability; if wrong, posterior estimates shift.

reviewed 2026-08-02 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving interpretation of latent class models for diagnostic tests by recognizing their measurands via directed acyclic graphs (DAGs)." pith.science (2026). https://pith.science/paper/EZG4WLGU

@misc{pith2026260714473,
  author       = {Pith},
  title        = {Pith review of: Improving interpretation of latent class models for diagnostic tests by recognizing their measurands via directed acyclic graphs (DAGs)},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EZG4WLGU}},
  note         = {Machine review of arXiv:2607.14473}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Summary: In the absence of a perfect diagnostic test for a target condition, multiple imperfect tests may be used to arrive at a clinical diagnosis. Latent class analysis can be used to model such data with the objective of estimating test accuracy and target condition prevalence. Such models typically assume two latent classes - target condition positive and target condition negative. However, as we will illustrate in this manuscript, this would be an oversimplification if the different tests do not share the target condition as their measurand. We show how a Directed Acyclic Graph (DAG) can be used to illustrate the relationships between the relevant variables - the observed imperfect test results, their latent measurands, the latent target condition of interest and observed covariates - revealing any conditional dependence relations. The DAG helps determine the number of latent classes, underlying the observed data, and their labels. We show how the likelihood function changes due to incorporating the measurand of each test. We study the impact on identifiability of the model. Using simulation studies we show how ignoring the measurand of an imperfect test, when it is distinct from the target condition, can lead to biased estimates of test accuracy and prevalence. We illustrate the value of the proposed approach by re-analyzing two datasets used in previously published latent class analyses of tests for pediatric tuberculosis and leptospirosis.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

15 extracted references

  1. [1]

    Albert, P. S. and Dodd, L. E. (2004). A Cautionary Note on the Robustness of Latent Class Models for Estimating Diagnostic Error without a Gold Standard - Albert - 2004 - Biometrics - Wiley Online Library.Biometrics60,427–435

  2. [2]

    Bossuyt, P. M. (2020). Testing COVID-19 tests faces methodological challenges.Journal of Clinical Epidemiology126,172–176

  3. [3]

    and Huynh, M

    Collins, J. and Huynh, M. (2014). Estimation of diagnostic test accuracy without full verification: A review of latent class methods.Statistics in Medicine33,4141–4169

  4. [4]

    and Joseph, L

    Dendukuri, N. and Joseph, L. (2001). Bayesian approaches to modeling the conditional dependence between multiple diagnostic tests.Biometricspages 158–167

  5. [5]

    Goodman, L. a. (1974). Exploratory latent structure analysis using both identifiable and unidentifiable models.Biometrika61,215

  6. [6]

    Greenland, S., Pearl, J., and Robins, J. (1999). Causal diagrams for epidemiologic research. Epidemiology1,37–48

  7. [7]

    Gustafson, P. (2005). On model expansion, model contraction, identifiability and prior information: Two illustrative scenarios involving mismeasured variables.Statistical Science20,111–140

  8. [8]

    MacLean, E., Kohli, M., K¨ oppel, L., Schiller, I., Sharma, S., Pai, M., et al. (2022). Bayesian latent class analysis produced diagnostic accuracy estimates that were more interpretable than composite reference standards for extrapulmonary tuberculosis tests.Diagnostic and Prognostic Research6:11,

  9. [9]

    (2022).rjags: Bayesian Graphical Models using MCMC

    Plummer, M. (2022).rjags: Bayesian Graphical Models using MCMC. R package version 4-13

  10. [10]

    Qu, Y., Tan, M., and Kutner, M. (1996). Random effects models in latent class analysis for evaluating accuracy of diagnostic tests.Biometrics3,797–810. 24Biometrics, January2023 R Core Team (2022).R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria

  11. [11]

    R., Maze, M

    Schofield, M. R., Maze, M. J., Crump, J. A., Rubach, M. P., Galloway, R., and Sharples, K. J. (2021). On the robustness of latent class models for diagnostic testing with no gold standard.Statistics in Medicine40,4751–4763

  12. [12]

    G., van Smeden, M., Dendukuri, N., Joseph, L., Nicol, M

    Schumacher, S. G., van Smeden, M., Dendukuri, N., Joseph, L., Nicol, M. P., Pai, M. P., and Zar, H. J. (2016). Diagnostic Test Accuracy in Childhood Pulmonary Tuberculosis: A Bayesian Latent Class Analysis.American journal of epidemiology. van Smeden, M., Naaktgeboren, C. A., Reitsma, J. B., Moons, K. G. M., and de Groot, J. A. H. (2014). Latent Class Mod...

  13. [13]

    Walter, S. D. and Irwig, L. M. (1988). Estimation of test error rates, disease prevalence and relative risk from misclassified data: a review.Journal of Clinical Epidemiology41, 923–937

  14. [14]

    Y., Dendukuri, N., and Joseph, L

    Wang, Z. Y., Dendukuri, N., and Joseph, L. (2016). Understanding the effects of conditional dependence in research studies involving imperfect diagnostic tests.Statistics in medicine

  15. [15]

    (2008-2010)

    Wasserman, L. (2008-2010). Directed graphical models.Carnegie Mellon University Chapter 18,. Received Nov2023 DAGs for interpretation of latent class analysis25 (a) Complete version (b) Reduced version Figure 1: Directed Acyclic Graph (DAG) for pediatric pulmonary tuberculosis tests Active TB disease refers to pediatic pulmonary tuberculosis 26Biometrics,...

This paper was first reviewed by deepseek-v4-flash on August 2, 2026.