{"id":"ebfa102a-6620-43fc-916f-a91684d468dc","arxiv_id":"2502.00171","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Grouping symptoms into clusters and modeling each cluster with a separate latent class improves cause-of-death assignment and interpretability in verbal autopsy data.","lead":"This paper introduces a Bayesian statistical model that groups related symptoms when predicting cause of death from verbal autopsy interviews. The grouped structure aims to improve prediction accuracy while making the symptom patterns easier for public health researchers to interpret.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Predictive-accuracy claim is not independently established: the simulation generates data from the proposed c-Tucker model at the fitted settings, and the PHMRC comparison shows only CSMF gains, not better top-cause accuracy.","rationale":"The paper is a well-specified methodological contribution: the conditional tensor models are clearly defined, the MCMC updates in Section 3 are coherent, and the PHMRC analysis offers a useful interpretability exercise. The load-bearing issue is not internal consistency but the evidential basis for the comparative claim in the abstract. A simulation that generates data from the very model being promoted cannot arbitrate between that model and alternatives; at best it checks estimation and recovery. The real-data comparison is the only independent evidence, and it shows mixed results: the c-Tucker model trails LCVA on top-cause accuracy and leads on CSMF accuracy. Since the abstract asserts overall better predictive accuracy, the paper needs a non-circular simulation or an external validation to support that statement. The label-shift assumption identified by the reader is a real limitation, but it is standard in VA modeling, is explicitly acknowledged, and the paper partially probes it in Scenario II; the more immediate soft spot is that the central comparative claim rests on data generated from the proposed model. A concrete test that regenerates the simulation from a competing or non-nested data-generating process would settle whether the claimed advantage is genuine. This does not change the reader's CONDITIONAL verdict; it sharpens the condition under which the paper should be accepted.","tokens_in":14528,"tokens_out":10073,"duration_ms":109218,"concrete_test":"Run a new simulation (same n, C, p, and informative-symptom structure as Section 4) drawing X from a model not nested in the proposed family, e.g., a standard PARAFAC with K=15 with cause-specific sparse profiles, or a latent Gaussian factor model with correlated binary symptoms. Fit c-Tucker (K=4, r=8, h=3), rIndep, PARAFAC K=10/15, and LCVA using the paper's own settings. If c-Tucker and rIndep no longer achieve the best or near-best top-cause and CSMF accuracy across the 50 replicates, the claimed advantage is an artifact of evaluating a model on data generated from itself.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's claim of 'better predictive accuracy than existing VA methods' is not supported by the evidence as it stands. In Section 4, Scenario I, all synthetic X are drawn from the c-Tucker model with the same K=3, r=5, h=3 used when fitting the proposed models; the comparisons in Figure 3 therefore show that c-Tucker can recover its own generative process, not that it outperforms competitors in general. Scenario II perturbs only the mixing weights ψ while keeping the same c-Tucker generative family, so it remains a within-family test. The only non-circular benchmark, Section 5.4 on PHMRC, shows c-Tucker's top-cause accuracy below LCVA (Figure 9); the claimed advantage reduces to CSMF accuracy on target datasets resampled from the same population as the training data. Thus the headline comparative claim depends on a favorable simulation design and is not confirmed by the real-data comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes two Bayesian hierarchical tensor decomposition models for verbal autopsy data: r-group independent PARAFACs and collapsed Tucker (c-Tucker). Symptoms are partitioned into groups, with cause-specific grouping indicators; each group has its own latent class sub-profile, and c-Tucker adds a higher-level mixture over group weights. The authors present Gibbs sampling updates, a simulation study, and an application to the PHMRC gold-standard dataset, claiming better predictive accuracy than LCVA and InSilicoVA along with a more parsimonious and interpretable latent representation.","tokens_in":14733,"tokens_out":5890,"duration_ms":59821,"significance":"If the comparative accuracy claim held, the tensor-decomposition framework would be a valuable addition to VA methodology because it combines flexible symptom dependence with interpretable symptom clusters. The paper has real strengths: the hierarchical representation in Section 2.3 is clearly specified, the MCMC updates in Section 3 are explicit and complete, and the PHMRC analysis includes comparisons with LCVA and InSilicoVA and exposes clinically coherent groupings (Figures 6-8). The central problem is that the headline claim of 'better predictive accuracy than existing VA methods' is not supported by the evidence: the simulation is a within-family recovery experiment, and the real-data results show c-Tucker is slightly worse than LCVA on top-cause accuracy (Figure 9). The methodological contribution is promising, but the evidence base needs revision before the comparative claims can be accepted.","major_comments":[{"comment":"The comparative claim in the Abstract that the proposed methods 'achieve better predictive accuracy than existing VA methods' is not supported by the evidence as presented. In Section 4, the synthetic data are generated from the proposed c-Tucker model with K=3, r=5, h=3 and the proposed models are fitted under the same K, r, h; the comparators are standard PARAFAC with K=5, 10, 15 and LCVA with K=10. Figure 3 therefore largely demonstrates that c-Tucker can recover its own generative process. The non-circular comparison on PHMRC (Figure 9) shows that c-Tucker has slightly lower top-cause accuracy than LCVA and only a CSMF-accuracy advantage. I recommend either adding simulations that generate data from competing or out-of-family models, or revising the abstract and conclusions to claim 'comparable top-cause accuracy and improved CSMF accuracy on PHMRC'.","section":"Abstract; Section 4; Section 5.4"},{"comment":"The label-shift assumption stated in Section 2.3--p(X|Y) is the same in training and target, with only p(Y) changing--is load-bearing for transfer to unlabeled target data. Scenario II of the simulation changes the mixing weights ψ(g) between training and target, but the data are still generated from the same c-Tucker family, so this tests only a restricted form of misspecification. The text in Section 4 stating that the proposed models are 'more robust to distribution shift' is stronger than warranted; realistic violations such as differential symptom reporting or questionnaire changes are not examined, and the Discussion acknowledges this limitation. Please add a sensitivity analysis with a genuinely different target distribution, or temper the robustness claims.","section":"Section 2.3, Eqs. (12)-(13); Section 4, Scenario II"},{"comment":"The interpretability analysis relies on posterior means of the latent parameters ϕ, ν, and ψ (e.g., Figure 7), but the paper does not address label switching or permutation invariance of the Bayesian mixture/tensor decomposition. Without a relabeling constraint or post-processing, posterior averages across MCMC iterations can mix different permutations of the latent classes, making the reported 'posterior mean' profiles and the cause dendrogram difficult to interpret. Please either justify that label switching is resolved in the sampler or use label-invariant summaries (e.g., clustering-based relabeling, or summaries of permutation-invariant quantities).","section":"Section 5.2, Figure 7"},{"comment":"The PHMRC results are conditional on model dimensions selected by the utilization-rate heuristics in Section 5.1 (r=8, K=4, h=3 for c-Tucker; r=6, K=5 for rIndep), with no sensitivity analysis around these choices. Since the headline comparison in Figure 9 depends on these settings, the reported advantage could be an artifact of tuning. Please report accuracy for nearby values of K and r (and h) and describe how sensitive the conclusions are.","section":"Section 5.1"}],"minor_comments":[{"comment":"There is a notational inconsistency: Eq. (3) uses sj for symptom group membership, while the hierarchical representation in Eqs. (7)-(8) and the MCMC update in Step 8 use scj, which is cause-specific. Please clarify whether the grouping is shared across causes or cause-specific and align the notation throughout.","section":"Section 2.2, Eqs. (3)-(4)"},{"comment":"The text says the analysis is illustrated 'in one synthetic dataset' even though Section 5 analyzes the PHMRC data. Since the PHMRC target datasets are resampled, please clarify whether this refers to one resampled target dataset or a simulated dataset.","section":"Section 5.2"},{"comment":"The boxplots in Figure 9 are informative, but numerical means and credible intervals for top-cause accuracy and CSMF accuracy across the 50 target datasets would make the comparison easier to assess, especially because the differences between c-Tucker and LCVA appear small.","section":"Section 5.4, Figure 9"},{"comment":"There is a typo, 'Specicically' instead of 'Specifically'.","section":"Introduction, paragraph 4"}],"recommendation":"major_revision","confidential_remarks":"The methodological core is sound and the revision path is clear. The main issue is that the abstract and parts of Section 4 overstate the predictive-accuracy evidence; the PHMRC comparison actually shows only a CSMF advantage. With a revised evidence base, a relabeling discussion, and a sensitivity analysis for model dimensions, the paper could be acceptable. The Section 5.2 'synthetic dataset' wording should also be corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful VA paper with a real interpretability contribution, but the headline accuracy claim is overstretched. The models (r-group independent PARAFACs and c-Tucker) come from Johndrow et al. 2017, and the authors say so. The new piece is the VA application: data-adaptive symptom grouping, label-shift handling, full MCMC, and the PHMRC analysis showing meaningful symptom and cause clusters.\n\nStrengths: the model is cleanly defined, the Gibbs steps are explicit and reproducible, and the real-data interpretation (stroke and AIDS symptom groups, cause dendrogram matching clinical categories) is a genuine advance. The paper ships the machinery to do this, and the PHMRC comparisons with LCVA and InSilicoVA are honest even when the proposed method is not on top.\n\nSoft spots, proportional: the abstract says \"better predictive accuracy than existing VA methods,\" but the evidence does not support that as a general claim. The simulation generates data from the proposed c-Tucker at the same K, r, h used for fitting, so it is a recovery test of the generative model, not a fair comparison. Scenario II only changes the mixing weights, so it is still within the same family. On PHMRC, c-Tucker is slightly worse than LCVA at top-cause accuracy and better at CSMF on resampled targets from the same population. That is a real but narrower result. The label-shift assumption is standard, but it is load-bearing and untested beyond a same-family perturbation; a cross-population validation would strengthen the paper substantially. The model-complexity selection is heuristic, which is minor. The discussion section is candid about limitations, which matters.\n\nVerdict: the central modeling contribution holds up, the interpretability results are the most valuable part, and the paper deserves a serious referee. It should not be desk-rejected, but the authors should be asked to temper the abstract and add a genuine out-of-family or cross-population check.","headline":"Solid VA modeling paper with a real interpretability payoff, but the abstract overclaims predictive accuracy; the simulation is a within-family recovery test and the real-data gain is CSMF-specific.","tokens_in":15232,"tokens_out":1654,"would_cite":true,"duration_ms":15750,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62H30","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Modeling symptom groups with a collapsed Tucker decomposition makes verbal autopsy cause assignment more accurate and more interpretable.","keywords":["verbal autopsy","tensor decomposition","collapsed Tucker decomposition","PARAFAC","latent class model","cause-of-death assignment","CSMF accuracy","label shift"],"falsifier":"Fit the c-Tucker model on one population's labeled verbal autopsies and apply it to a target population whose symptom-reporting patterns for the same underlying causes are known to differ (for example, a different questionnaire translation or care-seeking context), then compare the predicted cause-of-death fractions against medically certified causes. If CSMF accuracy falls well below the level reported on the PHMRC resampling experiment, the label-shift assumption is the point of failure.","tokens_in":14331,"feed_emoji":"🩺","tokens_out":10203,"duration_ms":88700,"temperature":0.7,"pith_summary":"Verbal autopsy (VA) data—questionnaires about symptoms collected from caregivers—are used to infer causes of death where medical certification is unavailable. The paper's central proposal is to model the conditional distribution of symptoms given cause not with a single latent class variable but with a tensor factorization that first partitions symptoms into groups and then models the joint distribution of group-level sub-profiles. Two variants are introduced: r-group independent PARAFACs, which assume groups are conditionally independent, and collapsed Tucker (c-Tucker) decomposition, which adds a higher-level latent variable so groups can share dependence. On the PHMRC gold-standard dataset, c-Tucker with four latent classes, eight symptom groups, and three higher-level components yields top-cause accuracy comparable to LCVA and better CSMF accuracy (the estimated population cause fractions) than LCVA and InSilicoVA, while learning interpretable symptom clusters and a cause-of-death dendrogram. The underlying claim is that this dimension-grouped parameterization balances flexibility and parsimony better than standard latent class models.","feed_headline":"Grouped symptoms sharpen verbal autopsy cause-of-death estimates","feed_subtitle":"Grouping symptoms jointly lifts population cause-fraction accuracy while keeping latent profiles readable.","key_machinery":"The central object is the conditional symptom probability tensor $p(X \\mid Y)$, a $p$-way binary tensor per cause. The key mechanism is the collapsed Tucker (c-Tucker) decomposition: each cause's distribution is a mixture over $r$ group-level latent classes, and the joint mixing weight $\\lambda_{c,k_1,\\dots,k_r}$ is itself written as a rank-$h$ PARAFAC, $\\lambda_{c,k_1,\\dots,k_r} = \\sum_{l=1}^h \\nu_{c l} \\prod_{s=1}^r \\psi_{c l s k_s}$. The factors $\\phi_{c k j}$ are Bernoulli probabilities for symptom $j$ under latent class $k$, and the grouping $s_{c j}$ is estimated from data rather than fixed. The r-group independent PARAFACs model is the special case $h=1$, and standard PARAFAC is $h=r=1$, so the approach interpolates between conditional independence and full Tucker structure. This decomposition carries the argument: it replaces one global latent class with many small group-level classes, cutting the profile count from $K^r$ to $pK$ and making the learned groups directly interpretable.","core_discovery":"The paper's central claim is that the probability tensor $p(X \\mid Y)$ of binary symptoms conditional on cause of death can be approximated by a dimension-grouped tensor decomposition that is both more flexible and more parsimonious than a single PARAFAC/latent class model. With symptom groups indexed by $s=1,\\dots,r$, each cause $c$ is described by group-specific latent class indicators $Z_{is}$ and symptom sub-profiles $\\phi_{c k j}$; the r-group independent PARAFACs model treats the group indicators as independent given cause, while the c-Tucker model couples them through a higher-level latent variable $H_i$ whose mixing weights $\\nu_c$ and group factors $\\psi_{c l s}$ form another PARAFAC. This lets the model express up to $K^r$ distinct latent symptom profiles using only $pK$ profile parameters. On the PHMRC dataset the fitted c-Tucker model ($K=4$, $r=8$, $h=3$) is reported to match LCVA on top-cause accuracy and exceed all compared methods on CSMF accuracy, and the posterior symptom groups reproduce recognizable clusters (e.g., stroke and AIDS each get distinct but overlapping symptom topics) while the cause-level dendrogram aligns with broad disease categories.","pith_inferences":["Editorial inference: the estimated symptom groups, especially the anchor symptoms with high group-assignment probability, could be used to shorten verbal autopsy questionnaires without much loss of classification accuracy.","Editorial inference: the Scenario II simulation suggests the grouped latent structure can absorb some cause-conditional distribution shift, so the c-Tucker model may serve as a building block for multi-source domain adaptation even though the paper does not implement it.","Editorial inference: the Adjusted Rand Index-based cause dendrogram could be used to decide when rare causes should be merged in a cause hierarchy, a decision VA users currently make ad hoc.","Editorial inference: because the empirical comparison uses only the PHMRC dataset, the claim that c-Tucker beats LCVA on CSMF accuracy would be strengthened by replication on independent gold-standard VA data such as WHO-2016."],"forward_implications":["With $r$ symptom groups and $K$ latent classes, c-Tucker represents a dependence structure that a standard PARAFAC would need $K^r$ latent profiles to express, so complex symptom dependence is captured with far fewer parameters.","The estimated symptom groups and the cause-level dendrogram give a descriptive summary of which symptoms cluster together for each cause, and the dendrogram aligns with broad medical categories such as infectious, circulatory, neoplastic, and external causes.","On the PHMRC resampled target datasets, c-Tucker achieves the highest CSMF accuracy among InSilicoVA, LCVA, PARAFAC, and r-group independent PARAFACs, meaning better estimates of population cause-of-death fractions.","In the simulation study, both grouped models maintain higher top-cause and CSMF accuracy than standard PARAFAC and LCVA even under a misspecified distribution shift, indicating that the flexible grouped structure is more robust to model error.","The model provides a practical selection rule for latent dimensions: fit with large $K$ and $r$, then choose the smallest values whose groups and classes are utilized in more than 5% of posterior samples."],"supporting_citations":[{"why":"Supplies the collapsed Tucker decomposition formulation that the c-Tucker model builds on.","marker":"Johndrow and others (2017)"},{"why":"Provides the Bayesian PARAFAC decomposition of categorical probability tensors underlying the latent-class representation.","marker":"Dunson and Xing (2009)"},{"why":"Defines the InSilicoVA conditional-independence baseline and the preprocessing that turns PHMRC data into 168 binary symptoms.","marker":"McCormick and others (2016)"},{"why":"Gives the LCVA sparse PARAFAC model used as the main benchmark and motivates the need for flexible latent classes.","marker":"Li and others (2024)"},{"why":"Contributes the PHMRC gold-standard verbal autopsy dataset on which the empirical evaluation is run.","marker":"Murray and others (2011)"},{"why":"Formalizes label shift, the assumption that only cause prevalence may differ between training and target data.","marker":"Storkey (2009)"}],"fun_headline_variants":["Grouped symptoms sharpen verbal autopsy cause estimates","Interpretable tensor decomposition boosts verbal autopsy accuracy","Bayesian tensor model simplifies and sharpens verbal autopsy","Grouped symptom profiles refine cause-of-death assignment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes label shift: the chance of reporting each symptom given a cause is identical in the training and target populations, so only the cause-of-death prevalence may differ; if symptom reporting changes, predicted causes will be biased.","fun_headline_variants_meta":{"raw":{"variants":["Grouped symptoms sharpen verbal autopsy cause estimates","Interpretable tensor decomposition boosts verbal autopsy accuracy","Bayesian tensor model simplifies and sharpens verbal autopsy","Grouped symptom profiles refine cause-of-death assignment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000519,"raw_usage":{"total_tokens":2544,"prompt_tokens":1007,"completion_tokens":1537,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":623,"completion_tokens_details":{"reasoning_tokens":1478}},"tokens_in":623,"tokens_out":1537,"duration_ms":13839,"temperature":1.0,"reasoning_tokens":1478,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T19:56:44.482619+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit the c-Tucker model on one population's labeled verbal autopsies and apply it to a target population whose symptom-reporting patterns for the same underlying causes are known to differ (for example, a different questionnaire translation or care-seeking context), then compare the predicted cause-of-death fractions against medically certified causes. If CSMF accuracy falls well below the level reported on the PHMRC resampling experiment, the label-shift assumption is the point of failure.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the collapsed Tucker decomposition formulation that the c-Tucker model builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Bayesian PARAFAC decomposition of categorical probability tensors underlying the latent-class representation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the InSilicoVA conditional-independence baseline and the preprocessing that turns PHMRC data into 168 binary symptoms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the PHMRC gold-standard verbal autopsy dataset on which the empirical evaluation is run."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Formalizes label shift, the assumption that only cause prevalence may differ between training and target data."}],"review_version":1}