{"id":"1ed1ece3-e577-4e5f-8644-bb746fa3c9f7","arxiv_id":"2607.29530","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An automated neurosymbolic pipeline recovers and proposes monotonic influence rules linking cognitive impairment to verbal fluency features extracted from raw audio.","lead":"The paper builds an automated pipeline that turns audio recordings of verbal fluency tests into a Bayesian network and extracts qualitative rules about how Alzheimer's-related impairment relates to speech and language markers. It reports recovering known clinical relationships and proposing new ones, but the network skeleton comes from an LLM and the clinical dataset is small and private.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Q3's nine 'novel' MIs are confounded with the LLM-generated skeleton: QuaKE only tests CI descendants, and 43% of LLM edges are extra, so graph prior, not clinical data, may produce them.","rationale":"Read in good faith, NeSyQuaKE is a coherent Type-3 neurosymbolic system with deterministic symbolic programs and an honest evaluation of the acoustic front-end. The PCEE gains and the appendix's transcription-error analysis are real support. However, the strongest claim—discovery of novel relationships—is supported only by QuaKE outputs, whose search space is the LLM skeleton. The paper itself supplies the evidence for the confound: 80% edge recovery and 43% extra edges. Because QuaKE evaluates marginal P(X|Y) only for CI descendants, the candidate set is not data-driven. The expert 'validation' of the 9 extra rules is post-hoc and cannot distinguish data signal from LLM prior. The reader's weakest_assumption identified the same issue; my emphasis is that the Q3 novelty set is exactly the set of variables the LLM graph made descendants. The verdict should remain conditional: the recovery claim (Q2) may survive, but the novel-hypothesis claim needs an ablation over skeletons and thresholds before acceptance. This is not a rejection of the pipeline; it is a request for a controlled comparison that separates prior from data.","tokens_in":14706,"tokens_out":4325,"duration_ms":48446,"concrete_test":"Using the released code and the same extracted features/discretization, rerun QuaKE on three skeletons: (a) the expert graph (Fig 2 left), (b) the LLM graph (Fig 2 right), and (c) the LLM graph with each non-expert edge removed one at a time. Compare the 9 Q3 MIs. If any of them disappear under (a) or under a single-edge removal, the 'novel relationships' are graph-prior artifacts. As a secondary check, repeat (a)-(c) with median rather than mean thresholds; if sign or existence of the 9 MIs changes, discretization is also load-bearing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"§3.3 restricts QuaKE to descendants of CI in the BN skeleton. The skeleton is not learned from data; §3.2 obtains it by prompting an LLM. §4 Q2 reports only 80% expert-edge recovery and 43% extra edges relative to the expert network. Therefore the Q3 'additional influences' (9 MIs beyond the expert set) are exactly the variables that the LLM graph made descendants of CI. Had the LLM omitted a true edge, a true monotonicity could be missed; had it added a false edge, a spurious monotonicity could be extracted. The post-hoc expert agreement with these 9 rules does not break this confound, because no blinded, pre-specified validation protocol is described. The PC baseline (1/15 expert edges, Appendix E) shows data-driven structure recovery is unusable in this sample, so no independent skeleton check exists. The mean-threshold discretization (Appendix C) is a second arbitrary choice that can change the computed marginal P(X|Y) and therefore the inferred MI signs. Thus the central novelty claim is not separable from the LLM prior.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes NeSyQuaKE, a Type-3 neurosymbolic pipeline that turns raw verbal-fluency audio into qualitative monotonic-influence (MI) statements about cognitive impairment (CI). The pipeline uses WhisperX for speech-to-text, MedGemma-27b with schema-constrained decoding for structured extraction, deterministic symbolic programs to compute eight clinical features, and an LLM-prompted Bayesian Network (BN) skeleton. A BDeu posterior is estimated from 162 HC/MCI recordings, and the QuaKE algorithm extracts MIs from CI (and other features) to their descendants in that skeleton. The paper reports that the grounding stage substantially lowers the Pairwise Conditional Estimation Error (PCEE) relative to a rule-based extractor (§4, Table 2), that the extracted MIs match all expert-elicited rules in both semantic and phonemic fluency tasks (§4, Table 3), and that nine additional influences are discovered, several later confirmed by the authors' domain experts. A feedback loop that repairs rejected words via an LLM is also reported to improve PCEE.","tokens_in":15049,"tokens_out":3510,"duration_ms":42013,"significance":"If the central claim holds, NeSyQuaKE would be a credible scalable method for automated, explainable hypothesis generation from clinical audio, combining foundation-model grounding with symbolic probabilistic reasoning. The paper has genuine strengths: the use of deterministic symbolic programs for feature computation is a transparent and verifiable design; the evaluation against manual annotations is appropriate; the code is made available; and the comparison to a PC-learned skeleton is an honest baseline that shows data-driven structure recovery is not viable at this sample size. These are important positive features. However, the hypothesis-generation claim—the paper's most novel contribution—is currently not separable from the LLM-supplied BN skeleton, and the quantitative evidence lacks uncertainty quantification. The reported recovery of known clinical knowledge and the discovery of novel influences therefore need to be shown to be robust to the skeleton prior and the discretization choice before the central claim can be accepted at face value.","major_comments":[{"comment":"The Q3 'additional hypotheses' claim is confounded with the LLM-generated skeleton. QuaKE tests monotonic influence only for descendants of CI in the graph G (§3.3), and G is obtained by prompting an LLM (§3.2). Section 4 reports that the LLM graph recovers only 80% of expert edges and contains 43% extra edges. Thus the nine 'additional' MIs are exactly the variables that the LLM graph placed as descendants of CI; a spurious LLM edge can manufacture an MI and an omitted edge can hide a true one. The post-hoc expert agreement with these nine rules is not a substitute for a blinded, pre-specified validation protocol. Please report which of the nine MIs survive when QuaKE runs on (i) the expert skeleton, (ii) the PC skeleton, and (iii) the LLM skeleton under different edge-inclusion thresholds/pooling rules. Without this ablation, the discovery claim is not separable from the LLM prior.","section":"§3.2, §3.3, §4 Q3"},{"comment":"All MI signs are computed from conditional probabilities over discretized features. The discretization thresholds are per-variable means from five trials; no sensitivity analysis is reported. Since the MI score in Eq. (1) depends on P(X_d ≤ k | Y), a shift in a threshold (e.g., to the median or to tertiles) can change the marginal conditional distribution and therefore the sign of an inferred MI. Please provide a threshold-sensitivity analysis and show that the expert-recovered and novel MIs in Table 3 are invariant across reasonable discretization choices. This is load-bearing because the core qualitative statements are defined on the discrete variables.","section":"Appendix C, §3.3"},{"comment":"No confidence intervals, significance tests, or bootstrap estimates are reported. The statement in Table 2 that all standard deviations are below 1e-4 reflects only run-to-run variation of the stochastic LLM component, not patient-level sampling uncertainty over the 162 recordings. For the central recovery and discovery claims, please report bootstrap CIs for the PCEE values and for the MI strengths δ_d, or a permutation test comparing the observed monotonicity score to the null of no monotonic influence. Without this, the reader cannot distinguish a robust clinical signal from noise in a small, noisy clinical dataset.","section":"Tables 2, 3; Fig. 3"}],"minor_comments":[{"comment":"Minor spacing typos appear in the author list ('Pranuthi T enali', 'V aishali Phatak') and in body headings ('F oundation Models', 'T o Do'). The method name is also rendered inconsistently as 'NeSyQuake' in several places (e.g., §4 Q1, Table 2).","section":"Author list and text"},{"comment":"The percent-difference formula is ambiguous: 'Percent Difference = |Corrected Count - Original Count| / Original Count' does not specify whether 'Original' is the raw extraction count or the manual annotation count. The table caption also says 'Phonetic Fluency' while the rest of the paper uses 'phonemic fluency'.","section":"Appendix D"},{"comment":"The heading 'Peter Clarke (PC) Algorithm' should be 'Peter Clark'.","section":"Appendix E"},{"comment":"The citation '(adf, 2024)' should be expanded to the standard Alzheimer's Association citation; also verify the reference formatting for 'Beatty et al.' and 'Laws et al.' against the journal style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The reviewer's main concern is legitimate: the Q3 discovery claim is not yet separable from the LLM-provided skeleton, and the absence of threshold sensitivity and uncertainty quantification weakens the quantitative claims. The request for a skeleton ablation and a threshold sensitivity analysis is, in my view, the correct path. I would not require full causal discovery—the PC baseline already shows that is infeasible at this sample size—but stability across reasonable skeletons and thresholds is necessary for the contribution to stand."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real applied contribution, not a paradigm shift. It packages WhisperX, MedGemma, deterministic symbolic programs, and QuaKE into an end-to-end pipeline for AD verbal fluency audio, and the engineering is clean. The code is shared, the symbolic programs are transparent, and the PCEE values against manual annotation are genuinely low. The Q4 feedback correction loop is a nice touch, and the paper is honest about its limits (pathological speech errors, no audio-grounded correction).\n\nThe soft spot is exactly where the reader put it: the LLM-skeleton confound. QuaKE only tests monotonicities for descendants of CI in the graph, and that graph comes from an LLM, not the data. With 43% extra edges relative to the expert network, the candidate set — and therefore the nine 'additional influences' — is substantially determined by that prior. The post-hoc agreement by the expert co-authors is confirmation after the fact, not a pre-specified blind check. The mean-threshold discretization and the BDeu prior get no sensitivity analysis. That is the weakest part of the paper.\n\nI do not think it is fatal. The signs themselves are estimated from patient data via QuaKE; the claim is not defined into existence. The paper is transparent about the LLM's role, and the PC baseline (1/15 expert edges) shows why data-driven structure learning is not viable at this sample size. The authors are also candid about the limits of the feedback loop.\n\nOne thing neither the reader nor the stress-test flagged: Table 3 and the text disagree about the direction of the word-length effect in the phonemic task. The text says NeSyQuaKE infers a positive monotonic influence of CI on word length; the table marks that rule with an × (opposite found). That is an internal inconsistency a referee will want resolved.\n\nBottom line: well-built applied system, modest novelty, one load-bearing methodological weakness, and one internal contradiction. It deserves peer review — a good referee will ask for sensitivity analyses and a clearer separation of LLM prior from data-driven discovery, but it is not a desk reject. I would bring it to a reading group for the neurosymbolic evaluation discussion and would likely cite it as an example of a careful applied pipeline.","headline":"A clean applied pipeline whose 'novel' monotonicities are partly inherited from the LLM-supplied skeleton, but the work is honest, transparent, and worth a serious referee.","tokens_in":15480,"tokens_out":4936,"would_cite":true,"duration_ms":47250,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that qualitative clinical knowledge about Alzheimer's disease markers can be extracted end-to-end from raw audio recordings of verbal fluency tests, without manual transcription.","keywords":["Alzheimer's disease","verbal fluency","qualitative influence statements","Bayesian network","monotonicity","neurosymbolic","speech-to-text","knowledge discovery"],"falsifier":"Reconstruct the Bayesian network skeleton from expert elicitation (or from a data-driven algorithm with more data) and rerun the qualitative knowledge extraction; if the extracted monotonicity set changes substantially, then the paper's recovered knowledge is not intrinsic to the data. Alternatively, test on an independent cohort: if the nine discovered influences do not replicate (e.g., cognitive impairment has no negative influence on cluster size in the phonemic task), the claim of novel relationships would be weakened.","tokens_in":14640,"feed_emoji":"🧠","tokens_out":3195,"duration_ms":27404,"temperature":0.7,"pith_summary":"The paper claims that qualitative clinical knowledge about Alzheimer's disease markers can be extracted end-to-end from raw audio recordings of verbal fluency tests, without manual transcription. It introduces a pipeline that grounds audio into discrete linguistic features via foundation models and deterministic programs, learns a Bayesian Network over these features, and applies a qualitative reasoning algorithm to output monotonic influence statements linking cognitive impairment to markers. The authors report that the extracted statements match all expert-elicited relationships and add nine new ones that experts partially confirmed. If true, this would make large-scale, explainable early screening of cognitive decline practical.","feed_headline":"Audio-only pipeline recovers known Alzheimer's language markers","feed_subtitle":"Speech-to-text and symbolic reasoning extract nine new qualitative links from one-minute fluency tests.","key_machinery":"The central mechanism is the QuaKE (Qualitative Knowledge Extraction) algorithm, which checks for first-order stochastic dominance between the cognitive-impairment variable and each descendant feature in the network, computing a monotonic influence score. The network's structure is elicited from an LLM prompted repeatedly, with edges pooled by frequency; parameters are learned with a Bayesian estimator. Symbolic programs compute each clinical variable (word count, cluster size, speech rate, etc.) deterministically from structured LLM output, ensuring transparency.","core_discovery":"On a dataset of 162 verbal fluency recordings from patients with mild cognitive impairment and healthy controls, the pipeline matched every monotonic influence initially elicited from domain experts for both semantic and phonemic tasks, and discovered nine additional influences, six in phonemic and three in semantic. Three semantic discoveries (negative influence of CI on cluster size, speech rate, word length) were validated by experts; others were plausible but uncertain. The symbol grounding stage had low pairwise conditional estimation error, with a 94.3% improvement over a rule-based baseline in semantic fluency and 77.2% in phonemic.","pith_inferences":["Since the Bayesian network skeleton comes from an LLM and only 80% of expert edges are recovered, the extracted monotonicities are conditioned on that prior; a different LLM or expert graph could change the set or sign of discovered rules.","The mean-threshold discretization is coarse; using median or multi-threshold discretization might alter borderline monotonicities.","The feedback loop correcting mistranscriptions improved grounding but not the monotonicity set, so transcript corrections may be unnecessary for qualitative discovery.","The positive influence of CI on word length in the phonemic task, left uncertain by experts, is a concrete candidate for prospective study with a larger cohort and audio-level verification."],"forward_implications":["Early Alzheimer's screening could be automated from a one-minute speech test, removing the transcription bottleneck.","The method generates testable hypotheses about less obvious linguistic markers, such as word-length effects in phonemic fluency.","The influence of cognitive impairment can be quantified by degree, giving an ordinal ranking of marker sensitivity.","The neurosymbolic architecture generalizes to other clinical audio tasks where structured variables and qualitative relationships are of interest."],"fun_headline_variants":["Audio pipeline recovers Alzheimer's markers, finds nine new links","One-minute speech test reveals new Alzheimer's language markers","Neurosymbolic model rediscovers known AD markers, adds nine","Speech AI uncovers novel Alzheimer's markers, validates three"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole qualitative analysis rests on a Bayesian network skeleton elicited from a language model: if that graph has spurious or missing edges, the set and signs of the monotonicity statements are determined as much by the model's prior as by the clinical data.","fun_headline_variants_meta":{"raw":{"variants":["Audio pipeline recovers Alzheimer's markers, finds nine new links","One-minute speech test reveals new Alzheimer's language markers","Neurosymbolic model rediscovers known AD markers, adds nine","Speech AI uncovers novel Alzheimer's markers, validates three"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000253,"raw_usage":{"total_tokens":1317,"prompt_tokens":575,"completion_tokens":742,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":319,"completion_tokens_details":{"reasoning_tokens":684}},"tokens_in":319,"tokens_out":742,"duration_ms":6954,"temperature":1.0,"reasoning_tokens":684,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T05:11:49.549159+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reconstruct the Bayesian network skeleton from expert elicitation (or from a data-driven algorithm with more data) and rerun the qualitative knowledge extraction; if the extracted monotonicity set changes substantially, then the paper's recovered knowledge is not intrinsic to the data. Alternatively, test on an independent cohort: if the nine discovered influences do not replicate (e.g., cognitive impairment has no negative influence on cluster size in the phonemic task), the claim of novel relationships would be weakened.","supporting_citations":[],"review_version":1}