REVIEW 4 major objections 4 minor 3 cited by
Heart2Mind: Human-Centered Contestable Psychiatric Disorder Diagnosis System using Wearable ECG Monitors
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper aims to show that a chest-strap ECG, a time-frequency transformer, and a clinician chatbot can together screen for schizophrenia and bipolar disorder, with the model reaching 91.7% accuracy on a 60-person dataset and the chatbot…
desk verdict A convincing contestable-AI prototype whose headline accuracy is not yet shown to be per-patient; the SAE mechanism and public code are worth engaging. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Multi-Scale Temporal-Frequency Transformer (MSTFT), a classifier that reads 300-beat sliding windows of R-R intervals and passes them through parallel branches: dilated causal convolutions with stochastic skips for temporal patterns, and separable convolutions acting as learnable wavelet transforms for frequency patterns. A cross-attention block fuses the two branches by treating temporal features as queries and frequency features as keys and values, followed by a gated multi-head self-attention block and a pooled classification head. Around this model, the paper builds two further mechanisms: Self-Adversarial Explanations (SAEs), which align averaged attention maps with gradient-based maps via dynamic time warping and threshold their absolute difference to find regions of disagreement; and a contestable LLM prompt that receives the baseline prediction, whole-signal HRV metrics, and discrepancy-region HRV metrics so the chatbot can justify, retain, or overturn the diagnosis. The classifier is the engine, the SAE discrepancy count is the built-in safeguard, and the LLM is the channel through which clinicians exercise contestation.
What would settle it
Evaluate MSTFT at the patient level: take all windows from each of the 60 participants, form one prediction per participant by majority vote, and build the 60-subject confusion matrix. If patient-level accuracy is close to chance while window-level accuracy stays high, the central screening claim is refuted.
Extended reading notes
Core claim
The central claim is that R-R interval time series from a single-lead wearable ECG carry enough autonomic information to separate people with schizophrenia/bipolar disorder from healthy controls, and that this information can be presented to clinicians in a form they can examine and dispute. The paper reports that MSTFT reaches 91.7% accuracy, 0.963 precision, 0.867 recall, and 0.940 AUC under leave-one-out cross-validation on HRV-ACC, beating 1D-CNN, a plain Transformer, and the published results of prior methods. It further reports that correct predictions have few discrepancies between attention-based and gradient-based explanations (means of 0.72–0.78 regions) while incorrect predictions have many (means of 7.0–7.67 regions), so the discrepancy count can flag unreliable predictions during inference. Finally, all three tested large language models retained every one of the 54 correct baseline predictions, and they overturned one, one, and three of the six incorrect predictions respectively, with the strongest model correcting half of the baseline errors.
Load-bearing premise
The load-bearing premise is that the tens of thousands of heavily overlapping 300-beat windows cut from each person's 70–120 minute recording can be treated as roughly independent observations, so the 91.7% accuracy is a per-window number rather than a demonstrated per-patient diagnostic accuracy.
Editorial extensions
If this is right
- A consumer chest-strap ECG can support objective, continuous screening for schizophrenia and bipolar disorder outside the clinic, potentially shortening the path from symptoms to treatment.
- MSTFT provides a new reference result on the HRV-ACC dataset; future HRV-based psychiatric classification work will need to report comparable leave-one-out metrics to be measured against it.
- The discrepancy count between explanation types can act as an inference-time uncertainty flag, sending only suspicious predictions to human review instead of requiring every case to be checked.
- Contestable LLMs give clinicians a natural-language channel to validate or override model decisions, matching the regulatory push toward contestability in AI-assisted healthcare.
- Because the three LLMs reach at least one correct overturn without medical fine-tuning, domain-specific tuning or ensembling is a direct route to higher correction rates.
Reading between the lines
- A stricter patient-level test would aggregate each participant's window predictions (for example by majority vote) and report the 60-subject confusion matrix; the current paper reports window-level metrics only, so the patient-level diagnostic rate remains an open question.
- The SAE discrepancy mechanism could be reused as a label-free uncertainty estimator for other transformer classifiers on physiological time series, since it only needs attention weights and gradients.
- Because the three LLMs overturn different subsets of errors, an ensemble that combines their votes would plausibly overturn more than any single model; this is a natural extension the paper does not test.
- Adding other wearable streams, such as electrodermal activity or motion, could turn the binary screening into a finer distinction among schizophrenia, bipolar disorder, and healthy states; the paper identifies this as a direction but does not pursue it.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Heart2Mind, a full-stack system for psychiatric disorder screening that combines a wearable ECG/RRI monitoring interface (CMI), a Multi-Scale Temporal-Frequency Transformer (MSTFT) classifier, and a Contestable Diagnosis Interface (CDI) built on Self-Adversarial Explanations (SAEs) and contestable LLMs. The central technical claims are that MSTFT achieves 91.7% accuracy on the HRV-ACC dataset under leave-one-out cross-validation, outperforming published baselines, and that SAEs detect unreliable model predictions by comparing attention-based and gradient-based explanations, which then enables LLMs to validate correct predictions and contest incorrect ones. The manuscript includes detailed architecture descriptions, hyperparameters, prompt templates, case studies from LLM outputs, and a public code repository.
Significance. If the headline accuracy were demonstrably per-patient diagnostic accuracy, the system would make a meaningful contribution: it combines a novel time-frequency transformer for RRI classification, a concrete implementation of contestable AI using explanation discrepancies, and an interactive LLM-based clinician interface, all in a reproducible open-source form. The strengths of the paper are its systems-level integration, the delivery of a working artifact, and the explicit engagement with emerging regulatory notions of contestability. However, the statistical evaluation does not currently support the per-patient screening claim: the reported metrics are computed over heavily overlapping windows rather than over patients, and the SAE discrepancy threshold is fitted on the same cases used to demonstrate its utility. These issues put the central quantitative claims on uncertain ground and need to be resolved before the results can be accepted.
major comments (4)
- [Section 5.1, Eq. (41) and Table 4b] The headline 91.7% LOOCV accuracy is computed over sliding windows of length T=300 with stride 1, so successive test inputs overlap in 299 of 300 samples and each participant contributes thousands of near-duplicate windows. LOOCV holds out one participant, but the metrics in Table 4b are aggregated over windows, not over the 60 participants. The effective sample size is therefore 60, not the tens of thousands of windows, and the reported accuracy, precision, recall, F1, and AUC do not establish per-patient diagnostic accuracy. A per-participant evaluation (e.g., majority vote over each subject's windows, or a mixed-effects model with participant as a random effect) with confidence intervals is the load-bearing check required to support the abstract's screening claim.
- [Section 5.1, preprocessing before windowing] The manuscript states that 'before windowing, we rescaled each participant's RRI signal to zero mean and unit variance.' In LOOCV, the held-out subject's full-signal mean and standard deviation are therefore used to normalize that subject's test windows, meaning test-window values depend on statistics of the test subject's entire recording. This is a form of information leakage and can inflate performance measures. The authors should either use normalization statistics derived only from training subjects or provide evidence that the reported performance is insensitive to this choice.
- [Section 5.3.1, Eq. (32) and Figure 10] The SAE discrepancy-detection claim is partially fitted rather than validated. The threshold rho = 0.5 in Eq. (32) is set empirically, and the paper's key operational criterion—that roughly 5-6 discrepancy regions indicate unreliable predictions—is inferred from the discrepancy counts of the same MSTFT checkpoint's correct and incorrect outputs (Figure 10). This is circular when the same cases are then used to claim that SAEs flag unreliable predictions. A held-out validation set, or a pre-specified threshold selection procedure, is needed before the discrepancy count can be claimed as an inference-time uncertainty quantification mechanism.
- [Section 5.2.1 and Table 4b] Several baselines (Buza et al., Książek et al.) are imported directly from their original publications, while other baselines are re-implemented, without a common preprocessing and evaluation pipeline across all methods. No confidence intervals or significance tests accompany the comparisons in Table 4a or 4b. The claim of outperforming state-of-the-art methods is therefore not statistically substantiated; the authors should report paired per-participant comparisons, confidence intervals, or significance tests, ideally using a shared windowing and normalization protocol.
minor comments (4)
- [Section 5.2.3] The text states that precision of 0.963 means 'fewer than 4% of health controls were wrongly flagged as positive.' This is an incorrect interpretation: precision is the fraction of predicted positives that are true positives, not the false positive rate among controls. The sentence should be rephrased.
- [Section 4.2.6] The sentence 'These pooled representations are concatenated to form a comprehensive feature vector' appears twice in consecutive lines; one instance should be deleted.
- [Section 4.3.1, Eq. (30)] The use of Dynamic Time Warping to align attention-based and gradient-based explanations is described only in one line. It would clarify the method to state whether DTW is used for temporal alignment, how the aligned maps are returned to the original time grid, and whether the alignment affects the threshold comparison in Eq. (32).
- [Section 5.3.2 and Table 5] The contestable LLM evaluation rests on only 6 erroneous predictions (3 FN and 3 FP cases), and the conclusion that 'all three LLMs successfully contested at least one erroneous prediction' is drawn from this very small sample. The paper should explicitly caveat the statistical fragility of these numbers, including in the abstract's claim of 'successfully challenging 50% of erroneous ones.'
Circularity Check
SAE discrepancy-flagging rule is inferred from the same evaluation cases it claims to detect; the MSTFT accuracy itself is an empirical result, not circular.
-
fitted input called prediction
[Section 5.3.1, 'SAEs Evaluation', Figure 10 and following paragraph; see also Eq. 32 and Section 6.2.1.]
"The consistent correlation between discrepancy count and prediction accuracy established a clear threshold effect: when discrepancies exceeded approximately 5 to 6 regions, prediction reliability became substantially compromised. This observation had important implications for clinical implementation."
The '5 to 6 regions' threshold is read off Figure 10, which plots discrepancy counts against the correct/incorrect labels of the single best MSTFT checkpoint used throughout the SAE evaluation. The proposed flagging rule ('flag potentially unreliable predictions') is therefore a restatement of the empirical pattern in the same labeled cases, not an independent prediction. Eq. 32 also sets the discrepancy threshold rho 'empirically' to 0.5, so the region counts on which the 5-6 threshold depends are themselves fitted. The paper's own limitation statement (Section 6.2.1: 'determining optimal thresholds for flagging potentially unreliable predictions remains challenging...
full rationale
The headline MSTFT result (91.7% LOOCV accuracy, Table 4b) is not circular: leave-one-out cross-validation trains on 59 participants and tests on a held-out subject, and no label or fitted parameter enters the feature construction by definition. The sliding-window design (Eq. 41, T=300 with stride 1) and per-participant z-scoring before windowing raise important statistical validity concerns (non-independent windows and test-subject normalization statistics), but those are evaluation flaws, not reductions of the prediction to its inputs. The contestable LLM evaluation is small (6 error cases) but is an empirical case study, not a self-definitional derivation. No load-bearing self-citation was found: Nguyen et al. [67] appears in related work/Table 2, and the MSTFT architecture is presented with its own equations rather than imported by citation. The one genuine circular step is the SAE discrepancy-flagging threshold: it is inferred from the same TP/TN/FP/FN cases used to demonstrate SAE success, and the paper itself concedes calibration is still needed. Because this affects a secondary component claim (SAE uncertainty flagging) while the central diagnostic accuracy remains an independent empirical result, the overall circularity is partial.
Assumptions & free parameters
free parameters (5)
- Discrepancy threshold rho =
0.5
- Window length T =
300
- Sliding window stride =
1 (implicit)
- Discrepancy flag threshold count =
5-6 regions
- Stochastic skip survival probability p_s =
0.8
assumptions (4)
- domain assumption RRI/HRV recordings from a Polar H10 chest strap are sufficiently informative to separate schizophrenia/bipolar patients from healthy controls, and the HRV-ACC labels are correct.
- ad hoc to paper Agreement between attention-based and gradient-based explanations indicates faithful model decision-making.
- domain assumption LLMs can interpret HRV metrics and reach clinically meaningful final decisions without domain fine-tuning or explicit medical guidelines.
- standard math Standard deep learning results (softmax attention, backpropagation, cross-attention, dilated convolutions) hold as implemented.
Cite this review
Pith. "Pith review of Heart2Mind: Human-Centered Contestable Psychiatric Disorder Diagnosis System using Wearable ECG Monitors." pith.science (2026). https://pith.science/paper/G6Q4SQXN
@misc{pith2026250511612,
author = {Pith},
title = {Pith review of: Heart2Mind: Human-Centered Contestable Psychiatric Disorder Diagnosis System using Wearable ECG Monitors},
year = {2026},
howpublished = {\url{https://pith.science/paper/G6Q4SQXN}},
note = {Machine review of arXiv:2505.11612}
}
read the original abstract
Psychiatric disorders affect millions globally, yet their diagnosis faces significant challenges in clinical practice due to subjective assessments and accessibility concerns, leading to potential delays in treatment. To help address this issue, we present Heart2Mind, a human-centered contestable psychiatric disorder diagnosis system using wearable electrocardiogram (ECG) monitors. Our approach leverages cardiac biomarkers, particularly heart rate variability (HRV) and R-R intervals (RRI) time series, as objective indicators of autonomic dysfunction in psychiatric conditions. The system comprises three key components: (1) a Cardiac Monitoring Interface (CMI) for real-time data acquisition from Polar H9/H10 devices; (2) a Multi-Scale Temporal-Frequency Transformer (MSTFT) that processes RRI time series through integrated time-frequency domain analysis; (3) a Contestable Diagnosis Interface (CDI) combining Self-Adversarial Explanations (SAEs) with contestable Large Language Models (LLMs). Our MSTFT achieves 91.7% accuracy on the HRV-ACC dataset using leave-one-out cross-validation, outperforming state-of-the-art methods. SAEs successfully detect inconsistencies in model predictions by comparing attention-based and gradient-based explanations, while LLMs enable clinicians to validate correct predictions and contest erroneous ones. This work demonstrates the feasibility of combining wearable technology with Explainable Artificial Intelligence (XAI) and contestable LLMs to create a transparent, contestable system for psychiatric diagnosis that maintains clinical oversight while leveraging advanced AI capabilities. Our implementation is publicly available at: https://github.com/Analytics-Everywhere-Lab/heart2mind.
Figures
Figures from the paper (11 more)
Forward citations
Cited by 3 Pith papers
-
ConGaIT: A Clinician-Centered Dashboard for Contestable AI in Parkinson's Disease Care
The paper presents ConGaIT, a dashboard embedding contestable AI mechanisms for PD gait analysis, and reports a high contestability score based on an unvalidated proxy evaluation.
-
Privacy-Preserving Multi-Stage Fall Detection Framework with Semi-supervised Federated Learning and Robotic Vision Confirmation
A multi-stage fall detection system combining federated IMU classification, BLE localization, and robot vision claims 99.99% accuracy, but the combined accuracy calculation is mathematically invalid.
-
Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models
A six-stage multi-agent MLLM pipeline with reverse image search, metadata analysis, and fact-checking tools is demonstrated on a single Ukraine missile-strike video, with no quantitative evaluation.
Reference graph
Works this paper leans on
-
[1]
Abdelrahman Abouelenin, Atabak Ashfaq, Adam Atkinson, Hany Awadalla, Nguyen Bach, Jianmin Bao, Alon Benhaim, Martin Cai, Vishrav Chaudhary, Congcong Chen, et al. 2025. Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras.arXiv preprint arXiv:2503.01743(2025)
arXiv 2025
-
[2]
ARSS Izzatunnisa Ainunhusna, Achmad Rizal, and Sony Sumaryo. 2020. Bipolar disorder classification based on electrocardiogram signal using support vector machine.INTERNATIONAL JOURNAL OF SCIENTIFIC and TECHNOLOGY RESEARCH9, 1 (2020), 4
2020
-
[3]
Ahmed Shihab Albahri, Ali M Duhaim, Mohammed A Fadhel, Alhamzah Alnoor, Noor S Baqer, Laith Alzubaidi, Osamah Shihab Albahri, Abdullah Hussein Alamoodi, Jinshuai Bai, Asma Salhi, et al. 2023. A systematic review of trustworthy and explainable artificial intelligence in healthcare: Assessment of quality, bias risk, and data fusion.Information Fusion96 (202...
2023
-
[4]
Kars Alfrink, Ianus Keller, Gerd Kortuem, and Neelke Doorn. 2023. Contestable AI by design: Towards a framework.Minds and Machines33, 4 (2023), 613–639
2023
-
[5]
Navya Alugubelli, Hussam Abuissa, and Attila Roka. 2022. Wearable devices for remote monitoring of heart rate and heart rate variability—what we know and what is coming.Sensors22, 22 (2022), 8903
2022
-
[6]
2013.Diagnostic and statistical manual of mental disorders: DSM-5
DSMTF American Psychiatric Association, DS American Psychiatric Association, et al . 2013.Diagnostic and statistical manual of mental disorders: DSM-5. Vol. 5. American psychiatric association Washington, DC
2013
-
[7]
Nancy C Andreasen, Stephan Arndt, D Del Miller, Michael Flaum, and Peg Napoulos. 1995. Correlational studies of the Scale for the Assessment of Negative Symptoms and the Scale for the Assessment of Positive Symptoms: an overview and update.Psychopathology 28, 1 (1995), 7–17
1995
-
[8]
Javiera T Arias and César A Astudillo. 2023. Enhancing Schizophrenia Prediction Using Class Balancing and SHAP Explainability Techniques on EEG Data. In2023 IEEE 13th International Conference on Pattern Recognition Systems (ICPRS). IEEE, 1–5
2023
Show all 114 references
-
[9]
İsmail Baydili, Burak Tasci, and Gülay Tasci. 2025. Artificial Intelligence in Psychiatry: A Review of Biological and Behavioral Data Analyses.Diagnostics15, 4 (2025), 434
2025
-
[10]
Eduardo E Benarroch. 1993. The central autonomic network: functional organization, dysfunction, and perspective. InMayo Clinic Proceedings, Vol. 68. Elsevier, 988–1001
1993
-
[11]
Beatrice R Benjamin, Mathias Valstad, Torbjørn Elvsåshagen, Erik G Jönsson, Torgeir Moberget, Adriano Winterton, Marit Haram, Margrethe C Høegh, Trine V Lagerberg, Nils Eiel Steen, et al. 2021. Heart rate variability is associated with disease severity in psychosis spectrum di...
2021
-
[12]
Gary G Berntson, J Thomas Bigger Jr, Dwain L Eckberg, Paul Grossman, Peter G Kaufmann, Marek Malik, Haikady N Nagaraja, Stephen W Porges, J Philip Saul, Peter H Stone, et al . 1997. Heart rate variability: origins, methods, and interpretive caveats. Heart2Mind: Human-Centered ...
1997
-
[13]
Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller, and Wojciech Samek. 2016. Layer-wise relevance propagation for neural networks with local renormalization layers. InArtificial Neural Networks and Machine Learning–ICANN 2016: 25th International Co...
2016
-
[14]
Lokesh Boggavarapu, Vineet Srivastava, Amit Maheswar Varanasi, Yingda Lu, and Runa Bhaumik. 2024. Evaluating Enhanced LLMs for Precise Mental Health Diagnosis from Clinical Notes.medRxiv(2024), 2024–12
2024
-
[15]
Louise Brådvik. 2018. Suicide risk and mental disorders. 2028 pages
2018
-
[16]
Elvira Bramon and Pak C Sham. 2001. The common genetic liability between schizophrenia and bipolar disorder: a review.Current psychiatry reports3, 4 (2001), 332–337
2001
-
[17]
Krisztian Buza, Kamil Ksiazek, Wilhelm Masarczyk, Przemysław Głomb, Piotr Gorczyca, and Magdalena Piegza. 2023. A Simple and Effective Classifier for the Detection of Psychotic Disorders based on Heart Rate Variability Time Series. InWorkshop on Bioinformatics and Computationa...
2023
-
[18]
Health Canada. 2025. Pre-market guidance for machine learning-enabled medical devices. https://www.canada.ca/en/health- canada/services/drugs-health-products/medical-devices/application-information/guidance-documents/pre-market-guidance- machine-learning-enabled-medical-devices.html
2025
-
[19]
Matteo Cella, Łukasz Okruszek, Megan Lawrence, Valerio Zarlenga, Zhimin He, and Til Wykes. 2018. Using wearable technology to detect the autonomic signature of illness severity in schizophrenia.Schizophrenia research195 (2018), 537–542
2018
-
[20]
Aditya Chattopadhay, Anirban Sarkar, Prantik Howlader, and Vineeth N Balasubramanian. 2018. Grad-cam++: Generalized gradient- based visual explanations for deep convolutional networks. In2018 IEEE winter conference on applications of computer vision (W ACV). IEEE, 839–847
2018
-
[21]
Chun-Fu Richard Chen, Quanfu Fan, and Rameswar Panda. 2021. Crossvit: Cross-attention multi-scale vision transformer for image classification. InProceedings of the IEEE/CVF international conference on computer vision. 357–366
2021
-
[22]
Tianqi Chen and Carlos Guestrin. 2016. Xgboost: A scalable tree boosting system. InProceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining. 785–794
2016
-
[23]
Hao-Fei Cheng, Ruotong Wang, Zheng Zhang, Fiona O’connell, Terrance Gray, F Maxwell Harper, and Haiyi Zhu. 2019. Explaining decision-making algorithms through UI: Strategies to help non-expert stakeholders. InProceedings of the 2019 chi conference on human factors in computing...
2019
-
[24]
Filippo Corponi, Bryan M Li, Gerard Anmella, Clàudia Valenzuela-Pascual, Isabella Pacchiarotti, Marc Valentí, Iria Grande, Antonio Benabarre, Marina Garriga, Eduard Vieta, et al. 2024. A Bayesian analysis of heart rate variability changes over acute episodes of bipolar disorde...
2024
-
[25]
Julia Cullen, Matthew J Reed, Alexandra Muir, Ross Murphy, Valery Pollard, Goran Zangana, Sean Krupej, Sylvia Askham, Patricia Holdsworth, and Lauren Davies. 2021. Experience of a smartphone ambulatory ECG clinic for emergency department patients with palpitation: a single-cen...
2021
-
[26]
Richard Dazeley, Peter Vamplew, Cameron Foale, Charlotte Young, Sunil Aryal, and Francisco Cruz. 2021. Levels of explainable artificial intelligence for human-aligned conversational explanations.Artificial Intelligence299 (2021), 103525
2021
-
[27]
Rui Duarte, Angela Stainthorpe, James Mahon, Janette Greenhalgh, Marty Richardson, Sarah Nevitt, Eleanor Kotas, Angela Boland, Howard Thom, Tom Marshall, et al. 2019. Lead-I ECG for detecting atrial fibrillation in patients attending primary care with an irregular pulse using ...
2019
-
[28]
Maria Faurholt-Jepsen, Lars Vedel Kessing, and Klaus Munkholm. 2017. Heart rate variability in bipolar disorder: A systematic review and meta-analysis.Neuroscience & Biobehavioral Reviews73 (2017), 68–80
2017
-
[29]
Food and Drug Administration
U.S. Food and Drug Administration. 2025. Clinical Decision Support Software - Guidance. https://www.fda.gov/regulatory-information/ search-fda-guidance-documents/clinical-decision-support-software
2025
-
[30]
Nils Freyer, Dominik Groß, and Myriam Lipprandt. 2024. The ethical requirement of explainability for AI-DSS in healthcare: a systematic review of reasons.BMC Medical Ethics25, 1 (2024), 104
2024
-
[31]
Maurizio Garbarino, Matteo Lai, Dan Bender, Rosalind W Picard, and Simone Tognetti. 2014. Empatica E3—A wearable wireless multi-sensor device for real-time computerized biofeedback and data acquisition. In2014 4th international conference on wireless mobile communication and h...
2014
-
[32]
MD Hill, SS Gill, H Le-Niculescu, O MacKie, R Bhagar, K Roseberry, OK Murray, HD Dainton, SK Wolf, A Shekhar, et al. 2024. Precision medicine for psychotic disorders: objective assessment, risk prediction, and pharmacogenomics.Molecular Psychiatry(2024), 1–22
2024
-
[33]
Katrina Hinde, Graham White, and Nicola Armstrong. 2021. Wearable devices suitable for monitoring twenty four hour heart rate variability in military populations.Sensors21, 4 (2021), 1061
2021
-
[34]
Tad Hirsch, Kritzia Merced, Shrikanth Narayanan, Zac E Imel, and David C Atkins. 2017. Designing contestability: Interaction design, machine learning, and mental health. InProceedings of the 2017 Conference on Designing Interactive Systems. 95–99. 111:38•Nguyen et al
2017
-
[35]
Oliver D Howes, Connor Cummings, George E Chapman, and Ekaterina Shatalina. 2023. Neuroimaging in schizophrenia: an overview of findings and their implications for synaptic changes.Neuropsychopharmacology48, 1 (2023), 151–167
2023
-
[36]
Tomoko Inoue, Toshikazu Shinba, Masanari Itokawa, Guanghao Sun, Maho Nishikawa, Mitsuhiro Miyashita, Kazuhiro Suzuki, Nobutoshi Kariya, Makoto Arai, and Takemi Matsui. 2022. The development and clinical application of a novel schizophrenia screening system using yoga-induced a...
2022
-
[37]
Sarthak Jain and Byron C Wallace. 2019. Attention is not Explanation. InProceedings of NAACL-HLT. 3543–3556
2019
-
[38]
Carmen Jimenez-Mesa, Javier Ramirez, Zhenghui Yi, Chao Yan, Raymond Chan, Graham K Murray, Juan Manuel Gorriz, and John Suckling. 2024. Machine learning in small sample neuroimaging studies: Novel measures for schizophrenia analysis.Human Brain Mapping45, 5 (2024), e26555
2024
-
[39]
Amir-Hossein Karimi, Gilles Barthe, Borja Balle, and Isabel Valera. 2020. Model-agnostic counterfactual explanations for consequential decisions. InInternational conference on artificial intelligence and statistics. PMLR, 895–905
2020
-
[40]
Smith K Khare, Vikram M Gadre, and U Rajendra Acharya. 2023. ECGPsychNet: An optimized hybrid ensemble model for automatic detection of psychiatric disorders using ECG signals.Physiological Measurement44, 11 (2023), 115004
2023
-
[41]
Hye-Geum Kim, Eun-Jin Cheon, Dai-Seg Bai, Young Hwan Lee, and Bon-Hoon Koo. 2018. Stress and heart rate variability: a meta-analysis and review of the literature.Psychiatry investigation15, 3 (2018), 235
2018
-
[42]
Joel EW Koh, Chui Ping Ooi, Nikki SJ Lim-Ashworth, Jahmunah Vicnesh, Hui Tian Tor, Oh Shu Lih, Ru-San Tan, U Rajendra Acharya, and Daniel Shuen Sheng Fung. 2022. Automated classification of attention deficit hyperactivity disorder and conduct disorder using entropy features wi...
2022
-
[43]
Xiangwei Kong, Shujie Liu, and Luhao Zhu. 2024. Toward Human-centered XAI in Practice: A survey.Machine Intelligence Research21, 4 (2024), 740–770
2024
-
[44]
Kamil Michał Książek, Wilhelm Masarczyk, Przemysław Głomb, Michał Romaszewski, Krisztian Buza, Przemysław Sekuła, Michał Cholewa, Katarzyna Kołodziej, Piotr Gorczyca, and Magdalena Piegza. 2025. Deep learning approach for automatic assessment of schizophrenia and bipolar disor...
2025
-
[45]
Kamil Michał Książek, Wilhelm Masarczyk, Przemysław Głomb, Michał Romaszewski, Iga Stokłosa, Piotr Ścisło, Paweł Dębski, Robert Pudlo, Krisztián Buza, Piotr Gorczyca, et al . 2023. The analysis of heart rate variability and accelerometer mobility data in the assessment of symp...
2023
-
[46]
K Ksikażek, W Masarczyk, P Głomb, M Romaszewski, I Stokłosa, P Ścisło, P Dkebski, R Pudlo, P Gorczyca, and M Piegza. [n. d.]. HRV-ACC: a dataset with RR intervals and accelerometer data for the diagnosis of psychotic disorders using a Polar H10 wearable sensor, 2023b.URL https...
-
[47]
Francesco Leofante, Hamed Ayoobi, Adam Dejl, Gabriel Freedman, Deniz Gorur, Junqi Jiang, Guilherme Paulino-Passos, Antonio Rago, Anna Rapberger, Fabrizio Russo, et al. 2024. Contestable AI needs computational argumentation. InProceedings of the 21st International Conference on...
2024
-
[48]
Hui Li and Xiao-Jun Wu. 2024. CrossFuse: A novel cross attention mechanism based infrared and visible image fusion approach. Information Fusion103 (2024), 102147
2024
-
[49]
Xiangchen Li, Yuting Song, Huang Wang, Xinyu Su, Mengyao Wang, Jing Li, Zhiqiang Ren, Daidi Zhong, and Zhiyong Huang. 2023. Evaluation of measurement accuracy of wearable devices for heart rate variability.Iscience26, 11 (2023)
2023
-
[50]
Percy Liang et al . 2023. Holistic Evaluation of Language Models.Transactions on Machine Learning Research(2023). Featured Certification, Expert Certification
2023
-
[51]
Paul Lichtenstein, Benjamin H Yip, Camilla Björk, Yudi Pawitan, Tyrone D Cannon, Patrick F Sullivan, and Christina M Hultman. 2009. Common genetic determinants of schizophrenia and bipolar disorder in Swedish families: a population-based study.The Lancet373, 9659 (2009), 234–239
2009
-
[52]
Yibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong, Jing Li, and Shiqi Wang. 2022. Rethinking attention-model explainability through faithfulness violation test. InInternational conference on machine learning. PMLR, 13807–13824
2022
-
[53]
Scott M Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions.Advances in neural information processing systems30 (2017)
2017
-
[54]
Henrietta Lyons, Eduardo Velloso, and Tim Miller. 2021. Conceptualising contestability: Perspectives on contesting algorithmic decisions.Proceedings of the ACM on Human-Computer Interaction5, CSCW1 (2021), 1–25
2021
-
[55]
Gennie Mansi, Naveena Karusala, and Mark Riedl. 2025. Legally-Informed Explainable AI.arXiv preprint arXiv:2504.10708(2025)
2025 arXiv
-
[56]
John McGrath, Sukanta Saha, David Chant, Joy Welham, et al. 2008. Schizophrenia: a concise overview of incidence, prevalence, and mortality.Epidemiologic reviews30, 1 (2008), 67–76
2008
-
[57]
Kathleen R Merikangas, Hagop S Akiskal, Jules Angst, Paul E Greenberg, Robert MA Hirschfeld, Maria Petukhova, and Ronald C Kessler. 2007. Lifetime and 12-month prevalence of bipolar spectrum disorder in the National Comorbidity Survey replication.Archives of general psychiatry...
2007
-
[58]
AI Meta. 2025. The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, April 2025. Heart2Mind: Human-Centered Contestable Psychiatric Disorder Diagnosis System using Wearable ECG Monitors•111:39
2025
-
[59]
Jacqueline Michelle Metsch, Anna Saranti, Alessa Angerschmid, Bastian Pfeifer, Vanessa Klemt, Andreas Holzinger, and Anne-Christin Hauschild. 2024. CLARUS: an interactive explainable AI platform for manual counterfactuals in graph neural networks.Journal of Biomedical Informat...
2024
-
[60]
Tim Miller. 2019. Explanation in artificial intelligence: Insights from the social sciences.Artificial intelligence267 (2019), 1–38
2019
-
[61]
Muzafar Mehraj Misgar and MPS Bhatia. 2024. Unveiling psychotic disorder patterns: A deep learning model analysing motor activity time-series data with explainable AI.Biomedical Signal Processing and Control91 (2024), 106000
2024
-
[62]
Nicola Montano, Alberto Porta, Chiara Cogliati, Giorgio Costantino, Eleonora Tobaldini, Karina Rabello Casali, and Ferdinando Iellamo
-
[63]
Mahdieh Montazeri, Mitra Montazeri, Kambiz Bahaadinbeigy, Mohadeseh Montazeri, and Ali Afraz. 2023. Application of machine learning methods in predicting schizophrenia and bipolar disorders: A systematic review.Health Science Reports6, 1 (2023), e962
2023
-
[64]
Ramaravind K Mothilal, Amit Sharma, and Chenhao Tan. 2020. Explaining machine learning classifiers through diverse counterfactual explanations. InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency. 607–617
2020
-
[65]
RJ Neuwirth. 2022. The EU Artificial Intelligence Act.The EU Artificial Intelligence Act106 (2022)
2022
-
[66]
Hung Nguyen, Tobias Clement, Loc Nguyen, Nils Kemmerzell, Binh Truong, Khang Nguyen, Mohamed Abdelaal, and Hung Cao. 2024. LangXAI: Integrating Large Vision Models for Generating Textual Explanations to Enhance Explainability in Visual Perception Tasks. InProceedings of the Th...
2024
-
[67]
Hung Nguyen, Alireza Rahimi, Veronica Whitford, Hélène Fournier, Irina Kondratova, René Richard, and Hung Cao. 2025. Human- centered Explainable Psychiatric Disorder Diagnosis System using Wearable ECG Monitors. InThe 29th Pacific-Asia Conference on Knowledge Discovery and Dat...
2025
-
[68]
Hung Truong Thanh Nguyen, Loc Phuc Truong Nguyen, and Hung Cao. 2024. XEdgeAI: A human-centered industrial inspection framework with data-centric Explainable Edge AI approach.Information Fusion(2024), 102782
2024
-
[69]
Phong X Nguyen, Hung Q Cao, Khang VT Nguyen, Hung Nguyen, and Takehisa Yairi. 2022. Secam: Tightly accelerate the image explanation via region-based segmentation.IEICE TRANSACTIONS on Information and Systems105, 8 (2022), 1401–1417
2022
-
[70]
Quoc Khanh Nguyen, Truong Thanh Hung Nguyen, Vo Thanh Khang Nguyen, Van Binh Truong, Tuong Phan, and Hung Cao
-
[71]
Truong Thanh Hung Nguyen, Van Binh Truong, Vo Thanh Khang Nguyen, Quoc Hung Cao, and Quoc Khanh Nguyen. 2023. Towards trust of explainable ai in thyroid nodule diagnosis. InInternational Workshop on Health Intelligence. Springer, 11–26
2023
-
[72]
Goverment of Canada. 2019. Directive on Automated Decision-Making- Canada.ca. https://www.tbs-sct.canada.ca/pol/doc-eng.aspx? id=32592
2019
-
[73]
Montréal Declaration on Responsible AI. 2023. Montréal Declaration on Responsible AI. https://montrealdeclaration-responsibleai.com/
2023
-
[74]
Michael J Owen, Sophie E Legge, Elliott Rees, James TR Walters, and Michael C O’Donovan. 2023. Genomic findings in schizophrenia and their implications.Molecular psychiatry28, 9 (2023), 3638–3647
2023
-
[75]
Cristiano Patrício, Isabel Rio-Torto, Jaime S Cardoso, Luís F Teixeira, and João C Neves. 2025. CBVLM: Training-free Explainable Concept-based Large Vision Language Models for Medical Image Classification.arXiv preprint arXiv:2501.12266(2025)
2025
-
[76]
Carsten Bøcker Pedersen, Ole Mors, Aksel Bertelsen, Berit Lindum Waltoft, Esben Agerbo, John J McGrath, Preben Bo Mortensen, and William W Eaton. 2014. A comprehensive nationwide study of the incidence rate and lifetime risk for treated mental disorders.JAMA psychiatry71, 5 (2...
2014
-
[77]
Vitali Petsiuk, Abir Das, and Kate Saenko. 2018. Rise: Randomized input sampling for explanation of black-box models.arXiv preprint arXiv:1806.07421(2018)
2018 arXiv
-
[78]
Vitali Petsiuk, Rajiv Jain, Varun Manjunatha, Vlad I Morariu, Ashutosh Mehra, Vicente Ordonez, and Kate Saenko. 2021. Black-box explanation of object detectors via saliency maps. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11443–11452
2021
-
[79]
Thomas Ploug and Søren Holm. 2020. The four dimensions of contestable AI diagnostics-A patient-centric approach to explainable AI. Artificial intelligence in medicine107 (2020), 101901
2020
-
[80]
Medicines & Healthcare products Regulatory Agency. 2024. Transparency for machine learning-enabled medical devices: guiding principles. https://www.gov.uk/government/publications/machine-learning-medical-devices-transparency-principles/transparency- for-machine-learning-enable...
2024
-
[81]
U Rajendra Acharya, K Paul Joseph, Natarajan Kannathal, Choo Min Lim, and Jasjit S Suri. 2006. Heart rate variability: a review. Medical and biological engineering and computing44 (2006), 1031–1051
2006
-
[82]
Ashvita Ramesh, Tanvi Nayak, Molly Beestrum, Giorgio Quer, and Jay A Pandit. 2023. Heart Rate Variability in Psychiatric Disorders: A Systematic Review.Neuropsychiatric Disease and Treatment(2023), 2217–2239
2023
-
[83]
Kavita Rawat and Trapti Sharma. 2025. PsyneuroNet architecture for multi-class prediction of neurological disorders.Biomedical Signal Processing and Control100 (2025), 107080. 111:40•Nguyen et al
2025
-
[84]
Protection Regulation. 2016. Regulation (EU) 2016/679 of the European Parliament and of the Council.Regulation (eu)679 (2016), 2016
2016
-
[85]
Why should i trust you?
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. “Why should i trust you?” Explaining the predictions of any classifier. InProceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 1135–1144
2016
-
[86]
Justus Robertson, Athanasios Vasileios Kokkinakis, Jonathan Hook, Ben Kirman, Florian Block, Marian F Ursu, Sagarika Patra, Simon Demediuk, Anders Drachen, and Oluseyi Olarewaju. 2021. Wait, but why?: assessing behavior explanation strategies for real-time strategy games. InPr...
2021
-
[87]
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2020. Grad-CAM: visual explanations from deep networks via gradient-based localization.International journal of computer vision128 (2020), 336–359
2020
-
[88]
Andrea Sgoifo, Nicola Montano, Carol Shively, Julian Thayer, and Andrew Steptoe. 2009. The inevitable link between heart and behavior. New insights from biomedical research and implications for clinical practice. (2009)
2009
-
[89]
Alessandro Silvani, Giovanna Calandra-Buonaura, Roger AL Dampney, and Pietro Cortelli. 2016. Brain–heart interactions: physiology and clinical implications.Philosophical transactions of the royal society A: Mathematical, physical and engineering sciences374, 2067 (2016), 20150181
2016
-
[90]
Ulf Simonsen, Frank Holden Christensen, and Niels Henrik Buus. 2009. The effect of tempol on endothelium-dependent vasodilatation and blood pressure.Pharmacology & therapeutics122, 2 (2009), 109–124
2009
-
[91]
K Simonyan, A Vedaldi, and A Zisserman. 2014. Deep inside convolutional networks: visualising image classification models and saliency maps. InProceedings of the International Conference on Learning Representations (ICLR). ICLR
2014
-
[92]
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023. Large language models encode clinical knowledge.Nature620, 7972 (2023), 172–180
2023
- [93]
-
[94]
Kathryn E Speer, Stuart Semple, Nenad Naumovski, and Andrew J McKune. 2020. Measuring heart rate variability using commercially available devices in healthy children: A validity and reliability study.European Journal of Investigation in Health, Psychology and Education10, 1 (2...
2020
-
[95]
Andrea Stautland, Ole B Fasmer, Petter Jakobsen, Berge Osnes, Jim Torresen, Tine Nordgreen, and Ketil J Oedegaard. 2022. P178. heart rate variability as biomarker for bipolar mania.Biological Psychiatry91, 9 (2022), S159
2022
-
[96]
Burak Tasci, Gulay Tasci, Sengul Dogan, and Turker Tuncer. 2024. A novel ternary pattern-based automatic psychiatric disorders classification using ECG signals.Cognitive Neurodynamics18, 1 (2024), 95–108
2024
-
[97]
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, et al. 2025. Gemma 3 technical report.arXiv preprint arXiv:2503.19786(2025)
2025 arXiv
-
[98]
Hardik Telangore, Nishant Sharma, Manish Sharma, and U Rajendra Acharya. 2025. A novel ECG-based approach for classifying psychiatric disorders: Leveraging wavelet scattering networks.Medical Engineering & Physics135 (2025), 104275
2025
-
[99]
Erhan Tiryaki, Akshay Sonawane, and Lakshman Tamil. 2021. Real-time CNN based ST depression episode detection using single-lead ECG. In2021 22nd International Symposium on Quality Electronic Design (ISQED). IEEE, 566–570
2021
-
[100]
Van Binh Truong, Truong Thanh Hung Nguyen, Vo Thanh Khang Nguyen, Quoc Khanh Nguyen, and Quoc Hung Cao. 2024. Towards Better Explanations for Object Detection. InProceedings of the 15th Asian Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 222), ...
2024
-
[101]
Gaetano Valenza, Mimma Nardelli, Antonio Lanata, Claudio Gentili, Gilles Bertschy, Markus Kosel, and Enzo Pasquale Scilingo. 2016. Predicting mood changes in bipolar disorder through heartbeat nonlinear dynamics.IEEE journal of biomedical and health informatics 20, 4 (2016), 1034–1043
2016
-
[102]
Gioacchino D De Sario Velasquez, Sahar Borna, Michael J Maniaci, Jordan D Coffey, Clifton R Haider, Bart M Demaerschalk, and Antonio Jorge Forte. 2024. Economic perspective of the use of wearables in health care: a systematic review.Mayo Clinic Proceedings: Digital Health2, 3 ...
2024
-
[103]
Chang Wang, Yaning Ren, Rui Zhang, Chen Wang, Xiangying Ran, Jiefen Shen, Zongya Zhao, Wei Tao, Yongfeng Yang, Wenjie Ren, et al. 2024. Schizophrenia classification and abnormalities reveal of brain region functional connection by deep-learning multiple sparsely connected netw...
2024
-
[104]
Danding Wang, Qian Yang, Ashraf Abdul, and Brian Y Lim. 2019. Designing theory-driven user-centric explainable AI. InProceedings of the 2019 CHI conference on human factors in computing systems. 1–15
2019
-
[105]
Zuxing Wang, Yazhu Zou, Jingwen Liu, Wei Peng, Mingmei Li, and Zhili Zou. 2025. Heart rate variability in mental disorders: an umbrella review of meta-analyses.Translational Psychiatry15, 1 (2025), 104
2025
-
[106]
I WHO. 2007. International classification of diseases (ICD)
2007
-
[107]
Yankun Wu, Yun-Ai Su, Linlin Zhu, Jitao Li, and Tianmei Si. 2024. Advances in functional MRI research in bipolar disorder: from the perspective of mood states.General Psychiatry37, 1 (2024), e101398
2024
-
[108]
Xiaohan Zang, Baimin Li, Lulu Zhao, Dandan Yan, and Licai Yang. 2022. End-to-end depression recognition based on a one-dimensional convolution neural network model using two-lead ECG signal.Journal of Medical and Biological Engineering42, 2 (2022), 225–233. Heart2Mind: Human-C...
2022
-
[109]
Matthew D Zeiler and Rob Fergus. 2014. Visualizing and understanding convolutional networks. InComputer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part I 13. Springer, 818–833
2014
-
[110]
Tian Hong Zhang, Xiao Chen Tang, Li Hua Xu, Yan Yan Wei, Ye Gang Hu, Hui Ru Cui, Ying Ying Tang, Tao Chen, Chun Bo Li, Lin Lin Zhou, et al. 2022. Imbalance model of heart rate variability and pulse wave velocity in psychotic and nonpsychotic disorders. Schizophrenia Bulletin48...
2022
-
[111]
Wencan Zhang and Brian Y Lim. 2022. Towards relatable explainable AI with the perceptual process. InProceedings of the 2022 CHI conference on human factors in computing systems. 1–24
2022
-
[112]
Bojan Žlahtič, Jernej Završnik, Helena Blažun Vošner, Peter Kokol, David Šuran, and Tadej Završnik. 2023. Agile Machine Learning Model Development Using Data Canyons in Medicine: A Step towards Explainable Artificial Intelligence and Flexible Expert-Based Model Improvement.App...
2023
-
[2009]
Heart rate variability explored in the frequency domain: a tool to investigate the link between heart and behavior.Neuroscience & Biobehavioral Reviews33, 2 (2009), 71–80
2009
-
[2024]
Efficient and Concise Explanations for Object Detection with Gaussian-Class Activation Mapping Explainer.arXiv preprint arXiv:2404.13417(2024)
2024 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.