REVIEW 4 major objections 3 minor 27 references
ConGaIT: A Clinician-Centered Dashboard for Contestable AI in Parkinson's Disease Care
T0 review · 4 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read ConGaIT demonstrates that contestable AI can be operationalized through a clinician-centered dashboard, with a Contestability Assessment Score of 0.970.
desk verdict A useful design pattern for contestable clinical dashboards, but the 0.970 CAS is a self-assessment plus LLM personas, not evidence that contestability works with clinicians. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Contest & Justify interaction pattern: a structured disagreement workflow with three argument types—Factual Error, Normative Conflict, and Reasoning Flaw—that a clinician can raise against a predictive insight. Behind it sit three supporting components: a CNN classifier that predicts Hoehn and Yahr stage from 10-second gait windows, layer-wise relevance propagation highlighting the sensors and time segments that drove the prediction, and a vision language model that produces normative justifications. A role-based feedback layer and an immutable audit trail complete the loop, and the CAS rubric turns contestability into a weighted, scored quantity.
What would settle it
Run the same CAS evaluation with, say, ten practising movement-disorder neurologists using ConGaIT on real gait cases; if their average Explanation Quality or Ease of Contestation scores pull the total CAS materially below the reported 0.970, the proxy-persona assumption and the operationalization claim would not generalize.
Extended reading notes
Core claim
The central discovery is that contestability can be operationalized as a measurable property of a user interface. ConGaIT demonstrates this through a tightly integrated dashboard in which every AI prediction carries visual explanations (layer-wise relevance propagation), a textual normative justification, and a structured path to dispute that justification. The paper's measurement is the Contestability Assessment Score, whose eight weighted criteria span human-centered, technical, legal, and organizational pillars; ConGaIT attains 0.970, with the only partial scores in traceability, ease of contestation, and explanation quality. The authors conclude that contestability can be embedded early in the design process and that their proxy evaluation with personas offers a reproducible template for doing so.
Load-bearing premise
The whole evaluation leans on the assumption that three large-language-model-generated clinician personas judge the interface the way real clinicians would.
Editorial extensions
If this is right
- Concrete interface patterns, not just model transparency, can satisfy contestability requirements in clinical AI.
- Structured argument types give clinicians a precise vocabulary for disagreement, making objections auditable and reusable for model refinement.
- A CAS score of 0.970 provides a benchmark that other clinical AI systems can be measured against on contestability.
- Combining visual and textual explanations with a contest flow addresses both explanation and justification, which the paper treats as distinct normative requirements.
- The proxy evaluation with language-model personas offers a reproducible early-design assessment, though it awaits clinical validation.
Reading between the lines
- The same Contest & Justify pattern could extend to other high-stakes settings—radiology, pathology, or automated hiring—where a human must legally or ethically be able to challenge an algorithmic decision.
- A natural next test is whether real clinicians' CAS scores and contestation behavior match the persona simulation; if they diverge, the pattern may need tailoring to actual clinical work practices.
- Because the audit trail and argument types are structured, the dashboard could feed contested cases back into model training, turning disagreement into a data asset.
- The CAS rubric's weighting is taken as given; a different weighting scheme could change which design features dominate the score.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents ConGaIT, a clinician-centered dashboard for Parkinson's disease gait monitoring that embeds contestable AI features: interactive visual explanations (LRP), role-based feedback, an immutable audit log, and a "Contest & Justify" workflow. The system is evaluated using the Contestability Assessment Score (CAS), where six observable criteria are scored by the authors through direct inspection and two subjective criteria are scored by LLM-generated clinician personas (GPT-4 and TinyTroupe), yielding a total CAS of 0.970. The paper claims this score demonstrates that contestability can be operationalized through human-centered design in compliance with emerging regulatory standards.
Significance. If substantiated, the work would offer a concrete interface pattern for contestable clinical AI and a systematic use of the CAS rubric, which is relevant to EU AI Act and GDPR Article 22 compliance. The paper also provides a public demo repository. However, the current evaluation is not independent or validated, so the significance is conditional on a substantially revised evidence base or a clearly repositioned claim.
major comments (4)
- [Section IV, Table I] The central claim of a 0.970 CAS is not supported by the evidence presented. Six of the eight criteria are scored by the authors "through direct inspection," and the remaining two are scored by LLM-generated personas rather than real clinicians; no inter-rater reliability or uncertainty is reported. The paper itself states in Section IV and the Conclusion that clinical usability testing and validation remain future work. Consequently, the abstract's claim that the score "demonstrates that contestability can be operationalized through human-centered design" is an overstatement that should be replaced with a claim about a preliminary or self-assessed evaluation, or supplemented with real clinician assessment.
- [Section IV, Discussion] The sentence describing the result as benefiting from "objective validation through CAS" is internally inconsistent with the preceding description of direct inspection and persona-based scoring. This contradiction must be corrected, since the majority of the final score rests on subjective, self-assigned judgments rather than objective measurement.
- [Table I] The scoring rationale is not reported. For each criterion, the authors should specify what interface evidence justifies the assigned score, why Traceability receives 9 out of 10 rather than a different value, and what rubric or behavioral definition was used to convert observations into scores. Without this, the numerical result is not verifiable or reproducible, which is especially important given the paper's emphasis on reproducibility.
- [Section IV, Conclusion] The use of LLM-generated clinician personas as a proxy for clinician judgment is a load-bearing assumption that is explicitly acknowledged as unvalidated. The paper concludes that contestability "can be embedded early in the design process" on the basis of this proxy, but no evidence is given that the personas reproduce real clinician evaluations. The conclusion should be downgraded to an illustration or pilot finding, and the proxy limitation should be made prominent in the abstract.
minor comments (3)
- [Abstract and Section III] The system name is written inconsistently as "Con-GaIT" in the abstract and "ConGaIT" in the body; please unify the spelling.
- [Section IV] The phrase "structured personas enabled reproducible, expert-informed evaluation" is misleading, since the personas are simulated and not validated as experts; consider rewording to "simulated expert-like personas."
- [Table I] The column header "Property Max" is unclear; it should be split into "Property" and "Maximum score" to avoid ambiguity with the weights column.
Circularity Check
No significant circularity: the CAS score is an external rubric and an arithmetic aggregation, not a prediction forced by construction.
full rationale
The central claim is an empirical evaluation claim: ConGaIT is designed to support contestability and then scored with the Contestability Assessment Score, whose source [7] is an external group with no overlap with the authors. The reported 0.970 is a weighted sum of criterion scores (Table I), where six criteria were assigned by direct inspection and two by LLM-simulated clinician personas. Nothing in the paper defines a design feature in terms of the score, and the score is not a fitted parameter later renamed as a prediction. The closest concern is that self-assessment and GPT-4/TinyTroupe personas could introduce favorable bias, but the paper itself flags this as a proxy: "future work will focus on clinical usability testing" and the conclusion calls it a "novel proxy evaluation." That is a validity and independence limitation, not a circular reduction under the patterns considered here. The authors' prior XAI papers are cited only as related work and are not load-bearing for the CAS result. Therefore the derivation chain is not circular.
Assumptions & free parameters
free parameters (3)
- Explanation Quality score (sp) =
42/50
- Traceability score (sp) =
9/10
- Observable criteria scores (Explainability, Openness, Safeguards, Adaptivity, Auditing) =
full marks
assumptions (3)
- domain assumption The Contestability Assessment Score (CAS) is a valid and reliable measure of contestability in clinical AI systems.
- ad hoc to paper LLM-generated clinician personas produce evaluations equivalent to those of real clinicians for subjective criteria.
- domain assumption The PhysioNet gait dataset and the CNN classifier trained on it are representative of real clinical gait data.
Cite this review
Pith. "Pith review of ConGaIT: A Clinician-Centered Dashboard for Contestable AI in Parkinson's Disease Care." pith.science (2026). https://pith.science/paper/YEZVAXWO
@misc{pith2026250722300,
author = {Pith},
title = {Pith review of: ConGaIT: A Clinician-Centered Dashboard for Contestable AI in Parkinson's Disease Care},
year = {2026},
howpublished = {\url{https://pith.science/paper/YEZVAXWO}},
note = {Machine review of arXiv:2507.22300}
}
read the original abstract
AI-assisted gait analysis holds promise for improving Parkinson's Disease (PD) care, but current clinical dashboards lack transparency and offer no meaningful way for clinicians to interrogate or contest AI decisions. We present Con-GaIT (Contestable Gait Interpretation & Tracking), a clinician-centered system that advances Contestable AI through a tightly integrated interface designed for interpretability, oversight, and procedural recourse. Grounded in HCI principles, ConGaIT enables structured disagreement via a novel Contest & Justify interaction pattern, supported by visual explanations, role-based feedback, and traceable justification logs. Evaluated using the Contestability Assessment Score (CAS), the framework achieves a score of 0.970, demonstrating that contestability can be operationalized through human-centered design in compliance with emerging regulatory standards. A demonstration of the framework is available at https://github.com/hungdothanh/Con-GaIT.
Figures
Reference graph
Works this paper leans on
-
[1]
Parkinson’s disease: clinical features and diagnosis,
J. Jankovic, “Parkinson’s disease: clinical features and diagnosis,” Jour- nal of neurology, neurosurgery & psychiatry, vol. 79, no. 4, pp. 368–376, 2008
work page 2008
-
[2]
W. Zhang, Y . Ling, Z. Chen, K. Ren, S. Chen, P. Huang, and Y . Tan, “Wearable sensor-based quantitative gait analysis in parkinson’s disease patients with different motor subtypes,” npj Digital Medicine , vol. 7, p. 169, Jun 2024
work page 2024
-
[3]
The prac- tical implementation of artificial intelligence technologies in medicine,
J. He, S. L. Baxter, J. Xu, J. Xu, X. Zhou, and K. Zhang, “The prac- tical implementation of artificial intelligence technologies in medicine,” Nature Medicine, vol. 25, pp. 30–36, Jan 2019
work page 2019
-
[4]
European Union, “Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence.” https://eur-lex.europa.eu/eli/reg/2024/1689/oj,
work page 2024
-
[5]
Regulation (eu) 2016/679 (general data protection regulation)
The European Parliament and the Council of the European Union, “Regulation (eu) 2016/679 (general data protection regulation).” Official Journal of the European Union, L 119/1, 2016. Cited for Article 22. Available at: https://eur-lex.europa.eu/eli/reg/2016/679/oj
work page 2016
-
[6]
Parkinsonism: onset, progression, and mortality,
M. M. Hoehn and M. D. Yahr, “Parkinsonism: onset, progression, and mortality,” Neurology, vol. 17, no. 5, pp. 427–427, 1967
work page 1967
-
[7]
Explainable ai systems must be contestable: Here’s how to make it happen,
C. Moreira, A. Palatkina, D. Braca, D. M. Walsh, P. J. Leihn, F. Chen, and N. C. Hubig, “Explainable ai systems must be contestable: Here’s how to make it happen,” arXiv preprint arXiv:2506.01662 , 2025
arXiv 2025
-
[8]
M. H. G. Monje, S. Grosjean, M. Srp, L. Antunes, R. Bouc ¸a-Machado, R. Cacho, S. Dom ´ınguez, J. Inocentes, T. Lynch, A. Tsakanika, D. Fo- tiadis, G. Rigas, E. R ˇuˇziˇcka, J. Ferreira, A. Antonini, N. Malpica, T. Mestre, A. S´anchez-Ferro, and iCARE PD Consortium, “Co-designing digital technologies for improving clinical care in people with parkin- son’...
work page 2023
Show all 27 references
-
[9]
Comprehensive real time remote monitoring for parkinson’s disease using quantitative digitography,
S. L. Hoffman, P. Schmiedmayer, A. S. Gala, K. B. Wilkins, L. Parisi, S. Karjagi, A. S. Negi, S. Revlock, C. Coriz, J. Revlock, et al. , “Comprehensive real time remote monitoring for parkinson’s disease using quantitative digitography,” npj Parkinson’s Disease, vol. 10, no. 1...
2024
-
[10]
Smart on fhir: a standards-based, interoperable apps platform for electronic health records,
J. C. Mandel, D. A. Kreda, K. D. Mandl, I. S. Kohane, and R. B. Ramoni, “Smart on fhir: a standards-based, interoperable apps platform for electronic health records,” Journal of the American Medical Informatics Association, vol. 23, no. 5, pp. 899–908, 2016
2016
-
[11]
Molnar, Interpretable machine learning
C. Molnar, Interpretable machine learning . Lulu. com, 2020
2020
-
[12]
Xedgeai: A human- centered industrial inspection framework with data-centric explainable edge ai approach,
H. T. T. Nguyen, L. P. T. Nguyen, and H. Cao, “Xedgeai: A human- centered industrial inspection framework with data-centric explainable edge ai approach,” Information Fusion, vol. 116, p. 102782, 2025
2025
-
[13]
Langxai: Integrating large vision models for generating textual explanations to enhance explainability in visual perception tasks,
H. Nguyen, T. Clement, L. Nguyen, N. Kemmerzell, B. Truong, K. Nguyen, M. Abdelaal, and H. Cao, “Langxai: Integrating large vision models for generating textual explanations to enhance explainability in visual perception tasks,” in Proceedings of the Thirty-Third International...
2024
-
[14]
Heart2mind: Human-centered contestable psychiatric disorder diagnosis system using wearable ecg monitors,
H. Nguyen, A. Rahimi, V . Whitford, H. Fournier, I. Kondratova, R. Richard, and H. Cao, “Heart2mind: Human-centered contestable psychiatric disorder diagnosis system using wearable ecg monitors,” arXiv preprint arXiv:2505.11612 , 2025
2025 arXiv
-
[15]
Beyond explain- ing: Opportunities and challenges of xai-based model improvement,
L. Weber, S. Lapuschkin, A. Binder, and W. Samek, “Beyond explain- ing: Opportunities and challenges of xai-based model improvement,” Information Fusion, vol. 92, pp. 154–176, 2023
2023
-
[16]
Opacity as a feature, not a flaw: The lobox governance ethic for role-sensitive explainability and institutional trust in ai,
F. Herrera and R. Calder ´on, “Opacity as a feature, not a flaw: The lobox governance ethic for role-sensitive explainability and institutional trust in ai,” arXiv preprint arXiv:2505.20304 , 2025
2025 arXiv
-
[17]
Ixaii: An interactive explain- able artificial intelligence interface for decision support systems,
P. Speckmann, M. Nadj, and C. Janiesch, “Ixaii: An interactive explain- able artificial intelligence interface for decision support systems,” arXiv preprint arXiv:2506.21310, 2025
2025 arXiv
-
[18]
Human-centered explainable psychiatric dis- order diagnosis system using wearable ecg monitors,
H. Nguyen, A. Rahimi, V . Whitford, H. Fournier, I. Kondratova, R. Richard, and H. Cao, “Human-centered explainable psychiatric dis- order diagnosis system using wearable ecg monitors,” in Pacific-Asia Conference on Knowledge Discovery and Data Mining , pp. 418–429, Springer, 2025
2025
-
[19]
Contestable ai by design: Towards a framework,
K. Alfrink, I. Keller, G. Kortuem, and N. Doorn, “Contestable ai by design: Towards a framework,” Minds and Machines , vol. 33, no. 4, pp. 613–639, 2023
2023
-
[20]
Contesting black-box ai decisions,
V . Dignum, L. Michael, J. C. Nieves, M. Slavkovik, J. Suarez, and A. Theodorou, “Contesting black-box ai decisions,” in Proc. of the 24th International Conference on Autonomous Agents and Multiagent Systems, pp. 2854–2858, 2025
2025
-
[21]
The lrp toolbox for artificial neural networks,
S. Lapuschkin, A. Binder, G. Montavon, K.-R. M ¨uller, and W. Samek, “The lrp toolbox for artificial neural networks,” Journal of Machine Learning Research, vol. 17, no. 114, pp. 1–5, 2016
2016
-
[22]
Explaining machine learning models for clinical gait analy- sis,
D. Slijepcevic, F. Horst, S. Lapuschkin, B. Horsak, A.-M. Raberger, A. Kranzl, W. Samek, C. Breiteneder, W. I. Sch ¨ollhorn, and M. Zep- pelzauer, “Explaining machine learning models for clinical gait analy- sis,” ACM Transactions for Computing on Healthcare , vol. 3, pp. 1–27...
2021
-
[23]
Gpt-4o system card,
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al., “Gpt-4o system card,” arXiv preprint arXiv:2410.21276 , 2024
2024 arXiv
-
[24]
Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals,
A. L. Goldberger, L. A. Amaral, L. Glass, J. M. Hausdorff, P. C. Ivanov, R. G. Mark, J. E. Mietus, G. B. Moody, C.-K. Peng, and H. E. Stanley, “Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals,” circulation, vol. 1...
2000
-
[25]
Gpt-4 technical report,
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
2023 arXiv
-
[26]
Tinytroupe: Llm- powered multiagent persona simulation for imagination enhancement and business insights
P. Salem, C. Olsen, P. Freire, Y . Ding, and P. Saxena, “Tinytroupe: Llm- powered multiagent persona simulation for imagination enhancement and business insights.” https://github.com/microsoft/tinytroupe, 2024. GitHub repository
2024
-
[2024]
Accessed: 2025-02-07
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.