Pith. sign in

REVIEW 4 major objections 3 minor 27 references

ConGaIT: A Clinician-Centered Dashboard for Contestable AI in Parkinson's Disease Care

T0 review · 4 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read ConGaIT demonstrates that contestable AI can be operationalized through a clinician-centered dashboard, with a Contestability Assessment Score of 0.970.

desk verdict A useful design pattern for contestable clinical dashboards, but the 0.970 CAS is a self-assessment plus LLM personas, not evidence that contestability works with clinicians. read the letter →

arxiv 2507.22300 v1 pith:YEZVAXWO submitted 2025-07-30 cs.HC

classification cs.HC
keywords contestableAIclinicaldashboardParkinson'sdiseasegaitanalysisexplainablehuman-computerinteractionContestabilityAssessmentScoreLLMpersonaevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ConGaIT claims that contestable AI—a clinician's ability to challenge and receive justification for an AI decision—can be embedded directly into a clinical dashboard rather than bolted on afterward. The system pairs Parkinson's disease gait monitoring with a Contest & Justify interaction in which clinicians flag a factual error, normative conflict, or reasoning flaw and receive a language-model justification grounded in Hoehn and Yahr staging. Evaluated with the Contestability Assessment Score, ConGaIT reports 0.970, earning full credit for explainability, openness to contestation, safeguards, and auditing. If the result holds, it gives interface designers a concrete pattern for meeting transparency and oversight expectations in high-stakes clinical AI.

What carries the argument

The load-bearing mechanism is the Contest & Justify interaction pattern: a structured disagreement workflow with three argument types—Factual Error, Normative Conflict, and Reasoning Flaw—that a clinician can raise against a predictive insight. Behind it sit three supporting components: a CNN classifier that predicts Hoehn and Yahr stage from 10-second gait windows, layer-wise relevance propagation highlighting the sensors and time segments that drove the prediction, and a vision language model that produces normative justifications. A role-based feedback layer and an immutable audit trail complete the loop, and the CAS rubric turns contestability into a weighted, scored quantity.

What would settle it

Run the same CAS evaluation with, say, ten practising movement-disorder neurologists using ConGaIT on real gait cases; if their average Explanation Quality or Ease of Contestation scores pull the total CAS materially below the reported 0.970, the proxy-persona assumption and the operationalization claim would not generalize.

Watch

Extended reading notes

Core claim

The central discovery is that contestability can be operationalized as a measurable property of a user interface. ConGaIT demonstrates this through a tightly integrated dashboard in which every AI prediction carries visual explanations (layer-wise relevance propagation), a textual normative justification, and a structured path to dispute that justification. The paper's measurement is the Contestability Assessment Score, whose eight weighted criteria span human-centered, technical, legal, and organizational pillars; ConGaIT attains 0.970, with the only partial scores in traceability, ease of contestation, and explanation quality. The authors conclude that contestability can be embedded early in the design process and that their proxy evaluation with personas offers a reproducible template for doing so.

Load-bearing premise

The whole evaluation leans on the assumption that three large-language-model-generated clinician personas judge the interface the way real clinicians would.

Editorial extensions

If this is right

  • Concrete interface patterns, not just model transparency, can satisfy contestability requirements in clinical AI.
  • Structured argument types give clinicians a precise vocabulary for disagreement, making objections auditable and reusable for model refinement.
  • A CAS score of 0.970 provides a benchmark that other clinical AI systems can be measured against on contestability.
  • Combining visual and textual explanations with a contest flow addresses both explanation and justification, which the paper treats as distinct normative requirements.
  • The proxy evaluation with language-model personas offers a reproducible early-design assessment, though it awaits clinical validation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same Contest & Justify pattern could extend to other high-stakes settings—radiology, pathology, or automated hiring—where a human must legally or ethically be able to challenge an algorithmic decision.
  • A natural next test is whether real clinicians' CAS scores and contestation behavior match the persona simulation; if they diverge, the pattern may need tailoring to actual clinical work practices.
  • Because the audit trail and argument types are structured, the dashboard could feed contested cases back into model training, turning disagreement into a data asset.
  • The CAS rubric's weighting is taken as given; a different weighting scheme could change which design features dominate the score.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The paper presents ConGaIT, a clinician-centered dashboard for Parkinson's disease gait monitoring that embeds contestable AI features: interactive visual explanations (LRP), role-based feedback, an immutable audit log, and a "Contest & Justify" workflow. The system is evaluated using the Contestability Assessment Score (CAS), where six observable criteria are scored by the authors through direct inspection and two subjective criteria are scored by LLM-generated clinician personas (GPT-4 and TinyTroupe), yielding a total CAS of 0.970. The paper claims this score demonstrates that contestability can be operationalized through human-centered design in compliance with emerging regulatory standards.

Significance. If substantiated, the work would offer a concrete interface pattern for contestable clinical AI and a systematic use of the CAS rubric, which is relevant to EU AI Act and GDPR Article 22 compliance. The paper also provides a public demo repository. However, the current evaluation is not independent or validated, so the significance is conditional on a substantially revised evidence base or a clearly repositioned claim.

major comments (4)
  1. [Section IV, Table I] The central claim of a 0.970 CAS is not supported by the evidence presented. Six of the eight criteria are scored by the authors "through direct inspection," and the remaining two are scored by LLM-generated personas rather than real clinicians; no inter-rater reliability or uncertainty is reported. The paper itself states in Section IV and the Conclusion that clinical usability testing and validation remain future work. Consequently, the abstract's claim that the score "demonstrates that contestability can be operationalized through human-centered design" is an overstatement that should be replaced with a claim about a preliminary or self-assessed evaluation, or supplemented with real clinician assessment.
  2. [Section IV, Discussion] The sentence describing the result as benefiting from "objective validation through CAS" is internally inconsistent with the preceding description of direct inspection and persona-based scoring. This contradiction must be corrected, since the majority of the final score rests on subjective, self-assigned judgments rather than objective measurement.
  3. [Table I] The scoring rationale is not reported. For each criterion, the authors should specify what interface evidence justifies the assigned score, why Traceability receives 9 out of 10 rather than a different value, and what rubric or behavioral definition was used to convert observations into scores. Without this, the numerical result is not verifiable or reproducible, which is especially important given the paper's emphasis on reproducibility.
  4. [Section IV, Conclusion] The use of LLM-generated clinician personas as a proxy for clinician judgment is a load-bearing assumption that is explicitly acknowledged as unvalidated. The paper concludes that contestability "can be embedded early in the design process" on the basis of this proxy, but no evidence is given that the personas reproduce real clinician evaluations. The conclusion should be downgraded to an illustration or pilot finding, and the proxy limitation should be made prominent in the abstract.
minor comments (3)
  1. [Abstract and Section III] The system name is written inconsistently as "Con-GaIT" in the abstract and "ConGaIT" in the body; please unify the spelling.
  2. [Section IV] The phrase "structured personas enabled reproducible, expert-informed evaluation" is misleading, since the personas are simulated and not validated as experts; consider rewording to "simulated expert-like personas."
  3. [Table I] The column header "Property Max" is unclear; it should be split into "Property" and "Maximum score" to avoid ambiguity with the weights column.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the CAS score is an external rubric and an arithmetic aggregation, not a prediction forced by construction.

full rationale

The central claim is an empirical evaluation claim: ConGaIT is designed to support contestability and then scored with the Contestability Assessment Score, whose source [7] is an external group with no overlap with the authors. The reported 0.970 is a weighted sum of criterion scores (Table I), where six criteria were assigned by direct inspection and two by LLM-simulated clinician personas. Nothing in the paper defines a design feature in terms of the score, and the score is not a fitted parameter later renamed as a prediction. The closest concern is that self-assessment and GPT-4/TinyTroupe personas could introduce favorable bias, but the paper itself flags this as a proxy: "future work will focus on clinical usability testing" and the conclusion calls it a "novel proxy evaluation." That is a validity and independence limitation, not a circular reduction under the patterns considered here. The authors' prior XAI papers are cited only as related work and are not load-bearing for the CAS result. Therefore the derivation chain is not circular.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central claim rests on the validity of the CAS rubric, the representativeness of LLM personas for clinical judgment, and the reliability of the backend classifier. The first two are especially fragile because they are asserted rather than demonstrated, and the CAS scores are effectively hand-picked.

free parameters (3)
  • Explanation Quality score (sp) = 42/50
    Assigned by LLM-generated personas; directly contributes 0.059 to the CAS total and is not independently validated.
  • Traceability score (sp) = 9/10
    Assigned by author direct inspection; contributes 0.108 to CAS and is not externally verified.
  • Observable criteria scores (Explainability, Openness, Safeguards, Adaptivity, Auditing) = full marks
    All set to the maximum by the authors through direct inspection; these hand-set values determine most of the CAS.
assumptions (3)
  • domain assumption The Contestability Assessment Score (CAS) is a valid and reliable measure of contestability in clinical AI systems.
    The paper uses CAS [7] as the sole evaluation instrument without questioning its psychometric properties or applicability to PD care dashboards.
  • ad hoc to paper LLM-generated clinician personas produce evaluations equivalent to those of real clinicians for subjective criteria.
    Section IV states explanation quality was assessed with three personas generated via GPT-4 and TinyTroupe; the paper itself calls this a proxy, so the assumption is load-bearing and unverified.
  • domain assumption The PhysioNet gait dataset and the CNN classifier trained on it are representative of real clinical gait data.
    The 96.58% test accuracy is reported without a detailed train/test split, hyperparameters, or error bounds, so the reliability of the underlying AI is taken as given.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ConGaIT: A Clinician-Centered Dashboard for Contestable AI in Parkinson's Disease Care." pith.science (2026). https://pith.science/paper/YEZVAXWO

@misc{pith2026250722300,
  author       = {Pith},
  title        = {Pith review of: ConGaIT: A Clinician-Centered Dashboard for Contestable AI in Parkinson's Disease Care},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YEZVAXWO}},
  note         = {Machine review of arXiv:2507.22300}
}
read the original abstract

AI-assisted gait analysis holds promise for improving Parkinson's Disease (PD) care, but current clinical dashboards lack transparency and offer no meaningful way for clinicians to interrogate or contest AI decisions. We present Con-GaIT (Contestable Gait Interpretation & Tracking), a clinician-centered system that advances Contestable AI through a tightly integrated interface designed for interpretability, oversight, and procedural recourse. Grounded in HCI principles, ConGaIT enables structured disagreement via a novel Contest & Justify interaction pattern, supported by visual explanations, role-based feedback, and traceable justification logs. Evaluated using the Contestability Assessment Score (CAS), the framework achieves a score of 0.970, demonstrating that contestability can be operationalized through human-centered design in compliance with emerging regulatory standards. A demonstration of the framework is available at https://github.com/hungdothanh/Con-GaIT.

Figures

Figures reproduced from arXiv: 2507.22300 by the authors.

Figure 2
Figure 2. Overview of the Gait Session Summary tab [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Overview of the Treatment Trend View tab [PITH_FULL_IMAGE:figures/full_fig_p002_3.png] view at source ↗
Figure 4
Figure 4. Overview of the Predictive Insight and Explanation tab [PITH_FULL_IMAGE:figures/full_fig_p002_4.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

27 extracted references · 21 canonical work pages

  1. [1]

    Parkinson’s disease: clinical features and diagnosis,

    J. Jankovic, “Parkinson’s disease: clinical features and diagnosis,” Jour- nal of neurology, neurosurgery & psychiatry, vol. 79, no. 4, pp. 368–376, 2008

  2. [2]

    Wearable sensor-based quantitative gait analysis in parkinson’s disease patients with different motor subtypes,

    W. Zhang, Y . Ling, Z. Chen, K. Ren, S. Chen, P. Huang, and Y . Tan, “Wearable sensor-based quantitative gait analysis in parkinson’s disease patients with different motor subtypes,” npj Digital Medicine , vol. 7, p. 169, Jun 2024

  3. [3]

    The prac- tical implementation of artificial intelligence technologies in medicine,

    J. He, S. L. Baxter, J. Xu, J. Xu, X. Zhou, and K. Zhang, “The prac- tical implementation of artificial intelligence technologies in medicine,” Nature Medicine, vol. 25, pp. 30–36, Jan 2019

  4. [4]

    Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence

    European Union, “Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence.” https://eur-lex.europa.eu/eli/reg/2024/1689/oj,

  5. [5]

    Regulation (eu) 2016/679 (general data protection regulation)

    The European Parliament and the Council of the European Union, “Regulation (eu) 2016/679 (general data protection regulation).” Official Journal of the European Union, L 119/1, 2016. Cited for Article 22. Available at: https://eur-lex.europa.eu/eli/reg/2016/679/oj

  6. [6]

    Parkinsonism: onset, progression, and mortality,

    M. M. Hoehn and M. D. Yahr, “Parkinsonism: onset, progression, and mortality,” Neurology, vol. 17, no. 5, pp. 427–427, 1967

  7. [7]

    Explainable ai systems must be contestable: Here’s how to make it happen,

    C. Moreira, A. Palatkina, D. Braca, D. M. Walsh, P. J. Leihn, F. Chen, and N. C. Hubig, “Explainable ai systems must be contestable: Here’s how to make it happen,” arXiv preprint arXiv:2506.01662 , 2025

  8. [8]

    Co-designing digital technologies for improving clinical care in people with parkin- son’s disease: What did we learn?,

    M. H. G. Monje, S. Grosjean, M. Srp, L. Antunes, R. Bouc ¸a-Machado, R. Cacho, S. Dom ´ınguez, J. Inocentes, T. Lynch, A. Tsakanika, D. Fo- tiadis, G. Rigas, E. R ˇuˇziˇcka, J. Ferreira, A. Antonini, N. Malpica, T. Mestre, A. S´anchez-Ferro, and iCARE PD Consortium, “Co-designing digital technologies for improving clinical care in people with parkin- son’...

Show all 27 references
  1. [9]

    Comprehensive real time remote monitoring for parkinson’s disease using quantitative digitography,

    S. L. Hoffman, P. Schmiedmayer, A. S. Gala, K. B. Wilkins, L. Parisi, S. Karjagi, A. S. Negi, S. Revlock, C. Coriz, J. Revlock, et al. , “Comprehensive real time remote monitoring for parkinson’s disease using quantitative digitography,” npj Parkinson’s Disease, vol. 10, no. 1...

  2. [10]

    Smart on fhir: a standards-based, interoperable apps platform for electronic health records,

    J. C. Mandel, D. A. Kreda, K. D. Mandl, I. S. Kohane, and R. B. Ramoni, “Smart on fhir: a standards-based, interoperable apps platform for electronic health records,” Journal of the American Medical Informatics Association, vol. 23, no. 5, pp. 899–908, 2016

  3. [11]

    Molnar, Interpretable machine learning

    C. Molnar, Interpretable machine learning . Lulu. com, 2020

  4. [12]

    Xedgeai: A human- centered industrial inspection framework with data-centric explainable edge ai approach,

    H. T. T. Nguyen, L. P. T. Nguyen, and H. Cao, “Xedgeai: A human- centered industrial inspection framework with data-centric explainable edge ai approach,” Information Fusion, vol. 116, p. 102782, 2025

  5. [13]

    Langxai: Integrating large vision models for generating textual explanations to enhance explainability in visual perception tasks,

    H. Nguyen, T. Clement, L. Nguyen, N. Kemmerzell, B. Truong, K. Nguyen, M. Abdelaal, and H. Cao, “Langxai: Integrating large vision models for generating textual explanations to enhance explainability in visual perception tasks,” in Proceedings of the Thirty-Third International...

  6. [14]

    Heart2mind: Human-centered contestable psychiatric disorder diagnosis system using wearable ecg monitors,

    H. Nguyen, A. Rahimi, V . Whitford, H. Fournier, I. Kondratova, R. Richard, and H. Cao, “Heart2mind: Human-centered contestable psychiatric disorder diagnosis system using wearable ecg monitors,” arXiv preprint arXiv:2505.11612 , 2025

  7. [15]

    Beyond explain- ing: Opportunities and challenges of xai-based model improvement,

    L. Weber, S. Lapuschkin, A. Binder, and W. Samek, “Beyond explain- ing: Opportunities and challenges of xai-based model improvement,” Information Fusion, vol. 92, pp. 154–176, 2023

  8. [16]

    Opacity as a feature, not a flaw: The lobox governance ethic for role-sensitive explainability and institutional trust in ai,

    F. Herrera and R. Calder ´on, “Opacity as a feature, not a flaw: The lobox governance ethic for role-sensitive explainability and institutional trust in ai,” arXiv preprint arXiv:2505.20304 , 2025

  9. [17]

    Ixaii: An interactive explain- able artificial intelligence interface for decision support systems,

    P. Speckmann, M. Nadj, and C. Janiesch, “Ixaii: An interactive explain- able artificial intelligence interface for decision support systems,” arXiv preprint arXiv:2506.21310, 2025

  10. [18]

    Human-centered explainable psychiatric dis- order diagnosis system using wearable ecg monitors,

    H. Nguyen, A. Rahimi, V . Whitford, H. Fournier, I. Kondratova, R. Richard, and H. Cao, “Human-centered explainable psychiatric dis- order diagnosis system using wearable ecg monitors,” in Pacific-Asia Conference on Knowledge Discovery and Data Mining , pp. 418–429, Springer, 2025

  11. [19]

    Contestable ai by design: Towards a framework,

    K. Alfrink, I. Keller, G. Kortuem, and N. Doorn, “Contestable ai by design: Towards a framework,” Minds and Machines , vol. 33, no. 4, pp. 613–639, 2023

  12. [20]

    Contesting black-box ai decisions,

    V . Dignum, L. Michael, J. C. Nieves, M. Slavkovik, J. Suarez, and A. Theodorou, “Contesting black-box ai decisions,” in Proc. of the 24th International Conference on Autonomous Agents and Multiagent Systems, pp. 2854–2858, 2025

  13. [21]

    The lrp toolbox for artificial neural networks,

    S. Lapuschkin, A. Binder, G. Montavon, K.-R. M ¨uller, and W. Samek, “The lrp toolbox for artificial neural networks,” Journal of Machine Learning Research, vol. 17, no. 114, pp. 1–5, 2016

  14. [22]

    Explaining machine learning models for clinical gait analy- sis,

    D. Slijepcevic, F. Horst, S. Lapuschkin, B. Horsak, A.-M. Raberger, A. Kranzl, W. Samek, C. Breiteneder, W. I. Sch ¨ollhorn, and M. Zep- pelzauer, “Explaining machine learning models for clinical gait analy- sis,” ACM Transactions for Computing on Healthcare , vol. 3, pp. 1–27...

  15. [23]

    Gpt-4o system card,

    A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al., “Gpt-4o system card,” arXiv preprint arXiv:2410.21276 , 2024

  16. [24]

    Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals,

    A. L. Goldberger, L. A. Amaral, L. Glass, J. M. Hausdorff, P. C. Ivanov, R. G. Mark, J. E. Mietus, G. B. Moody, C.-K. Peng, and H. E. Stanley, “Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals,” circulation, vol. 1...

  17. [25]

    Gpt-4 technical report,

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023

  18. [26]

    Tinytroupe: Llm- powered multiagent persona simulation for imagination enhancement and business insights

    P. Salem, C. Olsen, P. Freire, Y . Ding, and P. Saxena, “Tinytroupe: Llm- powered multiagent persona simulation for imagination enhancement and business insights.” https://github.com/microsoft/tinytroupe, 2024. GitHub repository

  19. [2024]

    Accessed: 2025-02-07

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.