Pith. sign in

REVIEW 83 references

Clinical acceptance of software based on artificial intelligence technologies (radiology)

T0 review · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper sets out a staged framework for clinical validation of AI radiology software using standard diagnostic metrics, reference dataset requirements, and a STARD-based reporting checklist.

arxiv 1908.00381 v2 pith:KJMVQOWB submitted 2019-08-01 cs.AI cs.CY

classification cs.AIcs.CY
keywords clinicalsoftwareacceptancealgorithmsartificialintelligenceradiologytechnologies
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This is a guideline, not a research study. It comes from a Moscow government clinical center and explains how AI software used in radiology should be tested before being registered as a medical device. The authors argue that software which helps doctors interpret images can affect patient health, so it should be regulated like other medical devices. The core proposal is a two-stage process. The first stage, analytical validation, checks whether the algorithm produces reliable outputs from real images. It has six steps: a questionnaire, a self-test on a small sample dataset, an interview with the developer, an online test, an evidence test on a larger reference dataset, and a final evaluation. The second stage, clinical acceptance, runs the software in a real radiology department to measure its effect on workflow, for example, by timing radiologists with and without the software. The document defines which metrics to use: sensitivity, specificity, accuracy, likelihood ratios, predictive values, ROC and AUC, Cohen's kappa, and the Dice similarity coefficient. It sets thresholds: values below 0.6 are unsuitable, 0.61 to 0.8 require revision, and 0.81 or higher are acceptable for clinical validation. It also specifies requirements for reference datasets, such as matching the prevalence of disease in the population and not being publicly available, so algorithms cannot be trained on the test data. The annexes include questionnaires, interview protocols, and report templates based on the STARD 2015 checklist.
Extended reading notes

Core claim

The central assertion is that AI-based algorithms and software used in radiology belong to medical devices and must undergo clinical trials, and that the proposed two-stage framework (analytical validation and clinical acceptance) is a valid way to evaluate their accuracy and safety. If correct, this means all AI radiology software in Russia must pass the described tests before registration.

Load-bearing premise

The framework assumes that diagnostic accuracy measured against a labeled reference dataset, plus the arbitrary acceptance thresholds (0.61 to 0.8 = revision required, 0.81 and above = admissible), is a sufficient basis for deciding clinical safety and effectiveness. This premise enters in the Quality metrics section (Table 5 and the ROC/AUC evaluation scale), where the cutoffs are stated without empirical or outcome-based justification.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The framework rests on standard statistical assumptions about diagnostic tests and on Russian regulatory law. The only ad hoc element is the set of numerical acceptance thresholds, which are presented without empirical support.

free parameters (1)
  • Diagnostic metric acceptance thresholds = 0.6, 0.8, 0.81
    The paper assigns unsuitable below 0.6, revision required between 0.61 and 0.8, and admissible at or above 0.81 for sensitivity, specificity, accuracy, AUC, kappa, and Dice. These cutoffs are chosen by hand with no empirical or regulatory derivation.
assumptions (4)
  • domain assumption A labeled reference dataset, tagged by qualified specialists, is a valid reference standard (ground truth) for evaluating AI diagnostic accuracy.
    Introduced in 'Diagnostic accuracy assessment': 'The biomedical dataset tagging... is used as a reference test.'
  • domain assumption Binary classification metrics (sensitivity, specificity, AUC) are sufficient to assess AI outputs, because 'most other tasks can almost always be reduced to a binary classification task.'
    Footnote 6 in the Quality metrics section.
  • domain assumption AI-based radiology software is a medical device under Russian law and therefore must undergo clinical tests.
    Introduction, citing Federal Law No. 323 and Roszdravnadzor letters.
  • ad hoc to paper The numeric thresholds for acceptance (0.61-0.8 revision, >=0.81 admissible) are valid without outcome-based calibration.
    Quality metrics section; no justification or citations are provided for these cutoffs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Clinical acceptance of software based on artificial intelligence technologies (radiology)." pith.science (2026). https://pith.science/paper/KJMVQOWB

@misc{pith2026190800381,
  author       = {Pith},
  title        = {Pith review of: Clinical acceptance of software based on artificial intelligence technologies (radiology)},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KJMVQOWB}},
  note         = {Machine review of arXiv:1908.00381}
}
read the original abstract

Aim: provide a methodological framework for the process of clinical tests, clinical acceptance, and scientific assessment of algorithms and software based on the artificial intelligence (AI) technologies. Clinical tests are considered as a preparation stage for the software registration as a medical product. The authors propose approaches to evaluate accuracy and efficiency of the AI algorithms for radiology.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

83 extracted references · 80 canonical work pages

  1. [1]

    Evidence-based medicine: evaluation of the effectiveness of diagnostic interventions

    Vlasov V.V, Rebrova O.Yu. Evidence-based medicine: evaluation of the effectiveness of diagnostic interventions. Deputy Head Doctor. 2010. No. 4. P . 50

  2. [2]

    Methodical recommendations on the order of assessing the quality, efficiency and safety of medical devices (in terms of software) for the state registration under the national system / M.: State Budgetary Institution All-Russian Scientific Research Institute of Medical Equipment of Roszdravnadzor, 2018. – 31 p

  3. [3]

    Deep Learning

    Nikolenko S., Kadurin A., Arkhangelskaya E. Deep Learning. SPb.: Peter, 2018

  4. [4]

    Best practices in medical imaging

    Morozov S.P ., Vetsheva N.N., Ledikhova N.V. Quality assessment of imaging studies / Series “Best practices in medical imaging” . – Issue 22. – М., 2019. – 50 p

  5. [5]

    The Role of the ACR Data Science Institute in Advancing Health Equity in Radiology

    Allen B, Dreyer K. The Role of the ACR Data Science Institute in Advancing Health Equity in Radiology. J Am Coll Radiol. 2019 Apr; 16(4 Pt B):644-648. DOI: 10.1016/j.jacr.2018.12.038

  6. [6]

    Opportunities, Applications and Risks

    Artificial Intelligence in Medical Imaging. Opportunities, Applications and Risks. Erik R. Ranschaert, Sergey Morozov, Paul R. Algra (Eds.); Springer, 2019

  7. [7]

    STARD 2015: An Updated List of Essential Items for Reporting Diagnostic Accuracy Studies

    Bossuyt PM, Reitsma JB, Bruns DE et al., For the STARD Group. STARD 2015: An Updated List of Essential Items for Reporting Diagnostic Accuracy Studies. Radiology. 2015:151516. PMID: 26509226

  8. [8]

    - https://bit.ly/2tiSybQ

    Guidance for Industry and Food and Drug Administration Staff - Computer- Assisted Detection Devices Applied to Radiology Images and Radiology Device Data - Premarket Notification. - https://bit.ly/2tiSybQ

Show all 83 references
  1. [9]

    – 20.07.2018

    RCR position statement on artificial intelligence. – 20.07.2018. https://www.rcr.ac.uk/posts/rcr-position-statement-artificial-intelligence

  2. [10]

    Canadian Association of Radiologists White Paper on Artificial Intelligence in Radiology

    Tang A, Tam R, Cadrin-Chênevert A et al. Canadian Association of Radiologists White Paper on Artificial Intelligence in Radiology. Can Assoc Radiol J. 2018 May; 69(2):120-135. DOI: 10.1016/j.carj.2018.02.002. Best practices in medical imaging 27 RADIOLOGY MOSCOW Annex 1 Criter...

  3. [11]

    Goals: - The software provides a preliminary automatic analysis of medical images to improve the quality and speed of the radiology workflow; - The software ensures a prioritization in the worklist according to the automatically revealed pathology; - The software provides a pr...

  4. [12]

    Scopus” and/or “Web of Science

    Certification: - The medical device has passed technical tests in an accredited laboratory; - Approvals of FDA and/or CE certification (class II); actual implementations of the currently working software in medical centers: at least 2 independent institutions; more than 6 mont...

  5. [13]

    Security: - Compliance with the requirements of the legislation of the Russian Federation in the field of personal data, information security and health protection; - Availability or readiness to deploy server capacities necessary for software operation within the Russian Federation

  6. [14]

    Evidence: - Once the development was completed, the accuracy of algorithms was assessed on independent data9; ______________________________ 9 The medical dataset for testing was different from the dataset used for training, calibration and validation of the algorithm (that is...

  7. [15]

    Standardization: - Automated analysis of diagnostic images in DICOM standard; - Support of the HealthLevel7 (HL7) / FHIR standard (in particular, the system should provide an exchange of messages on the completion of automatic image analysis, pathology detection and classifica...

  8. [16]

    seamless

    Integration: - Availability or readiness to develop means for “seamless” integration with information systems working in the field of health care of the given subject of the Russian Federation, medical information systems; - Availability of tools for integration with PACS and ...

  9. [17]

    status complete by radiographer

    Functionality: - Ability of stream processing and subsequent sending of series to PACS, extended with the results of the AI analysis; - Possibility to combine series of native images and those containing the results of the AI analysis; - Identification of findings (nosologies,...

  10. [18]

    Contract: - The legal entity that is the software developer must have a quality management system; - The legal entity that is the software developer must have a version preparation and control policies; - Regular system updates, including those for diagnostic accuracy informat...

  11. [19]

    The software provides a preliminary automatic analysis of medical images (DICOM files) to improve the quality and speed of the radiology workflow

    Goals 1.1. The software provides a preliminary automatic analysis of medical images (DICOM files) to improve the quality and speed of the radiology workflow. 1.2. The software ensures a prioritization in the worklist according to the automatically revealed pathology. 1.3. The ...

  12. [20]

    Scopus” and/or “Web of Science

    Certification 2.1. Approvals of FDA and/or CE certification (class II). If the answer to clause 2.1 is “no” , there should be positive answers to clauses 2.2 and 2.3. 2.2. Actual implementations of the currently working software in medical centers: - at least 2 independent ins...

  13. [21]

    Once the development was completed, the accuracy of algorithms was assessed on independent data, i.e

    Evidence 3.1. Once the development was completed, the accuracy of algorithms was assessed on independent data, i.e. biomedical dataset for testing differed from the one used for training, development and validation. That is, clinical tests were performed on data unknown to the...

  14. [22]

    Availability of a built-in accuracy assessment tool (if applicable)

    Functionality 4.1. Availability of a built-in accuracy assessment tool (if applicable). 4.2. Maximum processing time of a single radiology study does not exceed the set time period (which is defined individually based on the clinical scenario, infrastructure characteristics, e...

  15. [23]

    The legal entity that is the software developer must have a quality management system

    Contract 5.1. The legal entity that is the software developer must have a quality management system. 5.2. Regular system updates, including those for diagnostic accuracy information. 5.3. Software updates included in the price. 5.4. All medical data, related materials and soft...

  16. [24]

    Comments, 1.3. Anatomical area. Free response in the

    Solution 1.1. Solution name. Free response in the «Comments, 1.2. Imaging modality. Free response in the "Comments, 1.3. Anatomical area. Free response in the "Comments, 1.4. Application area. Oncology Pulmonology Cardiology Neurology Chronic diseases Medical emergencies 1.5. ...

  17. [25]

    Type of medical care

    Clinical application scenario 2.1. Type of medical care. Planned Emergency 2.2. Level of medical care. Primary Secondary Tertiary 2.3. Stage of medical care. Prehospital Hospital Outpatient results are provided to the user. Before viewing the worklist In the worklist After a s...

  18. [26]

    What class of medical software does the proposed AI service belong to? (Only one answer is possible) of the questionnaire

    Risks 3.1. What class of medical software does the proposed AI service belong to? (Only one answer is possible) of the questionnaire. Class 1 Class 2a Class 2b Class 3

  19. [27]

    User categories 4.1. Who can use the AI service? Specially trained healthcare professionals Patient supervised by specially trained healthcare professionals Patient without supervision of specially trained healthcare professionals

  20. [28]

    Automated analysis of medical images Yes No 5.2

    Functional capabilities 5.1. Automated analysis of medical images Yes No 5.2. Prioritization in the worklist according to the automatically revealed pathology. Yes No 5.3. Automated preparation of a draft radiology report based on the results of the analysis. Yes No 5.4. Preli...

  21. [29]

    Once the development was completed, the accuracy of algorithms was assessed on independent data, i.e

    Evidence 7.1. Once the development was completed, the accuracy of algorithms was assessed on independent data, i.e. medical database for development and validation. That is, clinical tests were performed on data unknown to the algorithms. Yes No 7.2. Diagnostic accuracy was te...

  22. [30]

    Availability of a built-in accuracy assessment tool

    Functionality 8.1. Availability of a built-in accuracy assessment tool. Yes No N/A 8.2. Processing time of a single radiology study (specify in sec.) Please, specify system requirements. sec 8.3. The result of software operation is series of images (DICOM format), with: – the ...

  23. [31]

    Regular system updates, including those for diagnostic accuracy information

    Contract 9.1. Regular system updates, including those for diagnostic accuracy information. Yes No 9.2. Software updates included in the price. Yes No 9.3. All medical data, related materials and software results are the property of the customer. Yes No Completed by (full name ...

  24. [32]

    What kind of clinical problem does the AI service solve?

  25. [33]

    Is the use of the AI service for the selected clinical problem justified in the context of evidence-based medicine?

  26. [34]

    How does the result provided by the AI service affect patient management strategy?

  27. [35]

    Does the quality of doctor’s work improve?

  28. [36]

    Does the speed of doctor’s work increase? Analytical validation A) Technical readiness

  29. [37]

    Timing, measuring the time required to process 1 study

  30. [38]

    Visual representation of the results obtained by the AI service on the radiology workstation

  31. [39]

    Best practices in medical imaging 38 RADIOLOGY MOSCOW B) Details on creating and applying the AI model

    Output data format (DICOM, JPEG, PNG, BMP , AVI (copy)). Best practices in medical imaging 38 RADIOLOGY MOSCOW B) Details on creating and applying the AI model

  32. [40]

    What databases were used in building the AI models (from open sources, external partners, in-house)?

  33. [41]

    Balancing classes in databases

  34. [42]

    Total amount of data used to build the AI model for each clinical problem

  35. [43]

    Architecture of the AI model

  36. [44]

    Libraries and resources used to build the model

  37. [45]

    Approaches to the model development (from scratch, transfer learning, retraining of existing architecture)

  38. [46]

    Type of machine learning (supervised / unsupervised / semi-supervised / weak supervision)

  39. [47]

    Validation of the model (independent validation set, cross-validation)

  40. [48]

    Available computational resources

    Time required to create AI models. Available computational resources

  41. [49]

    Are the AI models retrained? How is the versioning organized?

  42. [50]

    Does the AI service process data other than images?

  43. [51]

    Can the AI service evaluate technical quality of the study?

  44. [52]

    Is there any image preprocessing?

  45. [53]

    Can the AI service evaluate the quality of the study interpretation (i.e., control the radiologist’s work)?

  46. [54]

    Values of accuracy metrics of the AI service performance for each clinical problem (AUC, sens, spec, etc.)

  47. [55]

    Were the AI models compared to existing analogues that solve the same clinical problems? C) Details on the datasets and their tagging used for developing AI models

  48. [56]

    Were any criteria applied for selection of a database for training? What kind of criteria (gender, age, diseases, ICD, normal-to-abnormal ratio)?

  49. [57]

    Did the training database contain different studies of the same patient (to be trained to recognize dynamics)?

  50. [58]

    What was considered ground truth in each clinical scenario in the tagged database?

  51. [59]

    Were any quality control criteria applied to the collected database before tagging?

  52. [60]

    What is the business process of tagging?

  53. [61]

    Were the taggers selected in a particular way? Any criteria?

  54. [62]

    Any quality control of input data and tagging results?

  55. [63]

    How many people participated in the data tagging (how many people per 1 study; was the same study tagged by one specialist twice)?

  56. [64]

    Were the taggers trained before tagging?

  57. [65]

    Best practices in medical imaging 39 RADIOLOGY MOSCOW

    Tagging software: in-house or external? Please specify. Best practices in medical imaging 39 RADIOLOGY MOSCOW

  58. [66]

    Approach in case of inconsistency of taggers’ opinions

  59. [67]

    Business process used by the expert confirming the finding:

  60. [68]

    Digital footprint (a series of expert’s comments addressed to the tagger)

  61. [69]

    Was the tagging process automated (CAD for outlining)?

  62. [70]

    Clinical validation A) Practical application in clinical workflow

    Monitoring of the tagging process. Clinical validation A) Practical application in clinical workflow

  63. [71]

    Was the integration with PACS performed in a medical organization? If not, how was the workflow integration performed?

  64. [72]

    What is the business process used by the radiologist working with the AI service in practice?

  65. [73]

    Are the AI service operation instructions available for the doctor?

  66. [74]

    Questions to the doctor: how was he/she trained and whether AI service operation instructions and opportunities for feedback are available?

  67. [75]

    Availability of study triage

  68. [76]

    Possibility to localize the finding using AI

  69. [77]

    Possibility to classify the finding

  70. [78]

    Possibility to compare studies over time

  71. [79]

    Availability of a template to form a report and/or conclusion

  72. [80]

    At what stage of business process does the radiologist receive the processing results obtained by the AI service?

  73. [81]

    What kind of feedback? If available, is it mandatory?

    Availability of feedback from the radiologist on the use of AI. What kind of feedback? If available, is it mandatory?

  74. [82]

    What happens if the radiologist and AI have different opinions (overdiagnosis and underdiagnosis)? How is the final decision on the study made?

  75. [83]

    gold standard

    Availability of any metrics to evaluate the use of AI in clinical practice. Examples (e.g., decreasing the number of errors made by radiologists and reducing the time required to interpret one study). Best practices in medical imaging 40 RADIOLOGY MOSCOW Annex 5 Evaluation of ...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.