Pith. sign in

REVIEW 2 cited by

A Holistic Assessment of the Reliability of Machine Learning Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.10586 v2 pith:KN4QZC6W submitted 2023-07-20 cs.LG

classification cs.LG
keywords reliabilitysystemsadversarialalgorithmicassessmentholisticlearningmachine
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As machine learning (ML) systems increasingly permeate high-stakes settings such as healthcare, transportation, military, and national security, concerns regarding their reliability have emerged. Despite notable progress, the performance of these systems can significantly diminish due to adversarial attacks or environmental changes, leading to overconfident predictions, failures to detect input faults, and an inability to generalize in unexpected scenarios. This paper proposes a holistic assessment methodology for the reliability of ML systems. Our framework evaluates five key properties: in-distribution accuracy, distribution-shift robustness, adversarial robustness, calibration, and out-of-distribution detection. A reliability score is also introduced and used to assess the overall system reliability. To provide insights into the performance of different algorithmic approaches, we identify and categorize state-of-the-art techniques, then evaluate a selection on real-world tasks using our proposed reliability metrics and reliability score. Our analysis of over 500 models reveals that designing for one metric does not necessarily constrain others but certain algorithmic techniques can improve reliability across multiple metrics simultaneously. This study contributes to a more comprehensive understanding of ML reliability and provides a roadmap for future research and development.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predictive Uncertainty for Runtime Assurance of a Real-Time Computer Vision-Based Landing System

    cs.CV 2025-08 conditional novelty 4.0 of 10

    Runway keypoint detection with calibrated uncertainty estimates and a RAIM-based geometric check that flags inconsistent predictions.

  2. What Is AI Safety? What Do We Want It to Be?

    cs.CY 2025-05 conditional novelty 4.0 of 10

    AI safety is best understood as all research aimed at preventing or reducing harms from AI systems, covering social harms and catastrophic risks together.

Pith tools