Pith. sign in

REVIEW 1 cited by

AI Oversight and Human Mistakes: Evidence from Centre Court

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.16754 v3 pith:LE7EZT3K submitted 2024-01-30 cs.LG cs.CYecon.GNq-fin.EC

classification cs.LGcs.CYecon.GNq-fin.EC
keywords umpiresoversighterrorshumantypeballcallingcosts
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Powered by the increasing predictive capabilities of machine learning algorithms, artificial intelligence (AI) systems have the potential to overrule human mistakes in many settings. We provide the first field evidence that the use of AI oversight can impact human decision-making. We investigate one of the highest visibility settings where AI oversight has occurred: Hawk-Eye review of umpires in top tennis tournaments. We find that umpires lowered their overall mistake rate after the introduction of Hawk-Eye review, but also that umpires increased the rate at which they called balls in, producing a shift from making Type II errors (calling a ball out when in) to Type I errors (calling a ball in when out). We structurally estimate the psychological costs of being overruled by AI using a model of attention-constrained umpires, and our results suggest that because of these costs, umpires cared 37% more about Type II errors under AI oversight.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reliability, Resilience and Human Factors Engineering for Trustworthy AI Systems

    cs.AI 2024-11 conditional novelty 4.0 of 10

    A framework applying classical reliability, resilience, and human-factors engineering to AI systems, with a small subjective case study of OpenAI status incidents.

Pith tools