Decentralized poly-time algorithms achieve (1-1/e)-approximate assistance regret Õ(T^{3/4}) (or Õ(√T) with shared randomness) for online assistance games, and better approximation is intractable.
Learning to complement humans
6 Pith papers cite this work. Polarity classification is still indexing.
abstract
A rising vision for AI in the open world centers on the development of systems that can complement humans for perceptual, diagnostic, and reasoning tasks. To date, systems aimed at complementing the skills of people have employed models trained to be as accurate as possible in isolation. We demonstrate how an end-to-end learning strategy can be harnessed to optimize the combined performance of human-machine teams by considering the distinct abilities of people and machines. The goal is to focus machine learning on problem instances that are difficult for humans, while recognizing instances that are difficult for the machine and seeking human input on them. We demonstrate in two real-world domains (scientific discovery and medical diagnosis) that human-machine teams built via these methods outperform the individual performance of machines and people. We then analyze conditions under which this complementarity is strongest, and which training methods amplify it. Taken together, our work provides the first systematic investigation of how machine learning systems can be trained to complement human reasoning.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
GPT-4 exceeds the USMLE passing score by more than 20 points and outperforms both GPT-3.5 and the medically fine-tuned Med-PaLM on the MultiMedQA benchmarks.
In a competitive QA game, humans under-rely on correct AI suggestions 3.9% of the time and over-rely on incorrect ones 1.7% of the time, driven by confirmation bias and near-chance AI confidence when answers disagree.
Conditional risk calibration reduces to standard regression and is distinct from probability calibration.
MedMSA framework retrieves knowledge via language models then builds formal probabilistic models to produce uncertainty-weighted differential diagnoses from symptoms.
Recruiters perceive themselves as retaining agency over GenAI in hiring pipelines, yet GenAI invisibly architects core evaluation inputs, producing only marginal efficiency gains at the cost of deskilling.
citing papers explorer
-
Provably Optimal Learning Algorithms for Assistance Games
Decentralized poly-time algorithms achieve (1-1/e)-approximate assistance regret Õ(T^{3/4}) (or Õ(√T) with shared randomness) for online assistance games, and better approximation is intractable.
-
Capabilities of GPT-4 on Medical Challenge Problems
GPT-4 exceeds the USMLE passing score by more than 20 points and outperforms both GPT-3.5 and the medically fine-tuned Med-PaLM on the MultiMedQA benchmarks.
-
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
In a competitive QA game, humans under-rely on correct AI suggestions 3.9% of the time and over-rely on incorrect ones 1.7% of the time, driven by confirmation bias and near-chance AI confidence when answers disagree.
-
Calibrating conditional risk
Conditional risk calibration reduces to standard regression and is distinct from probability calibration.
-
Medical Model Synthesis Architectures: A Case Study
MedMSA framework retrieves knowledge via language models then builds formal probabilistic models to produce uncertainty-weighted differential diagnoses from symptoms.
-
Resume-ing Control: (Mis)Perceptions of Agency Around GenAI Use in Recruiting Workflows
Recruiters perceive themselves as retaining agency over GenAI in hiring pipelines, yet GenAI invisibly architects core evaluation inputs, producing only marginal efficiency gains at the cost of deskilling.