Pith. sign in

REVIEW 4 cited by

Predictable Artificial Intelligence

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.06167 v3 pith:7BJHURSE submitted 2023-10-09 cs.AI

classification cs.AI
keywords predictabilitypredictablechallengesecosystemsillustrateperformanceresearchsafety
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce the fundamental ideas and challenges of Predictable AI, a nascent research area that explores the ways in which we can anticipate key validity indicators (e.g., performance, safety) of present and future AI ecosystems. We argue that achieving predictability is crucial for fostering trust, liability, control, alignment and safety of AI ecosystems, and thus should be prioritised over performance. We formally characterise predictability, explore its most relevant components, illustrate what can be predicted, describe alternative candidates for predictors, as well as the trade-offs between maximising validity and predictability. To illustrate these concepts, we bring an array of illustrative examples covering diverse ecosystem configurations. Predictable AI is related to other areas of technical and non-technical AI research, but have distinctive questions, hypotheses, techniques and challenges. This paper aims to elucidate them, calls for identifying paths towards a landscape of predictably valid AI systems and outlines the potential impact of this emergent field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press

    cs.CL 2026-08 conditional novelty 7.0 of 10

    A six-year, open, daily-updated database of over two million unassimilated English borrowings in the Spanish press, with an honest evaluation of its detector's real-world precision.

  2. 11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A new spatial reasoning benchmark shows current multimodal models lag humans badly and lack the item-level predictability humans show.

  3. Model Performance-Guided Evaluation Data Selection for Effective Prompt Optimization

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An evaluation-set selection method that adds real-time model feedback to semantic sampling improves the accuracy and stability of three prompt optimization methods on two datasets.

  4. AI Behavioral Science

    cs.HC 2025-08 unverdicted novelty 3.0 of 10

    A position paper outlines a research agenda for AI Behavioral Science built on three pillars: assessing AI behavior, using AI as a behavioral science tool, and understanding human-AI ecosystems.

Pith tools