Pith. sign in

REVIEW 1 cited by

Measuring the Complexity of Domains Used to Evaluate AI Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.01985 v1 pith:AEQHTXNY submitted 2020-09-18 cs.AI

Measuring the Complexity of Domains Used to Evaluate AI Systems

classification cs.AI
keywords complexitysystemsdomainsmeasureapproximationscurrentlymeasuringpropose
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

There is currently a rapid increase in the number of challenge problem, benchmarking datasets and algorithmic optimization tests for evaluating AI systems. However, there does not currently exist an objective measure to determine the complexity between these newly created domains. This lack of cross-domain examination creates an obstacle to effectively research more general AI systems. We propose a theory for measuring the complexity between varied domains. This theory is then evaluated using approximations by a population of neural network based AI systems. The approximations are compared to other well known standards and show it meets intuitions of complexity. An application of this measure is then demonstrated to show its effectiveness as a tool in varied situations. The experimental results show this measure has promise as an effective tool for aiding in the evaluation of AI systems. We propose the future use of such a complexity metric for use in computing an AI system's intelligence.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Pre-Deployment Complexity Estimation for Federated Perception Systems

    cs.LG 2026-03 conditional novelty 4.0

    A pre-deployment complexity score for federated learning, built from entropy, sparsity, and intrinsic dimensionality plus client frequencies, predicts accuracy and communication rounds on three MNIST variants.