Pith. sign in

REVIEW

The Neural Testbed: Evaluating Joint Predictions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.04629 v4 pith:3ITQLI7X submitted 2021-10-09 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords predictionsjointagentsneuraltestbedacrossmarginalquality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Predictive distributions quantify uncertainties ignored by point estimates. This paper introduces The Neural Testbed: an open-source benchmark for controlled and principled evaluation of agents that generate such predictions. Crucially, the testbed assesses agents not only on the quality of their marginal predictions per input, but also on their joint predictions across many inputs. We evaluate a range of agents using a simple neural network data generating process. Our results indicate that some popular Bayesian deep learning agents do not fare well with joint predictions, even when they can produce accurate marginal predictions. We also show that the quality of joint predictions drives performance in downstream decision tasks. We find these results are robust across choice a wide range of generative models, and highlight the practical importance of joint predictions to the community.

Discussion (0). Continue with ORCID to comment.

Pith tools