A systematic critique showing temporal leakage and extrapolation flaws can undermine claims that LLM forecasters match or beat humans.
Alignment Problems With Current Forecasting Platforms
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We present alignment problems in current forecasting platforms, such as Good Judgment Open, CSET-Foretell or Metaculus. We classify those problems as either reward specification problems or principal-agent problems, and we propose solutions. For instance, the scoring rule used by Good Judgment Open is not proper, and Metaculus tournaments disincentivize sharing information and incentivize distorting one's true probabilities to maximize the chances of placing in the top few positions which earn a monetary reward. We also point out some partial similarities between the problem of aligning forecasters and the problem of aligning artificial intelligence systems.
fields
cs.LG 1years
2025 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Pitfalls in Evaluating Language Model Forecasters
A systematic critique showing temporal leakage and extrapolation flaws can undermine claims that LLM forecasters match or beat humans.