LiveBench is a contamination-limited LLM benchmark with auto-scored challenging tasks from recent sources across math, coding, reasoning and more, where top models score below 70%.
Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap
8 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
EngiBench shows LLMs accuracy drops with task complexity, degrades under perturbations, and stays below human performance on open-ended engineering problems.
LLMs display high variance and major accuracy drops on GSM-Symbolic variants of grade-school math problems, indicating they replicate training patterns rather than execute logical reasoning.
AI agents given only executable behavior and tests can reimplement fully scoped programs (e.g., gotree: 16k LoC, 2000/2001 tests), and the best model solves 56% of 25 MirrorCode tasks.
LPDS quantifies difficulty of logic-preserving problem variations and searches for the hardest ones, producing up to 5x larger performance drops than random sampling and better robustness gains from fine-tuning on difficult examples.
Frontier AI models score below 10% on Riemann-Bench, a private expert-authored benchmark of research-level mathematics with closed-form, programmatically verified solutions.
A 13-way text-scrambling benchmark makes open-weight LLMs drop up to 54% average accuracy, and a multi-problem prompt makes their accuracy on the last question decay.
AI for mathematics is best described as a supervision ladder — final answers, programs, process rewards, proof-assistant kernels — culminating in verified-discovery workflows.
citing papers explorer
-
LiveBench: A Challenging, Contamination-Limited LLM Benchmark
LiveBench is a contamination-limited LLM benchmark with auto-scored challenging tasks from recent sources across math, coding, reasoning and more, where top models score below 70%.
-
EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
EngiBench shows LLMs accuracy drops with task complexity, degrades under perturbations, and stays below human performance on open-ended engineering problems.
-
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
LLMs display high variance and major accuracy drops on GSM-Symbolic variants of grade-school math problems, indicating they replicate training patterns rather than execute logical reasoning.
-
MirrorCode: AI can rebuild entire programs from behavior alone
AI agents given only executable behavior and tests can reimplement fully scoped programs (e.g., gotree: 16k LoC, 2000/2001 tests), and the best model solves 56% of 25 MirrorCode tasks.
-
LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling
LPDS quantifies difficulty of logic-preserving problem variations and searches for the hardest ones, producing up to 5x larger performance drops than random sampling and better robustness gains from fine-tuning on difficult examples.
-
Riemann-Bench: A Benchmark for Moonshot Mathematics
Frontier AI models score below 10% on Riemann-Bench, a private expert-authored benchmark of research-level mathematics with closed-form, programmatically verified solutions.
-
Robust Reasoning Benchmark
A 13-way text-scrambling benchmark makes open-weight LLMs drop up to 54% average accuracy, and a multi-problem prompt makes their accuracy on the last question decay.
-
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
AI for mathematics is best described as a supervision ladder — final answers, programs, process rewards, proof-assistant kernels — culminating in verified-discovery workflows.