Frontier LLM agents recover at most 46% of the human speedup when asked to reimplement successive NanoGPT speedrun records, even with pseudocode, text, and mini-paper hints.
Level 2 hint generation prompt Given the current code, changelog, and next code, provide a detailed natural language description of the improvements made
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
Frontier LLM agents recover at most 46% of the human speedup when asked to reimplement successive NanoGPT speedrun records, even with pseudocode, text, and mini-paper hints.