TEMPO uses a leakage-first then performance reward with GRPO training to reduce post-cutoff knowledge leakage in LLM backtesting from 2-13% to 0.6-3.7% while preserving or improving task performance on three prediction tasks.
expects 900,000 b/d from Permian by yearend 2023, up from 650,000 b/d by yearend 2022
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting
TEMPO uses a leakage-first then performance reward with GRPO training to reduce post-cutoff knowledge leakage in LLM backtesting from 2-13% to 0.6-3.7% while preserving or improving task performance on three prediction tasks.