A new 254-problem ICPC benchmark with a multi-turn self-judge plus episodic retrieval method lifts o1's pass@1 from 19.1% to 42.2%, and a small human-in-the-loop study finds o1 can solve 17 of 18 previously unsolvable problems with a few hints.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Can Multi-turn Self-refined Single Agent LMs with Retrieval Solve Hard Coding Problems?
A new 254-problem ICPC benchmark with a multi-turn self-judge plus episodic retrieval method lifts o1's pass@1 from 19.1% to 42.2%, and a small human-in-the-loop study finds o1 can solve 17 of 18 previously unsolvable problems with a few hints.