Increasing LLM coding agents' reasoning effort raises cost and process complexity but does not reliably improve model quality across 140 controlled runs on networked anagram game data.
and McNeil, Barbara J
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
Lexical frequency is a stronger predictor of metaphor novelty than LM surprisal, with the surprisal-novelty link peaking early in training before declining as surprisal becomes more aligned with frequency.
citing papers explorer
-
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
Increasing LLM coding agents' reasoning effort raises cost and process complexity but does not reliably improve model quality across 140 controlled runs on networked anagram game data.
-
The Frequency Confound in Language-Model Surprisal and Metaphor Novelty
Lexical frequency is a stronger predictor of metaphor novelty than LM surprisal, with the surprisal-novelty link peaking early in training before declining as surprisal becomes more aligned with frequency.