REVIEW 6 cited by
When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In "Embers of Autoregression" (McCoy et al., 2023), we showed that several large language models (LLMs) have some important limitations that are attributable to their origins in next-word prediction. Here we investigate whether these issues persist with o1, a new system from OpenAI that differs from previous LLMs in that it is optimized for reasoning. We find that o1 substantially outperforms previous LLMs in many cases, with particularly large improvements on rare variants of common tasks (e.g., forming acronyms from the second letter of each word in a list, rather than the first letter). Despite these quantitative improvements, however, o1 still displays the same qualitative trends that we observed in previous systems. Specifically, o1 -- like previous LLMs -- is sensitive to the probability of examples and tasks, performing better and requiring fewer "thinking tokens" in high-probability settings than in low-probability ones. These results show that optimizing a language model for reasoning can mitigate but might not fully overcome the language model's probability sensitivity.
Forward citations
Cited by 6 Pith papers
-
LogiDebrief: A Signal-Temporal Logic based Automated Debriefing Approach with Large Language Models Integration
LogiDebrief automates 9-1-1 call debriefing by wrapping LLM yes/no checks in signal temporal logic specifications, and reports accurate results on real and simulated calls.
-
Thinking beyond the anthropomorphic paradigm benefits LLM research
Anthropomorphic language and assumptions are common and growing in LLM research, and the authors propose a framework for moving beyond them while keeping what is useful.
-
What is a Number, That a Large Language Model May Know It?
LLM similarity ratings over number pairs are best explained by combining Levenshtein string edit distance with a log-linear numerical distance, indicating entangled string and numeric representations.
-
Evaluating the Robustness of Analogical Reasoning in Large Language Models
GPT models solve original analogy tasks but fail many simple variants that humans handle easily, showing their analogy performance is not robust.
-
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence
Reliable deductive reasoning in AI requires replacing average-case statistical objectives with the exact learning criterion of universal correctness, a thesis supported by sample-complexity lower bounds showing statis...
-
Towards Contamination Resistant Benchmarks
The authors define contamination resistance as a benchmark property and show that most tested LLMs score near zero on Caesar-cipher encoding and decoding when the shift is not 3 and the text is random nonsense.
Discussion (0). Continue with ORCID to comment.