Transformers trained on hierarchically nested tasks favor the least complex hypothesis that fits the in-context examples, consistent with an in-context Bayesian Occam's razor.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
Transformers trained on hierarchically nested tasks favor the least complex hypothesis that fits the in-context examples, consistent with an in-context Bayesian Occam's razor.