LLMs are less accurate when asked to report facts in non-default measurement systems, and chain-of-thought restores accuracy only at a 180-300 percent increase in test-time compute.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
On Generalization across Measurement Systems: LLMs Entail More Test-Time Compute for Underrepresented Cultures
LLMs are less accurate when asked to report facts in non-default measurement systems, and chain-of-thought restores accuracy only at a 180-300 percent increase in test-time compute.