Even the best tested model, GPT4o-mini, falls more than 20 accuracy points short of the estimated human ceiling on a multiple-choice test and is fully correct in only about one-fourth of open-ended explanations.
English: Statement: A demonization effort that fails to take down Berlusconi, not even by using bad press, bad newspapers, think of Repubblica, but not only, also Corriere
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
They want to pretend not to understand: The Limits of Current LLMs in Interpreting Implicit Content of Political Discourse
Even the best tested model, GPT4o-mini, falls more than 20 accuracy points short of the estimated human ceiling on a multiple-choice test and is fully correct in only about one-fourth of open-ended explanations.