The VIVID benchmark shows that current LLMs, including GPT-4o, interpret Vietnamese idioms and proverbs at less than half of the maximum score, exposing a cultural competence gap.
Entries with clearly inaccurate or culturally inappropriate labels were flagged for removal, while uncer- tain cases were marked for collaborative dis- cussion
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP
The VIVID benchmark shows that current LLMs, including GPT-4o, interpret Vietnamese idioms and proverbs at less than half of the maximum score, exposing a cultural competence gap.