FarsEval-PKBETS is a new human-reviewed Persian benchmark of 4,000 questions where three tested models score below 50% correctly.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models
FarsEval-PKBETS is a new human-reviewed Persian benchmark of 4,000 questions where three tested models score below 50% correctly.