3LM provides open Arabic benchmarks for native and synthetic STEM multiple-choice questions and translated HumanEval/MBPP code tasks, with evaluations of 40 models.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
3LM: Bridging Arabic, STEM, and Code through Benchmarking
3LM provides open Arabic benchmarks for native and synthetic STEM multiple-choice questions and translated HumanEval/MBPP code tasks, with evaluations of 40 models.