C2LEVA is a bilingual, multi-task LLM benchmark that combines passive test-data renewal with active data watermarking to reduce contamination risk, and ranks 15 models.
In each cluster, theses whose cosine similarities to the centroid are of the lowest 25% portion are treated as outliers and discarded
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
C$^2$LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation
C2LEVA is a bilingual, multi-task LLM benchmark that combines passive test-data renewal with active data watermarking to reduce contamination risk, and ranks 15 models.