StackEval and StackUnseen offer a multi-language, multi-task coding benchmark from Stack Overflow, plus a human-validated analysis of LLM judges for coding answers.
Large Enough — mistral.ai
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
StackEval: Benchmarking LLMs in Coding Assistance
StackEval and StackUnseen offer a multi-language, multi-task coding benchmark from Stack Overflow, plus a human-validated analysis of LLM judges for coding answers.