Can large language models be trusted for evaluation? scalable meta-evaluation of llms as evaluators via agent debate
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
1 Pith paper cite this work. Polarity classification is still indexing.