Pith. sign in

Evaluating judges as evaluators: The jetts benchmark of llm-as-judges as test-time scaling evaluators

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

fields

cs.AI 3 cs.CL 1

years

2026 3 2025 1

representative citing papers

Counsel: A Meta-Evaluation Dataset for Agentic Tasks

cs.AI · 2026-06-19 · unverdicted · novelty 7.0

Counsel is a new dataset of LLM-generated process critiques on agent benchmarks paired with human labels on error location and reasoning quality, achieving 0.78 Krippendorff alpha.

citing papers explorer

Showing 4 of 4 citing papers.