arXiv preprint arXiv:2303.17557 , year =

Recognition, Recall · arXiv 2303.17557

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

representative citing papers

Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA

cs.CL · 2026-06-26 · unverdicted · novelty 6.0

Generation accuracy exceeds self-evaluation on three of four in-context QA benchmarks, with attention analysis showing evaluators attend 3-5x less to context and barely read the candidate answer.

citing papers explorer

Showing 1 of 1 citing paper.

Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA cs.CL · 2026-06-26 · unverdicted · none · ref 5
Generation accuracy exceeds self-evaluation on three of four in-context QA benchmarks, with attention analysis showing evaluators attend 3-5x less to context and barely read the candidate answer.

arXiv preprint arXiv:2303.17557 , year =

fields

years

verdicts

representative citing papers

citing papers explorer