Pith. sign in

Line goes up? inherent limitations of benchmarks for evaluating large language models.arXiv preprint arXiv:2502.14318, 2025

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 2 cs.CL 2

years

2026 4

roles

background 1

polarities

background 1

representative citing papers

Latent Performance Profiling of Large Language Models

cs.CL · 2026-05-28 · unverdicted · novelty 7.0

Introduces Latent Performance Profiling (LPP) as a task-agnostic framework deriving scalar metrics from LLM latent representations and dynamics to complement benchmark evaluations.

citing papers explorer

Showing 4 of 4 citing papers.