Pith. sign in

Scaling up active testing to large language models.arXiv preprint arXiv:2508.09093

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

years

2026 5

representative citing papers

Large Language Model Selection with Limited Annotations

cs.CL · 2026-05-24 · unverdicted · novelty 7.0

SELECT-LLM is the first active model selection framework for LLMs that uses expected information gain from pairwise output similarities to minimize required annotations, reporting up to 84.78% cost reduction across 23 datasets and 156 models.

Active Testing of Large Language Models via Approximate Neyman Allocation

cs.AI · 2026-05-11 · unverdicted · novelty 7.0 · 2 refs

Proposes surrogate semantic entropy stratification followed by approximate Neyman allocation for active testing of LLMs on generative benchmarks, reporting up to 28% MSE reduction and 22.9% average budget savings versus uniform sampling.

Prediction-Powered Active Testing

stat.ML · 2026-07-09 · accept · novelty 6.0

PPAT residualizes losses via a prediction-powered control variate inside LURE, yielding lower-variance unbiased risk estimates, tailored acquisition, and asymptotic CIs that cover with fewer labels.

citing papers explorer

Showing 5 of 5 citing papers.