Pith. sign in

REVIEW 2 cited by

Data Efficient Evaluation of Large Language Models and Text-to-Image Models via Adaptive Sampling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15527 v1 pith:K2VIYLOF submitted 2024-06-21 cs.LG cs.CL

classification cs.LGcs.CL
keywords benchmarksmodelssamplingevaluationacrosstext-to-imagesublimeadaptive
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Evaluating LLMs and text-to-image models is a computationally intensive task often overlooked. Efficient evaluation is crucial for understanding the diverse capabilities of these models and enabling comparisons across a growing number of new models and benchmarks. To address this, we introduce SubLIME, a data-efficient evaluation framework that employs adaptive sampling techniques, such as clustering and quality-based methods, to create representative subsets of benchmarks. Our approach ensures statistically aligned model rankings compared to full datasets, evidenced by high Pearson correlation coefficients. Empirical analysis across six NLP benchmarks reveals that: (1) quality-based sampling consistently achieves strong correlations (0.85 to 0.95) with full datasets at a 10\% sampling rate such as Quality SE and Quality CPD (2) clustering methods excel in specific benchmarks such as MMLU (3) no single method universally outperforms others across all metrics. Extending this framework, we leverage the HEIM leaderboard to cover 25 text-to-image models on 17 different benchmarks. SubLIME dynamically selects the optimal technique for each benchmark, significantly reducing evaluation costs while preserving ranking integrity and score distribution. Notably, a minimal sampling rate of 1% proves effective for benchmarks like MMLU. Additionally, we demonstrate that employing difficulty-based sampling to target more challenging benchmark segments enhances model differentiation with broader score distributions. We also combine semantic search, tool use, and GPT-4 review to identify redundancy across benchmarks within specific LLM categories, such as coding benchmarks. This allows us to further reduce the number of samples needed to maintain targeted rank preservation. Overall, SubLIME offers a versatile and cost-effective solution for the robust evaluation of LLMs and text-to-image models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RefactorAssist: Agentic Refinement for Reliable Code Refactoring

    cs.SE 2026-08 conditional novelty 6.0 of 10

    An agentic repair pipeline with static fixes and test-guided LLM iterations raises the cumulative pass rate of LLM-generated Java refactorings to about 94%, though it does not verify that the intended refactoring was ...

  2. Cer-Eval: Certifiable and Cost-Efficient Evaluation Framework for LLMs

    stat.ML 2025-05 conditional novelty 6.0 of 10

    Cer-Eval certifies LLM evaluation with confidence intervals while using 20 to 40 percent fewer test points on MMLU, AlpacaEval, and MATH.

Pith tools