Pith. sign in

REVIEW 4 cited by

BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.12974 v3 pith:OQ27JHWA submitted 2024-10-16 cs.CL

BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks

classification cs.CL
keywords benchmarkbenchmarkcardsbenchmarksllmsdatadocumentationlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation methods, as models exhibit varying capabilities across different domains. However, finding suitable benchmarks is difficult given the many available options. This complexity not only increases the risk of benchmark misuse and misinterpretation but also demands substantial effort from LLM users, seeking the most suitable benchmarks for their specific needs. To address these issues, we introduce \texttt{BenchmarkCards}, an intuitive and validated documentation framework that standardizes critical benchmark attributes such as objectives, methodologies, data sources, and limitations. Through user studies involving benchmark creators and users, we show that \texttt{BenchmarkCards} can simplify benchmark selection and enhance transparency, facilitating informed decision-making in evaluating LLMs. Data & Code: https://github.com/SokolAnn/BenchmarkCards

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

    cs.AI 2026-06 unverdicted novelty 7.0

    Introduces the first community-governed unified JSON schema and crowdsourced repository for AI evaluation results, with converters and a database spanning 22,235 models and 2,273 benchmarks.

  2. Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains

    cs.CY 2026-04 unverdicted novelty 6.0

    Benign fine-tuning of foundation models induces large, heterogeneous, and often contradictory changes in safety metrics across general and domain-specific benchmarks.

  3. AdaQE-CG: Adaptive Query Expansion for Web-Scale Generative AI Model and Data Card Generation

    cs.AI 2026-03 unverdicted novelty 6.0

    AdaQE-CG uses context-aware adaptive query expansion and inter-card knowledge transfer from a MetaGAI Pool to generate higher-quality model and data cards than prior methods, validated on the new expert-annotated Meta...

  4. Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends

    cs.RO 2026-06 unverdicted novelty 3.0

    Automation in embodied benchmark construction shifts costs from acquisition toward validation, auditability, version control, and long-term governance instead of simply lowering total cost.