Pith. sign in

REVIEW 3 major objections 4 minor 60 references

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

T0 review · 3 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Standardized commonsense benchmarks predict downstream performance unevenly: rankings transfer most consistently to False Beliefs and TRIP, while reworked benchmarks do not improve criterion validity.

desk verdict A solid, honest meta-evaluation whose headline 'narrow transfer' result is partly confounded by tiny filtered downstream sets; worth refereeing with a demand for reliability analyses. read the letter →

arxiv 2608.03340 v1 pith:B5FVRDY5 submitted 2026-08-04 cs.CL

classification cs.CL
keywords commonsensereasoningbenchmarkvaliditycriterionLLMevaluationmodelrankingspragmaticinferencedownstreamtransferrevision
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a high score on standardized commonsense benchmarks tells you anything about how a model will behave on tasks that require implicit social, pragmatic, temporal, or physical reasoning. The authors evaluate 23 open-weight models on four established benchmarks, four reworked versions, three non-commonsense controls, and eight downstream tasks, then compare model rankings and test whether benchmark scores predict downstream performance for models from unseen families. The answer is that benchmark rankings transfer unevenly: they generalize most consistently to False Beliefs and TRIP, provide partial or metric-specific signals for a few other tasks, and fail for CEI and Sarcasm. Revised, 'cleaner' benchmarks keep the original rankings and do not improve downstream prediction. So a model's commonsense-benchmark rank is task-dependent evidence, not broad evidence of commonsense competence.

What carries the argument

The central machinery is a cross-model criterion-validity comparison: Spearman correlations between benchmark and downstream rankings, partial Spearman correlations that residualize log model size, model family, and non-commonsense control scores, and leave-one-family-out cross-validation with ridge regression. A uniform log-likelihood multiple-choice scorer with label-prior calibration makes rankings comparable across heterogeneous tasks. The design tests whether benchmark-induced ordering of models transfers to external tasks and whether it adds information beyond general capability.

What would settle it

Run the same benchmark-to-downstream prediction with downstream tasks at their original, unfiltered sizes (keeping all CEI and TimeDial items) and with at least 1,000 items per task; if CEI and Sarcasm correlations turn positive and leave-one-family-out R2 rises above zero, the narrow-transfer result is an artifact of manual filtering. Alternatively, if the HellaSwag–TRIP partial correlation survives FDR correction when family and size are matched more tightly, the claim of no incremental validity beyond general capability fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that standardized commonsense benchmarks have task-dependent and uneven downstream usefulness. Across the suite, model rankings induced by WinoGrande, HellaSwag, Social IQA, and Physical IQA (and their reworked variants) correlate most strongly and consistently with False Beliefs and TRIP; Presupposition shows a smaller but robust positive signal; Implicature, Indirect Requests, and TimeDial show metric-specific gains; CEI and Sarcasm remain unpredicted. Once model size, family, and non-commonsense control performance are accounted for, no extended partial correlation survives correction, and leave-one-family-out prediction improves only for the same narrow subse

Load-bearing premise

The eight downstream tasks, several of which have very few items after manual filtering (CEI n=31, TimeDial n=60), are valid, representative operationalizations of the commonsense competence the benchmarks claim to measure; if those criteria are noisy or unrepresentative, the observed narrow transfer is an artifact of criterion quality rather than a property of the benchmarks.

Editorial extensions

If this is right

  • A model that tops a commonsense leaderboard will often be six rank positions away from the top model on a downstream task; top-three overlap averages below 25%.
  • Cleaning or paraphrasing a benchmark is not enough: reworked benchmarks preserve rankings but do not make scores more predictive of downstream behavior.
  • Benchmark scores should be reported as evidence about that benchmark, not as a measure of broad commonsense, pragmatic, or social reasoning ability.
  • Accuracy and macro-F1 can disagree: some tasks, such as Implicature, Indirect Requests, and TimeDial, show gains only on one metric, so predictive-validity claims need to state the metric.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If this narrow-transfer pattern holds beyond the eight tasks studied, leaderboard-based model selection for real applications should give way to task-specific evaluation; the paper's data suggest a wrong pick is common.
  • The strong transfer to False Beliefs and TRIP may reflect shared surface formats or shared simple-inference demands rather than a deep commonsense faculty; a testable extension is to rewrite those tasks in open-ended form and see whether transfer collapses.
  • The WinoWhat case, where the revision changes rankings but not downstream prediction, suggests that artifact filtering and construct validity can move independently; future revision studies should report both rather than internal quality alone.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper asks whether widely used commonsense benchmarks have criterion validity: do model scores on these benchmarks predict performance on downstream tasks that require implicit social, pragmatic, temporal, or physical reasoning? The authors evaluate 23 open-weight LLMs from six families on four benchmark/revision pairs (WinoGrande/WinoWhat, HellaSwag/GoldenSwag, Social IQA/filtered Social IQA, Physical IQA/Physical IQA-RUP), three non-commonsense controls, and eight downstream tasks grouped into social/emotional, pragmatic, and event/physical axes. Using Spearman correlations, partial correlations controlling for model size, family, and controls, bootstrap confidence intervals, and leave-one-family-out ridge regression, they report that reworked benchmarks largely preserve original model rankings and do not improve downstream prediction; that consistent transfer is limited to False Beliefs and TRIP, with metric-specific gains for Presupposition, Implicature, Indirect Requests, and TimeDial; and that axis-specific transfer is weak. The paper concludes that commonsense benchmarks have task-dependent and uneven downstream usefulness and should not be read as broad evidence of pragmatic or contextual reasoning competence.

Significance. If the result holds, it is practically important: benchmark scores need external validation before being used for model selection or capability claims, and dataset repair alone does not automatically create criterion validity. The study has clear strengths: uniform log-likelihood scoring across heterogeneous tasks, inclusion of non-commonsense controls, bootstrap-based uncertainty estimation, FDR correction, leave-one-family-out prediction, and public code/data. It engages directly with the current benchmarking-crisis literature. However, the evidential core is narrower than the abstract and first paragraphs suggest. The most robust positive associations are raw correlations to two of eight downstream tasks; partial correlations are uniformly non-significant after FDR correction; LOFOCV uses only six family folds; and some downstream tasks have very few items after manual filtering. These limitations are acknowledged in Section 6 but are not quantitatively addressed. The paper remains a clearly reported and valuable empirical contribution, but the headline claim about narrow transfer needs additional reliability analysis or a more cautious formulation.

major comments (3)
  1. [§3.2, Table 2, Appendix B] The central negative result—that transfer is limited to False Beliefs and TRIP—is vulnerable to criterion-score attenuation. After manual filtering, CEI has 31 items and TimeDial 60; Indirect Requests has about 64 (Table 2; Appendix B). Spearman correlations across 23 models computed on such short tasks have low reliability, and classical attenuation compresses correlations toward zero. The two tasks with the most consistent associations, False Beliefs (≈192) and TRIP (≈100), are among the larger sets, so the observed task-dependence may reflect test length rather than construct validity. The paper reports no split-half reliability, no item-level bootstrap by task, and no sensitivity analysis around the manual filtering decisions. Please add reliability estimates for each downstream task (e.g., split-half or rank split, bootstrap CIs for per-task model scores) and either use larger unfil
  2. [§4.3, Figure 3] Section 4.3 reports that after controlling for model size, family, and non-CS controls, no partial Spearman correlation survives FDR correction, while the positive LOFOCV gains for False Beliefs and TRIP are based on only six family folds and are presented descriptively. Given 23 models clustered in six families, the evidence for 'consistent cross-family predictive validity' is thinner than the wording in Sections 4.2–4.3 suggests. The manuscript flags this in Section 6, but the conclusion still leans on raw correlations. To make the positive claims load-bearing, please present family-level scatterplots or bootstrap distributions for the two key tasks, report the per-family LOFOCV predictions rather than only pooled R2, and explicitly state the effective sample size and its consequences. If this is infeasible, the abstract and conclusion should downgrade 'consistent cross-family' to 'sug
  3. [§4.4, Figure 4] The axis-specific analysis is underpowered and the axes themselves are constructed by the authors. Panel (b) omits the pragmatic axis because all CS benchmarks are considered potentially relevant, so the global comparison mixes different question types; with only a handful of tasks per axis the bootstrap CIs are wide. The conclusion that 'transfer does not follow domains' may be correct, but the current test cannot cleanly separate a true absence of domain specificity from low power or arbitrary axis assignment. A clearly justified axis assignment (or a preregistered alternative) plus an item-level task-by-task analysis would strengthen this claim. At minimum, the wording should reflect that the axis test is exploratory.
minor comments (4)
  1. [Figure 1 caption] 'common sense benchmarks' should be 'commonsense benchmarks' for consistency with the rest of the manuscript.
  2. [§3.2 / Appendix B] The manual filtering description gives only final n, not the number of removed items per task or the inter-annotator agreement if multiple annotators were involved. Adding a table with original n, removed n, final n, and annotation agreement would help readers assess the reliability of the filtered criteria.
  3. [§4.2 / Appendix C] The sentence 'the top-three models overlap never exceeds an average Jaccard score of .25' is confusing. Appendix C reports mean top-3 Jaccard values averaged over benchmark–task pairs, not a maximum. Please rephrase to indicate that the mean Jaccard overlap is around .25 and clarify whether this is averaged over accuracy, macro-F1, or both.
  4. [Eq. (2)] The sign convention for Δρ is stated in text but the equation itself would benefit from an explicit 'positive values favor the original benchmark' annotation. Also, the text uses 'ρOG' and 'ρRW' without defining these acronyms in the equation block.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is an empirical correlational study using external downstream criteria; same-group artifacts are test material, not load-bearing evidence.

full rationale

The paper makes no formal derivation; its central claims — reworked benchmarks preserve rankings, downstream transfer is narrow, gains concentrate on False Beliefs and TRIP, and axis-specificity is weak — are empirical estimates from Spearman correlations, partial correlations, and leave-one-family-out cross-validation. The downstream criteria (CEI, SARC7, False Beliefs, Implicature, Presupposition, Indirect Requests, TimeDial, TRIP) are external datasets, not constructed from the benchmark scores being validated; no equation defines a downstream score as a function of a benchmark score. The LOFOCV procedure explicitly withholds entire model families and performs standardization and fitting on the training families only, so the reported cross-family predictions are not fitted on the target data. The only same-group artifacts are WinoWhat (Gevers et al., 2025a), used as one of eight benchmark variants, and the prior critique (Gevers et al., 2025b), cited among several references documenting benchmark validity concerns. Neither is load-bearing: the finding that WinoWhat correlates less strongly with WinoGrande is an empirical result of the evaluation, not an input assumption, and the critique citation merely motivates the study rather than supplying a premise that forces the conclusions. The manual filtering described in Appendix B is a measurement-reliability concern (small n, potential selection effects), not a circularity, because the filtering does not make downstream scores a deterministic function of benchmark scores. No self-definitional, fitted-input, imported-uniqueness, ansatz-smuggling, or renaming pattern is present. The paper is self-contained against external benchmarks, so the honest finding is no significant circularity.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

No invented entities are introduced. The free parameters are limited to a fixed ridge penalty. The axioms are the stated domain assumptions that the benchmark-to-downstream transfer should hold if benchmarks measure reusable abilities, and that the chosen tasks and metric operationalize the constructs. These are explicit in Sections 2, 3.2, and 3.6.

free parameters (1)
  • ridge L2 penalty alpha = 1.0
    Fixed by hand in the LOFOCV ridge regression (Appendix B); not tuned to test data, but a modeling choice that affects the reported cross-validated R2 values.
assumptions (5)
  • domain assumption Commonsense reasoning is a broad body of everyday background knowledge about physical objects, events, social relations, beliefs, goals, emotions, norms, and likely consequences.
    Section 2; motivates why benchmarks and downstream tasks should relate.
  • domain assumption If a commonsense benchmark measures reusable reasoning ability, models that perform well on it should also rank highly on downstream tasks requiring related implicit inference.
    Section 3.6; this transfer assumption is the theoretical basis for the validity tests.
  • domain assumption Log-likelihood-based multiple-choice scoring is a valid uniform evaluation metric across benchmarks and downstream tasks.
    Section 3.2 and Limitations; the authors restrict to this metric for comparability and acknowledge it has flaws.
  • domain assumption The chosen downstream tasks (CEI, SARC7, False Beliefs, Implicature, Presupposition, Indirect Requests, TimeDial, TRIP) are valid criteria for the intended commonsense constructs.
    Section 3.2; relies on the published validity of these datasets, some of which are very small after filtering (e.g., CEI n=31).
  • domain assumption Spearman correlation across 23 models from 6 families is a meaningful basis for criterion-validity inference.
    Section 3.5; with only 6 families, family-level confounds can dominate, and the authors describe results as partly descriptive.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks." pith.science (2026). https://pith.science/paper/B5FVRDY5

@misc{pith2026260803340,
  author       = {Pith},
  title        = {Pith review of: Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B5FVRDY5}},
  note         = {Machine review of arXiv:2608.03340}
}
read the original abstract

Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified. To establish the practical usability of widely adopted commonsense benchmarks, we evaluate 23 models from six families on four established commonsense benchmarks, four reworked variants, three non-commonsense controls, and eight downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning. We compare model rankings, compute controlled correlations, and use leave-one-family-out cross-validation to assess the criterion validity of commonsense benchmarks. Our results show that revised benchmarks largely preserve original model rankings and do not improve downstream predictive power. Commonsense benchmarks show consistent cross-family predictive validity for only a narrow subset of downstream tasks, with smaller or metric-specific gains elsewhere. Overall, standardized commonsense benchmarks provide task-dependent rather than broad evidence of downstream commonsense competence.

Figures

Figures reproduced from arXiv: 2608.03340 by the authors.

Figure 1
Figure 1. Correlation matrix heatmap for original and [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 3
Figure 3. Change in LOFOCV [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Axis-specific criterion-validity tests. Panel (a) compares same- and cross-axis downstream prediction; [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 41 canonical work pages

  1. [1]

    What Are LLM Benchmarks? , year =

  2. [2]

    LLM Benchmarks and Leaderboards , year =

  3. [3]

    2026 , month = apr, url =

    Hao Wang and Qiuyang Mang and Alvin Cheung and Koushik Sen and Dawn Song , title =. 2026 , month = apr, url =

  4. [4]

    When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

    Alzahrani, Norah and Alyahya, Hisham and Alnumay, Yazeed and AlRashed, Sultan and Alsubaie, Shaykhah and Almushayqih, Yousef and Mirza, Faisal and Alotaibi, Nouf and Al-Twairesh, Nora and Alowisheq, Areeb and Bari, M Saiful and Khan, Haidar. When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards. Proceedings of the 62n...

  5. [5]

    Advances in Neural Information Processing Systems , volume=

    Measuring what matters: Construct validity in large language model benchmarks , author=. Advances in Neural Information Processing Systems , volume=

  6. [6]

    2026 , month = apr, url =

    Sha Sajadieh and Loredana Fattorini and Raymond Perrault and Yolanda Gil and Vanessa Parli and Lapo Santarlasci and Juan Pava and Nestor Maslej and Russ Altman and Erik Brynjolfsson and Carla Brodley and Jack Clark and Virginia Dignum and Vipin Kumar and James Landay and Terah Lyons and James Manyika and Juan Carlos Niebles and Yoav Shoham and Elham Tabas...

  7. [7]

    Or Not? , author=

    In Benchmarks We Trust... Or Not? , author=. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

  8. [8]

    How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the W inograd Schema Challenge and SWAG

    Trichelair, Paul and Emami, Ali and Trischler, Adam and Suleman, Kaheer and Cheung, Jackie Chi Kit. How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the W inograd Schema Challenge and SWAG. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language P...

Show all 60 references
  1. [9]

    Benchmark\^

    Qian, Qi and Huang, Chengsong and Xu, Jingwen and Lv, Changze and Wu, Muling and Liu, Wenhao and Wang, Xiaohua and Wang, Zhenghua and Huang, Zisu and Tian, Muzhao and others , journal=. Benchmark\^

  2. [10]

    Findings of the Association for Computational Linguistics: ACL 2025 , pages=

    Right answer, wrong score: Uncovering the inconsistencies of LLM evaluation in multiple-choice question answering , author=. Findings of the Association for Computational Linguistics: ACL 2025 , pages=

  3. [11]

    arXiv preprint arXiv:2602.10657 , year=

    Benchmarks Are Not That Out of Distribution: Word Overlap Predicts Performance , author=. arXiv preprint arXiv:2602.10657 , year=

  4. [12]

    arXiv preprint arXiv:2602.06221 , year=

    BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks , author=. arXiv preprint arXiv:2602.06221 , year=

  5. [13]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Which of these best describes multiple choice evaluation with llms? a) forced b) flawed c) fixable d) all of the above , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  6. [14]

    My answer is C

    “My answer is C”: First-token probabilities do not match text answers in instruction-tuned language models , author=. Findings of the Association for Computational Linguistics: ACL 2024 , pages=

  7. [15]

    Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

    Plausibly problematic questions in multiple-choice benchmarks for commonsense reasoning , author=. Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

  8. [16]

    PNAS nexus , volume=

    A large-scale evaluation of commonsense knowledge in humans and large language models , author=. PNAS nexus , volume=. 2026 , publisher=

  9. [17]

    Scientific Reports , volume=

    A noise audit of human-labeled benchmarks for machine commonsense reasoning , author=. Scientific Reports , volume=. 2024 , publisher=

  10. [18]

    arXiv preprint arXiv:2605.20520 , year=

    Open-world evaluations for measuring frontier AI capabilities , author=. arXiv preprint arXiv:2605.20520 , year=

  11. [19]

    arXiv preprint arXiv:2502.14359 , year=

    Triangulating llm progress through benchmarks, games, and cognitive tests , author=. arXiv preprint arXiv:2502.14359 , year=

  12. [20]

    Proceedings of the Teddington Conference on the Mechanization of Thought Processes , pages =

    Programs with Common Sense , author =. Proceedings of the Teddington Conference on the Mechanization of Thought Processes , pages =. 1959 , publisher =

  13. [21]

    Communications of the ACM , volume=

    Commonsense reasoning and commonsense knowledge in artificial intelligence , author=. Communications of the ACM , volume=. 2015 , publisher=

  14. [22]

    , author=

    The Winograd schema challenge. , author=. KR , volume=

  15. [23]

    Communications of the ACM , volume=

    Winogrande: An adversarial winograd schema challenge at scale , author=. Communications of the ACM , volume=. 2021 , publisher=

  16. [24]

    Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

    Hellaswag: Can a machine really finish your sentence? , author=. Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

  17. [25]

    Bisk, Yonatan and Zellers, Rowan and Gao, Jianfeng and Choi, Yejin and others , booktitle=

  18. [26]

    Sap, Maarten and Rashkin, Hannah and Chen, Derek and Le Bras, Ronan and Choi, Yejin , booktitle=. Social

  19. [27]

    Sap, Maarten and Le Bras, Ronan and Allaway, Emily and Bhagavatula, Chandra and Lourie, Nicholas and Rashkin, Hannah and Roof, Brendan and Smith, Noah A and Choi, Yejin , booktitle=

  20. [28]

    Proceedings of the National Academy of Sciences , volume=

    Evaluating large language models in theory of mind tasks , author=. Proceedings of the National Academy of Sciences , volume=. 2024 , publisher=

  21. [29]

    Nature human behaviour , volume=

    Testing theory of mind in large language models and humans , author=. Nature human behaviour , volume=. 2024 , publisher=

  22. [30]

    Proceedings of the Second Workshop on Insights from Negative Results in NLP , pages=

    Does commonsense help in detecting sarcasm? , author=. Proceedings of the Second Workshop on Insights from Negative Results in NLP , pages=

  23. [31]

    Chun, Jon and Sussman, Hannah and Mangine, Adrian and Kocaman, Murathan and Sidorko, Kirill and Koirala, Abhigya and McCloud, Andre and Eisenbeis, Gwen and Akanwe, Wisdom and Gassama, Moustapha and others , journal=

  24. [32]

    Speech acts , pages=

    Logic and conversation , author=. Speech acts , pages=. 1975 , publisher=

  25. [33]

    Linguistics and philosophy , volume=

    Common ground , author=. Linguistics and philosophy , volume=. 2002 , publisher=

  26. [34]

    The Stanford Encyclopedia of Philosophy , editor=

    Common Ground in Pragmatics , author=. The Stanford Encyclopedia of Philosophy , editor=. 2024 , publisher=

  27. [35]

    Proceedings of the 21st International Conference on Natural Language Processing (ICON) , pages=

    Survey on Computational Approaches to Implicature , author=. Proceedings of the 21st International Conference on Natural Language Processing (ICON) , pages=

  28. [36]

    Sravanthi, Settaluri and Doshi, Meet and Tankala, Pavan and Murthy, Rudra and Dabre, Raj and Bhattacharyya, Pushpak , booktitle=

  29. [37]

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Pragmatics in the era of large language models: A survey on datasets, evaluation, opportunities and challenges , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  30. [38]

    TIMEDIAL: Temporal commonsense reasoning in dialog , author=. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , pages=

  31. [39]

    Findings of the Association for Computational Linguistics: EMNLP 2021 , pages=

    Tiered reasoning for intuitive physics: Toward verifiable commonsense language understanding , author=. Findings of the Association for Computational Linguistics: EMNLP 2021 , pages=

  32. [40]

    Advances in Neural Information Processing Systems , volume=

    Fantastic bugs and where to find them in AI benchmarks , author=. Advances in Neural Information Processing Systems , volume=

  33. [41]

    arXiv preprint arXiv:2602.17594 , year=

    Ai gamestore: Scalable, open-ended evaluation of machine general intelligence with human games , author=. arXiv preprint arXiv:2602.17594 , year=

  34. [42]

    International conference on machine learning , pages=

    Calibrate before use: Improving few-shot performance of language models , author=. International conference on machine learning , pages=. 2021 , organization=

  35. [43]

    Are we done with mmlu? , author=. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

  36. [44]

    IEEE Access , year=

    Beyond the leaderboard: A survey of the science of evaluation, benchmarking, and methodologies for large language models , author=. IEEE Access , year=

  37. [45]

    Proceedings of the 29th Conference on Computational Natural Language Learning , pages=

    WinoWhat: A parallel corpus of paraphrased WinoGrande sentences with common sense categorization , author=. Proceedings of the 29th Conference on Computational Natural Language Learning , pages=

  38. [46]

    arXiv preprint arXiv:2504.07825 , year=

    What the hellaswag? on the validity of common-sense reasoning benchmarks , author=. arXiv preprint arXiv:2504.07825 , year=

  39. [47]

    Findings of the Association for Computational Linguistics: EACL 2026 , pages=

    Garbage in, reasoning out? why benchmark scores are unreliable and what to do about it , author=. Findings of the Association for Computational Linguistics: EACL 2026 , pages=

  40. [48]

    arXiv preprint arXiv:2406.11020 , year=

    Rupbench: Benchmarking reasoning under perturbations for robustness evaluation in large language models , author=. arXiv preprint arXiv:2406.11020 , year=

  41. [49]

    arXiv preprint arXiv:2311.12022 , year=

    Gpqa: A graduate-level google-proof q&a benchmark , author=. arXiv preprint arXiv:2311.12022 , year=

  42. [50]

    Transactions of the Association for Computational Linguistics , volume=

    BLiMP: The benchmark of linguistic minimal pairs for English , author=. Transactions of the Association for Computational Linguistics , volume=. 2020 , publisher=

  43. [51]

    Innovative Journal of Applied Science , pages=

    Evaluating Log-Likelihood for Confidence Estimation in LLM-Based Multiple-Choice Question Answering , author=. Innovative Journal of Applied Science , pages=

  44. [52]

    Expert Systems , volume =

    An Experimental Study Measuring the Generalization of Fine-Tuned Language Representation Models Across Commonsense Reasoning Benchmarks , author =. Expert Systems , volume =. 2023 , doi =

  45. [53]

    2026 , eprint =

    The Magic Correlations: Understanding Knowledge Transfer from Pretraining to Supervised Fine-Tuning , author =. 2026 , eprint =

  46. [54]

    Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics , year =

    On Emergent Social World Models: Evidence for Functional Integration of Theory of Mind and Pragmatic Reasoning in Language Models , author =. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics , year =

  47. [55]

    Journal of the Royal Statistical Society: Series B (Methodological) , volume =

    Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing , author =. Journal of the Royal Statistical Society: Series B (Methodological) , volume =. 1995 , doi =

  48. [56]

    AI Magazine , volume=

    Six principles for evaluating cognitive capabilities in AI models , author=. AI Magazine , volume=. 2026 , publisher=

  49. [57]

    Transactions of the Association for Computational Linguistics , volume=

    A statistical analysis of summarization evaluation metrics using resampling methods , author=. Transactions of the Association for Computational Linguistics , volume=. 2021 , publisher=

  50. [58]

    Proceedings of the 9th Widening NLP Workshop , pages=

    Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques , author=. Proceedings of the 9th Widening NLP Workshop , pages=

  51. [59]

    Transactions of the Association for Computational Linguistics , volume=

    Comparing humans and large language models on an experimental protocol inventory for theory of mind evaluation (EPITOME) , author=. Transactions of the Association for Computational Linguistics , volume=

  52. [60]

    Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

    Are natural language inference models IMPPRESsive? Learning IMPlicature and PRESupposition , author=. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.