Pith. sign in

REVIEW 2 major objections 3 minor 14 cited by

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

T0 review · 2 major / 3 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read The paper argues that evaluating generative AI systems is essentially a social science measurement challenge, and that adopting a four-level measurement framework would give the field a way to say precisely what a benchmark or safety test…

desk verdict A serious, useful synthesis of social-science measurement theory for GenAI evaluation, slightly overbroad in its claim to cover all measurement tasks, and best where the concepts are socially embedded. read the letter →

arxiv 2502.00561 v2 pith:JQDGEE2V submitted 2025-02-01 cs.CY

classification cs.CY
keywords generativeAIevaluationmeasurementtheoryvaliditysystematizationoperationalizationLLM-as-a-judgebenchmarkingsocialscience
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that evaluating generative AI systems is, at bottom, a social science measurement problem. It claims that the machine learning community jumps from vague concepts like stereotyping or refusal straight to benchmarks and judge models, without first fixing what exactly is being measured. The proposed remedy is a four-level measurement framework that separates conceptual definition from operational implementation and adds seven validity lenses for interrogating both. If the field adopted this discipline, evaluations would state precisely what they measure, comparisons across systems would be less misleading, and stakeholders beyond ML researchers could join the debate about what should count.

What carries the argument

The load-bearing mechanism is the four-level measurement framework adapted from social science measurement theory, together with the insistence that conceptual debates about what is being measured and why be kept separate from operational debates about how to measure it. The framework's core move is systematization: turning a broad, contested background concept into an explicit systematized concept that specifies observable phenomena and their relationships before any instrument is built. Operationalization then maps that definition to indicators, annotation guidelines, aggregation functions, and other measurement instruments; interrogation uses seven validity lenses to test the systematized concept, the instruments, and the measurements, and may send the process back for revision. The paper's running example of measuring text that stereotypes social groups in a chatbot's outputs shows how the framework distinguishes a judge model's annotation from the definition being measured, and why human annotations cannot be treated as automatic ground truth without first checking that they align with the systematized concept.

What would settle it

A concrete test would be to apply the framework to several existing benchmark tasks, fully systematizing each target concept, then compare the resulting measurements with those from the original instruments on the same systems; if the framework-guided instruments yield neither stronger convergent and discriminant evidence nor clearer predictions of external outcomes than the instruments they replace, the claim that this process brings rigor to GenAI evaluation loses its empirical footing. A quicker observational falsifier would be to find two well-specified real-world evaluation contexts where the same systematized concept and instruments are used but validity evidence cannot be re-established, contradicting the claim that validity is always context-dependent.

Watch

Extended reading notes

Core claim

The central claim is that the dominant way the ML community evaluates GenAI systems is measurement without a measurement theory: concepts of interest are abstract and contested, yet researchers typically move directly to datasets, prompts, and scoring rubrics without an explicit systematized definition, and rarely interrogate whether instruments and measurements are valid. The paper's proposed standard is a four-level framework, grounded in social science measurement theory, with levels for the background concept, the systematized concept, the measurement instruments, and the measurements themselves, linked by systematization, operationalization, application, and interrogation. Following this standard, evaluators would first develop an explicit definition that connects the concept to observable phenomena, then operationalize that definition into indicators and instruments, and finally use face, content, convergent, discriminant, predictive, hypothesis, and consequential validity to challenge every step. The authors directly address the objection that AI systems are not social systems by noting that the concepts at stake are deeply intertwined with people and society, not with the physical substrate of the system.

Load-bearing premise

The argument rests on the assumption that a measurement framework developed for studying human social and political concepts, with its definition of validity and its seven evidence lenses, can meaningfully be transplanted to the evaluation of machines, a transfer the paper asserts through argument and example but does not empirically demonstrate.

Editorial extensions

If this is right

  • Benchmark papers would routinely publish their systematized concept alongside datasets and rubrics, making it possible to tell when two benchmarks measure the same thing.
  • Claims like 'LLM-as-a-judge accuracy' would be reframed: agreement with humans would no longer be treated as ground truth unless the humans' annotations are shown to align with the systematized concept.
  • Validity would become a context-bound property, so instruments would be re-interrogated before each new deployment context rather than certified once.
  • Evaluation would open to more participants, since policymakers, users, and affected communities could contest the systematized concept and its consequences without needing to read code.
  • The framework would apply across measurement approaches, including benchmarks, red teaming, real-world evaluations, and user studies, not just tests that resemble psychometric instruments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If this position is adopted, a natural next step is a reporting standard: papers making evaluative claims about GenAI systems would disclose their systematized concept, indicators, aggregation functions, and validity evidence, an extension the paper gestures toward but does not formalize.
  • The reinterpretation of human-judge agreement has a sharp consequence: many published 'judge accuracy' numbers are uninterpretable as validity evidence, and re-examining them through this framework could change which models are trusted as judges.
  • The same logic applies to fairness metrics and other algorithmic measurements, since the paper notes the framework extends beyond GenAI; a testable extension would be to require an explicit systematized definition of fairness before selecting any metric.
  • A concrete prediction follows: benchmark suites built with explicit systematized concepts and validity interrogation will produce measurements that are more stable across minor changes to prompts and sampling than suites built without, because operationalization choices are constrained by the definition.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. This position paper argues that the evaluation of generative AI (GenAI) systems should be treated as a social science measurement problem. It adapts Adcock and Collier's (2001) four-level measurement framework—background concept, systematized concept, measurement instruments, and measurements—and recommends seven validity lenses drawn from Messick (1987) and Jacobs and Wallach (2021). The authors claim that separating systematization from operationalization clarifies what is being measured, that the validity lenses impose needed rigor on conceptual and operational debates, and that the framework applies to all measurement tasks involved in evaluating GenAI systems. The argument is developed through a running example (measuring stereotype prevalence in LLM outputs), abbreviated applications to mathematical reasoning, memorization, and refusal, and a discussion of adoption barriers and alternative views.

Significance. If the position is accepted, the paper could serve as a shared vocabulary and quality standard for GenAI evaluation, connecting benchmark design, red teaming, and user studies to a established body of social science measurement practice. The manuscript is careful and honest: it explicitly says the framework is not a panacea, that partial adoption is still valuable, and that it does not advocate transferring psychometric instruments to machines. Its main strengths are the detailed worked example, the transparent borrowing of an established framework, and the appendices showing how existing memorization results can be read through the validity lenses. The paper does not offer new formal results, code, or data, but as a position paper its contribution is conceptual unification and a concrete set of recommended actions.

major comments (2)
  1. [Sec. 2 and Appendix C] The central claim that the framework 'brings clarity to all measurement tasks involved in evaluating GenAI systems' (Sec. 2) is not established for constructs with checkable success criteria. In Appendix C, the mathematical reasoning example proposes as the systematized concept 'the accuracy of the model on highly challenging mathematical reasoning problems aimed at pre-university students, spanning algebra, number theory, combinatorics, and geometry'—this is already an operational quantity rather than a concept definition, so the separation of systematization from operationalization, which the paper presents as its main benefit, collapses by construction. The memorization example (Appendix C and E) is more developed, but it is largely a reinterpretation of existing measurement choices (token length, greedy versus probabilistic sampling, perplexity checks) through the validity lenses rather than a demonstration that the lenses generate evidence beyond standard benchmark validation. I recommend the authors either provide a worked systematization for at least one capability-oriented construct in which the background/systematized distinction and the validity lenses do substantive work, or revise the scope claim to concepts whose meanings are socially contested or whose measurement requires construct interpretation. Without one of these, the universal applicability claim is stronger than the evidence.
  2. [Sec. 3.2 and Appendix B] The paper's claim that validity lenses can inform conceptual debates (contra Adcock and Collier) is asserted rather than argued. Appendix B says 'we argue that three lenses—face validity, content validity, and convergent validity—are especially useful for shedding light on conceptual debates,' but no reasoning is given for this selection, and it contradicts Section 4.1.2, which lists face, content, and consequential validity as the three lenses. Because this deviation is one of the paper's stated contributions, the authors should either supply the promised argument or mark the selection as provisional.
minor comments (3)
  1. [Sec. 1, References] There is a missing space in 'Theprocess of evaluation' near the start of Section 1, and the bibliography entry for Hand (2004) misspells 'Measurement' as 'Meaurement'.
  2. [Sec. 3.1] The phrase 'researchers and practitioners appear to jump from background concepts to measurement instruments' is used as a motivation, but its evidentiary basis is not stated; please either cite a systematic analysis or qualify it as an informal observation about common practice.
  3. [Appendix C] In the mathematical reasoning example, the systematized concept should be expressed as a construct (e.g., the ability to solve pre-university-level olympiad problems in specified topic areas), with 'accuracy' reserved for the measurement level; as written, the example risks enacting the very conflation the paper criticizes.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the framework is imported from external measurement theory (Adcock and Collier; Messick), and the paper's self-citations are transparent, non-load-bearing, and do not reduce the argument to its inputs.

full rationale

The paper's central position is programmatic rather than derivational: it proposes adopting a variant of Adcock and Collier's (2001) four-level measurement framework and Messick's (1987) validity lenses. These sources are external to the authors, and the framework's definitions are not constructed from the paper's conclusions. The seven validity lenses are described as 'adapted from Messick (1987) by Jacobs & Wallach (2021)' (Section 3.2), which is a self-citation; however, the substantive content is attributed to Messick, the lenses are standard independently checkable categories, and the argument does not rely on treating Jacobs and Wallach as an unverified authority. The paper's worked examples (stereotyping text, mathematical reasoning, memorization, refusal) are instantiations of the framework, not predictions derived from fitted parameters; the claim that ML practice conflates systematization and operationalization is argued and supported by external critiques rather than by definition. The transferability of Adcock and Collier's framework from human social concepts to machine-measured constructs is asserted as a position and is a substantive limitation, but no statement in the paper equates the conclusion with an input by construction or makes a fitted value do the work of an independent prediction. Therefore no circular step is present.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper uses no free parameters and invents no entities. It relies on imported measurement theory and asserted generalizations about community practice. The main axioms are that Adcock and Collier's framework is fit for AI evaluation, that Messick's unified validity is the right standard, that ML practice is deficient in the ways claimed, and that separating systematization from operationalization yields the promised benefits. None of these are proven, but they are argued for and hedged.

assumptions (4)
  • domain assumption Adcock and Collier's four-level measurement framework is an appropriate and effective basis for standardizing GenAI evaluation.
    The paper adopts a variant of Adcock & Collier (2001) as its organizing structure (Section 3) without proving it is the best or only approach for AI evaluation.
  • domain assumption Messick's unified validity perspective, including consequential validity, is the appropriate lens for interrogating validity of AI measurements.
    The seven lenses come from Messick (1987) via Jacobs & Wallach (2021) and are assumed to transfer to GenAI evaluation (Section 3.2).
  • domain assumption ML researchers and practitioners routinely conflate concept definition with measurement and rarely interrogate validity.
    The paper asserts this generalization about current practice (Section 3.1) without systematic evidence, yet it motivates the framework.
  • domain assumption Separating systematization from operationalization will yield the claimed benefits of broader stakeholder participation and improved rigor.
    This causal claim is argued through examples but not empirically demonstrated (Sections 3.1, 4).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge." pith.science (2026). https://pith.science/paper/JQDGEE2V

@misc{pith2026250200561,
  author       = {Pith},
  title        = {Pith review of: Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JQDGEE2V}},
  note         = {Machine review of arXiv:2502.00561}
}
read the original abstract

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from learning from and drawing on the social sciences when developing and using measurement instruments for evaluating GenAI systems. Specifically, our position is that evaluating GenAI systems is a social science measurement challenge. We present a four-level framework, grounded in measurement theory from the social sciences, for measuring concepts related to the capabilities, behaviors, and impacts of GenAI systems. This framework has two important implications: First, it can broaden the expertise involved in evaluating GenAI systems by enabling stakeholders with different perspectives to participate in conceptual debates. Second, it brings rigor to both conceptual and operational debates by offering a set of lenses for interrogating validity.

Figures

Figures reproduced from arXiv: 2502.00561 by the authors.

Figure 1
Figure 1. A variant of the framework of Adcock & Collier (2001). The background concept, the systematized concept, the measurement instruments, and the measurements are linked by four processes: systematization, operationalization, application, and interrogation. also enable stakeholders with different perspectives—e.g., open-source developers, policymakers, customer, users, and members of marginalized communities, all of who… view at source ↗
Figure 2
Figure 2. The systematized concept for our running example of measuring text that stereotypes social groups. validity focuses on the extent to which the systematized concept looks reasonable. We might decide to seek input from members of different social groups—i.e., experiential experts—finding that they are not entirely comfortable with our systematized concept. We might therefore dig deeper, using content validity, which, … view at source ↗
Figure 3
Figure 3. Instantiating the levels in the framework described in Section 3 with four very different measurement tasks. of measuring the prevalence of text that stereotypes social groups in the outputs of an already deployed LLM-based chatbot, interrogating discriminant validity would involve comparing measurements obtained using our judge LLM to measurements of dissimilar concepts, such as hostile text or text with negative s… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Algorithmic Gender Prediction Is Illegitimate, But Gender Imputation Can Yield Valid Measurements

    cs.CY 2026-08 conditional novelty 7.0 of 10

    Gender prediction is illegitimate but gender imputation can still yield valid disparity measurements for traditional sexism, though not for oppositional sexism.

  2. Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs

    cs.CY 2025-02 conditional novelty 7.0 of 10

    The paper introduces DiffAware and CtxtAware metrics, an 8-benchmark suite with 16,000 questions, and shows ten LLMs are less able to recognize legitimate group differences than current fairness benchmarks suggest.

  3. What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

    cs.HC 2026-08 conditional novelty 6.0 of 10

    Chat UI and API access to the same chatbot produce different accuracy, consistency, citation, and refusal behaviors on safety benchmarks, and web search changes these patterns further.

  4. Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Base language models outperform post-trained models at generating individual survey responses that match human distributions, while post-trained models are more accurate when directly predicting population opinion dis...

  5. On the Convergent Validity of Offline Evaluation Designs for Recommender Systems

    cs.IR 2026-07 conditional novelty 6.0 of 10

    Sparse offline recommender rankings correlate only weakly—and sometimes negatively—with dense ground-truth rankings, and no evaluation design is uniformly best.

  6. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  7. Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering

    cs.SE 2026-06 unverdicted novelty 6.0 of 10

    Coding benchmarks misalign with agentic software engineering because they conflate model and harness, grade against single references, and provide no component-level iteration signals.

  8. Grounded Chess Reasoning in Language Models via Master Distillation

    cs.AI 2026-03 unverdicted novelty 6.0 of 10

    Master Distillation turns opaque chess-engine search into natural-language CoT, lifting a 4B model (C1) from near-zero to 48.1% accuracy that beats its teacher and most larger LMs.

  9. Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances

    cs.CY 2025-07 accept novelty 6.0 of 10

    A Fair Equality of Chances-based framework decomposes GenAI unfairness into harms/benefits, morally arbitrary factors, and morally decisive factors to improve measurement validity.

  10. Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Practitioners trying to measure representational harms in LLM-based systems often cannot use public measurement instruments, either because the instruments lack validity, specificity, interpretability, or actionabilit...

  11. Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions

    cs.AI 2026-08 accept novelty 5.0 of 10

    The paper sets out a research agenda for NLP to measure how prolonged language-model use changes human behavior over long time horizons, replacing single-session safety evaluations with longitudinal tracking.

  12. Against 'softmaxing' culture

    cs.HC 2025-06 unverdicted novelty 5.0 of 10

    A position paper arguing that AI evaluations should shift from defining culture to understanding when culture becomes relationally valid.

  13. Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects

    cs.CY 2025-05 conditional novelty 4.0 of 10

    A position paper argues that understanding AI's second-order effects requires moving from static benchmarks to an ecosystem of field testing, red teaming, and contextual evaluation.

  14. AI, Jobs, and the Automation Trap: Where Is HCI?

    cs.HC 2025-01 conditional novelty 3.0 of 10

    A position paper claiming AI patents overwhelmingly favor task automation over human augmentation, and proposing incentive changes to make human-centered AI more impactful.

Reference graph

Works this paper leans on

52 extracted references · 31 canonical work pages · cited by 14 Pith papers

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    Abebe, R., Barocas, S., Kleinberg, J., Levy, K., Raghavan, M., and Robinson, D. G. Roles for computing in social change . In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pp.\ 252--260, 2020

  3. [3]

    and Collier, D

    Adcock, R. and Collier, D. Measurement validity: A shared standard for qualitative and quantitative research . American Political Science Review, 95 0 (3): 0 529--546, 2001

  4. [4]

    M., Blaauw, G

    Amdahl, G. M., Blaauw, G. A., and Brooks, F. P. Architecture of the IBM S ystem/360. IBM Journal of Research and Development, 8 0 (2), 1964

  5. [5]

    Content Analysis in Communication Research

    Berelson, B. Content Analysis in Communication Research. Free Press, 1952

  6. [6]

    and Hancox-Li, L

    Blili-Hamelin, B. and Hancox-Li, L. Making Intelligence: Ethical Values in IQ and ML Benchmarks . In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, FAccT '23, pp.\ 271–284. Association for Computing Machinery, 2023. ISBN 9798400701924. URL https://doi.org/10.1145/3593013.3593996

  7. [7]

    Blodgett, S. L. Sociolinguistically Driven Approaches for Just Natural Language Processing. PhD thesis, University of Massachusetts Amherst, 2021. URL https://doi.org/10.7275/20410631

  8. [8]

    L., Barocas, S., Daum \'e III, H., and Wallach, H

    Blodgett, S. L., Barocas, S., Daum \'e III, H., and Wallach, H. Language (Technology) is Power: A Critical Survey of Bias in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp.\ 5454--5476. Association for Computational Linguistics, July 2020. doi:10.18653/v1/2020.acl-main.485. URL https://aclanthology.org...

Show all 52 references
  1. [9]

    L., Lopez, G., Olteanu, A., Sim, R., and Wallach, H

    Blodgett, S. L., Lopez, G., Olteanu, A., Sim, R., and Wallach, H. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Join...

  2. [10]

    Quantifying Memorization Across Neural Language Models

    Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., and Zhang, C. Quantifying Memorization Across Neural Language Models . In International Conference on Learning Representations, 2023

  3. [11]

    F., Corvi, E., Dow, P., Garcia-Gathright, J., Pangakis, N., Reed, S., Sheng, E., Vann, D., Vogel, M., Washington, H., and Wallach, H

    Chouldechova, A., Atalla, C., Barocas, S., Cooper, A. F., Corvi, E., Dow, P., Garcia-Gathright, J., Pangakis, N., Reed, S., Sheng, E., Vann, D., Vogel, M., Washington, H., and Wallach, H. A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, an...

  4. [12]

    Training verifiers to solve math word problems

    Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems . arXiv preprint arXiv:2110.14168, 2021

  5. [13]

    Cooper, A. F. and Grimmelmann, J. The Files are in the Computer: Copyright, Memorization, and Generative AI . arXiv preprint arXiv:2404.12590, 2024

  6. [14]

    F., Abrams, E., and NA, N

    Cooper, A. F., Abrams, E., and NA, N. Emergent Unfairness in Algorithmic Fairness-Accuracy Trade-Off Research . In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIES '21, pp.\ 46–54. Association for Computing Machinery, 2021. ISBN 9781450384735. doi:1...

  7. [15]

    F., Lee, K., Grimmelmann, J., Ippolito, D., Callison-Burch, C., Choquette-Choo, C

    Cooper, A. F., Lee, K., Grimmelmann, J., Ippolito, D., Callison-Burch, C., Choquette-Choo, C. A., Mireshghallah, N., Brundage, M., Mimno, D., Choksi, M. Z., Balkin, J. M., Carlini, N., Sa, C. D., Frankle, J., Ganguli, D., Gipson, B., Guadamuz, A., Harris, S. L., Jacobs, A. Z.,...

  8. [16]

    F., Choquette-Choo, C

    Cooper, A. F., Choquette-Choo, C. A., Bogen, M., Jagielski, M., Filippova, K., Liu, K. Z., Chouldechova, A., Hayes, J., Huang, Y., Mireshghallah, N., Shumailov, I., Triantafillou, E., Kairouz, P., Mitchell, N., Liang, P., Ho, D. E., Choi, Y., Koyejo, S., Delgado, F., Grimmelma...

  9. [17]

    Representational Harms through the Lens of Speech Act Theory , 2024

    Corvi, E., Washington, H., Reed, S., Atalla, C., Chouldechova, A., Dow, A., Garcia-Gathright, J., Pangakis, N., Sheng, E., Vann, D., Vogel, M., and Wallach, H. Representational Harms through the Lens of Speech Act Theory , 2024

  10. [18]

    Cronbach, L. J. and Meehl, P. E. Construct validity in psychological tests . Psychological bulletin, 52 0 (4): 0 281, 1955. doi:10.1037/h0040957

  11. [19]

    English, L. D. Mathematical reasoning: Analogies, metaphors, and images. Routledge, 2013

  12. [20]

    F., Denain, J.-S., Ho, A., Santos, E

    Glazer, E., Erdil, E., Besiroglu, T., Chicharro, D., Chen, E., Gunning, A., Olsson, C. F., Denain, J.-S., Ho, A., Santos, E. d. O., et al. FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI . arXiv preprint arXiv:2411.04872, 2024

  13. [21]

    Measuring Mathematical Problem Solving With the MATH Dataset

    Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. Measuring Mathematical Problem Solving With the MATH Dataset . In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021. UR...

  14. [22]

    Evaluation gaps in machine learning practice

    Hutchinson, B., Rostamzadeh, N., Greer, C., Heller, K., and Prabhakaran, V. Evaluation gaps in machine learning practice . In Proceedings of the 2022 ACM conference on Fairness, Accountability, and Transparency, pp.\ 1859--1876, 2022

  15. [23]

    Jacobs, A. Z. and Wallach, H. Measurement and fairness . In Proceedings of the 2021 ACM conference on Fairness, Accountability, and Transparency, pp.\ 375--385, 2021

  16. [24]

    and Kieran, C

    Jeannotte, D. and Kieran, C. A conceptual model of mathematical reasoning for school mathematics . Educational Studies in mathematics, 96: 0 1--16, 2017

  17. [25]

    Kapania, S., Ballard, S., Kessler, A., and Vaughan, J. W. Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline . Working paper, 2025

  18. [26]

    Algorithm = Logic + Control

    Kowalski, R. Algorithm = Logic + Control . Communications of the ACM, 22 0 (7), 1979

  19. [27]

    F., and Grimmelmann, J

    Lee, K., Cooper, A. F., and Grimmelmann, J. Talkin' 'Bout AI Generation: Copyright and the Generative- AI Supply Chain . arXiv preprint arXiv:2309.08133, 2023

  20. [28]

    IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering

    Li, R., Li, R., Wang, B., and Du, X. IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering . In Proceedings of the 38th Conference on Neural Information Processing Systems, 2024

  21. [29]

    L., Blodgett, S

    Liu, Y. L., Blodgett, S. L., Cheung, J., Liao, Q. V., Olteanu, A., and Xiao, Z. ECBD: Evidence-Centered Benchmark Design for NLP . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 16349--16365, Bangkok, Th...

  22. [30]

    C., Shoham, Y., Wald, R., and Clark, J

    Maslej, N., Fattorini, L., Perrault, R., Parli, V., Reuel, A., Brynjolfsson, E., Etchemendy, J., Ligett, K., Lyons, T., Manyika, J., Niebles, J. C., Shoham, Y., Wald, R., and Clark, J. The AI Index 2024 Annual Report . AI Index Steering Committee, Institute for Human-Centered ...

  23. [31]

    Validity

    Messick, S. Validity. ETS Research Report Series, 1987. doi:10.1002/j.2330-8516.1987.tb00244.x

  24. [32]

    Validity and washback in language testing

    Messick, S. Validity and washback in language testing . Language Testing, 13 0 (3): 0 241--256, 1996

  25. [33]

    K., Koopman, C., and Doty, N

    Mulligan, D. K., Koopman, C., and Doty, N. Privacy is an essentially contested concept: A multi-dimensional analytic for mapping privacy . Phil. Trans. R. Soc. A, 374 0 (2083): 0 20160118, 2016

  26. [34]

    K., Kroll, J

    Mulligan, D. K., Kroll, J. A., Kohli, N., and Wong, R. Y. This Thing Called Fairness: Disciplinary Confusion Realizing a Value in Technology . Proc. ACM Hum.-Comput. Interact., 3 0 (CSCW), Nov 2019

  27. [35]

    S tereo S et: Measuring stereotypical bias in pretrained language models

    Nadeem, M., Bethke, A., and Reddy, S. S tereo S et: Measuring stereotypical bias in pretrained language models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processin...

  28. [36]

    Nangia, N., Vania, C., Bhalerao, R., and Bowman, S. R. C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp.\ 1953--1967. Association fo...

  29. [37]

    F., Ippolito, D., Choquette-Choo, C

    Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramèr, F., and Lee, K. Scalable extraction of training data from (production) language models . arXiv preprint arXiv:2311.17035, 2023

  30. [38]

    Artificial intelligence Risk Management Framework: Generative Artificial Intelligence Profile , 2024

    National Institute for Standards and Technology . Artificial intelligence Risk Management Framework: Generative Artificial Intelligence Profile , 2024. URL https://doi.org/10.6028/NIST.AI.600-1. NIST Trustworthy and Responsible AI NIST AI 600-1

  31. [39]

    Learning to Reason with LLM s , 2024

    OpenAI. Learning to Reason with LLM s , 2024. URL https://openai.com/index/learning-to-reason-with-llms/

  32. [40]

    Red Teaming Language Models with Language Models

    Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G. Red Teaming Language Models with Language Models . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp.\ 3419--3448. Association...

  33. [41]

    Porada, I., Olteanu, A., Suleman, K., Trischler, A., and Cheung, J. C. K. Challenges to evaluating the generalization of coreference resolution models: A measurement modeling perspective . In Findings of the Association for Computational Linguistics ACL 2024, pp.\ 15380--15395, 2024

  34. [42]

    S., Deng, A., O'Brien, K., au2, J

    Prashanth, U. S., Deng, A., O'Brien, K., au2, J. S. V., Khan, M. A., Borkar, J., Choquette-Choo, C. A., Fuehne, J. R., Biderman, S., Ke, T., Lee, K., and Saphra, N. Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon , 2024. URL https://arxiv.org/a...

  35. [43]

    D., Bender, E

    Raji, I. D., Bender, E. M., Paullada, A., Denton, E., and Hanna, A. AI and the Everything in the Whole Wide World Benchmark . Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2021

  36. [44]

    A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L

    Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L. Gaps in the Safety Evaluation of Generative AI . AAAI/ACM C...

  37. [45]

    L., Gardner, M., and Singh, S

    Razeghi, Y., Logan IV, R. L., Gardner, M., and Singh, S. Impact of Pretraining Term Frequencies on Few-Shot Numerical Reasoning . In Findings of the Association for Computational Linguistics: EMNLP 2022, pp.\ 840--854. Association for Computational Linguistics, December 2022. ...

  38. [46]

    Roose, K. A . I . has a measurement problem . The New York Times, April 2024. URL https://www.nytimes.com/2024/04/15/technology/ai-models-measurement.html. Accessed: 2024-09-05

  39. [47]

    H., Reed, D

    Saltzer, J. H., Reed, D. P., and Clark, D. D. End-to-End Arguments in System Design . ACM Trans. Comput. Syst., 2 0 (4): 0 277–288, nov 1984. ISSN 0734-2071. doi:10.1145/357401.357402. URL https://doi.org/10.1145/357401.357402

  40. [48]

    Evaluating general-purpose AI with psychometrics

    Wang, X., Jiang, L., Hernandez-Orallo, J., Stillwell, D., Sun, L., Luo, F., and Xie, X. Evaluating general-purpose AI with psychometrics . arXiv preprint arXiv:2310.16379v2, 2023

  41. [49]

    A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., and Isaac, W

    Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., and Isaac, W. Sociotechnical Safety Evaluation of Generative AI Systems . arXiv:2310.11986, 2023

  42. [50]

    Evaluating Mathematical Reasoning Beyond Accuracy

    Xia, S., Li, X., Liu, Y., Wu, T., and Liu, P. Evaluating Mathematical Reasoning Beyond Accuracy . Proceedings of the AAAI Conference on Artificial Intelligence, 2025

  43. [51]

    The nature and origins of mass opinion

    Zaller, J. The nature and origins of mass opinion. Cambridge University, 1992

  44. [52]

    Y., Choi, Y., Mireshghallah, N., Le Bras, R., and Sap, M

    Zhou, X., Kim, H., Brahman, F., Jiang, L., Zhu, H., Lu, X., Xu, F., Lin, B. Y., Choi, Y., Mireshghallah, N., Le Bras, R., and Sap, M. HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human- AI Interactions . arXiv, 2024. URL http://arxiv.org/abs/2409.16427

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.