REVIEW 2 major objections 3 minor 14 cited by
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
T0 review · 2 major / 3 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read The paper argues that evaluating generative AI systems is essentially a social science measurement challenge, and that adopting a four-level measurement framework would give the field a way to say precisely what a benchmark or safety test…
desk verdict A serious, useful synthesis of social-science measurement theory for GenAI evaluation, slightly overbroad in its claim to cover all measurement tasks, and best where the concepts are socially embedded. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the four-level measurement framework adapted from social science measurement theory, together with the insistence that conceptual debates about what is being measured and why be kept separate from operational debates about how to measure it. The framework's core move is systematization: turning a broad, contested background concept into an explicit systematized concept that specifies observable phenomena and their relationships before any instrument is built. Operationalization then maps that definition to indicators, annotation guidelines, aggregation functions, and other measurement instruments; interrogation uses seven validity lenses to test the systematized concept, the instruments, and the measurements, and may send the process back for revision. The paper's running example of measuring text that stereotypes social groups in a chatbot's outputs shows how the framework distinguishes a judge model's annotation from the definition being measured, and why human annotations cannot be treated as automatic ground truth without first checking that they align with the systematized concept.
What would settle it
A concrete test would be to apply the framework to several existing benchmark tasks, fully systematizing each target concept, then compare the resulting measurements with those from the original instruments on the same systems; if the framework-guided instruments yield neither stronger convergent and discriminant evidence nor clearer predictions of external outcomes than the instruments they replace, the claim that this process brings rigor to GenAI evaluation loses its empirical footing. A quicker observational falsifier would be to find two well-specified real-world evaluation contexts where the same systematized concept and instruments are used but validity evidence cannot be re-established, contradicting the claim that validity is always context-dependent.
Extended reading notes
Core claim
The central claim is that the dominant way the ML community evaluates GenAI systems is measurement without a measurement theory: concepts of interest are abstract and contested, yet researchers typically move directly to datasets, prompts, and scoring rubrics without an explicit systematized definition, and rarely interrogate whether instruments and measurements are valid. The paper's proposed standard is a four-level framework, grounded in social science measurement theory, with levels for the background concept, the systematized concept, the measurement instruments, and the measurements themselves, linked by systematization, operationalization, application, and interrogation. Following this standard, evaluators would first develop an explicit definition that connects the concept to observable phenomena, then operationalize that definition into indicators and instruments, and finally use face, content, convergent, discriminant, predictive, hypothesis, and consequential validity to challenge every step. The authors directly address the objection that AI systems are not social systems by noting that the concepts at stake are deeply intertwined with people and society, not with the physical substrate of the system.
Load-bearing premise
The argument rests on the assumption that a measurement framework developed for studying human social and political concepts, with its definition of validity and its seven evidence lenses, can meaningfully be transplanted to the evaluation of machines, a transfer the paper asserts through argument and example but does not empirically demonstrate.
Editorial extensions
If this is right
- Benchmark papers would routinely publish their systematized concept alongside datasets and rubrics, making it possible to tell when two benchmarks measure the same thing.
- Claims like 'LLM-as-a-judge accuracy' would be reframed: agreement with humans would no longer be treated as ground truth unless the humans' annotations are shown to align with the systematized concept.
- Validity would become a context-bound property, so instruments would be re-interrogated before each new deployment context rather than certified once.
- Evaluation would open to more participants, since policymakers, users, and affected communities could contest the systematized concept and its consequences without needing to read code.
- The framework would apply across measurement approaches, including benchmarks, red teaming, real-world evaluations, and user studies, not just tests that resemble psychometric instruments.
Reading between the lines
- If this position is adopted, a natural next step is a reporting standard: papers making evaluative claims about GenAI systems would disclose their systematized concept, indicators, aggregation functions, and validity evidence, an extension the paper gestures toward but does not formalize.
- The reinterpretation of human-judge agreement has a sharp consequence: many published 'judge accuracy' numbers are uninterpretable as validity evidence, and re-examining them through this framework could change which models are trusted as judges.
- The same logic applies to fairness metrics and other algorithmic measurements, since the paper notes the framework extends beyond GenAI; a testable extension would be to require an explicit systematized definition of fairness before selecting any metric.
- A concrete prediction follows: benchmark suites built with explicit systematized concepts and validity interrogation will produce measurements that are more stable across minor changes to prompts and sampling than suites built without, because operationalization choices are constrained by the definition.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that the evaluation of generative AI (GenAI) systems should be treated as a social science measurement problem. It adapts Adcock and Collier's (2001) four-level measurement framework—background concept, systematized concept, measurement instruments, and measurements—and recommends seven validity lenses drawn from Messick (1987) and Jacobs and Wallach (2021). The authors claim that separating systematization from operationalization clarifies what is being measured, that the validity lenses impose needed rigor on conceptual and operational debates, and that the framework applies to all measurement tasks involved in evaluating GenAI systems. The argument is developed through a running example (measuring stereotype prevalence in LLM outputs), abbreviated applications to mathematical reasoning, memorization, and refusal, and a discussion of adoption barriers and alternative views.
Significance. If the position is accepted, the paper could serve as a shared vocabulary and quality standard for GenAI evaluation, connecting benchmark design, red teaming, and user studies to a established body of social science measurement practice. The manuscript is careful and honest: it explicitly says the framework is not a panacea, that partial adoption is still valuable, and that it does not advocate transferring psychometric instruments to machines. Its main strengths are the detailed worked example, the transparent borrowing of an established framework, and the appendices showing how existing memorization results can be read through the validity lenses. The paper does not offer new formal results, code, or data, but as a position paper its contribution is conceptual unification and a concrete set of recommended actions.
major comments (2)
- [Sec. 2 and Appendix C] The central claim that the framework 'brings clarity to all measurement tasks involved in evaluating GenAI systems' (Sec. 2) is not established for constructs with checkable success criteria. In Appendix C, the mathematical reasoning example proposes as the systematized concept 'the accuracy of the model on highly challenging mathematical reasoning problems aimed at pre-university students, spanning algebra, number theory, combinatorics, and geometry'—this is already an operational quantity rather than a concept definition, so the separation of systematization from operationalization, which the paper presents as its main benefit, collapses by construction. The memorization example (Appendix C and E) is more developed, but it is largely a reinterpretation of existing measurement choices (token length, greedy versus probabilistic sampling, perplexity checks) through the validity lenses rather than a demonstration that the lenses generate evidence beyond standard benchmark validation. I recommend the authors either provide a worked systematization for at least one capability-oriented construct in which the background/systematized distinction and the validity lenses do substantive work, or revise the scope claim to concepts whose meanings are socially contested or whose measurement requires construct interpretation. Without one of these, the universal applicability claim is stronger than the evidence.
- [Sec. 3.2 and Appendix B] The paper's claim that validity lenses can inform conceptual debates (contra Adcock and Collier) is asserted rather than argued. Appendix B says 'we argue that three lenses—face validity, content validity, and convergent validity—are especially useful for shedding light on conceptual debates,' but no reasoning is given for this selection, and it contradicts Section 4.1.2, which lists face, content, and consequential validity as the three lenses. Because this deviation is one of the paper's stated contributions, the authors should either supply the promised argument or mark the selection as provisional.
minor comments (3)
- [Sec. 1, References] There is a missing space in 'Theprocess of evaluation' near the start of Section 1, and the bibliography entry for Hand (2004) misspells 'Measurement' as 'Meaurement'.
- [Sec. 3.1] The phrase 'researchers and practitioners appear to jump from background concepts to measurement instruments' is used as a motivation, but its evidentiary basis is not stated; please either cite a systematic analysis or qualify it as an informal observation about common practice.
- [Appendix C] In the mathematical reasoning example, the systematized concept should be expressed as a construct (e.g., the ability to solve pre-university-level olympiad problems in specified topic areas), with 'accuracy' reserved for the measurement level; as written, the example risks enacting the very conflation the paper criticizes.
Circularity Check
No significant circularity: the framework is imported from external measurement theory (Adcock and Collier; Messick), and the paper's self-citations are transparent, non-load-bearing, and do not reduce the argument to its inputs.
full rationale
The paper's central position is programmatic rather than derivational: it proposes adopting a variant of Adcock and Collier's (2001) four-level measurement framework and Messick's (1987) validity lenses. These sources are external to the authors, and the framework's definitions are not constructed from the paper's conclusions. The seven validity lenses are described as 'adapted from Messick (1987) by Jacobs & Wallach (2021)' (Section 3.2), which is a self-citation; however, the substantive content is attributed to Messick, the lenses are standard independently checkable categories, and the argument does not rely on treating Jacobs and Wallach as an unverified authority. The paper's worked examples (stereotyping text, mathematical reasoning, memorization, refusal) are instantiations of the framework, not predictions derived from fitted parameters; the claim that ML practice conflates systematization and operationalization is argued and supported by external critiques rather than by definition. The transferability of Adcock and Collier's framework from human social concepts to machine-measured constructs is asserted as a position and is a substantive limitation, but no statement in the paper equates the conclusion with an input by construction or makes a fitted value do the work of an independent prediction. Therefore no circular step is present.
Assumptions & free parameters
assumptions (4)
- domain assumption Adcock and Collier's four-level measurement framework is an appropriate and effective basis for standardizing GenAI evaluation.
- domain assumption Messick's unified validity perspective, including consequential validity, is the appropriate lens for interrogating validity of AI measurements.
- domain assumption ML researchers and practitioners routinely conflate concept definition with measurement and rarely interrogate validity.
- domain assumption Separating systematization from operationalization will yield the claimed benefits of broader stakeholder participation and improved rigor.
Cite this review
Pith. "Pith review of Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge." pith.science (2026). https://pith.science/paper/JQDGEE2V
@misc{pith2026250200561,
author = {Pith},
title = {Pith review of: Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge},
year = {2026},
howpublished = {\url{https://pith.science/paper/JQDGEE2V}},
note = {Machine review of arXiv:2502.00561}
}
read the original abstract
The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from learning from and drawing on the social sciences when developing and using measurement instruments for evaluating GenAI systems. Specifically, our position is that evaluating GenAI systems is a social science measurement challenge. We present a four-level framework, grounded in measurement theory from the social sciences, for measuring concepts related to the capabilities, behaviors, and impacts of GenAI systems. This framework has two important implications: First, it can broaden the expertise involved in evaluating GenAI systems by enabling stakeholders with different perspectives to participate in conceptual debates. Second, it brings rigor to both conceptual and operational debates by offering a set of lenses for interrogating validity.
Figures
Forward citations
Cited by 14 Pith papers
-
Algorithmic Gender Prediction Is Illegitimate, But Gender Imputation Can Yield Valid Measurements
Gender prediction is illegitimate but gender imputation can still yield valid disparity measurements for traditional sexism, though not for oppositional sexism.
-
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
The paper introduces DiffAware and CtxtAware metrics, an 8-benchmark suite with 16,000 questions, and shows ten LLMs are less able to recognize legitimate group differences than current fairness benchmarks suggest.
-
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)
Chat UI and API access to the same chatbot produce different accuracy, consistency, citation, and refusal behaviors on safety benchmarks, and web search changes these patterns further.
-
Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation
Base language models outperform post-trained models at generating individual survey responses that match human distributions, while post-trained models are more accurate when directly predicting population opinion dis...
-
On the Convergent Validity of Offline Evaluation Designs for Recommender Systems
Sparse offline recommender rankings correlate only weakly—and sometimes negatively—with dense ground-truth rankings, and no evaluation design is uniformly best.
-
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.
-
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
Coding benchmarks misalign with agentic software engineering because they conflate model and harness, grade against single references, and provide no component-level iteration signals.
-
Grounded Chess Reasoning in Language Models via Master Distillation
Master Distillation turns opaque chess-engine search into natural-language CoT, lifting a 4B model (C1) from near-zero to 48.1% accuracy that beats its teacher and most larger LMs.
-
Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances
A Fair Equality of Chances-based framework decomposes GenAI unfairness into harms/benefits, morally arbitrary factors, and morally decisive factors to improve measurement validity.
-
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
Practitioners trying to measure representational harms in LLM-based systems often cannot use public measurement instruments, either because the instruments lack validity, specificity, interpretability, or actionabilit...
-
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
The paper sets out a research agenda for NLP to measure how prolonged language-model use changes human behavior over long time horizons, replacing single-session safety evaluations with longitudinal tracking.
-
Against 'softmaxing' culture
A position paper arguing that AI evaluations should shift from defining culture to understanding when culture becomes relationally valid.
-
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
A position paper argues that understanding AI's second-order effects requires moving from static benchmarks to an ecosystem of field testing, red teaming, and contextual evaluation.
-
AI, Jobs, and the Automation Trap: Where Is HCI?
A position paper claiming AI patents overwhelmingly favor task automation over human augmentation, and proposing incentive changes to make human-centered AI more impactful.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Abebe, R., Barocas, S., Kleinberg, J., Levy, K., Raghavan, M., and Robinson, D. G. Roles for computing in social change . In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pp.\ 252--260, 2020
work page 2020
-
[3]
Adcock, R. and Collier, D. Measurement validity: A shared standard for qualitative and quantitative research . American Political Science Review, 95 0 (3): 0 529--546, 2001
work page 2001
-
[4]
Amdahl, G. M., Blaauw, G. A., and Brooks, F. P. Architecture of the IBM S ystem/360. IBM Journal of Research and Development, 8 0 (2), 1964
work page 1964
-
[5]
Content Analysis in Communication Research
Berelson, B. Content Analysis in Communication Research. Free Press, 1952
work page 1952
-
[6]
Blili-Hamelin, B. and Hancox-Li, L. Making Intelligence: Ethical Values in IQ and ML Benchmarks . In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, FAccT '23, pp.\ 271–284. Association for Computing Machinery, 2023. ISBN 9798400701924. URL https://doi.org/10.1145/3593013.3593996
-
[7]
Blodgett, S. L. Sociolinguistically Driven Approaches for Just Natural Language Processing. PhD thesis, University of Massachusetts Amherst, 2021. URL https://doi.org/10.7275/20410631
doi:10.7275/20410631 2021
-
[8]
L., Barocas, S., Daum \'e III, H., and Wallach, H
Blodgett, S. L., Barocas, S., Daum \'e III, H., and Wallach, H. Language (Technology) is Power: A Critical Survey of Bias in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp.\ 5454--5476. Association for Computational Linguistics, July 2020. doi:10.18653/v1/2020.acl-main.485. URL https://aclanthology.org...
Show all 52 references
-
[9]
L., Lopez, G., Olteanu, A., Sim, R., and Wallach, H
Blodgett, S. L., Lopez, G., Olteanu, A., Sim, R., and Wallach, H. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Join...
2021
-
[10]
Quantifying Memorization Across Neural Language Models
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., and Zhang, C. Quantifying Memorization Across Neural Language Models . In International Conference on Learning Representations, 2023
2023
-
[11]
F., Corvi, E., Dow, P., Garcia-Gathright, J., Pangakis, N., Reed, S., Sheng, E., Vann, D., Vogel, M., Washington, H., and Wallach, H
Chouldechova, A., Atalla, C., Barocas, S., Cooper, A. F., Corvi, E., Dow, P., Garcia-Gathright, J., Pangakis, N., Reed, S., Sheng, E., Vann, D., Vogel, M., Washington, H., and Wallach, H. A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, an...
2024 arXiv
-
[12]
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems . arXiv preprint arXiv:2110.14168, 2021
2021 arXiv
-
[13]
Cooper, A. F. and Grimmelmann, J. The Files are in the Computer: Copyright, Memorization, and Generative AI . arXiv preprint arXiv:2404.12590, 2024
2024 arXiv
-
[14]
F., Abrams, E., and NA, N
Cooper, A. F., Abrams, E., and NA, N. Emergent Unfairness in Algorithmic Fairness-Accuracy Trade-Off Research . In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIES '21, pp.\ 46–54. Association for Computing Machinery, 2021. ISBN 9781450384735. doi:1...
2021
-
[15]
F., Lee, K., Grimmelmann, J., Ippolito, D., Callison-Burch, C., Choquette-Choo, C
Cooper, A. F., Lee, K., Grimmelmann, J., Ippolito, D., Callison-Burch, C., Choquette-Choo, C. A., Mireshghallah, N., Brundage, M., Mimno, D., Choksi, M. Z., Balkin, J. M., Carlini, N., Sa, C. D., Frankle, J., Ganguli, D., Gipson, B., Guadamuz, A., Harris, S. L., Jacobs, A. Z.,...
2023 arXiv
-
[16]
F., Choquette-Choo, C
Cooper, A. F., Choquette-Choo, C. A., Bogen, M., Jagielski, M., Filippova, K., Liu, K. Z., Chouldechova, A., Hayes, J., Huang, Y., Mireshghallah, N., Shumailov, I., Triantafillou, E., Kairouz, P., Mitchell, N., Liang, P., Ho, D. E., Choi, Y., Koyejo, S., Delgado, F., Grimmelma...
2024
-
[17]
Representational Harms through the Lens of Speech Act Theory , 2024
Corvi, E., Washington, H., Reed, S., Atalla, C., Chouldechova, A., Dow, A., Garcia-Gathright, J., Pangakis, N., Sheng, E., Vann, D., Vogel, M., and Wallach, H. Representational Harms through the Lens of Speech Act Theory , 2024
2024
-
[18]
Cronbach, L. J. and Meehl, P. E. Construct validity in psychological tests . Psychological bulletin, 52 0 (4): 0 281, 1955. doi:10.1037/h0040957
1955 doi
-
[19]
English, L. D. Mathematical reasoning: Analogies, metaphors, and images. Routledge, 2013
2013
-
[20]
F., Denain, J.-S., Ho, A., Santos, E
Glazer, E., Erdil, E., Besiroglu, T., Chicharro, D., Chen, E., Gunning, A., Olsson, C. F., Denain, J.-S., Ho, A., Santos, E. d. O., et al. FrontierMath: A benchmark for evaluating advanced mathematical reasoning in AI . arXiv preprint arXiv:2411.04872, 2024
2024 arXiv
-
[21]
Measuring Mathematical Problem Solving With the MATH Dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. Measuring Mathematical Problem Solving With the MATH Dataset . In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021. UR...
2021
-
[22]
Evaluation gaps in machine learning practice
Hutchinson, B., Rostamzadeh, N., Greer, C., Heller, K., and Prabhakaran, V. Evaluation gaps in machine learning practice . In Proceedings of the 2022 ACM conference on Fairness, Accountability, and Transparency, pp.\ 1859--1876, 2022
2022
-
[23]
Jacobs, A. Z. and Wallach, H. Measurement and fairness . In Proceedings of the 2021 ACM conference on Fairness, Accountability, and Transparency, pp.\ 375--385, 2021
2021
-
[24]
and Kieran, C
Jeannotte, D. and Kieran, C. A conceptual model of mathematical reasoning for school mathematics . Educational Studies in mathematics, 96: 0 1--16, 2017
2017
-
[25]
Kapania, S., Ballard, S., Kessler, A., and Vaughan, J. W. Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline . Working paper, 2025
2025
-
[26]
Algorithm = Logic + Control
Kowalski, R. Algorithm = Logic + Control . Communications of the ACM, 22 0 (7), 1979
1979
-
[27]
F., and Grimmelmann, J
Lee, K., Cooper, A. F., and Grimmelmann, J. Talkin' 'Bout AI Generation: Copyright and the Generative- AI Supply Chain . arXiv preprint arXiv:2309.08133, 2023
2023 arXiv
-
[28]
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering
Li, R., Li, R., Wang, B., and Du, X. IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering . In Proceedings of the 38th Conference on Neural Information Processing Systems, 2024
2024
-
[29]
L., Blodgett, S
Liu, Y. L., Blodgett, S. L., Cheung, J., Liao, Q. V., Olteanu, A., and Xiao, Z. ECBD: Evidence-Centered Benchmark Design for NLP . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp.\ 16349--16365, Bangkok, Th...
2024
-
[30]
C., Shoham, Y., Wald, R., and Clark, J
Maslej, N., Fattorini, L., Perrault, R., Parli, V., Reuel, A., Brynjolfsson, E., Etchemendy, J., Ligett, K., Lyons, T., Manyika, J., Niebles, J. C., Shoham, Y., Wald, R., and Clark, J. The AI Index 2024 Annual Report . AI Index Steering Committee, Institute for Human-Centered ...
2024
-
[31]
Validity
Messick, S. Validity. ETS Research Report Series, 1987. doi:10.1002/j.2330-8516.1987.tb00244.x
1987
-
[32]
Validity and washback in language testing
Messick, S. Validity and washback in language testing . Language Testing, 13 0 (3): 0 241--256, 1996
1996
-
[33]
K., Koopman, C., and Doty, N
Mulligan, D. K., Koopman, C., and Doty, N. Privacy is an essentially contested concept: A multi-dimensional analytic for mapping privacy . Phil. Trans. R. Soc. A, 374 0 (2083): 0 20160118, 2016
2016
-
[34]
K., Kroll, J
Mulligan, D. K., Kroll, J. A., Kohli, N., and Wong, R. Y. This Thing Called Fairness: Disciplinary Confusion Realizing a Value in Technology . Proc. ACM Hum.-Comput. Interact., 3 0 (CSCW), Nov 2019
2019
-
[35]
S tereo S et: Measuring stereotypical bias in pretrained language models
Nadeem, M., Bethke, A., and Reddy, S. S tereo S et: Measuring stereotypical bias in pretrained language models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processin...
2021 doi
-
[36]
Nangia, N., Vania, C., Bhalerao, R., and Bowman, S. R. C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp.\ 1953--1967. Association fo...
2020 doi
-
[37]
F., Ippolito, D., Choquette-Choo, C
Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramèr, F., and Lee, K. Scalable extraction of training data from (production) language models . arXiv preprint arXiv:2311.17035, 2023
2023 arXiv
-
[38]
Artificial intelligence Risk Management Framework: Generative Artificial Intelligence Profile , 2024
National Institute for Standards and Technology . Artificial intelligence Risk Management Framework: Generative Artificial Intelligence Profile , 2024. URL https://doi.org/10.6028/NIST.AI.600-1. NIST Trustworthy and Responsible AI NIST AI 600-1
2024 doi
-
[39]
Learning to Reason with LLM s , 2024
OpenAI. Learning to Reason with LLM s , 2024. URL https://openai.com/index/learning-to-reason-with-llms/
2024
-
[40]
Red Teaming Language Models with Language Models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G. Red Teaming Language Models with Language Models . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp.\ 3419--3448. Association...
2022 doi
-
[41]
Porada, I., Olteanu, A., Suleman, K., Trischler, A., and Cheung, J. C. K. Challenges to evaluating the generalization of coreference resolution models: A measurement modeling perspective . In Findings of the Association for Computational Linguistics ACL 2024, pp.\ 15380--15395, 2024
2024
-
[42]
S., Deng, A., O'Brien, K., au2, J
Prashanth, U. S., Deng, A., O'Brien, K., au2, J. S. V., Khan, M. A., Borkar, J., Choquette-Choo, C. A., Fuehne, J. R., Biderman, S., Ke, T., Lee, K., and Saphra, N. Recite, Reconstruct, Recollect: Memorization in LMs as a Multifaceted Phenomenon , 2024. URL https://arxiv.org/a...
2024 arXiv
-
[43]
D., Bender, E
Raji, I. D., Bender, E. M., Paullada, A., Denton, E., and Hanna, A. AI and the Everything in the Whole Wide World Benchmark . Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2021
2021
-
[44]
A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L
Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L. Gaps in the Safety Evaluation of Generative AI . AAAI/ACM C...
2024
-
[45]
L., Gardner, M., and Singh, S
Razeghi, Y., Logan IV, R. L., Gardner, M., and Singh, S. Impact of Pretraining Term Frequencies on Few-Shot Numerical Reasoning . In Findings of the Association for Computational Linguistics: EMNLP 2022, pp.\ 840--854. Association for Computational Linguistics, December 2022. ...
2022 doi
-
[46]
Roose, K. A . I . has a measurement problem . The New York Times, April 2024. URL https://www.nytimes.com/2024/04/15/technology/ai-models-measurement.html. Accessed: 2024-09-05
2024
-
[47]
H., Reed, D
Saltzer, J. H., Reed, D. P., and Clark, D. D. End-to-End Arguments in System Design . ACM Trans. Comput. Syst., 2 0 (4): 0 277–288, nov 1984. ISSN 0734-2071. doi:10.1145/357401.357402. URL https://doi.org/10.1145/357401.357402
1984
-
[48]
Evaluating general-purpose AI with psychometrics
Wang, X., Jiang, L., Hernandez-Orallo, J., Stillwell, D., Sun, L., Luo, F., and Xie, X. Evaluating general-purpose AI with psychometrics . arXiv preprint arXiv:2310.16379v2, 2023
2023 arXiv
-
[49]
A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., and Isaac, W
Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., and Isaac, W. Sociotechnical Safety Evaluation of Generative AI Systems . arXiv:2310.11986, 2023
-
[50]
Evaluating Mathematical Reasoning Beyond Accuracy
Xia, S., Li, X., Liu, Y., Wu, T., and Liu, P. Evaluating Mathematical Reasoning Beyond Accuracy . Proceedings of the AAAI Conference on Artificial Intelligence, 2025
2025
-
[51]
The nature and origins of mass opinion
Zaller, J. The nature and origins of mass opinion. Cambridge University, 1992
1992
-
[52]
Y., Choi, Y., Mireshghallah, N., Le Bras, R., and Sap, M
Zhou, X., Kim, H., Brahman, F., Jiang, L., Zhu, H., Lu, X., Xu, F., Lin, B. Y., Choi, Y., Mireshghallah, N., Le Bras, R., and Sap, M. HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human- AI Interactions . arXiv, 2024. URL http://arxiv.org/abs/2409.16427
2024 arXiv
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.