REVIEW 2 major objections 2 minor 1 cited by
Design and Report Benchmarks for Knowledge Work
T0 review · 2 major / 2 minor · reviewed 2026-05-25 · grok-4.3
Pith's one-line read Benchmark scores for knowledge-work AI support reliable claims only when tasks define the work activity, the tested setting, and the scored product.
desk verdict This paper gives a three-step method plus an O*NET-derived list of 18 activities to make benchmark claims about knowledge work more explicit, but stays conceptual with no check on whether the rules actually improve the mapping. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The three-step approach of defining the work activity under evaluation, specifying the tested setting, and scoring the appropriate work product.
What would settle it
A benchmark redesigned with explicit activity definition, setting specification, and product scoring that still yields scores unrelated to performance on the corresponding real-world work activity in actual deployment settings.
Extended reading notes
Core claim
The central claim is that benchmarked tasks for knowledge-work AI represent work claims only when the design process first defines the work activity under evaluation, then specifies the tested setting with its local materials, tools, roles and constraints, and finally scores the work product that must remain usable in downstream workflows; this mapping is derived from work studies and demonstrated through an inventory of eighteen activities plus case analyses of three benchmarks.
Load-bearing premise
Insights from studies of human knowledge work on roles, responsibilities, materials, tools, and downstream-usable artifacts can be translated directly into benchmark design rules without losing critical aspects or creating new mismatches.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that current benchmarks for knowledge-work AI (e.g., coding, research) follow traditional NLP task logic and thus do not reliably support real-world work claims. It contributes a three-step approach—defining the work activity under evaluation (via an inventory of 18 activities derived from the public O*NET database), specifying the tested setting (materials, tools, roles, constraints), and scoring the appropriate work product—after reviewing work studies on roles/responsibilities, local materials/tools, and downstream-usable artifacts, translating those into design/reporting guidance, and illustrating via post-hoc case analyses of GDPval, OfficeQA Pro, and APEX-SWE.
Significance. If the framework holds, it would provide a structured, explicit method to align benchmark scores with deployment claims for LLM agents in knowledge work, filling a recognized gap. The derivation of the 18-activity inventory from the public O*NET database is a reproducible, non-circular strength that supports the proposal's grounding in external work studies rather than self-referential fitting.
major comments (2)
- [paragraph on reviewing work studies and translating concerns into guidance] The translation from reviewed work-study concerns (roles, local materials/tools, downstream artifacts) into benchmark design rules is the load-bearing step for the central claim that the three-step approach makes work claims explicit. The manuscript acknowledges situated contextual factors in the work-studies review but provides no systematic mapping, coverage check, or independent validation that the resulting rules avoid omitting those factors or introducing new mismatches.
- [demonstration through three benchmark case analyses] The three case analyses (GDPval, OfficeQA Pro, APEX-SWE) demonstrate how design choices shape the strongest supported work claim and identify gaps, but they are retrospective applications to existing benchmarks. This does not test whether prospectively applying the three-step approach during benchmark design would produce measurably better fidelity between scores and work claims.
minor comments (2)
- The derivation of the 18 work activities from O*NET is referenced but not accompanied by an explicit aggregation procedure or example mappings from specific occupational tasks; an appendix with this detail would aid reproducibility.
- Terms such as 'work product' and 'tested setting' are used throughout without early formal definitions, which may reduce clarity for readers outside work-studies literature.
Simulated Author's Rebuttal
We thank the referee for the constructive review and the recognition of the framework's potential value. We respond point-by-point to the major comments below.
read point-by-point responses
-
Referee: The translation from reviewed work-study concerns (roles, local materials/tools, downstream artifacts) into benchmark design rules is the load-bearing step for the central claim that the three-step approach makes work claims explicit. The manuscript acknowledges situated contextual factors in the work-studies review but provides no systematic mapping, coverage check, or independent validation that the resulting rules avoid omitting those factors or introducing new mismatches.
Authors: We agree that the translation step would be strengthened by an explicit mapping. The current manuscript synthesizes the reviewed work-study literature directly into the three design rules and the 18-activity inventory (derived from O*NET), but does not include a tabular or systematic coverage check against the original concerns. We will add a dedicated subsection that maps each reviewed factor (roles/responsibilities, local materials/tools, downstream-usable artifacts) to the corresponding guidance on activity definition, setting specification, and product scoring, and will note any potential gaps or unaddressed contextual elements. revision: yes
-
Referee: The three case analyses (GDPval, OfficeQA Pro, APEX-SWE) demonstrate how design choices shape the strongest supported work claim and identify gaps, but they are retrospective applications to existing benchmarks. This does not test whether prospectively applying the three-step approach during benchmark design would produce measurably better fidelity between scores and work claims.
Authors: The case analyses are retrospective by design: the paper's primary contribution is a methodological proposal illustrated on three published benchmarks to show how the framework surfaces gaps between tasks, settings, products, and work claims. A prospective test—designing a new benchmark from scratch with the method and then measuring improved fidelity—would require a separate empirical study and is outside the scope of the current manuscript. We will revise the text to state this limitation explicitly and to frame the cases as illustrative rather than as a validation of prospective efficacy. revision: partial
- A prospective empirical test of the framework on newly designed benchmarks cannot be performed within the bounds of this methodological paper.
Circularity Check
No circularity; derivation draws on external work studies and O*NET
full rationale
The paper proposes a three-step approach (define work activity, specify tested setting, score work product) by reviewing external work studies on roles/responsibilities/materials/tools/artifacts and translating those into benchmark guidance. It derives an inventory of 18 activities from the public O*NET database. No self-citations, fitted parameters, or self-definitional reductions are present in the derivation chain. The case analyses apply the method to existing benchmarks without reducing claims to inputs by construction. The central claim remains independent of its own outputs.
Assumptions & free parameters
assumptions (1)
- domain assumption Knowledge work is organized through roles and responsibilities, local materials and tools, and artifacts that must remain usable in downstream workflows.
Cite this review
Pith. "Pith review of Design and Report Benchmarks for Knowledge Work." pith.science (2026). https://pith.science/paper/V3CLH7PV
@misc{pith2026260523262,
author = {Pith},
title = {Pith review of: Design and Report Benchmarks for Knowledge Work},
year = {2026},
howpublished = {\url{https://pith.science/paper/V3CLH7PV}},
note = {Machine review of arXiv:2605.23262}
}
read the original abstract
The development of LLM agents has led to a growing body of work on knowledge-work AI, including coding, research, and healthcare. However, current knowledge-work evaluation and benchmark design still largely follow the logic of traditional NLP tasks. As a result, higher benchmark performance does not reliably show that a system can carry out knowledge work in real-world deployment settings. This paper contributes a three-step approach for making explicit how benchmarked tasks represent the work claims attached to their scores: defining the work activity under evaluation, specifying the tested setting, and scoring the appropriate work product. We review work studies showing that knowledge work is organized through roles and responsibilities, local materials and tools, and artifacts that must remain usable in downstream workflows. We then translate these concerns into benchmark design and reporting guidance, covering how tasks should be mapped to work activities, how tested settings should specify materials, tools, roles, and constraints, and how scoring should focus on the work product left by the system. To name the work activity being evaluated and distinguish it from common benchmark tasks, we derive an inventory of 18 work activities from the O{*}NET occupational task database. We demonstrate the approach through three benchmark case analyses: GDPval, a non-code occupational deliverable benchmark; OfficeQA Pro, a grounded document-analysis benchmark scored by final answers; and APEX-SWE, a software-engineering benchmark with executable scored products. These cases show how benchmark design choices shape the strongest work claim a score can support, and where gaps arise between the benchmarked task, tested setting, scored product, and broader work claim.
Figures
Forward citations
Cited by 1 Pith paper
-
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents
Fetch-then-Explore, which stores fetched pages in a per-question workspace and extracts evidence on demand with grep/read, beats visit-and-read and browsing baselines on BrowseComp across three LLM backbones.
Reference graph
Works this paper leans on
-
[1]
Handa, Kunal and Tamkin, Alex and McCain, Miles and Huang, Saffron and Durmus, Esin and Heck, Sarah and Mueller, Jared and Hong, Jerry and Ritchie, Stuart and Belonax, Tim and Troy, Kevin K. and Amodei, Dario and Kaplan, Jared and Clark, Jack and Ganguli, Deep , year =. Which Economic Tasks are Performed with. 2503.04761 , archivePrefix =
-
[2]
American Psychologist , volume =
Messick, Samuel , title =. American Psychologist , volume =. 1995 , doi =
work page 1995
- [3]
- [4]
-
[5]
The System of Professions: An Essay on the Division of Expert Labor , author=. 1988 , publisher=
work page 1988
-
[6]
Standards for Educational and Psychological Testing , author =. 2014 , publisher =
work page 2014
-
[7]
Anthropic Economic Index Report: Uneven Geographic and Enterprise. 2025 , month = sep, url =
work page 2025
-
[8]
and Wei, Jason and Soskin Hicks, Rebecca and Bowman, Preston and Qui
Arora, Rahul K. and Wei, Jason and Soskin Hicks, Rebecca and Bowman, Preston and Qui. 2025 , eprint =
work page 2025
Show all 136 references
-
[9]
Administrative Science Quarterly , year =
Technicians in the Workplace: Ethnographic Evidence for Bringing Work into Organization Studies , author =. Administrative Science Quarterly , year =
-
[10]
, title =
Barres, Victor and Dong, Honghua and Ray, Soham and Si, Xujie and Narasimhan, Karthik R. , title =. 2025 , eprint =
2025
-
[11]
2504.12516 , archivePrefix =
Wei, Jason and Sun, Zhiqing and Papay, Spencer and McKinney, Scott and Han, Jeffrey and Fulford, Isa and Chung, Hyung Won and Passos, Alex Tachard and Fedus, William and Glaese, Amelia , year =. 2504.12516 , archivePrefix =
-
[12]
AEA Papers and Proceedings , year =
What Can Machines Learn, and What Does It Mean for Occupations and the Economy? , author =. AEA Papers and Proceedings , year =
-
[13]
, journal =
Brynjolfsson, Erik and Li, Danielle and Raymond, Lindsey R. , journal =. Generative. 2025 , doi =
2025
-
[14]
The Enterprise AI Playbook: Lessons from 51 Successful Enterprise AI Deployments , author=
-
[15]
Advances in Knowledge Discovery and Data Mining , pages =
Density-Based Clustering Based on Hierarchical Density Estimates , author =. Advances in Knowledge Discovery and Data Mining , pages =. 2013 , publisher =
2013
-
[16]
Organization Science , volume=
A Pragmatic View of Knowledge and Boundaries: Boundary Objects in New Product Development , author=. Organization Science , volume=. 2002 , publisher=
2002
-
[17]
Organization Science , volume =
Transferring, Translating, and Transforming: An Integrative Framework for Managing Knowledge Across Boundaries , author =. Organization Science , volume =. 2004 , doi =
2004
- [18]
-
[19]
The Effects of Generative
Cui, Zheyuan (Kevin) and Demirer, Mert and Jaffe, Sonia and Musolff, Leon and Peng, Sida and Salz, Tobias , journal =. The Effects of Generative. 2025 , note =
2025
-
[20]
2005 , publisher=
Thinking for a Living: How to Get Better Performance and Results from Knowledge Workers , author=. 2005 , publisher=
2005
-
[21]
Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality , author=
-
[22]
and Del Verme, Manuel and Marty, Tom and Vazquez, David and Chapados, Nicolas and Lacoste, Alexandre , booktitle =
Drouin, Alexandre and Gasse, Maxime and Caccia, Massimo and Laradji, Issam H. and Del Verme, Manuel and Marty, Tom and Vazquez, David and Chapados, Nicolas and Lacoste, Alexandre , booktitle =. 2024 , publisher =
2024
-
[23]
, title =
Drucker, Peter F. , title =. 1959 , publisher =
1959
-
[24]
, title =
Drucker, Peter F. , title =. California Management Review , volume =
-
[25]
doi:10.48550/arXiv.2303.10130 , url =
Eloundou, Tyna and Manning, Sam and Mishkin, Pamela and Rock, Daniel , year =. doi:10.48550/arXiv.2303.10130 , url =. 2303.10130 , archivePrefix=
-
[26]
2022 , howpublished =
2022
-
[27]
Strategic Management Journal , year =
Occupational, Industry, and Geographic Exposure to Artificial Intelligence , author =. Strategic Management Journal , year =
-
[28]
2001 , publisher=
Professionalism: The Third Logic , author=. 2001 , publisher=
2001
-
[29]
International Conference on Learning Representations (ICLR) , year =
Mialon, Gr. International Conference on Learning Representations (ICLR) , year =
-
[30]
and Chao, Patrick and Miserendino, Samuel and Chabot, Gildas and Li, David and Sharman, Michael and Barr, Alexandra and Glaese, Amelia and Tworek, Jerry , year =
Patwardhan, Tejal and Dias, Rachel and Proehl, Elizabeth and Kim, Grace and Wang, Michele and Watkins, Olivia and Fishman, Simon Posada and Aljubeh, Marwan and Thacker, Phoebe and Fauconnet, Laurance and Kim, Natalie S. and Chao, Patrick and Miserendino, Samuel and Chabot, Gil...
-
[31]
and Borsos, Zal
Agostinelli, Andrea and Denk, Timo I. and Borsos, Zal. 2023 , eprint=
2023
- [32]
-
[33]
How Well Can
Bianchi, Federico and Chia, Patrick John and Yuksekgonul, Mert and Tagliabue, Jacopo and Jurafsky, Dan and Zou, James , year=. How Well Can. 2402.05863 , archivePrefix=
-
[34]
and Li, Tianle and Li, Dacheng and Zhu, Hao and Zhang, Banghua and Jordan, Michael I
Chiang, Wei-Lin and Zheng, Lianmin and Sheng, Ying and Angelopoulos, Anastasios N. and Li, Tianle and Li, Dacheng and Zhu, Hao and Zhang, Banghua and Jordan, Michael I. and Gonzalez, Joseph E. and Stoica, Ion , year=. Chatbot Arena: An Open Platform for Evaluating. 2403.04132 ...
-
[35]
2307.13528 , archivePrefix=
Chern, I-Chun and Chern, Steffi and Chen, Shiqi and Yuan, Weizhe and Feng, Kehua and Zhou, Chunting and He, Junxian and Neubig, Graham and Liu, Pengfei , year=. 2307.13528 , archivePrefix=
-
[36]
2306.04757 , archivePrefix=
Chia, Yew Ken and Hong, Pengfei and Bing, Lidong and Poria, Soujanya , year=. 2306.04757 , archivePrefix=
-
[37]
2021 , publisher=
Chen, Zhiyu and Chen, Wenhu and Smiley, Charese and Shah, Sameena and Borova, Iana and Langdon, Dylan and Moussa, Reema and Beane, Matt and Huang, Ting-Hao and Routledge, Bryan and Wang, William Yang , booktitle=. 2021 , publisher=
2021
-
[38]
Think You Have Solved Question Answering? Try
Clark, Peter and Cowhey, Isaac and Etzioni, Oren and Khot, Tushar and Sabharwal, Ashish and Schoenick, Carissa and Tafjord, Oyvind , booktitle=. Think You Have Solved Question Answering? Try
-
[39]
2019 , publisher=
Clark, Christopher and Lee, Kenton and Chang, Ming-Wei and Kwiatkowski, Tom and Collins, Michael and Toutanova, Kristina , booktitle=. 2019 , publisher=
2019
-
[40]
2021 , eprint=
Training Verifiers to Solve Math Word Problems , author=. 2021 , eprint=
2021
-
[41]
2306.06070 , archivePrefix=
Deng, Xiang and Gu, Yu and Zheng, Boyuan and Chen, Shijie and Stevens, Sam and Wang, Boshi and Sun, Huan and Su, Yu , year=. 2306.06070 , archivePrefix=
-
[42]
2308.01861 , archivePrefix=
Du, Xueying and Liu, Mingwei and Wang, Kaixin and Wang, Hanlin and Liu, Junwei and Chen, Yixuan and Feng, Jiayi and Sha, Chaofeng and Peng, Xin and Lou, Yiling , year=. 2308.01861 , archivePrefix=
-
[43]
and Li, Irene and She, Tianwei and Li, Suyi and Radev, Dragomir R
Fabbri, Alexander R. and Li, Irene and She, Tianwei and Li, Suyi and Radev, Dragomir R. , journal=
-
[44]
Science , volume=
Human-level Play in the Game of. Science , volume=
-
[45]
Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics , pages=
Hierarchical Neural Story Generation , author=. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics , pages=. 2018 , publisher=
2018
-
[46]
2026 , publisher=
Fein, Daniel and Russo, Sebastian and Xiang, Violet and Jolly, Kabir and Rafailov, Rafael and Haber, Nick , booktitle=. 2026 , publisher=
2026
-
[47]
Measuring Coding Challenge Competence With
Hendrycks, Dan and Basart, Steven and Kadavath, Saurav and Mazeika, Mantas and Arora, Akul and Guo, Ethan and Burns, Collin and Puranik, Samir and He, Horace and Song, Dawn and Steinhardt, Jacob , year=. Measuring Coding Challenge Competence With. 2105.09938 , archivePrefix=
-
[48]
Hendrycks, Dan and Burns, Collin and Chen, Anya and Ball, Spencer , booktitle=
-
[49]
Measuring Mathematical Problem Solving With the
Hendrycks, Dan and Burns, Collin and Kadavath, Saurav and Arora, Akul and Basart, Steven and Tang, Eric and Song, Dawn and Steinhardt, Jacob , booktitle=. Measuring Mathematical Problem Solving With the
-
[50]
International Conference on Learning Representations , year=
Measuring Massive Multitask Language Understanding , author=. International Conference on Learning Representations , year=
-
[51]
and Solar-Lezama, Armando and Sen, Koushik and Stoica, Ion , booktitle=
Jain, Naman and Han, King and Gu, Alex and Li, Wen-Ding and Yan, Fanjia and Zhang, Tianjun and Wang, Sida I. and Solar-Lezama, Armando and Sen, Koushik and Stoica, Ion , booktitle=
-
[52]
and Lu, Xinghua , booktitle=
Jin, Qiao and Dhingra, Bhuwan and Liu, Zhengping and Cohen, William W. and Lu, Xinghua , booktitle=. 2019 , publisher=
2019
-
[53]
2020 , eprint=
What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams , author=. 2020 , eprint=
2020
-
[54]
and Zettlemoyer, Luke , booktitle=
Joshi, Mandar and Choi, Eunsol and Weld, Daniel S. and Zettlemoyer, Luke , booktitle=. 2017 , publisher=
2017
-
[55]
2019 , publisher=
Kim, Chris Dongjoo and Kim, Byeongchang and Lee, Hyunmin and Kim, Gunhee , booktitle=. 2019 , publisher=
2019
-
[56]
Koh, Jing Yu and Lo, Robert and Jang, Lawrence and Duvvur, Vikram and Lim, Ming Chong and Huang, Po-Yao and Neubig, Graham and Zhou, Shuyan and Salakhutdinov, Ruslan and Fried, Daniel , journal=
-
[57]
Koreeda, Yuta and Manning, Christopher D. , year=. 2110.01799 , archivePrefix=
-
[58]
Krithara, Anastasia and Nentidis, Anastasios and Bougiatiotis, Konstantinos and Paliouras, Georgios , journal=
-
[59]
Transactions of the Association for Computational Linguistics , volume=
Natural Questions: A Benchmark for Question Answering Research , author=. Transactions of the Association for Computational Linguistics , volume=
-
[60]
2012 , howpublished=
The Winograd Schema Challenge , author=. 2012 , howpublished=
2012
-
[61]
2305.11747 , archivePrefix=
Li, Junyi and Cheng, Xiaoxue and Zhao, Wayne Xin and Nie, Jian-Yun and Wen, Ji-Rong , year=. 2305.11747 , archivePrefix=
-
[62]
Transactions on Machine Learning Research , year=
Holistic Evaluation of Language Models , author=. Transactions on Machine Learning Research , year=
-
[63]
2109.07958 , archivePrefix=
Lin, Stephanie and Hilton, Jacob and Evans, Owain , year=. 2109.07958 , archivePrefix=
-
[64]
2308.03688 , archivePrefix=
Liu, Xiao and Yu, Hao and Zhang, Hanchen and Xu, Yifan and Lei, Xuanyu and Lai, Hanyu and Gu, Yu and Ding, Hangliang and Men, Kaiwen and Yang, Kejuan and Zhang, Shudan and Deng, Xiang and Zeng, Aohan and Du, Zhengxiao and Zhang, Chenhui and Shen, Sheng and Zhang, Tianjun and S...
-
[65]
2024 , howpublished=
2024
-
[66]
Miserendino, Samuel and Wang, Michele and Patwardhan, Tejal and Kuo, Charles E. and Dias, Rachel and Thacker, Phoebe and Thanneeru, Vishnu and Eapen, Suhas and Chastain, Eric and Barr, Alexandra and Thacker, Benjamin and Yau, Alvin and Li, David and Ludwinski, Pierce and Chabo...
-
[67]
Abstractive Text Summarization Using Sequence-to-Sequence
Nallapati, Ramesh and Zhou, Bowen and dos Santos, C. Abstractive Text Summarization Using Sequence-to-Sequence. Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning , pages=
-
[68]
Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages=
Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization , author=. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , pages=
2018
-
[69]
Rajpurkar, Pranav and Zhang, Jian and Lopyrev, Konstantin and Liang, Percy , booktitle=
-
[70]
2305.13117 , archivePrefix=
Schlichtkrull, Michael and Guo, Zhijiang and Vlachos, Andreas , year=. 2305.13117 , archivePrefix=
-
[71]
2405.07960 , archivePrefix=
Schmidgall, Samuel and Ziaei, Rojin and Harris, Carl and Reis, Eduardo and Jopling, Jeffrey and Moor, Michael , year=. 2405.07960 , archivePrefix=
-
[72]
2018 , publisher=
Thorne, James and Vlachos, Andreas and Christodoulopoulos, Christos and Mittal, Arpit , booktitle=. 2018 , publisher=
2018
-
[73]
2026 , howpublished=
Step Exams , author=. 2026 , howpublished=
2026
-
[74]
, booktitle=
Wang, Alex and Pruksachatkun, Yada and Nangia, Nikita and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel R. , booktitle=
-
[75]
2407.15711 , archivePrefix=
Yoran, Ori and Wolfson, Tomer and Ram, Ori and Berant, Jonathan , year=. 2407.15711 , archivePrefix=
-
[76]
Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin , booktitle=
-
[77]
2105.07624 , archivePrefix=
Zhu, Fengbin and Lei, Wenqiang and Huang, Youcheng and Wang, Chao and Zhang, Shuo and Lv, Jiancheng and Feng, Fuli and Chua, Tat-Seng , year=. 2105.07624 , archivePrefix=
-
[78]
Zhuo, Terry Yue and Vu, Minh Chien and Chim, Jenny and Hu, Han and Yu, Wenhao and Widyasari, Ratnadira and Yusuf, Imam Nur Bani and Zhan, Haolan and He, Junda and Paul, Indraneil and Brunner, Simon and Gong, Chen and Nguyen, Thong Hoang and Phan, Nam Dinh and Yan, Xingyao and ...
-
[79]
, title =
Grant, Robert M. , title =. Strategic Management Journal , volume =
-
[80]
Guha, Neel and Nyarko, Julian and Ho, Daniel E. and R. 2023 , eprint =
2023
-
[81]
doi:10.48550/arXiv.2509.14856 , url =
Guo, Hanyang and Zheng, Xunjin and Liao, Zihan and Yu, Hang and Di, Peng and Zhang, Ziyin and Dai, Hong-Ning , year =. doi:10.48550/arXiv.2509.14856 , url =. 2509.14856 , archivePrefix =
-
[82]
and Amodei, Dario and Kaplan, Jared and Clark, Jack and Ganguli, Deep , institution =
Handa, Kunal and Tamkin, Alex and McCain, Miles and Huang, Saffron and Durmus, Esin and Heck, Sarah and Mueller, Jared and Hong, Jerry and Ritchie, Stuart and Belonax, Tim and Troy, Kevin K. and Amodei, Dario and Kaplan, Jared and Clark, Jack and Ganguli, Deep , institution =....
2025
-
[83]
, journal =
Handel, Michael J. , journal =. The. 2016 , doi =
2016
-
[84]
New England Journal of Medicine , volume=
A Surgical Safety Checklist to Reduce Morbidity and Mortality in a Global Population , author=. New England Journal of Medicine , volume=
-
[85]
2024 , publisher =
Hu, Xueyu and Zhao, Ziyu and Wei, Shuang and Chai, Ziwei and Ma, Qianli and Wang, Guoyin and Wang, Xuwu and Su, Jing and Xu, Jingjing and Zhu, Ming and Cheng, Yao and Yuan, Jianbo and Li, Jiwei and Kuang, Kun and Yang, Yang and Yang, Hongxia and Wu, Fei , booktitle =. 2024 , p...
2024
-
[86]
International Standard Classification of Occupations 2008 (
2008
-
[87]
2311.11944 , archivePrefix =
Islam, Pranab and Kannappan, Anand and Kiela, Douwe and Qian, Rebecca and Scherrer, Nino and Vidgen, Bertie , year =. 2311.11944 , archivePrefix =
-
[88]
and Paullada, Amandalynne and Denton, Emily and Hanna, Alex , booktitle =
Raji, Inioluwa Deborah and Bender, Emily M. and Paullada, Amandalynne and Denton, Emily and Hanna, Alex , booktitle =. 2021 , url =
2021
-
[89]
and Wallach, Hanna , title =
Jacobs, Abigail Z. and Wallach, Hanna , title =. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , series =. 2021 , publisher =
2021
-
[90]
ACM Computing Surveys , volume=
Survey of Hallucination in Natural Language Generation , author=. ACM Computing Surveys , volume=
-
[91]
and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik R
Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik R. , booktitle =. 2024 , url =
2024
-
[92]
2024 , eprint =
Jing, Liqiang and Huang, Zhehui and Wang, Xiaoyang and Yao, Wenlin and Yu, Wenhao and Ma, Kaixin and Zhang, Hongming and Du, Xinya and Yu, Dong , title =. 2024 , eprint =
2024
-
[93]
Journal of Educational Measurement , volume =
Validating the Interpretations and Uses of Test Scores , author =. Journal of Educational Measurement , volume =. 2013 , doi =
2013
-
[94]
Academy of Management Annals , year =
Algorithms at Work: The New Contested Terrain of Control , author =. Academy of Management Annals , year =
-
[95]
doi:10.48550/arXiv.2601.08806 , url =
Kottamasu, Abhi and Mahapatra, Chirag and Lee, Sam and Pan, Ben and Barthwal, Aakash and Datta, Akul and Gupta, Anurag and Mehta, Pranav and Arun, Ajay and Alberti, Silas and Hiremath, Adarsh and Foody, Brendan and Vidgen, Bertie , year =. doi:10.48550/arXiv.2601.08806 , url =...
-
[96]
2026 , eprint =
Kumar, Deepak , title =. 2026 , eprint =. doi:10.48550/arXiv.2603.26130 , url =
2026 doi
-
[97]
Retrieval-Augmented Generation for Knowledge-Intensive
Lewis, Patrick and Perez, Ethan and Piktus, Aleksandra and Petroni, Fabio and Karpukhin, Vladimir and Goyal, Naman and K. Retrieval-Augmented Generation for Knowledge-Intensive. Advances in Neural Information Processing Systems , year =
-
[98]
2024 , eprint =
Ma, Zeyao and Zhang, Bohan and Zhang, Jing and Yu, Jifan and Zhang, Xiaokang and Zhang, Xiaohan and Luo, Sijia and Wang, Xi and Tang, Jie , title =. 2024 , eprint =
2024
-
[99]
1962 , publisher =
Machlup, Fritz , title =. 1962 , publisher =
1962
-
[100]
ACM Computing Surveys , volume=
The Interdisciplinary Study of Coordination , author=. ACM Computing Surveys , volume=. 1994 , publisher=
1994
-
[101]
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=
On Faithfulness and Factuality in Abstractive Summarization , author=. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=
-
[102]
2018 , doi =
McInnes, Leland and Healy, John and Saul, Nathaniel and Grossberger, Lukas , journal =. 2018 , doi =
2018
-
[103]
and Shaw, Alexander G
Merrill, Mike A. and Shaw, Alexander G. and Carlini, Nicholas and Li, Boxuan and Raj, Harsh and Bercovich, Ivan and Shi, Lin and Shin, Jeong Yeon and Walshe, Thomas and Buchanan, E. Kelly and Shen, Junhong and Ye, Guanghao and Lin, Haowei and Poulos, Jason and Marten, Ryan and...
2026
-
[104]
Educational Measurement , editor =
Messick, Samuel , title =. Educational Measurement , editor =. 1989 , publisher =
1989
-
[105]
Generative
Jaffe, Sonia and Shah, Neha Parikh and Butler, Jenna and Farach, Alex and Cambon, Alexia and Hecht, Brent and Schwarz, Michael and Teevan, Jaime , institution =. Generative. 2024 , month = jul, url =
2024
-
[106]
2026 , howpublished =
2026
-
[107]
doi:10.48550/arXiv.2603.08655 , url =
Opsahl-Ong, Krista and Singhvi, Arnav and Collins, Jasmine and Zhou, Ivan and Wang, Cindy and Baheti, Ashutosh and Oertell, Owen and Portes, Jacob and Havens, Sam and Elsen, Erich and Bendersky, Michael and Zaharia, Matei and Chen, Xing , year =. doi:10.48550/arXiv.2603.08655 ...
-
[108]
Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , year =
Petroni, Fabio and Piktus, Aleksandra and Fan, Angela and Lewis, Patrick and Yazdani, Majid and De Cao, Nicola and Thorne, James and Jernite, Yacine and Karpukhin, Vladimir and Maillard, Jean and Plachouras, Vassilis and Rockt. Proceedings of the 2021 Conference of the North A...
2021
-
[109]
1977 , address =
Porat, Marc Uri , title =. 1977 , address =
1977
-
[110]
Thirty-seventh Conference on Neural Information Processing Systems , year=
Toolformer: Language Models Can Teach Themselves to Use Tools , author=. Thirty-seventh Conference on Neural Information Processing Systems , year=
-
[111]
1987 , publisher=
Plans and Situated Actions: The Problem of Human-Machine Communication , author=. 1987 , publisher=
1987
-
[112]
Human-Machine Reconfigurations: Plans and Situated Actions , author =
-
[113]
doi:10.48550/arXiv.2509.16941 , url =
Deng, Xiang and Da, Jeff and Pan, Edwin and He, Yannis Yiming and Ide, Charles and Garg, Kanak and Lauffer, Niklas and Park, Andrew and Pasari, Nitin and Rane, Chetan and Sampath, Karmini and Krishnan, Maya and Kundurthy, Srivatsa and Hendryx, Sean and Wang, Zifan and Bharadwa...
-
[114]
Proceedings of the CHI Conference on Human Factors in Computing Systems , articleno =
Tankelevitch, Lev and Kewenig, Viktor and Simkute, Auste and Scott, Ava Elizabeth and Sarkar, Advait and Sellen, Abigail and Rintel, Sean , title =. Proceedings of the CHI Conference on Human Factors in Computing Systems , articleno =. 2024 , publisher =
2024
-
[115]
Measuring the Occupational Impact of
Tolan, Songul and Pesole, Annarosa and Mart. Measuring the Occupational Impact of. Journal of Artificial Intelligence Research , year =
-
[116]
2601.14242 , archivePrefix =
Vidgen, Bertie and Mann, Austin and Fennelly, Abby and Stanly, John Wright and Rothman, Lucas and Burstein, Marco and Benchek, Julien and Ostrofsky, David and Ravichandran, Anirudh and Sur, Debnil and Venugopal, Neel and Hsia, Alannah and Robinson, Isaac and Huang, Calix and V...
-
[117]
2024 , eprint =
Wang, Zilong and Cui, Yuedong and Zhong, Li and Zhang, Zimin and Yin, Da and Lin, Bill Yuchen and Shang, Jingbo , title =. 2024 , eprint =
2024
-
[118]
Frontiers of Computer Science , volume=
A Survey on Large Language Model Based Autonomous Agents , author=. Frontiers of Computer Science , volume=. 2024 , doi=
2024
-
[119]
SSRN Working Paper , year =
The Impact of Artificial Intelligence on the Labor Market , author =. SSRN Working Paper , year =
-
[120]
2007 , publisher=
Managing the Unexpected: Resilient Performance in an Age of Uncertainty , author=. 2007 , publisher=
2007
-
[121]
code-review-benchmark: Pull Requests with Golden Review Comments , author =
- [122]
-
[123]
2024 , eprint =
Xie, Tianbao and Zhang, Danyang and Chen, Jixuan and Li, Xiaochuan and Zhao, Siheng and Cao, Ruisheng and Hua, Toh Jing and Cheng, Zhoujun and Shin, Dongchan and Lei, Fangyu and Liu, Yitao and Xu, Yiheng and Zhou, Shuyan and Savarese, Silvio and Xiong, Caiming and Zhong, Victo...
2024
-
[124]
International Conference on Learning Representations , year=
ReAct: Synergizing Reasoning and Acting in Language Models , author=. International Conference on Learning Representations , year=
-
[125]
and Zhu, Hao and Zhou, Xuhui and Lo, Robert and Sridhar, Abishek and Cheng, Xianyi and Ou, Tianyue and Bisk, Yonatan and Fried, Daniel and Alon, Uri and Neubig, Graham , year =
Zhou, Shuyan and Xu, Frank F. and Zhu, Hao and Zhou, Xuhui and Lo, Robert and Sridhar, Abishek and Cheng, Xianyi and Ou, Tianyue and Bisk, Yonatan and Fried, Daniel and Alon, Uri and Neubig, Graham , year =. 2307.13854 , archivePrefix =
-
[126]
ArXiv , year=
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity , author=. ArXiv , year=. doi:10.48550/arXiv.2507.09089 , url=. 2507.09089 , archivePrefix=
2025 doi
-
[127]
and Geng, Gloria and Park, Danny and Zou, James and Ng, Andrew Y
Jiang, Yixing and Black, Kameron C. and Geng, Gloria and Park, Danny and Zou, James and Ng, Andrew Y. and Chen, Jonathan H. , year =. 2501.14654 , archivePrefix =
-
[128]
Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , year=
MTEB: Massive Text Embedding Benchmark , author=. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics , year=
-
[129]
2026 , month = feb, howpublished =
Why. 2026 , month = feb, howpublished =
2026
-
[130]
Advances in Neural Information Processing Systems , year=
BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models , author=. Advances in Neural Information Processing Systems , year=
-
[131]
Solved Issues
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study , author=. ArXiv , year=. doi:10.48550/arXiv.2503.15223 , url=. 2503.15223 , archivePrefix=
- [132]
-
[133]
2019 , pages=
Jaume, Guillaume and Ekenel, Hazim Kemal and Thiran, Jean-Philippe , booktitle=. 2019 , pages=
2019
-
[134]
2004 , pages=
Lin, Chin-Yew , booktitle=. 2004 , pages=
2004
-
[135]
Mathew, Minesh and Karatzas, Dimosthenis and Jawahar, C. V. , booktitle=. 2021 , pages=
2021
-
[136]
The Thirteenth International Conference on Learning Representations , year=
-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains , author=. The Thirteenth International Conference on Learning Representations , year=
Reviewed May 25, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.