REVIEW 3 major objections 6 minor 55 references
Frontier AI now scores above 87 percent on business case analysis against expert instructor rubrics, but fully complete answers remain rare.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 00:46 UTC pith:CQ5CLALP
load-bearing objection A real benchmark contribution with a load-bearing calibration gap: relative scores and the two-year trend are trustworthy, but the 87% absolute headline needs more human data before you cite it as a capability claim. the 3 major comments →
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On its own terms, the paper establishes that frontier LLMs already produce analytically strong drafts on open-ended business case work: GPT-5.4 scores 87.2 percent, Claude Sonnet 4.6 scores 88.4 percent, and Gemini 3 Flash Preview scores 81.6 percent under a partial-credit Standard scoring that credits each rubric criterion independently. Complete Answer scoring, which requires every criterion on a question to be satisfied, is far lower — 47.6, 49.6, and 32.0 percent respectively — so high partial credit coexists with rare full completeness. The paper further finds that the 23.3-point gain over two years is broad-based (numerical, subjective, and non-numerical questions all improve), that di
What carries the argument
The load-bearing mechanism is the case-grounded evaluation pipeline: licensed business school case narratives paired with exam-style questions and expert-written instructor solutions, with each solution converted into an equally-weighted checklist rubric. A fixed AI judge (Gemini 2.5 Flash) awards binary credit per criterion against the reference solution, producing two metrics — Standard scoring (the rubric-weighted fraction of satisfied criteria) and Complete Answer scoring (whether every criterion is satisfied). A blinded human-annotation protocol on ten assignments checks the automated grading, and O*NET mapping connects questions to occupational work activities. The benchmark's validity
Load-bearing premise
The scores stand or fall with a single AI judge's binary pass/fail verdict on each rubric criterion: it was checked against expert graders on only ten assignments (Spearman ρ=0.54, with the judge about 6.4 points lenient), and no expert baseline exists on the full 615 questions, so if the judge's 'criterion satisfied' calls drift from expert judgment at scale, the 87–88 percent headline overstates — or understates — true capability.
What would settle it
Grade a fresh sample of model answers with independent expert instructors using their own rubrics, on a larger subset than the ten assignments validated here. If human-assigned partial-credit scores run far below AI judge scores (beyond the observed 6.4-point leniency) or rank the models differently, the 87–88 percent headline overstates capability. Conversely, if expert graders score the same answers at or above the AI judge's levels, the paper's framing stands.
If this is right
- Business schools face a shifted design problem: producing a plausible case draft is cheap, so curricula should concentrate on verification, completeness, and recognizing what a fully sufficient answer requires.
- Entry-level analytical work appears most exposed where structured analytical tasks approach ceiling performance, while open-ended advisory work (identifying opportunities, advising on financial matters) remains the hardest slice.
- The remaining performance gap is a completeness gap, not a knowledge gap, so model development aimed at synthesis, trade-off analysis, and multi-criterion satisfaction is the natural next target.
- The within-family two-year trajectory is a lower bound on capability growth; today's Standard scores should not be read as a ceiling.
- The benchmark design — expert-written cases plus instructor-solution rubrics — generalizes to other professional domains that train judgment through narrative cases.
Where Pith is reading between the lines
- If the +6.4 percentage-point judge leniency observed on ten assignments holds across the full benchmark, true partial-credit capability could be several points below the reported 87–88 percent; a wider human-grader sample would settle the size of the correction.
- Because rubric grading only checks whether stated criteria are met, a model could earn high scores while padding answers with unsupported claims; auditing responses for extraneous or false content would be a natural stress test of whether high partial credit means high analytical quality.
- The finding that little difficulty is explained by discipline or question-type labels suggests that harder case questions could be mined directly from residual scores to build adversarial training and evaluation sets targeting completeness failures.
- The benchmark is single-turn and rubric-guided, so multi-turn clarification, information gathering, and accountability — where human advantage may remain — are left unmeasured; pairing the benchmark with interactive protocols would test whether the completeness gap shrinks when models can ask questions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces BusinessCaseBench, a benchmark of 615 open-ended analytical questions derived from 238 licensed business-school case studies across 18 disciplines, each paired with a checklist rubric constructed from the instructor case solution. Frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash Preview) are evaluated single-turn, with a fixed LLM-as-judge (Gemini 2.5 Flash) awarding binary credit per rubric criterion. Two metrics are reported: Standard scoring (partial credit, Eq. 1–2) and Complete Answer scoring (all criteria satisfied, Eq. 3–4). The central claims are that frontier models achieve high absolute rubric-graded scores (87–88% Standard; 47–50% Complete Answer), that capability within one model family improved by 23.3 percentage points over two years, and that residual difficulty is largely case- and question-specific rather than categorical. The paper also maps questions to O*NET work activities and reports an extended result for a later Anthropic model.
Significance. The benchmark addresses a real gap: existing LLM benchmarks under-represent open-ended analytical knowledge work, and business cases provide a natural expert-written format with instructor solutions. The paper ships reproducible code, pinned model identifiers, and aggregated outputs, and it includes several careful design elements: a contamination audit with InfiniGram, bootstrap confidence intervals, a three-judge rank-robustness check, and a blinded human annotation protocol. If the absolute score claims are validated, the benchmark would be a valuable instrument for tracking progress on economically relevant reasoning tasks. However, the current evidence for calibration of the automated judge to human expert standards is thin, and the central claims are stated in absolute terms. The relative rankings and within-family trajectory are more robust than the absolute levels.
major comments (3)
- [Methods, 'Human annotation interface and blinded validation protocol'; SI Appendix Table S8] The headline claims (87–88% Standard scoring, 47–50% Complete Answer scoring, 23.3pp within-family gain) are absolute levels, but the LLM-as-judge is validated on only ten assignments, with Spearman ρ=0.54, 40% of automated scores within 10pp of human scores, and a mean +6.4pp automated leniency. The three-judge robustness check in Methods only shows that three LLM judges agree on model rankings (W=1.0); it cannot detect a shared absolute bias against human standards. Since no human expert baseline is collected on the 615 benchmark questions, the reported absolute scores may be systematically over- or under-stated. Please either (a) expand human validation to a statistically meaningful stratified sample, (b) report bias-corrected scores using the measured +6.4pp leniency, or (c) explicitly reframe the central claims as relative/comparative rather than absolute.
- [Methods, 'Benchmark construction'; SI Appendix Section C3] Rubric construction and grading are both performed by Gemini-family models (rubrics via gemini-2.5-pro/flash; grading via a fixed Gemini 2.5 Flash judge). The rubric defines what counts as a 'satisfied criterion,' so any systematic rubric-generation bias directly shifts scores. The validation of rubric quality rests on the same ten assignments, with 100% acceptability rated by annotators who had already seen the automated rubric. This is not sufficient evidence that rubrics are well-calibrated across 615 questions and 18 disciplines. Please provide larger-sample evidence on rubric properties (e.g., criterion count, specificity, overlap with independently authored rubrics) or a rubric-quality audit on a random sample of the full benchmark.
- [SI Appendix D6, Tables S12–S14; Results, 'Difficulty is largely case- and question-specific'] The regression analysis reports adjusted R² = 0.026 for question-type tags, 0.036 for discipline, and 0.051 for both together. The text concludes that 'discipline is associated with far wider differences in model difficulty than question type' (Introduction) and that 'knowing which case a question comes from predicts its score far better than the coarse metadata labels do.' The 1-point adjusted-R² gap between discipline and question type is small, and the case-level adjusted R² (0.22) is descriptive, not predictive. The strength of the claim in the abstract and introduction seems disproportionate to the evidence. Please temper the wording or report variance components that directly quantify the relative contributions.
minor comments (6)
- [Abstract and Significance statement] The benchmark is described as 'validated' and 'discipline-spanning' early in the paper; given the limited human validation (10 assignments), consider qualifying this as 'directionally validated' or 'provisionally validated' until a larger human study is performed.
- [Methods, Eq. (1)–(4)] The definitions of Standard and Complete Answer scoring are clear, but the term 'Standard scoring' is not defined in the main text before first use in Fig. 2; consider adding a one-sentence definition in the Overview section.
- [Results, 'Aggregate performance'] Confidence intervals for the two leading models 'overlap throughout' (Fig. 2A). This is stated as if it implies statistical equivalence; consider adding a formal test of the difference or explicit language about uncertainty.
- [Extended Results, 'Claude Fable 5'] The phrase 'Mythos-class model' is unexplained. Either define the class or remove the term. Also, the claim that similarity 'may reflect substantial overlap in training data' is speculative; consider softening.
- [SI Appendix, Table S8] The comparison of automated–human agreement (40% within 10pp) to inter-annotator agreement (46.7% within 10pp) is made on n=10 assignments. Given the tiny sample, this comparison should be flagged as illustrative rather than as evidence of equivalence.
- [References] Some reference entries are malformed (e.g., 'and . Sachdeva' in the Gemini 2.5 entry, 'T. Computer' as an author). Please fix these in the final version.
Circularity Check
No significant circularity: the benchmark scores are empirical measurements produced by a fixed evaluation pipeline, not derivations that reduce to their own inputs.
full rationale
This is an empirical benchmark paper rather than a derivation chain. The headline scores (e.g., 87.2% for GPT-5.4) are measurements from a fixed pipeline: case narratives and questions are paired with instructor reference solutions, rubrics are derived from those solutions, and a single held-constant LLM-as-judge awards binary credit per rubric criterion. No parameter is fitted to the 615-question set or to the 23.3-point within-family trajectory, so the reported scores are not statistically forced by construction. The rubric-generation and judging both involve Gemini-family models, and the same instructor reference solution anchors both rubric construction and grading, but this is a same-family evaluation choice rather than a definitional identity: the solver models never see the reference solution, and the scores are externally probed through a blinded human-annotation protocol (10 assignments, Spearman ρ = 0.54) and a three-judge robustness check that preserves model rankings. The small validation sample and +6.4pp judge leniency are genuine measurement/calibration limitations, not circularity, because the automated scores are reported without being calibrated on the human data. Self-citations (DataDreamer; Dell'Acqua et al., which shares a co-author) are used only as tooling or as supporting references for downstream implications and are not load-bearing for the central benchmark results. No load-bearing step reduces to its own inputs, so the appropriate circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (3)
- Rubric criterion weights (equal, 1/k_j) =
1/k_j per criterion
- LLM-as-judge model (Gemini 2.5 Flash) =
google/gemini-2.5-flash
- Universal-hardness threshold =
70%
axioms (5)
- domain assumption Instructor case solutions are expert-written gold standards for high-quality analytical reasoning in business.
- domain assumption An equally-weighted checklist rubric with binary criteria captures answer quality for open-ended subjective tasks.
- domain assumption LLM-as-judge credit assignments are directionally valid measures of whether a response satisfies a criterion.
- domain assumption Scores are not substantially inflated by pre-training memorization of licensed case materials.
- domain assumption Single-turn full-case prompts faithfully represent the analytical knowledge work being measured.
read the original abstract
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The "case method" form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.
Reference graph
Works this paper leans on
-
[1]
doi:10.48550/arXiv.2109.00122 , urldate =
Chen, Zhiyu and Chen, Wenhu and Smiley, Charese and Shah, Sameena and Borova, Iana and Langdon, Dylan and Moussa, Reema and Beane, Matt and Huang, Ting-Hao and Routledge, Bryan and Wang, William Yang , month = may, year =. doi:10.48550/arXiv.2109.00122 , urldate =
-
[2]
doi:10.48550/arXiv.2311.06602 , urldate =
Koncel-Kedziorski, Rik and Krumdick, Michael and Lai, Viet and Reddy, Varshini and Lovering, Charles and Tanner, Chris , month = mar, year =. doi:10.48550/arXiv.2311.06602 , urldate =
-
[3]
Wang, Liya and Yi, David and Jose, Damien and Passarelli, John and Gao, James and Leventis, Jordan and Li, Kang , month = jun, year =. Enterprise. doi:10.48550/arXiv.2506.20274 , urldate =
-
[4]
doi:10.48550/arXiv.2303.17564 , urldate =
Wu, Shijie and Irsoy, Ozan and Lu, Steven and Dabravolski, Vadim and Dredze, Mark and Gehrmann, Sebastian and Kambadur, Prabhanjan and Rosenberg, David and Mann, Gideon , month = dec, year =. doi:10.48550/arXiv.2303.17564 , urldate =
-
[5]
doi:10.48550/arXiv.2402.12659 , urldate =
Xie, Qianqian and Han, Weiguang and Chen, Zhengyu and Xiang, Ruoyu and Zhang, Xiao and He, Yueru and Xiao, Mengxi and Li, Dong and Dai, Yongfu and Feng, Duanyu and Xu, Yijing and Kang, Haoqiang and Kuang, Ziyan and Yuan, Chenhan and Yang, Kailai and Luo, Zheheng and Zhang, Tianlin and Liu, Zhiwei and Xiong, Guojun and Deng, Zhiyang and Jiang, Yuechen and ...
-
[6]
doi:10.48550/arXiv.2402.18667 , urldate =
Xia, Congying and Xing, Chen and Du, Jiangshu and Yang, Xinyi and Feng, Yihao and Xu, Ran and Yin, Wenpeng and Xiong, Caiming , month = feb, year =. doi:10.48550/arXiv.2402.18667 , urldate =
-
[7]
doi:10.48550/arXiv.2506.17863 , urldate =
Liu, Haoran and Tahmasbi, Amir and Haque, Ehtesham Sam and Jain, Purak , month = jun, year =. doi:10.48550/arXiv.2506.17863 , urldate =
-
[8]
doi:10.48550/arXiv.2407.05291 , urldate =
Boisvert, Léo and Thakkar, Megh and Gasse, Maxime and Caccia, Massimo and Chezelles, Thibault Le Sellier De and Cappart, Quentin and Chapados, Nicolas and Lacoste, Alexandre and Drouin, Alexandre , month = feb, year =. doi:10.48550/arXiv.2407.05291 , urldate =
-
[9]
Drouin, Alexandre and Gasse, Maxime and Caccia, Massimo and Laradji, Issam H. and Verme, Manuel Del and Marty, Tom and Boisvert, Léo and Thakkar, Megh and Cappart, Quentin and Vazquez, David and Chapados, Nicolas and Lacoste, Alexandre , month = jul, year =. doi:10.48550/arXiv.2403.07718 , urldate =
-
[10]
doi:10.48550/arXiv.2411.02305 , urldate =
Huang, Kung-Hsiang and Prabhakar, Akshara and Dhawan, Sidharth and Mao, Yixin and Wang, Huan and Savarese, Silvio and Xiong, Caiming and Laban, Philippe and Wu, Chien-Sheng , month = feb, year =. doi:10.48550/arXiv.2411.02305 , urldate =
-
[11]
doi:10.48550/arXiv.2505.18878 , urldate =
Huang, Kung-Hsiang and Prabhakar, Akshara and Thorat, Onkar and Agarwal, Divyansh and Choubey, Prafulla Kumar and Mao, Yixin and Savarese, Silvio and Xiong, Caiming and Wu, Chien-Sheng , month = may, year =. doi:10.48550/arXiv.2505.18878 , urldate =
-
[12]
Li, Haohang and Cao, Yupeng and Yu, Yangyang and Javaji, Shashidhar Reddy and Deng, Zhiyang and He, Yueru and Jiang, Yuechen and Zhu, Zining and Subbalakshmi, Koduvayur and Xiong, Guojun and Huang, Jimin and Qian, Lingfei and Peng, Xueqing and Xie, Qianqian and Suchow, Jordan W. , month = dec, year =. doi:10.48550/arXiv.2412.18174 , urldate =
-
[13]
and Zhang, Hao and Stoica, Ion , month = sep, year =
Kwon, Woosuk and Li, Zhuohan and Zhuang, Siyuan and Sheng, Ying and Zheng, Lianmin and Yu, Cody Hao and Gonzalez, Joseph E. and Zhang, Hao and Stoica, Ion , month = sep, year =. Efficient. doi:10.48550/arXiv.2309.06180 , abstract =
-
[14]
Patel, Ajay and Raffel, Colin and Callison-Burch, Chris , year =. Proceedings of the 62nd. doi:10.18653/v1/2024.acl-long.208 , language =
-
[15]
Nori, Harsha and Daswani, Mayank and Kelly, Christopher and Lundberg, Scott and Ribeiro, Marco Tulio and Wilson, Marc and Liu, Xiaoxuan and Sounderajah, Viknesh and Carlson, Jonathan and Lungren, Matthew P. and Gross, Bay and Hames, Peter and Suleyman, Mustafa and King, Dominic and Horvitz, Eric , month = jul, year =. Sequential. doi:10.48550/arXiv.2506.2...
-
[16]
Xu, Frank F. and Song, Yufan and Li, Boxuan and Tang, Yuxuan and Jain, Kritanjali and Bao, Mengxue and Wang, Zora Z. and Zhou, Xuhui and Guo, Zhitong and Cao, Murong and Yang, Mingyang and Lu, Hao Yang and Martin, Amaad and Su, Zhe and Maben, Leander and Mehta, Raj and Chi, Wayne and Jang, Lawrence and Xie, Yiqing and Zhou, Shuyan and Neubig, Graham , mon...
-
[17]
Wang, Zora Zhiruo and Vijayvargiya, Sanidhya and Chen, Aspen and Zhang, Hanmo and Arangarajan, Venu Arvind and Chen, Jett and Chen, Valerie and Yang, Diyi and Fried, Daniel and Neubig, Graham , month = mar, year =. How. doi:10.48550/arXiv.2603.01203 , urldate =
-
[18]
Patwardhan, Tejal and Dias, Rachel and Proehl, Elizabeth and Kim, Grace and Wang, Michele and Watkins, Olivia and Fishman, Simón Posada and Aljubeh, Marwan and Thacker, Phoebe and Fauconnet, Laurance and Kim, Natalie S. and Chao, Patrick and Miserendino, Samuel and Chabot, Gildas and Li, David and Sharman, Michael and Barr, Alexandra and Glaese, Amelia an...
-
[19]
OpenAI and Achiam, Josh and Adler, Steven and Agarwal, Sandhini and Ahmad, Lama and Akkaya, Ilge and Aleman, Florencia Leoni and Almeida, ... , month = mar, year =. doi:10.48550/arXiv.2303.08774 , urldate =
-
[20]
Singh, Aaditya and Fry, Adam and Perelman, Adam and Tart, Adam and Ganesh, Adi and El-Kishky, Ahmed and McLaughlin, Aidan and Low, Aiden and Ostrow, ... , month = may, year =. doi:10.48550/arXiv.2601.03267 , urldate =
-
[21]
Gemini 3
Deepmind, Google , month = dec, year =. Gemini 3
-
[22]
Anthropic , month = may, year =. System
-
[23]
Redeploying
Anthropic , month = jun, year =. Redeploying
-
[24]
2026 , howpublished =
2026
-
[25]
arXiv preprint arXiv:2401.17377 , year=
Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens , author=. arXiv preprint arXiv:2401.17377 , year=
-
[26]
and Zhang, Hao and Gonzalez, Joseph E
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric P. and Zhang, Hao and Gonzalez, Joseph E. and Stoica, Ion , month = dec, year =. Judging. doi:10.48550/arXiv.2306.05685 , abstract =
-
[27]
doi:10.48550/arXiv.2406.04770 , abstract =
Lin, Bill Yuchen and Deng, Yuntian and Chandu, Khyathi and Brahman, Faeze and Ravichander, Abhilasha and Pyatkin, Valentina and Dziri, Nouha and Bras, Ronan Le and Choi, Yejin , month = oct, year =. doi:10.48550/arXiv.2406.04770 , abstract =
-
[28]
doi:10.48550/arXiv.2503.05142 , abstract =
Wei, Tianjun and Wen, Wei and Qiao, Ruizhi and Sun, Xing and Ma, Jianghong , month = mar, year =. doi:10.48550/arXiv.2503.05142 , abstract =
-
[29]
doi:10.48550/arXiv.2305.14251 , abstract =
Min, Sewon and Krishna, Kalpesh and Lyu, Xinxi and Lewis, Mike and Yih, Wen-tau and Koh, Pang Wei and Iyyer, Mohit and Zettlemoyer, Luke and Hajishirzi, Hannaneh , month = oct, year =. doi:10.48550/arXiv.2305.14251 , abstract =
-
[30]
Comanici, Gheorghe and Bieber, Eric and Schaekermann, Mike and Pasupat, Ice and Sachdeva, ... , month = dec, year =. Gemini 2.5:. doi:10.48550/arXiv.2507.06261 , abstract =
-
[31]
The Annals of Mathematical Statistics , author =
The. The Annals of Mathematical Statistics , author =. 1939 , pages =. doi:10.1214/aoms/1177732186 , language =
arXiv 1939
-
[32]
doi:10.48550/arXiv.2502.18443 , abstract =
Poznanski, Jake and Rangapur, Aman and Borchardt, Jon and Dunkelberger, Jason and Huff, Regan and Lin, Daniel and Rangapur, Aman and Wilhelm, Christopher and Lo, Kyle and Soldaini, Luca , month = jul, year =. doi:10.48550/arXiv.2502.18443 , abstract =
-
[33]
Hendrycks, Dan and Burns, Collin and Basart, Steven and Zou, Andy and Mazeika, Mantas and Song, Dawn and Steinhardt, Jacob , month = jan, year =. Measuring. doi:10.48550/arXiv.2009.03300 , abstract =
-
[34]
and Zettlemoyer, Luke , month = may, year =
Joshi, Mandar and Choi, Eunsol and Weld, Daniel S. and Zettlemoyer, Luke , month = may, year =. doi:10.48550/arXiv.1705.03551 , abstract =
-
[35]
doi:10.48550/arXiv.1905.07830 , abstract =
Zellers, Rowan and Holtzman, Ari and Bisk, Yonatan and Farhadi, Ali and Choi, Yejin , month = may, year =. doi:10.48550/arXiv.1905.07830 , abstract =
-
[36]
Chen, Mark and Tworek, Jerry and Jun, Heewoo and Yuan, Qiming and Pinto, Henrique Ponde de Oliveira and Kaplan, Jared and Edwards, Harri and Burda, Yuri and Joseph, Nicholas and Brockman, Greg and Ray, Alex and Puri, Raul and Krueger, Gretchen and Petrov, Michael and Khlaaf, Heidy and Sastry, Girish and Mishkin, Pamela and Chan, Brooke and Gray, Scott and...
-
[37]
Jimenez, Carlos E. and Yang, John and Wettig, Alexander and Yao, Shunyu and Pei, Kexin and Press, Ofir and Narasimhan, Karthik , month = nov, year =. doi:10.48550/arXiv.2310.06770 , abstract =
-
[38]
Merrill, Mike A. and Shaw, Alexander G. and Carlini, Nicholas and Li, Boxuan and Raj, Harsh and Bercovich, Ivan and Shi, Lin and Shin, Jeong Yeon and Walshe, Thomas and Buchanan, E. Kelly and Shen, Junhong and Ye, Guanghao and Lin, Haowei and Poulos, Jason and Wang, Maoyu and Nezhurina, Marianna and Jitsev, Jenia and Lu, Di and Mastromichalakis, Orfeas Me...
-
[39]
Hendrycks, Dan and Burns, Collin and Kadavath, Saurav and Arora, Akul and Basart, Steven and Tang, Eric and Song, Dawn and Steinhardt, Jacob , month = nov, year =. Measuring. doi:10.48550/arXiv.2103.03874 , abstract =
-
[40]
US-China Education Review A , author =
An. US-China Education Review A , author =
-
[41]
Harvard Business Impact Education , year =
Nitin Nohria , title =. Harvard Business Impact Education , year =
-
[42]
Handel, Michael J. , title =. Journal for Labour Market Research , year =. doi:10.1007/s12651-016-0199-8 , url =
-
[43]
2024 , month = dec, note =
Chen, Wilbur Xinyuan and Srinivasan, Suraj and Zakerinia, Saleh , title =. 2024 , month = dec, note =
2024
-
[44]
Csaszar, Felipe A. and Jacobides, Michael G. and Zemsky, Peter , title =. Strategic Organization , year =. doi:10.1177/14761270251385484 , url =
-
[45]
and Ketkar, Harsh and Kim, Hyunjin , title =
Csaszar, Felipe A. and Ketkar, Harsh and Kim, Hyunjin , title =. Strategy Science , volume =. 2024 , doi =. https://doi.org/10.1287/stsc.2024.0190 , abstract =
arXiv 2024
-
[46]
ArXiv , year=
Benchmark Data Contamination of Large Language Models: A Survey , author=. ArXiv , year=
-
[47]
Management Science , volume =
Sebastian Krakowski and Darek Haftor and Johannes Luger and Natallia Pashkevich and Sebastian Raisch , title =. Management Science , volume =. 2025 , doi =
2025
-
[48]
Allen and Rory M
Ryan T. Allen and Rory M. McDonald , title =. Strategy Science , volume =
-
[49]
Kellogg and Saran Rajendran and Lisa Krayer and François Candelon and Karim R
Fabrizio Dell'Acqua and Edward McFowland III and Ethan Mollick and Hila Lifshitz and Katherine C. Kellogg and Saran Rajendran and Lisa Krayer and François Candelon and Karim R. Lakhani , title =. Organization Science , volume =. 2026 , doi =
2026
-
[50]
2025 , eprint=
DataComp-LM: In search of the next generation of training sets for language models , author=. 2025 , eprint=
2025
-
[51]
Journal of machine learning research , volume=
Exploring the limits of transfer learning with a unified text-to-text transformer , author=. Journal of machine learning research , volume=. 2020 , note=
2020
-
[52]
Gao, Leo and Biderman, Stella and Black, Sid and Golding, Laurence and Hoppe, Travis and Foster, Charles and Phang, Jason and He, Horace and Thite, Anish and Nabeshima, Noa and Presser, Shawn and Leahy, Connor , journal=. The
-
[53]
Together Computer , title =
-
[54]
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Soldaini, Luca and Kinney, Rodney and Bhagia, Akshita and Schwenk, Dustin and Atkinson, David and Authur, Russell and Bogin, Ben and Chandu, Khyathi and Dumas, Jennifer and Elazar, Yanai and Hofmann, Valentin and Jha, Ananya and Kumar, Sachin and Lucy, Li and Lyu, Xinxi and Lambert, Nathan and Magnusson, Ian and Morrison, Jacob and Muennighoff, Niklas and...
2024
-
[55]
Brown and Benjamin Mann and Nick Ryder and Melanie Subbiah and Jared Kaplan
Tom B. Brown and Benjamin Mann and Nick Ryder and Melanie Subbiah and Jared Kaplan... , title =. CoRR , volume =. 2020 , url =. 2005.14165 , timestamp =
Pith/arXiv arXiv 2020
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.