Pith. sign in

REVIEW 3 major objections 6 minor 68 references

Research Community Perspectives on "Intelligence" and Large Language Models

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A survey of 303 researchers finds that the term 'intelligence' clusters around generalization, adaptability, and reasoning — and that most researchers do not apply it to current LLM-based systems.

desk verdict A useful, transparent survey of researchers' views on 'intelligence' and LLMs; the headline numbers are solid, but the 'consensus on three criteria' is a marginal-rate claim, not a joint one, and should be reported as such. read the letter →

arxiv 2505.20959 v1 pith:AXAFNNCX submitted 2025-05-27 cs.CL cs.CY

classification cs.CLcs.CY
keywords intelligencelargelanguagemodelssurveyresearchNLPcommunitybeliefsgeneralizationadaptabilityreasoningagendas
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to pin down what researchers actually mean when they call a system 'intelligent', because the term is widely used but rarely defined in NLP and AI research. It reports a survey of 303 researchers across fields including NLP, machine learning, cognitive science, linguistics, and neuroscience. The three criteria that draw the most agreement are generalization (86%), adaptability (83%), and reasoning (83%). Yet 71% of respondents disagree or strongly disagree that current LLM-based systems are intelligent, and only 16.2% list creating intelligent technology as a research goal. The survey's purpose is to give public discourse and the evaluation of marketing claims a concrete, evidence-based point of reference.

What carries the argument

The central machinery is the survey itself: a set of ten well-defined criteria, derived from a collection of 72 definitions of intelligence and the follow-up literature, plus three additional criteria presented without fixed definitions. Respondents selected which criteria were relevant to their notion of 'intelligence', which criteria current LLM-based systems lack, and rated the intelligence of current and future systems on a four-point Likert scale. Phi coefficients, a correlation measure for yes/no choices, reveal which criteria co-occur, while Fisher's exact tests and Cramer's V check whether demographic groups differ in their selections.

What would settle it

A decisive check would be a replication that lets respondents answer on a spectrum and gives them the option to say that 'intelligence' is not a useful concept; if generalization, adaptability, and reasoning no longer draw more than 80% agreement, the observed consensus is an artifact of the binary forced-choice design.

Watch

Extended reading notes

Core claim

The paper's central claim is that the research community's notion of 'intelligence' is more coherent than the absence of explicit definitions would suggest: across fields, occupations, and career stages, more than 80% of respondents select generalization, adaptability, and reasoning as relevant criteria for intelligence. At the same time, the majority position is that current LLM-based systems such as ChatGPT do not qualify as intelligent, with 71% disagreeing or strongly disagreeing. The paper further shows that this skepticism softens for future systems based on similar technology, and that respondents who say their research goal is creating intelligent technology are more likely to attribute intelligence to both current and future systems.

Load-bearing premise

The load-bearing premise is that each researcher has one coherent notion of 'intelligence' that a fixed list of ten binary criteria can capture; the paper itself states that this assumption may be false, and free-text responses about intelligence as a spectrum or as a problematic term suggest it can fail.

Editorial extensions

If this is right

  • If the consensus is real, claims that LLM-based systems are 'intelligent' should be anchored to generalization, adaptability, and reasoning rather than to benchmark scores alone.
  • Because respondents rate these three criteria as both central to intelligence and lacking in current systems, marketing that calls today's LLMs 'intelligent' is out of step with expert usage.
  • Researchers whose stated goal is building intelligent technology are more likely to call current systems intelligent, suggesting that research agendas shape the attribution of intelligence.
  • The finding that 62% of respondents believe no existing test adequately measures intelligence implies that composite benchmarks, the Turing test, and IQ-style tests are not accepted as measures of intelligence by the community.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the sample's skew toward Western academia reflects the field as a whole, the >80% agreement on the top criteria may be a regional consensus; a replication with a non-Western and industry-heavy sample could produce different top criteria.
  • The authors flag that the 16.2% figure for 'creating intelligent technology' as a research goal may be an overestimate due to self-selection; if so, the minority position on LLM intelligence is even smaller in the broader researcher population.
  • A testable implication of the agenda-attribution link is that disagreement over whether LLMs are intelligent may narrow as more research groups adopt the explicit goal of building intelligent systems.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper reports results of an online survey of 303 researchers on the meaning of 'intelligence' and its application to LLM-based systems. The main descriptive findings are that generalization (~86%), adaptability (~83%), and reasoning (~83%) are the most frequently selected criteria for intelligence; that ~71% of respondents disagree that current LLM-based systems are intelligent while ~60% disagree that future systems based on similar technology will be; and that only 16.2% of respondents list creating intelligent technology as a research goal, with those respondents more likely to attribute intelligence to current systems. The authors interpret the results as evidence of cross-field coherence in the notion of 'intelligence' and of majority skepticism about applying the term to LLMs.

Significance. If the findings hold, the study provides a valuable empirical snapshot of how researchers in NLP and adjacent fields use the term 'intelligence' two years after ChatGPT, and it gives concrete numbers to a debate that is often conducted with undefined terms. The paper's strengths include a publicly stated full questionnaire, openly available data and analysis code on GitHub, an honest limitations section that anticipates representativeness and construct-validity concerns, and a set of specific, falsifiable descriptive claims that can be compared with future surveys. The survey is a non-probability sample and the headline percentages lack confidence intervals, so the contribution is best understood as a time- and population-specific snapshot rather than an estimate for researchers in general.

major comments (3)
  1. [§5.3, Table 2, Figure 1] The headline claim that the community agrees on generalization, adaptability, and reasoning is computed from marginal selection rates (≈86%, ≈83%, ≈83%) rather than from the joint rate at which respondents selected all three criteria. Since reasoning is reported as not strongly correlated with the other two, the extent of simultaneous endorsement is not directly visible. The manuscript should report the three-way joint selection rate (and pairwise joint rates) before making the consensus claim. I note from the reported phi=.451 that the generalization-adaptability joint rate is near 77%, and with reasoning at 83% the three-way joint rate is bounded below by roughly 60%, so this is unlikely to be a pure artifact of overlapping margins; however, the exact number should be reported and discussed.
  2. [§5.3, RQ1] The cross-group coherence conclusion rests on null results of many Fisher exact tests, but no multiple-comparison correction is reported and some subgroup sizes are small (the Q8 research-area analysis in §5.4 reports HCI n=8). The two significant criterion effects (embodiment, phi_c=.218, p=.01; environment interaction, phi_c=.225, p=.003) would not survive a simple Bonferroni correction across the number of tests reported, while null results in underpowered groups cannot support a claim of true agreement. I ask the authors to apply a multiple-comparison correction, report effect sizes with intervals, or soften the wording to 'no significant differences were detected in this sample.'
  3. [§4.2, §7 Limitations, Table 9] Q7 has no 'none of these' option, and the paper's own limitations section concedes that the assumption of a single coherent notion of 'intelligence' per respondent might be false. This is not merely hypothetical: Table 9, comment C4 states that the respondent chose criteria 'under duress' because no none option existed, and several comments describe intelligence as a spectrum. The authors should report the distribution of the number of selected criteria, run a sensitivity analysis excluding respondents whose free-text answers reject the binary/coherent framing, and state how the top-3 pattern changes.
minor comments (6)
  1. [Abstract] There is a typo in 'Our results suggests'; it should be 'Our results suggest'.
  2. [§5.3] The use of approximate percentages (e.g., ≈86%) should be accompanied by exact counts and 95% confidence intervals; Table 3 also contains a rounding artifact (current-system row sums to 101%).
  3. [§4.2 / Table 7] The main analysis is based on 10 criteria, but Table 7 lists 13 response options for Q7; please clarify at the first mention of Q7 that the three additional criteria are treated separately and explain the 'well-defined' filter.
  4. [Figure 1] The caption should define the color and width of edges; currently only the phi>|0.1| threshold is stated.
  5. [References] The reference to Piaget spells the author's first name as 'John'; it should be 'Jean Piaget'.
  6. [Table 4 / §4.4] The Q14 option is worded 'Creating technology that qualifies for my notion of intelligence', while the abstract and §6 restate this as 'developing intelligent systems'; please make the rewording explicit or use the original wording.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's central claims are direct self-reported measurements, and its design-dependence concerns are disclosed validity limitations rather than circular reductions.

full rationale

The paper's derivation chain is: a literature-derived criteria list (Section 2.2, Table 1) followed by forced-choice survey responses (Q7-Q14), then marginal selection rates and correlation statistics (Section 5), and finally the conclusions. Each conclusion is a direct summary of self-reported data: the top-3 criteria are the three most frequently selected options (generalization ~86%, adaptability ~83%, reasoning ~83%, Section 5.3); the 71% skepticism is the Likert distribution of Q8 (Table 3); and the research-goal correlation is a phi coefficient between Q8 and Q14 (Section 5.5). No quantity in the paper is fitted to a target and then re-derived, and no equation equates an output to an input by construction. The criteria list is seeded from Legg et al. (2007) and related literature, which does emphasize adaptability, but respondents' selections are empirically variable (embodiment only 23%, environment interaction 45%), showing the instrument does not tautologically confirm its seed. The Limitations section explicitly flags the two design concerns a skeptic might raise, namely 'on the assumption that for each individual there is a coherent notion underlying their use of this term. It is possible that this assumption is false', and the binary/forced-choice structure (free-text comment C4 in Table 10 reports choosing 'under duress'). These are disclosed validity threats affecting interpretation, not circular reductions of the conclusions to the assumptions. The skeptical concern that marginal rates overstate joint consensus is a statistical robustness question, and the paper transparently reports the phi correlations (e.g., phi=.451 between generalization and adaptability) that would bear on it. Author self-citations (Jakobsen and Rogers 2022; Rogers and Luccioni 2024; Ray Choudhury et al. 2022; Rogers et al. 2020) are methodological precedents or background pointers, and none is load-bearing for the survey's central claims. The paper is self-contained against its own published data and questionnaire.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claims rest on survey design choices rather than fitted parameters or new theoretical entities. No free parameters or invented entities are introduced. The load-bearing assumptions are representativeness, coherence of respondents' notions, and completeness of the predefined criteria list.

assumptions (3)
  • domain assumption The 303 self-selected respondents are treated as representative of the research community for consensus claims.
    The paper states 'we find a high degree of consensus across research fields' and 'the community agrees', while the Limitations acknowledge the sample is skewed toward academics from the Western world and affected by self-selection bias. This makes representativeness load-bearing.
  • domain assumption Each respondent has a coherent, stable notion of 'intelligence' that can be elicited via forced-choice questions.
    Explicitly stated in Limitations: 'on the assumption that for each individual there is a coherent notion underlying their use of this term. It is possible that this assumption is false.' Free-text responses indicate spectrum and context-dependence.
  • domain assumption The 13 criteria offered in Q7 are a sufficient and unbiased space for capturing respondents' notions of intelligence.
    Criteria were derived from Legg et al. and selected literature; respondents' free-text comments mention missing dimensions such as social and emotional traits, self-awareness, and spectrum-based views, so the provided list shapes and constrains the consensus finding.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Research Community Perspectives on "Intelligence" and Large Language Models." pith.science (2026). https://pith.science/paper/AXAFNNCX

@misc{pith2026250520959,
  author       = {Pith},
  title        = {Pith review of: Research Community Perspectives on "Intelligence" and Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AXAFNNCX}},
  note         = {Machine review of arXiv:2505.20959}
}
read the original abstract

Despite the widespread use of ''artificial intelligence'' (AI) framing in Natural Language Processing (NLP) research, it is not clear what researchers mean by ''intelligence''. To that end, we present the results of a survey on the notion of ''intelligence'' among researchers and its role in the research agenda. The survey elicited complete responses from 303 researchers from a variety of fields including NLP, Machine Learning (ML), Cognitive Science, Linguistics, and Neuroscience. We identify 3 criteria of intelligence that the community agrees on the most: generalization, adaptability, & reasoning. Our results suggests that the perception of the current NLP systems as ''intelligent'' is a minority position (29%). Furthermore, only 16.2% of the respondents see developing intelligent systems as a research goal, and these respondents are more likely to consider the current systems intelligent.

Figures

Figures reproduced from arXiv: 2505.20959 by the authors.

Figure 1
Figure 1. Correlation between criteria that the survey [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The number of respondents by research area. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The percentage of respondents (x-axis) who [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Sankey diagram illustrating the flow of beliefs [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Number of respondent selecting an entity [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Correlation between criteria that the survey [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Geographical distribution of survey respondents. Most respondents are from the “western” world, although [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Sankey diagrams showing the different beliefs of researchers regarding the intelligence of current LLM [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

68 extracted references · 31 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    J.S. Albus. 1991. https://doi.org/10.1109/21.97471 Outline for a theory of intelligence . 21(3):473--509

  4. [4]

    Yonatan Belinkov and James Glass. 2019. https://doi.org/10.1162/tacl_a_00254 Analysis methods in neural language processing: A survey . Transactions of the Association for Computational Linguistics, 7:49--72

  5. [5]

    Emily M. Bender. 2024. https://faculty.washington.edu/ebender/papers/ACL_2024_Presidential_Address.pdf ACL Is Not an AI Conference

  6. [6]

    Bender and Batya Friedman

    Emily M. Bender and Batya Friedman. 2018. https://doi.org/10.1162/tacl_a_00041 Data Statements for Natural Language Processing : Toward Mitigating System Bias and Enabling Better Science . Transactions of the Association for Computational Linguistics, 6:587--604

  7. [7]

    Bender and Alexander Koller

    Emily M. Bender and Alexander Koller. 2020. https://doi.org/10.18653/v1/2020.acl-main.463 Climbing towards NLU : On meaning, form, and understanding in the age of data . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5185--5198, Online. Association for Computational Linguistics

  8. [8]

    Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023. https://doi.org/10.48550/arXiv.2303.12712 Sparks of Artificial General Intelligence : Early experiments with GPT-4 . Preprint, arXiv:2303.12712

Show all 68 references
  1. [9]

    Chang and Benjamin K

    Tyler A. Chang and Benjamin K. Bergen. 2023. https://doi.org/10.48550/ARXIV.2303.11504 Language Model Behavior : A Comprehensive Survey

  2. [10]

    Junzhe Chen, Xuming Hu, Shuodi Liu, Shiyu Huang, Wei-Wei Tu, Zhaofeng He, and Lijie Wen. 2024. https://arxiv.org/abs/2402.16499 Llmarena: Assessing capabilities of large language models in dynamic multi-agent environments . Preprint, arXiv:2402.16499

  3. [11]

    Fran c ois Chollet. 2019. On the measure of intelligence. arXiv preprint arXiv:1911.01547

  4. [12]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://doi.org/10.48550/arXiv.2110.14168 Training Verifiers to Solve Math Word Pr...

  5. [13]

    Frans De Waal. 2016. Are we smart enough to know how smart animals are? W. W. Norton & Company

  6. [14]

    Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan. 2024. https://doi.org/10.18653/v1/2024.naacl-long.482 Investigating data contamination in modern benchmarks for large language models . In Proceedings of the 2024 Conference of the North American Chapter ...

  7. [15]

    D. C. Dennett. 2018. From Bacteria to Bach and Back: The Evolution of Minds , fist published as a norton paperback edition. W. W. Norton & Company

  8. [16]

    Jesse Dodge, Maarten Sap, Ana Marasovi \'c , William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.98 Documenting large webtext corpora: A case study on the colossal clean crawled corpus . In Pro...

  9. [17]

    Yihong Dong, Xue Jiang, Huanyu Liu, Zhi Jin, Bin Gu, Mengfei Yang, and Ge Li. 2024. https://doi.org/10.18653/v1/2024.findings-acl.716 Generalization or memorization: Data contamination and trustworthy evaluation for large language models . In Findings of the Association for Co...

  10. [18]

    ESS ERIC . 2024. https://doi.org/10.21338/ESS11-2023 European Social Survey ( ESS ), Round 11 - 2023

  11. [19]

    Matt Gardner, William Merrill, Jesse Dodge, Matthew Peters, Alexis Ross, Sameer Singh, and Noah A. Smith. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.135 Competency problems: On finding and removing artifacts in language data . In Proceedings of the 2021 Conference on Em...

  12. [20]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daum \'e III, and Kate Crawford. 2020. https://arxiv.org/abs/1803.09010 Datasheets for Datasets . arXiv:1803.09010 [cs]

  13. [21]

    Yoav Goldberg. 2024. https://gist.github.com/yoavg/f952b7a6cafd2024f44c8bc444a64315 ACL is not an AI Conference (?)

  14. [22]

    Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018. https://doi.org/10.18653/v1/N18-2017 Annotation artifacts in natural language inference data . In Proceedings of the 2018 Conference of the North A merican Chapter of the As...

  15. [23]

    https://arxiv.org/abs/2009.03300 Measuring Massive Multitask Language Understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. https://arxiv.org/abs/2009.03300 Measuring Massive Multitask Language Understanding . Preprint, arXiv:2009.03300

  16. [24]

    Alon Jacovi, Avi Caciularu, Omer Goldman, and Yoav Goldberg. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.308 Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks . In Proceedings of the 2023 Conference on...

  17. [25]

    Terne Thorn Jakobsen and Anna Rogers. 2022. https://doi.org/10.18653/v1/2022.naacl-main.354 What Factors Should Paper-Reviewer Assignments Rely On ? Community Perspectives on Issues and Ideals in Conference Peer-Review . pages 4810--4823

  18. [26]

    Joohee Kim and Il Im. 2023. https://doi.org/10.1016/j.chb.2022.107512 Anthropomorphic response: Understanding interactions between humans and artificial intelligence agents . 139:107512

  19. [27]

    Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. 2017. Building machines that learn and think like people. Behavioral and brain sciences, 40:e253

  20. [28]

    Shane Legg and Marcus Hutter. 2006. https://doi.org/10.48550/arXiv.cs/0605024 A Formal Measure of Machine Intelligence . Preprint, arXiv:cs/0605024

  21. [29]

    Shane Legg, Marcus Hutter, and 1 others. 2007. A collection of definitions of intelligence. Frontiers in Artificial Intelligence and applications, 157:17

  22. [30]

    Jiaang Li, Yova Kementchedjhieva, Constanza Fierro, and Anders S gaard. 2024. https://doi.org/10.1162/tacl_a_00698 Do vision and language models share concepts? a vector space alignment study . Transactions of the Association for Computational Linguistics, 12:1232--1249

  23. [31]

    Gary F Marcus. 2003. The algebraic mind: Integrating connectionism and cognitive science. MIT press

  24. [32]

    Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D

    R. Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D. Hardy, and Thomas L. Griffiths. 2024. https://doi.org/10.1073/pnas.2322420121 Embers of autoregression show how large language models are shaped by the problem they are trained to solve . Proceedings of the National Academy ...

  25. [33]

    Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. https://doi.org/10.18653/v1/P19-1334 Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3...

  26. [34]

    Drew McDermott. 1976. https://doi.org/10.1145/1045339.1045340 Artificial intelligence meets natural stupidity . (57):4--9

  27. [35]

    William Merrill, Yoav Goldberg, Roy Schwartz, and Noah A. Smith. 2021. https://doi.org/10.1162/tacl_a_00412 Provable limitations of acquiring meaning from ungrounded form: What will future language models understand? Transactions of the Association for Computational Linguistic...

  28. [36]

    Julian Michael, Ari Holtzman, Alicia Parrish, Aaron Mueller, Alex Wang, Angelica Chen, Divyam Madaan, Nikita Nangia, Richard Yuanzhe Pang, Jason Phang, and Samuel R. Bowman. 2023. https://doi.org/10.18653/v1/2023.acl-long.903 What Do NLP Researchers Believe ? Results of the NL...

  29. [37]

    Margaret Mitchell, Alexandra Sasha Luccioni, Nathan Lambert, Marissa Gerchick, Angelina McMillan-Major, Ezinwanne Ozoani, Nazneen Rajani, Tristan Thrush, Yacine Jernite, and Douwe Kiela. 2023. https://arxiv.org/abs/2212.05129 Measuring data . Preprint, arXiv:2212.05129

  30. [38]

    Melanie Mitchell. 2021. https://doi.org/10.1145/3449639.3465421 Why AI is harder than we think . In Proceedings of the Genetic and Evolutionary Computation Conference , GECCO '21, page 3. Association for Computing Machinery

  31. [39]

    Krakauer

    Melanie Mitchell and David C. Krakauer. 2023. https://doi.org/10.1073/pnas.2215907120 The debate over understanding in AI 's large language models . Proceedings of the National Academy of Sciences, 120(13):e2215907120

  32. [40]

    Mortensen

    David R. Mortensen. 2024. https://changelinglab.github.io/blog/2024/cl_in_acl/ Is ACL an AI (or NLP or CL ) Conference ?

  33. [41]

    Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2023. https://doi.org/10.48550/arXiv.2301.05217 Progress measures for grokking via mechanistic interpretability . Preprint, arXiv:2301.05217

  34. [42]

    Catherine Olsson, Nelson Elhage, and Neel Nanda. 2022. https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html In-context Learning and Induction Heads

  35. [43]

    OpenAI. 2022. https://openai.com/blog/chatgpt Introducing ChatGPT

  36. [44]

    Brian Pasquini, Jeffrey Kennedy, Monica Gottfried, Giancarlo Anderson, and Colleen McClain. 2025. https://www.pewresearch.org/internet/2025/04/03/how-the-us-public-and-ai-experts-view-artificial-intelligence/ How the us public and ai experts view artificial intelligence

  37. [45]

    R Pfeifer. 2006. How the body shapes the way we think: A new view of intelligence

  38. [46]

    Rolf Pfeifer and Christian Scheier. 2001. Understanding intelligence. MIT press

  39. [47]

    Phillips

    Andrew W. Phillips. 2017. https://doi.org/10.5811/westjem.2016.11.32000 Proper Applications for Surveys as a Study Methodology . 18(1):8--11

  40. [48]

    John Piaget. 1952. The origins of intelligence in children. International University

  41. [49]

    Bender, Alex Hanna, and Amandalynne Paullada

    Inioluwa Deborah Raji, Emily Denton, Emily M. Bender, Alex Hanna, and Amandalynne Paullada. 2021. https://openreview.net/forum?id=j6NxpQbREA1 AI and the Everything in the Whole Wide World Benchmark

  42. [50]

    Understand

    Sagnik Ray Choudhury, Anna Rogers, and Isabelle Augenstein. 2022. https://aclanthology.org/2022.coling-1.8 Machine Reading , Fast and Slow : When Do Models “ Understand ” Language ? In Proceedings of the 29th International Conference on Computational Linguistics , pages 78--93...

  43. [51]

    Colin Robson. 1999. Real World Research: A Resource for Social Scientists and Practitioner-Researchers, repr edition. Blackwell

  44. [52]

    Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020. https://doi.org/10.1162/tacl_a_00349 A primer in BERT ology: What we know about how BERT works . Transactions of the Association for Computational Linguistics, 8:842--866

  45. [53]

    Anna Rogers and Sasha Luccioni. 2024. https://openreview.net/forum?id=M2cwkGleRL Position: Key Claims in LLM Research Have a Long Tail of Footnotes . In Forty-First International Conference on Machine Learning

  46. [54]

    Francesca Rossi. 2025. https://aaai.org/wp-content/uploads/2025/03/AAAI-2025-PresPanel-Report-FINAL.pdf The Future of AI Research

  47. [55]

    Naomi Saphra and Sarah Wiegreffe. 2024. https://doi.org/10.18653/v1/2024.blackboxnlp-1.30 Mechanistic? In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 480--498, Miami, Florida, US. Association for Computational Linguistics

  48. [56]

    David Schlangen. 2021. https://doi.org/10.18653/v1/2021.acl-short.85 Targeting the benchmark: On methodology in current natural language processing research . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International ...

  49. [57]

    Linda Smith and Michael Gasser. 2005. The development of embodied cognition: Six lessons from babies. Artificial life, 11(1-2):13--29

  50. [58]

    Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, ...

  51. [59]

    Joshua B Tenenbaum, Charles Kemp, Thomas L Griffiths, and Noah D Goodman. 2011. How to grow a mind: Statistics, structure, and abstraction. science, 331(6022):1279--1285

  52. [60]

    Alan M. Turing. 1950. https://doi.org/10.1093/mind/LIX.236.433 I.— COMPUTING MACHINERY AND INTELLIGENCE . LIX(236):433--460

  53. [61]

    Marcel Van Gerven. 2017. Computational foundations of natural intelligence. Frontiers in computational neuroscience, 11:299674

  54. [62]

    David Vilares and Carlos Gómez-Rodríguez. 2019. https://doi.org/10.48550/arXiv.1906.04701 HEAD-QA : A Healthcare Dataset for Complex Reasoning . Preprint, arXiv:1906.04701

  55. [63]

    Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2020. https://doi.org/10.48550/arXiv.1905.00537 SuperGLUE : A Stickier Benchmark for General-Purpose Language Understanding Systems . Preprint, arXiv:1905.00537

  56. [64]

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. https://doi.org/10.18653/v1/W18-5446 GLUE : A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding . In Proceedings of the 2018 EMNLP Workshop BlackboxNLP : Ana...

  57. [65]

    Joseph Weizenbaum. 1966. https://doi.org/10.1145/365153.365168 ELIZA —a computer program for the study of natural language communication between man and machine . 9(1):36--45

  58. [66]

    Geraint A Wiggins. 2020. Creativity, information, and consciousness: the information dynamics of thinking. Physics of life reviews, 34:1--39

  59. [67]

    Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Aky \"u rek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, and Yoon Kim. 2024. https://doi.org/10.18653/v1/2024.naacl-long.102 Reasoning or reciting? exploring the capabilities and limitations of language models through counter...

  60. [68]

    Maxwell Zeff. 2024. https://techcrunch.com/2024/12/26/microsoft-and-openai-have-a-financial-definition-of-agi-report/ Microsoft and OpenAI have a financial definition of AGI : Report

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.