REVIEW 4 major objections 4 minor 1 cited by
Demystifying AI Agents: The Final Generation of Intelligence
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Paper calls AI agents the final generation of intelligence
desk verdict A readable AI-milestones survey whose central '5.9-month doubling' claim is unsupported and contradicted by the paper's own Table VII; more position paper than research. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the modern AI agent, defined as a Transformer-based language model equipped with chain-of-thought prompting, RLHF alignment, retrieval-augmented generation or live tools, and enough hardware to scale. The paper treats this stack as the mechanism that turns pretraining into goal-directed action, and it uses the stack's benchmark scores as the evidence that intelligence is doubling. The second piece of machinery is the doubling estimate itself: an extrapolation of benchmark gains on a roughly 5.9-month cycle, which converts the historical record into a forward-looking prediction of human-level AI within a decade.
What would settle it
Watch the same benchmarks over the next two years: if MMLU or GSM8K scores plateau, slow below the 5.9-month doubling rate, or fail to rise when new models are released, the central claim is contradicted; saturation near 100 percent alone is enough to falsify a continued exponential.
Extended reading notes
Core claim
The central discovery, on the paper's own terms, is that the separate advances of the last decade have converged into a single new kind of system: an agent that can reason step by step, retrieve current information, call external tools, and act on goals. The paper calls this the final generation of intelligence "as we currently conceive it," and backs the label with benchmark data: MMLU from roughly 25 percent to above 86 percent in under three years, GSM8K from 17.7 percent to 92 percent with chain-of-thought, and an estimated doubling of capability every 5.9 months. If the paper is right, current systems are not an intermediate stage but a plateau-or-singularity boundary, and the next questions are ethical and managerial rather than primarily architectural.
Load-bearing premise
The load-bearing premise is that score gains on benchmarks like MMLU, GSM8K, and HumanEval measure one underlying quantity called "intelligence," and that the recent doubling rate will continue into the future unchanged.
Editorial extensions
If this is right
- If the six-month doubling holds, systems scoring at or below human average on broad tests today should match or exceed human performance on many cognitive benchmarks within roughly ten years.
- The "final generation" framing implies that the main remaining work is integration, safety, and governance of existing agent capabilities rather than invention of new core architectures.
- Planning that assumes linear progress will underestimate capability growth; the paper's compounding rate implies a capability roughly four times larger each year.
- The paper's own caveats tie the benefits to governance: whether agents cure diseases and personalize education or deepen inequality depends on deployment choices, not on the technology alone.
Reading between the lines
- The paper's doubling rate blends compute growth, algorithmic efficiency, and benchmark design; a capability metric separated from hardware scaling could double more slowly or faster, so the six-month figure should be read as a composite rather than a clock.
- The "final generation" label implies a plateau that the same tool-using agents could undermine: agents that design experiments or write code could accelerate progress beyond the historical doubling, making the label self-limiting.
- A testable extension the paper does not run is to hold one agent fixed on a private, unpublished task set across several model releases, separating genuine capability growth from benchmark saturation.
- The societal concerns listed in the paper—job displacement, bias, energy use, and accountability—do not actually depend on the doubling claim; they would be urgent even if progress flatlined tomorrow.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a position/survey paper that argues that recent AI systems, particularly large language model agents with tool integration, represent the 'final generation of intelligence' before a possible singularity or plateau. Its central evidence is an acceleration claim: benchmark performance such as MMLU allegedly rose from 25% to 86% in under three years, and 'intelligence metrics' are said to double approximately every 5.9 months, leading to a projection that AI will surpass human intelligence within a decade. The paper surveys prompting techniques, training methods, hardware evolution, architecture, and tool use as converging factors behind this acceleration.
Significance. If substantiated, the acceleration thesis would be a claim of major societal and scientific importance. The paper does provide a readable chronology of AI agent technologies and correctly identifies several real trends, such as the growth of context windows and the role of RLHF. However, the central quantitative argument is not supported: the 5.9-month doubling period is contradicted by the paper's own Table VII, the cited source for it reports compute-doubling rather than benchmark-doubling rates, and many specific empirical figures are uncited or misattributed. As a result, the paper's headline claims, including the prediction of human-surpassing AI within a decade, are not credible on the evidence presented.
major comments (4)
- [Section IX, Table VII] The text claims that AI performance is 'doubling approximately every 5.9 months' and implies a fourfold annual increase, but Table VII shows benchmark gains of only 2–3x for MMLU, 5.2x for GSM8K, and 2.3x for HumanEval over the 2020–2023 period. A true 5.9-month doubling across 36 months would imply a factor of roughly 2^(36/5.9) ≈ 68x, which is more than an order of magnitude larger than any row in the table. Moreover, MMLU and GSM8K are percentage-bounded benchmarks, so a literal doubling rate cannot persist beyond a few doublings; the paper never defines the unbounded 'intelligence metric' that is supposed to double. The projection that AI will surpass human intelligence within a decade is therefore not derivable from the paper's own evidence.
- [Section IX, reference [55]] The '5.9-month doubling' statement is attributed to Epoch AI via reference [55], which is the paper 'Compute Trends Across Three Eras of Machine Learning' by Sevilla et al. That paper analyzes growth in training compute, not doubling of benchmark performance or of any 'intelligence' measure. No other source or derivation is provided for a benchmark-doubling rate. Since the entire forward-looking conclusion rests on this number, the central extrapolation is unsupported by the cited literature.
- [Sections II-B, III, VII-A; Tables I-III] Many quantitative claims are presented without sufficient sourcing, and several citations do not support the asserted numbers. For example: Section II-B states that Self-Consistency reduced hallucination rates from 21% to 15.8% and cites [13], but Wang et al. 2022 does not report such hallucination-rate outcomes; the '63% to 82%' arithmetic-reasoning improvement is attributed to [14], a paper about training verifiers, not structured prompts; Table I reports 'Legal reasoning ∼40%/∼80%' and other accuracies with no study name or confidence interval; Table II reports zero-shot/few-shot gains for BERT, LLaMA, and CodeGen without sources containing those exact numbers; and Table VII gives no citations at all for its benchmark values. Because the paper's conclusion claims 'the evidence is stark,' the unreliability of this evidence is a load-bearing problem.
- [Section X] The core concepts of the thesis—'intelligence', 'final generation of intelligence', and 'surpassing human intelligence'—are never operationally defined. The paper does not state what measurements would confirm or falsify the claim that intelligence doubles every six months, nor what would count as having reached a 'final generation'. This makes the central claim unfalsifiable as presented, and it cannot be evaluated scientifically without such definitions.
minor comments (4)
- [Figures 1-4] The manuscript contains captions for Fig. 1, Fig. 2, Fig. 3, and Fig. 4, but the actual figure images are not included in the text, so the reader cannot inspect the 'Evolution of AI Capabilities' or the other visualized content.
- [Table IV] The TFLOPS values for the V100, A100, and H100 are listed without specifying the precision (FP16, FP32, sparse vs. dense) or the exact configuration, which makes the comparison ambiguous; for example, the H100's '1000+ TFLOPS' typically refers to sparse FP8 rather than dense throughput.
- [Section II-B, reference [13]] The text says 'Stanford researchers demonstrated' a self-consistency result, but reference [13] lists authors from multiple institutions including Google Research, not predominantly Stanford; the attribution is inaccurate.
- [Section VII-B, reference [55]] The bullet about error management states that tool integration failures can increase task error rates by 10% and cites [55], but reference [55] is the compute-trends paper and contains no such claim about tool integration error rates.
Circularity Check
No circularity: the paper performs no fitted derivation and its doubling claim is cited to external Epoch AI work, not constructed from its own inputs.
full rationale
The paper does not derive its central quantitative claim from its own fitted parameters or definitions. The claim that AI benchmark performance is 'effectively doubling approximately every 5.9 months' is explicitly attributed to Epoch AI via reference [55], and Table VII is presented as illustrative evidence rather than as the source from which the doubling rate is fit. There are no equations linking inputs to outputs, no parameters fitted to a subset of data and then renamed as a prediction, and no self-citations or author-imported uniqueness theorems. The projection in Section IX is conditional ('If capabilities continue to double every six months, we may see AI systems surpassing human intelligence within the next decade'), so the conclusion is not forced by construction. The mismatch between Table VII's 2–5x gains over roughly three years and the 5.9-month doubling rate, as well as the undefined meaning of 'intelligence' being doubled, are substantive correctness and support concerns, but they are not circularity: the paper's acceleration claim rests on an external citation rather than on a derivation that reduces to its own inputs. Accordingly, the circularity burden is minimal and the honest finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- capability doubling period =
~5.9 months
assumptions (3)
- domain assumption Benchmark scores measure general intelligence.
- domain assumption The historical doubling rate will continue into the future.
- domain assumption The cited vendor systems are representative of a general class of agents.
Cite this review
Pith. "Pith review of Demystifying AI Agents: The Final Generation of Intelligence." pith.science (2026). https://pith.science/paper/WZBKIN6R
@misc{pith2026250509932,
author = {Pith},
title = {Pith review of: Demystifying AI Agents: The Final Generation of Intelligence},
year = {2026},
howpublished = {\url{https://pith.science/paper/WZBKIN6R}},
note = {Machine review of arXiv:2505.09932}
}
read the original abstract
The trajectory of artificial intelligence (AI) has been one of relentless acceleration, evolving from rudimentary rule-based systems to sophisticated, autonomous agents capable of complex reasoning and interaction. This whitepaper chronicles this remarkable journey, charting the key technological milestones--advancements in prompting, training methodologies, hardware capabilities, and architectural innovations--that have converged to create the AI agents of today. We argue that these agents, exemplified by systems like OpenAI's ChatGPT with plugins and xAI's Grok, represent a culminating phase in AI development, potentially constituting the "final generation" of intelligence as we currently conceive it. We explore the capabilities and underlying technologies of these agents, grounded in practical examples, while also examining the profound societal implications and the unprecedented pace of progress that suggests intelligence is now doubling approximately every six months. The paper concludes by underscoring the critical need for wisdom and foresight in navigating the opportunities and challenges presented by this powerful new era of intelligence.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 1 Pith paper
-
Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
This position paper classifies testing methods for LLM applications into three layers and proposes AICL, a structured protocol for testable agent communication; neither the framework nor the protocol is empirically validated.
Reference graph
Works this paper leans on
-
[55]
Compute Trends Across Three Eras of Machine Learning,
J. Sevilla et al., “Compute Trends Across Three Eras of Machine Learning,” Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society (AIES ’22) , pp. 749–762, Aug. 2022. DOI: 10.1145/3514094.3534166. [Online]. Available: https://epochai.org/
arXiv 2022
- [50]
-
[22]
RankCSE: Unsupervised Sentence Representations via Learning to Rank,
K. Krishna et al., “RankCSE: Unsupervised Sentence Representations via Learning to Rank,” Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023) , pp. 8876–8890, Jul. 2023. DOI: 10.18653/v1/2023.acl-long.492
-
[13]
Self-Consistency Improves Chain of Thought Rea- soning in Language Models,
X. Wang et al., “Self-Consistency Improves Chain of Thought Rea- soning in Language Models,” arXiv:2203.11171, Mar. 2022. [Online]. Available: https://arxiv.org/abs/2203.11171
arXiv 2022
-
[14]
Training Verifiers to Solve Math Word Problems,
K. Cobbe et al., “Training Verifiers to Solve Math Word Problems,” arXiv:2110.14168, Oct. 2021. [Online]. Available: https://arxiv.org/abs/ 2110.14168
arXiv 2021
-
[1]
Empirical explorations with the logic theory machine: A case study in heuristics,
A. Newell, J. C. Shaw, and H. A. Simon, “Empirical explorations with the logic theory machine: A case study in heuristics,” Proceedings of the Western Joint Computer Conference , pp. 218–239, 1957. DOI: 10.1145/1455292.1455311
-
[2]
ELIZA–a computer program for the study of nat- ural language communication between man and machine,
J. Weizenbaum, “ELIZA–a computer program for the study of nat- ural language communication between man and machine,” Commu- nications of the ACM , vol. 9, no. 1, pp. 36–45, Jan. 1966. DOI: 10.1145/365153.365168
arXiv 1966
-
[3]
OpenAI, “GPT-4 Technical Report,” arXiv:2303.08774, Mar. 2023. [Online]. Available: https://arxiv.org/abs/2303.08774
arXiv 2023
Show all 58 references
-
[4]
Cramming more components onto integrated circuits,
G. E. Moore, “Cramming more components onto integrated circuits,” Electronics, vol. 38, no. 8, pp. 114–117, Apr. 1965
1965
-
[5]
Attention is All You Need,
A. Vaswani et al., “Attention is All You Need,” Advances in Neu- ral Information Processing Systems 30 (NIPS 2017) , pp. 5998– 6008, 2017. [Online]. Available: https://papers.nips.cc/paper/2017/hash/ 3f5ee243547dee91fbd053c1c4a845aa-Abstract.html
2017
-
[6]
Training language models to follow instruc- tions with human feedback,
L. Ouyang et al., “Training language models to follow instruc- tions with human feedback,” Advances in Neural Information Pro- cessing Systems 35 (NeurIPS 2022) , pp. 27730–27744, 2022. [On- line]. Available: https://proceedings.neurips.cc/paper files/paper/2022/ hash/b1efde53...
2022
-
[7]
Deep Blue,
M. Campbell, A. J. Hoane Jr., and F.-H. Hsu, “Deep Blue,” Artificial In- telligence, vol. 134, no. 1–2, pp. 57–83, Jan. 2002. DOI: 10.1016/S0004- 3702(01)00129-1
2002 doi
-
[8]
Language Models are Few-Shot Learners,
T. Brown et al., “Language Models are Few-Shot Learners,” Advances in Neural Information Processing Systems 33 (NeurIPS 2020) , pp. 1877– 1901, 2020. [Online]. Available: https://papers.nips.cc/paper/2020/hash/ 1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html
2020
-
[9]
Prompting Large Language Models for Legal Analysis Tasks,
L. Reynolds and K. D. Ashley, “Prompting Large Language Models for Legal Analysis Tasks,”Proceedings of the 19th International Conference on Artificial Intelligence and Law (ICAIL 2023), pp. 278–287, Jun. 2023. DOI: 10.1145/3594536.3595132
2023
-
[10]
Evaluating Large Language Models for Medical Diagno- sis,
X. Liu et al., “Evaluating Large Language Models for Medical Diagno- sis,” Stanford HAI & Stanford Medicine Research Update , Feb. 2022
2022
-
[11]
Learning Transferable Visual Models From Natural Language Supervision,
A. Radford et al., “Learning Transferable Visual Models From Natural Language Supervision,” Proceedings of the 38th International Confer- ence on Machine Learning (ICML 2021) , PMLR vol. 139, pp. 8748– 8763, Jul. 2021. [Online]. Available: https://proceedings.mlr.press/v139/ r...
2021
-
[12]
GitHub Copilot: Flying Ahead with Speed and Accu- racy,
A. Ziegler, “GitHub Copilot: Flying Ahead with Speed and Accu- racy,” GitHub Blog, Jun. 2022. [Online]. Available: https://github.blog/ 2022-06-21-github-copilot-is-generally-available-to-all-developers/
2022
-
[15]
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” Pro- ceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologie...
2019 doi
-
[16]
Improving language models by retrieving from trillions of tokens,
S. Borgeaud et al., “Improving language models by retrieving from trillions of tokens,” Proceedings of the 39th International Confer- ence on Machine Learning (ICML 2022) , PMLR vol. 162, pp. 2206– 2240, Jul. 2022. [Online]. Available: https://proceedings.mlr.press/v162/ borge...
2022
-
[17]
Llama 2: Open Foundation and Fine-Tuned Chat Models,
H. Touvron et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models,” arXiv:2307.09288, Jul. 2023. [Online]. Available: https://arxiv. org/abs/2307.09288
2023 arXiv
-
[18]
CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis,
E. Nijkamp et al., “CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis,” arXiv:2203.13474, Mar. 2022. [Online]. Available: https://arxiv.org/abs/2203.13474
2022 arXiv
-
[19]
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,
J. Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” Advances in Neural Information Pro- cessing Systems 35 (NeurIPS 2022) , pp. 24824–24837, 2022. [On- line]. Available: https://proceedings.neurips.cc/paper files/paper/2022/ hash/9d560961352...
2022
-
[20]
Large Language Models are Zero-Shot Reasoners,
T. Kojima et al., “Large Language Models are Zero-Shot Reasoners,” Advances in Neural Information Processing Systems 35 (NeurIPS 2022) , pp. 22199–22213, 2022. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2022/hash/ 8bb0d291acd4acf06ef112099c16f326-Abs...
2022
-
[21]
Building Watson: An Overview of the DeepQA Project,
D. Ferrucci et al., “Building Watson: An Overview of the DeepQA Project,” AI Magazine , vol. 31, no. 3, pp. 59–79, Fall 2010. DOI: 10.1609/aimag.v31i3.2303
2010 doi
-
[23]
TruthfulQA: Measuring How Models Mimic Human Falsehoods,
S. Lin, J. Hilton, and O. Evans, “TruthfulQA: Measuring How Models Mimic Human Falsehoods,” Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL 2022) , pp. 3214– 3252, May 2022. DOI: 10.18653/v1/2022.acl-long.229
2022 doi
-
[24]
Deep reinforcement learning from human pref- erences,
P. F. Christiano et al., “Deep reinforcement learning from human pref- erences,” Advances in Neural Information Processing Systems 30 (NIPS 2017), pp. 4299–4307, 2017. [Online]. Available: https://papers.nips.cc/ paper/2017/hash/d5e2c0adad503c91f91df240d0cd4e49-Abstract.html
2017
-
[25]
ChatGPT Sets Record for Fastest-Growing User Base,
K. Hu, “ChatGPT Sets Record for Fastest-Growing User Base,” Reuters, Feb. 2, 2023. [Online]. Available: https://www.reuters.com/technology/ chatgpt-sets-record-fastest-growing-user-base-analyst-note-2023-02-01/
2023
-
[26]
Constitutional AI: Harmlessness from AI Feedback,
Y . Bai et al., “Constitutional AI: Harmlessness from AI Feedback,” arXiv:2212.08073, Dec. 2022. [Online]. Available: https://arxiv.org/abs/ 2212.08073
2022 arXiv
-
[27]
BioBERT: a pre-trained biomedical language representa- tion model for biomedical text mining,
J. Lee et al., “BioBERT: a pre-trained biomedical language representa- tion model for biomedical text mining,” Bioinformatics, vol. 36, no. 4, pp. 1234–1240, Feb. 2020. DOI: 10.1093/bioinformatics/btz682
2020 doi
-
[28]
Mastering the game of Go with deep neural networks and tree search,
D. Silver et al., “Mastering the game of Go with deep neural networks and tree search,” Nature, vol. 529, no. 7587, pp. 484–489, Jan. 2016. DOI: 10.1038/nature16961
2016 doi
-
[29]
Retrieval-Augmented Generation for Knowledge- Intensive NLP Tasks,
P. Lewis et al., “Retrieval-Augmented Generation for Knowledge- Intensive NLP Tasks,” Advances in Neural Information Processing Systems 33 (NeurIPS 2020) , pp. 9459–9474,
2020
-
[30]
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,
V . Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” arXiv:1910.01108, Oct. 2019. [Online]. Available: https://arxiv.org/abs/1910.01108
1910 arXiv
-
[31]
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices,
Z. Sun et al., “MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL 2020) , pp. 2158– 2170, Jul. 2020. DOI: 10.18653/v1/2020.acl-main.194
2020 doi
-
[32]
NVIDIA A100 Tensor Core GPU Architecture,
NVIDIA, “NVIDIA A100 Tensor Core GPU Architecture,” NVIDIA Technical Brief , May 2020. [Online]. Available: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/ nvidia-ampere-architecture-whitepaper.pdf
2020
-
[33]
NVIDIA H100 Tensor Core GPU Architecture,
NVIDIA, “NVIDIA H100 Tensor Core GPU Architecture,” NVIDIA Technical Brief, Mar. 2022. [Online]. Available: https://resources.nvidia. com/en-us-tensor-core/nvidia-hopper-architecture-whitepaper
2022
-
[34]
Energy and Policy Consid- erations for Deep Learning in NLP,
E. Strubell, A. Ganesh, and A. McCallum, “Energy and Policy Consid- erations for Deep Learning in NLP,” Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019) , pp. 3645–3650, Jul. 2019. DOI: 10.18653/v1/P19-1355
2019 doi
-
[35]
PaLM: Scaling Language Modeling with Path- ways,
A. Chowdhery et al., “PaLM: Scaling Language Modeling with Path- ways,” Journal of Machine Learning Research (JMLR), vol. 24, no. 246, pp. 1-113, 2023. [Online]. Available: http://jmlr.org/papers/v24/22-1144. html
2023
-
[36]
TPU v4: An Optically Reconfigurable Super- computer for Machine Learning with Hardware Support for Spar- sity,
N. P. Jouppi et al., “TPU v4: An Optically Reconfigurable Super- computer for Machine Learning with Hardware Support for Spar- sity,” Proceedings of the 49th Annual International Symposium on Computer Architecture (ISCA ’22) , pp. 174–188, Jun. 2022. DOI: 10.1145/3470496.3527403
2022
-
[37]
Scaling Deep Learning Workloads Beyond the Limits of Today’s Hardware,
Cerebras Systems, “Scaling Deep Learning Workloads Beyond the Limits of Today’s Hardware,” Cerebras White Paper, Aug. 2021. [On- line]. Available: https://www.cerebras.net/wp-content/uploads/2021/08/ Scaling Deep Learning Workloads White Paper.pdf
2021
-
[38]
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,
M. Tan and Q. V . Le, “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” Proceedings of the 36th Interna- tional Conference on Machine Learning (ICML 2019) , PMLR vol. 97, pp. 6105–6114, Jun. 2019. [Online]. Available: https://proceedings.mlr. press/v9...
2019
-
[39]
Tesla AI Day 2022,
Tesla, “Tesla AI Day 2022,” Event Presentation, Sep. 30, 2022. [Online]. Available: https://www.tesla.com/AI
2022
-
[40]
A million spiking-neuron integrated circuit with a scalable communication network and interface,
P. A. Merolla et al., “A million spiking-neuron integrated circuit with a scalable communication network and interface,” Science, vol. 345, no. 6197, pp. 668–673, Aug. 2014. DOI: 10.1126/science.1254642
2014 doi
-
[41]
Apple unveils iPhone 14 Pro and iPhone 14 Pro Max,
Apple, “Apple unveils iPhone 14 Pro and iPhone 14 Pro Max,” Apple Newsroom , Sep. 7, 2022. [Online]. Available: https://www.apple.com/newsroom/2022/09/ apple-unveils-iphone-14-pro-and-iphone-14-pro-max/
2022
-
[42]
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity,
W. Fedus, B. Zoph, and N. Shazeer, “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity,” Journal of Machine Learning Research (JMLR), vol. 23, no. 120, pp. 1–39, 2022. [Online]. Available: http://jmlr.org/papers/v23/21-0998.html
2022
-
[43]
Introducing Claude 2.1,
Anthropic, “Introducing Claude 2.1,” Anthropic Blog , Nov. 21, 2023. [Online]. Available: https://www.anthropic.com/index/claude-2-1
2023
-
[44]
Introducing Gemini: our largest and most capable AI model,
D. Hassabis and E. Sheleff, “Introducing Gemini: our largest and most capable AI model,” Google DeepMind Blog , Dec. 6, 2023. [Online]. Available: https://deepmind.google/technologies/gemini/
2023
-
[45]
Gemini 1.5: Our next-generation model, break- through long context,
Google DeepMind, “Gemini 1.5: Our next-generation model, break- through long context,” Google DeepMind Blog, Feb. 15, 2024. [Online]. Available: https://deepmind.google/technologies/gemini/gemini-1-5/
2024
-
[46]
Lost in the Middle: How Language Models Use Long Contexts,
N. F. Liu, K. Lin, J. Hewitt, A. Singh, P. Liang, and P. S. H. E. Chen, “Lost in the Middle: How Language Models Use Long Contexts,” arXiv:2307.03172, Jul. 2023. [Online]. Available: https://arxiv.org/abs/ 2307.03172
2023 arXiv
-
[47]
WebGPT: Browser-assisted question-answering with human feedback,
R. Nakano et al., “WebGPT: Browser-assisted question-answering with human feedback,” arXiv:2112.09332, Dec. 2021. [Online]. Available: https://arxiv.org/abs/2112.09332
2021 arXiv
-
[48]
Auto-GPT: An Autonomous GPT-4 Experiment,
T. Grantham, “Auto-GPT: An Autonomous GPT-4 Experiment,” GitHub Repository , Mar. 2023. [Online]. Available: https://github.com/ Significant-Gravitas/Auto-GPT
2023
-
[49]
Experimental evidence on the productivity effects of generative artificial intelligence,
S. Noy and W. J. Zhang, “Experimental evidence on the productivity effects of generative artificial intelligence,” Science, vol. 381, no. 6654, pp. 187-192, Jul. 2023. DOI: 10.1126/science.adh1806
2023 doi
-
[51]
Introducing Adept: Useful General Intelligence That Enables Humans And Computers To Work Together Creatively,
Adept AI, “Introducing Adept: Useful General Intelligence That Enables Humans And Computers To Work Together Creatively,” Adept Blog ,
-
[52]
Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification,
J. Buolamwini and T. Gebru, “Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification,” Proceedings of Ma- chine Learning Research, vol. 81, pp. 77–91, 2018. [Online]. Available: http://proceedings.mlr.press/v81/buolamwini18a.html
2018
-
[53]
The future of employment: How susceptible are jobs to computerisation?
C. B. Frey and M. A. Osborne, “The future of employment: How susceptible are jobs to computerisation?” Technological Forecast- ing and Social Change , vol. 114, pp. 254–280, Jan. 2017. DOI: 10.1016/j.techfore.2016.08.019
2017 doi
-
[54]
Killer Robots,
R. Sparrow, “Killer Robots,” Journal of Applied Philosophy, vol. 24, no. 1, pp. 62–77, Feb. 2007. DOI: 10.1111/j.1468-5930.2007.00346.x
2007
-
[56]
Neural Architecture Search with Reinforcement Learning,
B. Zoph and Q. V . Le, “Neural Architecture Search with Reinforcement Learning,” Proceedings of the 5th International Conference on Learning Representations (ICLR 2017) , Apr. 2017. [Online]. Available: https:// arxiv.org/abs/1611.01578
2017 arXiv
-
[2020]
Available: https://papers.nips.cc/paper/2020/hash/ 6b493230205f780e1bc26945df7481e5-Abstract.html
[Online]. Available: https://papers.nips.cc/paper/2020/hash/ 6b493230205f780e1bc26945df7481e5-Abstract.html
2020
-
[2022]
Available: https://www.adept.ai/blog/introducing-adept
[Online]. Available: https://www.adept.ai/blog/introducing-adept
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.