REVIEW 4 major objections 2 minor 169 references
This paper introduces MAAC, a nine-dimension framework for evaluating text-based AI systems by the cognitive processes behind their outputs rather than by outcome benchmarks alone.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 00:42 UTC pith:F4ZRPMWD
load-bearing objection MAAC is a plausible taxonomy for process-oriented AI evaluation, but the abstract asserts operational status without showing the operational definitions — worth a look if the full text delivers. the 4 major comments →
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that process-level cognitive assessment of text-based AI is both theoretically grounded and operationally feasible through nine defined dimensions. Each dimension is tied to a specific cognitive theory, and the paper provides a coverage matrix, gap analysis, and a priori interdependency predictions to argue that the dimensions are broad, non-redundant, and testable. The framework is offered as a complement to outcome-based benchmarks, enabling evaluation of how AI systems reason, manage memory, handle complexity, and avoid false information.
What carries the argument
The central object is the MAAC framework itself: a set of nine named dimensions—Cognitive Load, Tool Execution, Content Quality, Memory Integration, Complexity Handling, Hallucination Control, Knowledge Transfer, Processing Efficiency, and Process-Outcome Alignment. Each dimension maps to a cognitive science theory (e.g., Baddeley's working memory model, Sweller's cognitive load theory), and the framework uses that mapping to define what to measure and how to interpret text outputs as evidence of underlying processes.
Load-bearing premise
The framework assumes that human cognitive constructs like working memory and cognitive load transfer meaningfully to text-based AI systems and that these internal processes can be reliably inferred from observable text outputs.
What would settle it
Run a controlled experiment where two AI systems produce identical text on a task designed to impose high working-memory demands, but one system is given explicit memory aids while the other is not; if MAAC's Memory Integration scores do not differ between the two systems, the framework's claim to measure underlying processes would be falsified.
If this is right
- If the framework holds, AI evaluation can move beyond pass/fail benchmarks to structured diagnostic profiles that pinpoint specific cognitive weaknesses.
- The a priori interdependency predictions give researchers concrete hypotheses to test empirically, linking dimensions like Cognitive Load and Memory Integration.
- MAAC could serve as a foundation for new evaluation tools that complement existing outcome-based tests by explaining the 'why' behind performance.
- The gap analysis suggests which current evaluation practices miss process-level information, pointing to where new assessment methods are needed.
- Process-Outcome Alignment as a dimension could help identify cases where a system gets the right answer for the wrong reasons.
Where Pith is reading between the lines
- A natural extension would be to apply MAAC to multimodal AI, though the current dimensions are text-specific and would need re-mapping to non-textual outputs.
- MAAC could also be used as an interpretability lens: a system's dimension scores might reveal where internal representations fail, offering a bridge between evaluation and model debugging.
- If the framework's predictions hold, it might inform AI design—for instance, architectural changes that reduce measured Cognitive Load could be prioritized even when outcomes are unchanged.
- A testable extension is to compare MAAC dimension scores against human cognitive measures on the same tasks; strong correlation would strengthen the theory, while divergence would indicate limits of the analogy.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MAAC, a Multi-Dimensional Assessment framework for text-based AI systems, proposing nine cognitively motivated dimensions (Cognitive Load, Tool Execution, Content Quality, Memory Integration, Complexity Handling, Hallucination Control, Knowledge Transfer, Processing Efficiency, Process-Outcome Alignment). It claims these dimensions are grounded in established cognitive science theories and that five theoretical analyses—dimension-to-theory mapping, a coverage matrix, a gap analysis, a worked diagnostic illustration, and a priori interdependency predictions—provide initial support for coherence and empirical testability. The stated goal is to complement outcome-based benchmarks with a process-level, multi-dimensional evaluation. This review is based only on the abstract, as the full text was not supplied.
Significance. If the framework succeeds, it would be a meaningful contribution to AI evaluation by shifting attention from task accuracy to underlying cognitive processes, with practical implications for diagnosing strengths and failure modes of text-based AI systems. The explicit list of a priori interdependency predictions is a commendable feature, as it creates a route to falsification. The coverage matrix and gap analysis, if executed rigorously, could also help position MAAC relative to existing benchmarks. However, the significance cannot be fully assessed from the abstract: the operational validity of the nine dimensions, the non-redundancy argument, and the empirical testability are asserted rather than demonstrated. The abstract indicates a theoretical framework, but the stronger claim of being an 'operational framework' requires evidence of concrete measurement procedures and validation, which is not visible here.
major comments (4)
- [Abstract, first sentence of the contribution claim] The abstract calls MAAC an 'operational framework,' but it does not specify how each of the nine dimensions is measured from observable text outputs. Without explicit operational definitions—e.g., which text features correspond to Cognitive Load or Memory Integration—dimension scores are underdetermined. For instance, Cognitive Load could plausibly be scored from output length, token-level perplexity, revision patterns, or response time, and these features may reflect training data or decoding parameters rather than any latent cognitive process. The abstract provides no such mappings, so the operational claim is unsupported.
- [Abstract, 'Five theoretical analyses provide initial support'] The listed analyses are all internal coherence checks: dimension-to-theory mapping, coverage matrix, gap analysis, an illustration, and a priori predictions. These can establish that the framework is internally consistent but do not, by themselves, establish that the dimensions correspond to distinct, real cognitive processes in AI systems or that those processes are inferable from text. In particular, a coverage matrix can be self-affirming if the categories are chosen to match the framework. The abstract should state whether any external validation—e.g., comparison with independent human judgments, behavioral experiments, or controlled manipulation of cognitive variables—is possible or planned.
- [Abstract, 'Process-Outcome Alignment' dimension] There is a potential circularity in including 'Process-Outcome Alignment' both as one of the nine dimensions and as part of the evaluation framework. If this dimension is scored from the same observable features that the framework uses to define the other dimensions, then its measurement would be tautological. To avoid this, the abstract should indicate that the alignment dimension is assessed using independent outcome measures (e.g., external ground-truth correctness or human judgment) rather than features derived from the MAAC dimensions themselves. Without such a distinction, the framework's validation is at risk of being circular.
- [Abstract, nine-dimension taxonomy] The list of dimensions mixes constructs of different types: 'Tool Execution' refers to an external action, 'Content Quality' and 'Hallucination Control' are outcome-adjacent quality judgments, while 'Memory Integration' and 'Complexity Handling' are more process-like. The abstract asserts that a coverage matrix assesses 'non-redundancy,' but no evidence is shown. The reader cannot verify whether the nine dimensions are mutually exclusive or collectively exhaustive, nor whether some dimensions are actually composites of others. The abstract should explain the selection criterion and provide at least a summary of the non-redundancy analysis.
minor comments (2)
- [Abstract] The abstract does not mention any limitations or scope conditions. For example, does the framework apply only to autoregressive language models, or also to retrieval-augmented and tool-using systems? Clarifying the intended scope would help readers assess the claims.
- [Abstract, 'Marr's tri-level hypothesis'] The references to Marr, Baddeley, Sweller, and unified theories of cognition are listed but not connected to specific dimensions. In a full paper, this mapping is presumably shown, but the abstract would benefit from one example—e.g., how Baddeley's working memory model justifies the Memory Integration dimension.
Circularity Check
No circularity identifiable from the abstract; MAAC's claims are presented as a theoretical framework with a priori predictions, not as fitted results or self-referential derivations.
full rationale
This review is based on the abstract only. The abstract claims that MAAC defines nine dimensions grounded in established cognitive science theory and offers five theoretical analyses: dimension-to-theory mapping, coverage matrix, gap analysis, a worked diagnostic illustration, and a priori predictions. None of these, as described, asserts that a quantity is derived from data and then predicted back from the same quantity. The 'worked diagnostic illustration' could in principle be circular if it merely scores dimensions using the same features the dimensions are meant to explain, but the abstract does not provide enough detail to show such a reduction, and it is not described as a validation of the framework. The 'a priori interdependency predictions' are explicitly deferred to future empirical testing, which is the opposite of fitting a parameter and calling it a prediction. The cited theories are classic external works rather than self-citations by the authors. There is therefore no quotable equation, definition, or fitted-input relation that would support a circularity finding. A legitimate open concern is construct validity—whether text outputs can reliably indicate latent cognitive processes—but that is an empirical and theoretical risk, not an instance of circular derivation. Under the hard rule requiring a specific quoted reduction, no circular step can be flagged.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Human cognitive science theories (Marr, Baddeley, Sweller) can be meaningfully applied to text-based AI systems.
- domain assumption AI cognitive processes are separable from their outputs and observable through text behavior.
- ad hoc to paper The nine chosen dimensions form a non-redundant and sufficiently complete decomposition of AI cognition.
invented entities (1)
-
MAAC nine-dimension taxonomy
no independent evidence
read the original abstract
Evaluating artificial intelligence systems has historically relied on outcome-based benchmarks that measure task accuracy, robustness, or fairness. While indispensable, these benchmarks provide limited diagnostic insight into the underlying cognitive processes that generate performance-leaving critical questions unanswered about how AI systems reason, integrate memory, manage complexity, or avoid generating false information. This paper introduces the Multi-Dimensional Assessment for AI Cognition (MAAC), a theoretically grounded framework for shifting evaluation from what text-based AI systems produce to how they think. MAAC defines nine cognitively motivated dimensions: Cognitive Load, Tool Execution, Content Quality, Memory Integration, Complexity Handling, Hallucination Control, Knowledge Transfer, Processing Efficiency, and Process-Outcome Alignment. Each dimension is grounded in established cognitive science theory-drawing on Marr's tri-level hypothesis, Baddeley's working memory model, Sweller's cognitive load theory, and unified theories of cognition. Five theoretical analyses provide initial support for the framework's coherence and empirical testability: dimension-to-theory mapping; a coverage matrix assessing breadth and non-redundancy; a formal gap analysis relative to current evaluation practice; a worked diagnostic illustration; and a set of a priori interdependency predictions for future empirical testing. MAAC provides a theoretical and operational framework for principled process-level cognitive assessment of text-based AI systems, complementing existing outcome-based benchmarks with cognitively grounded, multi-dimensional evaluation.
Reference graph
Works this paper leans on
-
[1]
author Akhtar, M. , author Reuel, A. , author Soni, P. , et al., year 2026 . title When AI benchmarks plateau: A systematic study of benchmark saturation . journal arXiv preprint arXiv:2602.16763 :10.48550/arXiv.2602.16763
-
[2]
author Anderson, J.R. , author Bothell, D. , author Byrne, M.D. , author Douglass, S. , author Lebiere, C. , author Qin, Y. , year 2004 . title An integrated theory of the mind . journal Psychological Review volume 111 , pages 1036--1060 . :10.1037/0033-295X.111.4.1036
-
[3]
author Arksey, H. , author O'Malley, L. , year 2005 . title Scoping studies: Towards a methodological framework . journal International Journal of Social Research Methodology volume 8 , pages 19--32 . :10.1080/1364557032000119616
-
[4]
author Baddeley, A. , year 1992 . title Working memory . journal Science volume 255 , pages 556--559 . :10.1126/science.1736359
-
[5]
author Baddeley, A. , year 2000 . title The episodic buffer: A new component of working memory? journal Trends in Cognitive Sciences volume 4 , pages 417--423 . :10.1016/S1364-6613(00)01538-2
-
[6]
author Baddeley, A. , year 2003 . title Working memory: Looking back and looking forward . journal Nature Reviews Neuroscience volume 4 , pages 829--839 . :10.1038/nrn1201
doi:10.1038/nrn1201 2003
-
[7]
author Barnett, S.M. , author Ceci, S.J. , year 2002 . title When and where do we apply what we learn? A taxonomy for far transfer . journal Psychological Bulletin volume 128 , pages 612--637 . :10.1037/0033-2909.128.4.612
-
[8]
author Bender, E.M. , author Gebru, T. , author McMillan-Major, A. , author Shmitchell, S. , year 2021 . title On the dangers of stochastic parrots: Can language models be too big? , in: booktitle Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pp. pages 610--623 . :10.1145/3442188.3445922
arXiv 2021
-
[9]
author Bommasani, R. , author Hudson, D.A. , author Adeli, E. , et al., year 2021 . title On the opportunities and risks of foundation models . journal arXiv preprint arXiv:2108.07258 :10.48550/arXiv.2108.07258
-
[10]
, author Mensch, A
author Borgeaud, S. , author Mensch, A. , author Hoffmann, J. , et al., year 2022 . title Improving language models by retrieving from trillions of tokens , in: booktitle Proceedings of the 39th International Conference on Machine Learning , pp. pages 2206--2240 . https://proceedings.mlr.press/v162/borgeaud22a.html
2022
-
[11]
author Campbell, D.J. , year 1988 . title Task complexity: A review and analysis . journal Academy of Management Review volume 13 , pages 40--52 . :10.5465/amr.1988.4306775
arXiv 1988
-
[12]
author Clark, A. , author Chalmers, D. , year 1998 . title The extended mind . journal Analysis volume 58 , pages 7--19 . :10.1093/analys/58.1.7
-
[13]
author Clark, K. , author Khandelwal, U. , author Levy, O. , author Manning, C.D. , year 2019 . title What does BERT look at? an analysis of BERT 's attention , in: booktitle Proceedings of the 2019 ACL Workshop BlackboxNLP , pp. pages 276--286 . :10.18653/v1/W19-4828
-
[14]
author Cronbach, L.J. , author Meehl, P.E. , year 1955 . title Construct validity in psychological tests . journal Psychological Bulletin volume 52 , pages 281--302 . :10.1037/h0040957
doi:10.1037/h0040957 1955
-
[15]
author Crossley, S.A. , author Kyle, K. , author McNamara, D.S. , year 2016 . title The tool for the automatic analysis of text cohesion ( TAACO ) . journal Behavior Research Methods volume 48 , pages 1227--1237 . :10.3758/s13428-015-0651-7
-
[16]
, year 2017
author DeVellis, R.F. , year 2017 . title Scale Development: Theory and Applications . edition 4th ed., publisher SAGE Publications
2017
-
[17]
author Doshi-Velez, F. , author Kim, B. , year 2017 . title Towards a rigorous science of interpretable machine learning . journal arXiv preprint arXiv:1702.08608 :10.48550/arXiv.1702.08608
-
[18]
author Eriksson, M. , author Purificato, E. , author Noroozian, A. , author Vinagre, J. , author Chaslot, G. , author Gomez, E. , author Fernandez-Llorca, D. , year 2025 . title Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation . journal arXiv preprint arXiv:2502.06559 :10.48550/arXiv.2502.06559
-
[19]
author Ethayarajh, K. , author Jurafsky, D. , year 2020 . title Utility is in the eye of the user: A critique of NLP leaderboards , in: booktitle Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pp. pages 4846--4853 . :10.18653/v1/2020.emnlp-main.393
-
[20]
, year 2018
author Furr, R.M. , year 2018 . title Psychometrics: An Introduction . edition 3rd ed., publisher SAGE Publications
2018
-
[21]
, author Ghahramani, Z
author Gal, Y. , author Ghahramani, Z. , year 2016 . title Dropout as a Bayesian approximation: Representing model uncertainty in deep learning , in: booktitle Proceedings of the 33rd International Conference on Machine Learning , pp. pages 1050--1059 . https://proceedings.mlr.press/v48/gal16.html
2016
-
[22]
author Gentner, D. , year 1983 . title Structure-mapping: A theoretical framework for analogy . journal Cognitive Science volume 7 , pages 155--170 . :10.1207/s15516709cog0702_3
-
[23]
author Griffiths, T.L. , author Lieder, F. , author Goodman, N.D. , year 2015 . title Rational use of cognitive resources: Levels of analysis between the computational and the algorithmic . journal Topics in Cognitive Science volume 7 , pages 217--229 . :10.1111/tops.12142
-
[24]
, author Pleiss, G
author Guo, C. , author Pleiss, G. , author Sun, Y. , author Weinberger, K.Q. , year 2017 . title On calibration of modern neural networks , in: booktitle Proceedings of the 34th International Conference on Machine Learning , pp. pages 1321--1330 . https://proceedings.mlr.press/v70/guo17a.html
2017
-
[25]
author Halford, G.S. , author Baker, R. , author McCredden, J.E. , author Bain, J.D. , year 2005 . title How many variables can humans process? journal Psychological Science volume 16 , pages 70--76 . :10.1111/j.0956-7976.2005.00782.x
arXiv 2005
-
[26]
, author Hasan, R
author Halliday, M.A.K. , author Hasan, R. , year 1976 . title Cohesion in English . publisher Longman
1976
-
[27]
author Hendrycks, D. , author Burns, C. , author Basart, S. , et al., year 2021 . title Measuring massive multitask language understanding , in: booktitle Proceedings of the International Conference on Learning Representations . :10.48550/arXiv.2009.03300
-
[28]
author Hoffmann, J. , author Borgeaud, S. , author Mensch, A. , et al., year 2022 . title Training compute-optimal large language models . journal arXiv preprint arXiv:2203.15556 :10.48550/arXiv.2203.15556
-
[29]
, editor Morrison, R.G
editor Holyoak, K.J. , editor Morrison, R.G. (Eds.), year 2012 . title The Oxford Handbook of Thinking and Reasoning . publisher Oxford University Press
2012
-
[30]
author Huang, L. , author Yu, W. , author Ma, W. , et al., year 2023 . title A survey on hallucination in large language models . journal arXiv preprint arXiv:2311.05232 :10.48550/arXiv.2311.05232
-
[31]
, year 1995
author Hutchins, E. , year 1995 . title Cognition in the Wild . publisher MIT Press
1995
-
[32]
author Ji, Z. , author Lee, N. , author Frieske, R. , et al., year 2023 . title Survey of hallucination in natural language generation . journal ACM Computing Surveys volume 55 , pages 1--38 . :10.1145/3571730
doi:10.1145/3571730 2023
-
[33]
author Just, M.A. , author Carpenter, P.A. , year 1992 . title A capacity theory of comprehension: Individual differences in working memory . journal Psychological Review volume 99 , pages 122--149 . :10.1037/0033-295X.99.1.122
-
[34]
, year 2011
author Kahneman, D. , year 2011 . title Thinking, Fast and Slow . publisher Farrar, Straus and Giroux
2011
-
[35]
author Kaplan, J. , author McCandlish, S. , author Henighan, T. , et al., year 2020 . title Scaling laws for neural language models . journal arXiv preprint arXiv:2001.08361 :10.48550/arXiv.2001.08361
-
[36]
author Kapoor, S. , author Stroebl, B. , author Siegel, Z.S. , author Nadgir, N. , author Narayanan, A. , year 2024 . title AI agents that matter . journal arXiv preprint arXiv:2407.01502 :10.48550/arXiv.2407.01502
-
[37]
, author Conway, A.R.A
author Kovacs, K. , author Conway, A.R.A. , year 2016 . title Process overlap theory: A unified account of the general factor of intelligence . journal Psychological Inquiry volume 27 , pages 151--177
2016
-
[38]
author Ku, A.Y. , author Campbell, D. , author Bai, X. , et al., year 2025 . title Levels of analysis for large language models . journal arXiv preprint arXiv:2503.13401 :10.48550/arXiv.2503.13401
-
[39]
author Lewis, P. , author Perez, E. , author Piktus, A. , et al., year 2020 . title Retrieval-augmented generation for knowledge-intensive NLP tasks , in: booktitle Advances in Neural Information Processing Systems , pp. pages 9459--9474 . :10.48550/arXiv.2005.11401
-
[40]
author Liang, P. , author Bommasani, R. , author Lee, T. , et al., year 2022 . title Holistic evaluation of language models . journal arXiv preprint arXiv:2211.09110 :10.48550/arXiv.2211.09110
-
[41]
author Lin, S. , author Hilton, J. , author Evans, O. , year 2022 . title TruthfulQA : Measuring how models mimic human falsehoods , in: booktitle Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pp. pages 3214--3252 . :10.18653/v1/2022.acl-long.229
-
[42]
author Ma, Z.R. , author Guo, Y.X. , author Xiao, Y. , year 2026 . title Beyond accuracy scores: Toward process-oriented evaluation of artificial intelligence clinical reasoning in clinical workflow integration . journal International Journal for Quality in Health Care volume 38 , pages mzag034 . :10.1093/intqhc/mzag034
-
[43]
author Magar, I. , author Schwartz, R. , year 2022 . title Data contamination: From memorization to exploitation , in: booktitle Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pp. pages 157--165 . :10.18653/v1/2022.acl-short.18
-
[44]
, year 1982
author Marr, D. , year 1982 . title Vision: A Computational Investigation into the Human Representation and Processing of Visual Information . publisher Henry Holt and Co
1982
-
[45]
author Maynez, J. , author Narayan, S. , author Bohnet, B. , author McDonald, R. , year 2020 . title On faithfulness and factuality in abstractive summarization , in: booktitle Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pp. pages 1906--1919 . :10.18653/v1/2020.acl-main.173
-
[46]
author McNamara, D.S. , author Louwerse, M.M. , author McCarthy, P.M. , author Graesser, A.C. , year 2010 . title Coh-Metrix : Capturing linguistic features of cohesion . journal Discourse Processes volume 47 , pages 292--330 . :10.1080/01638530902959943
-
[47]
author Mehta, S. , year 2025 . title Beyond accuracy: A multi-dimensional framework for evaluating enterprise agentic AI systems . journal arXiv preprint arXiv:2511.14136 :10.48550/arXiv.2511.14136
-
[48]
author Messick, S. , year 1995 . title Validity of psychological assessment . journal American Psychologist volume 50 , pages 741--749 . :10.1037/0003-066X.50.9.741
-
[49]
author Miller, G.A. , year 1956 . title The magical number seven, plus or minus two . journal Psychological Review volume 63 , pages 81--97 . :10.1037/h0043158
doi:10.1037/h0043158 1956
-
[50]
author Mitchell, M. , year 2021 . title Why AI is harder than we think , in: booktitle Proceedings of the Genetic and Evolutionary Computation Conference , pp. pages 4--10 . :10.1145/3449639.3465421
arXiv 2021
-
[51]
author Mokkink, L.B. , author Terwee, C.B. , author Patrick, D.L. , et al., year 2010 . title The COSMIN checklist for assessing the methodological quality of studies on measurement properties . journal Quality of Life Research volume 19 , pages 539--549 . :10.1007/s11136-010-9606-8
-
[52]
author Nakano, R. , author Hilton, J. , author Balaji, S. , et al., year 2021 . title WebGPT : Browser-assisted question-answering with human feedback . journal arXiv preprint arXiv:2112.09332 :10.48550/arXiv.2112.09332
-
[53]
, year 1990
author Newell, A. , year 1990 . title Unified Theories of Cognition . publisher Harvard University Press
1990
-
[54]
, author Simon, H.A
author Newell, A. , author Simon, H.A. , year 1972 . title Human Problem Solving . publisher Prentice-Hall
1972
-
[55]
, author Salomon, G
author Perkins, D.N. , author Salomon, G. , year 1992 . title Transfer of learning , in: booktitle International Encyclopedia of Education . edition 2nd ed.. publisher Pergamon Press
1992
-
[56]
author Peters, M.D.J. , author Godfrey, C. , author McInerney, P. , et al., year 2020 . title Chapter 11: Scoping reviews , in: editor Aromataris, E. , editor Munn, Z. (Eds.), booktitle JBI Manual for Evidence Synthesis . publisher JBI . :10.46658/JBIMES-20-12
-
[57]
author Prinsen, C.A.C. , author Mokkink, L.B. , author Bouter, L.M. , et al., year 2018 . title COSMIN guideline for systematic reviews of patient-reported outcome measures . journal Quality of Life Research volume 27 , pages 1147--1157 . :10.1007/s11136-018-1798-3
-
[58]
author Raji, I.D. , author Kumar, I.E. , author Horowitz, A. , author Selbst, A. , year 2022 . title The fallacy of AI functionality , in: booktitle Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pp. pages 959--972 . :10.1145/3531146.3533158
arXiv 2022
-
[59]
author Rogers, A. , author Kovaleva, O. , author Rumshisky, A. , year 2020 . title A primer in BERT ology: What we know about how BERT works . journal Transactions of the Association for Computational Linguistics volume 8 , pages 842--866 . :10.1162/tacl_a_00349
-
[60]
, author He, H
author Saparov, A. , author He, H. , year 2023 . title Language models are greedy reasoners: A systematic formal analysis of chain-of-thought , in: booktitle Proceedings of the Eleventh International Conference on Learning Representations (ICLR 2023) . https://openreview.net/forum?id=qFVVBzXxR2V
2023
-
[61]
author Sardana, N. , author Portes, J. , author Doubov, S. , author Frankle, J. , year 2024 . title Beyond Chinchilla-Optimal : Accounting for inference in language model scaling laws , in: booktitle Proceedings of the 41st International Conference on Machine Learning (ICML 2024) . :10.48550/arXiv.2401.00448
-
[62]
author Schick, T. , author Dwivedi-Yu, J. , author Dess \`i , R. , et al., year 2024 . title Toolformer: Language models can teach themselves to use tools . journal Advances in Neural Information Processing Systems volume 36 . :10.48550/arXiv.2302.04761
-
[63]
author Schwartz, R. , author Dodge, J. , author Smith, N.A. , author Etzioni, O. , year 2020 . title Green AI . journal Communications of the ACM volume 63 , pages 54--63 . :10.1145/3381831
doi:10.1145/3381831 2020
-
[64]
author Shi, W. , author Min, S. , author Yasunaga, M. , et al., year 2023 . title REPLUG : Retrieval-augmented black-box language models . journal arXiv preprint arXiv:2301.12652 :10.48550/arXiv.2301.12652
-
[65]
author Simon, H.A. , year 1956 . title Rational choice and the structure of the environment . journal Psychological Review volume 63 , pages 129--138 . :10.1037/h0042769
doi:10.1037/h0042769 1956
-
[66]
, year 1972
author Simon, H.A. , year 1972 . title Theories of bounded rationality . journal Decision and Organization volume 1 , pages 161--176
1972
-
[67]
author Srivastava, A. , author Rastogi, A. , author Rao, A. , et al., year 2022 . title Beyond the imitation game: Quantifying and extrapolating the capabilities of language models . journal arXiv preprint arXiv:2206.04615 :10.48550/arXiv.2206.04615
-
[68]
author Strubell, E. , author Ganesh, A. , author McCallum, A. , year 2019 . title Energy and policy considerations for deep learning in NLP , in: booktitle Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pp. pages 3645--3650 . :10.18653/v1/P19-1355
-
[69]
author Sweller, J. , year 1988 . title Cognitive load during problem solving: Effects on learning . journal Cognitive Science volume 12 , pages 257--285 . :10.1207/s15516709cog1202_4
-
[70]
, author van Merri \"e nboer, J.J.G
author Sweller, J. , author van Merri \"e nboer, J.J.G. , author Paas, F. , year 2019 . title Cognitive architecture and instructional design: 20 years later . journal Educational Psychology Review volume 31 , pages 261--292 . :10.1007/s10648-019-09465-5
-
[71]
author Terwee, C.B. , author Prinsen, C.A.C. , author Chiarotto, A. , et al., year 2018 . title COSMIN methodology for evaluating the content validity of patient-reported outcome measures: A delphi study . journal Quality of Life Research volume 27 , pages 1159--1170 . :10.1007/s11136-018-1829-0
-
[72]
author Turpin, M. , author Michael, J. , author Perez, E. , author Bowman, S.R. , year 2024 . title Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting . journal Advances in Neural Information Processing Systems volume 36 . :10.48550/arXiv.2305.04388
-
[73]
author Tversky, A. , author Kahneman, D. , year 1974 . title Judgment under uncertainty: Heuristics and biases . journal Science volume 185 , pages 1124--1131 . :10.1126/science.185.4157.1124
arXiv 1974
-
[74]
author Wang, A. , author Pruksachatkun, Y. , author Nangia, N. , et al., year 2019 . title SuperGLUE : A stickier benchmark for general-purpose language understanding systems , in: booktitle Advances in Neural Information Processing Systems . :10.48550/arXiv.1905.00537
-
[75]
author Wang, X. , author Wei, J. , author Schuurmans, D. , et al., year 2022 . title Self-consistency improves chain of thought reasoning in language models . journal arXiv preprint arXiv:2203.11171 :10.48550/arXiv.2203.11171
-
[76]
author Wei, J. , author Wang, X. , author Schuurmans, D. , et al., year 2022 . title Chain-of-thought prompting elicits reasoning in large language models . journal Advances in Neural Information Processing Systems volume 35 , pages 24824--24837 . :10.48550/arXiv.2201.11903
-
[77]
author Wood, R.E. , year 1986 . title Task complexity: Definition of the construct . journal Organizational Behavior and Human Decision Processes volume 37 , pages 60--82 . :10.1016/0749-5978(86)90044-0
-
[78]
author Yao, S. , author Yu, D. , author Zhao, J. , et al., year 2024 . title Tree of thoughts: Deliberate problem solving with large language models . journal Advances in Neural Information Processing Systems volume 36 . :10.48550/arXiv.2305.10601
-
[79]
author Zhang, T. , author Kishore, V. , author Wu, F. , author Weinberger, K.Q. , author Artzi, Y. , year 2020 . title BERTScore : Evaluating text generation with BERT , in: booktitle Proceedings of the International Conference on Learning Representations . :10.48550/arXiv.1904.09675
-
[80]
author Zhou, L. , author Pacchiardi, L. , author Mart \'i nez-Plumed, F. , et al., year 2026 . title General scales unlock AI evaluation with explanatory and predictive power . journal Nature volume 652 , pages 58--67 . :10.1038/s41586-026-10303-2
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.