Pith. sign in

REVIEW 4 major objections 2 minor 169 references

This paper introduces MAAC, a nine-dimension framework for evaluating text-based AI systems by the cognitive processes behind their outputs rather than by outcome benchmarks alone.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 00:42 UTC pith:F4ZRPMWD

load-bearing objection MAAC is a plausible taxonomy for process-oriented AI evaluation, but the abstract asserts operational status without showing the operational definitions — worth a look if the full text delivers. the 4 major comments →

arxiv 2608.00680 v1 pith:F4ZRPMWD submitted 2026-08-01 cs.AI

Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems

classification cs.AI
keywords AI evaluationcognitive assessmentprocess-oriented evaluationtext-based AIcognitive loadworking memoryhallucination controlevaluation framework
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

MAAC shifts AI evaluation from what a system produces to how it produces it. The paper defines nine cognitively motivated dimensions—such as Cognitive Load, Memory Integration, Hallucination Control, and Process-Outcome Alignment—each grounded in established cognitive science theories. It then offers five theoretical analyses to show the framework is coherent, non-redundant, and empirically testable. If valid, this would give researchers a structured way to diagnose why an AI system reasons well or poorly, complementing existing benchmark scores.

Core claim

The central claim is that process-level cognitive assessment of text-based AI is both theoretically grounded and operationally feasible through nine defined dimensions. Each dimension is tied to a specific cognitive theory, and the paper provides a coverage matrix, gap analysis, and a priori interdependency predictions to argue that the dimensions are broad, non-redundant, and testable. The framework is offered as a complement to outcome-based benchmarks, enabling evaluation of how AI systems reason, manage memory, handle complexity, and avoid false information.

What carries the argument

The central object is the MAAC framework itself: a set of nine named dimensions—Cognitive Load, Tool Execution, Content Quality, Memory Integration, Complexity Handling, Hallucination Control, Knowledge Transfer, Processing Efficiency, and Process-Outcome Alignment. Each dimension maps to a cognitive science theory (e.g., Baddeley's working memory model, Sweller's cognitive load theory), and the framework uses that mapping to define what to measure and how to interpret text outputs as evidence of underlying processes.

Load-bearing premise

The framework assumes that human cognitive constructs like working memory and cognitive load transfer meaningfully to text-based AI systems and that these internal processes can be reliably inferred from observable text outputs.

What would settle it

Run a controlled experiment where two AI systems produce identical text on a task designed to impose high working-memory demands, but one system is given explicit memory aids while the other is not; if MAAC's Memory Integration scores do not differ between the two systems, the framework's claim to measure underlying processes would be falsified.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the framework holds, AI evaluation can move beyond pass/fail benchmarks to structured diagnostic profiles that pinpoint specific cognitive weaknesses.
  • The a priori interdependency predictions give researchers concrete hypotheses to test empirically, linking dimensions like Cognitive Load and Memory Integration.
  • MAAC could serve as a foundation for new evaluation tools that complement existing outcome-based tests by explaining the 'why' behind performance.
  • The gap analysis suggests which current evaluation practices miss process-level information, pointing to where new assessment methods are needed.
  • Process-Outcome Alignment as a dimension could help identify cases where a system gets the right answer for the wrong reasons.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural extension would be to apply MAAC to multimodal AI, though the current dimensions are text-specific and would need re-mapping to non-textual outputs.
  • MAAC could also be used as an interpretability lens: a system's dimension scores might reveal where internal representations fail, offering a bridge between evaluation and model debugging.
  • If the framework's predictions hold, it might inform AI design—for instance, architectural changes that reduce measured Cognitive Load could be prioritized even when outcomes are unchanged.
  • A testable extension is to compare MAAC dimension scores against human cognitive measures on the same tasks; strong correlation would strengthen the theory, while divergence would indicate limits of the analogy.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 2 minor

Summary. The paper introduces MAAC, a Multi-Dimensional Assessment framework for text-based AI systems, proposing nine cognitively motivated dimensions (Cognitive Load, Tool Execution, Content Quality, Memory Integration, Complexity Handling, Hallucination Control, Knowledge Transfer, Processing Efficiency, Process-Outcome Alignment). It claims these dimensions are grounded in established cognitive science theories and that five theoretical analyses—dimension-to-theory mapping, a coverage matrix, a gap analysis, a worked diagnostic illustration, and a priori interdependency predictions—provide initial support for coherence and empirical testability. The stated goal is to complement outcome-based benchmarks with a process-level, multi-dimensional evaluation. This review is based only on the abstract, as the full text was not supplied.

Significance. If the framework succeeds, it would be a meaningful contribution to AI evaluation by shifting attention from task accuracy to underlying cognitive processes, with practical implications for diagnosing strengths and failure modes of text-based AI systems. The explicit list of a priori interdependency predictions is a commendable feature, as it creates a route to falsification. The coverage matrix and gap analysis, if executed rigorously, could also help position MAAC relative to existing benchmarks. However, the significance cannot be fully assessed from the abstract: the operational validity of the nine dimensions, the non-redundancy argument, and the empirical testability are asserted rather than demonstrated. The abstract indicates a theoretical framework, but the stronger claim of being an 'operational framework' requires evidence of concrete measurement procedures and validation, which is not visible here.

major comments (4)
  1. [Abstract, first sentence of the contribution claim] The abstract calls MAAC an 'operational framework,' but it does not specify how each of the nine dimensions is measured from observable text outputs. Without explicit operational definitions—e.g., which text features correspond to Cognitive Load or Memory Integration—dimension scores are underdetermined. For instance, Cognitive Load could plausibly be scored from output length, token-level perplexity, revision patterns, or response time, and these features may reflect training data or decoding parameters rather than any latent cognitive process. The abstract provides no such mappings, so the operational claim is unsupported.
  2. [Abstract, 'Five theoretical analyses provide initial support'] The listed analyses are all internal coherence checks: dimension-to-theory mapping, coverage matrix, gap analysis, an illustration, and a priori predictions. These can establish that the framework is internally consistent but do not, by themselves, establish that the dimensions correspond to distinct, real cognitive processes in AI systems or that those processes are inferable from text. In particular, a coverage matrix can be self-affirming if the categories are chosen to match the framework. The abstract should state whether any external validation—e.g., comparison with independent human judgments, behavioral experiments, or controlled manipulation of cognitive variables—is possible or planned.
  3. [Abstract, 'Process-Outcome Alignment' dimension] There is a potential circularity in including 'Process-Outcome Alignment' both as one of the nine dimensions and as part of the evaluation framework. If this dimension is scored from the same observable features that the framework uses to define the other dimensions, then its measurement would be tautological. To avoid this, the abstract should indicate that the alignment dimension is assessed using independent outcome measures (e.g., external ground-truth correctness or human judgment) rather than features derived from the MAAC dimensions themselves. Without such a distinction, the framework's validation is at risk of being circular.
  4. [Abstract, nine-dimension taxonomy] The list of dimensions mixes constructs of different types: 'Tool Execution' refers to an external action, 'Content Quality' and 'Hallucination Control' are outcome-adjacent quality judgments, while 'Memory Integration' and 'Complexity Handling' are more process-like. The abstract asserts that a coverage matrix assesses 'non-redundancy,' but no evidence is shown. The reader cannot verify whether the nine dimensions are mutually exclusive or collectively exhaustive, nor whether some dimensions are actually composites of others. The abstract should explain the selection criterion and provide at least a summary of the non-redundancy analysis.
minor comments (2)
  1. [Abstract] The abstract does not mention any limitations or scope conditions. For example, does the framework apply only to autoregressive language models, or also to retrieval-augmented and tool-using systems? Clarifying the intended scope would help readers assess the claims.
  2. [Abstract, 'Marr's tri-level hypothesis'] The references to Marr, Baddeley, Sweller, and unified theories of cognition are listed but not connected to specific dimensions. In a full paper, this mapping is presumably shown, but the abstract would benefit from one example—e.g., how Baddeley's working memory model justifies the Memory Integration dimension.

Circularity Check

0 steps flagged

No circularity identifiable from the abstract; MAAC's claims are presented as a theoretical framework with a priori predictions, not as fitted results or self-referential derivations.

full rationale

This review is based on the abstract only. The abstract claims that MAAC defines nine dimensions grounded in established cognitive science theory and offers five theoretical analyses: dimension-to-theory mapping, coverage matrix, gap analysis, a worked diagnostic illustration, and a priori predictions. None of these, as described, asserts that a quantity is derived from data and then predicted back from the same quantity. The 'worked diagnostic illustration' could in principle be circular if it merely scores dimensions using the same features the dimensions are meant to explain, but the abstract does not provide enough detail to show such a reduction, and it is not described as a validation of the framework. The 'a priori interdependency predictions' are explicitly deferred to future empirical testing, which is the opposite of fitting a parameter and calling it a prediction. The cited theories are classic external works rather than self-citations by the authors. There is therefore no quotable equation, definition, or fitted-input relation that would support a circularity finding. A legitimate open concern is construct validity—whether text outputs can reliably indicate latent cognitive processes—but that is an empirical and theoretical risk, not an instance of circular derivation. Under the hard rule requiring a specific quoted reduction, no circular step can be flagged.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 1 invented entities

Abstract-only review: no numeric free parameters or fitted values are available. The main burden lies in the domain assumption that human cognitive constructs are appropriate and measurable for AI systems, and in the ad hoc choice of nine dimensions as the correct decomposition.

axioms (3)
  • domain assumption Human cognitive science theories (Marr, Baddeley, Sweller) can be meaningfully applied to text-based AI systems.
    The framework's dimensions are grounded in theories of human cognition, but AI systems do not necessarily implement human-like working memory or cognitive load. This transfer is assumed, not demonstrated in the abstract.
  • domain assumption AI cognitive processes are separable from their outputs and observable through text behavior.
    The core premise of process-oriented evaluation is that internal cognitive states leave traces in text that can be assessed. This is an unproven modeling assumption.
  • ad hoc to paper The nine chosen dimensions form a non-redundant and sufficiently complete decomposition of AI cognition.
    The dimension set is introduced by the authors; the abstract mentions a coverage matrix but does not demonstrate independent justification or empirical validation of the particular partition.
invented entities (1)
  • MAAC nine-dimension taxonomy no independent evidence
    purpose: Provides the measurement axes for process-oriented cognitive evaluation of text-based AI systems.
    This taxonomy is introduced in the paper as the central construct. No external or empirical evidence is presented in the abstract to validate the dimensions or their operational definitions.

pith-pipeline@v1.3.0-daily-deepseek · 564 in / 6101 out tokens · 58017 ms · 2026-08-04T00:42:44.124164+00:00 · methodology

0 comments
read the original abstract

Evaluating artificial intelligence systems has historically relied on outcome-based benchmarks that measure task accuracy, robustness, or fairness. While indispensable, these benchmarks provide limited diagnostic insight into the underlying cognitive processes that generate performance-leaving critical questions unanswered about how AI systems reason, integrate memory, manage complexity, or avoid generating false information. This paper introduces the Multi-Dimensional Assessment for AI Cognition (MAAC), a theoretically grounded framework for shifting evaluation from what text-based AI systems produce to how they think. MAAC defines nine cognitively motivated dimensions: Cognitive Load, Tool Execution, Content Quality, Memory Integration, Complexity Handling, Hallucination Control, Knowledge Transfer, Processing Efficiency, and Process-Outcome Alignment. Each dimension is grounded in established cognitive science theory-drawing on Marr's tri-level hypothesis, Baddeley's working memory model, Sweller's cognitive load theory, and unified theories of cognition. Five theoretical analyses provide initial support for the framework's coherence and empirical testability: dimension-to-theory mapping; a coverage matrix assessing breadth and non-redundancy; a formal gap analysis relative to current evaluation practice; a worked diagnostic illustration; and a set of a priori interdependency predictions for future empirical testing. MAAC provides a theoretical and operational framework for principled process-level cognitive assessment of text-based AI systems, complementing existing outcome-based benchmarks with cognitively grounded, multi-dimensional evaluation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

169 extracted references · 8 canonical work pages

  1. [1]

    , author Reuel, A

    author Akhtar, M. , author Reuel, A. , author Soni, P. , et al., year 2026 . title When AI benchmarks plateau: A systematic study of benchmark saturation . journal arXiv preprint arXiv:2602.16763 :10.48550/arXiv.2602.16763

  2. [2]

    , author Bothell, D

    author Anderson, J.R. , author Bothell, D. , author Byrne, M.D. , author Douglass, S. , author Lebiere, C. , author Qin, Y. , year 2004 . title An integrated theory of the mind . journal Psychological Review volume 111 , pages 1036--1060 . :10.1037/0033-295X.111.4.1036

  3. [3]

    , author O'Malley, L

    author Arksey, H. , author O'Malley, L. , year 2005 . title Scoping studies: Towards a methodological framework . journal International Journal of Social Research Methodology volume 8 , pages 19--32 . :10.1080/1364557032000119616

  4. [4]

    , year 1992

    author Baddeley, A. , year 1992 . title Working memory . journal Science volume 255 , pages 556--559 . :10.1126/science.1736359

  5. [5]

    , year 2000

    author Baddeley, A. , year 2000 . title The episodic buffer: A new component of working memory? journal Trends in Cognitive Sciences volume 4 , pages 417--423 . :10.1016/S1364-6613(00)01538-2

  6. [6]

    , year 2003

    author Baddeley, A. , year 2003 . title Working memory: Looking back and looking forward . journal Nature Reviews Neuroscience volume 4 , pages 829--839 . :10.1038/nrn1201

  7. [7]

    , author Ceci, S.J

    author Barnett, S.M. , author Ceci, S.J. , year 2002 . title When and where do we apply what we learn? A taxonomy for far transfer . journal Psychological Bulletin volume 128 , pages 612--637 . :10.1037/0033-2909.128.4.612

  8. [8]

    , author Gebru, T

    author Bender, E.M. , author Gebru, T. , author McMillan-Major, A. , author Shmitchell, S. , year 2021 . title On the dangers of stochastic parrots: Can language models be too big? , in: booktitle Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pp. pages 610--623 . :10.1145/3442188.3445922

  9. [9]

    , author Hudson, D.A

    author Bommasani, R. , author Hudson, D.A. , author Adeli, E. , et al., year 2021 . title On the opportunities and risks of foundation models . journal arXiv preprint arXiv:2108.07258 :10.48550/arXiv.2108.07258

  10. [10]

    , author Mensch, A

    author Borgeaud, S. , author Mensch, A. , author Hoffmann, J. , et al., year 2022 . title Improving language models by retrieving from trillions of tokens , in: booktitle Proceedings of the 39th International Conference on Machine Learning , pp. pages 2206--2240 . https://proceedings.mlr.press/v162/borgeaud22a.html

  11. [11]

    , year 1988

    author Campbell, D.J. , year 1988 . title Task complexity: A review and analysis . journal Academy of Management Review volume 13 , pages 40--52 . :10.5465/amr.1988.4306775

  12. [12]

    , author Chalmers, D

    author Clark, A. , author Chalmers, D. , year 1998 . title The extended mind . journal Analysis volume 58 , pages 7--19 . :10.1093/analys/58.1.7

  13. [13]

    , author Khandelwal, U

    author Clark, K. , author Khandelwal, U. , author Levy, O. , author Manning, C.D. , year 2019 . title What does BERT look at? an analysis of BERT 's attention , in: booktitle Proceedings of the 2019 ACL Workshop BlackboxNLP , pp. pages 276--286 . :10.18653/v1/W19-4828

  14. [14]

    , author Meehl, P.E

    author Cronbach, L.J. , author Meehl, P.E. , year 1955 . title Construct validity in psychological tests . journal Psychological Bulletin volume 52 , pages 281--302 . :10.1037/h0040957

  15. [15]

    , author Kyle, K

    author Crossley, S.A. , author Kyle, K. , author McNamara, D.S. , year 2016 . title The tool for the automatic analysis of text cohesion ( TAACO ) . journal Behavior Research Methods volume 48 , pages 1227--1237 . :10.3758/s13428-015-0651-7

  16. [16]

    , year 2017

    author DeVellis, R.F. , year 2017 . title Scale Development: Theory and Applications . edition 4th ed., publisher SAGE Publications

  17. [17]

    , author Kim, B

    author Doshi-Velez, F. , author Kim, B. , year 2017 . title Towards a rigorous science of interpretable machine learning . journal arXiv preprint arXiv:1702.08608 :10.48550/arXiv.1702.08608

  18. [18]

    , author Purificato, E

    author Eriksson, M. , author Purificato, E. , author Noroozian, A. , author Vinagre, J. , author Chaslot, G. , author Gomez, E. , author Fernandez-Llorca, D. , year 2025 . title Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation . journal arXiv preprint arXiv:2502.06559 :10.48550/arXiv.2502.06559

  19. [19]

    , author Jurafsky, D

    author Ethayarajh, K. , author Jurafsky, D. , year 2020 . title Utility is in the eye of the user: A critique of NLP leaderboards , in: booktitle Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pp. pages 4846--4853 . :10.18653/v1/2020.emnlp-main.393

  20. [20]

    , year 2018

    author Furr, R.M. , year 2018 . title Psychometrics: An Introduction . edition 3rd ed., publisher SAGE Publications

  21. [21]

    , author Ghahramani, Z

    author Gal, Y. , author Ghahramani, Z. , year 2016 . title Dropout as a Bayesian approximation: Representing model uncertainty in deep learning , in: booktitle Proceedings of the 33rd International Conference on Machine Learning , pp. pages 1050--1059 . https://proceedings.mlr.press/v48/gal16.html

  22. [22]

    , year 1983

    author Gentner, D. , year 1983 . title Structure-mapping: A theoretical framework for analogy . journal Cognitive Science volume 7 , pages 155--170 . :10.1207/s15516709cog0702_3

  23. [23]

    , author Lieder, F

    author Griffiths, T.L. , author Lieder, F. , author Goodman, N.D. , year 2015 . title Rational use of cognitive resources: Levels of analysis between the computational and the algorithmic . journal Topics in Cognitive Science volume 7 , pages 217--229 . :10.1111/tops.12142

  24. [24]

    , author Pleiss, G

    author Guo, C. , author Pleiss, G. , author Sun, Y. , author Weinberger, K.Q. , year 2017 . title On calibration of modern neural networks , in: booktitle Proceedings of the 34th International Conference on Machine Learning , pp. pages 1321--1330 . https://proceedings.mlr.press/v70/guo17a.html

  25. [25]

    , author Baker, R

    author Halford, G.S. , author Baker, R. , author McCredden, J.E. , author Bain, J.D. , year 2005 . title How many variables can humans process? journal Psychological Science volume 16 , pages 70--76 . :10.1111/j.0956-7976.2005.00782.x

  26. [26]

    , author Hasan, R

    author Halliday, M.A.K. , author Hasan, R. , year 1976 . title Cohesion in English . publisher Longman

  27. [27]

    , author Burns, C

    author Hendrycks, D. , author Burns, C. , author Basart, S. , et al., year 2021 . title Measuring massive multitask language understanding , in: booktitle Proceedings of the International Conference on Learning Representations . :10.48550/arXiv.2009.03300

  28. [28]

    , author Borgeaud, S

    author Hoffmann, J. , author Borgeaud, S. , author Mensch, A. , et al., year 2022 . title Training compute-optimal large language models . journal arXiv preprint arXiv:2203.15556 :10.48550/arXiv.2203.15556

  29. [29]

    , editor Morrison, R.G

    editor Holyoak, K.J. , editor Morrison, R.G. (Eds.), year 2012 . title The Oxford Handbook of Thinking and Reasoning . publisher Oxford University Press

  30. [30]

    , author Yu, W

    author Huang, L. , author Yu, W. , author Ma, W. , et al., year 2023 . title A survey on hallucination in large language models . journal arXiv preprint arXiv:2311.05232 :10.48550/arXiv.2311.05232

  31. [31]

    , year 1995

    author Hutchins, E. , year 1995 . title Cognition in the Wild . publisher MIT Press

  32. [32]

    , author Lee, N

    author Ji, Z. , author Lee, N. , author Frieske, R. , et al., year 2023 . title Survey of hallucination in natural language generation . journal ACM Computing Surveys volume 55 , pages 1--38 . :10.1145/3571730

  33. [33]

    , author Carpenter, P.A

    author Just, M.A. , author Carpenter, P.A. , year 1992 . title A capacity theory of comprehension: Individual differences in working memory . journal Psychological Review volume 99 , pages 122--149 . :10.1037/0033-295X.99.1.122

  34. [34]

    , year 2011

    author Kahneman, D. , year 2011 . title Thinking, Fast and Slow . publisher Farrar, Straus and Giroux

  35. [35]

    , author McCandlish, S

    author Kaplan, J. , author McCandlish, S. , author Henighan, T. , et al., year 2020 . title Scaling laws for neural language models . journal arXiv preprint arXiv:2001.08361 :10.48550/arXiv.2001.08361

  36. [36]

    , author Stroebl, B

    author Kapoor, S. , author Stroebl, B. , author Siegel, Z.S. , author Nadgir, N. , author Narayanan, A. , year 2024 . title AI agents that matter . journal arXiv preprint arXiv:2407.01502 :10.48550/arXiv.2407.01502

  37. [37]

    , author Conway, A.R.A

    author Kovacs, K. , author Conway, A.R.A. , year 2016 . title Process overlap theory: A unified account of the general factor of intelligence . journal Psychological Inquiry volume 27 , pages 151--177

  38. [38]

    , author Campbell, D

    author Ku, A.Y. , author Campbell, D. , author Bai, X. , et al., year 2025 . title Levels of analysis for large language models . journal arXiv preprint arXiv:2503.13401 :10.48550/arXiv.2503.13401

  39. [39]

    , author Perez, E

    author Lewis, P. , author Perez, E. , author Piktus, A. , et al., year 2020 . title Retrieval-augmented generation for knowledge-intensive NLP tasks , in: booktitle Advances in Neural Information Processing Systems , pp. pages 9459--9474 . :10.48550/arXiv.2005.11401

  40. [40]

    , author Bommasani, R

    author Liang, P. , author Bommasani, R. , author Lee, T. , et al., year 2022 . title Holistic evaluation of language models . journal arXiv preprint arXiv:2211.09110 :10.48550/arXiv.2211.09110

  41. [41]

    , author Hilton, J

    author Lin, S. , author Hilton, J. , author Evans, O. , year 2022 . title TruthfulQA : Measuring how models mimic human falsehoods , in: booktitle Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pp. pages 3214--3252 . :10.18653/v1/2022.acl-long.229

  42. [42]

    , author Guo, Y.X

    author Ma, Z.R. , author Guo, Y.X. , author Xiao, Y. , year 2026 . title Beyond accuracy scores: Toward process-oriented evaluation of artificial intelligence clinical reasoning in clinical workflow integration . journal International Journal for Quality in Health Care volume 38 , pages mzag034 . :10.1093/intqhc/mzag034

  43. [43]

    , author Schwartz, R

    author Magar, I. , author Schwartz, R. , year 2022 . title Data contamination: From memorization to exploitation , in: booktitle Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pp. pages 157--165 . :10.18653/v1/2022.acl-short.18

  44. [44]

    , year 1982

    author Marr, D. , year 1982 . title Vision: A Computational Investigation into the Human Representation and Processing of Visual Information . publisher Henry Holt and Co

  45. [45]

    , author Narayan, S

    author Maynez, J. , author Narayan, S. , author Bohnet, B. , author McDonald, R. , year 2020 . title On faithfulness and factuality in abstractive summarization , in: booktitle Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pp. pages 1906--1919 . :10.18653/v1/2020.acl-main.173

  46. [46]

    , author Louwerse, M.M

    author McNamara, D.S. , author Louwerse, M.M. , author McCarthy, P.M. , author Graesser, A.C. , year 2010 . title Coh-Metrix : Capturing linguistic features of cohesion . journal Discourse Processes volume 47 , pages 292--330 . :10.1080/01638530902959943

  47. [47]

    , year 2025

    author Mehta, S. , year 2025 . title Beyond accuracy: A multi-dimensional framework for evaluating enterprise agentic AI systems . journal arXiv preprint arXiv:2511.14136 :10.48550/arXiv.2511.14136

  48. [48]

    , year 1995

    author Messick, S. , year 1995 . title Validity of psychological assessment . journal American Psychologist volume 50 , pages 741--749 . :10.1037/0003-066X.50.9.741

  49. [49]

    , year 1956

    author Miller, G.A. , year 1956 . title The magical number seven, plus or minus two . journal Psychological Review volume 63 , pages 81--97 . :10.1037/h0043158

  50. [50]

    , year 2021

    author Mitchell, M. , year 2021 . title Why AI is harder than we think , in: booktitle Proceedings of the Genetic and Evolutionary Computation Conference , pp. pages 4--10 . :10.1145/3449639.3465421

  51. [51]

    , author Terwee, C.B

    author Mokkink, L.B. , author Terwee, C.B. , author Patrick, D.L. , et al., year 2010 . title The COSMIN checklist for assessing the methodological quality of studies on measurement properties . journal Quality of Life Research volume 19 , pages 539--549 . :10.1007/s11136-010-9606-8

  52. [52]

    , author Hilton, J

    author Nakano, R. , author Hilton, J. , author Balaji, S. , et al., year 2021 . title WebGPT : Browser-assisted question-answering with human feedback . journal arXiv preprint arXiv:2112.09332 :10.48550/arXiv.2112.09332

  53. [53]

    , year 1990

    author Newell, A. , year 1990 . title Unified Theories of Cognition . publisher Harvard University Press

  54. [54]

    , author Simon, H.A

    author Newell, A. , author Simon, H.A. , year 1972 . title Human Problem Solving . publisher Prentice-Hall

  55. [55]

    , author Salomon, G

    author Perkins, D.N. , author Salomon, G. , year 1992 . title Transfer of learning , in: booktitle International Encyclopedia of Education . edition 2nd ed.. publisher Pergamon Press

  56. [56]

    , author Godfrey, C

    author Peters, M.D.J. , author Godfrey, C. , author McInerney, P. , et al., year 2020 . title Chapter 11: Scoping reviews , in: editor Aromataris, E. , editor Munn, Z. (Eds.), booktitle JBI Manual for Evidence Synthesis . publisher JBI . :10.46658/JBIMES-20-12

  57. [57]

    , author Mokkink, L.B

    author Prinsen, C.A.C. , author Mokkink, L.B. , author Bouter, L.M. , et al., year 2018 . title COSMIN guideline for systematic reviews of patient-reported outcome measures . journal Quality of Life Research volume 27 , pages 1147--1157 . :10.1007/s11136-018-1798-3

  58. [58]

    , author Kumar, I.E

    author Raji, I.D. , author Kumar, I.E. , author Horowitz, A. , author Selbst, A. , year 2022 . title The fallacy of AI functionality , in: booktitle Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pp. pages 959--972 . :10.1145/3531146.3533158

  59. [59]

    , author Kovaleva, O

    author Rogers, A. , author Kovaleva, O. , author Rumshisky, A. , year 2020 . title A primer in BERT ology: What we know about how BERT works . journal Transactions of the Association for Computational Linguistics volume 8 , pages 842--866 . :10.1162/tacl_a_00349

  60. [60]

    , author He, H

    author Saparov, A. , author He, H. , year 2023 . title Language models are greedy reasoners: A systematic formal analysis of chain-of-thought , in: booktitle Proceedings of the Eleventh International Conference on Learning Representations (ICLR 2023) . https://openreview.net/forum?id=qFVVBzXxR2V

  61. [61]

    , author Portes, J

    author Sardana, N. , author Portes, J. , author Doubov, S. , author Frankle, J. , year 2024 . title Beyond Chinchilla-Optimal : Accounting for inference in language model scaling laws , in: booktitle Proceedings of the 41st International Conference on Machine Learning (ICML 2024) . :10.48550/arXiv.2401.00448

  62. [62]

    , author Dwivedi-Yu, J

    author Schick, T. , author Dwivedi-Yu, J. , author Dess \`i , R. , et al., year 2024 . title Toolformer: Language models can teach themselves to use tools . journal Advances in Neural Information Processing Systems volume 36 . :10.48550/arXiv.2302.04761

  63. [63]

    , author Dodge, J

    author Schwartz, R. , author Dodge, J. , author Smith, N.A. , author Etzioni, O. , year 2020 . title Green AI . journal Communications of the ACM volume 63 , pages 54--63 . :10.1145/3381831

  64. [64]

    , author Min, S

    author Shi, W. , author Min, S. , author Yasunaga, M. , et al., year 2023 . title REPLUG : Retrieval-augmented black-box language models . journal arXiv preprint arXiv:2301.12652 :10.48550/arXiv.2301.12652

  65. [65]

    , year 1956

    author Simon, H.A. , year 1956 . title Rational choice and the structure of the environment . journal Psychological Review volume 63 , pages 129--138 . :10.1037/h0042769

  66. [66]

    , year 1972

    author Simon, H.A. , year 1972 . title Theories of bounded rationality . journal Decision and Organization volume 1 , pages 161--176

  67. [67]

    , author Rastogi, A

    author Srivastava, A. , author Rastogi, A. , author Rao, A. , et al., year 2022 . title Beyond the imitation game: Quantifying and extrapolating the capabilities of language models . journal arXiv preprint arXiv:2206.04615 :10.48550/arXiv.2206.04615

  68. [68]

    , author Ganesh, A

    author Strubell, E. , author Ganesh, A. , author McCallum, A. , year 2019 . title Energy and policy considerations for deep learning in NLP , in: booktitle Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pp. pages 3645--3650 . :10.18653/v1/P19-1355

  69. [69]

    , year 1988

    author Sweller, J. , year 1988 . title Cognitive load during problem solving: Effects on learning . journal Cognitive Science volume 12 , pages 257--285 . :10.1207/s15516709cog1202_4

  70. [70]

    , author van Merri \"e nboer, J.J.G

    author Sweller, J. , author van Merri \"e nboer, J.J.G. , author Paas, F. , year 2019 . title Cognitive architecture and instructional design: 20 years later . journal Educational Psychology Review volume 31 , pages 261--292 . :10.1007/s10648-019-09465-5

  71. [71]

    , author Prinsen, C.A.C

    author Terwee, C.B. , author Prinsen, C.A.C. , author Chiarotto, A. , et al., year 2018 . title COSMIN methodology for evaluating the content validity of patient-reported outcome measures: A delphi study . journal Quality of Life Research volume 27 , pages 1159--1170 . :10.1007/s11136-018-1829-0

  72. [72]

    , author Michael, J

    author Turpin, M. , author Michael, J. , author Perez, E. , author Bowman, S.R. , year 2024 . title Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting . journal Advances in Neural Information Processing Systems volume 36 . :10.48550/arXiv.2305.04388

  73. [73]

    , author Kahneman, D

    author Tversky, A. , author Kahneman, D. , year 1974 . title Judgment under uncertainty: Heuristics and biases . journal Science volume 185 , pages 1124--1131 . :10.1126/science.185.4157.1124

  74. [74]

    , author Pruksachatkun, Y

    author Wang, A. , author Pruksachatkun, Y. , author Nangia, N. , et al., year 2019 . title SuperGLUE : A stickier benchmark for general-purpose language understanding systems , in: booktitle Advances in Neural Information Processing Systems . :10.48550/arXiv.1905.00537

  75. [75]

    , author Wei, J

    author Wang, X. , author Wei, J. , author Schuurmans, D. , et al., year 2022 . title Self-consistency improves chain of thought reasoning in language models . journal arXiv preprint arXiv:2203.11171 :10.48550/arXiv.2203.11171

  76. [76]

    , author Wang, X

    author Wei, J. , author Wang, X. , author Schuurmans, D. , et al., year 2022 . title Chain-of-thought prompting elicits reasoning in large language models . journal Advances in Neural Information Processing Systems volume 35 , pages 24824--24837 . :10.48550/arXiv.2201.11903

  77. [77]

    , year 1986

    author Wood, R.E. , year 1986 . title Task complexity: Definition of the construct . journal Organizational Behavior and Human Decision Processes volume 37 , pages 60--82 . :10.1016/0749-5978(86)90044-0

  78. [78]

    , author Yu, D

    author Yao, S. , author Yu, D. , author Zhao, J. , et al., year 2024 . title Tree of thoughts: Deliberate problem solving with large language models . journal Advances in Neural Information Processing Systems volume 36 . :10.48550/arXiv.2305.10601

  79. [79]

    , author Kishore, V

    author Zhang, T. , author Kishore, V. , author Wu, F. , author Weinberger, K.Q. , author Artzi, Y. , year 2020 . title BERTScore : Evaluating text generation with BERT , in: booktitle Proceedings of the International Conference on Learning Representations . :10.48550/arXiv.1904.09675

  80. [80]

    , author Pacchiardi, L

    author Zhou, L. , author Pacchiardi, L. , author Mart \'i nez-Plumed, F. , et al., year 2026 . title General scales unlock AI evaluation with explanatory and predictive power . journal Nature volume 652 , pages 58--67 . :10.1038/s41586-026-10303-2

Showing first 80 references.