Pith. sign in

REVIEW 5 major objections 6 minor 1 cited by

Can OpenAI o1 outperform humans in higher-order cognitive thinking?

T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper claims that OpenAI's o1-preview model outperforms human comparison groups on five of seven established tests of higher-order thinking, with its main weakness showing up in problem-solving.

desk verdict The paper's own Table 5 contradicts its five-of-seven claim, but the assembled battery and the large systematic-thinking effects make it worth a referee's time. read the letter →

arxiv 2412.05753 v1 pith:GJ5ZXOBD submitted 2024-12-07 cs.CY cs.AI

classification cs.CYcs.AI
keywords OpenAIo1-previewhigher-orderthinkinglargelanguagemodelseducationalassessmentcriticalscientificreasoningbenchmarkcomparisonineducation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether OpenAI's o1-preview model can outperform human students and experts on established tests of higher-order thinking, and the introduction's answer is yes in five of seven domains: systematic thinking, computational thinking, data literacy, creative thinking, and scientific reasoning. The results tables actually show the model scoring above the human comparison groups on nearly every instrument, including the Ennis-Weir critical thinking essay test and LogiQA, with the clearest exception being the problem-solving subscale of the computational-thinking instrument, where the model scored 1.00 against a human mean of 3.68. The authors interpret the pattern as evidence that structured, well-defined assessments play to the model's strengths, while adaptive, ill-structured problem-solving remains a human strength. If the comparisons hold, AI tools could take on routine assessment and tutoring of higher-order skills, though the paper argues human oversight and better assessments are still needed.

What carries the argument

The machinery is a suite of published, normed assessment instruments, each administered to o1-preview as text prompts and scored with the same rubrics used for human test-takers: EWCTET, the Lake Urmia Vignette, the Computational Thinking Skills scale, the ATTA, Merk et al.'s and Chen et al.'s data literacy tests, the AUT and RAT, LogiQA, and TOSLS. The model's scores are then standardized against the human means and standard deviations reported in the source studies, with z-scores and one-sample t-tests used to express how far above or below the human distribution the model sits. The comparison only works if the instruments measure what they claim to measure, the human norms are representative, and the model encounters the items as novel prompts rather than as memorized training text.

What would settle it

Give o1-preview newly written parallel versions of the same seven assessments—same constructs and formats, but items that cannot be in its training data—and compare its scores on the originals; a large drop, or the model's ability to recite original items on demand, would show that the reported human-beating performance came from contamination rather than higher-order reasoning. A concrete spot check is the TOSLS Item 2 graph-selection question, which was converted from visual plots to text for the model: present the same item with novel graphs, and see whether the near-perfect score survives.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that a general-purpose reasoning model, o1-preview, can be prompted with the items of standard cognitive assessments and score at or above the human means those instruments were designed to spread out. The model scored 24.33 on the Ennis-Weir critical thinking essay test versus 18.39 for postgraduates; 46.10 on the Lake Urmia Vignette total versus 20.08 for undergraduates; 8.60 on Merk et al.'s "Use Data" dimension versus a 4.17 post-test mean; near-perfect 0.99 on TOSLS versus 0.85 for the best student cohort; 90% accuracy on LogiQA versus 86% for humans; and a perfect 20/20 on the ATTA versus 14.63 for experts. The paper claims this as outperformance in five of seven domains, while flagging that the model's problem-solving score on the computational-thinking scale was far below humans and that visual TOSLS items had to be converted to text, which may have affected results.

Load-bearing premise

The load-bearing premise is that o1-preview never memorized the publicly available test questions during training, so its high scores reflect reasoning rather than recall of answers.

Editorial extensions

If this is right

  • If the reported comparisons hold, o1-preview can already outperform typical university students on structured critical-thinking, data-literacy, and scientific-reasoning assessments, which suggests AI tools could take over routine scoring and tutoring of those skills.
  • The model's near-zero score on the computational-thinking problem-solving subscale is a direct counterexample to the idea that LLMs have generalized reasoning, so the paper implies that open-ended, ill-structured problem-solving should remain a human responsibility in AI-assisted classrooms.
  • Because the model saturates TOSLS and nearly saturates LogiQA, the paper's results imply that widely used tests of 'higher-order thinking' may need redesign if they are to keep measuring human cognitive development rather than machine pattern recognition.
  • The authors' recommendation that AI be used as a supplement, not a replacement, follows directly from the observed mix of very high scores on structured items and very low scores on problem-solving items.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the same logic implies that any LLM trained on public exam corpora will tend to saturate standardized reasoning tests, so the meaningful next comparison is on fresh, non-public items rather than on legacy benchmarks.
  • A test the authors did not run would be to administer paraphrased and answer-permuted versions of the same instruments; a large score drop on those versions would indicate the model is exploiting surface patterns rather than performing the reasoning the tests claim to measure.
  • The paper compares the model against human means drawn from different studies, cohorts, and years, so a direct head-to-head study with the same human sample and the same items would be needed to confirm the claimed five-of-seven superiority.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This manuscript evaluates OpenAI's o1-preview model on seven purported higher-order thinking domains — critical thinking, systematic thinking, computational thinking, data literacy, creative thinking, logical reasoning, and scientific reasoning — using established instruments (EWCTET, Village of Abeesee/LUV, CT Skills/ATTA, Merk/Chen data literacy tests, AUT/RAT, LogiQA, TOSLS). Human performance is taken from previously published studies, and the model is run for 10 trials per instrument. The authors report z-scores for model-versus-human comparisons and claim in the abstract and introduction that o1-preview outperforms human experts in five of seven domains. The paper concludes that AI can complement education in structured assessments but needs human oversight.

Significance. If the headline claim were supported, this would be a notable contribution to the debate on whether LLMs can match or exceed human performance on structured higher-order cognition assessments. The study's strengths include the use of multiple established instruments, transparent reporting of per-dimension means and z-scores in Tables 4–11, and an explicit acknowledgment of the model's weakness in problem-solving. However, the paper's central assertion is internally contradicted by its own reported results, and several load-bearing methodological choices (human-benchmark selection, contamination, and differential scoring across arms) undermine the comparability of the AI-human comparisons. As it stands, the evidence does not support the claimed five-of-seven outperformance.

major comments (5)
  1. [§1, Abstract, Table 5] The claim that o1-preview 'outperforms human experts in five out of seven domains, including ... computational thinking' is contradicted by Table 5, which shows that for Computational Thinking Skills the model's overall z-score is -0.15 (mean 3.84 vs. human 3.92) and the Problem-Solving z-score is -4.25 (mean 1.00 vs. human 3.68). Thus computational thinking is not outperformed overall. Moreover, the five named domains omit critical thinking (EWCTET z = 1.60 and 0.90 in Table 3) and logical reasoning (LogiQA z = 0.62 in Table 10), both of which have positive z-scores; if those count, the correct count would be six of seven, not five. This suggests the 'five of seven' list is not supported by the paper's own data and must be corrected or explicitly justified.
  2. [§2.2.6, §2.2.7, §3.6, §3.7] The paper never addresses the possibility that o1-preview's training data included the benchmark items. LogiQA, TOSLS, AUT, RAT, and EWCTET are all public, widely used evaluation instruments, and many are standard in LLM benchmarks. If the model has seen these items or their solutions, high scores may reflect memorization rather than reasoning. This is a load-bearing premise for the claim that o1-preview 'outperforms human experts' in higher-order thinking. The authors should either provide evidence of non-contamination, test on private held-out variants, or explicitly discuss this limitation and temper the conclusions accordingly.
  3. [§3.5, Table 9] The human AUT originality benchmark (1.74) was scored by trained expert raters, while o1-preview's AUT responses were scored using an automated AI-based tool (Organisciak et al., 2023). This changes the measurement procedure between the two arms of the comparison, so the z-score of 0.71 does not represent a like-for-like comparison. The authors should score both human and AI responses with the same rubric and same rater type (human or automated) to make the comparison valid.
  4. [§3.1, Tables 1 and 3; §3.7, Table 11] The human benchmarks are selected post hoc in ways that inflate the model's apparent advantage. For EWCTET, Table 3 uses 'Undergraduate Students (Highest after Treatment)' of 13.8, while Table 1 includes other undergraduate results such as 11.51 (Hollis) and 6.6 (Davidson); no rationale is given for choosing the highest available mean. For TOSLS, Table 11 reports a z-score of 1.78 against the 0.85 student cohort, but the text claims the model 'surpass[es] ... even biology experts employed at universities,' and Gormally et al. report a biology-expert mean of 0.91, which is omitted from Table 11. The comparison should be made against a pre-specified or clearly justified benchmark, and all reported human cohorts should be included.
  5. [§2.4, §3] The statistical analysis is not implemented as described. Section 2.4 states that 'a one-sample t-test was used' and that 'results were supplemented with confidence intervals and effect sizes,' but no t-statistics, confidence intervals, or effect sizes appear anywhere in Section 3. The z-scores are computed by comparing the model's 10-trial mean to the human mean scaled by the human SD, which ignores sampling error in the human mean and treats the model's performance as a fixed point. With seven domains and multiple dimensions, no multiple-comparison correction is applied. The authors should report the actual inferential statistics or revise the methods section to describe the descriptive z-score analysis that was performed.
minor comments (6)
  1. [Abstract] Typographical errors: 'TOSLS,, exceeding' has a double comma, and 'o1-preview models's' should be 'o1-preview model's.'
  2. [§3.7] The reported '0.99 ± 0.12' is difficult to interpret for a proportion that is bounded at 1.0. Please report the raw values of the five trials or use a scale on which the standard deviation is meaningful.
  3. [§3.6] In Table 10, the first row is labeled 'Model' but refers to human participants; relabel to 'Human' or 'Human participants' for clarity.
  4. [§3.2, Table 4] The dimension name 'Implemented Challenges' in Table 4 differs from 'Implementation Challenges' in the text (Section 2.2.2); unify the naming.
  5. [§3.5] The RAT human benchmark of 44.12% is not clearly derived from the preceding sentences, which report M = 27.38 and M = 23.80 for high- and low-proficiency bilinguals on a Chinese RAT. Clarify which human result is the benchmark and from which scale it comes.
  6. [§6] The Data availability statement says data are 'available within the article,' but the paper does not include prompts, raw model responses, or scoring scripts. Consider providing a repository with these materials for reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

Benchmark comparison is self-contained; no circular derivation identified.

full rationale

The paper is an empirical benchmark comparison, not a derivation. Each reported result (EWCTET, Lake Urmia Vignette, Computational Thinking Skills, Merk and Chen data literacy tests, AUT/RAT, LogiQA, and TOSLS) compares an external test score or accuracy rate achieved by o1-preview against published human data. No model parameter is fitted to a subset of the benchmark and then used to predict a closely related quantity; no target result is defined in terms of an input. The only self-citations appear in a list of references supporting the general importance of higher-order thinking domains (refs. [2,3]), and they are not load-bearing for any specific empirical claim. The AUT originality scoring asymmetry—an AI-based tool for the model versus expert raters for humans—is a measurement-comparability limitation, not a circular reduction. Internal inconsistencies such as the 'five out of seven' claim conflicting with Table 5, and possible benchmark contamination, are validity concerns outside the circularity framework. No circular step can be exhibited from the paper's own equations or definitions.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The comparison rests on unstated assumptions about score comparability. The most consequential are that o1-preview did not memorize public benchmark items, that self-report instruments measure the same construct for AI as for humans, and that human means from different studies can be pooled as a single baseline.

assumptions (5)
  • standard math One-sample z-scores using the human SD as the denominator are valid significance tests for comparing a single AI mean to a human mean.
    Section 2.4 says a one-sample t-test was used, but the reported z-scores use the human SD and treat the AI as a single observation; they do not incorporate the AI trial count or sampling error, so 'significantly outperforming' is not established.
  • domain assumption Human benchmark means from different studies, years, and populations can serve as a single comparison distribution without meta-analytic correction.
    Sections 2.1 and 3.1 to 3.7 pool means from Hatcher, Hollis, Davis, Merk, Chen, Urban, Xia, Gormally, and others; these samples differ in educational level, language, intervention, and study conditions.
  • domain assumption A self-report Likert instrument (Korkmaz CT Skills) measures the same computational thinking construct when completed by a language model as when completed by humans.
    Section 2.2.3 and Table 5 use the CT Skills scale, which asks humans to rate their own perceived skills; asking o1 to produce Likert responses and comparing them to human self-reports conflates self-perception with performance.
  • domain assumption o1-preview's training data did not include the public benchmark items in a usable form.
    Sections 2.2.6, 2.2.7, 3.6, and 3.7 use LogiQA and TOSLS, both public and widely cited; the paper does not test for or discuss contamination.
  • domain assumption Converting a visual TOSLS item into text with GPT-4o preserves the measured scientific reasoning construct.
    Section 3.7 modifies the graph-based Item 2 into text; the paper itself notes this modality shift may have contributed to errors, so the modified item is not the same instrument.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can OpenAI o1 outperform humans in higher-order cognitive thinking?." pith.science (2026). https://pith.science/paper/GJ5ZXOBD

@misc{pith2026241205753,
  author       = {Pith},
  title        = {Pith review of: Can OpenAI o1 outperform humans in higher-order cognitive thinking?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GJ5ZXOBD}},
  note         = {Machine review of arXiv:2412.05753}
}
read the original abstract

This study evaluates the performance of OpenAI's o1-preview model in higher-order cognitive domains, including critical thinking, systematic thinking, computational thinking, data literacy, creative thinking, logical reasoning, and scientific reasoning. Using established benchmarks, we compared the o1-preview models's performance to human participants from diverse educational levels. o1-preview achieved a mean score of 24.33 on the Ennis-Weir Critical Thinking Essay Test (EWCTET), surpassing undergraduate (13.8) and postgraduate (18.39) participants (z = 1.60 and 0.90, respectively). In systematic thinking, it scored 46.1, SD = 4.12 on the Lake Urmia Vignette, significantly outperforming the human mean (20.08, SD = 8.13, z = 3.20). For data literacy, o1-preview scored 8.60, SD = 0.70 on Merk et al.'s "Use Data" dimension, compared to the human post-test mean of 4.17, SD = 2.02 (z = 2.19). On creative thinking tasks, the model achieved originality scores of 2.98, SD = 0.73, higher than the human mean of 1.74 (z = 0.71). In logical reasoning (LogiQA), it outperformed humans with average 90%, SD = 10% accuracy versus 86%, SD = 6.5% (z = 0.62). For scientific reasoning, it achieved near-perfect performance (mean = 0.99, SD = 0.12) on the TOSLS,, exceeding the highest human scores of 0.85, SD = 0.13 (z = 1.78). While o1-preview excelled in structured tasks, it showed limitations in problem-solving and adaptive reasoning. These results demonstrate the potential of AI to complement education in structured assessments but highlight the need for ethical oversight and refinement for broader applications.

Figures

Figures reproduced from arXiv: 2412.05753 by the authors.

Figure 1
Figure 1. for an overview of its performance) [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Human-Centered Design for AI-based Automatically Generated Assessment Reports: A Systematic Review

    cs.HC 2024-12 conditional novelty 6.0 of 10

    Most K-12 STEM automatic assessment reports underuse low-cognitive-load presentation formats such as text and visual aids, according to a coding review and latent class analysis of 29 systems.

Reference graph

Works this paper leans on

69 extracted references · 67 canonical work pages · cited by 1 Pith paper

  1. [1]

    Introducing OpenAI o1-preview

    OpenAI. Introducing OpenAI o1-preview. https://openaicom/index/introducing-openai-o1-preview/. 2024

  2. [2]

    Can generative AI and ChatGPT outperform humans on cognitive-demanding problem-solving tasks in science? Science & Education

    Zhai X, Nyaaba M, Ma W. Can generative AI and ChatGPT outperform humans on cognitive-demanding problem-solving tasks in science? Science & Education. 2024; p. 1–22

  3. [3]

    Artificial intelligence in education research during 2013–2023: A review based on bibliometric analysis

    Guo S, Zheng Y, Zhai X. Artificial intelligence in education research during 2013–2023: A review based on bibliometric analysis. Education and Information Technologies. 2024; p. 1–23

  4. [4]

    Defining higher order thinking

    Lewis A, Smith D. Defining higher order thinking. Theory into practice. 1993;32(3):131–137

  5. [5]

    Skills for the 21st Century: teaching higher-order thinking

    Collins R. Skills for the 21st Century: teaching higher-order thinking. Curriculum & Leadership Journal. 2014;12(14):1–8

  6. [6]

    ChatGPT user experience: Implications for education

    Zhai X. ChatGPT user experience: Implications for education. Available at SSRN 4312418. 2022

  7. [7]

    Evaluation of OpenAI o1: Opportunities and Challenges of AGI; 2024

    Zhong T, Liu Z, Pan Y, Zhang Y, Zhou Y, Liang S, et al.. Evaluation of OpenAI o1: Opportunities and Challenges of AGI; 2024. Available from: https://arxiv.org/abs/2409.18486

  8. [8]

    Fast Analysis of the OpenAI O1-Preview Model in Solving Random K-SAT Problem: Does the LLM Solve the Problem Itself or Call an External SAT Solver? arXiv preprint arXiv:240911232

    Marino R. Fast Analysis of the OpenAI O1-Preview Model in Solving Random K-SAT Problem: Does the LLM Solve the Problem Itself or Call an External SAT Solver? arXiv preprint arXiv:240911232. 2024

Show all 69 references
  1. [9]

    Can GPT-O1 Kill All Bugs? arXiv preprint arXiv:240910033

    Hu H, Shang Y, Xu G, He C, Zhang Q. Can GPT-O1 Kill All Bugs? arXiv preprint arXiv:240910033. 2024

  2. [10]

    Let’s Verify Step by Step

    Lightman H, Kosaraju V, Burda Y, Edwards H, Baker B, Lee T, et al. Let’s Verify Step by Step. arXiv preprint arXiv:230520050. 2023

  3. [11]

    Self-Reflection in LLM Agents: Effects on Problem-Solving Performance

    Renze M, Guven E. Self-Reflection in LLM Agents: Effects on Problem-Solving Performance. arXiv preprint arXiv:240506682. 2024

  4. [12]

    Relevant or Random: Can LLMs Truly Perform Analogical Reasoning? arXiv preprint arXiv:240412728

    Qin C, Xia W, Wang T, Jiao F, Hu Y, Ding B, et al. Relevant or Random: Can LLMs Truly Perform Analogical Reasoning? arXiv preprint arXiv:240412728. 2024

  5. [13]

    Semantic Structure-Mapping in LLM and Human Analogical Reasoning

    Musker S, Duchnowski A, Milli` ere R, Pavlick E. Semantic Structure-Mapping in LLM and Human Analogical Reasoning. arXiv preprint arXiv:240613803. 2024

  6. [14]

    Exploring collaboration mechanisms for llm agents: A social psychology view

    Zhang J, Xu X, Deng S. Exploring collaboration mechanisms for llm agents: A social psychology view. arXiv preprint arXiv:231002124. 2023

  7. [15]

    Industry 6.0: New Generation of Industry driven by Generative AI and Swarm of Heterogeneous Robots

    Lykov A, Cabrera MA, Konenkov M, Serpiva V, Gbagbe KF, Alabbas A, et al. Industry 6.0: New Generation of Industry driven by Generative AI and Swarm of Heterogeneous Robots. arXiv preprint arXiv:240910106. 2024

  8. [16]

    Enhancing LLM Problem Solving with REAP: Reflection, Explicit Problem Deconstruction, and Advanced Prompting

    Lingo R, Arroyo M, Chhajer R. Enhancing LLM Problem Solving with REAP: Reflection, Explicit Problem Deconstruction, and Advanced Prompting. arXiv preprint arXiv:240909415. 2024;. December 10, 2024 20/24

  9. [17]

    Ethical Alignment of LLMs in Healthcare: Does GPT-o1 Adopt a Deontological or Utilitarian Approach? medRxiv

    Sorin V, Glicksberg BS, Korfiatis P, Nadkarni GN, Klang E. Ethical Alignment of LLMs in Healthcare: Does GPT-o1 Adopt a Deontological or Utilitarian Approach? medRxiv. 2024

  10. [18]

    Cornell Critical Thinking Tests Level X & Level Z Manual

    Ennis R. Cornell Critical Thinking Tests Level X & Level Z Manual. Midwest Publications; 1985

  11. [19]

    The Impact of Teaching Critical Thinking on Iranian Students’ Writing Performance and Their Critical Thinking Dispositions

    Taghinezhad A, Riasati MJ, Rassaei E, Behjat F. The Impact of Teaching Critical Thinking on Iranian Students’ Writing Performance and Their Critical Thinking Dispositions. Brain-Broad Research in Artificial Intelligence and Neuroscience. 2018;9(SI):64–80

  12. [20]

    The Ennis-Weir Critical Thinking Essay Test: An Instrument for Testing and Teaching

    Werner P. The Ennis-Weir Critical Thinking Essay Test: An Instrument for Testing and Teaching. Journal of Reading. 1991;34(6):494–495

  13. [21]

    Critical-Inquiry-Based-Learning: Model of Learning to Promote Critical Thinking Ability of Pre-service Teachers

    Prayogi S, Yuanita L, Wasis. Critical-Inquiry-Based-Learning: Model of Learning to Promote Critical Thinking Ability of Pre-service Teachers. In: Mathematics, Informatics, Science and Education International Conference (MISEIC). vol. 947 of Journal of Physics Conference Series; 2018

  14. [22]

    The Relationship Between Critical Thinking Skills and In-Class Questioning Behaviours of English Language Teaching Students

    Seker H, Komur S. The Relationship Between Critical Thinking Skills and In-Class Questioning Behaviours of English Language Teaching Students. European Journal of Teacher Education. 2008;31(4):389–402. doi:10.1080/02619760802420784

  15. [23]

    Stand-Alone Versus Integrated Critical Thinking Courses

    Hatcher DL. Stand-Alone Versus Integrated Critical Thinking Courses. The Journal of General Education. 2006;55(3-4):247–272. doi:10.2307/27798054

  16. [24]

    Assessing EFL Student Progress in Critical Thinking with the Ennis-Weir Critical Thinking Essay Test

    Davidson BW, Dunham RL. Assessing EFL Student Progress in Critical Thinking with the Ennis-Weir Critical Thinking Essay Test. In: Annual International Conference of the Japan Association for Language Teaching, 21st, Nagoya, Japan and International Conference on Critical Thinki...

  17. [25]

    Validity and Reliability Testing of the International Critical Thinking Essay Test form A (ICTET-A)

    Hollis H, Rachitskiy M, van der Leer L, Elder L. Validity and Reliability Testing of the International Critical Thinking Essay Test form A (ICTET-A). Psychological Reports. in press

  18. [26]

    Systems thinking assessments: Approaches that examine engagement in systems thinking

    Dugan KE, Mosyjowski EA, Daly SR, Lattuca LR. Systems thinking assessments: Approaches that examine engagement in systems thinking. In: 2021 ASEE Virtual Annual Conference Content Access; 2021

  19. [27]

    Investigating student approaches to scenario-based assessments of systems thinking

    Norris MB, Grohs JR, Knight DB. Investigating student approaches to scenario-based assessments of systems thinking. In: Frontiers in Education. vol. 7. Frontiers Media SA; 2022. p. 1055403

  20. [28]

    Developing and Validating a Biological System Thinking Test for Middle School Students

    Li R, Li G. Developing and Validating a Biological System Thinking Test for Middle School Students. International Journal of Science and Mathematics Education. 2024; p. 1–21

  21. [29]

    Assessing systems thinking: A tool to measure complex reasoning through ill-structured problems

    Grohs JR, Kirk GR, Soledad MM, Knight DB. Assessing systems thinking: A tool to measure complex reasoning through ill-structured problems. Thinking Skills and Creativity. 2018;28:110–130

  22. [30]

    The Lake Urmia vignette: a tool to assess understanding of complexity in socio-environmental systems

    Davis K, Ghaffarzadegan N, Grohs J, Grote D, Hosseinichimeh N, Knight D, et al. The Lake Urmia vignette: a tool to assess understanding of complexity in socio-environmental systems. System Dynamics Review. 2020;36(2):191–222. December 10, 2024 21/24

  23. [31]

    A validity and reliability study of the computational thinking scales (CTS)

    Korkmaz ¨O, C ¸ akir R,¨Ozden MY. A validity and reliability study of the computational thinking scales (CTS). Computers in human behavior. 2017;72:558–569

  24. [32]

    Assessing computational thinking: Development and validation of the algorithmic thinking test for adults

    Lafuente Mart ´ ınez M, L´ evˆ eque O, Ben ´ ıtez I, Hardebolle C, Zufferey JD. Assessing computational thinking: Development and validation of the algorithmic thinking test for adults. Journal of Educational Computing Research. 2022;60(6):1436–1463

  25. [33]

    Data literacy for researchers and data librarians

    Koltay T. Data literacy for researchers and data librarians. Journal of Librarianship and Information Science. 2017;49(1):3–14

  26. [34]

    Creating an understanding of data literacy for a data-driven society

    Wolff A, Gooch D, Montaner JJC, Rashid U, Kortuem G. Creating an understanding of data literacy for a data-driven society. The Journal of Community Informatics. 2016;12(3)

  27. [35]

    Data literacy assessments: A systematic literature review

    Cui Y, Chen F, Lutsyk A, Leighton JP, Cutumisu M. Data literacy assessments: A systematic literature review. Assessment in Education: Principles, Policy & Practice. 2023;30(1):76–96

  28. [36]

    Effects of an asynchronous online data literacy intervention on pre-service and in-service educators’ beliefs, self-efficacy, and practices

    Reeves TD, Chiang JL. Effects of an asynchronous online data literacy intervention on pre-service and in-service educators’ beliefs, self-efficacy, and practices. Computers & Education. 2019;136:13–33

  29. [37]

    Fostering aspects of pre-service teachers’ data literacy: Results of a randomized controlled trial

    Merk S, Poindl S, Wurster S, Bohl T. Fostering aspects of pre-service teachers’ data literacy: Results of a randomized controlled trial. Teaching and Teacher Education. 2020;91:103043

  30. [38]

    Validating a novel digital performance-based assessment of data literacy: Psychometric and eye-tracking analyses

    Chen F, Cui Y, Lutsyk-King A, Gao Y, Liu X, Cutumisu M, et al. Validating a novel digital performance-based assessment of data literacy: Psychometric and eye-tracking analyses. Education and Information Technologies. 2024;29(8):9417–9444

  31. [39]

    The (dis) pleasures of creativity: Spontaneous eye blink rate during divergent and convergent thinking depends on individual differences in positive and negative affect

    De Rooij A, Vromans RD. The (dis) pleasures of creativity: Spontaneous eye blink rate during divergent and convergent thinking depends on individual differences in positive and negative affect. The Journal of Creative Behavior. 2020;54(2):436–452

  32. [40]

    How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment

    Ashkinaze J, Mendelsohn J, Qiwei L, Budak C, Gilbert E. How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment. arXiv preprint arXiv:240113481. 2024

  33. [41]

    The nature of human intelligence

    Guilford JP. The nature of human intelligence. New York: Macgraw Hill. 1967

  34. [42]

    Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models

    Organisciak P, Acar S, Dumas D, Berthiaume K. Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models. Thinking Skills and Creativity. 2023;49:101356

  35. [43]

    Applying automated originality scoring to the verbal form of Torrance tests of creative thinking

    Acar S, Berthiaume K, Grajzel K, Dumas D, Flemister C, Organisciak P. Applying automated originality scoring to the verbal form of Torrance tests of creative thinking. Gifted Child Quarterly. 2023;67(1):3–17

  36. [44]

    A distracted muse: The positive effect of dual-task distraction on creative potential

    Collins MJD. A distracted muse: The positive effect of dual-task distraction on creative potential. Creativity Research Journal. 2020;32(4):357–367

  37. [45]

    From book to bludgeon: A closer look at unsolicited malevolent responses on the alternate uses task

    Dumas DG, Strickland AL. From book to bludgeon: A closer look at unsolicited malevolent responses on the alternate uses task. Creativity Research Journal. 2018;30(4):439–450. December 10, 2024 22/24

  38. [46]

    Fixation, flexibility, and forgetting during alternate uses tasks

    George T, Wiley J. Fixation, flexibility, and forgetting during alternate uses tasks. Psychology of Aesthetics, Creativity, and the Arts. 2019;13(3):305

  39. [47]

    The associative basis of the creative process

    Mednick S. The associative basis of the creative process. Psychological review. 1962;69(3):220

  40. [48]

    An Objective Measuring Tool of Creativity: The Development of Chinese Remote Association Test

    LI Lm, LUO Ll, LIU W. An Objective Measuring Tool of Creativity: The Development of Chinese Remote Association Test. Journal of Northeastern University (Social Science). 2015;17(1):19

  41. [49]

    Normative data for 102 Spanish remote associate problems and age-related differences in performance

    Pel´ aez-Alfonso JL, Pelegrina S, Lechuga MT. Normative data for 102 Spanish remote associate problems and age-related differences in performance. Psicol´ ogica Journal. 2020;41(1):39–65

  42. [50]

    LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

    Liu Z, Chen Y, Liu Z, Fu J, Cheng X, Li Y, et al. LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning. arXiv preprint. 2020;doi:10.48550/arXiv.2007.08124

  43. [51]

    A concise introduction to logic

    Hurley PJ. A concise introduction to logic. Nelson Education; 2014

  44. [52]

    Introduction to Categories and Categorical Logic

    Abramsky S, Tzevelekos N. Introduction to Categories and Categorical Logic. In: Coecke B, editor. New Structures for Physics. vol. 813 of Lecture Notes in Physics. Springer-Verlag; 2011. p. 3–94

  45. [53]

    Roberta: A robustly optimized bert pretraining approach

    Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, et al. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:190711692. 2019

  46. [54]

    Measuring scientific reasoning–a review of test instruments

    Opitz A, Heene M, Fischer F. Measuring scientific reasoning–a review of test instruments. Educational Research and Evaluation. 2017;23(3-4):78–101

  47. [55]

    Developing a test of scientific literacy skills (TOSLS): Measuring undergraduates’ evaluation of scientific information and arguments

    Gormally C, Brickman P, Lutz M. Developing a test of scientific literacy skills (TOSLS): Measuring undergraduates’ evaluation of scientific information and arguments. CBE—Life Sciences Education. 2012;11(4):364–377

  48. [56]

    Milestones: Cognitive

    Chen Z. Milestones: Cognitive. In: Benson JB, editor. Encyclopedia of Infant and Early Childhood Development (Second Edition). second edition ed. Oxford: Elsevier; 2020. p. 330–338. Available from: https: //www.sciencedirect.com/science/article/pii/B9780128093245218257

  49. [57]

    The theory of the estimation of test reliability

    Kuder GF, Richardson MW. The theory of the estimation of test reliability. Psychometrika. 1937;2(3):151–160

  50. [58]

    Coefficient alpha and the internal structure of tests

    Cronbach LJ. Coefficient alpha and the internal structure of tests. psychometrika. 1951;16(3):297–334

  51. [59]

    Student performance on the Test of Scientific Literacy Skills (TOSLS) does not change with assignment of a low-stakes grade

    Segarra V A, Hughes NM, Ackerman KM, Grider MH, Lyda T, Vigueira PA. Student performance on the Test of Scientific Literacy Skills (TOSLS) does not change with assignment of a low-stakes grade. BMC research notes. 2018;11:1–5

  52. [60]

    Test of scientific literacy skills (TOSLS) indicates limited scientific thinking gains as a result of science and mathematics general education

    Propsom PM, Tobin WM, Roberts JR. Test of scientific literacy skills (TOSLS) indicates limited scientific thinking gains as a result of science and mathematics general education. Interdisciplinary Faculty Scholarship. 2023

  53. [61]

    Enhancement of students’ biological literacy and critical thinking of biology through socio-biological case-based learning

    Suwono H, Pratiwi H, Susanto H, Susilo H. Enhancement of students’ biological literacy and critical thinking of biology through socio-biological case-based learning. Jurnal Pendidikan IPA Indonesia. 2017;6(2):213–220. December 10, 2024 23/24

  54. [62]

    Comparing self-report assessments and scenario-based assessments of systems thinking competence

    Davis KA, Grote D, Mahmoudi H, Perry L, Ghaffarzadegan N, Grohs J, et al. Comparing self-report assessments and scenario-based assessments of systems thinking competence. Journal of Science Education and Technology. 2023;32(6):793–813

  55. [63]

    What influences computational thinking? A theoretical and empirical study based on the influence of learning engagement on computational thinking in higher education

    Liu S, Peng C, Srivastava G. What influences computational thinking? A theoretical and empirical study based on the influence of learning engagement on computational thinking in higher education. Computer Applications in Engineering Education. 2023;31(6):1690–1704

  56. [64]

    STEM professional development program for gifted education teachers: STEM lesson plan design competence, self-efficacy, computational thinking and entrepreneurial skills

    S ¸ahin E, Sarı U, S ¸en¨OF. STEM professional development program for gifted education teachers: STEM lesson plan design competence, self-efficacy, computational thinking and entrepreneurial skills. Thinking Skills and Creativity. 2024;51:101439

  57. [65]

    Scaffolding Computational Thinking with ChatGPT

    Liao J, Zhong L, Zhe L, Xu H, Liu M, Xie T. Scaffolding Computational Thinking with ChatGPT. IEEE Transactions on Learning Technologies. 2024

  58. [66]

    ChatGPT improves creative problem-solving performance in university students: An experimental study

    Urban M, Dˇ echtˇ erenko F, Lukavsk` y J, Hrabalov´ a V, Svacha F, Brom C, et al. ChatGPT improves creative problem-solving performance in university students: An experimental study. Computers & Education. 2024;215:105031

  59. [67]

    Bilingualism and creativity: Benefits from cognitive inhibition and cognitive flexibility

    Xia T, An Y, Guo J. Bilingualism and creativity: Benefits from cognitive inhibition and cognitive flexibility. Frontiers in Psychology. 2022;13:1016777

  60. [68]

    Exploring the effect of red and blue on cognitive task performances

    Xia T, Song L, Wang TT, Tan L, Mo L. Exploring the effect of red and blue on cognitive task performances. Frontiers in Psychology. 2016;7:784

  61. [69]

    A quantitative study on the scientific literacy skills of prospective biology teachers

    Firdaus L, Ibrohim I, Lestari SR, Masiah M, Primawati SN, Hunaepi H. A quantitative study on the scientific literacy skills of prospective biology teachers. Jurnal Penelitian Pendidikan IPA. 2023;9(1):80–86. December 10, 2024 24/24

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.