REVIEW 5 major objections 6 minor 1 cited by
Can OpenAI o1 outperform humans in higher-order cognitive thinking?
T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that OpenAI's o1-preview model outperforms human comparison groups on five of seven established tests of higher-order thinking, with its main weakness showing up in problem-solving.
desk verdict The paper's own Table 5 contradicts its five-of-seven claim, but the assembled battery and the large systematic-thinking effects make it worth a referee's time. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a suite of published, normed assessment instruments, each administered to o1-preview as text prompts and scored with the same rubrics used for human test-takers: EWCTET, the Lake Urmia Vignette, the Computational Thinking Skills scale, the ATTA, Merk et al.'s and Chen et al.'s data literacy tests, the AUT and RAT, LogiQA, and TOSLS. The model's scores are then standardized against the human means and standard deviations reported in the source studies, with z-scores and one-sample t-tests used to express how far above or below the human distribution the model sits. The comparison only works if the instruments measure what they claim to measure, the human norms are representative, and the model encounters the items as novel prompts rather than as memorized training text.
What would settle it
Give o1-preview newly written parallel versions of the same seven assessments—same constructs and formats, but items that cannot be in its training data—and compare its scores on the originals; a large drop, or the model's ability to recite original items on demand, would show that the reported human-beating performance came from contamination rather than higher-order reasoning. A concrete spot check is the TOSLS Item 2 graph-selection question, which was converted from visual plots to text for the model: present the same item with novel graphs, and see whether the near-perfect score survives.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that a general-purpose reasoning model, o1-preview, can be prompted with the items of standard cognitive assessments and score at or above the human means those instruments were designed to spread out. The model scored 24.33 on the Ennis-Weir critical thinking essay test versus 18.39 for postgraduates; 46.10 on the Lake Urmia Vignette total versus 20.08 for undergraduates; 8.60 on Merk et al.'s "Use Data" dimension versus a 4.17 post-test mean; near-perfect 0.99 on TOSLS versus 0.85 for the best student cohort; 90% accuracy on LogiQA versus 86% for humans; and a perfect 20/20 on the ATTA versus 14.63 for experts. The paper claims this as outperformance in five of seven domains, while flagging that the model's problem-solving score on the computational-thinking scale was far below humans and that visual TOSLS items had to be converted to text, which may have affected results.
Load-bearing premise
The load-bearing premise is that o1-preview never memorized the publicly available test questions during training, so its high scores reflect reasoning rather than recall of answers.
Editorial extensions
If this is right
- If the reported comparisons hold, o1-preview can already outperform typical university students on structured critical-thinking, data-literacy, and scientific-reasoning assessments, which suggests AI tools could take over routine scoring and tutoring of those skills.
- The model's near-zero score on the computational-thinking problem-solving subscale is a direct counterexample to the idea that LLMs have generalized reasoning, so the paper implies that open-ended, ill-structured problem-solving should remain a human responsibility in AI-assisted classrooms.
- Because the model saturates TOSLS and nearly saturates LogiQA, the paper's results imply that widely used tests of 'higher-order thinking' may need redesign if they are to keep measuring human cognitive development rather than machine pattern recognition.
- The authors' recommendation that AI be used as a supplement, not a replacement, follows directly from the observed mix of very high scores on structured items and very low scores on problem-solving items.
Reading between the lines
- Beyond the paper, the same logic implies that any LLM trained on public exam corpora will tend to saturate standardized reasoning tests, so the meaningful next comparison is on fresh, non-public items rather than on legacy benchmarks.
- A test the authors did not run would be to administer paraphrased and answer-permuted versions of the same instruments; a large score drop on those versions would indicate the model is exploiting surface patterns rather than performing the reasoning the tests claim to measure.
- The paper compares the model against human means drawn from different studies, cohorts, and years, so a direct head-to-head study with the same human sample and the same items would be needed to confirm the claimed five-of-seven superiority.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript evaluates OpenAI's o1-preview model on seven purported higher-order thinking domains — critical thinking, systematic thinking, computational thinking, data literacy, creative thinking, logical reasoning, and scientific reasoning — using established instruments (EWCTET, Village of Abeesee/LUV, CT Skills/ATTA, Merk/Chen data literacy tests, AUT/RAT, LogiQA, TOSLS). Human performance is taken from previously published studies, and the model is run for 10 trials per instrument. The authors report z-scores for model-versus-human comparisons and claim in the abstract and introduction that o1-preview outperforms human experts in five of seven domains. The paper concludes that AI can complement education in structured assessments but needs human oversight.
Significance. If the headline claim were supported, this would be a notable contribution to the debate on whether LLMs can match or exceed human performance on structured higher-order cognition assessments. The study's strengths include the use of multiple established instruments, transparent reporting of per-dimension means and z-scores in Tables 4–11, and an explicit acknowledgment of the model's weakness in problem-solving. However, the paper's central assertion is internally contradicted by its own reported results, and several load-bearing methodological choices (human-benchmark selection, contamination, and differential scoring across arms) undermine the comparability of the AI-human comparisons. As it stands, the evidence does not support the claimed five-of-seven outperformance.
major comments (5)
- [§1, Abstract, Table 5] The claim that o1-preview 'outperforms human experts in five out of seven domains, including ... computational thinking' is contradicted by Table 5, which shows that for Computational Thinking Skills the model's overall z-score is -0.15 (mean 3.84 vs. human 3.92) and the Problem-Solving z-score is -4.25 (mean 1.00 vs. human 3.68). Thus computational thinking is not outperformed overall. Moreover, the five named domains omit critical thinking (EWCTET z = 1.60 and 0.90 in Table 3) and logical reasoning (LogiQA z = 0.62 in Table 10), both of which have positive z-scores; if those count, the correct count would be six of seven, not five. This suggests the 'five of seven' list is not supported by the paper's own data and must be corrected or explicitly justified.
- [§2.2.6, §2.2.7, §3.6, §3.7] The paper never addresses the possibility that o1-preview's training data included the benchmark items. LogiQA, TOSLS, AUT, RAT, and EWCTET are all public, widely used evaluation instruments, and many are standard in LLM benchmarks. If the model has seen these items or their solutions, high scores may reflect memorization rather than reasoning. This is a load-bearing premise for the claim that o1-preview 'outperforms human experts' in higher-order thinking. The authors should either provide evidence of non-contamination, test on private held-out variants, or explicitly discuss this limitation and temper the conclusions accordingly.
- [§3.5, Table 9] The human AUT originality benchmark (1.74) was scored by trained expert raters, while o1-preview's AUT responses were scored using an automated AI-based tool (Organisciak et al., 2023). This changes the measurement procedure between the two arms of the comparison, so the z-score of 0.71 does not represent a like-for-like comparison. The authors should score both human and AI responses with the same rubric and same rater type (human or automated) to make the comparison valid.
- [§3.1, Tables 1 and 3; §3.7, Table 11] The human benchmarks are selected post hoc in ways that inflate the model's apparent advantage. For EWCTET, Table 3 uses 'Undergraduate Students (Highest after Treatment)' of 13.8, while Table 1 includes other undergraduate results such as 11.51 (Hollis) and 6.6 (Davidson); no rationale is given for choosing the highest available mean. For TOSLS, Table 11 reports a z-score of 1.78 against the 0.85 student cohort, but the text claims the model 'surpass[es] ... even biology experts employed at universities,' and Gormally et al. report a biology-expert mean of 0.91, which is omitted from Table 11. The comparison should be made against a pre-specified or clearly justified benchmark, and all reported human cohorts should be included.
- [§2.4, §3] The statistical analysis is not implemented as described. Section 2.4 states that 'a one-sample t-test was used' and that 'results were supplemented with confidence intervals and effect sizes,' but no t-statistics, confidence intervals, or effect sizes appear anywhere in Section 3. The z-scores are computed by comparing the model's 10-trial mean to the human mean scaled by the human SD, which ignores sampling error in the human mean and treats the model's performance as a fixed point. With seven domains and multiple dimensions, no multiple-comparison correction is applied. The authors should report the actual inferential statistics or revise the methods section to describe the descriptive z-score analysis that was performed.
minor comments (6)
- [Abstract] Typographical errors: 'TOSLS,, exceeding' has a double comma, and 'o1-preview models's' should be 'o1-preview model's.'
- [§3.7] The reported '0.99 ± 0.12' is difficult to interpret for a proportion that is bounded at 1.0. Please report the raw values of the five trials or use a scale on which the standard deviation is meaningful.
- [§3.6] In Table 10, the first row is labeled 'Model' but refers to human participants; relabel to 'Human' or 'Human participants' for clarity.
- [§3.2, Table 4] The dimension name 'Implemented Challenges' in Table 4 differs from 'Implementation Challenges' in the text (Section 2.2.2); unify the naming.
- [§3.5] The RAT human benchmark of 44.12% is not clearly derived from the preceding sentences, which report M = 27.38 and M = 23.80 for high- and low-proficiency bilinguals on a Chinese RAT. Clarify which human result is the benchmark and from which scale it comes.
- [§6] The Data availability statement says data are 'available within the article,' but the paper does not include prompts, raw model responses, or scoring scripts. Consider providing a repository with these materials for reproducibility.
Circularity Check
Benchmark comparison is self-contained; no circular derivation identified.
full rationale
The paper is an empirical benchmark comparison, not a derivation. Each reported result (EWCTET, Lake Urmia Vignette, Computational Thinking Skills, Merk and Chen data literacy tests, AUT/RAT, LogiQA, and TOSLS) compares an external test score or accuracy rate achieved by o1-preview against published human data. No model parameter is fitted to a subset of the benchmark and then used to predict a closely related quantity; no target result is defined in terms of an input. The only self-citations appear in a list of references supporting the general importance of higher-order thinking domains (refs. [2,3]), and they are not load-bearing for any specific empirical claim. The AUT originality scoring asymmetry—an AI-based tool for the model versus expert raters for humans—is a measurement-comparability limitation, not a circular reduction. Internal inconsistencies such as the 'five out of seven' claim conflicting with Table 5, and possible benchmark contamination, are validity concerns outside the circularity framework. No circular step can be exhibited from the paper's own equations or definitions.
Assumptions & free parameters
assumptions (5)
- standard math One-sample z-scores using the human SD as the denominator are valid significance tests for comparing a single AI mean to a human mean.
- domain assumption Human benchmark means from different studies, years, and populations can serve as a single comparison distribution without meta-analytic correction.
- domain assumption A self-report Likert instrument (Korkmaz CT Skills) measures the same computational thinking construct when completed by a language model as when completed by humans.
- domain assumption o1-preview's training data did not include the public benchmark items in a usable form.
- domain assumption Converting a visual TOSLS item into text with GPT-4o preserves the measured scientific reasoning construct.
Cite this review
Pith. "Pith review of Can OpenAI o1 outperform humans in higher-order cognitive thinking?." pith.science (2026). https://pith.science/paper/GJ5ZXOBD
@misc{pith2026241205753,
author = {Pith},
title = {Pith review of: Can OpenAI o1 outperform humans in higher-order cognitive thinking?},
year = {2026},
howpublished = {\url{https://pith.science/paper/GJ5ZXOBD}},
note = {Machine review of arXiv:2412.05753}
}
read the original abstract
This study evaluates the performance of OpenAI's o1-preview model in higher-order cognitive domains, including critical thinking, systematic thinking, computational thinking, data literacy, creative thinking, logical reasoning, and scientific reasoning. Using established benchmarks, we compared the o1-preview models's performance to human participants from diverse educational levels. o1-preview achieved a mean score of 24.33 on the Ennis-Weir Critical Thinking Essay Test (EWCTET), surpassing undergraduate (13.8) and postgraduate (18.39) participants (z = 1.60 and 0.90, respectively). In systematic thinking, it scored 46.1, SD = 4.12 on the Lake Urmia Vignette, significantly outperforming the human mean (20.08, SD = 8.13, z = 3.20). For data literacy, o1-preview scored 8.60, SD = 0.70 on Merk et al.'s "Use Data" dimension, compared to the human post-test mean of 4.17, SD = 2.02 (z = 2.19). On creative thinking tasks, the model achieved originality scores of 2.98, SD = 0.73, higher than the human mean of 1.74 (z = 0.71). In logical reasoning (LogiQA), it outperformed humans with average 90%, SD = 10% accuracy versus 86%, SD = 6.5% (z = 0.62). For scientific reasoning, it achieved near-perfect performance (mean = 0.99, SD = 0.12) on the TOSLS,, exceeding the highest human scores of 0.85, SD = 0.13 (z = 1.78). While o1-preview excelled in structured tasks, it showed limitations in problem-solving and adaptive reasoning. These results demonstrate the potential of AI to complement education in structured assessments but highlight the need for ethical oversight and refinement for broader applications.
Figures
Forward citations
Cited by 1 Pith paper
-
Human-Centered Design for AI-based Automatically Generated Assessment Reports: A Systematic Review
Most K-12 STEM automatic assessment reports underuse low-cognitive-load presentation formats such as text and visual aids, according to a coding review and latent class analysis of 29 systems.
Reference graph
Works this paper leans on
-
[1]
OpenAI. Introducing OpenAI o1-preview. https://openaicom/index/introducing-openai-o1-preview/. 2024
work page 2024
-
[2]
Zhai X, Nyaaba M, Ma W. Can generative AI and ChatGPT outperform humans on cognitive-demanding problem-solving tasks in science? Science & Education. 2024; p. 1–22
work page 2024
-
[3]
Guo S, Zheng Y, Zhai X. Artificial intelligence in education research during 2013–2023: A review based on bibliometric analysis. Education and Information Technologies. 2024; p. 1–23
work page 2013
-
[4]
Defining higher order thinking
Lewis A, Smith D. Defining higher order thinking. Theory into practice. 1993;32(3):131–137
work page 1993
-
[5]
Skills for the 21st Century: teaching higher-order thinking
Collins R. Skills for the 21st Century: teaching higher-order thinking. Curriculum & Leadership Journal. 2014;12(14):1–8
work page 2014
-
[6]
ChatGPT user experience: Implications for education
Zhai X. ChatGPT user experience: Implications for education. Available at SSRN 4312418. 2022
work page 2022
-
[7]
Evaluation of OpenAI o1: Opportunities and Challenges of AGI; 2024
Zhong T, Liu Z, Pan Y, Zhang Y, Zhou Y, Liang S, et al.. Evaluation of OpenAI o1: Opportunities and Challenges of AGI; 2024. Available from: https://arxiv.org/abs/2409.18486
arXiv 2024
-
[8]
Marino R. Fast Analysis of the OpenAI O1-Preview Model in Solving Random K-SAT Problem: Does the LLM Solve the Problem Itself or Call an External SAT Solver? arXiv preprint arXiv:240911232. 2024
work page 2024
Show all 69 references
-
[9]
Can GPT-O1 Kill All Bugs? arXiv preprint arXiv:240910033
Hu H, Shang Y, Xu G, He C, Zhang Q. Can GPT-O1 Kill All Bugs? arXiv preprint arXiv:240910033. 2024
2024
-
[10]
Let’s Verify Step by Step
Lightman H, Kosaraju V, Burda Y, Edwards H, Baker B, Lee T, et al. Let’s Verify Step by Step. arXiv preprint arXiv:230520050. 2023
2023
-
[11]
Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
Renze M, Guven E. Self-Reflection in LLM Agents: Effects on Problem-Solving Performance. arXiv preprint arXiv:240506682. 2024
2024
-
[12]
Relevant or Random: Can LLMs Truly Perform Analogical Reasoning? arXiv preprint arXiv:240412728
Qin C, Xia W, Wang T, Jiao F, Hu Y, Ding B, et al. Relevant or Random: Can LLMs Truly Perform Analogical Reasoning? arXiv preprint arXiv:240412728. 2024
2024
-
[13]
Semantic Structure-Mapping in LLM and Human Analogical Reasoning
Musker S, Duchnowski A, Milli` ere R, Pavlick E. Semantic Structure-Mapping in LLM and Human Analogical Reasoning. arXiv preprint arXiv:240613803. 2024
2024
-
[14]
Exploring collaboration mechanisms for llm agents: A social psychology view
Zhang J, Xu X, Deng S. Exploring collaboration mechanisms for llm agents: A social psychology view. arXiv preprint arXiv:231002124. 2023
2023
-
[15]
Industry 6.0: New Generation of Industry driven by Generative AI and Swarm of Heterogeneous Robots
Lykov A, Cabrera MA, Konenkov M, Serpiva V, Gbagbe KF, Alabbas A, et al. Industry 6.0: New Generation of Industry driven by Generative AI and Swarm of Heterogeneous Robots. arXiv preprint arXiv:240910106. 2024
2024
-
[16]
Enhancing LLM Problem Solving with REAP: Reflection, Explicit Problem Deconstruction, and Advanced Prompting
Lingo R, Arroyo M, Chhajer R. Enhancing LLM Problem Solving with REAP: Reflection, Explicit Problem Deconstruction, and Advanced Prompting. arXiv preprint arXiv:240909415. 2024;. December 10, 2024 20/24
2024
-
[17]
Ethical Alignment of LLMs in Healthcare: Does GPT-o1 Adopt a Deontological or Utilitarian Approach? medRxiv
Sorin V, Glicksberg BS, Korfiatis P, Nadkarni GN, Klang E. Ethical Alignment of LLMs in Healthcare: Does GPT-o1 Adopt a Deontological or Utilitarian Approach? medRxiv. 2024
2024
-
[18]
Cornell Critical Thinking Tests Level X & Level Z Manual
Ennis R. Cornell Critical Thinking Tests Level X & Level Z Manual. Midwest Publications; 1985
1985
-
[19]
The Impact of Teaching Critical Thinking on Iranian Students’ Writing Performance and Their Critical Thinking Dispositions
Taghinezhad A, Riasati MJ, Rassaei E, Behjat F. The Impact of Teaching Critical Thinking on Iranian Students’ Writing Performance and Their Critical Thinking Dispositions. Brain-Broad Research in Artificial Intelligence and Neuroscience. 2018;9(SI):64–80
2018
-
[20]
The Ennis-Weir Critical Thinking Essay Test: An Instrument for Testing and Teaching
Werner P. The Ennis-Weir Critical Thinking Essay Test: An Instrument for Testing and Teaching. Journal of Reading. 1991;34(6):494–495
1991
-
[21]
Critical-Inquiry-Based-Learning: Model of Learning to Promote Critical Thinking Ability of Pre-service Teachers
Prayogi S, Yuanita L, Wasis. Critical-Inquiry-Based-Learning: Model of Learning to Promote Critical Thinking Ability of Pre-service Teachers. In: Mathematics, Informatics, Science and Education International Conference (MISEIC). vol. 947 of Journal of Physics Conference Series; 2018
2018
-
[22]
The Relationship Between Critical Thinking Skills and In-Class Questioning Behaviours of English Language Teaching Students
Seker H, Komur S. The Relationship Between Critical Thinking Skills and In-Class Questioning Behaviours of English Language Teaching Students. European Journal of Teacher Education. 2008;31(4):389–402. doi:10.1080/02619760802420784
2008 doi
-
[23]
Stand-Alone Versus Integrated Critical Thinking Courses
Hatcher DL. Stand-Alone Versus Integrated Critical Thinking Courses. The Journal of General Education. 2006;55(3-4):247–272. doi:10.2307/27798054
2006 doi
-
[24]
Assessing EFL Student Progress in Critical Thinking with the Ennis-Weir Critical Thinking Essay Test
Davidson BW, Dunham RL. Assessing EFL Student Progress in Critical Thinking with the Ennis-Weir Critical Thinking Essay Test. In: Annual International Conference of the Japan Association for Language Teaching, 21st, Nagoya, Japan and International Conference on Critical Thinki...
1996
-
[25]
Validity and Reliability Testing of the International Critical Thinking Essay Test form A (ICTET-A)
Hollis H, Rachitskiy M, van der Leer L, Elder L. Validity and Reliability Testing of the International Critical Thinking Essay Test form A (ICTET-A). Psychological Reports. in press
-
[26]
Systems thinking assessments: Approaches that examine engagement in systems thinking
Dugan KE, Mosyjowski EA, Daly SR, Lattuca LR. Systems thinking assessments: Approaches that examine engagement in systems thinking. In: 2021 ASEE Virtual Annual Conference Content Access; 2021
2021
-
[27]
Investigating student approaches to scenario-based assessments of systems thinking
Norris MB, Grohs JR, Knight DB. Investigating student approaches to scenario-based assessments of systems thinking. In: Frontiers in Education. vol. 7. Frontiers Media SA; 2022. p. 1055403
2022
-
[28]
Developing and Validating a Biological System Thinking Test for Middle School Students
Li R, Li G. Developing and Validating a Biological System Thinking Test for Middle School Students. International Journal of Science and Mathematics Education. 2024; p. 1–21
2024
-
[29]
Assessing systems thinking: A tool to measure complex reasoning through ill-structured problems
Grohs JR, Kirk GR, Soledad MM, Knight DB. Assessing systems thinking: A tool to measure complex reasoning through ill-structured problems. Thinking Skills and Creativity. 2018;28:110–130
2018
-
[30]
The Lake Urmia vignette: a tool to assess understanding of complexity in socio-environmental systems
Davis K, Ghaffarzadegan N, Grohs J, Grote D, Hosseinichimeh N, Knight D, et al. The Lake Urmia vignette: a tool to assess understanding of complexity in socio-environmental systems. System Dynamics Review. 2020;36(2):191–222. December 10, 2024 21/24
2020
-
[31]
A validity and reliability study of the computational thinking scales (CTS)
Korkmaz ¨O, C ¸ akir R,¨Ozden MY. A validity and reliability study of the computational thinking scales (CTS). Computers in human behavior. 2017;72:558–569
2017
-
[32]
Assessing computational thinking: Development and validation of the algorithmic thinking test for adults
Lafuente Mart ´ ınez M, L´ evˆ eque O, Ben ´ ıtez I, Hardebolle C, Zufferey JD. Assessing computational thinking: Development and validation of the algorithmic thinking test for adults. Journal of Educational Computing Research. 2022;60(6):1436–1463
2022
-
[33]
Data literacy for researchers and data librarians
Koltay T. Data literacy for researchers and data librarians. Journal of Librarianship and Information Science. 2017;49(1):3–14
2017
-
[34]
Creating an understanding of data literacy for a data-driven society
Wolff A, Gooch D, Montaner JJC, Rashid U, Kortuem G. Creating an understanding of data literacy for a data-driven society. The Journal of Community Informatics. 2016;12(3)
2016
-
[35]
Data literacy assessments: A systematic literature review
Cui Y, Chen F, Lutsyk A, Leighton JP, Cutumisu M. Data literacy assessments: A systematic literature review. Assessment in Education: Principles, Policy & Practice. 2023;30(1):76–96
2023
-
[36]
Effects of an asynchronous online data literacy intervention on pre-service and in-service educators’ beliefs, self-efficacy, and practices
Reeves TD, Chiang JL. Effects of an asynchronous online data literacy intervention on pre-service and in-service educators’ beliefs, self-efficacy, and practices. Computers & Education. 2019;136:13–33
2019
-
[37]
Fostering aspects of pre-service teachers’ data literacy: Results of a randomized controlled trial
Merk S, Poindl S, Wurster S, Bohl T. Fostering aspects of pre-service teachers’ data literacy: Results of a randomized controlled trial. Teaching and Teacher Education. 2020;91:103043
2020
-
[38]
Validating a novel digital performance-based assessment of data literacy: Psychometric and eye-tracking analyses
Chen F, Cui Y, Lutsyk-King A, Gao Y, Liu X, Cutumisu M, et al. Validating a novel digital performance-based assessment of data literacy: Psychometric and eye-tracking analyses. Education and Information Technologies. 2024;29(8):9417–9444
2024
-
[39]
The (dis) pleasures of creativity: Spontaneous eye blink rate during divergent and convergent thinking depends on individual differences in positive and negative affect
De Rooij A, Vromans RD. The (dis) pleasures of creativity: Spontaneous eye blink rate during divergent and convergent thinking depends on individual differences in positive and negative affect. The Journal of Creative Behavior. 2020;54(2):436–452
2020
-
[40]
How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment
Ashkinaze J, Mendelsohn J, Qiwei L, Budak C, Gilbert E. How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment. arXiv preprint arXiv:240113481. 2024
2024
-
[41]
The nature of human intelligence
Guilford JP. The nature of human intelligence. New York: Macgraw Hill. 1967
1967
-
[42]
Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models
Organisciak P, Acar S, Dumas D, Berthiaume K. Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models. Thinking Skills and Creativity. 2023;49:101356
2023
-
[43]
Applying automated originality scoring to the verbal form of Torrance tests of creative thinking
Acar S, Berthiaume K, Grajzel K, Dumas D, Flemister C, Organisciak P. Applying automated originality scoring to the verbal form of Torrance tests of creative thinking. Gifted Child Quarterly. 2023;67(1):3–17
2023
-
[44]
A distracted muse: The positive effect of dual-task distraction on creative potential
Collins MJD. A distracted muse: The positive effect of dual-task distraction on creative potential. Creativity Research Journal. 2020;32(4):357–367
2020
-
[45]
From book to bludgeon: A closer look at unsolicited malevolent responses on the alternate uses task
Dumas DG, Strickland AL. From book to bludgeon: A closer look at unsolicited malevolent responses on the alternate uses task. Creativity Research Journal. 2018;30(4):439–450. December 10, 2024 22/24
2018
-
[46]
Fixation, flexibility, and forgetting during alternate uses tasks
George T, Wiley J. Fixation, flexibility, and forgetting during alternate uses tasks. Psychology of Aesthetics, Creativity, and the Arts. 2019;13(3):305
2019
-
[47]
The associative basis of the creative process
Mednick S. The associative basis of the creative process. Psychological review. 1962;69(3):220
1962
-
[48]
An Objective Measuring Tool of Creativity: The Development of Chinese Remote Association Test
LI Lm, LUO Ll, LIU W. An Objective Measuring Tool of Creativity: The Development of Chinese Remote Association Test. Journal of Northeastern University (Social Science). 2015;17(1):19
2015
-
[49]
Normative data for 102 Spanish remote associate problems and age-related differences in performance
Pel´ aez-Alfonso JL, Pelegrina S, Lechuga MT. Normative data for 102 Spanish remote associate problems and age-related differences in performance. Psicol´ ogica Journal. 2020;41(1):39–65
2020
- [50]
-
[51]
A concise introduction to logic
Hurley PJ. A concise introduction to logic. Nelson Education; 2014
2014
-
[52]
Introduction to Categories and Categorical Logic
Abramsky S, Tzevelekos N. Introduction to Categories and Categorical Logic. In: Coecke B, editor. New Structures for Physics. vol. 813 of Lecture Notes in Physics. Springer-Verlag; 2011. p. 3–94
2011
-
[53]
Roberta: A robustly optimized bert pretraining approach
Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, et al. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:190711692. 2019
2019
-
[54]
Measuring scientific reasoning–a review of test instruments
Opitz A, Heene M, Fischer F. Measuring scientific reasoning–a review of test instruments. Educational Research and Evaluation. 2017;23(3-4):78–101
2017
-
[55]
Developing a test of scientific literacy skills (TOSLS): Measuring undergraduates’ evaluation of scientific information and arguments
Gormally C, Brickman P, Lutz M. Developing a test of scientific literacy skills (TOSLS): Measuring undergraduates’ evaluation of scientific information and arguments. CBE—Life Sciences Education. 2012;11(4):364–377
2012
-
[56]
Milestones: Cognitive
Chen Z. Milestones: Cognitive. In: Benson JB, editor. Encyclopedia of Infant and Early Childhood Development (Second Edition). second edition ed. Oxford: Elsevier; 2020. p. 330–338. Available from: https: //www.sciencedirect.com/science/article/pii/B9780128093245218257
2020
-
[57]
The theory of the estimation of test reliability
Kuder GF, Richardson MW. The theory of the estimation of test reliability. Psychometrika. 1937;2(3):151–160
1937
-
[58]
Coefficient alpha and the internal structure of tests
Cronbach LJ. Coefficient alpha and the internal structure of tests. psychometrika. 1951;16(3):297–334
1951
-
[59]
Student performance on the Test of Scientific Literacy Skills (TOSLS) does not change with assignment of a low-stakes grade
Segarra V A, Hughes NM, Ackerman KM, Grider MH, Lyda T, Vigueira PA. Student performance on the Test of Scientific Literacy Skills (TOSLS) does not change with assignment of a low-stakes grade. BMC research notes. 2018;11:1–5
2018
-
[60]
Test of scientific literacy skills (TOSLS) indicates limited scientific thinking gains as a result of science and mathematics general education
Propsom PM, Tobin WM, Roberts JR. Test of scientific literacy skills (TOSLS) indicates limited scientific thinking gains as a result of science and mathematics general education. Interdisciplinary Faculty Scholarship. 2023
2023
-
[61]
Enhancement of students’ biological literacy and critical thinking of biology through socio-biological case-based learning
Suwono H, Pratiwi H, Susanto H, Susilo H. Enhancement of students’ biological literacy and critical thinking of biology through socio-biological case-based learning. Jurnal Pendidikan IPA Indonesia. 2017;6(2):213–220. December 10, 2024 23/24
2017
-
[62]
Comparing self-report assessments and scenario-based assessments of systems thinking competence
Davis KA, Grote D, Mahmoudi H, Perry L, Ghaffarzadegan N, Grohs J, et al. Comparing self-report assessments and scenario-based assessments of systems thinking competence. Journal of Science Education and Technology. 2023;32(6):793–813
2023
-
[63]
What influences computational thinking? A theoretical and empirical study based on the influence of learning engagement on computational thinking in higher education
Liu S, Peng C, Srivastava G. What influences computational thinking? A theoretical and empirical study based on the influence of learning engagement on computational thinking in higher education. Computer Applications in Engineering Education. 2023;31(6):1690–1704
2023
-
[64]
STEM professional development program for gifted education teachers: STEM lesson plan design competence, self-efficacy, computational thinking and entrepreneurial skills
S ¸ahin E, Sarı U, S ¸en¨OF. STEM professional development program for gifted education teachers: STEM lesson plan design competence, self-efficacy, computational thinking and entrepreneurial skills. Thinking Skills and Creativity. 2024;51:101439
2024
-
[65]
Scaffolding Computational Thinking with ChatGPT
Liao J, Zhong L, Zhe L, Xu H, Liu M, Xie T. Scaffolding Computational Thinking with ChatGPT. IEEE Transactions on Learning Technologies. 2024
2024
-
[66]
ChatGPT improves creative problem-solving performance in university students: An experimental study
Urban M, Dˇ echtˇ erenko F, Lukavsk` y J, Hrabalov´ a V, Svacha F, Brom C, et al. ChatGPT improves creative problem-solving performance in university students: An experimental study. Computers & Education. 2024;215:105031
2024
-
[67]
Bilingualism and creativity: Benefits from cognitive inhibition and cognitive flexibility
Xia T, An Y, Guo J. Bilingualism and creativity: Benefits from cognitive inhibition and cognitive flexibility. Frontiers in Psychology. 2022;13:1016777
2022
-
[68]
Exploring the effect of red and blue on cognitive task performances
Xia T, Song L, Wang TT, Tan L, Mo L. Exploring the effect of red and blue on cognitive task performances. Frontiers in Psychology. 2016;7:784
2016
-
[69]
A quantitative study on the scientific literacy skills of prospective biology teachers
Firdaus L, Ibrohim I, Lestari SR, Masiah M, Primawati SN, Hunaepi H. A quantitative study on the scientific literacy skills of prospective biology teachers. Jurnal Penelitian Pendidikan IPA. 2023;9(1):80–86. December 10, 2024 24/24
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.