{"id":"36580d5e-9c29-41f5-b932-86ee32529f29","arxiv_id":"2505.22112","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"With chain-of-thought prompting and text inputs, GPT-4o, Gemini-1.5 Pro, and Claude-3.5 Sonnet reach or exceed human-level set-shifting on the WCST, but not with visual inputs or direct answers.","lead":"The authors tested three visual AI models on the Wisconsin Card Sorting Test and found they match or beat human accuracy only when the task is explained in text and the model is asked to reason step by step. The result matters because it pins down when a benchmark executive function can and cannot be performed by current AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'human-level' claim depends on an untested asymmetry: models were given an explicit rule-exclusivity instruction while humans may not have been, and removing it degrades model performance substantially.","rationale":"The reader identified the rule-exclusivity instruction asymmetry as the weakest assumption; the paper's own Table III quantifies that this is not a cosmetic difference. Removing the constraint caused a 2.2-category drop for Gemini, a 1.1-category drop for GPT-4o, and a 0.3-category drop for Claude. Since human instructions are only described as 'adapted to be more intuitive,' it is unknown whether humans received the same constraint. This directly affects the load-bearing claim that VLLMs 'achieve or surpass human-level set-shifting capabilities.' The absence of inferential statistics compounds the issue by making the comparison descriptive only. I do not see a separate concern that is more decisive than this one: the ALIEN control is a reasonable check against memorization, the visual-accuracy analysis explains the VI/TI gap, and the impairment simulations are appropriately hedged as potentially stereotypical. Verdict remains conditional: the conclusion can be accepted only after the instruction equivalence is verified and the human comparison is given statistical support.","tokens_in":16087,"tokens_out":3487,"duration_ms":42380,"concrete_test":"Check the exact human instruction text (supplementary Figure s-5 or the deposited materials) for the rule-exclusivity sentence. Then compare the human CC distribution (n=30, M=4.73, SD=0.45) against each model's CoT-TI condition without the exclusivity constraint (n=10, Table III: Gemini M=2.6, GPT-4o M=3.5, Claude M=4.7) using Welch's t-test or bootstrap 95% CIs. If all three no-constraint model distributions remain statistically indistinguishable from or above the human distribution, the asymmetry is not decisive; if any model falls significantly below humans while its with-constraint version was above, the human-level claim is an artifact of unmatched instructions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim requires that the model and human protocols implement the same WCST task. Section IV.D states that all previous model results included the explicit rule-exclusivity constraint 'There will be no combination of these characteristics to define the rule.' Section III.A only says human instructions were 'carefully adapted to be more intuitive' and does not state whether this constraint was included. This matters because the constraint removes composite-rule hypotheses from the model's hypothesis space. Table III shows that removing it lowers CoT-TI categories completed from 4.8 to 2.6 for Gemini-1.5 Pro, from 4.6 to 3.5 for GPT-4o, and from 5.0 to 4.7 for Claude-3.5 Sonnet. If human participants did not receive the same exclusivity hint, the models faced a smaller search space, and the observed 'human-level or surpass' performance may reflect an instruction asymmetry rather than cognitive flexibility. The paper also reports no inferential statistics: with n=10 model runs and n=30 humans, the headline comparison of Claude's perfect 5.00 against the human 4.73 is not tested against sampling variability. The concern is therefore load-bearing for the abstract's central conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates three state-of-the-art VLLMs (GPT-4o, Gemini-1.5 Pro, Claude-3.5 Sonnet) on a 64-trial Wisconsin Card Sorting Test under a 2x2 design (visual/textual input crossed with straight-to-answer/chain-of-thought prompting), compares model performance to 30 human participants, and reports additional experiments: a novel ALIEN Task to control for memorization, removal of an explicit rule-exclusivity constraint, and role-playing prompts intended to simulate goal-maintenance, inhibitory-control, and adaptive-updating impairments. The central claim is that VLLMs achieve or surpass human-level set-shifting under CoT prompting with text-based inputs, and that role-played impairments mimic prefrontal dysfunction patterns.","tokens_in":16287,"tokens_out":3593,"duration_ms":44892,"significance":"If the headline claim holds, this would be a notable result: a systematic neuropsychology-inspired evaluation of multimodal LLMs with a strong control task (ALIEN) and multiple standard WCST metrics. The ALIEN Task is a genuine strength, as it addresses memorization-based alternative explanations, and the paper makes code and data available. However, the headline 'human-level or surpass' conclusion depends on two conditions that are not currently met: a fair model-human instruction match and inferential statistics that support the comparison. The impairment-simulation interpretation is also substantially overreach relative to the evidence and is partly retracted by the paper's own caveat.","major_comments":[{"comment":"The model-human comparison is not demonstrably fair. Section IV.D states that all previous model results included the explicit rule-exclusivity statement 'There will be no combination of these characteristics to define the rule,' while Section III.A only says the human instructions were 'carefully adapted to be more intuitive' and does not specify whether this constraint was included. Standard WCST instructions do not normally provide this exclusivity hint. This matters because Table III shows that removing the constraint lowers CoT-TI categories completed for all three models (Gemini-1.5 Pro 4.8 to 2.6; GPT-4o 4.6 to 3.5; Claude-3.5 Sonnet 5.0 to 4.7). Unless the human protocol contained the same exclusivity information, the models faced a smaller hypothesis space, and the abstract's 'achieve or surpass human-level' claim is not established. The authors should state the full human instruction text and, ideally, compare the no-constraint model condition directly against human performance.","section":"§IV.D vs. §III.A"},{"comment":"The paper reports no inferential statistics for the model-versus-human comparison. With n=10 model runs and n=30 human participants, the claim that Claude-3.5 Sonnet 'surpasses' the human baseline (CC 5.00 vs. 4.73, with zero variance for the model) is not tested against sampling variability, and the same applies to the 'near-human' characterizations for GPT-4o and Gemini-1.5 Pro. The authors should add confidence intervals, effect sizes, and appropriate two-sample tests or bootstrap intervals for at least the primary metric CC, and preferably for the secondary metrics as well. Given the number of models, conditions, and metrics, a multiple-comparison correction or an explicitly preregistered analysis plan should be reported.","section":"§IV.A, Table I, Figure 2"},{"comment":"The impairment-simulation results do not support the claim that VLLMs 'may possess a cognitive architecture, at least regarding the ability of set-shifting, similar to the brain.' The role-playing prompts explicitly instruct the model to behave as if it had a given deficit (e.g., impaired goal maintenance), so observing the requested performance decrements is largely a check of instruction following, not evidence about the model's internal architecture. The paper itself acknowledges in the Discussion that such simulations 'may reflect stereotypical representations of clinical populations rather than authentic mechanisms.' The authors should either reframe Section IV.E as an exploration of prompt-conditioned behavior, or provide convergent evidence beyond the prompted role-play to support the architectural claim.","section":"§IV.E and Discussion"}],"minor_comments":[{"comment":"'We employs a standard version of the WCST-64' should be 'We employed' or 'We use.'","section":"§III.A"},{"comment":"The human baseline row is labeled 'Human STA-VI,' but human participants did not receive a straight-to-answer or chain-of-thought prompt; a label such as 'Human (Visual)' would be clearer and would avoid implying a factorial condition that humans did not have.","section":"Table I"},{"comment":"The figure shows distributions without error bars, confidence intervals, or individual data points; adding within-condition standard errors or a strip plot would make the reported variability more interpretable.","section":"Figure 2"},{"comment":"The 'Decline' column reports raw mean differences without uncertainty; because the standard deviations increase substantially without the exclusivity constraint, the authors should report effect sizes or confidence intervals for these declines.","section":"§IV.D, Table III"},{"comment":"'The test's established validity make it a core benchmark' should be 'makes it a core benchmark.'","section":"§I"},{"comment":"The statement that GPT-4o 'almost always misidentified 5 cards as 6 cards' would benefit from an example or a note in the supplementary material, since Table II reports Count accuracy as 0%.","section":"§IV.C"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a strong empirical core, especially the ALIEN control and the systematic 2x2 design, and the headline claim is plausible from the reported means. The blocking issues are the instruction asymmetry for the model-human comparison, the absence of inferential statistics, and the overinterpretation of the role-play impairment results. These are fixable within the scope of the manuscript, so I do not recommend rejection, but the current version should not be accepted without addressing them."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my honest take. The paper is a solid empirical mapping, not a breakthrough. It does something genuinely new: a clean 2x2 (visual/text input × STA/CoT prompting) across three frontier VLLMs on the WCST, with 10 runs per condition, plus a re-skinned ALIEN control and an impairment-simulation probe. The ALIEN task is the best part — performance patterns survive a complete surface-feature change, which is real evidence against memorization. The role-playing impairment results are qualitative but line up with the clinical literature, and the Discussion is appropriately cautious about stereotype contamination. Code and data are linked. Citation pattern is fair: they extend Loconte et al. and Qu et al. rather than ignoring them.\n\nNow the soft spots, and they matter. First, there are no inferential statistics anywhere. Means and SDs are all you get. With n=10 model runs and n=30 humans, 'Claude 5.00 vs humans 4.73' is untested against sampling variability — a bootstrap would take ten minutes. Second, and more load-bearing, the human-model comparison is probably unfair. Section IV.D says every model condition included an explicit rule-exclusivity hint ('There will be no combination of these characteristics to define the rule'). Section III.A only says human instructions were 'adapted to be more intuitive' and never mentions that hint. Table III shows removing the hint drops Gemini from 4.8 to 2.6 and GPT-4o from 4.6 to 3.5, so it's not a trivial detail. Third, the headline result is text input, while the human baseline is visual; the paper's own modality effect shows text is far easier, so comparing CoT-TI to a visual human baseline stacks the deck. The title and abstract 'human-level' claim is not supported as written.\n\nIf the authors add inferential stats and either give humans the same exclusivity instruction or stop claiming an apples-to-apples human comparison, the core result would be solid. As it stands, it's a useful empirical contribution with a headline that overreaches.\n\nI'd send it to peer review — the ALIEN control and the factorial design deserve referee time — but the reviewers need to push on comparison fairness. For my own citation, I'd cite the methodology with a caveat, not the human-level claim.","headline":"Useful systematic WCST mapping for VLLMs, but the 'human-level' claim rests on an instruction asymmetry and no inferential statistics.","tokens_in":16819,"tokens_out":3479,"would_cite":true,"duration_ms":36293,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Under chain-of-thought prompting with text inputs, visual large language models achieve or surpass human-level set-shifting on the Wisconsin Card Sorting Test.","keywords":["cognitive flexibility","Wisconsin Card Sorting Test","visual large language models","set-shifting","chain-of-thought prompting","rule exclusivity","cognitive impairment simulation","prefrontal cortex"],"falsifier":"Give a fresh set of healthy adults the same 64-trial text-based WCST with the exact instruction set the models received, including the explicit sentence that the rule is never a combination of features, and compare CC, PE, and FMS distributions with the CoT-TI model results. If the human mean climbs well above the reported baseline, the core claim rests on an instruction mismatch; if it does not, the human-level comparison is supported.","tokens_in":15884,"feed_emoji":"🧠","tokens_out":8459,"duration_ms":90486,"temperature":0.7,"pith_summary":"This paper asks whether visual large language models can do what the Wisconsin Card Sorting Test measures: discover an unstated sorting rule, apply it consistently, and switch when the rule silently changes. The authors tested GPT-4o, Gemini-1.5 Pro, and Claude-3.5 Sonnet under four conditions that vary input type (image vs. text) and prompting strategy (direct vs. chain-of-thought). They report that with chain-of-thought prompting and text-based card descriptions, all three models complete as many or more sorting categories than a 30-person human baseline, with Claude-3.5 Sonnet reaching a perfect score in all repetitions. They also report that removing an explicit rule-exclusivity hint degrades performance, and that role-playing prompts can make the models mimic prefrontal-type deficits in goal maintenance, inhibition, and updating. If correct, the result suggests that a key component of human executive function can be approximated by current VLLMs under favorable conditions, with the important caveat that the human comparison may not be instruction-matched.","feed_headline":"Chain-of-thought prompts lift VLLMs to human-level rule switching","feed_subtitle":"Text-based chain-of-thought prompts let three top vision-language models reach or beat human set-shifting scores.","key_machinery":"The load-bearing instrument is the WCST-64, a short-form card-sorting test in which the correct rule is one of three features, color, shape, or number, and the rule switches silently after ten consecutive correct matches. Its power for this study is that its standard metrics, Categories Completed (CC), Perseverative Errors (PE), Non-Perseverative Errors (NPE), Trials to First Category (TFC), Conceptual Level Responses (CLR), and Failure to Maintain Set (FMS), separate rule discovery, set maintenance, and set shifting within a single procedure. The second key mechanism is chain-of-thought prompting, which asks each model to state observations, hypotheses, and choice justifications before selecting a card; the paper's central comparison is the jump from near-chance performance in the direct-answer visual condition to near-human performance in the CoT-TI condition. A third control point is the explicit rule-exclusivity sentence, 'There will be no combination of these characteristics to define the rule,' because removing it is the manipulation that most clearly degrades model performance. The ALIEN Task serves as a surface-replacement control, preserving the same logical structure while changing all terminology and stimuli.","core_discovery":"The central claim is that VLLMs can exhibit human-level, and in one case superhuman, cognitive flexibility as measured by the WCST, provided they are prompted to reason step by step and given card information as text. In the best condition, CoT-TI, Claude-3.5 Sonnet completed every category in every repetition ($CC = 5.00$, $\\sigma = 0.00$), at or above the human baseline; Gemini-1.5 Pro and GPT-4o also matched or approached the baseline. The same performance pattern held on the ALIEN Task, a re-skinned version with alien-themed terminology, which the authors use to argue that the behavior reflects genuine set-shifting rather than memorization of the classic test. They also find that the visual-versus-text gap stems less from poor low-level recognition, which is mostly accurate, than from cascading errors after occasional visual misperceptions, and that removing the explicit statement that the rule is not a combination of features lowers performance, especially for Gemini-1.5 Pro. Finally, role-playing instructions that simulate impaired goal maintenance, inhibitory control, or adaptive updating produce distinct performance-decrement patterns that the authors map onto neuropsychological profiles of prefrontal lesion patients.","pith_inferences":["The fairness of the headline comparison is the natural thing to test next: the models were helped by an explicit sentence ruling out compound rules, while human instructions are only described as 'adapted to be more intuitive,' so a human comparison using the identical instruction text is needed before 'human-level' is taken literally.","The role-played impairments may reflect learned clinical stereotypes rather than authentic internal mechanisms; comparing model error patterns trial-by-trial with actual patient data would help distinguish simulation from mechanism.","A reverse experiment, giving healthy humans chain-of-thought-style instructions, would clarify how much of the CoT-TI advantage comes from the instruction format itself rather than from any model-specific reasoning ability.","The visual-input findings suggest a useful evaluation protocol: benchmark VLLMs on text-description versions of established cognitive tests to isolate reasoning from perception, then add images to measure how perceptual errors cascade into higher-level rule following."],"forward_implications":["If the central claim is correct, current VLLMs can serve as behavioral stand-ins for healthy set-shifting in structured rule-learning experiments, matching or exceeding a typical adult human sample.","The near-chance performance under direct answering and image inputs means benchmark scores on cognitive tests depend heavily on prompt and modality details, so claims of human-level ability must state those conditions.","The drop when the rule-exclusivity hint is removed implies that VLLMs' flexibility is partly scaffolded by explicit constraints, so real-world tasks with ambiguous rule spaces should be expected to be harder for them.","The ALIEN Task result, if accepted, implies that the behavior transfers to novel surface content, which is evidence against rote memorization of the WCST.","The impairment-simulation results, if accepted, offer a prompt-level method for generating distinct executive-dysfunction profiles without modifying model weights."],"supporting_citations":[{"why":"Supplies the original Wisconsin Card Sorting Test paradigm that the study adapts for VLLMs.","marker":"[13]"},{"why":"Supplies the WCST-64 short form and its standardized scoring, which the experiment follows.","marker":"[15]"},{"why":"Supplies chain-of-thought prompting, the manipulation that produces the human-level performance.","marker":"[32]"},{"why":"Identifies GPT-4o as one of the three VLLMs under test.","marker":"[10]"},{"why":"Identifies Gemini-1.5 Pro as one of the three VLLMs under test.","marker":"[11]"},{"why":"Identifies Claude-3.5 Sonnet as one of the three VLLMs under test.","marker":"[12]"},{"why":"Supplies the clinical scoring metrics used to quantify cognitive flexibility.","marker":"[40]"},{"why":"Provides the lesion-based predictions of goal-maintenance and set-shifting deficits that the role-playing results are compared against.","marker":"[44]"},{"why":"Provides the classic frontal-lobe lesion findings that anchor the interpretation of the simulated impairment patterns.","marker":"[45]"}],"fun_headline_variants":["Visual LLMs match human set-shifting with CoT text prompts","Step-by-step text prompts lift VLLMs to human-level flexibility","CoT prompting makes VLLMs match or beat humans on WCST","Text-based reasoning gives VLLMs human-level rule-switching"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the human baseline is a fair comparison: the 30 human participants faced the same inferential difficulty as the models, even though the models received an explicit hint that the sorting rule is never a combination of features.","fun_headline_variants_meta":{"raw":{"variants":["Visual LLMs match human set-shifting with CoT text prompts","Step-by-step text prompts lift VLLMs to human-level flexibility","CoT prompting makes VLLMs match or beat humans on WCST","Text-based reasoning gives VLLMs human-level rule-switching"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000286,"raw_usage":{"total_tokens":1699,"prompt_tokens":982,"completion_tokens":717,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":598,"completion_tokens_details":{"reasoning_tokens":642}},"tokens_in":598,"tokens_out":717,"duration_ms":7707,"temperature":1.0,"reasoning_tokens":642,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:14:12.427313+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give a fresh set of healthy adults the same 64-trial text-based WCST with the exact instruction set the models received, including the explicit sentence that the rule is never a combination of features, and compare CC, PE, and FMS distributions with the CoT-TI model results. If the human mean climbs well above the reported baseline, the core claim rests on an instruction mismatch; if it does not, the human-level comparison is supported.","supporting_citations":[{"cited_title":"A simple objective technique for measuring flexibility in thinking,","cited_arxiv_id":null,"evidence_quote":"Supplies the original Wisconsin Card Sorting Test paradigm that the study adapts for VLLMs."},{"cited_title":"The wcst-64: A standardized short-form of the wisconsin card sorting test,","cited_arxiv_id":null,"evidence_quote":"Supplies the WCST-64 short form and its standardized scoring, which the experiment follows."},{"cited_title":"Hello gpt-4o,","cited_arxiv_id":null,"evidence_quote":"Identifies GPT-4o as one of the three VLLMs under test."},{"cited_title":"Announcements: Claude 3.5 sonnet,","cited_arxiv_id":null,"evidence_quote":"Identifies Claude-3.5 Sonnet as one of the three VLLMs under test."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the clinical scoring metrics used to quantify cognitive flexibility."},{"cited_title":"Wisconsin card sorting test performance in patients with focal frontal and posterior brain damage: effects of lesion location and test structure on separable cognitive processes,","cited_arxiv_id":null,"evidence_quote":"Provides the lesion-based predictions of goal-maintenance and set-shifting deficits that the role-playing results are compared against."},{"cited_title":"Effects of different brain lesions on card sorting: The role of the frontal lobes,","cited_arxiv_id":null,"evidence_quote":"Provides the classic frontal-lobe lesion findings that anchor the interpretation of the simulated impairment patterns."}],"review_version":1}